dataaaaa!
a platform to stack them all
1088 Analytics automation resources collected and tagged on dataaaaa — 679 articles, 347 podcasts, 24 release notes, 20 projects and 18 events. The 30 most recent are listed below, newest first.
Otrium deployed Omni to enable AI self-service over a governed semantic model, cutting ad hoc data tickets. Data engineers synced dbt docs and built curated datasets with joins and access controls, enabling accurate querying across LLM interfaces like Claude via an MCP server.
BigQuery introduces agent-ready augmented analytics utilizing table-valued functions (TVFs). This feature enables data engineering workflows to integrate analytical insights and automated data processing directly within BigQuery data pipelines.
Snowflake announced the general availability of Automations in Snowflake CoWork. The platform update expands pipeline engineering capabilities alongside second-generation Openflow deployments on GCP and online constraint management for hybrid tables.
Traditional BI struggled with data democratization. Now AI agents enable automated querying and dashboard generation, raising critical
Mnemiq is an open-source text-to-SQL engine tunable to custom schemas. It uses a deterministic validation layer ensuring generated queries are read-only, compile to the native dialect, and pass an EXPLAIN plan, refusing execution with stated reasons rather than returning plausible bad data.
AI tools automate intermediary pipeline work. Analysts face a three-bucket split: absorbed routine tasks, compounding systems, and at-risk go-between roles. Long-term value requires engineering scalable tools that function past early user growth and minimize ongoing maintenance costs.
🎙️ Maxq Analytics — The open source agent nao queries data warehouses via natural language across Slack and Claude MCP. Because text-to-SQL is not enough, it integrates a context layer above the semantic layer to streamline consumption.
Graphene is an open-source analytics framework tailored for AI coding agents. It pairs an ANSI-compliant semantic layer with a code-based dashboard format to ensure accurate metrics and joins. Governed by a CLI, it brings version control, CI testing, and automated iteration to data pipelines.
OpenAI launched the Data agent in ChatGPT Work, connecting directly to Snowflake, Databricks, BigQuery, and Redshift. It integrates with dbt and semantic layers to enforce data models, while respecting existing row- and column-level access controls.
ChatGPT Work integrates with data warehouses, data lakes, and sources like Amazon Redshift and Azure CosmosDB. Its Data agent connects to visualization tools and context layers, letting teams query data architectures and build agentic dashboards through plain-language natural query interfaces.
ClickHouse launched an integration for the Data agent in ChatGPT Work via a Remote MCP server. Authenticated with OAuth, the managed server exposes tools that allow agents to inspect schemas, execute SELECT queries, and construct dashboards directly from ClickHouse Cloud data.
Autonomous data engineering shifts teams from manual pipelines to self-healing data products. Across a five-stage maturity model, AI advances from copilots to autonomous agents that detect failures, apply fixes, and adapt to schema changes while humans govern policy and business semantics.
🎙️ The Joe Reis Show — AI agents and prompt-driven workflows are disrupting data engineering, shifting tooling from SQL IDEs to agentic analytics harnesses. This rise resurfaces foundational challenges around data discovery, modeling, and documentation while creating new risks of "token slop" across pipelines.
xlDuckDb is an add-in enabling DuckDB SQL execution directly inside 64-bit Excel 365. It queries dynamic ranges and named tables using the DuckDbQuery function, letting analytical SQL run on spreadsheet data and returning results as native cells via dynamic arrays.
Spotify avoids adding Bayesian A/B testing pipelines, noting both frameworks overlap. Default Bayesian setups often mirror frequentist peeking errors. Selecting priors, likelihoods, and stopping rules requires custom statistical engineering based on specific program error bounds and cost metrics.
Customer reporting requires robust data modeling, standardized metrics, and query-level multi-tenant isolation. Instead of relying on vulnerable UI filters, data teams must implement row-level security at the query layer via runtime identity tokens to prevent multi-tenant data exposure.
GTM AI transformation requires unifying siloed pipelines across campaign touches, call transcripts, and product signals. Building a centralized data foundation creates a single source of truth, enabling cross-functional models to automate connected, end-to-end workflows that drive revenue.
Manual reporting breaks as ecommerce scales because transactional data stays isolated from ad platforms. Engineering automated ELT pipelines into a cloud data warehouse unifies Shopify, web, and ad sources, enabling reliable multi-touch attribution and full customer lifecycle analysis.
Google Cloud announces the Data Agent Kit for agentic analytics. Targeted at data analytics workflows, the release focuses on deploying autonomous agents to handle analytical tasks.
Raw LLMs and AI tools fail at business scale due to lacking an enforced semantic layer, schema maintenance, and row-level permissions. Managed AI analytics services bridge this gap by building governed, operated data apps that ensure metric consistency, reliable pipelines, and audited access.
🎙️ Tony Zeljkovic — Deploying analytics agents requires more than connecting an LLM to a database. Nao co-founders discuss building open-source agentic analytics, delegating workflows, and shifting analytics engineers toward context engineering with semantic layers and metadata to replace traditional dashboards.
Free-text-to-SQL agents fail at scale because generating queries from scratch over uncurated lakehouses causes metric sprawl. Even advanced RAG setups need complex context layers like schema embeddings and query log ingestion to improve accuracy, yet maintenance costs compound without a unified
Natural language database interfaces often fail by relying on single Text-to-SQL queries. A new data investigation paradigm introduces D^2, an engine that autonomously executes sequences of SQL queries and reasons over intermediate results to solve complex, multi-step analytical problems.
AI agents are not eliminating data teams. Complex enterprise architectures and messy pipelines still require human engineers. Instead of shrinking, senior data engineering demand is growing 23% year-over-year, while AI primarily automates boilerplate SQL and schema drafting.
🎙️ The Joe Reis Show — AI agents are hyped to replace data teams, mirroring past cycles like self-service BI. However, core engineering challenges persist: engineers must still prepare underlying data, resolve conflicting definitions, and validate outputs, keeping demand for experienced practitioners essential.
ClickHouse released data-agent-mnist, an open benchmark harness evaluating LLM analytics agents against data warehouses using 201 real-world questions. Claude Fable 5.1 led correctness at 76.6%, while DeepSeek V4 Flash ran the suite for $1 versus $52 for Fable 5.1 with an 11pp drop.
Google introduced TabFM in BigQuery to integrate predictive machine learning directly into its cloud data analytics platform. This tooling expands in-engine predictive analytics capabilities for data engineering and analytics workloads.
AI dashboards create false trust because visual polish masks bad math and wrong KPIs. To enable verification, data platforms should expose the underlying SQL queries, source tables, and row counts directly alongside generated charts.
GoalFlow treats non-linear funnels as graphs defined in YAML. Using NetworkX, it validates dependencies and generates pipeline jobs, typed schemas, and backfill scripts directly from the specification.
🎙️ DataGen — Automate BI and analytics development using AI coding agents like Claude Code. Engineering teams scale workflows by treating agents like a harness for productivity, while managing technical risks like token costs and shadow IT.
See all 1088 Analytics automation resources