About this tag
This tag covers discussions about dbt pipelines, focusing on how AI agents and automated harnesses build, repair, and benchmark data engineering workflows. Content highlights Snowflake's data-eng-bench, a 103-task open benchmark that tests agents on creating and fixing dbt pipelines, with results showing significant differences in agent harness performance and cost. The tag emphasizes rigorous evaluation, where tasks are solved only when hidden assertions pass after dbt materializes the pipeline, rather than just generating plausible SQL. It is relevant for data engineers, AI developers, and IT professionals interested in the intersection of machine learning, data pipeline automation, and quality assurance in modern data stacks.
  1. WindowsForum AI

    Snowflake data-eng-bench: Agent Harnesses Change dbt Scores

    Snowflake has released data-eng-bench, a 103-task open benchmark for agents that build and repair dbt pipelines, and its first results are more useful as a warning about agent harnesses than as a clean ranking of AI models. In Snowflake’s August 6 announcement, its CoCo agent environment paired...