AI Agent Examples

Vercel: A Data Science Agent Rebuilt Around a File System

6 min read
On this page (6)

Vercel built an internal data science agent that answers questions about customers, products and usage for every team in the company, according to a conference talk by Andrew Qu, Chief of Software at Vercel, published on September 14, 2026.

TL;DR: Vercel's data team was a bottleneck. Every question from marketing or sales meant a data scientist stopped work to write a query. Andrew Qu describes the versions of an agent that took over that work: one large prompt, a chain of narrow agents, one agent that managed its own state, and finally one agent that works inside a sandbox with a file system. Qu reports that the file-system rebuild doubled the eval score (09:35), that a recurring job turned common queries into about 100 skills (10:42), and that Vercel now runs about 20 agents with decent product-market fit (16:03). The talk gives no absolute eval figures after the rebuild.

Fact Detail
Source View source
Source type YouTube talk (AI Engineer), 17 min
Source published 2026-09-14
Speakers Andrew Qu, Chief of Software, Vercel
Industry Developer platform and cloud hosting
Function Data and analytics, used by marketing, sales, finance and legal
Techniques Single agent loop, file-system tools, sandbox, skills library, evals
Tools named Snowflake, a semantic layer, a bash tool package on npm, Claude Code (as the reference design)

Independent summary of public material. Vercel is not affiliated with Taskade.

The system Vercel built

The agent answers data questions for people outside the data team. Andrew Qu started the project about a year before the talk (02:22). He asked the marketing, sales, finance and legal teams what they disliked most about their jobs. The strongest answer came from the data team. The team was small, the company grew faster, and every question about a customer or a product made a data scientist drop their work to write a query, analyze the result and report back.

Qu worked with Vercel's VP of Data to replace that loop with an agent. Inside Vercel, the agent has the internal name D0. By the time of the talk, people across the company sent it thousands of queries a day (10:19), from customer metrics and sales metrics to npm download counts.

Architecture of the system

The agent went through these designs. Each one fixed a limit of the one before it.

Version Design Limit Qu reports
1 One large prompt with a dump of the Snowflake schema. A person copied the SQL and ran it by hand (03:54) A test of whether models can write valid SQL, not a product
2 A chain of agents (query, planning, execution, reporting), each with its own system prompt and a small set of tools (05:02) Each agent saw only a summary of the work before it
3 One agent that managed its own state across planning, execution and reporting, with up to 100 steps per run (06:38) The first trusted users judged it poor. It passed about 30% of the evals (07:20)
4 One agent in a sandbox that holds the whole semantic layer as files. The agent lists, reads, writes and greps files and runs bash, plus a few Vercel-specific tools (08:56) The design that went company-wide

The change to version 4 came from a comparison. Qu watched a general-purpose coding agent answer the same questions with only file and bash tools. He concluded that models already know how to use those tools well, and that a rigid, hand-designed tool set held his own agent back. He ties this shift to the release of a new model, so the redesign and the model change arrived together.

After the rollout, a recurring job reads the most recent queries and turns repeated query shapes into skills. Qu reports about 100 skills (10:42), from aggregations to lookups of specific data. A new run then starts with that context instead of only the system prompt and the semantic layer.

Results the source reports

  • Andrew Qu reports that the eval score "basically doubled" after the move to the file-system design (09:35).
  • Qu reports thousands of queries a day from across Vercel (10:19), and about 100 skills distilled from them (10:42).
  • Qu reports about 20 internal agents with decent product-market fit (16:03), for example marketing retrospectives, the choice of who to contact, and a first redline of a contract for the legal team.
  • Qu reports that the data team now has time to improve Snowflake performance and to add missing data sources.

Critical assessment

All results come from the team that built the agent, and no outside party checked them. The talk gives two eval points, about 30% before the rebuild (07:20) and "doubled" after it (09:35), but it does not describe the eval set, its size or how it changed over time. Qu ties the shift to a new model release, so the doubled score mixes a design change with a model change. The talk also promotes a Vercel product, and the speaker presents the agent as the origin of that product. Treat the architecture lessons as one team's experience, not as a controlled comparison.

Build this in Taskade

  • Build an AI agent with web search, file analysis and persistent memory that answers questions from your team. Start at AI agents.
  • Start from a template for data work in data agents.
  • Pull rows from a spreadsheet into an automation with the Google Sheets integration, one of 100+ integrations.
  • Connect an outside MCP server through an automation with the MCP Client connector. The connector works on every plan.

More like this