AI Agent Examples

Clay: Go-to-Market Agents at 350 Million Runs a Month

5 min read
On this page (6)

Clay runs a go-to-market research agent across whole sales markets, more than 350 million times a month (00:51), according to a conference talk by Jeff Barg, Head of AI at Clay, published on June 24, 2026.

TL;DR: Clay's agent, Claygent, researches companies and people so that sales teams can score accounts and pick the right time to reach out. Jeff Barg reports more than 350 million agent runs a month (00:51) and trillions of tokens a week (04:45). He describes four production problems (04:55): reliable infrastructure, throughput under rate limits, cost, and quality. He reports that adaptive throttling gave 4 to 10 times the throughput of a simple design in internal tests (07:06), and that prompt caching can save up to 70% of cost with some model providers (07:59). The talk is short and gives no method for these figures.

Fact Detail
Source View source
Source type YouTube talk (LangChain Interrupt 2026), 12 min
Source published 2026-06-24
Speakers Jeff Barg, Head of AI, Clay
Industry Sales and marketing software
Function Sales
Techniques Durable execution, checkpoints, adaptive throttling, prompt caching, step limits, offline and online evals
Tools named Claygent, LangGraph, LangSmith, AWS Lambda, Amazon ECS, Snowflake, Salesforce, Gong

Independent summary of public material. Clay is not affiliated with Taskade.

The system Clay built

Clay helps sales teams build lists of companies and people, enrich those lists, and send them into the CRM and outbound campaigns. Jeff Barg says that Clay enriches lists with more than 150 data providers and AI agents (00:37), and that its own data set holds more than 40 million companies and 900 million contacts (00:57).

Barg argues that better targeting matters more than a better email. The strongest Clay users scan their whole market, add signals such as funding news, use agents to score each account and pick the right time to reach out, and then learn from the results. That loop needs an agent run for each account in the market.

Clay's agent, Claygent, does this research. For example, it checks whether a company is a good account to contact at this time.

Architecture of the system

Barg describes four problems and the fix for each (04:55):

  1. Infrastructure. Barg says that "most of our agents are actually just spending their time waiting" (05:36) on browsers, APIs or model inference. Clay first ran Claygent on AWS Lambda, which charges for wall time, so the cost was too high. On Amazon ECS, the team had to handle random host failures. The fix was a durable workflow design with queues and checkpoints at set steps.
  2. Throughput. Clay has dedicated model capacity, but the load comes in spikes. The team built back pressure that works like TCP congestion control: send as much traffic as possible, then reduce it step by step when rate limits appear. A fairness layer stops one large account from using all the capacity.
  3. Cost. Clay built its own agent harness. Barg says that the team designs agents around the caching rules of each model provider. The team also limits retries and tool calls, because an agent that must return after a set number of steps often gives better results (08:15).
  4. Quality. Agents get access to web data and Clay's own data set, and a full team works on that access. The harness is tuned for go-to-market tasks with offline and online evals. Clay also built an agent builder, so that users can test an agent before they run it across a whole market.

The next product is Audiences. It collects a company's own data from tools such as Snowflake, Salesforce and Gong, adds outside signals, and gives the result to Clay's agents as context. Barg says that Audiences is also the base for agent memory, so that agents can suggest plays from what worked before.

Results the source reports

  • Jeff Barg reports more than 350 million go-to-market agent runs a month (00:51).
  • He reports that the agent processes trillions of tokens a week (04:45).
  • He reports that adaptive throttling gave 4 to 10 times the throughput of a simple design in internal experiments (07:06).
  • He reports that caching can save up to 70% of cost with providers such as Anthropic (07:59).
  • He reports that a step limit on research often gives better results than a run to completion, and that this result depends on the use case and must be checked with evals (08:15).

Critical assessment

The talk is short, so each problem gets little time. The 70% figure (07:59) is an upper bound for some providers, not a measured saving for Clay. The 4 to 10 times figure (07:06) comes from internal experiments with no stated method. The run count measures volume, not business results, and the talk gives no figure for meetings, pipeline or revenue. The talk took place at a conference run by the vendor of tools that Barg names. No outside party checked the numbers. The four-problem list (04:55) is concrete and transfers to any team that runs agents at high volume.

Build this in Taskade

More like this