JPMorgan Chase built Jarvis, a shared assembly line that its teams use to build, test, deploy and govern AI agents, according to a conference talk by David Odomirok, Tony Wu and Jane Xue of JPMorgan Chase, published on September 29, 2026.
TL;DR: Jarvis is not one agent. It is a line of replaceable skills that takes an agent idea through development, evaluation and production on the bank's own infrastructure. David Odomirok reports that the one agent the team took months to build the year before now takes minutes (00:23), and that the team went from one agent to a fleet (00:27). Jane Xue describes the trust model: offline evals with positive and negative test cases, online checks for hallucination and relevancy, and subject matter experts who review results and teach the agents new business skills. The talk gives no accuracy figures and does not define "minutes".
| Fact | Detail |
|---|---|
| Source | View source |
| Source type | YouTube talk (LangChain Interrupt NYC), 18 min |
| Source published | 2026-09-29 |
| Speakers | David Odomirok, Executive Director. Tony Wu, Solutions Tech Lead. Jane Xue, Executive Director. All at JPMorgan Chase |
| Industry | Banking and financial services |
| Function | Engineering platform for agents, used by business teams |
| Techniques | Reusable skills, deep agent setup, offline and online evals, LLM judge, human review, governance |
| Tools named | Deep agents, AI coding tools, CI/CD, private and public cloud, generative UI protocols (AG-UI, A2UI) |
Independent summary of public material. JPMorgan Chase is not affiliated with Taskade.
The system JPMorgan Chase built
Jarvis is an agent-building assembly line. David Odomirok explains the reason for it. A financial institution cannot adopt whatever tool is new that week, and it keeps its data inside its own walls. Teams build on the bank's infrastructure, systems and standards. Without a shared line, good agent ideas took so long to build that the moment passed before they shipped.
The team gives Jarvis three principles (01:51):
- Speed. Take an idea from concept to production in days, not months.
- Standardization. One common framework and one quality bar for every agent.
- Evaluations and governance. Build governance in at the start, not at the end.
Three roles run the line: product (David Odomirok), data science (Jane Xue) and engineering (Tony Wu) (02:29).
Architecture of the system
Tony Wu describes the line as a set of skills that covers three life cycles: software development, agent development, and production with operations (03:05). The team writes those skills with AI coding tools. The line includes:
- A deep agent setup with skills, adapted to the enterprise security rules of the bank. The team calls this work "last-mile engineering".
- Extra harness parts on top of an off-the-shelf harness: long-term and short-term memory, technical skills, communication between agents, and a sandbox.
- CI/CD, deployment to the cloud, support from operations and product owners, evaluation and governance.
Every skill on the line is replaceable. A team can swap the agent stack, move from a private cloud to a public cloud, or change the generative UI protocol by replacing one skill. Agents can connect to existing systems, such as quant models and portfolio tools, in an embedded, direct-control or standalone mode. Agents can also work as a team through orchestration skills.
Tony Wu describes a new agent as a "newborn baby" for the subject matter expert (SME) that it supports. It starts at perhaps 20% to 30% of its potential ability (08:26) and grows into a partner through skills, short-term and long-term memory, and evaluation.
Jane Xue describes how the team keeps trust when agents ship quickly:
| Layer | What it does |
|---|---|
| Offline evals | Positive test cases show good behavior. Negative test cases mark a boundary. For an investment analytics agent, a positive case pulls portfolio analytics. A negative case is a request for investment ideas, which the agent must refuse even when asked directly. The team runs the cases until the agent reaches its accuracy and consistency targets, and the result feeds the bank's AI governance process |
| Online evals | Checks on live answers without ground truth. A hallucination check makes sure that every number and claim is grounded in the input. A relevancy check makes sure that the answer addresses the question. Results feed ongoing performance monitoring |
| Human review | SMEs sample results on a schedule. LLM judges escalate unclear cases to a human, who makes the final call. Good and bad results become new test cases |
Each agent ships with its governance skills, so tracing, monitoring and online evals run from its first answer.
Results the source reports
- David Odomirok reports that the one agent the team built the year before "took us months, now takes us minutes" (00:23).
- Odomirok reports that the team moved from one agent to a fleet of agents in a year (00:27).
- Jane Xue reports that human feedback becomes reusable business skills, and that the SME becomes the team lead for a group of agents.
- The speakers give no counts of agents in production and no eval scores.
Critical assessment
"Months to minutes" (00:23) compares the first build of one agent with a later build on a finished line. The talk does not say what "built" includes in each case, such as testing and governance review. The talk gives no figures for accuracy, hallucination rate or adoption. It took place at a conference run by the vendor of the agent framework that the team uses. No outside party checked the claims. The trust layers are the most reusable part of the source, because they apply to any regulated team that ships agents.
Build this in Taskade
- Build a team of AI agents with persistent memory and multi-agent collaboration. Start at AI agents.
- Turn a prompt into a live internal app with Taskade Genesis at AI apps.
- Start from a template in finance agents.
- Connect an outside MCP server through an automation with the MCP Client connector. The connector works on every plan.