download dots

Qwen vs Claude

Two families met in the same test on Aug 25, 2026. Qwen 3.7 Plus carried all 32 questions of a real customer's client sign-up form and saved every answer at a fraction of the usual cost, and Qwen 3.7 Max built a complete match tracker on its first run. Claude Opus 5 built the most complete sign-up form and the best-executed match tracker of that test, at by far the highest cost, while Claude Sonnet 5 wrote the plan instead of the app. Claude leads on follow-up changes. This page is a routing matrix, not a scoreboard.

Last updated: September 2026

Quick Comparison Table

Feature Qwen (Alibaba) Claude (Anthropic)
Weights Qwen 3.7 Plus and Max run on Alibaba's own service; no Qwen 3.7 weights on the official Hugging Face page. Other Qwen generations publish open weights under Apache 2.0 Closed, never released
Live models qwen3.7-plus (text, image, and video in, text out), qwen3.7-max (source) Claude Sonnet 5, Claude Opus 5, Claude Haiku 4.5, plus the Fable 5.1 tier (source)
Context window 1,000,000 tokens on Qwen 3.7 Plus, max output 131,072; Qwen 3.7 Max billed on one tier up to 1M (source) 1M tokens on Sonnet 5 and Opus 5, max output 128K; Haiku 4.5 is 200K with 64K out (source)
Multimodal ✅ Qwen 3.7 Plus takes text, images, and video ✅ Text and image input on every current model
Price per 1M tokens (as of Sep 2026) Published on the vendor's site (source); batch calls at half price, cache hits discounted Haiku 4.5 $1 in / $5 out · Sonnet 5 $2 / $10 · Opus 5 $5 / $25; batch 50% off, cache hits at a tenth of input (source)
Headline public benchmark Vendor-published; TSK-1 is the controlled evidence below Vendor-published; TSK-1 is the controlled evidence below
What we found Plus carried all 32 questions of the client sign-up form at a fraction of the usual cost; Max built a complete match tracker on its first run (both Aug 25, 2026) Opus 5 built the most complete sign-up form and the best-executed match tracker of the same Aug 25, 2026 test, at by far the highest cost; leads on follow-up changes
Best for Value on long, detailed forms Polished apps that keep improving

TL;DR: Same test, different wins. On Aug 25, 2026, Qwen 3.7 Plus carried all 32 questions of a real customer's client sign-up form at a fraction of the usual cost, but it started both follow-up edits and finished neither. In that same test, Claude Opus 5 built the most complete sign-up form and the best-executed match tracker, at by far the highest cost, while Claude Sonnet 5 wrote the plan instead of the app. Claude leads on follow-up changes. Route by task inside Taskade Genesis rather than standardizing on one.


What TSK-1 Found

Both families met in the same test on Aug 25, 2026: a real customer's request, an app we open and use ourselves, and a finished app read back against the brief. That day was Qwen's first hands-on test. Claude also carries evidence from Jul 30 through Aug 6, 2026, and every finding below is dated.

Qwen is the value story. Qwen 3.7 Plus took all 32 questions of the client sign-up form, saved every answer, and did it at a fraction of the usual cost. Its apps are plainer than the leaders' apps, and it saved less around the app. Qwen 3.7 Max built a complete match tracker on its first run.

Claude is the finish story, and Aug 25, 2026 showed both sides of it. Claude Opus 5 built the most complete client sign-up form of the test: a five-page form, every answer saved, and the scoring automation graded the submission on its own. Opus 5 also built the best-executed match tracker of the test, without a single failed build action, at by far the highest cost. Claude Sonnet 5 stopped after asking its questions on the sign-up form and wrote the plan instead of the app. Claude Haiku 4.5 was the fastest build of the test, but it left out the weekly recap automation the brief asked for. Claude also leads on follow-up changes, which is exactly where Qwen 3.7 Plus fell short.

  • Qwen: Aug 25, 2026: Plus carried all 32 questions of the sign-up form at a fraction of the usual cost; Max built a complete match tracker on its first run; Plus started both follow-up edits and finished neither.
  • Claude: Jul 30, 2026: Opus 5 set the design high-water mark and Sonnet 5 opened its own app to check the assistant's answers. Aug 1, 2026: Opus 5 created the fullest workspace of that test. Aug 3, 2026: Sonnet 5 produced our most carefully finished result, with eight minor items left.

See the full evidence at /tsk/qwen, /tsk/claude, and the TSK-1 hub.


Qwen 3.7 Plus vs Claude Sonnet 5

On the long form, Qwen 3.7 Plus finished and Claude Sonnet 5 did not. On everything after the first build, Sonnet 5 is the safer pick. The two met on the same 32-question client sign-up form on Aug 25, 2026, a real customer's request in which every question has to appear in the app and save correctly. Qwen 3.7 Plus took all 32 questions, saved every answer, and did it at a fraction of the usual cost.

Claude Sonnet 5 stopped after asking its questions on that same form and wrote the plan instead of the app. Its earlier record on the form is also mixed. On Aug 3, 2026, it built a contacts database instead of the requested sign-up form after two stalls. On Aug 6, 2026, four attempts stalled before the form was completed. Losing details in a long request is the Claude family's main risk. Against Sonnet 5, Qwen 3.7 Plus wins the long form on value and on completeness. Against the Claude family as a whole, it does not: in the same Aug 25 test, Claude Opus 5 built the most complete client sign-up form of the test, with every answer saved and the scoring automation grading the submission on its own, at by far the highest cost. Qwen's win is a value win. Opus 5's win is a completeness win that you pay for.

The picture flips once the app exists. Qwen 3.7 Plus started both follow-up edits and finished neither: a new page was written but never linked into the app, so the user would never have found it. Claude handles follow-up changes especially well, and it leads on that measure among the families we test. Sonnet 5 also opened its own app and checked the assistant's answers (Jul 30, 2026), the only model to verify its work that way, and on Aug 3, 2026 it produced our most carefully finished result, with eight minor items left.

Qwen 3.7 Plus also has an honest ceiling. Its apps are plainer than the leaders' apps, and it saved less around the app. On a harder field-inspection brief it left the inspector with nothing to do on day one, because the checklist could not start until a weekly automation had run. Value on the form is real. The follow-up edit is where it stopped short.


Qwen 3.7 Max vs Claude Opus 5

Both built a complete match tracker in the same Aug 25, 2026 test. Claude Opus 5 executed it best, at by far the highest cost. Qwen 3.7 Max made it work, with one silent gap. Qwen 3.7 Max built a complete match tracker with a clean dashboard on the first run, and it saved a real match once a hero was picked. With the hero left blank, the save button did nothing and said nothing: a first-use trap for a real user, and the kind of thing a business owner discovers from a customer. That is why we open every app and use it.

Claude Opus 5 built the best-executed match tracker of that same test, without a single failed build action, at by far the highest cost. That is the trade in one sentence: the cleanest build in the test and the most expensive one. On Jul 30, 2026, it set the design high-water mark with a polished, consistent interface. On Aug 1, 2026, it created the fullest workspace of that test, with an assistant that understood its contents. On the Qwen side, Qwen 3.7 Plus built its match dashboard quickly on Aug 25, 2026, but with the thinnest workspace behind it of any model in the test. Both families can put a working screen in front of you. What sits behind the screen, and what it costs, is where they differ most.

Claude Haiku 4.5 adds two more data points. On Jul 31, 2026, it found a problem while building that would have left the app blank, repaired it, and then finished, the only model in that test to recover on its own. On Aug 25, 2026, it was the fastest build of the test, but it left out the weekly recap automation the brief asked for. Read the finished app back against the brief before you trust it.


Choose Qwen If…

  • The job is a long, detailed form and the budget is tight. Aug 25, 2026 is the standing evidence: Qwen 3.7 Plus carried all 32 questions of the client sign-up form, saved every answer, and did it at a fraction of the usual cost, on the day Claude Sonnet 5 wrote the plan instead of the app.
  • You want a working dashboard on the first run. Qwen 3.7 Max built a complete match tracker with a clean dashboard the first time, and Qwen 3.7 Plus built a working match dashboard quickly (Aug 25, 2026).
  • Your input includes images or video. Alibaba lists text, image, and video input for Qwen 3.7 Plus.
  • Volume is high and much of it can wait. Alibaba bills batch calls at half price and discounts cache hits.
  • You want a family with an open-weight option. The two tiers on this page run on Alibaba's service, but other Qwen generations publish weights under Apache 2.0.

Choose Claude If…

  • The finish matters. Opus 5 built the best-executed match tracker of the Aug 25, 2026 test and set the design high-water mark (Jul 30, 2026). Sonnet 5 produced our most carefully finished result (Aug 3, 2026).
  • The form has to be complete and cost comes second. Opus 5 built the most complete client sign-up form of the Aug 25, 2026 test, with the scoring automation grading the submission on its own, at by far the highest cost.
  • You will change the app after the first build. Claude leads on follow-up changes, and that is exactly the step where Qwen 3.7 Plus started two edits and finished neither.
  • You want a full workspace behind the app. Opus 5 created the fullest workspace of the Aug 1, 2026 test, with an assistant that understood its contents. Qwen 3.7 Plus had the thinnest workspace of the Aug 25, 2026 test.
  • You want a model that checks and repairs its own work. Sonnet 5 opened its own app and checked the answers (Jul 30, 2026). Haiku 4.5 repaired a problem while building and finished (Jul 31, 2026).
  • You need a published rate card you can plan against. Haiku 4.5 at $1 in and $5 out, Sonnet 5 at $2 and $10, and Opus 5 at $5 and $25 per million tokens, with the full 1M-token context billed at standard rates (as of Sep 2026).

The Taskade Angle: Route, Don't Standardize

Most comparison pages end with "pick one". The evidence points the other way: one family owns value on a long form, the other owns finish and follow-up changes. Serious teams run both and route between them.

Taskade routes across 15+ frontier models from OpenAI, Anthropic, and open-weight providers inside one workspace, with the AI allowance included in the subscription rather than a separate API account per lab. Paid plans start at Pro $10/mo billed annually, with Business $25, Max $100 and Enterprise $250 per month billed annually. You set the model per agent or per automation step, so a form-heavy build on Qwen 3.7 Plus and a follow-up edit on Claude Sonnet 5 can each get the model that leads there. Leave a step on TSK-1 Auto and it adapts the depth instead: fast when the step is quick, deeper reasoning when it is not.

Four patterns that hold up:

  • Value model builds the form, finish model polishes it. The 32-question form on Qwen 3.7 Plus, then the visual pass on Claude Opus 5.
  • First build on one model, every change on another. Qwen builds it, Claude changes it. That is the split the Aug 25 follow-up finding argues for.
  • Every step lands in the same project graph. Whichever model runs a step, the result becomes shared workspace memory for the next agent.
  • Scheduled automations read from the same place. Model choice becomes a per-step setting.

Final Word: Value vs Finish

Qwen is the value pick: all 32 questions of the client sign-up form saved at a fraction of the usual cost, plus a complete match tracker on the first run from Qwen 3.7 Max, all on Aug 25, 2026. Its apps are plainer, its workspace is thinner, and its follow-up edits did not finish. Claude is the finish pick: the most complete sign-up form and the best-executed match tracker of that same test from Opus 5, at by far the highest cost, plus the lead on follow-up changes. Its long-request risk showed up the same day, when Sonnet 5 wrote the plan instead of the app. Both families reach a million tokens of context, so the choice is about the job, not the window.

Neither is the winner. The winner is the setup that puts value where the form lives and finish where the customer looks.

▲ Memory feeds Intelligence. ■ Intelligence triggers Execution. ● Execution creates Memory. Two families. One workspace.

Build with Qwen and Claude in one workspace →


TSK-1 Benchmark

Same request, both models

In the TSK-1 Benchmark, every model receives the same app request, word for word: build a working app that keeps what people enter, runs an automation, and answers questions about its own data, then take a follow-up change. Here is how Qwen and Claude did, tested inside Taskade Genesis.

Alibaba · Tested Aug 2026

Qwen

Interface
Emerging
Task
Strong
Memory
Emerging
Adapt
Limited

Best for: Value on long, detailed forms

Qwen is the value newcomer. In its first hands-on test, Qwen 3.7 Plus finished the 32-question client sign-up form with every answer saved, at a fraction of the usual cost, and Qwen 3.7 Max built a clean match tracker. Follow-up changes are where it still stalls: two asked-for edits were started and not finished.

  • · Qwen 3.7 PlusTook all 32 questions of the client sign-up form, saved every answer, and did it at a fraction of the usual cost.
  • · Qwen 3.7 PlusBuilt a working match dashboard quickly, but with the thinnest workspace behind it of any model in the test.

All Qwen results

Anthropic · Tested Aug 2026

Claude

Interface
Strong
Task
Strong
Memory
Strong
Adapt
Leading

Best for: Polished apps that keep improving

Claude produces the most polished finished apps we test and handles follow-up changes especially well. Its main risk is losing details in a long request.

  • · Claude Sonnet 5Stopped after asking its questions on the client sign-up form, and wrote the plan instead of the app.
  • · Claude Opus 5The most complete client sign-up build of the test: a five-page form, every answer saved, and the scoring automation graded the submission on its own.

All Claude results

Interface, Task, Memory and Adapt are the four qualities TSK-1 grades: how finished the app feels, how closely it follows the request, whether it keeps your data, and how cleanly it handles follow-up changes. Read the method and every published test on the TSK-1 hub.

Open a live app built the same way

These are App Kits from the official Taskade account, not benchmark builds. Each one is a working app with projects, agents and automations, the same shape every benchmark request asks for. Open one, then clone it into your own workspace.

Browse all App Kits →

Verify the comparison yourself

This is our take. We’re biased — we make Taskade. Read the alternatives from the source:

When you are ready, build with Taskade Genesis or browse live apps from the Taskade community.

More Competitors & Alternatives

View All Alternatives ↗

Cursor

Codex vs Cursor in 2026: OpenAI's agentic coding system versus the AI-native code editor, with a per-task routing matrix, what Cursor's compute-based pricing actually buys, and the third path for people who want the finished app — Taskade Genesis.

Learn More

Cursor

Taskade Genesis vs Cursor in 2026. Cursor is one of the most-used AI-native code editors and ships new versions fast, the best-in-class agentic IDE for working engineers. Taskade Genesis is for the rest of the team (operators, founders, PMs), shipping deployed apps from one prompt with AI agents, workspace data, and 100+ bidirectional integrations included — and an AI allowance that comes with the subscription instead of being metered at API rates.

Learn More

Windsurf

Windsurf is now Devin Desktop — Cognition folded the IDE into the Devin product line and windsurf.com redirects to devin.ai. Taskade Genesis ships a deployed AI app workspace with built-in agents and 100+ integrations, so anyone on the team can use what gets built, not just the engineer who ran the prompt.

Learn More

Lovable

Codex Sites vs Lovable in 2026: OpenAI's Business-only, workspace-private app builder versus Lovable's full-stack code generator — with real 2026 pricing, an honest look at credit metering on both sides, and the prompt-to-app builder that publishes to the open web for everyone, Taskade Genesis.

Learn More

Lovable

The best Lovable alternatives in 2026, compared for people who ship business systems rather than codebases. Lovable is an excellent design-first builder that returns a React + Vite project you host and maintain. This page ranks eight alternatives by what you are actually building, states Lovable's real 2026 pricing with sources, and explains where Taskade Genesis fits: a running system with data, AI agents, automations, and app sign-in, with no deployment step.

Learn More

Lovable

Taskade vs Lovable, head-to-head for 2026. Taskade Genesis turns one prompt into a living app with AI agents, automations, and 100+ integrations you publish to the open web. Lovable generates React and Supabase code you deploy yourself.

Learn More

Bolt.new

Taskade Genesis vs Bolt.new in May 2026, after Bolt V2 (October 2025) Bolt Cloud + databases + hosting + Expo mobile, $40M ARR in 5 months, and StackBlitz's $105.5M Series B at ~$700M valuation. Bolt has the only browser-native WebContainers runtime in the category. Genesis ships deployed apps with AI Agents v2, 100+ bidirectional integrations, and Workspace DNA, flat $10/mo (billed annually) Pro, no token meter on bug fixes.

Learn More

Bolt.new

Taskade vs Bolt.new, head-to-head for 2026. Taskade Genesis ships a deployed app with AI agents, automations, and 100+ integrations from one prompt. Bolt.new generates React code in a browser sandbox you deploy yourself.

Learn More

V0

Taskade Genesis vs v0 by Vercel in 2026 — after the v0.dev to v0.app rebrand, Figma and custom design-system import, the built-in Git panel, and agentic workflows. v0 ships best-in-class React/Next.js and shadcn code with the cleanest Figma-to-code path, now entering at Plus $30/user/mo with no annual billing. Taskade Genesis ships full deployed apps with a workspace backend, AI agents, and 100+ integrations on flat $10/mo billed annually.

Learn More

Imagine it. Run it live.

One prompt. Memory, intelligence, and execution — already wired, already running.