TL;DR: Same test, different wins. On Aug 25, 2026, Qwen 3.7 Plus carried all 32 questions of a real customer's client sign-up form at a fraction of the usual cost, but it started both follow-up edits and finished neither. In that same test, Claude Opus 5 built the most complete sign-up form and the best-executed match tracker, at by far the highest cost, while Claude Sonnet 5 wrote the plan instead of the app. Claude leads on follow-up changes. Route by task inside Taskade Genesis rather than standardizing on one.
What TSK-1 Found
Both families met in the same test on Aug 25, 2026: a real customer's request, an app we open and use ourselves, and a finished app read back against the brief. That day was Qwen's first hands-on test. Claude also carries evidence from Jul 30 through Aug 6, 2026, and every finding below is dated.
Qwen is the value story. Qwen 3.7 Plus took all 32 questions of the client sign-up form, saved every answer, and did it at a fraction of the usual cost. Its apps are plainer than the leaders' apps, and it saved less around the app. Qwen 3.7 Max built a complete match tracker on its first run.
Claude is the finish story, and Aug 25, 2026 showed both sides of it. Claude Opus 5 built the most complete client sign-up form of the test: a five-page form, every answer saved, and the scoring automation graded the submission on its own. Opus 5 also built the best-executed match tracker of the test, without a single failed build action, at by far the highest cost. Claude Sonnet 5 stopped after asking its questions on the sign-up form and wrote the plan instead of the app. Claude Haiku 4.5 was the fastest build of the test, but it left out the weekly recap automation the brief asked for. Claude also leads on follow-up changes, which is exactly where Qwen 3.7 Plus fell short.
- Qwen: Aug 25, 2026: Plus carried all 32 questions of the sign-up form at a fraction of the usual cost; Max built a complete match tracker on its first run; Plus started both follow-up edits and finished neither.
- Claude: Jul 30, 2026: Opus 5 set the design high-water mark and Sonnet 5 opened its own app to check the assistant's answers. Aug 1, 2026: Opus 5 created the fullest workspace of that test. Aug 3, 2026: Sonnet 5 produced our most carefully finished result, with eight minor items left.
See the full evidence at /tsk/qwen, /tsk/claude, and the TSK-1 hub.
Qwen 3.7 Plus vs Claude Sonnet 5
On the long form, Qwen 3.7 Plus finished and Claude Sonnet 5 did not. On everything after the first build, Sonnet 5 is the safer pick. The two met on the same 32-question client sign-up form on Aug 25, 2026, a real customer's request in which every question has to appear in the app and save correctly. Qwen 3.7 Plus took all 32 questions, saved every answer, and did it at a fraction of the usual cost.
Claude Sonnet 5 stopped after asking its questions on that same form and wrote the plan instead of the app. Its earlier record on the form is also mixed. On Aug 3, 2026, it built a contacts database instead of the requested sign-up form after two stalls. On Aug 6, 2026, four attempts stalled before the form was completed. Losing details in a long request is the Claude family's main risk. Against Sonnet 5, Qwen 3.7 Plus wins the long form on value and on completeness. Against the Claude family as a whole, it does not: in the same Aug 25 test, Claude Opus 5 built the most complete client sign-up form of the test, with every answer saved and the scoring automation grading the submission on its own, at by far the highest cost. Qwen's win is a value win. Opus 5's win is a completeness win that you pay for.
The picture flips once the app exists. Qwen 3.7 Plus started both follow-up edits and finished neither: a new page was written but never linked into the app, so the user would never have found it. Claude handles follow-up changes especially well, and it leads on that measure among the families we test. Sonnet 5 also opened its own app and checked the assistant's answers (Jul 30, 2026), the only model to verify its work that way, and on Aug 3, 2026 it produced our most carefully finished result, with eight minor items left.
Qwen 3.7 Plus also has an honest ceiling. Its apps are plainer than the leaders' apps, and it saved less around the app. On a harder field-inspection brief it left the inspector with nothing to do on day one, because the checklist could not start until a weekly automation had run. Value on the form is real. The follow-up edit is where it stopped short.
Qwen 3.7 Max vs Claude Opus 5
Both built a complete match tracker in the same Aug 25, 2026 test. Claude Opus 5 executed it best, at by far the highest cost. Qwen 3.7 Max made it work, with one silent gap. Qwen 3.7 Max built a complete match tracker with a clean dashboard on the first run, and it saved a real match once a hero was picked. With the hero left blank, the save button did nothing and said nothing: a first-use trap for a real user, and the kind of thing a business owner discovers from a customer. That is why we open every app and use it.
Claude Opus 5 built the best-executed match tracker of that same test, without a single failed build action, at by far the highest cost. That is the trade in one sentence: the cleanest build in the test and the most expensive one. On Jul 30, 2026, it set the design high-water mark with a polished, consistent interface. On Aug 1, 2026, it created the fullest workspace of that test, with an assistant that understood its contents. On the Qwen side, Qwen 3.7 Plus built its match dashboard quickly on Aug 25, 2026, but with the thinnest workspace behind it of any model in the test. Both families can put a working screen in front of you. What sits behind the screen, and what it costs, is where they differ most.
Claude Haiku 4.5 adds two more data points. On Jul 31, 2026, it found a problem while building that would have left the app blank, repaired it, and then finished, the only model in that test to recover on its own. On Aug 25, 2026, it was the fastest build of the test, but it left out the weekly recap automation the brief asked for. Read the finished app back against the brief before you trust it.
Choose Qwen If…
- The job is a long, detailed form and the budget is tight. Aug 25, 2026 is the standing evidence: Qwen 3.7 Plus carried all 32 questions of the client sign-up form, saved every answer, and did it at a fraction of the usual cost, on the day Claude Sonnet 5 wrote the plan instead of the app.
- You want a working dashboard on the first run. Qwen 3.7 Max built a complete match tracker with a clean dashboard the first time, and Qwen 3.7 Plus built a working match dashboard quickly (Aug 25, 2026).
- Your input includes images or video. Alibaba lists text, image, and video input for Qwen 3.7 Plus.
- Volume is high and much of it can wait. Alibaba bills batch calls at half price and discounts cache hits.
- You want a family with an open-weight option. The two tiers on this page run on Alibaba's service, but other Qwen generations publish weights under Apache 2.0.
Choose Claude If…
- The finish matters. Opus 5 built the best-executed match tracker of the Aug 25, 2026 test and set the design high-water mark (Jul 30, 2026). Sonnet 5 produced our most carefully finished result (Aug 3, 2026).
- The form has to be complete and cost comes second. Opus 5 built the most complete client sign-up form of the Aug 25, 2026 test, with the scoring automation grading the submission on its own, at by far the highest cost.
- You will change the app after the first build. Claude leads on follow-up changes, and that is exactly the step where Qwen 3.7 Plus started two edits and finished neither.
- You want a full workspace behind the app. Opus 5 created the fullest workspace of the Aug 1, 2026 test, with an assistant that understood its contents. Qwen 3.7 Plus had the thinnest workspace of the Aug 25, 2026 test.
- You want a model that checks and repairs its own work. Sonnet 5 opened its own app and checked the answers (Jul 30, 2026). Haiku 4.5 repaired a problem while building and finished (Jul 31, 2026).
- You need a published rate card you can plan against. Haiku 4.5 at $1 in and $5 out, Sonnet 5 at $2 and $10, and Opus 5 at $5 and $25 per million tokens, with the full 1M-token context billed at standard rates (as of Sep 2026).
The Taskade Angle: Route, Don't Standardize
Most comparison pages end with "pick one". The evidence points the other way: one family owns value on a long form, the other owns finish and follow-up changes. Serious teams run both and route between them.
Taskade routes across 15+ frontier models from OpenAI, Anthropic, and open-weight providers inside one workspace, with the AI allowance included in the subscription rather than a separate API account per lab. Paid plans start at Pro $10/mo billed annually, with Business $25, Max $100 and Enterprise $250 per month billed annually. You set the model per agent or per automation step, so a form-heavy build on Qwen 3.7 Plus and a follow-up edit on Claude Sonnet 5 can each get the model that leads there. Leave a step on TSK-1 Auto and it adapts the depth instead: fast when the step is quick, deeper reasoning when it is not.
Four patterns that hold up:
- Value model builds the form, finish model polishes it. The 32-question form on Qwen 3.7 Plus, then the visual pass on Claude Opus 5.
- First build on one model, every change on another. Qwen builds it, Claude changes it. That is the split the Aug 25 follow-up finding argues for.
- Every step lands in the same project graph. Whichever model runs a step, the result becomes shared workspace memory for the next agent.
- Scheduled automations read from the same place. Model choice becomes a per-step setting.
Final Word: Value vs Finish
Qwen is the value pick: all 32 questions of the client sign-up form saved at a fraction of the usual cost, plus a complete match tracker on the first run from Qwen 3.7 Max, all on Aug 25, 2026. Its apps are plainer, its workspace is thinner, and its follow-up edits did not finish. Claude is the finish pick: the most complete sign-up form and the best-executed match tracker of that same test from Opus 5, at by far the highest cost, plus the lead on follow-up changes. Its long-request risk showed up the same day, when Sonnet 5 wrote the plan instead of the app. Both families reach a million tokens of context, so the choice is about the job, not the window.
Neither is the winner. The winner is the setup that puts value where the form lives and finish where the customer looks.
▲ Memory feeds Intelligence. ■ Intelligence triggers Execution. ● Execution creates Memory. Two families. One workspace.
Build with Qwen and Claude in one workspace →
Related reading
- Qwen vs DeepSeek: Apache 2.0 versus MIT, and which Qwen tiers are open.
- Qwen vs Kimi: Two open-weight ladders, head to head.
- Kimi vs Claude: Open weights versus a closed frontier family.
- DeepSeek vs Claude: The MIT value pick against the finish leader.
- 10 Best Open-Source AI LLMs in 2026: Where the open Qwen line sits in the field.
- Multi-Model AI Access: How Taskade routes across providers.
- TSK-1 Qwen profile: Full evidence for the Qwen family.
- TSK-1 Claude profile: Full evidence for the Claude family.
- TSK-1 hub: The complete model test dataset.



