Overview
In the last week of September, the two biggest AI labs released their new everyday workhorse models one day apart. Anthropic shipped Claude Sonnet 5.5 on September 28. OpenAI followed with GPT-6.1 Sol at its DevDay event on September 29.
Both cost the same through the API: $2 per million input tokens and $10 per million output tokens. Both are pitched as strong at coding. So which one should your team, or the developers building your software, actually use?
The honest answer is "it depends on the job." Below is how we think about it at Design Develop Now, where our developers write production code with AI tools like Claude Code and Cursor every day.
The short version
- Pick Claude Sonnet 5.5 for fast, well-scoped coding work: bug fixes, feature additions, UI polish, and quick iteration inside an existing codebase.
- Pick GPT-6.1 Sol for very large codebases and long-running agent work where a huge context window and cheap cached input matter.
- Use both if you're building agents or AI features into a product. Routing different tasks to different models is now normal practice, and it protects you from any one vendor's outages or price changes.
For a non-technical owner, here's the takeaway: the model matters less than the team using it. Both are excellent. What decides your project's outcome is how well the developers scope the work, test the output, and review what the AI writes.
Side-by-side: what each company says
Every number here comes from the vendor's own launch post. There is no independent head-to-head of these two exact models yet. Anthropic's comparisons were run against the older GPT-6 Sol, and OpenAI's against Claude Opus 5.5.
| Claude Sonnet 5.5 | GPT-6.1 Sol | |
|---|---|---|
| Released | September 28, 2026 | September 29, 2026 |
| API price (input / output, per 1M tokens) | $2 / $10 | $2 / $10 |
| Cached input (per 1M tokens) | $0.20 | $0.10 |
| API model name | claude-sonnet-5-5 | gpt-6.1-sol |
| Headline coding claim | 70.6% on Terminal-Bench 4.0, up from 10.3% for Sonnet 5 | Matches GPT-6 Astra on DeepSWE v1.1 at about one-fifth the cost |
| Speed claim | 30%+ faster output than Sonnet 5 | "Ultrafast" mode coming, up to 8x faster in Codex |
| Where it's offered | Claude Platform, AWS, Google Cloud, Microsoft Azure | ChatGPT Work, Codex, OpenAI API |
Where Claude Sonnet 5.5 stands out
Anthropic positions Sonnet 5.5 as the fast, lower-cost partner to its flagship Opus 5.5. The company says it's strongest at well-scoped everyday tasks and fixing bugs, and that it uses far fewer tokens than Sonnet 5 to finish the same work, cutting cost per task by up to 30%.
Three things matter for real projects:
- Real-world coding sessions. On CursorBench 4.0, built from actual Cursor coding sessions, Anthropic reports Sonnet 5.5 at 55.5%, within about two points of Opus 5.5. That's the closest benchmark to what developers do all day.
- Fewer wasted steps. Early testers quoted by Anthropic, including Lovable and Base44, reported fewer tool calls and fewer iterations to finish a build. Fewer steps means faster delivery and a smaller bill.
- Design sense. Anthropic highlights its eye for user interfaces. For agencies building client-facing apps, that's practical: less time fixing spacing and layout after the code works.
One caveat. Anthropic itself says Opus 5.5 remains clearly stronger for complex, open-ended work that needs sustained judgment. For architecture decisions on a large system, Sonnet isn't the top pick, even from its own maker.
Where GPT-6.1 Sol stands out
OpenAI calls GPT-6.1 Sol "near-Astra intelligence for a fifth of the price", meaning it approaches its flagship GPT-6 Astra on coding and computer use at much lower cost.
- Long, hard engineering tasks. On DeepSWE v1.1, which tests long-horizon tasks in real codebases, OpenAI says it matches Astra at roughly one-fifth the cost.
- Cheap repeated context. Cached input is $0.10 per million tokens. If an agent re-reads the same large codebase or document set over and over, that adds up to real savings.
- Business workflows. On AutomationBench, which tests multi-step workflows across tools like CRMs and support desks, OpenAI reports it 2.2 points ahead of Claude Opus 5.5 at medium effort.
The caveat here: GPT-6.1 Sol launched in ChatGPT Work and Codex, and OpenAI notes it is not yet in the regular ChatGPT chat. Check where your team actually works before standardizing on it.
Which model we'd start with, by type of project
This is our starting point, not a law. We test on the client's actual code before committing.
| Project type | Where we'd start | Why |
|---|---|---|
| Next.js and React web apps | Claude Sonnet 5.5 | Strong on scoped feature work and UI polish |
| Debugging an existing codebase | Claude Sonnet 5.5 | Positioned for bug fixing, fast to iterate with |
| Large refactors and legacy modernization | GPT-6.1 Sol, with Opus 5.5 for planning | Long-horizon tasks, big context, cheap caching |
| Mobile app development | Either; test both | Depends on framework and codebase size |
| API, backend and database work | Either; test both | Both strong; pick the one your team reviews best |
| Design-to-code and interface work | Claude Sonnet 5.5 | Anthropic emphasizes interface quality |
| AI agents and workflow automation | Both, routed by task | Avoids lock-in; each wins different steps |
What the benchmarks don't tell you
Benchmarks measure a model working alone on a test. Real software is built by people working with the model. A few things matter more than a few points on a chart:
- Review discipline. AI-written code still needs a senior developer reading it. Speed without review just ships bugs faster.
- Clear task scoping. Both models do best with a well-defined task. Vague instructions waste tokens on either platform.
- Security and data handling. Both companies offer business plans with stricter data terms. Anthropic notes Sonnet 5.5 supports zero data retention. If you're in healthcare or legal, confirm the terms before sending client data to any model.
- Your existing tools. If your team lives in GitHub, Cursor, or a specific cloud, pick the model that plugs into that workflow cleanly.
Frequently asked questions
Is Claude better than ChatGPT for coding in 2026? Neither wins everywhere. Sonnet 5.5 is a strong default for everyday coding; GPT-6.1 Sol is compelling for long, large-codebase tasks. Most serious teams use both.
Are they the same price? Yes, $2 input and $10 output per million tokens. Real cost depends on how many tokens each uses to finish a task, which is why both companies talk about "cost per task."
Should a small business care which model its developer uses? Mostly, care that your developer uses AI responsibly: reviewed code, tested features, and no client data sent where it shouldn't go.
Build with the right model, not just the newest one
New models now arrive every few weeks. The advantage goes to teams that can test quickly and switch without rewriting everything. That's how we approach custom software development at DDN.
If you're planning an app, an AI feature, or an internal tool and want a straight answer on which model fits, book a 15-minute discovery call.
Keep exploring
Related services
Need help applying this?
Design Develop Now builds websites, apps, and SEO-ready digital systems for businesses that need practical execution.
Start a project