The most useful AI upgrade may be the one that makes a difficult job affordable enough to repeat. Anthropic’s new Claude Opus 5.5 is aimed squarely at that calculation.
Released on 22 September, the model is positioned for extended coding tasks and knowledge work. Anthropic says it reaches the level of Claude Fable 5.1 on much of that work, while reporting stronger results than earlier Opus models across several evaluations. Those comparisons come from the developer’s release material and should be read with the stated test conditions. Anthropic’s announcement
The price of a successful result
The official developer documentation lists standard API pricing at US$4 per million input tokens and US$20 per million output tokens. It also lists a one-million-token context window. These are API specifications, not the price or usage allowance of a consumer subscription. Opus 5.5 documentation

For a business, the meaningful comparison goes beyond the price printed beside a model’s name. Suppose one system produces a report cheaply but requires several retries and substantial correction. Another might charge more per response and still deliver the less expensive finished job.
That is why our interpretation of this launch centres on cost per acceptable result. Useful tests would include checking a spreadsheet, repairing a software bug or producing a sourced research brief, then counting the human work required afterwards.
A benchmark can help identify promising models. It cannot establish that the winner will understand a particular company’s documents, permissions and unusual workflows.
Upgrading also changes behaviour
Anthropic’s documentation flags migration changes for developers moving from Opus 5. Adaptive thinking stays enabled, and some existing tool-use patterns need adjustment. Applications that display progress between tool calls may also need changes to preserve that experience. Migration details
These details matter because a model sits inside a larger product. A capable replacement can still disrupt a workflow if the surrounding software expects the previous model’s behaviour.
The practical test is therefore straightforward: run representative tasks, measure the complete cost and inspect the errors. Include difficult examples and incomplete information, rather than only demonstrations that already work well.
Opus 5.5 adds pressure to an AI market where capability alone is becoming an incomplete selling point. Dependable work, delivered at a manageable total cost, is the prize developers and customers should be watching.
Sources
Featured image: Representative image of knowledge work, not a Claude product screen. Credit: Glenn Carstens-Peters.


Leave a Reply