MATINTELLECT

🔥 On September 22 Anthropic quietly rolled out Opus 5.5 - and I sat down to work with it right away. A day later I can honestly say: this is the first time since Opus 4 and Opus 5 that it feels like the model thinks differently. Not faster. Differently

What exactly changed in Opus 5.5:
The main thing isn't the numbers - it's how the model behaves inside a task. It decides on its own when to call a subagent, checks the result on its own in non-trivial ways, doesn't wait for you to lead it by the hand. I gave it tasks and just went off to keep working. Came back - done, no screw-ups that need redoing. That "ok, it'll figure it out" feeling - which wasn't there in Opus 4 or in Opus 5

The model natively drives tools: runs tests, reads logs, edits files, calls external APIs - all within a single task without you being involved. It doesn't just execute steps, it monitors the process and corrects course on its own if something goes wrong. Anthropic's official case: a migration of a 680,000-line codebase completed in less than a day. I haven't reached that scale, but the feeling of autonomy - I definitely recognize it

Benchmarks Terminal-Bench 4.0 - agentic coding:
🟢 Opus 5.5 - 66.4% - first place among all public models
🟢 GPT-6 Astra - 57.9% - the closest competitor
🟢 Opus 5 - 52.3% - the previous version, to show the gap

OSWorld 2.0 (computer use): 81.8% - also the leader. GDPval-AA v2.1 (knowledge work): 1,846 Elo. It solves terminal tasks in roughly 40% fewer steps than Opus 5 - and you feel it in real work: fewer interventions, more autonomy

Price and speed got better too:
💡 Input/Output: $4/$20 per MTok - was $5/$25, minus 20%
💡 Cache reads: $0.20/MTok - was $0.50, minus 60%
💡 Generation speed: +30% over Opus 5
💡 Limits on Pro/Max/Team/Enterprise raised - if you run agents for hours, this matters

Available on AWS, Google Cloud, Azure. Sonnet 5.5 and Haiku 5.5 with the same key improvements are promised in the coming weeks

Important release context: the day before launch Dario Amodei published an open letter calling on the industry to slow down the pace of developing frontier AI systems. Opus 5.5 is exactly that kind of product: not a radical leap with howling marketing, but a quiet, deep iteration. According to Anthropic's internal behavioral audit - the highest alignment result among all of the company's tested models. Before release, Frontier Design and METR ran independent verification

💻 Alignment audit: the best result among all Anthropic models

💭 With Opus 4 I felt power. With Opus 5 - speed. With Opus 5.5, for the first time in the history of these models, I feel that the model thinks rather than just executes commands. Set a task, walked away - came back, done. It's a difference you can't describe with a benchmark score, but a day of real work conveys it precisely

Instagram | YouTube | Threads

Share:

No comments yet

Leave a Comment

Fields marked with an asterisk (*) are required