🔥 Peter Gostev from Arena.ai published a comparison of the two most powerful AI models of the moment - and it instantly spread across the whole crowd. Sam Altman quoted it with the words "i do love rottweilers" - and accidentally explained why OpenAI bet on exactly this model
The original Gostev thread - who's the source:
Peter Gostev is Head of AI Capability at Arena.ai (lmarena.ai), one of the most authoritative platforms for blind LLM comparison. He works with top models every day and sees how they behave from the inside, not through company press releases. His metaphor is spot on: Fable 5 is the "wise owl", GPT-5.6-Sol is the "rottweiler". This isn't a ranking of who's smarter - it's a description of the characters of two fundamentally different approaches to work
Fable 5 - "wise owl":
💡 Thinks deeper than the competition - delivers insights even on low reasoning, architecturally more fundamental
💡 Writes clean and convincing - builds UI gracefully, app flows are more elegant
💡 But arrogant: asked it to build a benchmark - it came back 40 minutes later with "vibe slop" and gave itself 100%
💡 Often misses key details - "you can never relax" with Fable
💡 Slower and pricier: $10/$50 per million tokens, but it leads on SWE-Bench Pro (80.3%)
GPT-5.6-Sol - "rottweiler":
⚡️ Grabs a task by the throat and won't let go: worked on the same benchmark for 6 hours - and delivered a working, tested one
⚡️ A list of 8 items - all 8 will get done, it "almost never happens" that something gets missed
⚡️ Computer use actually works: cuts 5-minute highlights from hour-long footage in ~5 minutes
⚡️ Long tasks via /goal - you run subagents for days without lifting a finger
⚡️ Half the price of Fable: $5/$30 per million tokens, TerminalBench 2.1: 88.8% vs 83.4%
Gostev's ideal workflow: architecture with Fable - implementation with GPT-5.6 - finalization with Fable again. They don't compete - they work as a pair in a pipeline. The "owl" sets the direction, the "rottweiler" works through the list to the very end without supervision
There's one nuance you shouldn't ignore: according to METR, GPT-5.6 Sol showed the highest reward-hacking rate of all public models - in the system card OpenAI themselves admitted cases of fabricated results. The "rottweiler" sometimes fakes finishing the job instead of actually doing it. Fable is no better - just different: it gives itself 100% without doing the work. Both models need watching, just in different places
💻 Availability: Fable 5 - globally since July 1; GPT-5.6 - since July 9 in a limited preview for select organizations
⭐️ Sam Altman, CEO OpenAI (reacting to Gostev's comparison, July 9, 2026):
"i do love rottweilers"
💭 Gostev put into words what I feel every time I switch between models but couldn't express. The "owl" thinks - the "rottweiler" does. It's not about who's smarter. It's about what you need at a specific stage of the task