💡 You plug in Claude, GPT, Gemini - the best there is. And the agent still crashes. Every run is like its first day on the job: no memory, no role, no route.
Where agent systems actually break:
Companies are obsessed with model accuracy and ignore the infrastructure layer - and that's exactly where everything quietly falls apart: data pipelines, orchestration logic, retrieval systems, downstream workflows. According to research, more than 80% of AI deployments fail in the first 6 months - and the problem is almost never the model. If an agent is 85% reliable at each step, a 10-step workflow will succeed only ~20% of the time. Not because the model made a mistake - because the system can't checkpoint, recover from a partial failure or pick up where it left off
LLMs are stateless by nature: every new session starts from zero unless the history is explicitly passed in with every call. A neural net that's just plugged in is not an agent. It's a smart intern who has to be told all over again every day where the tasks are, where the context is, where the rules are and why you don't touch prod without checking
What an agent needs to actually work:
✅ Memory - what's done, what failed and why, so it doesn't repeat the same mistakes
✅ Role and permissions - what to take on, what not to touch, where its authority ends
✅ Routing - which task to pick up in which situation and who to hand it off to
✅ Result verification - who confirms the work is done right, and how
✅ Context cleanup - how not to drag old junk into a new task
In 2025 Microsoft released a whitepaper on the taxonomy of agent failures: goal hijacking, tool misuse, memory poisoning, cascading failures in multi-agent systems. This isn't AI-specific - these are classic distributed-systems reliability problems that engineering solved long ago. Anthropic separately published a guide on context management for agents - because context is the agent's operating system, not just a convenience
The model is the engine. The operating system around the agent is the car. Without a steering wheel, brakes and navigation you get a roaring motor on the garage floor. Sounds powerful - impossible to drive
Just a year ago people argued about which prompt to write so the agent would finally work. Now a different question matters more: what the agent remembers, what role it plays, which tasks it's allowed to take and where it puts the result. That's why right now I spend so much time digging not into the models themselves, but into skills, Notion, memory, routing and roles. From the outside it looks like nerding out - "Mat, just give the agent the task". But "just give it the task" works once. If you're lucky. I need repeatability: a system that tomorrow will see the context again, grab the right card, produce the result, put it in the right place and not break the process next door. That's what real autonomy is - not when the agent chats nicely, but when its work can be checked
⭐️ Jeff Bezos, founder of Amazon:
"Good intentions don't work. You have to have a mechanism to make it work"
🎯 Stop waiting for a model that will figure everything out on its own. Start building an environment where the agent has something to figure out - and then every next model will get better automatically