💡 Claude sometimes confidently wraps up an answer on buggy code - and says "done" itself. I used to catch this by eye or when something broke an hour later in a live system. Now before every final answer Codex automatically reads all the edits and gives a verdict.
What Codex is and how the logic works:
Codex is an autonomous cloud agent from OpenAI for checking code, announced in May 2025. It runs in an isolated container in read-only mode: it physically can't change anything, only read and analyze. The internet is off while it works - the agent sees only the code and dependencies from the repository
Two AIs, trained by different teams on different data, systematically make mistakes in different places. Claude wrote the code - Codex looks at the same edits independently. What the first one missed, the second one catches. This isn't about who's smarter - it's the second-pair-of-eyes principle that always works in development. At the end Codex writes in Russian: "VERDICT: clean" or "VERDICT: issues found" with a concrete list in the format file:line - the point
How the system is built - three parts:
📌 Auto-run at the finish - when Claude is about to finish its answer, a script fires. It checks three things: are there real edits in the session, did anything actually change, wasn't Codex just run. If everything checks out - it runs the review and returns the analysis before Claude's final answer
📌 Manual run via a skill - at any moment I ask Claude to run the review right now. Handy when I want to rerun it after fixes, or check a piece of code before the work is done
📌 AGENTS.md in every project - a file with instructions for Codex in the repository root: what's allowed and forbidden, style rules, architectural constraints. An open standard - the counterpart of Anthropic's CLAUDE.md. Codex reads the file itself before every review
What Codex looks for in every review:
✅ Bugs and logic errors that are easy to miss when you write fast
✅ Security holes - leftover passwords, unsafe input, SQL injections
✅ Violations of the project patterns spelled out in AGENTS.md
✅ Edge cases the author missed
In March 2026 Codex Security scanned 1.2 million commits of open-source projects and found about 800 critical and over 10,000 high-severity vulnerabilities in Chromium, OpenSSL and PHP. That's the principle in action: even in code that's been reviewed for years - a second independent reviewer finds what everyone before them missed
Available on the ChatGPT Plus subscription $20/month - the base plan, not a separate tier. Codex works in the background in the cloud and returns a result in 1-30 minutes. On the same complex task it spends ~1.5M tokens vs ~6.2M for Claude Code - faster and lighter for a parallel review
💻 Performance: SWE-bench Pro ~57-59% - on par with Claude Code on hard engineering tasks
⭐️ Eric S. Raymond, open-source developer, "The Cathedral and the Bazaar":
"Given enough eyeballs, all bugs are shallow"
🎯 It used to be that "a second look at the code" meant waiting for a colleague or rereading it yourself an hour later. Now it's 20 seconds while Codex reads the edits and gives a verdict. $20 a month that saves hours of debugging - the simplest investment in my stack I've made in a long time