A realistic comparison of Opus and Codex
Watch on YouTube →
Overview
Theo Browne compares Opus 4.6 and Codex 5.3 for coding tasks, finding Codex superior for problem-solving and code quality due to its thoroughness, despite Opus's speed and better front-end design capabilities. He highlights Codex's ability to handle complex migrations, like upgrading Round (ping.gg), and its use in Cursor's long-running feature, while noting Opus's utility in configuring local computer settings and swift code.
Key takeaways
- Codex 5.3 is superior for complex problem-solving and code quality due to its thoroughness, making it ideal for large codebases and migrations.
- Opus 4.6 is faster and better for front-end design, local computer configuration, and tasks where quick, less cautious solutions are acceptable.
- OpenAI's practice of rerouting queries from Codex 5.3 to 5.2 without transparency raises concerns about trust and predictability.
- Theo Browne uses a combination of Codex and Opus, leveraging their strengths by having Codex build and Opus refine UIs, or vice versa.
- Codex's diligence can lead to over-engineering, as seen in Cursor's long-running feature where it generated 85,000 lines of tests.
- Codex tends to explore patterns in existing codebases, making it better at maintaining consistency, while Opus relies more on its training data and is better at starting new projects.
Chapters
0:00
Introduction to Opus vs. Codex and Sponsor: Arcjet
- Theo Browne introduces Opus 4.6 and Codex 5.3 as leading models for coding tasks.
- He mentions Arcjet as a sponsor, highlighting its security features like bot protection and SQL injection prevention.
- Arcjet offers customizable rate limiting using a token bucket system, which Theo Browne found useful for T3 chat.
5:15
Price Comparison: Opus vs. Codex
- Opus 4.6 costs $25 per million tokens out and $5 per million tokens in, but these rates are absurdly expensive.
- Codex 5.3 API pricing is assumed to be similar to 5.2 ($1.75 per million in, $14 per million out), roughly half the price of Opus.
- OpenAI hasn't released Codex 5.3 over API due to security concerns.
10:46
Dashboard Usage and Initial Task Comparison
- Theo Browne's Claude usage is at 3% of weekly and hourly limits after a few days, while Codex is at 8% of weekly usage.
- He tasks Claude Code to review a branch against main and compares it to Codex UI, which he finds more pleasant than the CLI.
- Codex is running on high settings, with extra high rarely needed, and medium sometimes outperforming high.
15:49
Price Analysis and Subscription Plans
- Codeex wins on price per token and subscriptions, but Opus may win on price per run, though this is unconfirmed.
- OpenAI's $20 plan lacks fast inference, unlike Codex's, but Claude's $20 plan can be quickly maxed out.
- Theo Browne finds both $200 subscriptions hard to max out, especially with Codex's app doubling usage.
18:32
Intelligence Comparison: Codex vs. Opus on Hard Problems
- Codex wins in overall intelligence and problem-solving capabilities, but Opus excels in specific scenarios.
- Theo Browne references T3 Chat, where he integrated T3 Canvas, an image generation product.
- Codex ported a shoddy UI and failed during generation, while Opus performed better but missed features.
22:30
Detail Handling and Model Behavior
- Codex doesn't miss details as often as Opus, leading to the saying, "I keep having to bring in Codex to clean up all the Opus slop."
- Opus ignores blockers and trims scope, while Codex embraces and fixes them directly.
- Opus fixed an error in Shu by updating the SDK, breaking other examples, while Codex is more thorough.
25:59
Security Flaws and Environment Variable Management
- Opus failed to manage environment variables in a basic send.js file, requiring manual guidance.
- Theo Browne contrasts this with Opus's ability to handle large code changes, highlighting its inconsistent performance.
- Opus made user ID a nullable field on image generations, a security flaw Codex would likely never introduce.
30:15
Code Migration and Patching with Codex
- Codex successfully migrated Round (ping.gg), an old project with outdated dependencies, by patching packages.
- Codex bumped the first things it wanted to bump, would see things break and it would then patch package the things that were broken temporarily.
- This involved adding and deleting six or seven patches throughout the migration.
34:33
Cursor's Long-Running Feature and Model Thoroughness
- Theo Browne tested Cursor's long-running feature (24-72 hour runs) with Codex 5.3 and Opus 4.6.
- Codex 5.3 generated 85,000 lines of code (mostly tests) over 20 hours, indicating it got stuck in a fix-everything loop.
- Opus completed the same task in 8 minutes and 16 seconds, but with less thoroughness.
42:34
Model Selection and Security Considerations
- Theo Browne prefers Codex for solving real problems and ensuring nothing is left behind.
- Opus is like a caffeine-fueled engineer who just wants to ship, while Codex is a disgruntled engineer scared of outages.
- Codex is more cautious about unsafe tasks, while Opus is more willing to perform them.
45:29
OpenAI's Rerouting and Transparency Issues
- OpenAI reroutes queries from Codex 5.3 to 5.2 if they suspect risky cyber abuse, but doesn't always inform users.
- Theo Browne is okay with rerouting on specific queries if transparent, but not account-wide rerouting.
- Anthropic, unlike OpenAI, tends to ban accounts instead of rerouting queries.
50:14
Front-End Design and Swift Language Performance
- Opus excels in front-end design, often used to fix UIs built by Codex or create mock UIs for Codex to implement.
- Codex struggles with Swift, particularly AppKit and niche UI characteristics, while Opus performs better.
- Theo Browne uses Opus to configure his computers and manage his network, valuing its quickness.
55:13
Security Flaws and Code Audits
- Opus introduced a security flaw by making user ID a nullable field on image generations, which Codex would likely avoid.
- Theo Browne uses both models to audit code, with each finding different issues.
- Codex missed that Opus used V.any, which lets you store arbitrary shaped blobs on a lot of things, in the DB schema, while Opus caught it.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Theo - t3․gg.