Video summary
Opus 5 is my new go-to model
Main summary
Key takeaways
Summary of Technological Concepts / Product Features / Analysis
What the video claims about Opus 5 (Anthropic)
Performance vs competitors (benchmarks)
The speaker claims Opus 5 ranks highly across multiple evaluations, often beating Fable 5 and outperforming or competing closely with other top models. Specific benchmark examples mentioned:
- Frontier-style / terminal coding benchmarks: Opus 5 allegedly “crushes” on agentic terminal coding, focusing on maintainability/mergeability (not just unit-test pass rates).
- Arc AGI 3: Opus 5 allegedly scores extremely high because the benchmark penalizes step count and reasoning length, not only correctness.
- Agentic Search Through Browse: Opus 5 reportedly scores close to 5.6 Soul.
- Humanities Last Exam (tools vs no tools): Opus is described as strong with tools, and slightly worse without tools.
- Deep Squeeze / other code benches: The speaker prefers this bench as more realistic, but notes that results can vary depending on which model/version and benchmark details.
Safety / alignment claims
The video emphasizes Anthropic’s positioning that Opus 5 is the most aligned and safest model Anthropic has made, including claims such as:
- Better adherence to their “constitution”
- Lower deceptive behavior
- Harder to trick into misuse
- Less prone to reckless irreversible side effects
Use-case targeting
Opus 5 is described as:
- Built to be used every day
- More efficient in practice than some alternatives
- A default model option on Anthropic products (e.g., Claude Pro / Claude Max), with availability depending on plan
Pricing / cost vs “token efficiency” nuance
Marketing vs practical cost
The speaker argues that “half the price” doesn’t fully translate to real savings because:
- Even though token prices are lower for Opus 5,
- Opus may use more tokens per task than Fable.
An example cited (from Artificial Analysis):
- Opus 5: ~37K tokens/task
- Fable: ~33K tokens/task
The speaker estimates a more realistic real-world discount of ~20–25% off rather than 50%.
Operational effects of higher token use
The video notes that using more tokens can lead to:
- Slower responses
- Faster consumption of the context window
- More chances to “go off track” earlier
Real-world coding experience (the “review” aspect)
The speaker says they’ve:
- Coded all day with Opus 5
- Made it their new default/go-to model for tasks expected to be merged into their codebase
They also compare model “personalities” as practical behavior traits:
- Opus 5: described as diligent and instruction-following, asking clarifying questions, and being “thorough,” with fewer “gimmick workarounds.”
- Fable 5: described as having more “taste” and producing more pleasant code, but sometimes cutting corners, requiring follow-up to catch edge cases.
- 5.6 Soul: described as robotic and tool-executing, highly directive and great for automation/one-off tasks—but can write too much code or do things “at all costs.”
Plan/code workflow experiment
The speaker describes a multi-model “plan review” process:
- Opus and Fable both reviewed each other’s plans and then revised them.
- A separate “Sol” review allegedly preferred Opus’s plan.
Overall takeaway: Opus is portrayed as better at practical execution quality, not only bench performance.
Subscriptions / limits (Claude plan feature analysis)
The video discusses how Anthropic may allocate usage limits differently by plan:
- Fable: reportedly gets a smaller share of weekly usage limits (the speaker claims “half” weekly limits for certain plans).
- Opus: reportedly gets 100% of its limit, making it more cost-effective for heavy usage.
The speaker also gives personal math indicating that Opus consumes weekly limits much more slowly than Fable for similar kinds of work.
“Guide/tutorial” style elements
How to evaluate models yourself
The speaker recommends a practical evaluation approach:
- Run the same prompt on Opus and Fable (optionally also Soul).
- Have the models review each other’s output.
- Use an independent model/thread (e.g., Soul/Codex) to audit results.
This is framed as the best way to understand differences since the speaker says they don’t fully trust third-party coverage.
Sponsors (agent messaging + app backend)
SendTM
A messaging platform that routes messages to SMS/RCS/WhatsApp based on user number preferences. Includes:
- An MCP server for agents
- Agent tooling for number lookup, analytics, and messaging via MCP bindings
Convex
A backend/data platform for building apps, emphasizing:
- “Agent-friendly” setup using a folder that describes endpoints/data
- Easy synchronization across clients
- Rapid demo building (e.g., a Kanban board) from a single agent request
Main Speakers / Sources
- Main speaker: The YouTube creator (referred to in subtitles as Theo; a conversation partner “Theo” is also mentioned by name).
- External sources referenced:
- Anthropic: official article and Opus 5 claims about alignment, benchmarks, and plan behavior
- Hank Green: responded on Twitter about Opus 5 “feelings”
- Ars Technica: criticized a claim related to token efficiency
- Peter: previously quoted comparison framing between models
- Theo / Hank Green / Thoric / Sol: mentioned accounts/models; “Sol” appears as a model used in reviews