Video summary

Opus 5 is my new go-to model

Main summary

Key takeaways

Technology

Summary of Technological Concepts / Product Features / Analysis

What the video claims about Opus 5 (Anthropic)

Performance vs competitors (benchmarks)

The speaker claims Opus 5 ranks highly across multiple evaluations, often beating Fable 5 and outperforming or competing closely with other top models. Specific benchmark examples mentioned:

  • Frontier-style / terminal coding benchmarks: Opus 5 allegedly “crushes” on agentic terminal coding, focusing on maintainability/mergeability (not just unit-test pass rates).
  • Arc AGI 3: Opus 5 allegedly scores extremely high because the benchmark penalizes step count and reasoning length, not only correctness.
  • Agentic Search Through Browse: Opus 5 reportedly scores close to 5.6 Soul.
  • Humanities Last Exam (tools vs no tools): Opus is described as strong with tools, and slightly worse without tools.
  • Deep Squeeze / other code benches: The speaker prefers this bench as more realistic, but notes that results can vary depending on which model/version and benchmark details.

Safety / alignment claims

The video emphasizes Anthropic’s positioning that Opus 5 is the most aligned and safest model Anthropic has made, including claims such as:

  • Better adherence to their “constitution”
  • Lower deceptive behavior
  • Harder to trick into misuse
  • Less prone to reckless irreversible side effects

Use-case targeting

Opus 5 is described as:

  • Built to be used every day
  • More efficient in practice than some alternatives
  • A default model option on Anthropic products (e.g., Claude Pro / Claude Max), with availability depending on plan

Pricing / cost vs “token efficiency” nuance

Marketing vs practical cost

The speaker argues that “half the price” doesn’t fully translate to real savings because:

  • Even though token prices are lower for Opus 5,
  • Opus may use more tokens per task than Fable.

An example cited (from Artificial Analysis):

  • Opus 5: ~37K tokens/task
  • Fable: ~33K tokens/task

The speaker estimates a more realistic real-world discount of ~20–25% off rather than 50%.

Operational effects of higher token use

The video notes that using more tokens can lead to:

  • Slower responses
  • Faster consumption of the context window
  • More chances to “go off track” earlier

Real-world coding experience (the “review” aspect)

The speaker says they’ve:

  • Coded all day with Opus 5
  • Made it their new default/go-to model for tasks expected to be merged into their codebase

They also compare model “personalities” as practical behavior traits:

  • Opus 5: described as diligent and instruction-following, asking clarifying questions, and being “thorough,” with fewer “gimmick workarounds.”
  • Fable 5: described as having more “taste” and producing more pleasant code, but sometimes cutting corners, requiring follow-up to catch edge cases.
  • 5.6 Soul: described as robotic and tool-executing, highly directive and great for automation/one-off tasks—but can write too much code or do things “at all costs.”

Plan/code workflow experiment

The speaker describes a multi-model “plan review” process:

  • Opus and Fable both reviewed each other’s plans and then revised them.
  • A separate “Sol” review allegedly preferred Opus’s plan.

Overall takeaway: Opus is portrayed as better at practical execution quality, not only bench performance.


Subscriptions / limits (Claude plan feature analysis)

The video discusses how Anthropic may allocate usage limits differently by plan:

  • Fable: reportedly gets a smaller share of weekly usage limits (the speaker claims “half” weekly limits for certain plans).
  • Opus: reportedly gets 100% of its limit, making it more cost-effective for heavy usage.

The speaker also gives personal math indicating that Opus consumes weekly limits much more slowly than Fable for similar kinds of work.


“Guide/tutorial” style elements

How to evaluate models yourself

The speaker recommends a practical evaluation approach:

  1. Run the same prompt on Opus and Fable (optionally also Soul).
  2. Have the models review each other’s output.
  3. Use an independent model/thread (e.g., Soul/Codex) to audit results.

This is framed as the best way to understand differences since the speaker says they don’t fully trust third-party coverage.


Sponsors (agent messaging + app backend)

SendTM

A messaging platform that routes messages to SMS/RCS/WhatsApp based on user number preferences. Includes:

  • An MCP server for agents
  • Agent tooling for number lookup, analytics, and messaging via MCP bindings

Convex

A backend/data platform for building apps, emphasizing:

  • Agent-friendly” setup using a folder that describes endpoints/data
  • Easy synchronization across clients
  • Rapid demo building (e.g., a Kanban board) from a single agent request

Main Speakers / Sources

  • Main speaker: The YouTube creator (referred to in subtitles as Theo; a conversation partner “Theo” is also mentioned by name).
  • External sources referenced:
    • Anthropic: official article and Opus 5 claims about alignment, benchmarks, and plan behavior
    • Hank Green: responded on Twitter about Opus 5 “feelings”
    • Ars Technica: criticized a claim related to token efficiency
    • Peter: previously quoted comparison framing between models
    • Theo / Hank Green / Thoric / Sol: mentioned accounts/models; “Sol” appears as a model used in reviews

Original video