Video summary

Why Opus 4.8 Pulled Me Back to Claude

Main summary

Key takeaways

Technology

Tech/Product Takeaways (What This Video Says About Opus 4.8)

Release context

  • Opus 4.8 is framed as a major Anthropic jump—it’s suggested it “could have been Opus 5” because it feels meaningfully better than Opus 4.7.

Overall positioning

  • The speaker claims Opus 4.8 is “top of the pack” on their internal benchmarks.
  • In some areas, results are described as being close to GPT 5.5, and clearly better than 4.7.

Core strengths

  • Writing

    • Rated as the best writing model they’ve tested.
    • Writing benchmark score: 79.6/100 (vs GPT 5.5: 73).
    • Behavior notes: Fast and expressive, with fewer “AI tells”, especially on high reasoning tasks.
    • Voice imitation: Strong at continuing a user’s writing voice from context (e.g., “Continue” + your paragraph keeps your style).
    • Voice/persona work: Especially good for personal/interpersonal and emotionally intelligent writing; described as “best at pushing my own frame.”
  • Knowledge work

    • Strong at producing structured artifacts like slide decks.
    • Example: it generated a beginner slide deck with depth and good styling—the first time the presenter felt an auto-generated deck had “real” richness.
    • Also described as good for mixed threads (e.g., switching between coding + writing in the same conversation).
  • Coding

    • A powerhouse on extra-high reasoning settings.
    • Senior engineer benchmark: 63
      • vs Opus 4.7: about 30 points lower
      • vs GPT 5.5: 62
    • Benchmark method (as described):
      • The model gets a “vibe-coded slop” codebase.
      • It must rewrite from first principles.
      • Output is compared to rewrites by human senior engineers.
    • LFG bench (more realistic style coding tasks):
      • Examples include SaaS/e-commerce and 3D game landscapes.
      • Output quality: code is readable and the model balances engineering quality with creative detail.
      • Example comparison: a 3D “cozy island” scene is described as richer/lush on 4.8, while GPT 5.5 is more diverse but less “vibrant,” more straightforward.

Main caveats

  • Reasoning sensitivity

    • Performance improves substantially with reasoning settings.
    • Best results are emphasized for high and extra high (especially for hardest coding + important writing).
    • Medium is noted as weaker, particularly for writing.
  • Daily-driver limitation: UI/product ecosystem

    • Despite the model strength, the presenter prefers Codex’s desktop app over Claude’s desktop app, due to perceived UI/UX fragmentation.
    • Specific complaint:
      • Claude’s app has multiple tabs (chat/code/cowork) that feel like they’re run by different teams (“shipping an org chart”).
    • Codex is described as faster, simpler, and it includes an in-app browser, which helps knowledge work.
  • Practical recommendation

    • Don’t rely on Claude as your only interface yet.
    • Use Opus 4.8 as part of your toolkit/arsenal.
    • Expect best results using Claude desktop/code with high/extra-high reasoning.

Review / Benchmark Framework Mentioned

  • “Day zero vibe check” and internal testing at Every for about a week.
  • “Reach test”: a simple measure of whether you naturally want to use it (“do you reach for it, and when?”)
    • Speaker: Gold / green (limited by harness/daily-driver UX)
    • Kieran Klassen (GM of Quora): Straight Gold (“paradigm shift”)
    • Katie Parrot (senior staff writer): Green (mostly writing/knowledge work)
  • Senior Engineer benchmark
    • Rewrite from first principles; compared against human senior engineers.
  • LFG bench
    • Real-world style coding tasks (e.g., SaaS/e-commerce, 3D environments).
  • Writing benchmark
    • Multiple writing genres (e.g., intro, promo email, middle paragraphs).
  • Knowledge-work test
    • Specifically slide deck generation quality and depth.

Tutorials / Guides Explicitly Referenced

  • No step-by-step tutorial for Opus itself, but the speaker provides testing guidance:
    • Use High / Extra High reasoning for the hardest coding and most important writing.
    • For best experience, try it in Claude desktop app and Claude code.
    • Compare workflow against Codex desktop app.

Main Speakers / Sources

  • Speaker: Narrator from Every (host/tester; internal voices referenced).
  • Mentioned internal testers/sources:
    • Kieran Klassen — GM of Quora (internal tester; “most human model” comment; “Paradigm shift” grade)
    • Katie Parrot — senior staff writer (internal tester; “reach test” rating)
  • Primary product/company context: Anthropic (Claude Opus models).

Original video