Video summary

Claude Fable 5 - is this Mythos model worth the wait?

Main summary

Key takeaways

Technology

Overview

  • Anthropic launched “Mythos” as a new model class, specifically “Claude Fable 5” (often framed in the video as “baby Mythos”).
  • The creator received early access and tests it with a focus on:
    • whether it matches marketing hype
    • whether it’s practical for their own “backlog”
  • The model is described as state-of-the-art on benchmarks and strong on long, complex, autonomous tasks, but the reviewer flags cost, token intensity, and quality issues in some real-world workflows—especially specs and design.

Product positioning & safety / restrictions

  • “First Mythos-class intelligence model to reach GA”, but the video emphasizes that Fable 5 isn’t full “Mythos (capital M)”:
    • It includes guardrails for cybersecurity and biology (and possibly other sensitive areas).
  • Safety classifiers target categories like:
    • cyber security
    • biology
    • chemistry
    • (and “distillation” is mentioned)
  • Fallback behavior (key feature):
    • If classified into restricted categories, it falls back to Opus 48 instead of hard-blocking.
    • The API supports graceful fallback to Opus 4.8 / Opus pricing for continued handling.
  • Retention / misuse handling:
    • A 30-day retention policy is used for misuse checking, explicitly stated as not used to train Claude.
    • The reviewer notes that most sessions don’t trigger fallback (claims ~95% did not).

Benchmarks & performance claims

  • Anthropic claims Fable 5 exceeds benchmarks by a significant margin.
  • A cited metric: ~80% on SweetBench Pro (compared to other recent models).
  • Compared against: Opus 4.8, GPT-5 (“GPT55” as written), Gemini 3.1 Pro—the video claims Fable 5 is far ahead on the tested benchmark(s).
  • Strength areas Anthropic emphasizes:
    • long, complex tasks
    • autonomy, including asynchronous / days-long task style
    • agentic execution, including sub-agents / dynamic workflows
    • vision
    • “effort”/thoroughness (verifies more)

Cost / efficiency notes (from the reviewer)

  • Pricing is described as very expensive:
    • $10 per input token
    • $50 per output token
  • The reviewer claims token usage is ~2x compared to other models (both faster and higher consumption are implied).
  • They explicitly recommend paying attention to:
    • cost
    • efficiency
  • Testing was often done at higher “effort levels,” which increases token burn.

What it’s good at (reviewer’s hands-on findings)

  1. Vision + document formatting

    • The reviewer is genuinely impressed.
    • Strong at:
      • PDF/document parsing
      • layout and formatting
    • Example mentioned: formatting a handwriting sheet for a child’s work—spacing and readability improved versus other models (an Opus vs Fable 5 comparison is referenced).
  2. Long-horizon / agentic workflows (partially validated)

    • The reviewer could get hours-long runs.
    • They did not personally run it “for many days”, so they couldn’t fully verify the “days” promise, but it seemed capable.
  3. Hard technical problems (where detail matters)

    • The reviewer suggests using it for:
      • detailed engineering / technical reasoning
      • ambitious execution
    • They also note that multi-agent orchestration produced some successes.

What doesn’t work as well (reviewer’s critiques)

  1. “Too engineer-like” thoroughness hurts spec/PRD readability

    • Output is:
      • complete and detailed
      • but hard to “zoom out” and parse
    • Specs/reviews can look “long and intelligent” yet become nearly impossible to read as high-level product documents.
  2. Bad one-shot UI/design quality

    • When asked to design a skills registry UI, the result was fundamentally terrible design (not just generic “AI slop”).
    • Even with more detailed prompting, the design remained disappointing.
    • Conclusion: don’t rely on Fable 5 for front-end design.
  3. May be too conservative about execution/MVP ambition

    • When asked to ship an MVP that gives value to a customer, it produced something too minimal / narrow.
    • This is framed as potentially related to safeguards or general tuning toward caution.
  4. Multi-agent orchestration can stall

    • The reviewer experienced stalls/errors after leaving dev agents running for ~3 hours.
    • They suspect some issues might involve Claude Code / tooling rather than the model, but emphasize that long-running agentic systems must work reliably.

Recommended usage pattern (implied “best fit”)

  • Use Fable 5 for:
    • vision + document formatting
    • hard technical tasks
    • long-horizon, agentic execution
    • cases where thoroughness/detail is needed
  • Prefer other models (e.g., Sonnet / Opus) for:
    • spec writing / PRDs / strategy docs (it over-investigates and outputs dense text)
    • front-end design
  • Potentially use it as:
    • an “advisor” while cheaper models execute (described below under product features).

Launches / API & product features mentioned

  • CloudManage Agents: described as a hosted harness/sandbox for running long-running agentic work.
    • The reviewer is unsure of the best use cases but notes Fable ships with them.
  • New “advisor strategy”:
    • Use Fable 5 as a senior adviser.
    • Use cheaper models as an execution layer.
    • Works in both API and “cloud code.”
  • Fallback API parameter:
    • Allows optional fallback to Opus pricing when requests get blocked/safelisted.

Final takeaway (reviewer’s conclusion)

  • The model is not dismissed (“not a hater”).
  • Best described as:
    • valuable for hard problems and vision/document formatting
    • less ideal for specs/strategy and design
    • potentially expensive and token-intensive
    • useful but requires careful model routing (which model + effort level + task type)

Main speaker/source

  • Creator/host: “How I AI” (the narrator who runs the tests and gives personal evaluation)
  • Source referenced: Anthropic (model release claims, benchmark and safety/fallback details)

Original video