Video summary

Paste This Into Claude, Never Hit a Token Limit Again

Main summary

Key takeaways

Technology

Core problem addressed

Powerful Claude models (e.g., Fable 5.1) can burn tokens very quickly, causing you to hit token limits and/or pay much higher costs—especially for complex tasks where Claude spins up many internal sub-agents.

Main solution / highest-leverage prompt instruction

Run sub-agents on a cheaper model while keeping Fable 5.1 as the orchestrator.

  • Issue: By default, sub-agents spawned for complex tasks may also run on Fable 5.1, leading to huge token usage.
    • Example: A customer intelligence report used 67 sub-agents, and the speaker estimates ~80% of tokens came from sub-agents.
    • Cost estimate: ~$350 in that example.
  • Fix: Add an instruction like:
    • “When using sub-agents, always run them on Sonnet 5.”
  • Result: With the same task, the speaker reports:
    • Almost identical output quality
    • ~$44 cost
    • Estimated 80–85% cost reduction

Why it works (conceptually)

  • Without sub-agents, the orchestrator may have to read/process everything itself, risking:
    • hitting the model’s large context window limit sooner (described as 1M token context window)
    • spending too many tokens in the “dumb zone” (quality drops as token usage in a chat increases)
  • With sub-agents, Claude delegates “dirty work” (reading/research/transcripts) and only returns summaries, letting the orchestrator analyze while still operating in the “smart zone.”

Anthropic guidance is referenced: Fable is best as orchestration, while Sonnet/Opus are better suited for delegated research/reading.

Optional strengthening instruction

Explicitly require sub-agents for research/transcript reading to ensure they’re actually used.

Five additional best practices to avoid token limits

1. Start with a clean context window

A “fresh chat” already consumes tokens (example: 84,000 tokens before prompting).

Main contributors they say you can control:

  • MCP tools/connectors
  • skills/plugins

Actions:

  • In customization settings, disable connectors/MCPs/skills/plugins you don’t use (they’re careful to say you can disable rather than delete).

Further cost-saving approach:

  • Use Composio as a meta-connector so Claude connects to one connector rather than loading many separate ones.

2. Pre-plan tasks and prompts (avoid iterative wandering)

Better models perform best when given the entire job upfront.

  • Iterating by adding prompts can make the model re-read the full chat each iteration, increasing token cost exponentially.
  • Goal: get to ~80% quality immediately, then do only 1–2 small iterations.
  • Mindset suggestion: with good upfront prompts, you can often use medium effort rather than high/extra.

3. Use a prompting framework (Anthropic-based)

The speaker describes 4 things to include:

  • Goal / entire task
  • Intent / “why”
  • Guardrails (what could go wrong + rules)
    • This is where you can also include the Sonnet sub-agent instruction.
  • Define “done” (what the finished output looks like; optionally with an example)
    • Prevents overcommitting

If you’re unclear on requirements, they propose:

  • An Anthropic/Claude “interview me” skill (asks questions, then drafts the full prompt)
  • A WhisperFlow voice transcription “brain dump,” then structure it via a Prompt Master skill

4. Switch effort levels mid-chat

Keep the same model, but reduce reasoning/compute for later tweaks (saves tokens).

  • Example: after generating a dashboard/layout, switch to low effort for final UI changes.
  • Important update mentioned: with Claude Opus 5 + Fable 5.1, switching effort mid-conversation doesn’t break cache, so Claude doesn’t reread the entire conversation.

5. Use one chat for one task

Don’t reuse a chat for unrelated tasks.

  • Later tasks force Claude to ingest irrelevant prior context, increasing tokens and reducing output quality (“dump zone” effect).

Rule of thumb:

  • After roughly 300–500 tokens spent in a single task chat, start a new chat even if the task isn’t fully done.

Tools/approaches mentioned:

  • /{slash} usage to monitor token usage
  • /{slash} context for better overview in Claude Code
  • /{slash} compact to summarize chat and continue
    • Speaker notes limitations: may not preserve enough context
  • Speaker’s own /{slash} refresh skill:
    • builds a more comprehensive context prompt for continuing in a new session
    • can work in Code Work

Three smaller “avoid these” tips (Anthropic)

  • Don’t request redundant verification like “double-check your work / verify twice” (models already do this; leads to extra token use).
  • Avoid excessive emphasis (e.g., extreme “maximally thorough/critical,” yelling in caps).
  • Avoid prescribed reasoning styles / fixed step process instructions such as:
    • “think step-by-step”
    • rigid numbered steps

Reported test result: not using these rules led to ~15% less token spent and ~5% more accuracy.

Main speakers/sources

  • Speaker: Ben (described as a founder running an AI agency and an AI community)
  • Referenced sources: Anthropic (articles/guidance on orchestration and prompting, plus test observations)

Original video