Video summary
Paste This Into Claude, Never Hit a Token Limit Again
Main summary
Key takeaways
Core problem addressed
Powerful Claude models (e.g., Fable 5.1) can burn tokens very quickly, causing you to hit token limits and/or pay much higher costs—especially for complex tasks where Claude spins up many internal sub-agents.
Main solution / highest-leverage prompt instruction
Run sub-agents on a cheaper model while keeping Fable 5.1 as the orchestrator.
- Issue: By default, sub-agents spawned for complex tasks may also run on Fable 5.1, leading to huge token usage.
- Example: A customer intelligence report used 67 sub-agents, and the speaker estimates ~80% of tokens came from sub-agents.
- Cost estimate: ~$350 in that example.
- Fix: Add an instruction like:
- “When using sub-agents, always run them on Sonnet 5.”
- Result: With the same task, the speaker reports:
- Almost identical output quality
- ~$44 cost
- Estimated 80–85% cost reduction
Why it works (conceptually)
- Without sub-agents, the orchestrator may have to read/process everything itself, risking:
- hitting the model’s large context window limit sooner (described as 1M token context window)
- spending too many tokens in the “dumb zone” (quality drops as token usage in a chat increases)
- With sub-agents, Claude delegates “dirty work” (reading/research/transcripts) and only returns summaries, letting the orchestrator analyze while still operating in the “smart zone.”
Anthropic guidance is referenced: Fable is best as orchestration, while Sonnet/Opus are better suited for delegated research/reading.
Optional strengthening instruction
Explicitly require sub-agents for research/transcript reading to ensure they’re actually used.
Five additional best practices to avoid token limits
1. Start with a clean context window
A “fresh chat” already consumes tokens (example: 84,000 tokens before prompting).
Main contributors they say you can control:
- MCP tools/connectors
- skills/plugins
Actions:
- In customization settings, disable connectors/MCPs/skills/plugins you don’t use (they’re careful to say you can disable rather than delete).
Further cost-saving approach:
- Use Composio as a meta-connector so Claude connects to one connector rather than loading many separate ones.
2. Pre-plan tasks and prompts (avoid iterative wandering)
Better models perform best when given the entire job upfront.
- Iterating by adding prompts can make the model re-read the full chat each iteration, increasing token cost exponentially.
- Goal: get to ~80% quality immediately, then do only 1–2 small iterations.
- Mindset suggestion: with good upfront prompts, you can often use medium effort rather than high/extra.
3. Use a prompting framework (Anthropic-based)
The speaker describes 4 things to include:
- Goal / entire task
- Intent / “why”
- Guardrails (what could go wrong + rules)
- This is where you can also include the Sonnet sub-agent instruction.
- Define “done” (what the finished output looks like; optionally with an example)
- Prevents overcommitting
If you’re unclear on requirements, they propose:
- An Anthropic/Claude “interview me” skill (asks questions, then drafts the full prompt)
- A WhisperFlow voice transcription “brain dump,” then structure it via a Prompt Master skill
4. Switch effort levels mid-chat
Keep the same model, but reduce reasoning/compute for later tweaks (saves tokens).
- Example: after generating a dashboard/layout, switch to low effort for final UI changes.
- Important update mentioned: with Claude Opus 5 + Fable 5.1, switching effort mid-conversation doesn’t break cache, so Claude doesn’t reread the entire conversation.
5. Use one chat for one task
Don’t reuse a chat for unrelated tasks.
- Later tasks force Claude to ingest irrelevant prior context, increasing tokens and reducing output quality (“dump zone” effect).
Rule of thumb:
- After roughly 300–500 tokens spent in a single task chat, start a new chat even if the task isn’t fully done.
Tools/approaches mentioned:
- /{slash} usage to monitor token usage
- /{slash} context for better overview in Claude Code
- /{slash} compact to summarize chat and continue
- Speaker notes limitations: may not preserve enough context
- Speaker’s own /{slash} refresh skill:
- builds a more comprehensive context prompt for continuing in a new session
- can work in Code Work
Three smaller “avoid these” tips (Anthropic)
- Don’t request redundant verification like “double-check your work / verify twice” (models already do this; leads to extra token use).
- Avoid excessive emphasis (e.g., extreme “maximally thorough/critical,” yelling in caps).
- Avoid prescribed reasoning styles / fixed step process instructions such as:
- “think step-by-step”
- rigid numbered steps
Reported test result: not using these rules led to ~15% less token spent and ~5% more accuracy.
Main speakers/sources
- Speaker: Ben (described as a founder running an AI agency and an AI community)
- Referenced sources: Anthropic (articles/guidance on orchestration and prompting, plus test observations)