Video summary

GPT-5.5, Claude 4.7, Gemini: Хватит платить за все нейронки!

Main summary

Key takeaways

Product Review

Product reviewed (video’s subject)

The video is not a single product review; it’s a guide to which AI model to use for which task in 2026 (e.g., GPT-5.5, Claude 4.7/4.8, Gemini 3.1 Pro, Grok, DeepSeek, plus media models like GPT Image 2, Gemini/Google image & video tools, Runway, etc.).


Key message / overall recommendation

Don’t try to find “the best model.” In 2026, models are described as highly specialized, so the money-saving strategy is to use a small stack:

  • 1 paid “base” model for ~80% of tasks (often Claude for tech, or Gemini for Google users)
  • Free/cheap add-ons for specific needs (e.g., Grok for real-time X news, DeepSeek for bulk code under privacy)

This combination is claimed to save ~$2,000/year vs relying on a single universal model.


Unique points mentioned (by model/tool)

1) GPT Chat (GPT-5.5, released Apr 23, 2026)

Positioning / features

  • Still described as the most versatile model; “not the best everywhere, but not the worst.”
  • Has:
    • Best overall intelligence rank (cited)
    • 58.6 on SVE bench (code bug/applied tasks); compared to “IQ level”
    • Strong plugin ecosystem + custom GPT
    • Canvas editor for editing long documents inside chat
  • Suggested “morning coffee” entry tasks:
    • starting tasks you don’t know where to begin
    • quick email editing
    • screenshot analysis
    • voice while on the road

Not recommended for

  • Serious code
  • Long documentation
  • Deep, single-topic in-depth work (specialists beat it)

Pros

  • Broad coverage; least likely to be “wrong for the task”
  • Strong tooling (plugins/custom GPT, canvas)

Cons

  • Versatility becomes a minus in 2026 due to specialization
  • Not the top choice for deep specialist tasks

Numerical score(s)

  • 58.6 (SVE bench)

2) Claude (Anthropic) 4.7 / 4.8 (Claude Optimized for text+code)

Positioning / philosophy

  • Presented as a different brand/philosophy:
    • OpenAI = generalist
    • Anthropic = specialist in text and code
  • Claimed market outcome: developer tools (examples given: Cursor, Windsurf) work mainly on Claude.

Performance / capabilities

  • Claude 4.7: 87.6 on SVE bench (called best among all models at the time; “key detail” is Claude is a specialist)
  • Up to 128,000 tokens in one go (unique among “top models” per the video)
    • described as being able to write a full book in one answer without losing style/logic
  • Correction behavior:
    • Claude is described as the only top model likely to explicitly say “you’re wrong”
    • 4.8 improves this further (described as less “overly agreeable”)

Tone / UX

  • Less “praise/agree/nod” behavior than GPT/Gemini
  • “Annoying for first 2 hours,” then “priceless” after users realize its value for serious work

Not recommended for

  • “Everyday trifles”
  • Image generation (not available / not here)
  • No voice communication (voice only in English, per subtitles)

Pros

  • Best-in-class for professional code and serious text
  • Large context output (128k tokens)
  • Direct correction rather than flattery
  • Better fit for long legal docs, complex analysis

Cons

  • Not good for quick casual tasks (as portrayed)
  • Limited voice (English only) and absent multimodal/image capability (per subtitles)

Numerical score(s)

  • 87.6 (SVE bench) for Claude 4.7
  • 128,000 tokens context/output claim

3) Gemini (Google) 3.1 Pro

Performance / features

  • Claims:
    • Best in “pure thinking tests”: 94.1 on GPQA Diamond
    • 1 million token context window (upload a novel / many PDFs / long video)
    • Large context use case: answer across the entire uploaded array
  • Price advantage (claimed):
    • $2 per million input tokens
    • $12 per million output tokens
  • Compared to others:
    • “three times cheaper than Claude”
    • “almost one and a half times cheaper than GPT 5.5” (for comparable quality)

Best use cases

  • Users in the Google ecosystem (Gmail, Docs, Drive, Calendar, YouTube) get a stronger assistant because it can use that context.
  • Also recommended for:
    • large documents/research
    • anything tied to Google services
    • video analysis
    • low-cost, high-volume tasks

Not recommended

  • “High level code” (Claude still better)
  • For “text aesthetics,” Claude is said to be ahead

Pros

  • Huge context (1M tokens)
  • Strong scientific/data-thinking benchmark (GPQA)
  • Very cost-effective, especially for high volume
  • Strong ecosystem integration (per video)

Cons

  • Not the best for high-end coding (in this comparison)
  • Less favored on text aesthetics

Numerical score(s)

  • 94.1 (GPQA Diamond)
  • 1,000,000 tokens context window (claim)
  • Pricing: $2 input / $12 output per 1M tokens

4) Grok (real-time X feed access)

Unique feature

  • Direct access to the X feed in real time
  • Practical result: answers about what’s hot from the last ~10 minutes, not old training data.

Best use cases

  • Real-time IT community trends
  • News, trends, current events, social-media sentiment analysis
  • “Anything relevant to my niche” (implied current events research)

Not recommended

  • Tasks where neutral tone is important (corporate documents)
  • Working with children

Tone / UX

  • Can be joking or sharp; may deliver politically incorrect truth
  • For some, a plus; for others, a reason to avoid

Performance mentions

  • Pure tests: code is near the top:
    • 75% on SVE bench
  • Math:
    • 50.7% on Humanitest Last Exem (described as very difficult)

Pros

  • Best for fresh, real-time information
  • Good enough coding performance (near GPT level, below Claude)

Cons

  • Not reliably neutral; potentially risky tone for sensitive use

Numerical score(s)

  • 75% (SVE bench)
  • 50.7% (Humanitest Last Exem)

5) DeepSeek (open-source, local run)

Key features

  • Open source and “fully open,” downloadable for local execution
  • Claimed benefits:
    • No API fees
    • No corporate data leakage to other servers
    • Full control

Performance / pricing

  • Cited: >80% on SVE bench
  • Pricing stated:
    • $74 per million input tokens
    • Compared with Claude OPUS 4.7 at $15 per million
    • (Video implies it’s “science fiction” expensive in API terms, but makes the case that local running changes the economics.)

Best use cases

  • Large volumes of code
  • Any tasks where privacy is required
  • Bulk data processing

Not recommended

  • Creative texts
  • Emotional storytelling
  • Multimodal image tasks (implied not good / not supported)

Pros

  • Strong privacy/control (local)
  • High code performance (especially in the “bulk code” scenario)
  • Economics reframed around owning a gaming PC

Cons

  • Not for creative/emotional writing and image/multimodal needs (per video)
  • API pricing described as high, but local use is the workaround

Numerical score(s)

  • >80% on SVE bench (DeepSeek)
  • $74 per million input tokens (as stated)

Media models mentioned (not a single “product”)

The video also claims specialized leaders for image/video generation.

Images

  • GPT Image 2: #1 in rankings; “242-point lead” over previous leader
    • best for images with text inside (posters, covers, UI, mockups)
  • Google Nanobona Pro: best portraits / reference work
    • up to 14 photos of a face + Google search during generation
  • Nanobanana 2 Lite (4K): “in 5 seconds” (per subtitles)

Video

  • “42”: best physics
    • videos up to 25 seconds
  • Veo 3.1: cinematic 4K with native sound generated
  • CLН 3.0: first 4K at 60 fps, plus “free plan of 66 credits per day”
  • ECDEN 2.0: up to 12 input files, leader for multi-scene narrative
  • Runway 4.5: directorial control of camera and effects

Overarching media verdict

  • No universal approach in media; each tool covers classes of tasks and trying to do everything with one model fails.

The “$2,000/year” strategy (core recommendation)

The video’s promised pattern:

  • Use one paid base account for about 80% of tasks:
    • often Claude (for techies) or Gemini (for Google ecosystem)
  • Add free access for targeted tasks:
    • Grok via Perplexity (basic plan)
    • DeepSeek via web interface (“webinfйс”)
  • Cost framing:
    • $12/month instead of $80 (as stated)

Important caveat: Which model is “leader” can change every 2–3 weeks, so users should keep checking updates.


Pros vs Cons (as presented overall)

Pros of the strategy

  • Higher quality per task via specialization
  • Significant cost savings (~$2,000/year claim)
  • Better fit for real workflows (docs, code, real-time info, bulk local privacy)

Cons / risks

  • Requires model switching and monitoring (leaders change frequently)
  • Not every model covers every modality (e.g., Claude lacks voice/image per subtitles; Grok tone can be risky)

Verdict (concise)

Recommended approach: Build a small stack rather than paying for a single “universal” model.

  • Pick Claude if your work is heavy coding + serious text.
  • Pick Gemini if you’re in the Google ecosystem and need huge context + cost efficiency.
  • Add Grok for real-time X/news and DeepSeek for privacy/bulk code (often locally).

Overall, the video strongly supports the “right model for the right job” plan and claims it can save ~$2,000/year.


Speakers / perspectives

  • Single speaker throughout (no separate speakers identified in the subtitles). All comparisons and recommendations appear to come from the same tester/narrator describing their 3-week evaluation across models.

Original video