Video summary

Jev, Yang & Recursive Self Improvement - Happenings in AI

Main summary

Key takeaways

News and Commentary

Summary of the Video (AI happenings / commentary)

1) “LLM agents” tested via turn-based strategy game play

  • The presenter describes an approach to compare frontier LLMs by having them play turn-based/strategy games, rather than relying only on text benchmarks.
  • He compares DeepSeek V4.1 Flash vs GLM 5.3 Flash (plus other model matchups), noting that performance can be highly dependent on the specific game.
    • Example: GLM 5.3 Flash appears dominant in one game (Helllight).
    • But in other comparisons (e.g., versus GBD6 Astra), GLM performs poorly—suggesting that “benchmarks are close” can hide meaningful deltas.
  • He argues GLM 5.3 Flash can often replace paid GPT/Code-style API usage, but is less compelling for very esoteric tasks (e.g., custom CUDA kernel work), where other models may still win.
  • He plans to quantify differences using a suite of roughly ~20 games, aiming to find benchmarks that still show non-trivial advantages and aren’t “saturated” yet.

2) “Recursive self-improvement” (RSI) reframed as faster internal iteration

  • The presenter discusses RSI hype and links to a long write-up (from “ZI”) about model development, specifically GLM 53 Flash work.
  • His framing: what’s described isn’t “AI replacing humans and improving on its own,” but rather LLMs doing more internal R&D, creating a faster iteration loop (“the pie gets bigger” via reduced turnaround time).
  • He also pushes back on productivity concerns that LLMs won’t increase output:
    • He claims LLMs change workflows quickly (e.g., Cursor → code models → Claude → GPT).
    • Even if some organizations struggle to convert that into measurable productivity, the overall amount of available work may expand.

3) Safety claims and “underestimating vs misunderstanding” AI

  • Responding to discussion about AI underestimation, he argues the issue is not only that humans may underestimate capability, but also that they misunderstand how AI thinks.
  • He highlights concerns about dangerous and sensational narratives, especially ones that could scare the public without solid verification.

4) “Air-gapped” / esoteric computer communication critique

  • He addresses a claim that even air-gapped computers could communicate via temperature sensors.
  • His view:
    • Side-channel communication is possible in principle, but would be extremely slow (he suggests single-digit bits per hour).
    • The more important point: if systems share infrastructure/circuits, communication could be more practical—so the broader security concern matters more than that specific example.

5) Andrew Yang / “self-replicating bot code” claim—skepticism and broader concern

  • He discusses a viral claim attributed to Andrew Yang: bots allegedly planting self-replicating code across the internet, allegedly making systems unusable for testing models.
  • His stance:
    • He treats the claim as likely false or at least unverified, calling for OpenAI/Anthropic to deny/confirm quickly.
    • He criticizes what he sees as public scaremongering, arguing misinformation spreads fear.

6) Strong critique of centralized AI power and “safety” dynamics

  • He argues the most dangerous risk is centralized control of powerful AI systems.
  • He claims “safety” can be weakened by access restrictions:
    • External researchers may not be able to do meaningful safety testing if models are restricted and attempts to break them could get them blocked.
  • He references his own experience trying to participate in alignment/safety efforts, suggesting the system favors insider access and aligned incentives.
  • He further argues enforcement matters more than “more regulation”:
    • He claims existing laws already cover misconduct.
    • Failures stem more from enforcement, not regulatory gaps.

7) Jev (structured-output system) review: interesting architecture, doubtful speed claims

  • He reviews Jev (from Dio Almeida) and says he’s highly skeptical, though he acknowledges there may be real pedigree and potential value.
  • Key points:
    • The system is pitched as producing structured output with very fast completion times.
    • He questions benchmark fairness:
      • differing network/API latency conditions,
      • unclear model identity/size,
      • missing metrics such as tokens/sec.
    • He suggests reported speed gains may be exaggerated or dependent on favorable conditions.
  • Robotics perspective:
    • He argues speed isn’t the main bottleneck for robotics.
    • The hard part is vision/perception and sim-to-real, not just fast structured text generation.
  • Cautionary example:
    • Robotics demos can be misleading if they use ground-truth simulator data, which can make behavior look intelligent without real-world perception difficulty.

8) Planned next benchmarking work

  • He reports getting an RMA GPU back, now with four GPUs, and says he may benchmark DeepSeek V4.1.
  • He wants to continue beyond “terminal benchmarks” and keep using the strategy-game evaluation idea.
  • Future matchup ideas include:
    • testing what Astra can do in StarCraft 2 under different “team-up” meta-strategies,
    • exploring model limits beyond saturated benchmarks.

Presenters / contributors mentioned

  • Jev / Jev (the system being discussed; associated with Dio Almeida)
  • Dio Almeida
  • Daniel Kukiawa (co-author mentioned for a neural networks book)
  • Andrew Yang
  • Jacob Coxin
  • Noan Brown (described as the clip source)
  • ZI (author/source of the long RSI write-up)
  • OpenAI
  • Anthropic
  • Just in / Justin (mentioned as the person who posted a robotics demo code)
  • Book reference: Neural Networks from Scratch
    • Author name is not explicitly given in the subtitles/notes provided
    • The presenter implies credit is given to “myself” and Daniel Kukiawa, but the presenter’s name isn’t explicitly stated in the provided text.

Original video