Video summary
Jev, Yang & Recursive Self Improvement - Happenings in AI
Main summary
Key takeaways
Summary of the Video (AI happenings / commentary)
1) “LLM agents” tested via turn-based strategy game play
- The presenter describes an approach to compare frontier LLMs by having them play turn-based/strategy games, rather than relying only on text benchmarks.
- He compares DeepSeek V4.1 Flash vs GLM 5.3 Flash (plus other model matchups), noting that performance can be highly dependent on the specific game.
- Example: GLM 5.3 Flash appears dominant in one game (Helllight).
- But in other comparisons (e.g., versus GBD6 Astra), GLM performs poorly—suggesting that “benchmarks are close” can hide meaningful deltas.
- He argues GLM 5.3 Flash can often replace paid GPT/Code-style API usage, but is less compelling for very esoteric tasks (e.g., custom CUDA kernel work), where other models may still win.
- He plans to quantify differences using a suite of roughly ~20 games, aiming to find benchmarks that still show non-trivial advantages and aren’t “saturated” yet.
2) “Recursive self-improvement” (RSI) reframed as faster internal iteration
- The presenter discusses RSI hype and links to a long write-up (from “ZI”) about model development, specifically GLM 53 Flash work.
- His framing: what’s described isn’t “AI replacing humans and improving on its own,” but rather LLMs doing more internal R&D, creating a faster iteration loop (“the pie gets bigger” via reduced turnaround time).
- He also pushes back on productivity concerns that LLMs won’t increase output:
- He claims LLMs change workflows quickly (e.g., Cursor → code models → Claude → GPT).
- Even if some organizations struggle to convert that into measurable productivity, the overall amount of available work may expand.
3) Safety claims and “underestimating vs misunderstanding” AI
- Responding to discussion about AI underestimation, he argues the issue is not only that humans may underestimate capability, but also that they misunderstand how AI thinks.
- He highlights concerns about dangerous and sensational narratives, especially ones that could scare the public without solid verification.
4) “Air-gapped” / esoteric computer communication critique
- He addresses a claim that even air-gapped computers could communicate via temperature sensors.
- His view:
- Side-channel communication is possible in principle, but would be extremely slow (he suggests single-digit bits per hour).
- The more important point: if systems share infrastructure/circuits, communication could be more practical—so the broader security concern matters more than that specific example.
5) Andrew Yang / “self-replicating bot code” claim—skepticism and broader concern
- He discusses a viral claim attributed to Andrew Yang: bots allegedly planting self-replicating code across the internet, allegedly making systems unusable for testing models.
- His stance:
- He treats the claim as likely false or at least unverified, calling for OpenAI/Anthropic to deny/confirm quickly.
- He criticizes what he sees as public scaremongering, arguing misinformation spreads fear.
6) Strong critique of centralized AI power and “safety” dynamics
- He argues the most dangerous risk is centralized control of powerful AI systems.
- He claims “safety” can be weakened by access restrictions:
- External researchers may not be able to do meaningful safety testing if models are restricted and attempts to break them could get them blocked.
- He references his own experience trying to participate in alignment/safety efforts, suggesting the system favors insider access and aligned incentives.
- He further argues enforcement matters more than “more regulation”:
- He claims existing laws already cover misconduct.
- Failures stem more from enforcement, not regulatory gaps.
7) Jev (structured-output system) review: interesting architecture, doubtful speed claims
- He reviews Jev (from Dio Almeida) and says he’s highly skeptical, though he acknowledges there may be real pedigree and potential value.
- Key points:
- The system is pitched as producing structured output with very fast completion times.
- He questions benchmark fairness:
- differing network/API latency conditions,
- unclear model identity/size,
- missing metrics such as tokens/sec.
- He suggests reported speed gains may be exaggerated or dependent on favorable conditions.
- Robotics perspective:
- He argues speed isn’t the main bottleneck for robotics.
- The hard part is vision/perception and sim-to-real, not just fast structured text generation.
- Cautionary example:
- Robotics demos can be misleading if they use ground-truth simulator data, which can make behavior look intelligent without real-world perception difficulty.
8) Planned next benchmarking work
- He reports getting an RMA GPU back, now with four GPUs, and says he may benchmark DeepSeek V4.1.
- He wants to continue beyond “terminal benchmarks” and keep using the strategy-game evaluation idea.
- Future matchup ideas include:
- testing what Astra can do in StarCraft 2 under different “team-up” meta-strategies,
- exploring model limits beyond saturated benchmarks.
Presenters / contributors mentioned
- Jev / Jev (the system being discussed; associated with Dio Almeida)
- Dio Almeida
- Daniel Kukiawa (co-author mentioned for a neural networks book)
- Andrew Yang
- Jacob Coxin
- Noan Brown (described as the clip source)
- ZI (author/source of the long RSI write-up)
- OpenAI
- Anthropic
- Just in / Justin (mentioned as the person who posted a robotics demo code)
- Book reference: Neural Networks from Scratch
- Author name is not explicitly given in the subtitles/notes provided
- The presenter implies credit is given to “myself” and Daniel Kukiawa, but the presenter’s name isn’t explicitly stated in the provided text.