Video summary

i ran bonsai 2 for 24 hours so you don't have to

Main summary

Key takeaways

Product Review

Product reviewed

Bonsai 2 (PrismML) — described in the video as a ternary quantization method for Qwen 3.8 27B, not a standalone “new model.” The review focuses on whether Bonsai 2 is a good choice for a software engineering agent / coding use case—especially under tight VRAM (8GB) and limited system RAM.


Main features / what Bonsai 2 is

  • Ternary quantization approach

    “Bonsai 2 is an impressive ternary quantization method”

  • Tested against Qwen 3.8 “Bonsai” variants and other Qwen quantizations under controlled resource budgets.

  • Evaluated using the creator’s Bob Bench with a consistent harness, definitions, and timeout.

Benchmark setup & comparisons (apples-to-apples)

  • Evaluation uses 86 tasks per point (wall-clock vs benchmark graph).
  • Runs on a 4090 with same harness, task definitions, grading, and timeout.
  • Bonsai 2 tested across two memory budgets:
    1. Full 4090 VRAM + large context (256K) + higher KV cache quant
      • Used as a high-VRAM comparison baseline (“tons of VRAM” consumption).
    2. ~8GB VRAM / no spare system RAM available scenario.

Key numerical results mentioned

  • Bonsai 2 vs Qwen 3.6 (similar “family” / Q4-ish comparison)

    • Bonsai 2: 75% in 13 hours
    • Qwen 3.6: 90% in much less time
    • Conclusion: Bonsai 2 is “lobotomized” relative to comparable Q4 quant of the same model.
  • 8GB VRAM / no spare system RAM scenario

    • Bonsai 2: 71% in 11 hours
    • Beats Qwen 3.5 9B by 4 points in roughly similar time.
    • The common “98% intelligence retained” claim is not measured by the reviewer (so it may not hold in practice).
  • When system RAM is available (e.g., 10–12GB system RAM)

    • Bonsai 2 is utterly dominated by Qwen 3.6 35B A3B
    • Reviewer states the larger Qwen gets about +10 benchmark points under those conditions.
    • Takeaway: if you can afford system RAM for a larger/higher-quant variant, Bonsai 2 becomes less attractive.

Pros (as stated)

  • Best/strongest option tested within a strict regime:
    • Only ~8GB VRAM
    • No spare system RAM
    • In that constrained case, Bonsai 2 is described as: “the strongest model I have tested fully on Bob Bench wide alpha.”
  • Narrowly beats a baseline:
    • Beats Qwen 3.5 9B by about +4 points at similar time.

Cons (core criticism)

  • Does not perform well as a coding/software engineering agent model compared to better alternatives.
  • Hype may outpace reality:
    • “98% intelligence retained” claims are argued to be unsupported by the reviewer’s measurements.
  • Often significantly worse than comparable Qwen Q4 quant:
    • Example: 75% vs 90%, and much slower.
  • Falls apart when system RAM is available:
    • With 8GB VRAM + extra system RAM, Qwen 35B A3B is said to be far better and faster.
  • Marketing confusion / mischaracterization:
    • Reviewer explicitly says Bonsai 2 is not a new model, but a quantization method for Qwen 3.8 27B—comparing it to branding a product rather than releasing a new model.

Deployment guidance / when to use it (per video)

The reviewer considers 4 environments:

  1. Phone (severely constrained) Implies Bonsai 2 might make more sense, but overall the reviewer is skeptical about edge devices handling hard coding problems.

  2. Very old gaming PC (mainly VRAM-constrained / limited system RAM) Unclear/murky, but may be close versus Qwen 3.5 9B if there is effectively zero spare system RAM.

  3. Dedicated AI rig with limited constraints (e.g., 16–32GB RAM, older GPU) If you have system RAM, Bonsai 2 is not preferred.

  4. High-end setup (5090 / dual GPU) Explicitly says don’t use Bonsai 2—run larger quant models instead.

Final stance on purpose

  • Not recommended for coding/engineering workloads.
  • Possibly hinted use case: chatbots on a phone (but still not presented as a strong coding solution).

Pricing / affordability claim (practical comparison)

  • Suggests an alternative “cheap coding setup”:
    • Used 1070 Ti (8GB VRAM) for about $70–$75
    • Then run Qwen 35B A3B, which he claims will “blow Bonsai 2 out of the water.”

Overall verdict (concise recommendation)

Bonsai 2 is only compelling in a very narrow constraint: ~8GB VRAM with essentially no spare system RAM. In those tight conditions it can be the best among the tested options, but for software engineering / coding—especially when any additional system RAM is available—Bonsai 2 is outperformed by Qwen 3.6 35B A3B (and often by Qwen 3.5 9B).

Recommendation: skip Bonsai 2 for coding/agent use unless you are truly stuck in the 8GB VRAM + no system RAM regime.


Unique points mentioned (all)

  • Bonsai 2 is a ternary quantization method, not a new model.
  • Tested with Bob Bench; 86 tasks per point; consistent harness/timeout/definitions.
  • Under high VRAM baseline: 75% in 13 hours vs Qwen 3.6 at 90% faster → performance gap.
  • “98% retained intelligence” claim is not measured by the reviewer.
  • In strict 8GB VRAM / no system RAM: 71% in 11 hours, beats Qwen 3.5 9B by ~4 points.
  • With extra system RAM (10–12GB): Bonsai 2 is dominated by Qwen 3.6 35B A3B, about +10 points and faster.
  • Reviewer is skeptical of edge compute for the hardest coding problems (e.g., phones/low-RAM systems).
  • For phones: maybe usable for chatbots, not coding.
  • For cheap coding: suggests 1070 Ti (~$70–$75) + Qwen 35B A3B is a better buy.
  • Criticism: hype/marketing may be ahead of real performance.
  • Skepticism: small quantization changes shouldn’t be presented as a “major new model.”

Speaker views

  • Single main speaker/reviewer throughout:
    • Emphasizes benchmark results, comparative performance vs Qwen quantizations, and deployment recommendations by hardware class.
    • Also delivers a strong marketing/positioning critique (Bonsai 2 as quantization rather than a new model).

Original video