Video summary
i ran bonsai 2 for 24 hours so you don't have to
Main summary
Key takeaways
Product reviewed
Bonsai 2 (PrismML) — described in the video as a ternary quantization method for Qwen 3.8 27B, not a standalone “new model.” The review focuses on whether Bonsai 2 is a good choice for a software engineering agent / coding use case—especially under tight VRAM (8GB) and limited system RAM.
Main features / what Bonsai 2 is
-
Ternary quantization approach
“Bonsai 2 is an impressive ternary quantization method”
-
Tested against Qwen 3.8 “Bonsai” variants and other Qwen quantizations under controlled resource budgets.
- Evaluated using the creator’s Bob Bench with a consistent harness, definitions, and timeout.
Benchmark setup & comparisons (apples-to-apples)
- Evaluation uses 86 tasks per point (wall-clock vs benchmark graph).
- Runs on a 4090 with same harness, task definitions, grading, and timeout.
- Bonsai 2 tested across two memory budgets:
- Full 4090 VRAM + large context (256K) + higher KV cache quant
- Used as a high-VRAM comparison baseline (“tons of VRAM” consumption).
- ~8GB VRAM / no spare system RAM available scenario.
- Full 4090 VRAM + large context (256K) + higher KV cache quant
Key numerical results mentioned
-
Bonsai 2 vs Qwen 3.6 (similar “family” / Q4-ish comparison)
- Bonsai 2: 75% in 13 hours
- Qwen 3.6: 90% in much less time
- Conclusion: Bonsai 2 is “lobotomized” relative to comparable Q4 quant of the same model.
-
8GB VRAM / no spare system RAM scenario
- Bonsai 2: 71% in 11 hours
- Beats Qwen 3.5 9B by 4 points in roughly similar time.
- The common “98% intelligence retained” claim is not measured by the reviewer (so it may not hold in practice).
-
When system RAM is available (e.g., 10–12GB system RAM)
- Bonsai 2 is utterly dominated by Qwen 3.6 35B A3B
- Reviewer states the larger Qwen gets about +10 benchmark points under those conditions.
- Takeaway: if you can afford system RAM for a larger/higher-quant variant, Bonsai 2 becomes less attractive.
Pros (as stated)
- Best/strongest option tested within a strict regime:
- Only ~8GB VRAM
- No spare system RAM
- In that constrained case, Bonsai 2 is described as: “the strongest model I have tested fully on Bob Bench wide alpha.”
- Narrowly beats a baseline:
- Beats Qwen 3.5 9B by about +4 points at similar time.
Cons (core criticism)
- Does not perform well as a coding/software engineering agent model compared to better alternatives.
- Hype may outpace reality:
- “98% intelligence retained” claims are argued to be unsupported by the reviewer’s measurements.
- Often significantly worse than comparable Qwen Q4 quant:
- Example: 75% vs 90%, and much slower.
- Falls apart when system RAM is available:
- With 8GB VRAM + extra system RAM, Qwen 35B A3B is said to be far better and faster.
- Marketing confusion / mischaracterization:
- Reviewer explicitly says Bonsai 2 is not a new model, but a quantization method for Qwen 3.8 27B—comparing it to branding a product rather than releasing a new model.
Deployment guidance / when to use it (per video)
The reviewer considers 4 environments:
-
Phone (severely constrained) Implies Bonsai 2 might make more sense, but overall the reviewer is skeptical about edge devices handling hard coding problems.
-
Very old gaming PC (mainly VRAM-constrained / limited system RAM) Unclear/murky, but may be close versus Qwen 3.5 9B if there is effectively zero spare system RAM.
-
Dedicated AI rig with limited constraints (e.g., 16–32GB RAM, older GPU) If you have system RAM, Bonsai 2 is not preferred.
-
High-end setup (5090 / dual GPU) Explicitly says don’t use Bonsai 2—run larger quant models instead.
Final stance on purpose
- Not recommended for coding/engineering workloads.
- Possibly hinted use case: chatbots on a phone (but still not presented as a strong coding solution).
Pricing / affordability claim (practical comparison)
- Suggests an alternative “cheap coding setup”:
- Used 1070 Ti (8GB VRAM) for about $70–$75
- Then run Qwen 35B A3B, which he claims will “blow Bonsai 2 out of the water.”
Overall verdict (concise recommendation)
Bonsai 2 is only compelling in a very narrow constraint: ~8GB VRAM with essentially no spare system RAM. In those tight conditions it can be the best among the tested options, but for software engineering / coding—especially when any additional system RAM is available—Bonsai 2 is outperformed by Qwen 3.6 35B A3B (and often by Qwen 3.5 9B).
Recommendation: skip Bonsai 2 for coding/agent use unless you are truly stuck in the 8GB VRAM + no system RAM regime.
Unique points mentioned (all)
- Bonsai 2 is a ternary quantization method, not a new model.
- Tested with Bob Bench; 86 tasks per point; consistent harness/timeout/definitions.
- Under high VRAM baseline: 75% in 13 hours vs Qwen 3.6 at 90% faster → performance gap.
- “98% retained intelligence” claim is not measured by the reviewer.
- In strict 8GB VRAM / no system RAM: 71% in 11 hours, beats Qwen 3.5 9B by ~4 points.
- With extra system RAM (10–12GB): Bonsai 2 is dominated by Qwen 3.6 35B A3B, about +10 points and faster.
- Reviewer is skeptical of edge compute for the hardest coding problems (e.g., phones/low-RAM systems).
- For phones: maybe usable for chatbots, not coding.
- For cheap coding: suggests 1070 Ti (~$70–$75) + Qwen 35B A3B is a better buy.
- Criticism: hype/marketing may be ahead of real performance.
- Skepticism: small quantization changes shouldn’t be presented as a “major new model.”
Speaker views
- Single main speaker/reviewer throughout:
- Emphasizes benchmark results, comparative performance vs Qwen quantizations, and deployment recommendations by hardware class.
- Also delivers a strong marketing/positioning critique (Bonsai 2 as quantization rather than a new model).