Ternary Bonsai 2 27B GGUF 1-bit 2-bit Performance, Memory, Reasoning Evaluation
Clip title: Bonsai 2 27B tested - 16GB Local LLM setup Author / channel: Luke’s Dev Lab URL: https://www.youtube.com/watch?v=ZzLHGHMXkEw
Summary
This video from Luke’s Dev Lab reviews Ternary Bonsai 2 by Prism ML, a new 27B-class reasoning model utilizing ternary transformer weights for extreme quantization. The presenter aims to test Prism ML’s bold claims: that the model is approximately 9.3 times smaller than its FP16 counterpart while retaining an impressive 98.2% of its intelligence, and achieves 47 tokens per second on an Apple M5 Max laptop. The review specifically examines the 1-bit (TQ1_0) and 2-bit (Q2_0) GGUF versions of the model, running them through a comprehensive suite of benchmarks for performance, memory recall, reasoning, and various coding challenges using tools like HumanEval Remix, Kanban, Sand Physics, Dungeon Crawler, Blender, and Godot. The testing environment consists of an Ubuntu Server with an RTX 2000 Ada GPU (16GB VRAM target) and Llama.cpp for model serving.
In terms of raw performance, the 2-bit (Q2_0) version generally showed better prefill throughput, particularly with cached KV, reaching nearly 147,000 tokens/second compared to the 1-bit (TQ1_0) version’s 116,000 tokens/second. However, their decode speeds were almost identical, both averaging around 24-25 tokens per second. The memory recall test, using a “Needle-in-a-Haystack” approach with a 256,608-token context, revealed significant weaknesses in both models; they struggled considerably to retrieve information, especially at 50% context depth, with overall pass rates of 60% (Q1) and 67% (Q2). Reasoning capabilities were a surprising highlight, as both models achieved a 100% pass rate on ‘easy’, ‘medium’, and ‘hard’ difficulty questions, though their performance dropped sharply for ‘expert’ level challenges. For Python coding tasks (HumanEval Remix), both models scored identically with 73 correct answers out of 100, demonstrating a consistent, albeit not outstanding, ability.
However, the models largely faltered in more complex coding and creative tasks. For frontend web development (Kanban), the Q1 version produced a functional but visually clunky app, while Q2 failed entirely. Sand Physics, a mathematical/algorithmic task, resulted in visual glitches for Q1 and a complete failure for Q2 due to code truncation. Dungeon Crawler saw moderate success, generating distinct map styles, but both Q1 and Q2 exhibited issues with raycasting and player movement, consuming a high number of tokens for their output. The Blender and Godot tests, designed for 3D asset creation and game development, were particularly disappointing; neither model managed to save a usable project file or create a functional, playable outcome, despite consuming hundreds of thousands of tokens and multiple “compaction” attempts to optimize context.
The presenter concludes that Ternary Bonsai 2, much like its predecessor, is a significant “letdown.” He explicitly states that Prism ML’s claims of retaining nearly 98.2% of FP16 intelligence are “clearly not true” based on his practical tests. While acknowledging the model’s small footprint allows it to run comfortably on 16GB GPUs, he strongly advises users to opt for the larger, unquantized base model (Qwen 27B) instead. Despite the base model’s potentially slower token generation speed (e.g., 5 tokens/second vs. 25 tokens/second for Bonsai), it often yields dramatically superior results for a similar total wall-clock time, as Bonsai consumes excessive tokens to produce mediocre outputs.
Video Description & Links
Description
In this video we’re going to be look at version 2 of Bonsai from Prism ML. Does this version live up to its big claims?
I run through a few tests which are:
- Performance
- Memory
- Reasoning
- OpenAI Human Eval
- Kanban
- Sand Physics
- Dungeon Crawler
- Blender
- Godot
Model: https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf
localllm localai homelab llamacpp homelab openai qwen 27b #3.8 bonsai ternary prismml
Chapters: 0:00 Intro 0:09 Model 1:34 Tests Overview 2:09 System Specs 2:55 Performance 4:17 Memory 5:18 Reasoning 6:03 HumanEval Remix 6:33 Kanban 8:02 Sand Physics 10:20 Dungeon Crawler 11:37 Blender 13:01 Godot 15:15 Conclusion
Tags
ai, llm, local llm, comparison, benchmarks, qwen, bonsai, ternary, prism, prism-ml, 27b
URLs
Related Concepts
- ternary transformer weights
- extreme quantization
- 1-bit quantization
- 2-bit quantization
- GGUF format
- reasoning model — Wikipedia
- memory efficiency
- inference speed
- Apple M5 chip
- local LLM
- Ternary Bonsai 2
- 1-bit quantization (TQ1_0)
- 2-bit quantization (Q2_0)
- Context window recall
- Local LLM deployment
Related Entities
- Prism ML
- Luke’s Dev Lab
- Gemini 2.5 Flash
- Llama.cpp — Wikipedia
- Ubuntu Server — Wikipedia
- Kanban — Wikipedia
- Dungeon Crawler — Wikipedia