Qwen3.8 27B Turbo Fable Cold Fusion LLM: 16GB Local Performance Benchmark

Clip title: DavidAU Qwen3.8 27B TURBO Fable Cold Fusion tested - 16GB Local LLM setup Author / channel: Luke’s Dev Lab URL: https://www.youtube.com/watch?v=jNVl7TMd2DQ

Summary

This video provides a detailed benchmark of the large language model “Quen 3.8 27B Turbo Fable Cold Fusion 735 882 Heretic Uncensored Neo Coder Max MTP GGUF” from David AU. The presenter, from Luke’s Dev Lab, evaluates the model’s performance across various computational and creative tasks, using a system featuring an RTX 2000 Ada with 16GB VRAM. Initial claims suggest the model is superior to its predecessor, uses fewer tokens, and performs well even in a 4-bit quantized version (Q4_K_M with MTP was used for testing).

The benchmark suite covered performance metrics like prefill and decode speeds, memory retrieval (“Needle-in-a-Haystack”), reasoning, and coding challenges using OpenAI’s HumanEval. The performance tests showed decent prefill speeds (around 120-123 tok/s non-cached, up to 32,000 tok/s cached) and decode speeds (7.2-8.1 tok/s on the RTX 2000, estimated double on gaming GPUs). However, the memory test was notably poor, with only a 47% overall pass rate and accuracy dropping to 0% at 50% and 100% context depth. Reasoning skills were “about standard” with an 83% pass rate, while HumanEval (Python challenges) scored a 62% overall pass rate, primarily because the model failed to answer 63 out of 164 questions, indicating a tendency to “overthink.”

For agent-based coding tasks, the model demonstrated a mixed but interesting performance. It excelled in front-end web development (Kanban) and algorithmic dungeon generation (Dungeon Crawler), fixing initial issues efficiently and with low token usage (19.6% and 10.7% context respectively for single-shot fixes). This highlights its strength in quickly identifying and resolving problems within existing codebases. In contrast, Sand Physics delivered an “okay” but imperfect result after two prompts, with minor behavioral issues persisting. More complex 3D tasks were challenging: Blender produced a visually acceptable asset (a lantern) with very low context usage (18.9%), but it lacked refinement in lighting, textures, and material properties like glass.

The most demanding task, creating a 3D platformer game in Godot, proved to be beyond the model’s current capabilities. Despite extensive prompting and compacting the context twice, the model continuously got stuck in loops, failed to implement basic physics (the player character consistently fell through the floor), and could not effectively diagnose or resolve the issue, ultimately consuming 44.5% of the context without a working solution.

In conclusion, the Quen model is an intriguing offering. While it struggles significantly with deep memory recall and complex game engine development, its standout feature is its efficiency in quickly identifying and fixing problems in certain coding scenarios, particularly in front-end web development and algorithmic generation. This suggests it could be a valuable tool for bug fixing and improving existing codebases, rather than for highly creative or multi-faceted development tasks that require deeper understanding or more robust physics implementation.

Description

Lots of big claims on this latest finetune from DavidAU, super smart yet burning much less tokens, can it be true?

I run through a few tests which are:

  1. Performance
  2. Memory
  3. Reasoning
  4. OpenAI Human Eval
  5. Kanban
  6. Sand Physics
  7. Dungeon Crawler
  8. Blender
  9. Godot

Model: https://huggingface.co/DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NEO-CODER-MAX-MTP-GGUF HumanEval: https://github.com/openai/human-eval

localllm localai homelab llamacpp homelab openai qwen 27b #3.8 davidau fable fusion uncensored cold

Chapters: 0:00 Intro 0:22 Model 1:19 Tests Overview 1:52 System Spec 2:27 Performance 3:35 Memory 4:46 Reasoning 5:12 OpenAI HumanEval 5:52 Kanban 6:39 Sand Physics 7:47 Dungeon Crawler 8:58 Blender 10:56 Godot 13:50 Conclusion

Tags

ai, llm, local llm, comparison, benchmarks, qwen, 3.8, 27b, davidau, fable, cold, fusion, uncesored, heretic

URLs