Qwen 3.8-Max Performance Benchmarks and Agentic Task Evaluation

Clip title: Qwen 3.8 Max tested with ClinePass Author / channel: Luke’s Dev Lab URL: https://www.youtube.com/watch?v=KZ6uQMQtJW4

Summary

Luke then proceeds to evaluate Qwen 3.8-Max through a series of benchmarks, starting with fundamental performance metrics. The model demonstrated excellent prefill speeds, ranging from approximately 640 tokens/second for short prompts to an impressive 16,000-18,000 tokens/second for longer ones. Its decode throughput was consistently strong, maintaining around 260 tokens/second. In memory retrieval tests using the “Needle-in-a-Haystack” method, Qwen 3.8-Max achieved a perfect 100% pass rate across a 256K context depth, successfully finding embedded data at various points within the large context. Furthermore, its reasoning capabilities were robust, with a 94% overall pass rate (45/48 questions), including perfect scores on easy, medium, and hard difficulty levels, and a 75% success rate on expert-level problems. The OpenAI HumanEval benchmark, consisting of 164 hand-written Python challenges, resulted in an 84% pass rate and a 95% answer rate, indicating proficiency in solving coding problems.

Moving to more complex, agentic tasks, Qwen 3.8-Max showed impressive versatility. It successfully generated a functional and aesthetically pleasing Kanban board (front-end web development) and a dynamic “Falling Sand” physics sandbox in zero-shot attempts. For the Dungeon Crawler, Blender (creating a 3D lantern asset), and Godot (developing a 3D platformer game) tasks, the model generally produced good initial results, requiring only one additional prompt each to refine or add specific features like screenshots or fixing minor bugs. The Blender lantern, while functionally complete, was noted for potentially benefiting from more detailed geometry or refined lighting. The Godot platformer was fully functional, featuring player movement, jumping, FOV changes, collectible keys, and respawn points. A bonus 3D dungeon crawler with raycasting and a minimap was also created in Godot after a couple of prompts.

In summary, Qwen 3.8-Max proved to be a powerful, frontier-level AI model with strong performance across a wide spectrum of tasks. While some complex tasks required minor human intervention for optimal results, its ability to generate functional code, develop interactive applications, and create 3D assets in a highly autonomous manner is significant. A key observation was the surprisingly consistent context usage (ranging from approximately 78% to 114% of the available context) across different coding challenges, regardless of task complexity. The presenter concludes that while his tests might not fully capture the model’s extensive capabilities, it demonstrated exceptional potential for tackling intricate challenges, making ClinePass a valuable platform for exploring such advanced open-weight models.

Description

In this video I am excited to try out the best that the Qwen team has to offer.

I run through a few tests which are:

  1. Performance
  2. Memory
  3. Agency
  4. OpenAI Human Eval
  5. Kanban
  6. Sand Physics
  7. Dungeon Crawler
  8. Blender
  9. Godot

Model: https://qwen.ai/blog?id=qwen3.8 HumanEval: https://github.com/openai/human-eval

homelab llamacpp homelab openai qwen #3.8 max cline clinepass

Chapters: 0:00 Intro 0:07 Model 0:37 ClinePass 1:38 Tests Overview 2:20 Performance 3:11 Memory 4:15 Reasoning 4:38 OpenAI HumanEval 5:05 Kanban 6:01 Sand Physics 6:46 Dungeon Crawler 7:23 Blender 8:41 Godot Platformer 10:33 Godot Dungeon Crawler 12:15 Conclusion

Tags

ai, llm, local llm, comparison, benchmarks, qwen, 3.8, max, frontier, cline, clinepass

URLs