Qwen 3.8-27B LLM: Local Deployment, Performance Benchmarks, and Optimized Serving

Clip title: Qwen3.8-27B & How to Serve it Fast Author / channel: Sam Witteveen URL: https://www.youtube.com/watch?v=PTuGGdDuyPI

Summary

The video provides a comprehensive overview of the recently released Qwen 3.8-27B large language model, highlighting its capabilities and practical implementation for local AI enthusiasts. While the colossal 2.4 trillion parameter Qwen 3.8 Max model is acknowledged for its power, the video quickly shifts focus to the more accessible 27 billion parameter Qwen 3.8-27B. This smaller version is presented as the highly anticipated model for local deployment, offering significant advantages over its predecessor, Qwen 3.6-27B, and posing a strong challenge to other leading AI models.

Benchmarking results demonstrate Qwen 3.8-27B’s superior performance across various tasks, including coding, agentic operations, and multimodal intelligence. It notably outperforms Meta’s Muse Glimmer 30B and in some vision-related metrics, even surpasses Opus 4.6 Max. The Artificial Analysis Intelligence Index scores Qwen 3.8-27B at 52, placing it on par with models like GLM 5.2 and DeepSeek V4 Pro, and substantially ahead of Qwen 3.6-27B. Furthermore, its Agentic Index score indicates strong reasoning abilities, outperforming GLM 5.2 and certain GPT-5.6 models, which is particularly impressive for a model designed to run on consumer hardware.

A crucial aspect emphasized in the video is that achieving optimal performance with Qwen 3.8-27B locally is not just about choosing the right model version, but also about correctly configuring its “reasoning” settings. The presenter illustrates how different reasoning levels—“off,” “low,” “medium,” and “X-high”—impact token consumption and output quality. While “no thinking” is fast, and “low” thinking offers improved results, “X-high” thinking often leads to an excessive internal monologue, consuming a massive number of tokens and frequently failing to complete tasks. The “medium” reasoning setting is identified as the optimal balance for most use cases.

The video also delves into various practical implementations, testing different quantizations such as BF16 (the original full-resolution version), FP8, and NVFP4 from community contributions like Unsloth, as well as MLX versions for Apple Silicon. Initial tests showed BF16 at around 30 tokens/second, improving to 80-120 tokens/second with FP8 and speculative decoding (MTP=3) without significant quality loss. However, the highest speeds—averaging 200 tokens/second—were achieved by running Qwen 3.8-27B through SGLang using a specific RadixArk NVFP4 quantization and DSpark speculative decoding within a Docker container. This setup efficiently utilizes the model’s full 262K context window, providing a highly responsive local AI experience.

In conclusion, Qwen 3.8-27B is positioned as the leading model for local AI on prosumer hardware, setting a new benchmark for performance and accessibility. The presenter expresses excitement for future developments, including potential “ThinkingCap” versions that offer similar intelligence with fewer tokens, further fine-tunes that enhance capabilities, and fused kernels to unlock even faster inference speeds on existing hardware. The model represents a significant stride towards making powerful AI more broadly available for local applications.

Description

In this video, I look at the long awaited Qwen3.8-27B model. Both what it can do and how to serve it at the maximum tokens per second

DellProPrecision DellProMax DellTech NVIDIA

🤗 HF: https://huggingface.co/collections/Qwen/qwen38 SGLang: https://lmsysorg.mintlify.app/cookbook/autoregressive/Qwen/Qwen3.8-27B

🕵️ Interested in building LLM Agents? Fill out the form below

👨‍💻Github: https://github.com/samwit/llm-tutorials

⏱️Time Stamps: 00:00 Intro 00:50 ThinkingCap 01:25 Qwen3.8 - 27B 01:59 Different Versions on Hugging Face 02:12 Benchmarks 03:14 Artificial Analysis Benchmark 04:10 Qwen3.8-27B on Hugging Face 06:40 Demo 14:17 SGLang

Tags

Qwen 3.8-27B, Qwen3.8-27B, Qwen 3.8, Qwen, Qwen 27B, Alibaba Qwen, new Qwen model, local AI, local LLM, run LLM locally, open weights, self hosted AI, best local LLM, Muse Glimmer 30B, Meta Muse Glimmer, Qwen 3.6-27B, Opus 4.6 Max, best open source LLM, local coding model, agentic coding, computer use AI, LLM benchmarks, Artificial Analysis, LLM quantization, GGUF, AWQ, MLX, fine tunes, tokens per second, how to run Qwen, Hugging Face, DGX Spark, GB10, AI workstation

URLs