Qwen 3.8-27B Open-Source Multimodal LLM: Agentic Capabilities & DeepSeek Harness
Clip title: Qwen 3.8-27B in Deepseek Harness: This Is the Open-Source AI Stack to Beat Author / channel: Prompt Engineering URL: https://www.youtube.com/watch?v=UK2xbBmxywg
Summary
The video introduces Qwen 3.8-27B, an advanced open-source large language model (LLM) that can run efficiently on consumer-grade hardware. Positioned as a “frontier-adjacent agentic coding” model, it stands out for its compact size (27 billion parameters, fitting within 17GB for a 4-bit quantized version on a 24GB GPU) while offering impressive capabilities. The presenter highlights its strength in image understanding, drawing bounding boxes, and complex reasoning, and demonstrates how to integrate it with the DeepSeek Harness platform to build powerful agentic systems with enhanced observability.
A key demonstration showcased Qwen 3.8’s multimodal capabilities, including the generation of a voxel-based 3D world from a text description and real-time tracking of the International Space Station with high accuracy. The model’s proficiency in computer vision was particularly emphasized through examples such as providing a highly detailed description of an image (a raccoon with a trash hat) and accurately counting cars in an aerial parking lot, even distinguishing colors and drawing precise bounding boxes around them. This ability to not only identify but also “point at” objects without external tools is presented as a novel and powerful feature for a vision-language model.
The video also details the technical aspects of deploying Qwen 3.8, including obtaining the model from Hugging Face (with a specific MLX version for Apple Silicon users) and running it on various platforms like Linux/Windows via vLLM or SGLang. A significant portion is dedicated to configuring Qwen 3.8 within DeepSeek Harness, an agentic development environment built on a plugin system. Harness’s unique “Trajectory” feature provides unprecedented transparency, allowing users to trace every step, system prompt, tool call, and thought process of the AI agent, making debugging and auditing of complex workflows highly efficient.
A critical takeaway from the experiments is the impact of “reasoning effort” settings on Qwen 3.8’s performance. While higher reasoning efforts (Low, Medium, Xhigh) significantly improve the quality and detail of outputs, such as generating a sophisticated website from a simple prompt, they also lead to a substantial increase in token consumption and processing time. Conversely, setting reasoning effort to “Off” results in quick but often generic or “AI slop” outputs. This highlights an important trade-off for users between computational cost and desired output quality. The presenter concludes by expressing deep impressiveness with Qwen 3.8’s overall capabilities, especially its visual analysis and the innovative, transparent design of the DeepSeek Harness platform.
Video Description & Links
Description
Qwen 3.8 27B Locally + DeepSeek Harness: Setup, Reasoning Levels, and Vision Bounding Boxes
LINKS: https://recipes.vllm.ai/Qwen/Qwen3.8-27B https://huggingface.co/mlx-community/Qwen3.8-27B-4bit https://github.com/deepseek-ai/deepseek-harness
Deepseek Harness video: https://www.youtube.com/watch?v=NPO2CwHnfmI
Thanks to @NVIDIADeveloper for DGX Spark. Check it out here: https://nvda.ws/3XIkwsh
My voice to text App: whryte.com
00:00 Qwen 3.8 27B 01:27 DeepSeek Harness 02:05 Server Setup on DGX 02:47 Harness Model Config 03:50 Vision Demo OpenCV 04:51 Counting and Patching 06:45 Bounding Box Results 08:08 ISS Tracker Example 08:36 Reasoning Effort Levels 09:08 Website Prompt Comparison 11:22 Token Cost Breakdown 12:31 Trajectory and Auditing
Tags
prompt engineering, Prompt Engineer, LLMs, AI, artificial Intelligence, Llama, GPT-4, fine-tuning LLMs
URLs
- https://recipes.vllm.ai/Qwen/Qwen3.8-27B
- https://huggingface.co/mlx-community/Qwen3.8-27B-4bit
- https://github.com/deepseek-ai/deepseek-harness
- https://www.youtube.com/watch?v=NPO2CwHnfmI
- https://nvda.ws/3XIkwsh
Related Concepts
- Qwen 3.8-27B
- open-source LLM
- multimodal AI — Wikipedia
- agentic coding — Wikipedia
- consumer-grade hardware
- 4-bit quantization
- 24GB GPU
- DeepSeek Harness
- vision-language model — Wikipedia
- vLLM — Wikipedia
- SGLang — Wikipedia
- Hugging Face — Wikipedia
- Apple Silicon — Wikipedia