Swift 1.5 Qwen3.8-27B GSQ-RCO IQ3_S 16GB LLM Performance Benchmark

Clip title: Swift 1.5 Qwen3.8-27B · GSQ-RCO tested - 16GB Local LLM setup Author / channel: Luke’s Dev Lab URL: https://www.youtube.com/watch?v=aNOUkWk9piU

Summary

This video provides a comprehensive benchmark of the UkisAI Swift-1.5-Qwen3.8-27B-GSQ-RCO-GGUF model, specifically focusing on its IQ3_S quantization performance on a local Ubuntu server setup with an RTX 2000 Ada (16GB VRAM) GPU. The presenter, Luke, details a wide array of tests designed to evaluate the model’s capabilities in performance, memory, reasoning, and practical coding tasks across various domains including web development (Kanban), simulations (Sand Physics), game development (Dungeon Crawler, Blender, Godot), and Python challenges (HumanEval Remix). The overarching goal is to assess how well this highly quantized model performs on real-world and complex tasks, particularly in scenarios where it can fit entirely within the GPU’s VRAM.

In terms of raw performance, the model demonstrated a significant boost when its context fit entirely within the 16GB VRAM. For a 128K context (requiring overflow to system RAM), the decode throughput was approximately 9.9 tokens/second. However, when the context was reduced to 48K, allowing it to reside fully on the GPU, the decode speed tripled to an impressive 29.9-30.0 tokens/second. Prefill throughput also saw substantial increases, ranging from ~167-234 tokens/second (cold) and ~1100-43000 tokens/second (cached) when overflowing, to ~342-390 tokens/second (cold) and ~2160-88700 tokens/second (cached) when GPU-resident. The memory test, “Needle-in-a-Haystack” with a 256K context, yielded a perfect 100% pass rate, consistently retrieving information regardless of its depth within the context, although the output formatting was sometimes verbose.

The model also showed commendable abilities in logical reasoning and coding. In the reasoning benchmark (48 questions across varying difficulties), it achieved a 73% overall pass rate, scoring 100% on medium questions, 67% on hard, and 33% on expert. Interestingly, it had a minor unusual miss on one easy question (92% pass rate for easy). For the HumanEval Remix, a suite of 100 Python challenges, the model achieved an 86% pass rate and a 100% answer rate, indicating it consistently provided code solutions without refusing prompts.

The agent tests, which involve generating code for practical applications, highlighted the model’s problem-solving capabilities. It successfully created a functional Kanban board, a cellular automata (Sand Physics) simulation with correct liquid/acid flow (after an initial fix with re-prompting), and a 2D dungeon map (also refined after re-prompting to be less empty). Most impressively, in the Blender test, it generated a detailed 3D “Starfall Lantern” asset with minimal context usage (40.3% and zero compactions), and in Godot, it produced a playable 3D platformer game, which, despite initial issues with player movement and respawn logic, was successfully debugged and fixed through further prompting. Overall, the Swift 1.5 GSQ-RCO model demonstrates very strong performance for its quantization level, particularly in complex agentic tasks, making it a promising option for local AI development.

Description

In this video I’m going to test Swift 1.5 Qwen3.8-27B · GSQ-RCO by UkisAI, the claim is - “Swift 1.5 uses 58.5% fewer thinking tokens while scoring 0.35% higher than the base, for a 9.18× speed-up on several tasks”, let’s see if that holds up.

I run through a few tests which are:

  1. Performance
  2. Memory
  3. Reasoning
  4. OpenAI Human Eval
  5. Kanban
  6. Sand Physics
  7. Dungeon Crawler
  8. Blender
  9. Godot

Model: https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27B-GSQ-RCO-GGUF

localllm localai homelab llamacpp homelab openai qwen 27b #3.8 ukisai swift #1.5 gsq rco

Chapters: 0:00 Intro 0:10 Model 0:36 Tests Overview 1:09 System Specs 1:35 Performance 3:03 Memory 3:53 Reasoning 4:39 HumanEval Remix 4:57 Kanban 6:09 Sand Physics 6:55 Dungeon Crawler 7:41 Blender 9:18 Godot 11:38 Conclusion

Tags

ai, llm, local llm, comparison, benchmarks, qwen, 27b, swift, ukisai, 1.5, gsq rco, gsq, rco

URLs