Nail-Qwen 35B A3B

Nail-Qwen 35B A3B is a specialized large language model variant based on the Qwen architecture, optimized for efficient local deployment. It is characterized by its specific quantization formats and performance metrics on consumer-grade hardware.

Key Specifications

  • Base Architecture: Qwen 3.6 series
  • Parameter Count: 35B total parameters
  • Active Parameters: 3B (Mixture of Experts / Sparse activation)
  • Quantization: Q4_K_XL (GGUF format)
  • Hardware Target: 16GB VRAM GPUs
  • Format: GGUF-MTP

Evaluation & Performance

Detailed benchmarks and setup guides are available in the following lab note:

Summary of Findings

Based on the evaluation by Luke’s Dev Lab:

  • Quantization: Tested specifically with Q4_K_XL quantization to balance speed and accuracy.
  • Hardware Constraints: Successfully runs on a 16GB GPU setup, making it accessible for mid-range hardware.
  • Capabilities: Comprehensive testing covers:
    • General performance metrics
    • Logical reasoning tasks
    • Code generation and debugging
  • Source: Video analysis by Luke’s Dev Lab provides a detailed walkthrough of the test suite and results.

References