Nail-Qwen 35B A3B
Nail-Qwen 35B A3B is a specialized large language model variant based on the Qwen architecture, optimized for efficient local deployment. It is characterized by its specific quantization formats and performance metrics on consumer-grade hardware.
Key Specifications
- Base Architecture: Qwen 3.6 series
- Parameter Count: 35B total parameters
- Active Parameters: 3B (Mixture of Experts / Sparse activation)
- Quantization: Q4_K_XL (GGUF format)
- Hardware Target: 16GB VRAM GPUs
- Format: GGUF-MTP
Evaluation & Performance
Detailed benchmarks and setup guides are available in the following lab note:
Summary of Findings
Based on the evaluation by Luke’s Dev Lab:
- Quantization: Tested specifically with Q4_K_XL quantization to balance speed and accuracy.
- Hardware Constraints: Successfully runs on a 16GB GPU setup, making it accessible for mid-range hardware.
- Capabilities: Comprehensive testing covers:
- General performance metrics
- Logical reasoning tasks
- Code generation and debugging
- Source: Video analysis by Luke’s Dev Lab provides a detailed walkthrough of the test suite and results.