Strata
Strata is an open-source runtime designed to enable the efficient execution of large language models (LLMs) on consumer-grade hardware. It leverages advanced quantization and sparse Mixture of Experts (MoE) architectures to bypass traditional VRAM limitations.
Key Capabilities
- 125B Model on 12GB VRAM: Successfully runs the Qwen3.8-Flash-Next (125 billion parameters) on an RTX 5070 GPU with only 12GB of VRAM.
- Sparse MoE Optimization: Utilizes sparse Mixture of Experts to activate only a subset of parameters during inference, drastically reducing memory footprint.
- Consumer Hardware Support: Makes high-capacity AI accessible without requiring enterprise-grade A100/H100 clusters.
Technical Details
- Architecture: Sparse Mixture of Experts (MoE).
- Target Hardware: Consumer GPUs (e.g., RTX 5070).
- Primary Use Case: Local inference of massive models on limited resources.