FreeToken Evaluation: Large Language Models on Limited VRAM vs. Llama.cpp
Clip title: FreeToken - first look and test Author / channel: No place like localhost URL: https://www.youtube.com/watch?v=vWGhX3aeFcg
Summary
This video provides a practical evaluation of FreeToken, a new project claiming to run exceptionally large language models efficiently on systems with limited VRAM. The presenter aims to test these ambitious claims, comparing FreeToken’s performance and VRAM utilization against an established solution, llama.cpp, using a 37GB model on a single 12GB GPU setup. The central question is whether FreeToken can truly deliver performant inference under VRAM constraints.
To establish a baseline, the presenter first attempts to run a 37GB Qwen 3.6 model (Q8 quantized) on an RTX 4070 with 12GB of VRAM using llama.cpp. Although llama.cpp successfully loads the model by offloading the bulk of the data to system RAM, performance is significantly compromised. GPU utilization remains low, around 18-19%, resulting in a slow generation speed of approximately 23-24 tokens per second. This performance bottleneck is attributed to the CPU having to constantly transfer data between the slower system RAM and the GPU, leading to the GPU spending most of its time idle.
The video then transitions to FreeToken. The setup involves cloning the FreeToken repository, building it, and using ft bench bw to profile the hardware. A critical step is converting the Hugging Face model into FreeToken’s proprietary .ftw format using ft checkpoint. This conversion process, along with subsequent model serving, presented unexpected challenges, requiring the presenter to manually copy various JSON and Jinja template files from the original Hugging Face repository to address chat templating and tokenization issues. Once these hurdles were overcome, FreeToken served the same 37GB model, still restricted to the 12GB RTX 4070. Impressively, FreeToken achieved a generation speed of 35-36 tokens per second, marking a significant 50% improvement over llama.cpp in the same VRAM-constrained environment, with the GPU now operating at nearly 100% utilization.
In conclusion, FreeToken largely lives up to its claims, demonstrating the ability to host large Mixture-of-Experts (MoE) models on limited VRAM by effectively maximizing GPU utilization. While the initial setup can be complex due to manual file copying requirements, and it necessitates substantial system RAM as a backing store, the performance gains are notable. FreeToken also exposes an OpenAI-compatible API, successfully integrated and tested with OpenCode. However, the project currently has limitations, such as supporting only text-only input for multimodal checkpoints and primarily benefiting MoE models (dense models still require VRAM equal to their size). Despite these caveats, FreeToken provides a legitimate and faster alternative for running large models on consumer-grade hardware, though the presenter anticipates that llama.cpp will likely integrate similar optimizations in the near future given its rapid development.
Video Description & Links
Description
FreeToken is a new project that is making some bold claims about running large or very large models on constrained hardware. Let’s take a look at it, and see what we can do with only 12GB of VRAM at our disposal!
FreeToken: https://github.com/FlashML-org/FreeToken
00:00 Intro 00:21 llama-server baseline 03:34 FreeToken’s turn 09:56 OpenCode 11:19 Verdict 12:05 Caveats 14:21 Outro
Thanks for watching!
URLs
Related Concepts
- Llama.cpp — Wikipedia
- Large Language Models — Wikipedia
- VRAM — Wikipedia
- GPU — Wikipedia
- Model Quantization
- Memory Optimization — Wikipedia
- FreeToken
- VRAM Optimization
- Mixture-of-Experts (MoE)
- Tokenization
- Inference Performance
- Consumer-Grade Hardware
Related Entities
- Llama.cpp — Wikipedia
- Gemini 2.5 Flash
- RTX 4070 — Wikipedia
- OpenCode — Wikipedia
- Hugging Face — Wikipedia
- Discord — Wikipedia
- Patreon — Wikipedia
- CUDA — Wikipedia