NemoClaw Knowledge Wiki

Tag: llamacpp

11 items with this tag.

  • Jul 22, 2026

    Bonsai 27B vs. Qwen 35B: LLM Performance and Replacement Feasibility Benchmarks

    • localai
    • llm
    • bonsai
    • quantization
    • 1bit
    • qwen
    • llamacpp
    • selfhosted
    • ai
    • rtx3060
  • Jul 22, 2026

    Ternary Bonsai 27B vs. Qwen 27B: LLM Performance Benchmarking Summary

    • localllm
    • localai
    • homelab
    • llamacpp
    • openai
    • qwen
    • 27b
    • bonsai
    • prism-ml
  • Jul 17, 2026

    fahd-mirza

    • local-ai
    • llm-inference
    • fine-tuning
    • speculative-decoding
    • quantization
    • llamacpp
    • unsloth
    • export-controls
    • coding-agents
    • kv-cache
    • structured-data-extraction
    • pdf-processing
    • moe
    • ram-inference
    • model-comparison
    • coding-challenge
  • Jul 13, 2026

    Developing Persistent, Intelligent Memory for Local AI with a Librarian System

    • localai
    • llamacpp
    • edgeai
    • aimemory
    • llm
    • obsidian
    • okf
  • Jul 11, 2026

    gguf-format

    • gguf-format
    • llm-serialization
    • local-inference
    • llamacpp
    • ollama
    • model-distribution
    • binary-format
  • Jun 20, 2026

    Ollama, LM Studio, and llama.cpp: Local AI Tool Comparison and Use Cases

    • ollama
    • lmstudio
    • localllm
    • llamacpp
    • localai
  • Jun 03, 2026

    Adaptive PFlash and Hermes Agent: Self-Tuning LLM Prefill for Long Contexts

    • llamacpp
    • lucebox
    • lucedflash
    • speculativedecoding
    • pflash
  • May 22, 2026

    llama.cpp Router Mode: Native Hot-Swappable Local LLM Switching

    • llamacpp
  • May 20, 2026

    MTP + Ngram Stacked Speculative Decoding in Llama.cpp for LLM Inference

    • llamacpp
    • mtp
    • multitokenprediction
    • speculativedecoding
    • ngrammod
  • May 11, 2026

    Higgsfield: Enabling LLMs like Claude for Media Generation

    • higgsfield
    • llm
    • llamacpp
  • May 10, 2026

    Achieving Fast 35B MoE AI Model Performance on 6GB VRAM with Llama.cpp

    • LocalAI
    • LLM
    • llamacpp
    • Qwen
    • AIonGPU
    • LowVRAM

Created with Quartz v4.5.2 © 2026

  • GitHub
  • Discord Community