NemoClaw Knowledge Wiki

Tag: mtp

4 items with this tag.

  • Jul 12, 2026

    qwen-36-35b-a3b

    • ai
    • llm
    • moe
    • qwen
    • local-inference
    • llama-cpp
    • vram-optimization
    • quantization
    • gguf
    • low-vram
    • coding-agent
    • mtp
  • Jul 04, 2026

    DeepSeek DFlash Accelerates Gemma 12B LLM Text Generation up to 5x

    • deepseek
    • dspark
    • mtp
    • speculativedecoding
    • dflash
  • Jun 30, 2026

    DeepSpec DSparK: Local Qwen3 LLM Acceleration through Speculative Decoding

    • deepseek
    • dspark
    • mtp
    • speculativedecoding
  • May 20, 2026

    MTP + Ngram Stacked Speculative Decoding in Llama.cpp for LLM Inference

    • llamacpp
    • mtp
    • multitokenprediction
    • speculativedecoding
    • ngrammod

Created with Quartz v4.5.2 © 2026

  • GitHub
  • Discord Community