Video summary

Qwen 3.8 27b x NInfer = 1.5x Faster@256K Context(on one 24GB 4090)

Main summary

Key takeaways

Technology

Video summary (tech concepts, product features, analysis)

The video compares Qwen 3.8 27B inference speed and long-context performance using a new inference engine/framework called NInfer, versus llama.cpp (with mention of MTP). It also contrasts NInfer’s design with vLLM and SGLang, focusing on differences in engineering philosophy and caching/concurrency strategies.


What NInfer is (core idea)

NInfer is described as an inference engine written “from zero”, engineered to be extremely optimized for a single fixed execution path, rather than broad compatibility.

  • It locks to one model family (primarily Qwen 3.8) and one hardware target
  • In the described setup, it’s effectively a 4090-specific fork (“14 90 fork”), optimized for:
    • Qwen 3.8 27B + RTX 4090

How it differs from other engines

  • llama.cpp: prioritizes broad compatibility (e.g., many model formats such as GGUF and many backends). The tradeoff is potential compromises/bugs that can appear under specific conditions.
  • vLLM: focuses on high throughput under concurrency using techniques like:
    • PagedAttention
    • continuous batching
    • prefix caching
  • SGLang: focuses on cache reuse across workflows using:
    • RadixAttention, which builds prefixes into a tree to avoid recomputation—particularly useful for agent-like repetitive behavior.

Key framing: NInfer is positioned as “pure hardware juicing” (like a Formula 1 car on one track), not a general-purpose engine.


Deployment experience described

The creator describes moving from llama.cpp to NInfer through a tuning-and-testing cycle:

  • Previously deployed Qwen 3.8 27B on llama.cpp
    • Spent dozens of hours tuning to reach approximately:
      • ~70–80 tokens/sec at ~200k context
  • After adopting NInfer:
    • Used DeepSeek V4 Flash to run about:
      • 5 hours of configuration/metrics testing for NInfer parameters
    • Produced a best configuration, then compared results against llama.cpp

Reported performance results (RTX 4090, 24GB)

Context size / memory fit

  • NInfer can run up to 256K (≈262K) native context on a 24GB RTX 4090.
  • At 256K, the video claims about ~1.25GB of VRAM remains, enabling additional stacking (e.g., running a vision model alongside).

KV cache approach

The video describes Qwen 3.8 as using a hybrid architecture, enabling large context on 24GB because:

  • KV cache is quantized, not full precision
  • Mentioned approach: RK4V4-E8
    • described as “rotating jumping compression”
    • intended to save VRAM with only small precision loss

It further claims NInfer’s optimization reduces KV cache demands enough for 256K to fit in 24GB.

Tokens/sec numbers

  • Reported average generation speed:
    • ~100 tokens/sec
  • In “real use,” the creator reports:
    • ~130 tokens/sec for small tasks
    • just over ~100 tokens/sec for long chains

Relative speed claim vs llama.cpp + MTP

  • Estimated ~1.5× speedup
  • llama.cpp often lands around ~50–70 tok/s, depending on conditions

Cold-start / prefill trade-off

The creator notes a downside with NInfer: cold prefill is worse.

  • First token latency (TTFT / cold request):
    • llama.cpp: ~5 seconds
    • NInfer: ~12–13 seconds for cold starts
  • Prefill throughput trade-off (cold-start case):
    • not “2× slower”
    • roughly ~20–30% slower prefill throughput

After warm-up and caching behavior improves, performance is said to improve, with TTFT dropping to a few seconds later.


Power / cost analysis (measured)

Power draw

  • Highest-load wall power (during decoding):
    • ~530–540W (NInfer setup)
  • During preview/cache hits:
    • ~200–300W

Power savings and electricity estimate

  • Estimated power savings:
    • ~25–30%
  • Electricity usage estimate:
    • Old (llama.cpp): ~6–7 kWh/day
    • New (NInfer): ~4.5–5 kWh/day
  • Annualized electricity cost estimate:
    • ~$140/year

The creator also claims qualitative benefits tied to measurements:

  • less heat
  • fans spin less

Stability / operational behavior

Startup validation

NInfer is described as more stable in production because it:

  • validates capacity at startup
  • if you exceed budget, it doesn’t start rather than crashing mid-run
  • avoids mid-session OOM behavior

Comparison to llama.cpp

Llama.cpp is described as potentially less stable on long sessions due to KV slot management differences, including:

  • mismatches between unified KV behavior and expectations
  • default slot settings (e.g., “default is four”)
  • needing to pin KV slots to avoid instability as context grows

Limitations mentioned

  • Main downside (beyond cold-start latency):
    • Concurrency trade-off
  • The creator wants full 256K context and is concerned that achieving more concurrency would require giving up context.
  • The video frames this as acceptable because:
    • concurrency on a single GPU with extremely long context isn’t the goal
    • priority is speed and long-chain correctness, not multi-request throughput

Overall conclusion from the creator

  • The creator moved their production setup to NInfer.
  • Claims include:
    • speed improvements made generation more healthy and usable
    • no quality drop in their tests
    • fewer trouble-free sessions than llama.cpp

They also mention broader benchmarking:

  • DeepSeek V4.x “Flash” is still described as even faster
  • but NInfer is preferred for the creator’s specific needs: long-context performance, stability, and their 4090 setup

Main speakers / sources

  • Primary speaker: the YouTube video author/host (creator/operator describing their own deployment and measurements)
  • Referenced tools/engines:
    • NInfer
    • llama.cpp
    • vLLM
    • SGLang
    • DeepSeek V4 Flash (used to test/config-search)
  • Model referenced:
    • Qwen 3.8 27B (and Qwen 3.6 27B mentioned historically)

Rate this summary

Your feedback will help improve summaries.

Improve this summary

Reprocess with a stronger model when the summary feels incomplete or inaccurate.

Pro

Translate summary in another language

Pro

Ask questions to this video

Chat for follow-up questions, clarifications, and source-backed answers.

Coming soon

Share this summary

Original video