Adaptive Inference Scheduling That Hot-Swaps Prefill and Decode Roles Without Reloading Weights
Narwhal automatically reassigns prefill and decode roles across a fixed GPU fleet as demand shifts, keeping model weights resident while optimizing for latency-aware service levels.
- LinkLoot access
- Free
- Provider costs
- Unknown
What you get from it
What it does
Narwhal is an adaptive, disaggregated inference framework that dynamically reassigns prefill and decode roles across a fixed GPU fleet without reloading model weights. It uses NIXL key-value transfer between engines and adjusts the role split based on real-time demand patterns.
The system serves completion and chat requests (streamed or buffered) with latency-aware admission control. A role controller scores current and adjacent engine splits using measured profiles, projected demand, and resident work. Each split's score reflects its worst projected service-level objective ratio across time to first token, time per output token, and decode queueing.
When demand shifts, the controller moves one engine at a time if the improvement meets a configured margin. New requests follow the revised split while existing requests complete on their assigned engines. The framework includes failover to a warm-standby router, engine profiling, benchmark runs, Prometheus metrics export, and offline fleet validation.
Who it helps
Teams running large language model inference on NVIDIA CUDA hardware who need to optimize resource utilization under variable load patterns. Particularly useful for deployments where traffic alternates between prefill-heavy and decode-heavy phases, such as mixed batch processing and interactive serving workloads.
Getting started
Install on Linux with Python 3.11 or newer:
python -m pip install narwhal-inference
Narwhal supports single-GPU development through multi-node production deployments. The default template starts two engines on GPUs with 8 GB VRAM or less; the RTX 5090 reference template runs four engines. Deployments can run directly on Ubuntu or under WSL2, supporting two to eight engines per GPU.
Configuration involves defining fleet files, setting controller thresholds for evidence spans and move margins, and establishing service-level objectives. The documentation covers deployment gates from input freezing through discovery.
Limits and costs
Requires NVIDIA CUDA hardware and Apache-2.0 licensed software. Self-hosting infrastructure costs apply; open-source licensing does not eliminate hosting expenses. Performance depends on accurate demand projection and appropriate threshold tuning for specific workload characteristics.
Source links
Discussion
Share practical experience, questions, or warnings with the community.
Sign in to join the discussion and vote on comments.
Sign in