Topic

#disaggregated-inference

Loot, blog posts and adjacent themes connected to this topic. Follow the tag to keep it in your orbit.

#disaggregated-inference
1Shown loot
0Shown articles
6Linked neighbor tags
Topic paths

If you want to go deeper, the adjacent tags are the fastest way to compare and branch into related workflows.

Loot

More from this topic

Explore all loot

Adaptive Inference Scheduling That Hot-Swaps Prefill and Decode Roles Without Reloading Weights

No votes yet
Text: AI-generated
AI-generated · Automatically published by LinkLoot. Narwhal automatically reassigns prefill and decode roles across a fixed GPU fleet as demand shifts, keeping model weights resident while optimizing for latency-aware service levels. AI-generated: This Loot was created and published automatically by LinkLoot and was not substantively reviewed by a human editor. What it does Narwhal is an adaptive, disaggregated inference framework that dynamically reassigns prefill and decode roles across a fixed GPU fleet without reloading model weights. It uses NIXL key-value transfer between engines and adjusts the role split based on real-time demand patterns. The system serves completion and chat requests (streamed or buffered) with latency-aware admission control. A role controller scores current and adjacent engine splits using measured profiles, projected demand, and resident work. Each split's score reflects its worst projected service-level objective ratio across time to first token, time per output token, and decode queueing. When demand shifts, the controller moves one engine at a time if the improvement meets a configured margin. New requests follow the revised split while existing requests complete on their assigned engines. The framework includes failover to a warm-standby router, engine profiling, benchmark runs, Prometheus metrics export, and offline fleet validation. Who it helps Teams running large language model inference on NVIDIA CUDA hardware who need to optimize resource utilization under variable load patterns. Particularly useful for deployments where traffic alternates between prefill-heavy and decode-heavy phases, such as mixed batch processing and interactive serving workloads. Getting started Install on Linux with Python 3.11 or newer: Narwhal supports single-GPU development through multi-node production deployments. The default template starts two engines on GPUs with 8 GB VRAM or less; the RTX 5090 reference template runs four engines. Deployments can run directly on Ubuntu or under WSL2, supporting two to eight engines per GPU. Configuration involves defining fleet files, setting controller thresholds for evidence spans and move margins, and establishing service-level objectives. The documentation covers deployment gates from input freezing through discovery. Limits and costs Requires NVIDIA CUDA hardware and Apache-2.0 licensed software. Self-hosting infrastructure costs apply; open-source licensing does not eliminate hosting expenses. Performance depends on accurate demand projection and appropriate threshold tuning for specific workload characteristics. Source links Official repository Project documentation Apache-2.0 License
LinkLoot access
Free
Provider costs
Unknown
Review open
0
Blog

Related reads

Browse blog
No blog posts for #disaggregated-inference yet

There is no published article with this tag right now. Browse the blog for adjacent themes or follow the tag for future updates.