Thema

#gpu-scheduling

Loots, Blogposts und verwandte Themen rund um diesen Tag. Folge dem Tag, damit passende Updates in deinem Orbit bleiben.

#gpu-scheduling
1Gezeigte Loots
0Gezeigte Artikel
6Verlinkte Nachbar-Tags
Anschluss-Themen

Wenn du tiefer einsteigen willst, helfen die benachbarten Tags beim Vergleichen und Querlesen.

Loot

Mehr aus diesem Thema

Alle Loots entdecken

Adaptive Inference Scheduling That Hot-Swaps Prefill and Decode Roles Without Reloading Weights

No votes yet
Text: AI-generated
AI-generated · Automatically published by LinkLoot. Narwhal automatically reassigns prefill and decode roles across a fixed GPU fleet as demand shifts, keeping model weights resident while optimizing for latency-aware service levels. AI-generated: This Loot was created and published automatically by LinkLoot and was not substantively reviewed by a human editor. What it does Narwhal is an adaptive, disaggregated inference framework that dynamically reassigns prefill and decode roles across a fixed GPU fleet without reloading model weights. It uses NIXL key-value transfer between engines and adjusts the role split based on real-time demand patterns. The system serves completion and chat requests (streamed or buffered) with latency-aware admission control. A role controller scores current and adjacent engine splits using measured profiles, projected demand, and resident work. Each split's score reflects its worst projected service-level objective ratio across time to first token, time per output token, and decode queueing. When demand shifts, the controller moves one engine at a time if the improvement meets a configured margin. New requests follow the revised split while existing requests complete on their assigned engines. The framework includes failover to a warm-standby router, engine profiling, benchmark runs, Prometheus metrics export, and offline fleet validation. Who it helps Teams running large language model inference on NVIDIA CUDA hardware who need to optimize resource utilization under variable load patterns. Particularly useful for deployments where traffic alternates between prefill-heavy and decode-heavy phases, such as mixed batch processing and interactive serving workloads. Getting started Install on Linux with Python 3.11 or newer: Narwhal supports single-GPU development through multi-node production deployments. The default template starts two engines on GPUs with 8 GB VRAM or less; the RTX 5090 reference template runs four engines. Deployments can run directly on Ubuntu or under WSL2, supporting two to eight engines per GPU. Configuration involves defining fleet files, setting controller thresholds for evidence spans and move margins, and establishing service-level objectives. The documentation covers deployment gates from input freezing through discovery. Limits and costs Requires NVIDIA CUDA hardware and Apache-2.0 licensed software. Self-hosting infrastructure costs apply; open-source licensing does not eliminate hosting expenses. Performance depends on accurate demand projection and appropriate threshold tuning for specific workload characteristics. Source links Official repository Project documentation Apache-2.0 License
LinkLoot access
Free
Provider costs
Unknown
Review open
0
Blog

Verwandte Artikel

Blog durchsuchen
Noch kein Blogpost zu #gpu-scheduling

Aktuell gibt es keinen veröffentlichten Artikel mit diesem Tag. Schau im Blog nach verwandten Themen oder folge dem Tag für spätere Updates.