Effortless adaptive inference for flexible AI deployment.
Project details
Narwhal is an innovative framework for adaptive, disaggregated inference that optimally manages roles between GPUs without needing to reload model weights. Designed for versatility, it scales seamlessly from single GPU setups to complex multi-node deployments, making it ideal for dynamic AI workloads.
Narwhal is an adaptive, disaggregated inference framework for vLLM. It hot-swaps engines between prefill and decode roles as demand changes, without reloading model weights. It runs on NVIDIA and AMD GPUs and scales from a single GPU to multi-node fleets.
The role controller scores the current prefill/decode split and each adjacent split one engine move away. A split's score is its worst projected SLO ratio across time to first token (TTFT), time per output token (TPOT), and decode queueing. The controller moves only when an adjacent split improves the score by at least a configured margin. New requests follow the revised split, while resident requests finish on their assigned engines.
On Kimi-K3 with six TP8 engines, Narwhal was compared against Dynamo Planner and Ray Serve LLM on identical hardware, using AIPerf to replay a fixed schedule of rapidly changing workloads. With prefix caching on, Narwhal delivered 76.18% of chat/document requests within the latency SLO, against 56.76% for Dynamo Planner and 51.66% for Ray Serve LLM, and had the highest completion and SLO-qualified rates on both workloads. The main trade-off was an 11-second p95 TTFT tail on mixed traffic. Read the full evaluation.
Install on Linux with Python 3.11+:
pip install narwhal-inference
Try it on one GPU with narwhal dev init and narwhal dev up, or follow the gated deployment guide to bring up a multi-node fleet.
Narwhal's scheduling algorithms derive from Arrow: Adaptive Scheduling Mechanisms for Disaggregated LLM Inference Architecture by Wu et al. (2025).
Narwhal is open source under Apache-2.0. See the documentation and GitHub repo.
Comments
0Start the conversation
Share the first comment.