vllm-project/vllm

A high-throughput and memory-efficient inference and serving engine for LLMs

View on GitHub
Python
Stars 86.8k
Forks 19.7k
License Apache-2.0
Open Issues 5905
Updated 1d ago