
vllm-project/vllm
Inference
Rank of 3202from #2
Known for
Serving a model to real traffic: batching, paged attention and an OpenAI-shaped endpoint.
The serving layer behind a large share of production inference. Low star count relative to its reach, because the people who depend on it are operators rather than browsers.
76index score · Open Source
How well does it score?
Compared with everyone else on this board.
—votes · crowd
Reading votes…
Votes don’t change the score.