vllm-project/vllm

Inference
Rank of 3202from #2

Known for

Serving a model to real traffic: batching, paged attention and an OpenAI-shaped endpoint.

The serving layer behind a large share of production inference. Low star count relative to its reach, because the people who depend on it are operators rather than browsers.

76index score · Open Source

How well does it score?

x
92k · top 14%
x
22k · top 8%
x
1.4k · top 5%
x
top 48%

Compared with everyone else on this board.

—votes · crowd
Reading votes…

Votes don’t change the score.

Buy me a coffee