brief products
Hugging Face adds one-command vLLM servers
Hugging Face Jobs now launches a private, OpenAI-compatible vLLM endpoint in one command, billed per second with no servers or Kubernetes to manage.
Posted June 26. The endpoint is OpenAI-compatible, so existing client code points at it with a token swap; Hugging Face bills by the second and frames it for experimentation rather than production, where its separate Inference Endpoints product handles auto-scaling and access controls.
sources 1 cited
1 huggingface.co Deploy a vLLM server on Hugging Face Jobs