vLLM Quick Start
Deploy vLLM on Fugoku bare-metal GPUs for high-throughput LLM inference with PagedAttention
vLLM Quick Start
vLLM is a high-throughput, low-latency inference engine for large language models. Its PagedAttention scheduler eliminates KV-cache fragmentation, giving it industry-leading tokens/second under concurrent load. This guide deploys vLLM on a Fugoku bare-metal GPU server.
Overview
vLLM exposes an OpenAI-compatible HTTP API (default port 8000) and supports most Hugging Face causal-LM and multi-modal checkpoints out of the box. Key features:
- Continuous batching and PagedAttention
- Tensor parallel and pipeline parallel across multi-GPU plans
- Hot-swappable LoRA adapters
- Guided JSON / regex / choice outputs via Outlines
- Quantization: GPTQ, AWQ, FP8, INT4
Prerequisites
- A Fugoku account with API access
- An SSH key registered with Fugoku (see SSH Keys)
- A Hugging Face model ID (public or with a token configured on the server)
- ~50 GB free disk for model weights and the vLLM Docker image
Recommended GPU Plans
| Model size | Recommended plans | Notes |
|---|---|---|
| ≤13B (Llama 3.1 8B, Mistral 7B) | gpu-a100-1, gpu-h100-1 | Single GPU is sufficient |
| 13B–70B (Llama 3.1 70B, Mixtral 8x22B) | gpu-a100-4, gpu-h100-2 | Tensor parallel across GPUs |
| 70B+ (Llama 3.1 405B) | gpu-h100-4, gpu-h100-8 | NVLink bandwidth matters |
Regions: lagos-1, london-1, frankfurt-1. Recommended image: pytorch-2.0-cuda12 (full stack) or ubuntu-22.04-cuda12 (minimal).
Deployment
Docker (recommended)
The fastest path: provision a GPU server, then run vLLM's official container with --gpus all and a Hugging Face model.
# Provision a single A100
fugoku create instance \
--name vllm-prod \
--plan gpu-a100-1 \
--image pytorch-2.0-cuda12 \
--region london-1 \
--ssh-key laptop
# SSH in
fugoku ssh vllm-prod
# Run vLLM (Llama 3.1 8B Instruct, FP16)
docker run --rm -it \
--gpus all \
--network host \
--shm-size=16g \
-v ~/.cache/huggingface:/root/.cache/huggingface \
vllm/vllm-openai:latest \
--model meta-llama/Llama-3.1-8B-Instruct \
--host 0.0.0.0 \
--port 8000 \
--dtype float16Once you see Application startup complete, the OpenAI-compatible API is live on port 8000.
Fugoku Model API
Deploy vLLM as a managed model endpoint via the Fugoku Models API. The platform provisions the GPU server, configures the container, and handles health checks automatically.
# Deploy via API
curl -X POST https://api.fugoku.com/v1/models \
-H "X-Fugoku-API-Key: $FUGOKU_API_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"clusterId": "cluster-abc123",
"name": "vllm-llama70b",
"framework": "vllm",
"modelId": "meta-llama/Llama-3.1-70B-Instruct",
"replicas": 1,
"gpuType": "A100"
}'Check deployment status in the console's AI Models tab:
Prerequisites: You need a GPU cluster provisioned in the Fugoku console. See AI Models API for full reference.
Testing
Verify the endpoint with a chat-completions request. Replace <server-ip> with the public IP from fugoku servers get.
curl http://<server-ip>:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "meta-llama/Llama-3.1-8B-Instruct",
"messages": [
{"role": "user", "content": "Explain PagedAttention in one paragraph."}
],
"max_tokens": 200,
"temperature": 0.7
}'You can also list available models:
curl http://<server-ip>:8000/v1/modelsPerformance
- Batch size: vLLM batches requests automatically; tune
--max-num-seqs(default 256) to balance latency vs throughput. - KV cache memory: Use
--gpu-memory-utilization 0.9(default) to leave headroom for activations. - Quantization: For 70B+ models on
gpu-h100-2, use AWQ (--quantization awq) or FP8 to fit in VRAM. - Tensor parallel: Pass
--tensor-parallel-size 2ongpu-h100-2plans to split the model across both GPUs. - Prefix caching: Enable with
--enable-prefix-cachingto reuse KV cache across requests sharing a common system prompt (massive win for RAG). - Throughput: Expect 2,000–5,000 tokens/s on a single
gpu-h100-1for 7B-class models in FP16 under steady load.
Troubleshooting
CUDA out of memory during startup
- Reduce
--gpu-memory-utilization(try0.85) - Use a quantized checkpoint (AWQ, GPTQ, FP8)
- Increase GPU count via tensor parallel
Model download fails / 401 Unauthorized
- Public models: ensure outbound HTTPS works from the GPU server
- Gated models (Llama, Mistral): export
HUGGING_FACE_HUB_TOKENon the server before launching Docker
Slow first-token latency
- Lower
--max-num-seqsto reduce queue wait - Enable
--enable-prefix-cachingif prompts share prefixes - Check
nvidia-smifor memory pressure or thermal throttling
Port 8000 unreachable from outside
- Open the firewall:
fugoku firewalls add-rule vllm-prod --protocol tcp --port 8000 --source 0.0.0.0/0 - Confirm
fugoku servers get vllm-prodshows the public IP andrunningstate
Container exits immediately
- Check
docker logs <container>— usually a missing model name or OOM at weight load - Re-run with
--shm-size=16gif you seeDataLoader workererrors
Next Steps
- SGLang Quick Start — for structured outputs and multi-turn agents
- HuggingFace TGI Quick Start — reference HF implementation
- Ollama Quick Start — simplest local-first setup
- Model Deployment — production serving patterns
- GPU Plans — full plan catalog