FugokuFugoku Docs
Mask

vLLM Quick Start

Deploy vLLM on Fugoku bare-metal GPUs for high-throughput LLM inference with PagedAttention

vLLM Quick Start

vLLM is a high-throughput, low-latency inference engine for large language models. Its PagedAttention scheduler eliminates KV-cache fragmentation, giving it industry-leading tokens/second under concurrent load. This guide deploys vLLM on a Fugoku bare-metal GPU server.

Overview

vLLM exposes an OpenAI-compatible HTTP API (default port 8000) and supports most Hugging Face causal-LM and multi-modal checkpoints out of the box. Key features:

  • Continuous batching and PagedAttention
  • Tensor parallel and pipeline parallel across multi-GPU plans
  • Hot-swappable LoRA adapters
  • Guided JSON / regex / choice outputs via Outlines
  • Quantization: GPTQ, AWQ, FP8, INT4

Prerequisites

  • A Fugoku account with API access
  • An SSH key registered with Fugoku (see SSH Keys)
  • A Hugging Face model ID (public or with a token configured on the server)
  • ~50 GB free disk for model weights and the vLLM Docker image
Model sizeRecommended plansNotes
≤13B (Llama 3.1 8B, Mistral 7B)gpu-a100-1, gpu-h100-1Single GPU is sufficient
13B–70B (Llama 3.1 70B, Mixtral 8x22B)gpu-a100-4, gpu-h100-2Tensor parallel across GPUs
70B+ (Llama 3.1 405B)gpu-h100-4, gpu-h100-8NVLink bandwidth matters

Regions: lagos-1, london-1, frankfurt-1. Recommended image: pytorch-2.0-cuda12 (full stack) or ubuntu-22.04-cuda12 (minimal).

Deployment

The fastest path: provision a GPU server, then run vLLM's official container with --gpus all and a Hugging Face model.

# Provision a single A100
fugoku create instance \
  --name vllm-prod \
  --plan gpu-a100-1 \
  --image pytorch-2.0-cuda12 \
  --region london-1 \
  --ssh-key laptop

# SSH in
fugoku ssh vllm-prod

# Run vLLM (Llama 3.1 8B Instruct, FP16)
docker run --rm -it \
  --gpus all \
  --network host \
  --shm-size=16g \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  vllm/vllm-openai:latest \
  --model meta-llama/Llama-3.1-8B-Instruct \
  --host 0.0.0.0 \
  --port 8000 \
  --dtype float16

Once you see Application startup complete, the OpenAI-compatible API is live on port 8000.

Fugoku Model API

Deploy vLLM as a managed model endpoint via the Fugoku Models API. The platform provisions the GPU server, configures the container, and handles health checks automatically.

# Deploy via API
curl -X POST https://api.fugoku.com/v1/models \
  -H "X-Fugoku-API-Key: $FUGOKU_API_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "clusterId": "cluster-abc123",
    "name": "vllm-llama70b",
    "framework": "vllm",
    "modelId": "meta-llama/Llama-3.1-70B-Instruct",
    "replicas": 1,
    "gpuType": "A100"
  }'

Check deployment status in the console's AI Models tab:

Prerequisites: You need a GPU cluster provisioned in the Fugoku console. See AI Models API for full reference.

Testing

Verify the endpoint with a chat-completions request. Replace <server-ip> with the public IP from fugoku servers get.

curl http://<server-ip>:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "meta-llama/Llama-3.1-8B-Instruct",
    "messages": [
      {"role": "user", "content": "Explain PagedAttention in one paragraph."}
    ],
    "max_tokens": 200,
    "temperature": 0.7
  }'

You can also list available models:

curl http://<server-ip>:8000/v1/models

Performance

  • Batch size: vLLM batches requests automatically; tune --max-num-seqs (default 256) to balance latency vs throughput.
  • KV cache memory: Use --gpu-memory-utilization 0.9 (default) to leave headroom for activations.
  • Quantization: For 70B+ models on gpu-h100-2, use AWQ (--quantization awq) or FP8 to fit in VRAM.
  • Tensor parallel: Pass --tensor-parallel-size 2 on gpu-h100-2 plans to split the model across both GPUs.
  • Prefix caching: Enable with --enable-prefix-caching to reuse KV cache across requests sharing a common system prompt (massive win for RAG).
  • Throughput: Expect 2,000–5,000 tokens/s on a single gpu-h100-1 for 7B-class models in FP16 under steady load.

Troubleshooting

CUDA out of memory during startup

  • Reduce --gpu-memory-utilization (try 0.85)
  • Use a quantized checkpoint (AWQ, GPTQ, FP8)
  • Increase GPU count via tensor parallel

Model download fails / 401 Unauthorized

  • Public models: ensure outbound HTTPS works from the GPU server
  • Gated models (Llama, Mistral): export HUGGING_FACE_HUB_TOKEN on the server before launching Docker

Slow first-token latency

  • Lower --max-num-seqs to reduce queue wait
  • Enable --enable-prefix-caching if prompts share prefixes
  • Check nvidia-smi for memory pressure or thermal throttling

Port 8000 unreachable from outside

  • Open the firewall: fugoku firewalls add-rule vllm-prod --protocol tcp --port 8000 --source 0.0.0.0/0
  • Confirm fugoku servers get vllm-prod shows the public IP and running state

Container exits immediately

  • Check docker logs <container> — usually a missing model name or OOM at weight load
  • Re-run with --shm-size=16g if you see DataLoader worker errors

Next Steps

On this page