SGLang Quick Start
Deploy SGLang on Fugoku bare-metal GPUs for high-performance LLM inference with RadixAttention
SGLang Quick Start
SGLang is a high-performance inference framework built around the RadixAttention prefix-sharing algorithm. It excels at multi-turn conversations, agentic tool use, structured generation, and complex prompt templates where the same prefix is reused across many requests.
Overview
SGLang exposes an OpenAI-compatible HTTP API (default port 30000) plus a native SGLang frontend for advanced features. Key capabilities:
- RadixAttention for automatic prefix cache reuse across requests
- Native structured outputs (JSON mode, regex, EBNF grammars)
- Tensor + pipeline parallelism across multi-GPU plans
- Speculative decoding for low-latency single-stream generation
- Quantization: FP8, AWQ, GPTQ
Prerequisites
- A Fugoku account with API access
- An SSH key registered with Fugoku (see SSH Keys)
- A Hugging Face model ID
- ~50 GB free disk for model weights and the SGLang Docker image
Recommended GPU Plans
| Model size | Recommended plans | Notes |
|---|---|---|
| ≤13B | gpu-a100-1, gpu-h100-1 | Single GPU is sufficient |
| 13B–70B | gpu-a100-2, gpu-h100-2 | Tensor parallel across 2 GPUs |
| 70B+ | gpu-h100-4, gpu-h100-8 | NVLink bandwidth matters |
Regions: lagos-1, london-1, frankfurt-1. Recommended image: pytorch-2.0-cuda12 or ubuntu-22.04-cuda12.
Deployment
Docker (recommended)
SGLang publishes official images at lmsysorg/sglang. Launch with python3 -m sglang.launch_server inside the container.
# Provision a single H100
fugoku create instance \
--name sglang-prod \
--plan gpu-h100-1 \
--image pytorch-2.0-cuda12 \
--region london-1 \
--ssh-key laptop
# SSH in
fugoku ssh sglang-prod
# Run SGLang (Llama 3.1 8B Instruct)
docker run --rm -it \
--gpus all \
--network host \
--shm-size=16g \
-v ~/.cache/huggingface:/root/.cache/huggingface \
lmsysorg/sglang:latest \
python3 -m sglang.launch_server \
--model-path meta-llama/Llama-3.1-8B-Instruct \
--host 0.0.0.0 \
--port 30000 \
--dtype float16Once you see The server is fired up and ready to serve!, the API is live on port 30000.
Fugoku Model API
Deploy SGLang as a managed model endpoint via the Fugoku Models API. The platform provisions the GPU server, configures the container, and handles health checks automatically.
# Deploy via API
curl -X POST https://api.fugoku.com/v1/models \
-H "X-Fugoku-API-Key: $FUGOKU_API_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"clusterId": "cluster-abc123",
"name": "sglang-llama70b",
"framework": "sglang",
"modelId": "meta-llama/Llama-3.1-70B-Instruct",
"replicas": 1,
"gpuType": "A100"
}'Check deployment status in the console's AI Models tab:
Prerequisites: You need a GPU cluster provisioned in the Fugoku console. See AI Models API for full reference.
Testing
SGLang supports both the OpenAI-compatible endpoint and its native structured-output endpoint.
# OpenAI-compatible chat
curl http://<server-ip>:30000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "meta-llama/Llama-3.1-8B-Instruct",
"messages": [
{"role": "user", "content": "Explain RadixAttention in two sentences."}
],
"max_tokens": 200,
"temperature": 0.7
}'
# Structured JSON output via the native endpoint
curl http://<server-ip>:30000/generate \
-H "Content-Type: application/json" \
-d '{
"text": "List three French cities as a JSON object with name and population keys.",
"sampling_params": {
"max_new_tokens": 200,
"json_schema": "{\"type\": \"object\", \"properties\": {\"cities\": {\"type\": \"array\", \"items\": {\"type\": \"object\", \"properties\": {\"name\": {\"type\": \"string\"}, \"population\": {\"type\": \"integer\"}}}}}}"
}
}'
# Health check
curl http://<server-ip>:30000/healthPerformance
- Prefix caching: RadixAttention is automatic — no flag needed. Reused system prompts and tool schemas hit cache with no recompute.
- Tensor parallel: Pass
--tp Nmatching the GPU count on multi-GPU plans (e.g.--tp 2ongpu-h100-2). - Speculative decoding: Use
--speculative-algo EAGLEwith a draft model for 1.5–2× latency reduction on single-stream code completion. - Quantization: Enable FP8 with
--quantization fp8ongpu-h100-*plans for 2× throughput at minimal quality cost. - Structured outputs: Combine with
--json-schema-constraintor grammar constraints for guaranteed-valid JSON. - Throughput: Expect 2,500–6,000 tokens/s on a single
gpu-h100-1for 7B-class models when prefixes are shared.
Troubleshooting
CUDA out of memory at startup
- Lower
--mem-fraction-static(try0.85) - Use FP8 quantization:
--quantization fp8(H100 only) - Enable tensor parallel across more GPUs
Model download fails
- Set
HUGGING_FACE_HUB_TOKENfor gated models - Confirm outbound HTTPS from the GPU server
Structured output is invalid JSON
- Verify the
json_schemais itself valid JSON - Use
regexorebnfconstraints as a stricter alternative - Increase
max_new_tokens— truncation can break JSON validity
Port 30000 unreachable
- Open firewall:
fugoku firewalls add-rule sglang-prod --protocol tcp --port 30000 --source 0.0.0.0/0 - Confirm
--network hostis passed or the right-pmapping is used
Container exits immediately
- Inspect
docker logs— usually a missing--model-pathor OOM at weight load - Verify the image tag with
docker pull lmsysorg/sglang:latestfirst
Next Steps
- vLLM Quick Start — highest throughput for plain chat workloads
- HuggingFace TGI Quick Start — HF reference implementation
- Ollama Quick Start — simplest local-first setup
- Model Deployment — production serving patterns
- GPU Plans — full plan catalog