FugokuFugoku Docs
Mask

SGLang Quick Start

Deploy SGLang on Fugoku bare-metal GPUs for high-performance LLM inference with RadixAttention

SGLang Quick Start

SGLang is a high-performance inference framework built around the RadixAttention prefix-sharing algorithm. It excels at multi-turn conversations, agentic tool use, structured generation, and complex prompt templates where the same prefix is reused across many requests.

Overview

SGLang exposes an OpenAI-compatible HTTP API (default port 30000) plus a native SGLang frontend for advanced features. Key capabilities:

  • RadixAttention for automatic prefix cache reuse across requests
  • Native structured outputs (JSON mode, regex, EBNF grammars)
  • Tensor + pipeline parallelism across multi-GPU plans
  • Speculative decoding for low-latency single-stream generation
  • Quantization: FP8, AWQ, GPTQ

Prerequisites

  • A Fugoku account with API access
  • An SSH key registered with Fugoku (see SSH Keys)
  • A Hugging Face model ID
  • ~50 GB free disk for model weights and the SGLang Docker image
Model sizeRecommended plansNotes
≤13Bgpu-a100-1, gpu-h100-1Single GPU is sufficient
13B–70Bgpu-a100-2, gpu-h100-2Tensor parallel across 2 GPUs
70B+gpu-h100-4, gpu-h100-8NVLink bandwidth matters

Regions: lagos-1, london-1, frankfurt-1. Recommended image: pytorch-2.0-cuda12 or ubuntu-22.04-cuda12.

Deployment

SGLang publishes official images at lmsysorg/sglang. Launch with python3 -m sglang.launch_server inside the container.

# Provision a single H100
fugoku create instance \
  --name sglang-prod \
  --plan gpu-h100-1 \
  --image pytorch-2.0-cuda12 \
  --region london-1 \
  --ssh-key laptop

# SSH in
fugoku ssh sglang-prod

# Run SGLang (Llama 3.1 8B Instruct)
docker run --rm -it \
  --gpus all \
  --network host \
  --shm-size=16g \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  lmsysorg/sglang:latest \
  python3 -m sglang.launch_server \
    --model-path meta-llama/Llama-3.1-8B-Instruct \
    --host 0.0.0.0 \
    --port 30000 \
    --dtype float16

Once you see The server is fired up and ready to serve!, the API is live on port 30000.

Fugoku Model API

Deploy SGLang as a managed model endpoint via the Fugoku Models API. The platform provisions the GPU server, configures the container, and handles health checks automatically.

# Deploy via API
curl -X POST https://api.fugoku.com/v1/models \
  -H "X-Fugoku-API-Key: $FUGOKU_API_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "clusterId": "cluster-abc123",
    "name": "sglang-llama70b",
    "framework": "sglang",
    "modelId": "meta-llama/Llama-3.1-70B-Instruct",
    "replicas": 1,
    "gpuType": "A100"
  }'

Check deployment status in the console's AI Models tab:

Prerequisites: You need a GPU cluster provisioned in the Fugoku console. See AI Models API for full reference.

Testing

SGLang supports both the OpenAI-compatible endpoint and its native structured-output endpoint.

# OpenAI-compatible chat
curl http://<server-ip>:30000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "meta-llama/Llama-3.1-8B-Instruct",
    "messages": [
      {"role": "user", "content": "Explain RadixAttention in two sentences."}
    ],
    "max_tokens": 200,
    "temperature": 0.7
  }'

# Structured JSON output via the native endpoint
curl http://<server-ip>:30000/generate \
  -H "Content-Type: application/json" \
  -d '{
    "text": "List three French cities as a JSON object with name and population keys.",
    "sampling_params": {
      "max_new_tokens": 200,
      "json_schema": "{\"type\": \"object\", \"properties\": {\"cities\": {\"type\": \"array\", \"items\": {\"type\": \"object\", \"properties\": {\"name\": {\"type\": \"string\"}, \"population\": {\"type\": \"integer\"}}}}}}"
    }
  }'

# Health check
curl http://<server-ip>:30000/health

Performance

  • Prefix caching: RadixAttention is automatic — no flag needed. Reused system prompts and tool schemas hit cache with no recompute.
  • Tensor parallel: Pass --tp N matching the GPU count on multi-GPU plans (e.g. --tp 2 on gpu-h100-2).
  • Speculative decoding: Use --speculative-algo EAGLE with a draft model for 1.5–2× latency reduction on single-stream code completion.
  • Quantization: Enable FP8 with --quantization fp8 on gpu-h100-* plans for 2× throughput at minimal quality cost.
  • Structured outputs: Combine with --json-schema-constraint or grammar constraints for guaranteed-valid JSON.
  • Throughput: Expect 2,500–6,000 tokens/s on a single gpu-h100-1 for 7B-class models when prefixes are shared.

Troubleshooting

CUDA out of memory at startup

  • Lower --mem-fraction-static (try 0.85)
  • Use FP8 quantization: --quantization fp8 (H100 only)
  • Enable tensor parallel across more GPUs

Model download fails

  • Set HUGGING_FACE_HUB_TOKEN for gated models
  • Confirm outbound HTTPS from the GPU server

Structured output is invalid JSON

  • Verify the json_schema is itself valid JSON
  • Use regex or ebnf constraints as a stricter alternative
  • Increase max_new_tokens — truncation can break JSON validity

Port 30000 unreachable

  • Open firewall: fugoku firewalls add-rule sglang-prod --protocol tcp --port 30000 --source 0.0.0.0/0
  • Confirm --network host is passed or the right -p mapping is used

Container exits immediately

  • Inspect docker logs — usually a missing --model-path or OOM at weight load
  • Verify the image tag with docker pull lmsysorg/sglang:latest first

Next Steps

On this page