FugokuFugoku Docs
Mask

HuggingFace TGI Quick Start

Deploy HuggingFace Text Generation Inference on Fugoku bare-metal GPUs

HuggingFace TGI Quick Start

Text Generation Inference (TGI) is Hugging Face's reference production server for transformer models. It is tightly integrated with the transformers library, ships with Prometheus metrics, and supports most decoder models on Hugging Face Hub out of the box.

Overview

TGI exposes an OpenAI-compatible /v1/chat/completions endpoint plus a streaming /generate endpoint, with built-in sharded inference, token streaming via Server-Sent Events, and production-grade observability. Key features:

  • HF-native model loading (no extra conversion step)
  • Tensor parallel sharding for multi-GPU plans
  • Quantization via bitsandbytes (INT4/INT8) and GPTQ
  • Prometheus metrics endpoint on port 9000
  • Token streaming via SSE

Prerequisites

  • A Fugoku account with API access
  • An SSH key registered with Fugoku (see SSH Keys)
  • A Hugging Face model ID (public or with a token configured on the server)
  • ~50 GB free disk for model weights and the TGI Docker image
Model sizeRecommended plansNotes
≤13Bgpu-a100-1, gpu-h100-1Single GPU is sufficient
13B–70Bgpu-a100-4, gpu-h100-2Tensor parallel across GPUs
70B+gpu-h100-4, gpu-h100-8NVLink bandwidth matters

Regions: lagos-1, london-1, frankfurt-1. Recommended image: pytorch-2.0-cuda12 or ubuntu-22.04-cuda12.

Deployment

TGI publishes official images at ghcr.io/huggingface/text-generation-inference. The --sharded flag enables multi-GPU tensor parallelism automatically based on --num-shard.

# Provision a single H100
fugoku create instance \
  --name tgi-prod \
  --plan gpu-h100-1 \
  --image pytorch-2.0-cuda12 \
  --region london-1 \
  --ssh-key laptop

# SSH in
fugoku ssh tgi-prod

# Run TGI (Mistral 7B Instruct)
docker run --rm -it \
  --gpus all \
  --network host \
  --shm-size=16g \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  -e HUGGING_FACE_HUB_TOKEN=$HF_TOKEN \
  ghcr.io/huggingface/text-generation-inference:latest \
  --model-id mistralai/Mistral-7B-Instruct-v0.3 \
  --port 8000 \
  --hostname 0.0.0.0 \
  --dtype float16

Once you see Server listening on 0.0.0.0:8000, the API is live.

Fugoku Model API

Deploy TGI as a managed model endpoint via the Fugoku Models API. The platform provisions the GPU server, configures the container, and handles health checks automatically.

# Deploy via API
curl -X POST https://api.fugoku.com/v1/models \
  -H "X-Fugoku-API-Key: $FUGOKU_API_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "clusterId": "cluster-abc123",
    "name": "tgi-llama70b",
    "framework": "tgi",
    "modelId": "meta-llama/Llama-3.1-70B-Instruct",
    "replicas": 1,
    "gpuType": "A100"
  }'

Check deployment status in the console's AI Models tab:

Prerequisites: You need a GPU cluster provisioned in the Fugoku console. See AI Models API for full reference.

Testing

TGI exposes both the OpenAI-compatible chat endpoint and the native /generate streaming endpoint.

# OpenAI-compatible chat
curl http://<server-ip>:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "mistralai/Mistral-7B-Instruct-v0.3",
    "messages": [
      {"role": "user", "content": "What is TGI in two sentences?"}
    ],
    "max_tokens": 150,
    "stream": false
  }'

# Native streaming generation
curl http://<server-ip>:8000/generate \
  -H "Content-Type: application/json" \
  -d '{
    "inputs": "The capital of France is",
    "parameters": {"max_new_tokens": 50}
  }'

# Health check
curl http://<server-ip>:8000/health

Performance

  • Quantization: Use --quantize bitsandbytes or --quantize gptq to fit larger models in fewer GPUs.
  • Tensor parallel: Pass --num-shard N matching the GPU count on multi-GPU plans (e.g. --num-shard 2 on gpu-h100-2).
  • Batching: TGI batches continuously by default; tune --max-batch-size and --max-waiting-tokens for your latency budget.
  • Observability: Scrape Prometheus metrics from port 9000 — TGI exposes per-request latency, token throughput, and queue depth.
  • Throughput: Expect 1,500–3,500 tokens/s on a single gpu-h100-1 for 7B-class models in FP16.

Troubleshooting

CUDA out of memory at startup

  • Enable quantization: --quantize bitsandbytes (4-bit) or gptq
  • Increase --num-shard to spread across more GPUs
  • Lower --max-batch-size and --max-waiting-tokens

Model download hangs or 401

  • Set HUGGING_FACE_HUB_TOKEN for gated models (Llama, Mistral)
  • Check outbound HTTPS from the GPU server

Slow streaming / chunked output

  • Ensure stream: true in the request body for SSE responses
  • Confirm the client uses curl -N (no buffering) or a proper SSE library
  • Check nvidia-smi for GPU saturation

Port 8000 unreachable

  • Open firewall: fugoku firewalls add-rule tgi-prod --protocol tcp --port 8000 --source 0.0.0.0/0
  • Verify the container has --network host or the right -p mapping

Container restart loop

  • Check docker logs for OOM kills — increase plan size or reduce --max-input-length / --max-batch-size

Next Steps

On this page