HuggingFace TGI Quick Start
Deploy HuggingFace Text Generation Inference on Fugoku bare-metal GPUs
HuggingFace TGI Quick Start
Text Generation Inference (TGI) is Hugging Face's reference production server for transformer models. It is tightly integrated with the transformers library, ships with Prometheus metrics, and supports most decoder models on Hugging Face Hub out of the box.
Overview
TGI exposes an OpenAI-compatible /v1/chat/completions endpoint plus a streaming /generate endpoint, with built-in sharded inference, token streaming via Server-Sent Events, and production-grade observability. Key features:
- HF-native model loading (no extra conversion step)
- Tensor parallel sharding for multi-GPU plans
- Quantization via
bitsandbytes(INT4/INT8) and GPTQ - Prometheus metrics endpoint on port
9000 - Token streaming via SSE
Prerequisites
- A Fugoku account with API access
- An SSH key registered with Fugoku (see SSH Keys)
- A Hugging Face model ID (public or with a token configured on the server)
- ~50 GB free disk for model weights and the TGI Docker image
Recommended GPU Plans
| Model size | Recommended plans | Notes |
|---|---|---|
| ≤13B | gpu-a100-1, gpu-h100-1 | Single GPU is sufficient |
| 13B–70B | gpu-a100-4, gpu-h100-2 | Tensor parallel across GPUs |
| 70B+ | gpu-h100-4, gpu-h100-8 | NVLink bandwidth matters |
Regions: lagos-1, london-1, frankfurt-1. Recommended image: pytorch-2.0-cuda12 or ubuntu-22.04-cuda12.
Deployment
Docker (recommended)
TGI publishes official images at ghcr.io/huggingface/text-generation-inference. The --sharded flag enables multi-GPU tensor parallelism automatically based on --num-shard.
# Provision a single H100
fugoku create instance \
--name tgi-prod \
--plan gpu-h100-1 \
--image pytorch-2.0-cuda12 \
--region london-1 \
--ssh-key laptop
# SSH in
fugoku ssh tgi-prod
# Run TGI (Mistral 7B Instruct)
docker run --rm -it \
--gpus all \
--network host \
--shm-size=16g \
-v ~/.cache/huggingface:/root/.cache/huggingface \
-e HUGGING_FACE_HUB_TOKEN=$HF_TOKEN \
ghcr.io/huggingface/text-generation-inference:latest \
--model-id mistralai/Mistral-7B-Instruct-v0.3 \
--port 8000 \
--hostname 0.0.0.0 \
--dtype float16Once you see Server listening on 0.0.0.0:8000, the API is live.
Fugoku Model API
Deploy TGI as a managed model endpoint via the Fugoku Models API. The platform provisions the GPU server, configures the container, and handles health checks automatically.
# Deploy via API
curl -X POST https://api.fugoku.com/v1/models \
-H "X-Fugoku-API-Key: $FUGOKU_API_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"clusterId": "cluster-abc123",
"name": "tgi-llama70b",
"framework": "tgi",
"modelId": "meta-llama/Llama-3.1-70B-Instruct",
"replicas": 1,
"gpuType": "A100"
}'Check deployment status in the console's AI Models tab:
Prerequisites: You need a GPU cluster provisioned in the Fugoku console. See AI Models API for full reference.
Testing
TGI exposes both the OpenAI-compatible chat endpoint and the native /generate streaming endpoint.
# OpenAI-compatible chat
curl http://<server-ip>:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "mistralai/Mistral-7B-Instruct-v0.3",
"messages": [
{"role": "user", "content": "What is TGI in two sentences?"}
],
"max_tokens": 150,
"stream": false
}'
# Native streaming generation
curl http://<server-ip>:8000/generate \
-H "Content-Type: application/json" \
-d '{
"inputs": "The capital of France is",
"parameters": {"max_new_tokens": 50}
}'
# Health check
curl http://<server-ip>:8000/healthPerformance
- Quantization: Use
--quantize bitsandbytesor--quantize gptqto fit larger models in fewer GPUs. - Tensor parallel: Pass
--num-shard Nmatching the GPU count on multi-GPU plans (e.g.--num-shard 2ongpu-h100-2). - Batching: TGI batches continuously by default; tune
--max-batch-sizeand--max-waiting-tokensfor your latency budget. - Observability: Scrape Prometheus metrics from port
9000— TGI exposes per-request latency, token throughput, and queue depth. - Throughput: Expect 1,500–3,500 tokens/s on a single
gpu-h100-1for 7B-class models in FP16.
Troubleshooting
CUDA out of memory at startup
- Enable quantization:
--quantize bitsandbytes(4-bit) orgptq - Increase
--num-shardto spread across more GPUs - Lower
--max-batch-sizeand--max-waiting-tokens
Model download hangs or 401
- Set
HUGGING_FACE_HUB_TOKENfor gated models (Llama, Mistral) - Check outbound HTTPS from the GPU server
Slow streaming / chunked output
- Ensure
stream: truein the request body for SSE responses - Confirm the client uses
curl -N(no buffering) or a proper SSE library - Check
nvidia-smifor GPU saturation
Port 8000 unreachable
- Open firewall:
fugoku firewalls add-rule tgi-prod --protocol tcp --port 8000 --source 0.0.0.0/0 - Verify the container has
--network hostor the right-pmapping
Container restart loop
- Check
docker logsfor OOM kills — increase plan size or reduce--max-input-length/--max-batch-size
Next Steps
- vLLM Quick Start — higher throughput, broader feature set
- SGLang Quick Start — for multi-turn and structured generation
- Ollama Quick Start — simplest local-first setup
- Model Deployment — production serving patterns
- GPU Plans — full plan catalog