AI Inference Frameworks
Quick starts for deploying vLLM, HuggingFace TGI, SGLang, and Ollama on Fugoku bare-metal GPUs
AI Inference Frameworks
Deploy open-source large language models on Fugoku's bare-metal H100 and A100 GPUs using the most popular open-source inference frameworks. This guide helps you choose the right engine for your throughput, latency, and operational requirements.
Why These Frameworks
All four frameworks serve Hugging Face compatible models behind an OpenAI-compatible HTTP API. They differ in scheduling, batching strategy, and feature set. Fugoku's bare-metal GPUs (no virtualization overhead, NVLink interconnect, 10 Gbps networking) give these engines the headroom they need to saturate tensor cores.
Common reasons teams pick a framework:
- High-throughput production serving — vLLM and SGLang lead on tokens/second under load.
- Hugging Face native integration — TGI is the reference implementation maintained by Hugging Face.
- Low-latency interactive use — Ollama is the simplest path for a single-user or small-team setup.
- Structured generation and complex prompts — SGLang's RadixAttention excels at multi-turn, branching, and JSON-mode workloads.
Comparison
| Feature | vLLM | TGI | SGLang | Ollama |
|---|---|---|---|---|
| Throughput | Very high (PagedAttention) | High | Very high (RadixAttention) | Moderate |
| Latency | Low | Low | Low | Low (single-user) |
| OpenAI-compatible | Yes | Yes | Yes | Yes |
| Multi-GPU | Tensor + Pipeline parallel | Tensor parallel | Tensor + Pipeline parallel | Limited |
| Quantization | GPTQ, AWQ, FP8, INT4 | GPTQ, AWQ, bitsandbytes | GPTQ, AWQ, FP8 | GGUF (Q4/Q5/Q8) |
| LoRA hot-swap | Yes | No | Yes | No |
| Structured outputs | Guided decoding | Guided decoding | Native JSON mode | No |
| Setup complexity | Low | Low | Low–Medium | Very low |
| Best for | Production LLM serving | Hugging Face ecosystem | Multi-turn, structured gen | Local dev, edge |
Recommended GPU Plans
| Framework | Small models (≤13B) | Medium models (13B–70B) | Large models (70B+) |
|---|---|---|---|
| vLLM | gpu-a100-1, gpu-h100-1 | gpu-a100-4, gpu-h100-2 | gpu-h100-4, gpu-h100-8 |
| TGI | gpu-a100-1, gpu-h100-1 | gpu-a100-4, gpu-h100-2 | gpu-h100-4, gpu-h100-8 |
| SGLang | gpu-a100-1, gpu-h100-1 | gpu-a100-2, gpu-h100-2 | gpu-h100-4, gpu-h100-8 |
| Ollama | gpu-a100-1, gpu-h100-1 | gpu-a100-2, gpu-h100-2 | gpu-h100-4 |
Recommended images for all frameworks: pytorch-2.0-cuda12 (full ML stack) or ubuntu-22.04-cuda12 (minimal CUDA runtime for container-based engines).
Selection Guide
Choose vLLM when:
- You need maximum tokens/second across many concurrent users
- You run LoRA adapters and want hot-swapping without restarting
- You need guided decoding / JSON schema enforcement
- You want the broadest model coverage (most Hugging Face models "just work")
Choose TGI when:
- You already use Hugging Face Inference Endpoints or
transformers - You want the canonical reference implementation maintained by Hugging Face
- You need built-in Prometheus metrics out of the box
- You're running standard transformer models without exotic features
Choose SGLang when:
- You serve multi-turn conversations, agents, or complex prompt templates
- You need native JSON mode and structured outputs at high throughput
- You want the fastest single-stream latency for code completion
- You're building RAG or tool-use pipelines with branching control flow
Choose Ollama when:
- You want the absolute fastest path from zero to a running model
- You primarily serve a single user or a small team
- You prefer a simple CLI and curated model library over raw Hugging Face repos
- You're prototyping locally and will graduate to vLLM or TGI for production
Deployment Patterns
All four frameworks follow the same deployment shape on Fugoku:
- Provision a GPU server (
gpu-a100-1or larger) using thepytorch-2.0-cuda12image - SSH into the server with
fugoku ssh - Pull and run the framework's Docker container, exposing the OpenAI-compatible API
- Open a firewall rule for the serving port (default
8000) - Test with
curlfrom your laptop
Each framework page below walks through Docker, CLI, and API deployment with copy-pasteable commands.
Next Steps
- vLLM Quick Start — highest throughput, broadest model coverage
- HuggingFace TGI Quick Start — canonical reference implementation
- SGLang Quick Start — best for agents, structured generation, multi-turn
- Ollama Quick Start — simplest path, local-first experience
- GPU Plans — full plan catalog and pricing
- Model Deployment — production serving patterns