FugokuFugoku Docs
Mask

AI Inference Frameworks

Quick starts for deploying vLLM, HuggingFace TGI, SGLang, and Ollama on Fugoku bare-metal GPUs

AI Inference Frameworks

Deploy open-source large language models on Fugoku's bare-metal H100 and A100 GPUs using the most popular open-source inference frameworks. This guide helps you choose the right engine for your throughput, latency, and operational requirements.

Why These Frameworks

All four frameworks serve Hugging Face compatible models behind an OpenAI-compatible HTTP API. They differ in scheduling, batching strategy, and feature set. Fugoku's bare-metal GPUs (no virtualization overhead, NVLink interconnect, 10 Gbps networking) give these engines the headroom they need to saturate tensor cores.

Common reasons teams pick a framework:

  • High-throughput production serving — vLLM and SGLang lead on tokens/second under load.
  • Hugging Face native integration — TGI is the reference implementation maintained by Hugging Face.
  • Low-latency interactive use — Ollama is the simplest path for a single-user or small-team setup.
  • Structured generation and complex prompts — SGLang's RadixAttention excels at multi-turn, branching, and JSON-mode workloads.

Comparison

FeaturevLLMTGISGLangOllama
ThroughputVery high (PagedAttention)HighVery high (RadixAttention)Moderate
LatencyLowLowLowLow (single-user)
OpenAI-compatibleYesYesYesYes
Multi-GPUTensor + Pipeline parallelTensor parallelTensor + Pipeline parallelLimited
QuantizationGPTQ, AWQ, FP8, INT4GPTQ, AWQ, bitsandbytesGPTQ, AWQ, FP8GGUF (Q4/Q5/Q8)
LoRA hot-swapYesNoYesNo
Structured outputsGuided decodingGuided decodingNative JSON modeNo
Setup complexityLowLowLow–MediumVery low
Best forProduction LLM servingHugging Face ecosystemMulti-turn, structured genLocal dev, edge
FrameworkSmall models (≤13B)Medium models (13B–70B)Large models (70B+)
vLLMgpu-a100-1, gpu-h100-1gpu-a100-4, gpu-h100-2gpu-h100-4, gpu-h100-8
TGIgpu-a100-1, gpu-h100-1gpu-a100-4, gpu-h100-2gpu-h100-4, gpu-h100-8
SGLanggpu-a100-1, gpu-h100-1gpu-a100-2, gpu-h100-2gpu-h100-4, gpu-h100-8
Ollamagpu-a100-1, gpu-h100-1gpu-a100-2, gpu-h100-2gpu-h100-4

Recommended images for all frameworks: pytorch-2.0-cuda12 (full ML stack) or ubuntu-22.04-cuda12 (minimal CUDA runtime for container-based engines).

Selection Guide

Choose vLLM when:

  • You need maximum tokens/second across many concurrent users
  • You run LoRA adapters and want hot-swapping without restarting
  • You need guided decoding / JSON schema enforcement
  • You want the broadest model coverage (most Hugging Face models "just work")

Choose TGI when:

  • You already use Hugging Face Inference Endpoints or transformers
  • You want the canonical reference implementation maintained by Hugging Face
  • You need built-in Prometheus metrics out of the box
  • You're running standard transformer models without exotic features

Choose SGLang when:

  • You serve multi-turn conversations, agents, or complex prompt templates
  • You need native JSON mode and structured outputs at high throughput
  • You want the fastest single-stream latency for code completion
  • You're building RAG or tool-use pipelines with branching control flow

Choose Ollama when:

  • You want the absolute fastest path from zero to a running model
  • You primarily serve a single user or a small team
  • You prefer a simple CLI and curated model library over raw Hugging Face repos
  • You're prototyping locally and will graduate to vLLM or TGI for production

Deployment Patterns

All four frameworks follow the same deployment shape on Fugoku:

  1. Provision a GPU server (gpu-a100-1 or larger) using the pytorch-2.0-cuda12 image
  2. SSH into the server with fugoku ssh
  3. Pull and run the framework's Docker container, exposing the OpenAI-compatible API
  4. Open a firewall rule for the serving port (default 8000)
  5. Test with curl from your laptop

Each framework page below walks through Docker, CLI, and API deployment with copy-pasteable commands.

Next Steps

On this page