Ollama Quick Start
Deploy Ollama on Fugoku bare-metal GPUs for the simplest local-LLM experience
Ollama Quick Start
Ollama packages open-source models as versioned, self-contained bundles and exposes an OpenAI-compatible API. It is the fastest path from a fresh GPU server to a running model — single binary, no Python stack required.
Overview
Ollama runs as a single Go binary on Linux. It handles model download, GPU offload, and a local model registry behind one process. Key features:
- Curated library of GGUF-quantized models (
llama3.1,mistral,codellama,qwen2.5,gemma2, …) - OpenAI-compatible REST API on port
11434 - Streaming via SSE and NDJSON
- Simple Modelfile format for customization
- Single-binary install — no Python, no Docker required (though Docker is supported)
Prerequisites
- A Fugoku account with API access
- An SSH key registered with Fugoku (see SSH Keys)
- ~50 GB free disk for Ollama + model weights
Recommended GPU Plans
| Model size | Recommended plans | Notes |
|---|---|---|
| ≤13B (Llama 3.1 8B Q4) | gpu-a100-1, gpu-h100-1 | Plenty of headroom |
| 13B–70B (Llama 3.1 70B Q4) | gpu-a100-2, gpu-h100-2 | Single-stream latency sweet spot |
| 70B+ (Llama 3.1 405B Q4) | gpu-h100-4 | Multi-GPU via --num-gpu |
Regions: lagos-1, london-1, frankfurt-1. Recommended image: ubuntu-22.04-cuda12 (Ollama brings its own CUDA runtime) or pytorch-2.0-cuda12.
Deployment
Docker (recommended)
Ollama publishes an official image at ollama/ollama. The image bundles the CUDA runtime so the ubuntu-22.04-cuda12 base is fine but not strictly required.
# Provision a single A100
fugoku create instance \
--name ollama-prod \
--plan gpu-a100-1 \
--image ubuntu-22.04-cuda12 \
--region london-1 \
--ssh-key laptop
# SSH in
fugoku ssh ollama-prod
# Run Ollama with NVIDIA runtime
docker run -d \
--gpus all \
--network host \
--name ollama \
-v ollama-data:/root/.ollama \
--restart unless-stopped \
ollama/ollama:latest
# Pull a model (Llama 3.1 8B Q4)
docker exec -it ollama ollama pull llama3.1:8b
# Verify
docker exec -it ollama ollama listOnce Ollama prints llama3.1:8b, the API is live on port 11434.
Self-Managed via Docker
Ollama is deployed directly on a provisioned GPU server rather than through the Fugoku Models API (which currently supports vLLM, TGI, and SGLang). Use Docker for the simplest path:
# Provision a GPU server (see "Docker" section above for full example)
fugoku create server \
--name ollama-prod \
--plan gpu-a100-1 \
--image ubuntu-22.04-cuda12 \
--region london-1 \
--ssh-key laptop
# SSH in and run Ollama container
fugoku ssh ollama-prod
docker run -d \
--gpus all \
--network host \
--name ollama \
-v ollama-data:/root/.ollama \
--restart unless-stopped \
ollama/ollama:latest
# Pull a model
docker exec -it ollama ollama pull llama3.1:8bFor managed model deployments with automatic scaling, health checks, and model lifecycle management, use one of the supported frameworks: vLLM, TGI, or SGLang. See AI Models API for details.
Testing
Ollama exposes both its native /api/chat endpoint and the OpenAI-compatible /v1/chat/completions endpoint.
# OpenAI-compatible chat (preferred for client libraries)
curl http://<server-ip>:11434/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "llama3.1:8b",
"messages": [
{"role": "user", "content": "Write a haiku about GPUs."}
],
"max_tokens": 100,
"temperature": 0.7
}'
# Native streaming chat
curl http://<server-ip>:11434/api/chat \
-H "Content-Type: application/json" \
-d '{
"model": "llama3.1:8b",
"messages": [{"role": "user", "content": "Hello!"}],
"stream": true
}'
# List installed models
curl http://<server-ip>:11434/api/tagsPerformance
- Quantization: Ollama serves GGUF checkpoints. The
:8btag is typically Q4_0; pass:8b-instruct-q6_Kif you need higher quality at the cost of throughput. - VRAM headroom: Quantized 7B models use ~5 GB VRAM — a
gpu-a100-1(40 GB) has plenty of headroom for large context windows. - Multi-GPU: Pass
num_gpuper model in the request, or setOLLAMA_NUM_PARALLELfor concurrent request handling. - Context length: Defaults to 2048 tokens; pull a model with longer context (
OLLAMA_CONTEXT_LENGTH) for RAG workloads. - Throughput: Single-stream tokens/s on
gpu-h100-1for 7B Q4: ~120–180 tokens/s generation, ~2,000 tokens/s prompt eval.
Troubleshooting
could not detect NVIDIA GPU at startup
- Confirm the NVIDIA Container Toolkit is installed:
docker info | grep -i nvidia - Ensure the Docker daemon was restarted after installing the toolkit
- Verify
nvidia-smiworks on the host before launching the container
Model pull is slow or hangs
- Check outbound HTTPS from the GPU server
- Try a smaller tag first (
llama3.2:1b) to verify connectivity before pulling multi-GB models
Port 11434 unreachable
- Open firewall:
fugoku firewalls add-rule ollama-prod --protocol tcp --port 11434 --source 0.0.0.0/0 - Confirm
--network hostis set or the right-p 11434:11434mapping is used
Slow first-token latency
- First request pays model load cost (~5–15s); subsequent requests are fast
- Pre-load the model with
ollama runover SSH during setup to avoid the cold start on the first user request - Check
ollama psto confirm the model is loaded in VRAM
Out of memory with large models
- Use a smaller or more quantized tag (e.g.
:8b-instruct-q4_0instead of:8b-instruct-q8_0) - Upgrade to a larger plan (
gpu-h100-2orgpu-h100-4) - Lower
OLLAMA_CONTEXT_LENGTHto reduce KV cache usage
Next Steps
- vLLM Quick Start — graduate to higher throughput when you outgrow Ollama
- HuggingFace TGI Quick Start — for HF-native deployments
- SGLang Quick Start — for multi-turn and structured generation
- Model Deployment — production serving patterns
- GPU Plans — full plan catalog