FugokuFugoku Docs
Mask

Ollama Quick Start

Deploy Ollama on Fugoku bare-metal GPUs for the simplest local-LLM experience

Ollama Quick Start

Ollama packages open-source models as versioned, self-contained bundles and exposes an OpenAI-compatible API. It is the fastest path from a fresh GPU server to a running model — single binary, no Python stack required.

Overview

Ollama runs as a single Go binary on Linux. It handles model download, GPU offload, and a local model registry behind one process. Key features:

  • Curated library of GGUF-quantized models (llama3.1, mistral, codellama, qwen2.5, gemma2, …)
  • OpenAI-compatible REST API on port 11434
  • Streaming via SSE and NDJSON
  • Simple Modelfile format for customization
  • Single-binary install — no Python, no Docker required (though Docker is supported)

Prerequisites

  • A Fugoku account with API access
  • An SSH key registered with Fugoku (see SSH Keys)
  • ~50 GB free disk for Ollama + model weights
Model sizeRecommended plansNotes
≤13B (Llama 3.1 8B Q4)gpu-a100-1, gpu-h100-1Plenty of headroom
13B–70B (Llama 3.1 70B Q4)gpu-a100-2, gpu-h100-2Single-stream latency sweet spot
70B+ (Llama 3.1 405B Q4)gpu-h100-4Multi-GPU via --num-gpu

Regions: lagos-1, london-1, frankfurt-1. Recommended image: ubuntu-22.04-cuda12 (Ollama brings its own CUDA runtime) or pytorch-2.0-cuda12.

Deployment

Ollama publishes an official image at ollama/ollama. The image bundles the CUDA runtime so the ubuntu-22.04-cuda12 base is fine but not strictly required.

# Provision a single A100
fugoku create instance \
  --name ollama-prod \
  --plan gpu-a100-1 \
  --image ubuntu-22.04-cuda12 \
  --region london-1 \
  --ssh-key laptop

# SSH in
fugoku ssh ollama-prod

# Run Ollama with NVIDIA runtime
docker run -d \
  --gpus all \
  --network host \
  --name ollama \
  -v ollama-data:/root/.ollama \
  --restart unless-stopped \
  ollama/ollama:latest

# Pull a model (Llama 3.1 8B Q4)
docker exec -it ollama ollama pull llama3.1:8b

# Verify
docker exec -it ollama ollama list

Once Ollama prints llama3.1:8b, the API is live on port 11434.

Self-Managed via Docker

Ollama is deployed directly on a provisioned GPU server rather than through the Fugoku Models API (which currently supports vLLM, TGI, and SGLang). Use Docker for the simplest path:

# Provision a GPU server (see "Docker" section above for full example)
fugoku create server \
  --name ollama-prod \
  --plan gpu-a100-1 \
  --image ubuntu-22.04-cuda12 \
  --region london-1 \
  --ssh-key laptop

# SSH in and run Ollama container
fugoku ssh ollama-prod

docker run -d \
  --gpus all \
  --network host \
  --name ollama \
  -v ollama-data:/root/.ollama \
  --restart unless-stopped \
  ollama/ollama:latest

# Pull a model
docker exec -it ollama ollama pull llama3.1:8b

For managed model deployments with automatic scaling, health checks, and model lifecycle management, use one of the supported frameworks: vLLM, TGI, or SGLang. See AI Models API for details.

Testing

Ollama exposes both its native /api/chat endpoint and the OpenAI-compatible /v1/chat/completions endpoint.

# OpenAI-compatible chat (preferred for client libraries)
curl http://<server-ip>:11434/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "llama3.1:8b",
    "messages": [
      {"role": "user", "content": "Write a haiku about GPUs."}
    ],
    "max_tokens": 100,
    "temperature": 0.7
  }'

# Native streaming chat
curl http://<server-ip>:11434/api/chat \
  -H "Content-Type: application/json" \
  -d '{
    "model": "llama3.1:8b",
    "messages": [{"role": "user", "content": "Hello!"}],
    "stream": true
  }'

# List installed models
curl http://<server-ip>:11434/api/tags

Performance

  • Quantization: Ollama serves GGUF checkpoints. The :8b tag is typically Q4_0; pass :8b-instruct-q6_K if you need higher quality at the cost of throughput.
  • VRAM headroom: Quantized 7B models use ~5 GB VRAM — a gpu-a100-1 (40 GB) has plenty of headroom for large context windows.
  • Multi-GPU: Pass num_gpu per model in the request, or set OLLAMA_NUM_PARALLEL for concurrent request handling.
  • Context length: Defaults to 2048 tokens; pull a model with longer context (OLLAMA_CONTEXT_LENGTH) for RAG workloads.
  • Throughput: Single-stream tokens/s on gpu-h100-1 for 7B Q4: ~120–180 tokens/s generation, ~2,000 tokens/s prompt eval.

Troubleshooting

could not detect NVIDIA GPU at startup

  • Confirm the NVIDIA Container Toolkit is installed: docker info | grep -i nvidia
  • Ensure the Docker daemon was restarted after installing the toolkit
  • Verify nvidia-smi works on the host before launching the container

Model pull is slow or hangs

  • Check outbound HTTPS from the GPU server
  • Try a smaller tag first (llama3.2:1b) to verify connectivity before pulling multi-GB models

Port 11434 unreachable

  • Open firewall: fugoku firewalls add-rule ollama-prod --protocol tcp --port 11434 --source 0.0.0.0/0
  • Confirm --network host is set or the right -p 11434:11434 mapping is used

Slow first-token latency

  • First request pays model load cost (~5–15s); subsequent requests are fast
  • Pre-load the model with ollama run over SSH during setup to avoid the cold start on the first user request
  • Check ollama ps to confirm the model is loaded in VRAM

Out of memory with large models

  • Use a smaller or more quantized tag (e.g. :8b-instruct-q4_0 instead of :8b-instruct-q8_0)
  • Upgrade to a larger plan (gpu-h100-2 or gpu-h100-4)
  • Lower OLLAMA_CONTEXT_LENGTH to reduce KV cache usage

Next Steps

On this page