Performance Benchmarks
GPU training throughput, network latency matrix, and storage IOPS benchmarks for Fugoku infrastructure
Performance Benchmarks
Real-world performance measurements from Fugoku production infrastructure. All benchmarks are measured on bare-metal hardware with NVMe storage and 10 Gbps networking unless otherwise specified.
GPU Benchmarks
LLM Training Throughput
Measured with Llama 3 70B in INT4/AWQ quantization using vLLM with continuous batching.
| GPU | Tokens per second | Batch size | Precision |
|---|---|---|---|
| 1x A100 80GB | ~28 tok/s | 1 | INT4/AWQ |
| 1x H100 80GB | ~52 tok/s | 1 | INT4/AWQ |
| 4x A100 80GB | ~110 tok/s | 4 | INT4/AWQ |
| 8x A100 80GB | ~210 tok/s | 8 | INT4/AWQ |
Test conditions:
- Model: Llama 3 70B (INT4/AWQ quantized)
- Framework: vLLM 0.4.0 with continuous batching
- Context length: 4096 tokens
- Input: 1000 prompt tokens, 1000 output tokens
Inference Latency
Measured with ResNet-50 image classification model using TensorRT optimization.
| GPU | p50 Latency | p95 Latency | p99 Latency |
|---|---|---|---|
| 1x A100 80GB | 18ms | 23ms | 31ms |
| 1x H100 80GB | 14ms | 18ms | 24ms |
Test conditions:
- Model: ResNet-50 (FP16)
- Batch size: 1
- Image size: 224x224
- Framework: TensorRT 8.6
Distributed Training Scaling
Measured with ResNet-50 training across multiple GPUs using PyTorch DistributedDataParallel.
| GPUs | Throughput | Scaling efficiency |
|---|---|---|
| 1x A100 80GB | 1.0x | 100% |
| 2x A100 80GB | 1.92x | 96% |
| 4x A100 80GB | 3.80x | 95% |
| 8x A100 80GB | 7.56x | 95% |
Test conditions:
- Model: ResNet-50
- Batch size: 256 per GPU
- Optimizer: AdamW
- Precision: FP16 with mixed precision
- Interconnect: NVLink / PCIe
Network Benchmarks
Inter-Region Latency
Round-trip latency measured via ICMP ping between Fugoku regions.
| Source | Destination | Latency |
|---|---|---|
| Lagos | Nairobi | ~15ms |
| Lagos | Accra | ~20ms |
| Lagos | Johannesburg | ~35ms |
| Lagos | Frankfurt | ~120ms |
| Frankfurt | London | ~8ms |
| Frankfurt | New York | ~65ms |
| Frankfurt | Singapore | ~150ms |
| New York | Los Angeles | ~45ms |
Measurement method:
- 100 ICMP pings per pair
- 95th percentile reported
- Measured during off-peak hours
Intra-Region Bandwidth
| Region | Server-to-server bandwidth | Latency |
|---|---|---|
| Lagos | 10 Gbps | < 0.5ms |
| Frankfurt | 10 Gbps | < 0.5ms |
| London (coming soon) | 10 Gbps | < 0.5ms |
Storage Benchmarks
NVMe RAID Performance
Measured with fio on production NVMe RAID arrays.
| Test | Block Size | Queue Depth | Result |
|---|---|---|---|
| Sequential Read | 1 MB | 32 | 5.2 GB/s |
| Sequential Write | 1 MB | 32 | 4.8 GB/s |
| Random Read (4K) | 4 KB | 64 | 850K IOPS |
| Random Write (4K) | 4 KB | 64 | 720K IOPS |
| Mixed Read/Write (4K) | 4 KB | 64 | 780K IOPS |
Test conditions:
- Storage: 4x NVMe SSD in RAID 10
- Direct I/O, no caching
- File system: ext4 with default mount options
- Server: m4.metal.large (24 cores, 384 GB RAM)
Block Volume Performance
| Test | Block Size | Queue Depth | Result |
|---|---|---|---|
| Sequential Read | 1 MB | 32 | 3.1 GB/s |
| Sequential Write | 1 MB | 32 | 2.8 GB/s |
| Random Read (4K) | 4 KB | 64 | 420K IOPS |
| Random Write (4K) | 4 KB | 64 | 380K IOPS |
Test conditions:
- Storage: Attached block volume (100 GB)
- Direct I/O, no caching
- Instance: vm-standard (4 vCPU, 16 GB RAM)
Benchmarking Your Workload
We offer custom benchmark services for enterprise customers:
- Workload analysis — We profile your application to understand resource requirements
- Infrastructure sizing — We recommend optimal server types and quantities
- Performance testing — We run your workload on our infrastructure and measure results
- Optimization report — We provide tuning recommendations for maximum performance
Contact benchmarks@fugoku.com to request custom benchmarks.
Running Your Own Benchmarks
You can reproduce these benchmarks on your Fugoku instances:
GPU Benchmark
pip install torch torchvision transformers vllm
python -c "
import torch
import time
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained('meta-llama/Llama-3-70B-Instruct', torch_dtype=torch.float16, device_map='auto')
tokenizer = AutoTokenizer.from_pretrained('meta-llama/Llama-3-70B-Instruct')
prompt = 'Hello, my name is'
inputs = tokenizer(prompt, return_tensors='pt').to('cuda')
start = time.time()
outputs = model.generate(**inputs, max_new_tokens=100)
elapsed = time.time() - start
tokens = outputs.shape[1] - inputs.input_ids.shape[1]
print(f'Generated {tokens} tokens in {elapsed:.2f}s ({tokens/elapsed:.1f} tok/s)')
"Storage Benchmark
sudo apt install -y fio
fio --name=seq_read --rw=read --bs=1M --iodepth=32 --size=10G --filename=/tmp/test --direct=1
fio --name=seq_write --rw=write --bs=1M --iodepth=32 --size=10G --filename=/tmp/test --direct=1
fio --name=rand_read --rw=randread --bs=4k --iodepth=64 --size=10G --filename=/tmp/test --direct=1
fio --name=rand_write --rw=randwrite --bs=4k --iodepth=64 --size=10G --filename=/tmp/test --direct=1Network Benchmark
sudo apt install -y iperf3
iperf3 -s
iperf3 -c <server-ip> -t 30 -P 4Next Steps:
- Deploy a GPU instance to run your own benchmarks
- Review pricing for GPU and storage
- Contact support for custom benchmark requests