FugokuFugoku Docs
Mask

Performance Benchmarks

GPU training throughput, network latency matrix, and storage IOPS benchmarks for Fugoku infrastructure

Performance Benchmarks

Real-world performance measurements from Fugoku production infrastructure. All benchmarks are measured on bare-metal hardware with NVMe storage and 10 Gbps networking unless otherwise specified.

GPU Benchmarks

LLM Training Throughput

Measured with Llama 3 70B in INT4/AWQ quantization using vLLM with continuous batching.

GPUTokens per secondBatch sizePrecision
1x A100 80GB~28 tok/s1INT4/AWQ
1x H100 80GB~52 tok/s1INT4/AWQ
4x A100 80GB~110 tok/s4INT4/AWQ
8x A100 80GB~210 tok/s8INT4/AWQ

Test conditions:

  • Model: Llama 3 70B (INT4/AWQ quantized)
  • Framework: vLLM 0.4.0 with continuous batching
  • Context length: 4096 tokens
  • Input: 1000 prompt tokens, 1000 output tokens

Inference Latency

Measured with ResNet-50 image classification model using TensorRT optimization.

GPUp50 Latencyp95 Latencyp99 Latency
1x A100 80GB18ms23ms31ms
1x H100 80GB14ms18ms24ms

Test conditions:

  • Model: ResNet-50 (FP16)
  • Batch size: 1
  • Image size: 224x224
  • Framework: TensorRT 8.6

Distributed Training Scaling

Measured with ResNet-50 training across multiple GPUs using PyTorch DistributedDataParallel.

GPUsThroughputScaling efficiency
1x A100 80GB1.0x100%
2x A100 80GB1.92x96%
4x A100 80GB3.80x95%
8x A100 80GB7.56x95%

Test conditions:

  • Model: ResNet-50
  • Batch size: 256 per GPU
  • Optimizer: AdamW
  • Precision: FP16 with mixed precision
  • Interconnect: NVLink / PCIe

Network Benchmarks

Inter-Region Latency

Round-trip latency measured via ICMP ping between Fugoku regions.

SourceDestinationLatency
LagosNairobi~15ms
LagosAccra~20ms
LagosJohannesburg~35ms
LagosFrankfurt~120ms
FrankfurtLondon~8ms
FrankfurtNew York~65ms
FrankfurtSingapore~150ms
New YorkLos Angeles~45ms

Measurement method:

  • 100 ICMP pings per pair
  • 95th percentile reported
  • Measured during off-peak hours

Intra-Region Bandwidth

RegionServer-to-server bandwidthLatency
Lagos10 Gbps< 0.5ms
Frankfurt10 Gbps< 0.5ms
London (coming soon)10 Gbps< 0.5ms

Storage Benchmarks

NVMe RAID Performance

Measured with fio on production NVMe RAID arrays.

TestBlock SizeQueue DepthResult
Sequential Read1 MB325.2 GB/s
Sequential Write1 MB324.8 GB/s
Random Read (4K)4 KB64850K IOPS
Random Write (4K)4 KB64720K IOPS
Mixed Read/Write (4K)4 KB64780K IOPS

Test conditions:

  • Storage: 4x NVMe SSD in RAID 10
  • Direct I/O, no caching
  • File system: ext4 with default mount options
  • Server: m4.metal.large (24 cores, 384 GB RAM)

Block Volume Performance

TestBlock SizeQueue DepthResult
Sequential Read1 MB323.1 GB/s
Sequential Write1 MB322.8 GB/s
Random Read (4K)4 KB64420K IOPS
Random Write (4K)4 KB64380K IOPS

Test conditions:

  • Storage: Attached block volume (100 GB)
  • Direct I/O, no caching
  • Instance: vm-standard (4 vCPU, 16 GB RAM)

Benchmarking Your Workload

We offer custom benchmark services for enterprise customers:

  1. Workload analysis — We profile your application to understand resource requirements
  2. Infrastructure sizing — We recommend optimal server types and quantities
  3. Performance testing — We run your workload on our infrastructure and measure results
  4. Optimization report — We provide tuning recommendations for maximum performance

Contact benchmarks@fugoku.com to request custom benchmarks.

Running Your Own Benchmarks

You can reproduce these benchmarks on your Fugoku instances:

GPU Benchmark

pip install torch torchvision transformers vllm

python -c "
import torch
import time

from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained('meta-llama/Llama-3-70B-Instruct', torch_dtype=torch.float16, device_map='auto')
tokenizer = AutoTokenizer.from_pretrained('meta-llama/Llama-3-70B-Instruct')

prompt = 'Hello, my name is'
inputs = tokenizer(prompt, return_tensors='pt').to('cuda')
start = time.time()
outputs = model.generate(**inputs, max_new_tokens=100)
elapsed = time.time() - start
tokens = outputs.shape[1] - inputs.input_ids.shape[1]
print(f'Generated {tokens} tokens in {elapsed:.2f}s ({tokens/elapsed:.1f} tok/s)')
"

Storage Benchmark

sudo apt install -y fio

fio --name=seq_read --rw=read --bs=1M --iodepth=32 --size=10G --filename=/tmp/test --direct=1
fio --name=seq_write --rw=write --bs=1M --iodepth=32 --size=10G --filename=/tmp/test --direct=1

fio --name=rand_read --rw=randread --bs=4k --iodepth=64 --size=10G --filename=/tmp/test --direct=1
fio --name=rand_write --rw=randwrite --bs=4k --iodepth=64 --size=10G --filename=/tmp/test --direct=1

Network Benchmark

sudo apt install -y iperf3

iperf3 -s

iperf3 -c <server-ip> -t 30 -P 4

Next Steps:

On this page