AI Models
Deploy and manage AI model endpoints
AI Models
Deploy and manage AI model serving endpoints. Models are deployed via supported inference frameworks on GPU instances.
List Models
GET /v1/modelsReturns deployed models for the current project.
curl https://api.fugoku.com/v1/models \
-H "X-Fugoku-API-Key: $FUGOKU_API_TOKEN"Response:
{
"data": [],
"meta": { "total": 0 }
}Deploy Model
POST /v1/modelsRequest Body:
| Field | Type | Required | Description |
|---|---|---|---|
clusterId | string | Yes | Cluster ID to deploy on |
name | string | Yes | Model name (1–100 chars) |
framework | string | Yes | Inference framework |
modelId | string | Yes | Model identifier (e.g. meta-llama/Llama-3-70B) |
replicas | integer | No | Number of replicas (default 1) |
gpuType | string | No | GPU type preference |
port | integer | No | Custom serving port |
Supported Frameworks:
| Framework | Description |
|---|---|
vllm | vLLM inference engine |
tgi | HuggingFace Text Generation Inference |
triton | NVIDIA Triton Inference Server |
sglang | SGLang serving framework |
curl -X POST https://api.fugoku.com/v1/models \
-H "X-Fugoku-API-Key: $FUGOKU_API_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"clusterId": "cluster-abc123",
"name": "llama-3-70b",
"framework": "vllm",
"modelId": "meta-llama/Llama-3-70B",
"replicas": 2,
"gpuType": "A100"
}'Response (201):
{
"data": {
"id": "model-a1b2c3",
"type": "ai-models",
"attributes": {
"id": "model-a1b2c3",
"clusterId": "cluster-abc123",
"framework": "vllm",
"modelId": "meta-llama/Llama-3-70B",
"replicas": 2,
"status": "deploying",
"skyServiceName": null,
"endpointUrl": null,
"createdAt": "2026-08-22T00:00:00.000Z"
}
}
}Model Status Values
| Status | Description |
|---|---|
deploying | Model is being deployed |
active | Model is serving requests |
stopped | Model is stopped |
failed | Deployment failed |
Model deployment integrates with the AI Gateway (LiteLLM) and SkyPilot for actual GPU provisioning. The endpoint URL becomes available once the model reaches active status.