FugokuFugoku Docs
Mask

AI Models

Deploy and manage AI model endpoints

AI Models

Deploy and manage AI model serving endpoints. Models are deployed via supported inference frameworks on GPU instances.

List Models

GET /v1/models

Returns deployed models for the current project.

curl https://api.fugoku.com/v1/models \
  -H "X-Fugoku-API-Key: $FUGOKU_API_TOKEN"

Response:

{
  "data": [],
  "meta": { "total": 0 }
}

Deploy Model

POST /v1/models

Request Body:

FieldTypeRequiredDescription
clusterIdstringYesCluster ID to deploy on
namestringYesModel name (1–100 chars)
frameworkstringYesInference framework
modelIdstringYesModel identifier (e.g. meta-llama/Llama-3-70B)
replicasintegerNoNumber of replicas (default 1)
gpuTypestringNoGPU type preference
portintegerNoCustom serving port

Supported Frameworks:

FrameworkDescription
vllmvLLM inference engine
tgiHuggingFace Text Generation Inference
tritonNVIDIA Triton Inference Server
sglangSGLang serving framework
curl -X POST https://api.fugoku.com/v1/models \
  -H "X-Fugoku-API-Key: $FUGOKU_API_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "clusterId": "cluster-abc123",
    "name": "llama-3-70b",
    "framework": "vllm",
    "modelId": "meta-llama/Llama-3-70B",
    "replicas": 2,
    "gpuType": "A100"
  }'

Response (201):

{
  "data": {
    "id": "model-a1b2c3",
    "type": "ai-models",
    "attributes": {
      "id": "model-a1b2c3",
      "clusterId": "cluster-abc123",
      "framework": "vllm",
      "modelId": "meta-llama/Llama-3-70B",
      "replicas": 2,
      "status": "deploying",
      "skyServiceName": null,
      "endpointUrl": null,
      "createdAt": "2026-08-22T00:00:00.000Z"
    }
  }
}

Model Status Values

StatusDescription
deployingModel is being deployed
activeModel is serving requests
stoppedModel is stopped
failedDeployment failed

Model deployment integrates with the AI Gateway (LiteLLM) and SkyPilot for actual GPU provisioning. The endpoint URL becomes available once the model reaches active status.

On this page