API reference · API 1.0.0

AI models

The AI model catalogue and node caches.

get /api/v1/ai-models

List AI models

Operation ai_models_list · bearer token

The model catalogue, in id order: every model the platform records, whether or not its weights are present. The filters are exact and combine: repo finds the row of one Hugging Face repo (the id an import needs); status, category and gateway_tier narrow the list.

Parameters

NameInTypeDescription
repo query string | null

Exact repo id: finds a model's id to import it.

max length 200
status query string | null

Exact status.

max length 20
category query string | null
max length 24
gateway_tier query string | null
max length 40
limit query integer

Page size.

min 1 · max 200 · default 50
cursor query string | null

next_cursor from the previous page; omit it for the first page. A cursor this list did not issue is a 400 invalid_cursor.

Responses

  • 200

    One page of the catalogue.

    application/json → AiModelPage
    Headers: X-Request-ID
  • 401

    Not authenticated: no Authorization: Bearer header (not_authenticated), or the credential is refused (token_invalid, token_expired, token_revoked, principal_disabled, token_ip_not_allowed).

    application/jsonapplication/problem+json → Problem
  • 403

    The caller lacks a permission the operation needs (forbidden; required lists the keys, any one of which would do), or the licence refuses a change (licence_locked, licence_restricted, licence_required, with state and remedy; never retry a licence_* code).

    application/jsonapplication/problem+json → Problem
    Headers: X-Request-ID
  • 429

    More than 600 requests in a minute with this token (rate_limited), or the first request with this Idempotency-Key is still running (idempotency_request_in_progress). Retry after Retry-After.

    application/jsonapplication/problem+json → Problem
  • 503

    The platform cannot answer right now (unavailable; retry after Retry-After), or API tokens are not configured on it (api_tokens_unconfigured; an operator must act).

    application/jsonapplication/problem+json → Problem
  • default

    Problem details (RFC 9457)

    application/jsonapplication/problem+json → Problem
    Headers: X-Request-ID
post /api/v1/ai-models

Add an AI model to the catalogue

Operation ai_models_create · bearer token

Records a catalogue row. The model starts planned, with no central copy and no node cache: weights arrive only through the store actions, and pulling them is not part of v1. status, location, offline_ready and nas_volume are not accepted. Send an Idempotency-Key to make a retry safe: a retry with the same key and request gets the first answer and changes nothing.

Parameters

NameInTypeDescription
Idempotency-Key header string

Makes a retry safe: for 24 hours, the same key with the same request (method, path, query and body) answers with the stored response (Idempotent-Replayed: true) instead of doing the work again. The same key with a different request is a 409 idempotency_key_reused; while the first request is still running it is a 429 idempotency_request_in_progress with Retry-After. 1-255 printable characters without spaces (400 invalid_idempotency_key otherwise); a UUID is a good key. Keys are scoped to the calling account.

min length 1 · max length 255 · pattern ^[!-~]+$

Request bodyrequired

application/json

Schema AiModelCreate

A catalogue row. The model is created planned, with no central copy and no node cache: weights arrive only through the store actions, and pulling them is not part of v1. display_name defaults to the repo's last part.

NameTypeDescription
architecture string | null

The model's architecture, as recorded: dense or moe (mixture of experts). Other values may appear.

max length 40
benchmarks map of number (double) | integer | string | null

Published benchmark scores: benchmark name to score, e.g. {"SWE-bench Verified": 69.6}. As published by the vendor; not measured here.

category string | null

What the model is for: coding, agentic, reasoning, vision, embedding, rerank, general, speech or video. Other values may appear.

max length 24
context_window string | null

The context window in human form, as recorded, e.g. 128K or 256K→1M.

max length 20
description string | null

A longer introduction, for the model's detail view.

max length 20000
display_name string | null

The name the catalogue shows. On create it defaults to the last part of repo.

min length 1 · max length 255
frontier_equiv string | null

Which hosted frontier model it is roughly comparable to, and from when, in a few words. An estimate, not a measurement.

max length 120
gated boolean | null

The Hugging Face repo needs an accept-click (access approval) before a pull.

gateway_tier string | null

The AI gateway serving tier the model is meant for, e.g. code or general (GET /ai/gateway/tiers). Recorded only: which model backs a tier is decided there, not here.

max length 40
license string | null

The model's licence, as recorded, e.g. Apache-2.0, MIT.

max length 80
min_target string | null

The smallest hardware it runs on, in human form, e.g. 1× 3090 or 2×DGX.

max length 40
model_card_url string | null

The model card's web address.

max length 300
notes string | null

Free-form notes of the platform's operators.

max length 20000
org string | null

The Hugging Face organisation the repo is published under. It may be a quantizer rather than the lab that built the model (that is vendor).

max length 120
param_count_b number (double) | null

The parameter count in billions, as a number to sort by.

min 0 · max 9999999
params string | null

The parameter count in human form, e.g. "480B (35B active)".

max length 60
published string | null

When the model was released upstream, as recorded, e.g. 2024-11.

max length 20
quant string | null

The weights' precision or quantisation, as recorded, e.g. BF16, FP8, AWQ, NVFP4.

max length 40
reporequired string

A Hugging Face repo id org/name: each part starts with a letter or digit and holds only letters, digits, ., _ and - (no ..), at most 96 characters. Frozen after create; unique.

serving_node string | null

Where the model is meant to be served: an AI node's hostname or a DGX cluster, as recorded. Recorded only: loading a model is not part of this API.

max length 80
size_gb number (double) | null

The weights' size on disk in GB, as recorded.

min 0 · max 99999999
strong_axis string | null

What the model is strongest at, in a few words, e.g. agentic coding.

max length 40
summary string | null

A one-line introduction, for a catalogue card.

max length 2000
vendor string | null

The lab that built the model. The Hugging Face organisation (org) may be a quantizer.

max length 200
vendor_country string | null

The vendor's home country, as recorded, e.g. China, France, USA.

max length 40

Responses

  • 201

    The model, as recorded.

    application/json → AiModel
  • 401

    Not authenticated: no Authorization: Bearer header (not_authenticated), or the credential is refused (token_invalid, token_expired, token_revoked, principal_disabled, token_ip_not_allowed).

    application/jsonapplication/problem+json → Problem
  • 403

    The caller lacks a permission the operation needs (forbidden; required lists the keys, any one of which would do), or the licence refuses a change (licence_locked, licence_restricted, licence_required, with state and remedy; never retry a licence_* code).

    application/jsonapplication/problem+json → Problem
    Headers: X-Request-ID
  • 409

    The repo is already in the catalogue (repo_taken).

    application/jsonapplication/problem+json → Problem
    Headers: X-Request-ID
  • 422

    The repo is not a strict Hugging Face id, or the body is invalid.

    application/jsonapplication/problem+json → Problem
    Headers: X-Request-ID
  • 429

    More than 600 requests in a minute with this token (rate_limited), or the first request with this Idempotency-Key is still running (idempotency_request_in_progress). Retry after Retry-After.

    application/jsonapplication/problem+json → Problem
  • 503

    The platform cannot answer right now (unavailable; retry after Retry-After), or API tokens are not configured on it (api_tokens_unconfigured; an operator must act).

    application/jsonapplication/problem+json → Problem
  • default

    Problem details (RFC 9457)

    application/jsonapplication/problem+json → Problem
    Headers: X-Request-ID
get /api/v1/ai-models/load-targets

List load targets

Operation ai_models_load_targets_list · bearer token

The nodes and DGX clusters a model can be served on, with their VRAM budget. Live values when monitoring answers, static fallbacks otherwise.

Parameters

NameInTypeDescription
limit query integer

Page size.

min 1 · max 200 · default 50
cursor query string | null

next_cursor from the previous page; omit it for the first page. A cursor this list did not issue is a 400 invalid_cursor.

Responses

  • 200

    One page of load targets.

    application/json → LoadTargetPage
    Headers: X-Request-ID
  • 401

    Not authenticated: no Authorization: Bearer header (not_authenticated), or the credential is refused (token_invalid, token_expired, token_revoked, principal_disabled, token_ip_not_allowed).

    application/jsonapplication/problem+json → Problem
  • 403

    The caller lacks a permission the operation needs (forbidden; required lists the keys, any one of which would do), or the licence refuses a change (licence_locked, licence_restricted, licence_required, with state and remedy; never retry a licence_* code).

    application/jsonapplication/problem+json → Problem
    Headers: X-Request-ID
  • 429

    More than 600 requests in a minute with this token (rate_limited), or the first request with this Idempotency-Key is still running (idempotency_request_in_progress). Retry after Retry-After.

    application/jsonapplication/problem+json → Problem
  • 503

    The platform cannot answer right now (unavailable; retry after Retry-After), or API tokens are not configured on it (api_tokens_unconfigured; an operator must act).

    application/jsonapplication/problem+json → Problem
  • default

    Problem details (RFC 9457)

    application/jsonapplication/problem+json → Problem
    Headers: X-Request-ID
get /api/v1/ai-models/runs/{run_id}

One store run

Operation ai_models_runs_get · bearer token

One run of the model store, as recorded: a node cache or uncache started through this API, or a run started in the portal. The same run is the operation model-store-run:<id>; GET /operations/{id} reports it in the shape every long-running operation shares.

Parameters

NameInTypeDescription
run_idrequired path string

The run's id.

pattern ^[1-9][0-9]{0,8}$

Responses

  • 200

    The store run.

    application/json → StoreRun
    Headers: X-Request-ID
  • 401

    Not authenticated: no Authorization: Bearer header (not_authenticated), or the credential is refused (token_invalid, token_expired, token_revoked, principal_disabled, token_ip_not_allowed).

    application/jsonapplication/problem+json → Problem
  • 403

    The caller lacks a permission the operation needs (forbidden; required lists the keys, any one of which would do), or the licence refuses a change (licence_locked, licence_restricted, licence_required, with state and remedy; never retry a licence_* code).

    application/jsonapplication/problem+json → Problem
    Headers: X-Request-ID
  • 404

    No such run (run_not_found).

    application/jsonapplication/problem+json → Problem
    Headers: X-Request-ID
  • 429

    More than 600 requests in a minute with this token (rate_limited), or the first request with this Idempotency-Key is still running (idempotency_request_in_progress). Retry after Retry-After.

    application/jsonapplication/problem+json → Problem
  • 503

    The platform cannot answer right now (unavailable; retry after Retry-After), or API tokens are not configured on it (api_tokens_unconfigured; an operator must act).

    application/jsonapplication/problem+json → Problem
  • default

    Problem details (RFC 9457)

    application/jsonapplication/problem+json → Problem
    Headers: X-Request-ID
get /api/v1/ai-models/storage

Model storage

Operation ai_models_storage_get · bearer token

The central-store shares and each node's local disk with the models cached on it, as last scanned. Nothing is scanned by this read.

Responses

  • 200

    The storage, as last scanned.

    application/json → Storage
    Headers: X-Request-ID
  • 401

    Not authenticated: no Authorization: Bearer header (not_authenticated), or the credential is refused (token_invalid, token_expired, token_revoked, principal_disabled, token_ip_not_allowed).

    application/jsonapplication/problem+json → Problem
  • 403

    The caller lacks a permission the operation needs (forbidden; required lists the keys, any one of which would do), or the licence refuses a change (licence_locked, licence_restricted, licence_required, with state and remedy; never retry a licence_* code).

    application/jsonapplication/problem+json → Problem
    Headers: X-Request-ID
  • 429

    More than 600 requests in a minute with this token (rate_limited), or the first request with this Idempotency-Key is still running (idempotency_request_in_progress). Retry after Retry-After.

    application/jsonapplication/problem+json → Problem
  • 503

    The platform cannot answer right now (unavailable; retry after Retry-After), or API tokens are not configured on it (api_tokens_unconfigured; an operator must act).

    application/jsonapplication/problem+json → Problem
  • default

    Problem details (RFC 9457)

    application/jsonapplication/problem+json → Problem
    Headers: X-Request-ID
get /api/v1/ai-models/{model_id}

One AI model

Operation ai_models_get · bearer token

One model of the catalogue, by id, as recorded. This read contacts no node and scans no storage.

Parameters

NameInTypeDescription
model_idrequired path string

The model's id (id on a model).

pattern ^[1-9][0-9]{0,8}$

Responses

  • 200

    The model.

    application/json → AiModel
    Headers: X-Request-ID
  • 401

    Not authenticated: no Authorization: Bearer header (not_authenticated), or the credential is refused (token_invalid, token_expired, token_revoked, principal_disabled, token_ip_not_allowed).

    application/jsonapplication/problem+json → Problem
  • 403

    The caller lacks a permission the operation needs (forbidden; required lists the keys, any one of which would do), or the licence refuses a change (licence_locked, licence_restricted, licence_required, with state and remedy; never retry a licence_* code).

    application/jsonapplication/problem+json → Problem
    Headers: X-Request-ID
  • 404

    No such model (model_not_found).

    application/jsonapplication/problem+json → Problem
    Headers: X-Request-ID
  • 429

    More than 600 requests in a minute with this token (rate_limited), or the first request with this Idempotency-Key is still running (idempotency_request_in_progress). Retry after Retry-After.

    application/jsonapplication/problem+json → Problem
  • 503

    The platform cannot answer right now (unavailable; retry after Retry-After), or API tokens are not configured on it (api_tokens_unconfigured; an operator must act).

    application/jsonapplication/problem+json → Problem
  • default

    Problem details (RFC 9457)

    application/jsonapplication/problem+json → Problem
    Headers: X-Request-ID
patch /api/v1/ai-models/{model_id}

Change an AI model's metadata

Operation ai_models_update · bearer token

Metadata only. repo is frozen; status, location, offline_ready and nas_volume are read-only (set by the store actions): each may be sent with its current value, and a different value is a 422.

The body is a JSON Merge Patch (RFC 7396), sent as application/merge-patch+json or application/json: a member that is omitted is left unchanged, and a member set to null clears that field where clearing is allowed (the schema marks those fields nullable; null for any other field is a 422).

Parameters

NameInTypeDescription
model_idrequired path string

The model's id (id on a model).

pattern ^[1-9][0-9]{0,8}$

Request bodyrequired

application/jsonapplication/merge-patch+json

Schema AiModelPatch

JSON Merge Patch (RFC 7396): a member that is omitted keeps its current value; a member sent as null clears the field when the field is nullable (the schema marks it so), and null for any other field is a 422. A "" is a value (an empty string), not a clear. Of the metadata: display_name and gated are not nullable. repo is frozen and status, location, offline_ready, nas_volume are read-only: each may be sent with its current value and nothing else.

NameTypeDescription
architecture string | null

The model's architecture, as recorded: dense or moe (mixture of experts). Other values may appear.

max length 40
benchmarks map of number (double) | integer | string | null

Published benchmark scores: benchmark name to score, e.g. {"SWE-bench Verified": 69.6}. As published by the vendor; not measured here.

category string | null

What the model is for: coding, agentic, reasoning, vision, embedding, rerank, general, speech or video. Other values may appear.

max length 24
context_window string | null

The context window in human form, as recorded, e.g. 128K or 256K→1M.

max length 20
description string | null

A longer introduction, for the model's detail view.

max length 20000
display_name string | null

The name the catalogue shows. On create it defaults to the last part of repo.

min length 1 · max length 255
frontier_equiv string | null

Which hosted frontier model it is roughly comparable to, and from when, in a few words. An estimate, not a measurement.

max length 120
gated boolean | null

The Hugging Face repo needs an accept-click (access approval) before a pull.

gateway_tier string | null

The AI gateway serving tier the model is meant for, e.g. code or general (GET /ai/gateway/tiers). Recorded only: which model backs a tier is decided there, not here.

max length 40
license string | null

The model's licence, as recorded, e.g. Apache-2.0, MIT.

max length 80
location string | null

Read-only.

min_target string | null

The smallest hardware it runs on, in human form, e.g. 1× 3090 or 2×DGX.

max length 40
model_card_url string | null

The model card's web address.

max length 300
nas_volume string | null

Read-only.

notes string | null

Free-form notes of the platform's operators.

max length 20000
offline_ready boolean | null

Read-only.

org string | null

The Hugging Face organisation the repo is published under. It may be a quantizer rather than the lab that built the model (that is vendor).

max length 120
param_count_b number (double) | null

The parameter count in billions, as a number to sort by.

min 0 · max 9999999
params string | null

The parameter count in human form, e.g. "480B (35B active)".

max length 60
published string | null

When the model was released upstream, as recorded, e.g. 2024-11.

max length 20
quant string | null

The weights' precision or quantisation, as recorded, e.g. BF16, FP8, AWQ, NVFP4.

max length 40
repo string | null

Frozen.

serving_node string | null

Where the model is meant to be served: an AI node's hostname or a DGX cluster, as recorded. Recorded only: loading a model is not part of this API.

max length 80
size_gb number (double) | null

The weights' size on disk in GB, as recorded.

min 0 · max 99999999
status string | null

Read-only.

strong_axis string | null

What the model is strongest at, in a few words, e.g. agentic coding.

max length 40
summary string | null

A one-line introduction, for a catalogue card.

max length 2000
vendor string | null

The lab that built the model. The Hugging Face organisation (org) may be a quantizer.

max length 200
vendor_country string | null

The vendor's home country, as recorded, e.g. China, France, USA.

max length 40

Responses

  • 200

    The model after the change.

    application/json → AiModel
    Headers: X-Request-ID
  • 401

    Not authenticated: no Authorization: Bearer header (not_authenticated), or the credential is refused (token_invalid, token_expired, token_revoked, principal_disabled, token_ip_not_allowed).

    application/jsonapplication/problem+json → Problem
  • 403

    The caller lacks a permission the operation needs (forbidden; required lists the keys, any one of which would do), or the licence refuses a change (licence_locked, licence_restricted, licence_required, with state and remedy; never retry a licence_* code).

    application/jsonapplication/problem+json → Problem
    Headers: X-Request-ID
  • 404

    No such model (model_not_found).

    application/jsonapplication/problem+json → Problem
    Headers: X-Request-ID
  • 422

    repo was changed (immutable_field), a read-only field was changed (read_only_field), or the body is invalid.

    application/jsonapplication/problem+json → Problem
    Headers: X-Request-ID
  • 429

    More than 600 requests in a minute with this token (rate_limited), or the first request with this Idempotency-Key is still running (idempotency_request_in_progress). Retry after Retry-After.

    application/jsonapplication/problem+json → Problem
  • 503

    The platform cannot answer right now (unavailable; retry after Retry-After), or API tokens are not configured on it (api_tokens_unconfigured; an operator must act).

    application/jsonapplication/problem+json → Problem
  • default

    Problem details (RFC 9457)

    application/jsonapplication/problem+json → Problem
    Headers: X-Request-ID
delete /api/v1/ai-models/{model_id}

Remove an AI model from the catalogue

Operation ai_models_delete · bearer token

Removes the catalogue ROW and nothing else: no disk is touched. It is refused while the row is the record of weights that exist (a central copy, or status owned / serving), while any node holds a cache, and while a store run is pending or running. Removing a node cache first is allowed (DELETE .../node-caches/{node}); removing a central copy is not possible through v1. Needs the admin permission; NOT destroy-gated, because it never destroys weights.

Parameters

NameInTypeDescription
model_idrequired path string

The model's id (id on a model).

pattern ^[1-9][0-9]{0,8}$

Responses

  • 204

    The catalogue row was removed.

    Headers: X-Request-ID
  • 401

    Not authenticated: no Authorization: Bearer header (not_authenticated), or the credential is refused (token_invalid, token_expired, token_revoked, principal_disabled, token_ip_not_allowed).

    application/jsonapplication/problem+json → Problem
  • 403

    The caller lacks a permission the operation needs (forbidden; required lists the keys, any one of which would do), or the licence refuses a change (licence_locked, licence_restricted, licence_required, with state and remedy; never retry a licence_* code).

    application/jsonapplication/problem+json → Problem
    Headers: X-Request-ID
  • 404

    No such model (model_not_found).

    application/jsonapplication/problem+json → Problem
    Headers: X-Request-ID
  • 409

    The model still has weights or work: an active store run (model_has_active_run), a node cache (model_has_node_caches) or a central copy (model_has_central_copy). blockers counts each: {"active_runs": n, "node_caches": n, "central_copy": 1}.

    application/jsonapplication/problem+json → Problem
    Headers: X-Request-ID
  • 429

    More than 600 requests in a minute with this token (rate_limited), or the first request with this Idempotency-Key is still running (idempotency_request_in_progress). Retry after Retry-After.

    application/jsonapplication/problem+json → Problem
  • 503

    The platform cannot answer right now (unavailable; retry after Retry-After), or API tokens are not configured on it (api_tokens_unconfigured; an operator must act).

    application/jsonapplication/problem+json → Problem
  • default

    Problem details (RFC 9457)

    application/jsonapplication/problem+json → Problem
    Headers: X-Request-ID
get /api/v1/ai-models/{model_id}/node-caches

List a model's node caches

Operation ai_models_node_caches_list · bearer token

The model's node-cache records, in node order: one per node that holds, held or is getting a local copy; state says which (cached once the copy is complete). Empty when no node ever cached the model; 404 when the model does not exist.

Parameters

NameInTypeDescription
model_idrequired path string

The model's id (id on a model).

pattern ^[1-9][0-9]{0,8}$
limit query integer

Page size.

min 1 · max 200 · default 50
cursor query string | null

next_cursor from the previous page; omit it for the first page. A cursor this list did not issue is a 400 invalid_cursor.

Responses

  • 200

    One page of the model's node caches.

    application/json → NodeCachePage
    Headers: X-Request-ID
  • 401

    Not authenticated: no Authorization: Bearer header (not_authenticated), or the credential is refused (token_invalid, token_expired, token_revoked, principal_disabled, token_ip_not_allowed).

    application/jsonapplication/problem+json → Problem
  • 403

    The caller lacks a permission the operation needs (forbidden; required lists the keys, any one of which would do), or the licence refuses a change (licence_locked, licence_restricted, licence_required, with state and remedy; never retry a licence_* code).

    application/jsonapplication/problem+json → Problem
    Headers: X-Request-ID
  • 404

    No such model (model_not_found).

    application/jsonapplication/problem+json → Problem
    Headers: X-Request-ID
  • 429

    More than 600 requests in a minute with this token (rate_limited), or the first request with this Idempotency-Key is still running (idempotency_request_in_progress). Retry after Retry-After.

    application/jsonapplication/problem+json → Problem
  • 503

    The platform cannot answer right now (unavailable; retry after Retry-After), or API tokens are not configured on it (api_tokens_unconfigured; an operator must act).

    application/jsonapplication/problem+json → Problem
  • default

    Problem details (RFC 9457)

    application/jsonapplication/problem+json → Problem
    Headers: X-Request-ID
get /api/v1/ai-models/{model_id}/node-caches/{node}

One node cache

Operation ai_models_node_caches_get · bearer token

The model's cache on one node. A record whose copy was removed (state absent) answers 404, like a node that never had one.

Parameters

NameInTypeDescription
model_idrequired path string

The model's id (id on a model).

pattern ^[1-9][0-9]{0,8}$
noderequired path string

The AI node's hostname, e.g. prod-ai-03.

pattern ^[A-Za-z0-9][A-Za-z0-9._-]{0,79}$

Responses

  • 200

    The node cache.

    application/json → NodeCache
    Headers: X-Request-ID
  • 401

    Not authenticated: no Authorization: Bearer header (not_authenticated), or the credential is refused (token_invalid, token_expired, token_revoked, principal_disabled, token_ip_not_allowed).

    application/jsonapplication/problem+json → Problem
  • 403

    The caller lacks a permission the operation needs (forbidden; required lists the keys, any one of which would do), or the licence refuses a change (licence_locked, licence_restricted, licence_required, with state and remedy; never retry a licence_* code).

    application/jsonapplication/problem+json → Problem
    Headers: X-Request-ID
  • 404

    No such model (model_not_found) or no cache of it on that node (node_cache_not_found).

    application/jsonapplication/problem+json → Problem
    Headers: X-Request-ID
  • 429

    More than 600 requests in a minute with this token (rate_limited), or the first request with this Idempotency-Key is still running (idempotency_request_in_progress). Retry after Retry-After.

    application/jsonapplication/problem+json → Problem
  • 503

    The platform cannot answer right now (unavailable; retry after Retry-After), or API tokens are not configured on it (api_tokens_unconfigured; an operator must act).

    application/jsonapplication/problem+json → Problem
  • default

    Problem details (RFC 9457)

    application/jsonapplication/problem+json → Problem
    Headers: X-Request-ID
put /api/v1/ai-models/{model_id}/node-caches/{node}

Cache a model on a node

Operation ai_models_node_caches_put · bearer token

Copies the model from the central store to the node's local disk (a fast, offline-ready serving copy) and answers 202 with operation model-store-run:<id>. The model must have a central copy. When the node already holds a cached copy the answer is 200 with the cache and nothing is dispatched.

The run is dispatched to the runner and followed by the portal, which survives an API restart; its time budget starts when the runner job starts, not while it queues. On an estate that fakes dispatch (sandbox, or dispatch mode dryrun or simulate) the operation says dispatch_mode: dryrun: the run holds a fake pipeline id, nothing reaches the runner and nothing is simulated — a store run is never simulated, so under simulate it behaves exactly as under dryrun and the operation ends failed. Poll GET /operations/{id} (also the Location header) until succeeded or failed. Idempotency-Key is honoured: a retry with the same key and request gets the same answer.

Parameters

NameInTypeDescription
model_idrequired path string

The model's id (id on a model).

pattern ^[1-9][0-9]{0,8}$
noderequired path string

The AI node's hostname, e.g. prod-ai-03.

pattern ^[A-Za-z0-9][A-Za-z0-9._-]{0,79}$
Idempotency-Key header string

Makes a retry safe: for 24 hours, the same key with the same request (method, path, query and body) answers with the stored response (Idempotent-Replayed: true) instead of doing the work again. The same key with a different request is a 409 idempotency_key_reused; while the first request is still running it is a 429 idempotency_request_in_progress with Retry-After. 1-255 printable characters without spaces (400 invalid_idempotency_key otherwise); a UUID is a good key. Keys are scoped to the calling account.

min length 1 · max length 255 · pattern ^[!-~]+$

Responses

  • 200

    The node already holds a cached copy; nothing was dispatched.

    application/json → NodeCache
  • 202

    The copy started: poll the operation.

    application/json → Operation
  • 400

    The Idempotency-Key is not 1-255 printable characters (invalid_idempotency_key).

    application/jsonapplication/problem+json → Problem
    Headers: X-Request-ID
  • 401

    Not authenticated: no Authorization: Bearer header (not_authenticated), or the credential is refused (token_invalid, token_expired, token_revoked, principal_disabled, token_ip_not_allowed).

    application/jsonapplication/problem+json → Problem
  • 403

    The caller lacks a permission the operation needs (forbidden; required lists the keys, any one of which would do), or the licence refuses a change (licence_locked, licence_restricted, licence_required, with state and remedy; never retry a licence_* code).

    application/jsonapplication/problem+json → Problem
    Headers: X-Request-ID
  • 404

    No such model (model_not_found).

    application/jsonapplication/problem+json → Problem
    Headers: X-Request-ID
  • 409

    A store run of this model is pending or running (run_in_progress; operation_id names it), the key was used for another request (idempotency_key_reused), or — cache — the model has no central copy to cache from (no_central_copy), or — uncache — it is loaded on the node (model_loaded_on_node).

    application/jsonapplication/problem+json → Problem
    Headers: X-Request-ID
  • 422

    The node is not an AI node (unknown_node) or has no management address (node_has_no_mgmt_ip), the stored repo is not a strict Hugging Face id (unsafe_repo), or the model's share is not a current one (unknown_volume).

    application/jsonapplication/problem+json → Problem
    Headers: X-Request-ID
  • 429

    More than 600 requests in a minute with this token (rate_limited), or the first request with this Idempotency-Key is still running (idempotency_request_in_progress). Retry after Retry-After.

    application/jsonapplication/problem+json → Problem
  • 502

    The runner pipeline could not be triggered (dispatch_failed; the run is recorded as failed and operation_id names it).

    application/jsonapplication/problem+json → Problem
    Headers: X-Request-ID
  • 503

    The platform cannot answer right now (unavailable; retry after Retry-After), or API tokens are not configured on it (api_tokens_unconfigured; an operator must act).

    application/jsonapplication/problem+json → Problem
  • default

    Problem details (RFC 9457)

    application/jsonapplication/problem+json → Problem
    Headers: X-Request-ID
delete /api/v1/ai-models/{model_id}/node-caches/{node}

Remove a model's cache from a node

Operation ai_models_node_caches_delete · bearer token

Deletes the node's local copy (recoverable: the central copy is untouched) and answers 202 with operation model-store-run:<id>. Refused while the model is loaded on that node, and — because that cannot be proven otherwise — while monitoring cannot be read.

The run is dispatched to the runner and followed by the portal, which survives an API restart; its time budget starts when the runner job starts, not while it queues. On an estate that fakes dispatch (sandbox, or dispatch mode dryrun or simulate) the operation says dispatch_mode: dryrun: the run holds a fake pipeline id, nothing reaches the runner and nothing is simulated — a store run is never simulated, so under simulate it behaves exactly as under dryrun and the operation ends failed. Poll GET /operations/{id} (also the Location header) until succeeded or failed. Idempotency-Key is honoured: a retry with the same key and request gets the same answer.

Parameters

NameInTypeDescription
model_idrequired path string

The model's id (id on a model).

pattern ^[1-9][0-9]{0,8}$
noderequired path string

The AI node's hostname, e.g. prod-ai-03.

pattern ^[A-Za-z0-9][A-Za-z0-9._-]{0,79}$
Idempotency-Key header string

Makes a retry safe: for 24 hours, the same key with the same request (method, path, query and body) answers with the stored response (Idempotent-Replayed: true) instead of doing the work again. The same key with a different request is a 409 idempotency_key_reused; while the first request is still running it is a 429 idempotency_request_in_progress with Retry-After. 1-255 printable characters without spaces (400 invalid_idempotency_key otherwise); a UUID is a good key. Keys are scoped to the calling account.

min length 1 · max length 255 · pattern ^[!-~]+$

Responses

  • 202

    The removal started: poll the operation.

    application/json → Operation
  • 400

    The Idempotency-Key is not 1-255 printable characters (invalid_idempotency_key).

    application/jsonapplication/problem+json → Problem
    Headers: X-Request-ID
  • 401

    Not authenticated: no Authorization: Bearer header (not_authenticated), or the credential is refused (token_invalid, token_expired, token_revoked, principal_disabled, token_ip_not_allowed).

    application/jsonapplication/problem+json → Problem
  • 403

    The caller lacks a permission the operation needs (forbidden; required lists the keys, any one of which would do), or the licence refuses a change (licence_locked, licence_restricted, licence_required, with state and remedy; never retry a licence_* code).

    application/jsonapplication/problem+json → Problem
    Headers: X-Request-ID
  • 404

    No such model (model_not_found), or no cache of it on that node (node_cache_not_found).

    application/jsonapplication/problem+json → Problem
    Headers: X-Request-ID
  • 409

    A store run of this model is pending or running (run_in_progress; operation_id names it), the key was used for another request (idempotency_key_reused), or — cache — the model has no central copy to cache from (no_central_copy), or — uncache — it is loaded on the node (model_loaded_on_node).

    application/jsonapplication/problem+json → Problem
    Headers: X-Request-ID
  • 422

    The node is not an AI node (unknown_node) or has no management address (node_has_no_mgmt_ip), the stored repo is not a strict Hugging Face id (unsafe_repo), or the model's share is not a current one (unknown_volume).

    application/jsonapplication/problem+json → Problem
    Headers: X-Request-ID
  • 429

    More than 600 requests in a minute with this token (rate_limited), or the first request with this Idempotency-Key is still running (idempotency_request_in_progress). Retry after Retry-After.

    application/jsonapplication/problem+json → Problem
  • 502

    The runner pipeline could not be triggered (dispatch_failed; the run is recorded as failed and operation_id names it).

    application/jsonapplication/problem+json → Problem
    Headers: X-Request-ID
  • 503

    Whether the model is loaded on the node cannot be read right now (loaded_state_unknown); nothing was dispatched. Also: the platform cannot answer right now (unavailable; retry after Retry-After), or API tokens are not configured on it (api_tokens_unconfigured; an operator must act).

    application/jsonapplication/problem+json → Problem
  • default

    Problem details (RFC 9457)

    application/jsonapplication/problem+json → Problem
    Headers: X-Request-ID

Rendered from openapi-v1.json, platform release 1.0.187. Your installation serves the contract of its own version at /api/v1/openapi.json.