LLM Gateway /docs

Base URL/v1
AuthAuthorization: Bearer AI_API_KEY
Content typeapplication/json
CompatibilityOpenAI SDK drop-in

This gateway sits between your apps and many LLM providers. Clients speak the OpenAI API; the gateway resolves the model through routes (v1-style ordered lists), aliases, an auto-router, a fallback chain and a per-provider cooldown, then adapts the request to each upstream (OpenAI, Anthropic, Gemini, Azure).

Quick start

# 1. Generate a downstream key in the dashboard (/dashboard -> API Keys)
# 2. Call the gateway
curl https://YOUR_WORKER.dev/v1/chat/completions \
  -H "Authorization: Bearer AI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4o-mini",
    "messages": [{"role":"user","content":"Hello!"}]
  }'

Authentication

Every /v1/* request must carry a downstream key issued by the dashboard:

HeaderValue
AuthorizationBearer <key> — the literal placeholder used in these docs is AI_API_KEY
Content-Typeapplication/json (POST)

If the secret AI_API_KEY is set on the Worker it is also accepted as a master key (no expiry, no rate limit, no model allowlist).

Key lifecycle

StateHTTPCode
Missing / unknown key401invalid_api_key
Expired (exp reached)403key_expired
Revoked in dashboard403key_disabled
Over per-minute limit429rate_limit_exceeded
Over monthly quota429quota_exceeded

Endpoints

MethodPathDescription
POST/v1/chat/completionsChat completion, streaming and non-streaming
GET/v1/modelsList gateway models (aliases included)
GET/v1Endpoint discovery
GET/healthLiveness probe (no auth)
GET/docsThis page
GET/dashboardAdmin console

POST /v1/chat/completions

FieldTypeNotes
modelstringGateway id, alias, or any auto-routed upstream model. Falls back to the configured default model.
messagesarrayOpenAI message objects (system / user / assistant / tool), text and image parts.
streambooleanSSE chunks in OpenAI format, terminated with data: [DONE].
temperature, top_p, max_tokens, stopmixedPassed through; gateway model params act as defaults.
tools, tool_choicearrayFunction calling, translated for Anthropic and Gemini.
userstringForwarded to the upstream when present.
curl https://YOUR_WORKER.dev/v1/chat/completions \
  -H "Authorization: Bearer AI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"smart","messages":[{"role":"user","content":"Hi"}],"temperature":0.4}'

Response (non-streaming):

{
  "id": "chatcmpl-9f2c...",
  "object": "chat.completion",
  "created": 1712000000,
  "model": "gpt-4o-mini",
  "choices": [{"index":0,"message":{"role":"assistant","content":"Hello!"},
               "finish_reason":"stop"}],
  "usage": {"prompt_tokens":11,"completion_tokens":9,"total_tokens":20}
}

Streaming ("stream": true):

data: {"id":"chatcmpl-1","object":"chat.completion.chunk","model":"gpt-4o-mini",
       "choices":[{"index":0,"delta":{"content":"Hel"},"finish_reason":null}]}

data: {"id":"chatcmpl-1","object":"chat.completion.chunk",
       "choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}

data: [DONE]

Response headers: x-gateway-request-id, x-gateway-provider, x-gateway-model, x-gateway-latency-ms, x-gateway-attempts (how many candidates were tried), and x-gateway-degraded: cooldown when every healthy target was cooling down.

GET /v1/models

{
  "object": "list",
  "data": [
    {"id":"gpt-4o-mini","object":"model","created":1712000000,"owned_by":"gateway",
     "aliases":["mini","fast"],"targets":["openai/gpt-4o-mini"]}
  ]
}

Model routing

AliasesOne public id maps to one logical model with an ordered target list.
Auto-routerUnknown model name? Enabled providers are matched exactly first, then fuzzily (openai/gpt-4o-mini ↔ gpt-4o-mini), priority order.
Routesv1-style ordered route lists in Settings → routes (e.g. "fast": ["groq/llama-3.3-70b-versatile", …]) — called with "model": "fast", tried in exactly that order.
auto model"model": "auto" picks a usable target. Default strategy random (v1: uniform distribution, prompt never inspected); switch to priority (v2) in Settings.
FallbackIf all targets fail, the configured fallbacks models are tried next.
CooldownAfter N failures inside a window the provider is skipped for a while (circuit breaker).
{
  "id": "smart",
  "aliases": ["gpt-4o", "best"],
  "targets": [
    {"provider": "openai",  "model": "gpt-4o"},
    {"provider": "azure",   "model": "gpt4o-deploy"}
  ],
  "fallbacks": ["cheap"],
  "params": {"temperature": 0.7}
}

Retry policy (v1 fast retry): up to maxAttempts attempts per request (default 4), every candidate tried in turn, including on 400 / 409 / 413 / 422 — providers disagree about geo blocks, unsupported parameters and model names — and the first client error is relayed once every candidate failed, with an x-gateway-attempts header. 401 / 403 put the provider into cooldown and rotate to the next target. 404, 408, 429, 5xx and network errors rotate to the next target after an exponential backoff (250 ms × 2^n, capped at 2 s). Upstream response headers slower than headerTimeoutMs (default 15000 ms) abort the attempt as a timeout — the timer is cleared as soon as the headers arrive, so long streams are never cut off. The last failure is returned as 502 upstream_error (or the original status for client errors).

Errors

All errors use the OpenAI envelope:

{"error":{"message":"...","type":"rate_limit_error","param":null,"code":"rate_limit_exceeded"}}
StatusCodeMeaning
400invalid_messages / missing_model / invalid_jsonMalformed request
401invalid_api_keyUnknown downstream key
403key_expired / key_disabled / model_not_allowedKey or model not permitted
404model_not_foundNo configured target, and no enabled provider matched (exact or fuzzy)
429rate_limit_exceeded / quota_exceededKey limits hit
502upstream_errorAll candidates failed upstream
503no_available_providerNo usable provider for this model

Dashboard & admin API

/dashboard manages providers, models, keys, settings and usage. It talks to /api/admin/*, protected by a session cookie (POST /api/admin/login with the admin token, or a first-run setup token).

Method & pathPurpose
GET /api/admin/bootstrapEverything the dashboard needs in one call
POST /api/admin/login · /logoutSession (first call also performs first-run setup)
PUT /api/admin/settingsDefault model, auto-route, cooldown, timeout, CORS
GET/POST /api/admin/providersList / create upstream providers
PUT/DELETE /api/admin/providers/{id}Update / remove a provider
POST /api/admin/providers/{id}/testSend a 1-token probe to the provider
GET /api/admin/providers/{id}/modelsFetch the provider's model list
GET/POST /api/admin/modelsList / upsert gateway models
PUT/DELETE /api/admin/models/{id}Update / remove a gateway model
GET/POST /api/admin/keysList / generate downstream keys
PATCH/DELETE /api/admin/keys/{token}Edit limits or revoke
GET /api/admin/usage?days=7Usage rows per key per day
DELETE /api/admin/cooldowns/{provider}Clear a cooldown manually

Configuration reference

Binding / secretRequiredPurpose
GATEWAY_KV (KV namespace)yesProviders, models, keys, usage, cooldowns
ADMIN_TOKEN (secret)noDashboard admin token (otherwise set it on first login)
AI_API_KEY (secret)noExtra master downstream key
SettingDefaultDescription
defaultModel""Used when the request has no model
autoRoutetrueRoute unknown names through providers advertising them (exact, then fuzzy) and enable the built-in auto model
routes{}v1-style ordered route lists: {"fast": ["groq/llama-3.3-70b-versatile", …]} — called with "model": "fast", tried in that exact order
autoStrategyrandomrandom (v1): uniform pick for model: "auto" · priority (v2): provider priority order
requestTimeoutMs120000Upstream timeout until response headers arrive
headerTimeoutMs15000v1 fast retry: abort a chat attempt when upstream response headers are slower than this
maxAttempts4v1 fast retry: total upstream attempts per request, including the first
retryBaseDelayMs / retryMaxDelayMs250 / 2000Exponential backoff between attempts (250 ms × 2^n, capped)
cooldown.threshold / windowMs / durationMs3 / 60000 / 120000Circuit breaker sensitivity
maxFallbackDepth3How deep the fallback chain may recurse
rateLimitEnabled / per-key rpmtrue / 0Requests per minute per key (0 = unlimited)
per-key quota0Requests per calendar month (0 = unlimited)
per-key expiresInDaysnoneExpiry timestamp for the key

Open Dashboard