/v1Authorization: Bearer AI_API_KEYapplication/jsonThis gateway sits between your apps and many LLM providers. Clients speak the OpenAI API; the gateway resolves the model through routes (v1-style ordered lists), aliases, an auto-router, a fallback chain and a per-provider cooldown, then adapts the request to each upstream (OpenAI, Anthropic, Gemini, Azure).
# 1. Generate a downstream key in the dashboard (/dashboard -> API Keys)
# 2. Call the gateway
curl https://YOUR_WORKER.dev/v1/chat/completions \
-H "Authorization: Bearer AI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4o-mini",
"messages": [{"role":"user","content":"Hello!"}]
}'
Every /v1/* request must carry a downstream key issued by the dashboard:
| Header | Value |
|---|---|
Authorization | Bearer <key> — the literal placeholder used in these docs is AI_API_KEY |
Content-Type | application/json (POST) |
If the secret AI_API_KEY is set on the Worker it is also accepted as a master key
(no expiry, no rate limit, no model allowlist).
| State | HTTP | Code |
|---|---|---|
| Missing / unknown key | 401 | invalid_api_key |
Expired (exp reached) | 403 | key_expired |
| Revoked in dashboard | 403 | key_disabled |
| Over per-minute limit | 429 | rate_limit_exceeded |
| Over monthly quota | 429 | quota_exceeded |
| Method | Path | Description |
|---|---|---|
| POST | /v1/chat/completions | Chat completion, streaming and non-streaming |
| GET | /v1/models | List gateway models (aliases included) |
| GET | /v1 | Endpoint discovery |
| GET | /health | Liveness probe (no auth) |
| GET | /docs | This page |
| GET | /dashboard | Admin console |
| Field | Type | Notes |
|---|---|---|
model | string | Gateway id, alias, or any auto-routed upstream model. Falls back to the configured default model. |
messages | array | OpenAI message objects (system / user / assistant / tool), text and image parts. |
stream | boolean | SSE chunks in OpenAI format, terminated with data: [DONE]. |
temperature, top_p, max_tokens, stop | mixed | Passed through; gateway model params act as defaults. |
tools, tool_choice | array | Function calling, translated for Anthropic and Gemini. |
user | string | Forwarded to the upstream when present. |
curl https://YOUR_WORKER.dev/v1/chat/completions \
-H "Authorization: Bearer AI_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"smart","messages":[{"role":"user","content":"Hi"}],"temperature":0.4}'
Response (non-streaming):
{
"id": "chatcmpl-9f2c...",
"object": "chat.completion",
"created": 1712000000,
"model": "gpt-4o-mini",
"choices": [{"index":0,"message":{"role":"assistant","content":"Hello!"},
"finish_reason":"stop"}],
"usage": {"prompt_tokens":11,"completion_tokens":9,"total_tokens":20}
}
Streaming ("stream": true):
data: {"id":"chatcmpl-1","object":"chat.completion.chunk","model":"gpt-4o-mini",
"choices":[{"index":0,"delta":{"content":"Hel"},"finish_reason":null}]}
data: {"id":"chatcmpl-1","object":"chat.completion.chunk",
"choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}
data: [DONE]
Response headers: x-gateway-request-id, x-gateway-provider,
x-gateway-model, x-gateway-latency-ms, x-gateway-attempts
(how many candidates were tried), and x-gateway-degraded: cooldown
when every healthy target was cooling down.
{
"object": "list",
"data": [
{"id":"gpt-4o-mini","object":"model","created":1712000000,"owned_by":"gateway",
"aliases":["mini","fast"],"targets":["openai/gpt-4o-mini"]}
]
}
openai/gpt-4o-mini ↔ gpt-4o-mini), priority order."fast": ["groq/llama-3.3-70b-versatile", …]) — called with "model": "fast", tried in exactly that order."model": "auto" picks a usable target. Default strategy random (v1: uniform distribution, prompt never inspected); switch to priority (v2) in Settings.fallbacks models are tried next.{
"id": "smart",
"aliases": ["gpt-4o", "best"],
"targets": [
{"provider": "openai", "model": "gpt-4o"},
{"provider": "azure", "model": "gpt4o-deploy"}
],
"fallbacks": ["cheap"],
"params": {"temperature": 0.7}
}
Retry policy (v1 fast retry): up to maxAttempts attempts per request
(default 4), every candidate tried in turn, including on
400 / 409 / 413 / 422 — providers disagree about geo blocks, unsupported parameters and model
names — and the first client error is relayed once every candidate failed, with an
x-gateway-attempts header. 401 / 403 put the provider into cooldown and rotate to
the next target. 404, 408, 429, 5xx and network errors rotate to the next target after an
exponential backoff (250 ms × 2^n, capped at 2 s). Upstream response headers slower than
headerTimeoutMs (default 15000 ms) abort the attempt as a timeout — the timer is cleared
as soon as the headers arrive, so long streams are never cut off. The last failure is returned
as 502 upstream_error (or the original status for client errors).
All errors use the OpenAI envelope:
{"error":{"message":"...","type":"rate_limit_error","param":null,"code":"rate_limit_exceeded"}}
| Status | Code | Meaning |
|---|---|---|
| 400 | invalid_messages / missing_model / invalid_json | Malformed request |
| 401 | invalid_api_key | Unknown downstream key |
| 403 | key_expired / key_disabled / model_not_allowed | Key or model not permitted |
| 404 | model_not_found | No configured target, and no enabled provider matched (exact or fuzzy) |
| 429 | rate_limit_exceeded / quota_exceeded | Key limits hit |
| 502 | upstream_error | All candidates failed upstream |
| 503 | no_available_provider | No usable provider for this model |
/dashboard manages providers, models, keys, settings and usage. It talks to
/api/admin/*, protected by a session cookie (POST /api/admin/login with the
admin token, or a first-run setup token).
| Method & path | Purpose |
|---|---|
GET /api/admin/bootstrap | Everything the dashboard needs in one call |
POST /api/admin/login · /logout | Session (first call also performs first-run setup) |
PUT /api/admin/settings | Default model, auto-route, cooldown, timeout, CORS |
GET/POST /api/admin/providers | List / create upstream providers |
PUT/DELETE /api/admin/providers/{id} | Update / remove a provider |
POST /api/admin/providers/{id}/test | Send a 1-token probe to the provider |
GET /api/admin/providers/{id}/models | Fetch the provider's model list |
GET/POST /api/admin/models | List / upsert gateway models |
PUT/DELETE /api/admin/models/{id} | Update / remove a gateway model |
GET/POST /api/admin/keys | List / generate downstream keys |
PATCH/DELETE /api/admin/keys/{token} | Edit limits or revoke |
GET /api/admin/usage?days=7 | Usage rows per key per day |
DELETE /api/admin/cooldowns/{provider} | Clear a cooldown manually |
| Binding / secret | Required | Purpose |
|---|---|---|
GATEWAY_KV (KV namespace) | yes | Providers, models, keys, usage, cooldowns |
ADMIN_TOKEN (secret) | no | Dashboard admin token (otherwise set it on first login) |
AI_API_KEY (secret) | no | Extra master downstream key |
| Setting | Default | Description |
|---|---|---|
defaultModel | "" | Used when the request has no model |
autoRoute | true | Route unknown names through providers advertising them (exact, then fuzzy) and enable the built-in auto model |
routes | {} | v1-style ordered route lists: {"fast": ["groq/llama-3.3-70b-versatile", …]} — called with "model": "fast", tried in that exact order |
autoStrategy | random | random (v1): uniform pick for model: "auto" · priority (v2): provider priority order |
requestTimeoutMs | 120000 | Upstream timeout until response headers arrive |
headerTimeoutMs | 15000 | v1 fast retry: abort a chat attempt when upstream response headers are slower than this |
maxAttempts | 4 | v1 fast retry: total upstream attempts per request, including the first |
retryBaseDelayMs / retryMaxDelayMs | 250 / 2000 | Exponential backoff between attempts (250 ms × 2^n, capped) |
cooldown.threshold / windowMs / durationMs | 3 / 60000 / 120000 | Circuit breaker sensitivity |
maxFallbackDepth | 3 | How deep the fallback chain may recurse |
rateLimitEnabled / per-key rpm | true / 0 | Requests per minute per key (0 = unlimited) |
per-key quota | 0 | Requests per calendar month (0 = unlimited) |
per-key expiresInDays | none | Expiry timestamp for the key |