ArkorAlpha

Chat Completions API

The full request and response specification for *.arkor.app inference endpoints.

Every Arkor endpoint serves the same OpenAI-compatible surface at https://<slug>.arkor.app. This page is the wire-level reference; it applies equally to one-click endpoints and project endpoints.

Routes

MethodPathPurpose
POST/v1/chat/completionsChat completion (JSON or SSE streaming)
POST/v1/messagesAnthropic-native Messages API, for Anthropic-dialect models only
GET/v1/modelsList the model ids the endpoint serves
GET/v1/chat/completions/runs/{id}Replay a stored run as SSE
GET/healthzLiveness check, returns { "ok": true }

GET /v1/models returns the standard OpenAI list shape. For a multi-model endpoint it lists the current public catalog (the list changes as models are published); for every other endpoint it lists the single pinned model. The route sits behind the same API-key auth as chat completions, and responses may be cached for up to 30 seconds.

Each model is served on exactly one wire dialect. Most speak the OpenAI chat completions surface this page documents; Anthropic models (claude-* ids) are instead served via Anthropic's native Messages API at POST /v1/messages on the same endpoint — point an Anthropic SDK at the base URL with the same API key. GET /v1/models lists both kinds without distinguishing them, and sending a model to the wrong surface returns 400 with a message naming the right route.

Authentication

Endpoints in fixed API key mode accept the key in either header:

curl
curl https://<slug>.arkor.app/v1/chat/completions \
  -H "Authorization: Bearer $ARKOR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"messages":[{"role":"user","content":"Hello"}]}'

x-api-key: <key> works identically. Endpoints in no auth mode skip the check entirely.

Auth failures return 401 with a machine-readable code:

CodeMeaning
missing_keyNo key presented
invalid_keyKey does not match any enabled key
no_keysThe endpoint requires a key but has none configured

These code values appear on the OpenAI-shaped routes only. The same auth failures on POST /v1/messages render in Anthropic's error envelope instead — { "type": "error", "error": { "type": "authentication_error", "message": "..." } } — which carries no code field.

Revoked keys can keep authenticating for up to about 30 seconds while edge caches refresh.

Request body

POST /v1/chat/completions takes the OpenAI chat completions shape. Requests larger than 1 MiB are rejected with 413.

FieldTypeNotes
messagesarray, requiredAt least one chat message. content is a plain string or OpenAI's content-part array with text parts only ({ "type": "text", "text": "..." }); multimodal parts such as image_url, input_audio, or file are rejected with 400
modelstringOn single-model endpoints: accepted and ignored (the endpoint pins the model). On multi-model endpoints: selects the model by id (matched case-insensitively); omitting it uses the endpoint's default model, and an unknown id returns 404 with code model_not_found. GET /v1/models lists the valid ids
streambooleanDefaults to false (JSON response)
stream_options.include_usagebooleanOpt in to a final usage chunk when streaming
nintegerOnly 1 (the OpenAI default) is accepted; any other value returns 400 — the pipeline serves a single choice
temperaturenumber0 to 2
top_pnumber0 to 1
max_tokensintegerPositive
max_completion_tokensintegerOpenAI's newer name for the output-token cap. Both are accepted; when a request carries both, max_completion_tokens wins
stopstring or arrayStop sequences. Non-empty strings; a list holds up to 16
presence_penaltynumber-2 to 2
frequency_penaltynumber-2 to 2
seedintegerBest-effort deterministic sampling
logprobsbooleanReturn log-probabilities for output tokens
top_logprobsinteger0 to 20. Values above 0 need logprobs: true (enforced by the model server)
logit_biasobjectStringified token id → bias, -100 to 100
tools / tool_choiceOpenAI shapesFunction calling. Only automatic selection (tool_choice omitted or "auto") needs a tool-parser-enabled model; a named function, "required", and "none" work without one
response_formatOpenAI shapeIncluding json_schema structured outputs
structured_outputsobjectvLLM extension: json / regex / choice / grammar / json_object constraint fields, mutually exclusive
enable_thinkingbooleanReasoning mode, on supported models
reasoning_effortstringOne of none, minimal, low, medium, high, xhigh, max (anything else is a 400). Gemma 4 only distinguishes thinking on/off: none turns thinking off, every other value turns it on
userstringOpenAI's end-user attribution field. Feeds the same Usage attribution as X-Arkor-End-User-Id (the header wins when both are present), so the value is stored in usage history and grouped by. Subject to the same limits as the attribution headers — up to 256 characters of printable Latin-1; anything else is dropped silently. Never used for authorization

Messages use the four OpenAI roles system, user, assistant, and tool — the newer developer role is not accepted and returns 400. Assistant messages may carry tool_calls (in which case content is optional), a tool message answers one via its required tool_call_id, and the optional name participant label is accepted on system/user/assistant messages.

Unknown fields are stripped rather than rejected, so OpenAI SDK extras do not cause errors. Tool definitions, JSON schemas, and structured_outputs constraints are forwarded verbatim to the model server.

Responses and streaming

With stream: false (the default) you get a standard chat.completion JSON object. With stream: true you get OpenAI-style SSE chunks terminated by data: [DONE]. Usage numbers stream only when stream_options.include_usage was requested.

Models running with a reasoning parser emit their reasoning as delta.reasoning_content in streaming chunks, separate from the visible answer.

When run retention is enabled, responses persisted by Arkor carry an X-Arkor-Run-Id header identifying the stored run (see Run history). Responses served by an external third-party provider (e.g. Gemini or Claude models on a multi-model endpoint) are never persisted and carry no run id, regardless of the retention setting.

Replaying stored runs

GET /v1/chat/completions/runs/{id} re-emits a stored run's SSE stream. It supports resuming via Last-Event-ID or a ?from= query parameter. Unknown, expired, or foreign run ids return 410 with code run_gone.

Attribution headers

Two optional request headers label the request in Usage breakdowns:

  • X-Arkor-End-User-Id — your application's user id.
  • X-Arkor-Session-Id — a conversation or session id.

Values are limited to 256 characters of printable Latin-1. Invalid values are dropped silently; they never cause a rejection and are never used for authorization.

The OpenAI user body field feeds the same end-user attribution; when both the field and the X-Arkor-End-User-Id header are present, the header wins. On the Anthropic Messages route, the standard metadata.user_id field plays the same role with the same precedence (the header wins). The limits above apply to these body fields too: a value over 256 characters or outside printable Latin-1 (for example Japanese text) is dropped silently, so the request succeeds but carries no end-user attribution. Whichever value survives is stored with the request's usage row, so do not send identifiers you do not want recorded.

Errors

StatusWhen
400Invalid JSON, schema violation, or conflicting response_format / structured_outputs
401Auth failure (see codes above)
403The selected model requires an Arkor account (code model_requires_signup). Premium catalog models still appear in GET /v1/models on anonymous endpoints, but dispatching them there is rejected with a signup prompt
404Unknown subdomain, disabled endpoint, or expired endpoint (single uniform message); on multi-model endpoints also an unknown model id (code model_not_found)
410Stored run gone (replay route)
413Body over 1 MiB
429Rate limited (code rate_limit_exceeded), for either of two reasons: the endpoint owner's inference quota is exceeded (anonymous one-click endpoints have stricter per-minute and per-day limits than account-owned ones), or — on a multi-model request served by an external provider — the vendor's own rate limit was exhausted. Either way the response carries a Retry-After header (seconds, forwarded from the vendor in the external case); wait at least that long before retrying
499Client closed the connection mid-request
502Upstream model server error
503Provider temporarily unavailable (code provider_env_unavailable); safe to retry
504Every inference candidate timed out before responding (code inference_upstream_timeout); safe to retry
529Every attempted external provider reported itself overloaded (code overloaded, mirroring Anthropic's convention). The vendor's Retry-After is forwarded when present; retry later

On the OpenAI-compatible routes (POST /v1/chat/completions and GET /v1/models) error bodies use OpenAI's nested envelope:

Error body (OpenAI-compatible routes)
{ "error": { "message": "...", "type": "...", "param": null, "code": "..." } }

The Arkor-specific run replay route (GET /v1/chat/completions/runs/{id}) keeps a flat envelope instead: { "error": string, "code"?: string }.

CORS

Endpoints send permissive CORS (Access-Control-Allow-Origin: *, no credentials), so browsers can call them directly. Allowed request headers include Authorization, Content-Type, x-api-key, Last-Event-ID, and the attribution headers above.