Build
API reference.
The OpenAI-compatible endpoints, their parameters, and what each response carries.
On this page
The base URL is https://api.usdf.fi/v1. Requests and responses follow the OpenAI chat completions format. The Responses API is not served.
Two ways to pay for a request:
- No account: send the same body with no key to the same path under
https://api.usdf.fi/x402and pay per request in USDF. - A prepaid balance: send
Authorization: Bearer sk_…to the path below.
The model endpoints and the reads around them, with what each needs:
| Endpoint | What it does | Auth |
|---|---|---|
POST /v1/chat/completions | Chat, streaming, tools and vision. Also submits batch jobs. | Key, or pay per request under /x402 |
POST /v1/completions | Legacy text completions. | Key, or pay per request under /x402 |
POST /v1/embeddings, /images/generations, /images/edits, /audio/transcriptions, /audio/speech, /rerank, /videos/generations | The other modalities. See Modalities. | Key, or pay per request under /x402 |
GET /v1/models | The models that serve right now. | None |
GET /v1/pricing | The full price sheet, every listed model. | None |
GET /v1/pricing/sheets, /{version} | Every price sheet ever served. | None |
GET /v1/pricing/observations | Upstream list prices as observed. | None |
GET /v1/catalog, /signer, /signature | The signed model catalog. | None |
GET /v1/batches/{id}, /result | A batch job, and its result. | Key |
GET /v1/tools | Tools, their prices and input schemas. | None |
POST /v1/tools/{id} | Runs one tool on the balance. | Key |
GET /v1/receipts | Your latest receipts. | Key or session |
GET /v1/receipts/{id} | One receipt, with its place in the usage log. | The key that made the request |
GET /x402/v1/receipts/{id} | The public receipt of any request. | None |
GET /v1/limits | Every rate limit, live. | None |
GET /v1/status | Upstream health, latency and incidents. | None |
GET /health | The gateway's own health. | None |
Public data, with no key. Each one backs a page of the site:
| Endpoint | Returns | On the site |
|---|---|---|
GET /v1/index, /history.json, /history.csv, /tiers, /methodology | The daily compute index in USDF per million tokens, per tier and as a composite, from the gateway's own settled traffic. Each row is signed. | Compute index |
GET /v1/rankings | Models ranked by tokens served across every account, over ?days=. | Rankings |
GET /v1/network | Totals across every account served, over ?days=. | Network |
GET /v1/stats | All-time requests, metered and settled spend, active keys and funded accounts. | Transparency |
GET /v1/ecosystem | The figures the site quotes, each with the endpoint it is read from. | Docs |
GET /v1/holders/windows | Each day's holder distribution window, with its pool and Merkle root. | Earn |
GET /v1/susdf | The sUSDF vault: its address, TVL, share price, yield and revenue deposits. | Earn |
GET /v1/privacy/attestation | The attestation report behind private inference. | Private inference |
Every limit these reads share is in Rate limits. Account, agent, platform and token endpoints are on their own pages: Accounts and keys, Agents, Platforms and providers, USDF token and Receipts and verification.
POST /v1/chat/completions takes the standard body. Samples use <model>, the first available chat model.
Chat completion
curl https://api.usdf.fi/v1/chat/completions \
-H "Authorization: Bearer $USDF_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "<model>",
"messages": [{"role": "user", "content": "Hello"}],
"max_tokens": 256
}'Response, trimmed
x-request-id: <request_id>
x-cost-units: <units>
{
"id": "<request_id>",
"object": "chat.completion",
"model": "<model>",
"choices": [{ "index": 0, "message": { "role": "assistant", "content": "…" }, "finish_reason": "stop" }],
"usage": { "prompt_tokens": …, "completion_tokens": …, "total_tokens": … },
"receipt": { "request_id": "<request_id>", "status": "ok", "cost": { … }, … }
}The parameters the gateway reads:
| Parameter | Type | Required | Notes |
|---|---|---|---|
model | String | Yes | A model id. Suffixes such as :cheapest are in Routing. |
messages | Array | Yes | Non-empty. Images go in message content. |
max_tokens, max_completion_tokens | Integer | No | Positive, up to the model's max_output. Sets the hold. Above it: 400 context_length_exceeded. |
stream | Boolean | No | Server-sent events. |
stream_options.include_usage | Boolean | No | A final usage chunk. |
temperature | Number | No | 0 to 2. |
n | Integer | No | Must be 1. |
tools, tool_choice, parallel_tool_calls | Array, string or object | No | Function calling. |
response_format | Object | No | JSON mode or a JSON schema. |
reasoning_effort | String | No | For reasoning models. |
top_p, top_k, min_p, stop, seed, frequency_penalty, presence_penalty, repetition_penalty, logit_bias, logprobs, top_logprobs, user | As OpenAI | No | Passed through to the provider. |
provider | Object | No | Routing and privacy. See Provider preferences. |
batch, webhook_url | Boolean, string | No | Submit as a job. See Batch jobs. |
- The deprecated
functionsandfunction_callare refused with400 unsupported_parameter. Usetools. - Fields that would change what the provider bills are dropped.
- Up to 16 images per request. More gets
400 too_many_images. - The response
idis the request id, also in thex-request-idheader. A response that is not streamed carries the cost inx-cost-units; a stream carries it in its receipt line.
How the hold is sized
The prompt is held at two bytes per token, plus eight tokens per message and 16 per request, plus max_tokens. A model with long-context rates is held at one byte per token, so the hold never lands in a cheaper tier.
Without max_tokens, the model's default output budget is held, shrinking to what the balance covers. The charge is the provider's reported usage; the rest of the hold is released.
Set "stream": true for server-sent events in the standard chunk format.
The order of events:
Server-sent events
data: {"id":"<request_id>","object":"chat.completion.chunk",…}
data: {"id":"<request_id>","choices":[],"usage":{…}}
: receipt {"request_id":"<request_id>","cost":{"units":"…","usd":"…","usdf":"…"},…}
data: [DONE]The receipt arrives as a comment line just before [DONE]. Standard clients drop comment lines, so read the receipt by id afterwards with GET /v1/receipts/{id}. Every chunk's id is the request id.
Stream, then read the receipt
curl -N https://api.usdf.fi/v1/chat/completions \
-H "Authorization: Bearer $USDF_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "<model>",
"stream": true,
"stream_options": {"include_usage": true},
"messages": [{"role": "user", "content": "Hello"}]
}'POST /v1/completions serves only models whose endpoints include completions. prompt is a string, or an array of one string.
Text completion
curl https://api.usdf.fi/v1/completions \
-H "Authorization: Bearer $USDF_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "<model>",
"prompt": "The capital of France is",
"max_tokens": 16
}'GET /v1/models lists what serves now. GET /v1/pricing is the full sheet, every listed model with its prices and routes.
The fields of GET /v1/pricing:
| Field | Meaning |
|---|---|
version, unit | The sheet's version and its default unit. |
rate_lock | days and valid_until: how long the sheet's rates are promised. |
models[].id, name, owned_by | The model id to send, its name and its maker. |
modality, unit | What the model does, and what its price is per. |
endpoints | Where it serves: chat, completions, embeddings and others. |
context, max_output | The context window and the largest output. |
price | input, output, cached_input, cache_write_5m, cache_write_1h, the batch rates batch_input, batch_output, batch_cached_input, batch_cache_write_5m, batch_cache_write_1h, and long_context with its threshold. |
price_private | The private-tier rate. |
available, last_verified, rate_locked_until | Whether it serves now, when it last passed a check, and how long its price is promised. |
routes[] | Each provider route, with enabled and its data_policy. |
Each price sheet promises its rates for 30 days. Every sheet ever served is at GET /v1/pricing/sheets, and the signed catalog at GET /v1/catalog. Browse them on Models and Pricing.
A batch job runs at the upstream's batch rate. Submit it with a :batch model suffix or "batch": true, without stream. Add webhook_url to be told when it ends.
A job runs on a model whose price carries batch_input in GET /v1/pricing.
Submit a job
curl https://api.usdf.fi/v1/chat/completions \
-H "Authorization: Bearer $USDF_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "<model with batch rates>:batch",
"messages": [{"role": "user", "content": "Hello"}],
"max_tokens": 256,
"webhook_url": "https://example.com/usdf-hooks"
}'Response, 202
{
"object": "batch.job",
"id": "batch_…",
"status": "submitted",
"rate": "batch",
"hold": { "units": "<units>", "usd": "<usd>", "usdf": "<usdf>" },
"price": { … },
"poll": "https://api.usdf.fi/v1/batches/batch_…",
"result": "https://api.usdf.fi/v1/batches/batch_…/result"
}Poll the job, then fetch its result. The result is kept 7 days, then answers 410.
Poll and fetch
curl https://api.usdf.fi/v1/batches/<batch id> \ -H "Authorization: Bearer $USDF_API_KEY" curl https://api.usdf.fi/v1/batches/<batch id>/result \ -H "Authorization: Bearer $USDF_API_KEY"
How a job ends:
| Status | Meaning | Charged |
|---|---|---|
done | The result is ready. | Yes, at the batch rates |
failed | The upstream failed it. Its result answers 409 batch_failed. | No |
expired | The upstream did not finish in time. 409 batch_expired. | No |
The hold is taken at the standard rate and the charge is made at the batch_* rates when the result arrives.
A job is refused, and nothing is charged, in these cases:
| Case | Error |
|---|---|
| The model has no batch rates | 400 batch_unavailable |
The body asks for stream | 400 batch_unavailable |
| The key is in no-logs mode: a job's result is stored until fetched | 400 batch_unavailable |
| The request asks for the private tier, or the key is private tier | 400 private_batch_unavailable |
| The request is paid per request | 400 batch_unavailable |
| The gateway is not taking jobs right now | 503 batch_unavailable, with Retry-After |
See Batch.
GET /v1/tools is public. Each tool lists its price (flat per call, or a base plus a price per result), its enabled flag, and its input_schema.
POST /v1/tools/{id} with a key runs one. The body is the tool's input. A call takes a hold, charges what it actually cost, and writes a receipt. A failed call is not charged. A call to a tool whose enabled is false is refused with 409 tool_unavailable before any hold.
The tools, read live:
| Tool | Price | What it does |
|---|---|---|
| Loading… | ||
The sample sends the required fields of the tool's input_schema.
Call a tool
curl https://api.usdf.fi/v1/tools/<tool id> \
-H "Authorization: Bearer $USDF_API_KEY" \
-H "Content-Type: application/json" \
-d '{}'The headers the gateway reads and sends:
| Header | Direction | Meaning |
|---|---|---|
Authorization | Request | A key, a session, a read token, or Payment for MPP. |
Content-Type | Request | application/json, or multipart/form-data for uploads. |
Idempotency-Key | Request | Makes an agent's fund or pay call safe to repeat. |
x-approval-id | Request | Releases a request an agent's owner approved. |
x-erc8004-agent | Request | Names the ERC-8004 agent behind a paid request. |
PAYMENT-SIGNATURE | Request | An x402 payment. |
x-request-id | Response | The request id, also the receipt id. |
x-cost-units | Response | What the request cost, in units, on a response that is not streamed. A chat or text completion paid per request carries receipt.paid instead. |
PAYMENT-REQUIRED | Response | The x402 payment requirements, on a 402. |
WWW-Authenticate: Payment | Response | The MPP challenges, on a 402. |
PAYMENT-RESPONSE | Response | The settled x402 payment. |
Payment-Receipt | Response | The settled MPP payment. |
Retry-After | Response | On a 429 or 503: how many seconds to wait. |
Three public reads show how the gateway is doing:
| Endpoint | Returns |
|---|---|
GET /health | ok, the workers, the upstreams and the chain block. |
GET /v1/status | Each upstream's state, latency, open incidents and the refund rate. |
GET /v1/status/history | Per-day uptime over 30-day, 60-day or 90-day windows. |
The same figures are on Status.
curl https://api.usdf.fi/health