Build
Embeddings, images, audio, rerank and video.
Every modality on the same key, the same balance and the same receipts.
Every modality is paid the two ways chat is. With no account, send the request to the same path under /x402 and pay per request in USDF; see Endpoints paid per request. With a key, it is held, charged on what the provider reports, and receipted against the balance. The samples on this page use a key.
The models that serve each one now, and what each is billed per, read live from GET /v1/pricing:
| Modality | Endpoint | Billed per | Models available now |
|---|---|---|---|
| Loading… | |||
Privacy asks work as on chat: provider.zdr, provider.data_collection and provider.private. On a multipart endpoint, send provider as a JSON string form field.
POST /v1/embeddings turns text into vectors.
| Parameter | Type | Required | Notes |
|---|---|---|---|
model | String | Yes | An embedding model. |
input | String or array | Yes | One string, or up to 1,024 strings. |
encoding_format | String | No | float or base64. The SDKs' default, base64, is honoured. |
dimensions | Integer | No | Where the model supports it. |
The hold is one token per UTF-8 byte of the input. The charge is the provider's prompt_tokens.
Embed two texts
curl https://api.usdf.fi/v1/embeddings \
-H "Authorization: Bearer $USDF_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "<embedding model>",
"input": ["The first text", "The second text"]
}'Response, trimmed
{
"object": "list",
"data": [{ "object": "embedding", "index": 0, "embedding": [ … ] }, … ],
"model": "<embedding model>",
"usage": { "prompt_tokens": …, "total_tokens": … },
"receipt": { "request_id": "<request_id>", "cost": { … }, … }
}Generate
POST /v1/images/generations takes JSON.
| Parameter | Type | Required | Notes |
|---|---|---|---|
model | String | Yes | An image model. |
prompt | String | Yes | What to draw. |
n | Integer | No | 1 to 4 images. Default 1. |
size | String | No | WxH, 128 to 1920 pixels a side. Default 1024x1024. |
response_format | String | No | b64_json only. |
Generate an image
curl https://api.usdf.fi/v1/images/generations \
-H "Authorization: Bearer $USDF_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "<image model>",
"prompt": "A lighthouse at dawn, watercolor",
"n": 1,
"size": "1024x1024",
"response_format": "b64_json"
}'Response, trimmed
{
"created": …,
"data": [{ "b64_json": "…" }],
"model": "<image model>",
"receipt": { "request_id": "<request_id>", "billed": { … }, "cost": { … }, … }
}Edit
POST /v1/images/edits takes multipart: image and prompt (both required), mask, model, n and size. Models whose endpoints include image_edits serve it.
Edit an image
curl https://api.usdf.fi/v1/images/edits \ -H "Authorization: Bearer $USDF_API_KEY" \ -F model=<image edit model> \ -F prompt="The same scene at night" \ -F image=@input.png
Images are billed per image unit: 1024 x 1024 output pixels at the model's default steps, measured from the bytes returned. The response carries created, data, model and the receipt.
A bad size gets 400 invalid_size. A body that is not multipart gets not_multipart, and one that cannot be parsed bad_multipart.
POST /v1/audio/speech turns text into audio.
| Parameter | Type | Required | Notes |
|---|---|---|---|
model | String | Yes | A speech model. |
input | String | Yes | Up to the model's context in characters. Longer gets 400 input_too_long. |
voice | String | No | One of the model's own voices: the names depend on the model. Left out, the model's default voice is used. |
response_format | String | No | mp3, opus, flac, wav or pcm. |
speed | Number | No | 0.25 to 4. |
Speech is billed per million characters. The response body is the audio. The receipt id is in x-request-id and the cost in x-cost-units; the full receipt is at GET /v1/receipts/{id}.
Synthesize speech
curl https://api.usdf.fi/v1/audio/speech \
-H "Authorization: Bearer $USDF_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "<speech model>",
"input": "Hello from USDF.",
"response_format": "mp3"
}' \
-D - --output speech.mp3POST /v1/audio/transcriptions turns audio into text. It takes multipart.
| Parameter | Type | Required | Notes |
|---|---|---|---|
file | File | Yes | WAV (PCM or float), FLAC, MP3, or Ogg (Opus or Vorbis). Others get 400 unsupported_audio. Over 4 hours: 400 audio_too_long. |
model | String | Yes | A transcription model. |
language | String | No | The spoken language. |
prompt | String | No | Context for the model. |
response_format | String | No | json, text or verbose_json. |
temperature | Number | No | 0 to 1. |
timestamp_granularities[] | String | No | segment or word. |
The hold is the longest audio the file can decode to. The charge is per minute, from the duration the provider reports.
Transcribe a file
curl https://api.usdf.fi/v1/audio/transcriptions \ -H "Authorization: Bearer $USDF_API_KEY" \ -F model=<transcription model> \ -F file=@meeting.mp3 \ -F response_format=json
POST /v1/rerank scores documents against a query.
| Parameter | Type | Required | Notes |
|---|---|---|---|
model | String | Yes | A rerank model. |
query or queries | String or array | Yes | What to rank against. |
documents | Array | Yes | Queries times documents: at most 1,024 pairs. |
instruction | String | No | Up to 2,048 characters. |
Rerank is billed per 1M tokens, on the provider's input_tokens.
Rerank two documents
curl https://api.usdf.fi/v1/rerank \
-H "Authorization: Bearer $USDF_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "<rerank model>",
"query": "How is USDF backed?",
"documents": [
"USDF wraps USDG one to one.",
"The gateway serves many models."
]
}'Response, trimmed
{
"scores": [ … ],
"input_tokens": …,
"results": [{ "index": 0, "relevance_score": … }, … ],
"receipt": { "request_id": "<request_id>", "cost": { … }, … }
}POST /v1/videos/generations starts a job and answers 202 at once. Poll the job, then fetch the clip.
| Parameter | Type | Required | Notes |
|---|---|---|---|
model | String | Yes | A video model. |
prompt | String | Yes | Up to the model's context in characters. |
seconds | Integer | No | The clip length. Each model has its own range, and some make one fixed length. Left out, the model's default is used. A length outside the range gets 400 invalid_seconds, and the message names the range. |
resolution | String | No | The model's one resolution, or left out. size is refused with 400 size_unsupported. |
aspect_ratio | String | No | One of the model's own. |
image | String | No | A first frame, as a PNG or JPEG data URL, where the model takes one. |
negative_prompt | String | No | What to avoid. |
seed | Integer | No | For repeatable output. |
Make a clip
curl https://api.usdf.fi/v1/videos/generations \
-H "Authorization: Bearer $USDF_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "<video model>",
"prompt": "A paper boat drifting down a river, slow pan",
"seconds": <seconds>
}'
# Poll until "status" is "completed" or "failed".
curl https://api.usdf.fi/v1/videos/<id> \
-H "Authorization: Bearer $USDF_API_KEY"
# The clip, billed on this first fetch. Content-Type says mp4 or webm.
curl https://api.usdf.fi/v1/videos/<id>/content \
-H "Authorization: Bearer $USDF_API_KEY" \
-D - --output clipResponse, 202
{
"id": "<request_id>",
"status": "queued",
"seconds": <seconds>,
"resolution": "<resolution>",
"aspect_ratio": "…",
"hold": { "units": "<units>", "usd": "<usd>", "usdf": "<usdf>" },
"poll": "https://api.usdf.fi/v1/videos/<request_id>",
"content": "https://api.usdf.fi/v1/videos/<request_id>/content"
}Poll GET /v1/videos/{id} until status is completed or failed. GET /v1/videos/{id}/content returns the clip as video/mp4 or video/webm. Only the key that made a job reads it.
A clip is billed per second, at the model's one resolution, on its first fetch. One never fetched is billed when it expires, 30 minutes after it was made. A key has at most two unbilled jobs; a third gets 429 too_many_video_jobs. A clip is at most 32 MiB. A gateway restart ends a job unbilled.
Paid per request, the 202 carries a read_token. Read the job with Authorization: Bearer <read_token>.
Images to edit and audio to transcribe are uploaded as multipart bodies.
- A body is at most 25 MiB. It is held in memory and never written to disk.
- A request with no key, or a malformed one, is refused with
401before its body is read. A larger declared body gets413. - One gateway instance holds eight uploads at a time, two per key. Over that:
429 too_many_uploadswithRetry-After. - Paid per request, send the multipart body whole both times: the quote is read from the file.