Speech-to-text, long audio (OpenAI transcription wire)

$0.100 per call · USDC via x402 · POST /v1/pro/audio/transcriptions

OpenAI's transcription wire on the ten-minute tier: POST multipart/form-data with a `file` part. Send POST /v1/pro/audio/transcriptions with the required field file and pay $0.100 per call over x402 or MPP (there is no free tier). It returns a JSON object with text, duration and model.

Same model as /v1/audio/transcriptions with a longer cap, for a recording a four-minute route refuses.

Category: AI & compute · Tags: stt speech-to-text transcription audio whisper openai

TRY IN PLAYGROUND →

Parameters

NameTypeRequiredDescription
filestringyesThe audio file, as a multipart part named `file`
languagestringnoOptional ISO-639-1 hint
diarizestringno"true" for speaker labels and word timestamps (ElevenLabs Scribe v2)

Example request

curl -i -X POST https://agent402.tools/v1/pro/audio/transcriptions \
  -H "Content-Type: application/json" \
  -d '{"file":"<audio bytes, multipart part named file>"}'

Without payment this returns HTTP 402 Payment Required with the exact price for v1-audio-transcriptions-pro; any x402 v2 or MPP client pays it and retries.

Example response

{
  "text": "Example transcript.",
  "duration": 420.5,
  "model": "gpt-transcribe"
}
FieldTypeAlways presentIn the example
textstringyesExample transcript.
durationnumberyes420.5
modelstringyesgpt-transcribe

From an MCP client

catalog.call {
  "slug": "v1-audio-transcriptions-pro",
  "params": {
    "file": "<audio bytes, multipart part named file>"
  }
}

The hosted connector at https://agent402.tools/mcp needs a payment for v1-audio-transcriptions-pro; the stdio package pays it from a wallet or from AGENT402_CREDITS_KEY. Local install: npx -y agent402-mcp.

Errors and behavior

Paid call (JavaScript agent)

import { wrapFetchWithPayment } from "@x402/fetch";
import { x402Client } from "@x402/core/client";
import { registerExactEvmScheme } from "@x402/evm/exact/client";
import { privateKeyToAccount } from "viem/accounts";

const client = new x402Client();
client.setSpendControls?.(false); // keep your own spending ceiling in code
registerExactEvmScheme(client, { signer: privateKeyToAccount(KEY) });
const payFetch = wrapFetchWithPayment(fetch, client);

const res = await payFetch("https://agent402.tools/v1/pro/audio/transcriptions", {
  method: "POST",
  headers: { "Content-Type": "application/json" },
  body: JSON.stringify({
    "file": "<audio bytes, multipart part named file>"
  }),
});

Related tools

Speech-to-text

$0.030 · POST /api/transcribe

Transcribe audio to text using OpenAI (gpt-transcribe). Provide a URL to an audio file (mp3, wav, m4a, etc.) and get bac…

Speech-to-text (Pro)

$0.100 · POST /api/transcribe-pro

Transcribe audio to text using OpenAI (gpt-transcribe) - the same model as /api/transcribe with a longer cap. Provide a …

Speech-to-text (OpenAI transcription wire)

$0.030 · POST /v1/audio/transcriptions

OpenAI's own transcription wire: POST multipart/form-data with a `file` part and get the transcript back. Point any Whis…

Decisions (OpenAI wire)

$0.001 · POST /v1/decisions

OpenAI's Decisions API on its own wire: point an OpenAI SDK's base URL here and client.decisions.create works unchanged,…

Text embeddings

$0.002 · POST /api/embed

Generate a text embedding vector using OpenAI text-embedding-3-small (1536 dimensions). Ideal for semantic search, RAG, …

Text embeddings (Large)

$0.010 · POST /api/embed-large

Generate a text embedding vector using OpenAI text-embedding-3-large (3072 dimensions). Higher accuracy than the small m…