Speech-to-text
POST /api/transcribeTranscribe audio to text using OpenAI (gpt-transcribe). Send POST /api/transcribe with the required field url and pay $0.030 per call over x402 or MPP (there is no free tier). It returns a JSON object with model, provider, text, language and duration.
Provide a URL to an audio file (mp3, wav, m4a, etc.) and get back the transcript. Add diarize:true for speaker labels and word timestamps (ElevenLabs Scribe v2, same price). No API key needed; pay per call via x402. Max 4 minutes of audio, 25 MB file size; /api/transcribe-pro takes the same models to 10 minutes.
Parameters
| Name | Type | Required | Description |
|---|---|---|---|
url | string | yes | URL of the audio file to transcribe (mp3, wav, m4a, ogg, flac, webm) Also accepted as link, uri, href, page. |
language | string | no | Optional ISO-639-1 language code (e.g. 'en', 'es', 'fr') for better accuracy |
diarize | boolean | no | true: transcribe with ElevenLabs Scribe v2 and add speaker labels and word timestamps (`words`, `speakers`). Same price and cap. Default false |
Example request
curl -i -X POST https://agent402.tools/api/transcribe \
-H "Content-Type: application/json" \
-d '{"url":"https://agent402.tools/fixtures/sample-speech.wav"}'
Without payment this returns HTTP 402 Payment Required with the exact price for transcribe; any x402 v2 or MPP client pays it and retries.
Example response
{
"model": "gpt-transcribe",
"provider": "openai",
"text": "Hi, this is the Agent402 sample recording. Can you hear me clearly? Yes, loud and clear. There are two of us on this clip, so the transcript can tell us apart.",
"language": "en",
"duration": 12.6
}
| Field | Type | Always present | In the example |
|---|---|---|---|
model | string | yes | gpt-transcribe |
provider | string | yes | openai |
text | string | yes | Hi, this is the Agent402 sample recording. Can you hear me clearly? Yes, loud... |
language | string | yes | en |
duration | number | yes | 12.6 |
From an MCP client
catalog.call {
"slug": "transcribe",
"params": {
"url": "https://agent402.tools/fixtures/sample-speech.wav"
}
}
The hosted connector at https://agent402.tools/mcp needs a payment for transcribe; the stdio package pays it from a wallet or from AGENT402_CREDITS_KEY. Local install: npx -y agent402-mcp.
Errors and behavior
urlis required. An input the tool rejects returns an HTTP 4xx whose body carrieserror,tool,expected,requiredandexample, so the caller can correct it.- A paid call that ends in any status of 400 or above is not charged over x402, MPP or a prepaid credits key: settlement is cancelled when the tool fails. The exception is a Tempo push credential, a transfer the buyer sent before the call: it settles before the tool runs, so if the tool then fails the payment is recorded as a refund owed to the paying wallet.
- Wallet-only: this tool runs a model, so it has no proof-of-work tier. A prepaid card-credits key issued earlier (
Authorization: Bearer a402_...) also pays it. - Model-backed: the answer is generated by a model, so the same input can produce different wording.
- A
GETorHEADto /api/transcribe returns the same 402 quote, so the price can be read without a body. - An
Idempotency-Keyheader makes a retried paid call replay the first 200 instead of charging again (an answer larger than 1 MB is not replayed).
Paid call (JavaScript agent)
import { wrapFetchWithPayment } from "@x402/fetch";
import { x402Client } from "@x402/core/client";
import { registerExactEvmScheme } from "@x402/evm/exact/client";
import { privateKeyToAccount } from "viem/accounts";
const client = new x402Client();
client.setSpendControls?.(false); // keep your own spending ceiling in code
registerExactEvmScheme(client, { signer: privateKeyToAccount(KEY) });
const payFetch = wrapFetchWithPayment(fetch, client);
const res = await payFetch("https://agent402.tools/api/transcribe", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({
"url": "https://agent402.tools/fixtures/sample-speech.wav"
}),
});
Part of these workflows
Speech-to-text is one step in this skill pack, each sold as a single call:
- Subtitle pipeline - Audio URL → finished subtitles in one call: transcribe the audio, emit the transcript as SRT/WebVTT/JSON cues, and report the text statistics - length, reading time, word count.
Related tools
Speech-to-text (Pro)
POST /api/transcribe-proTranscribe audio to text using OpenAI (gpt-transcribe) - the same model as /api/transcribe with a longer cap. Provide a …
Speech-to-text (OpenAI transcription wire)
POST /v1/audio/transcriptionsOpenAI's own transcription wire: POST multipart/form-data with a `file` part and get the transcript back. Point any Whis…
Speech-to-text, long audio (OpenAI transcription wire)
POST /v1/pro/audio/transcriptionsOpenAI's transcription wire on the ten-minute tier: POST multipart/form-data with a `file` part. Same model as /v1/audio…
Decisions (OpenAI wire)
POST /v1/decisionsOpenAI's Decisions API on its own wire: point an OpenAI SDK's base URL here and client.decisions.create works unchanged,…
Text embeddings
POST /api/embedGenerate a text embedding vector using OpenAI text-embedding-3-small (1536 dimensions). Ideal for semantic search, RAG, …
Text embeddings (Large)
POST /api/embed-largeGenerate a text embedding vector using OpenAI text-embedding-3-large (3072 dimensions). Higher accuracy than the small m…