Speech-to-Text
Payment wires: every paid endpoint accepts x402 and MPP (Machine Payments Protocol) on the same 402 - see Paying with x402 and Paying with MPP. Agent402 is the applied layer of Agentic Finance: agents that pay and get paid on their own.
Acceptable use. Hosted-instance traffic is governed by the Terms of Service - including a generative-content acceptable-use policy - and by the upstream model providers' usage policies. Wallets used for prohibited content are blocked before settlement. Outputs are generated by third-party models from your inputs; you are responsible for how you use them.
Two tiers of audio transcription, paywalled via x402. Provide a URL to an audio file, get back the transcript with language detection and duration. The operator's OPENAI_API_KEY handles the default path; diarize: true runs on OPENROUTER_API_KEY.
Tiers
| Endpoint | Price | Model | Max duration |
|---|---|---|---|
POST /api/transcribe |
$0.03 | gpt-transcribe |
4 min |
POST /api/transcribe-pro |
$0.10 | gpt-transcribe |
10 min |
Both tiers run the same model and differ only in the duration cap.
Speaker labels and word timestamps
Add "diarize": true (or a diarize=true form field on /v1/audio/transcriptions) to transcribe with ElevenLabs Scribe v2 instead, at the same price and cap. The answer adds speakers (how many were heard) and words, each with start, end and a speaker number:
{ "model": "elevenlabs/scribe-v2", "provider": "openrouter", "text": "Hello there. Hi.",
"language": "eng", "duration": 1.0, "speakers": 2,
"words": [ { "word": "Hello", "start": 0, "end": 0.3, "speaker": 0 },
{ "word": "there.", "start": 0.3, "end": 0.5, "speaker": 0 },
{ "word": "Hi.", "start": 0.6, "end": 0.9, "speaker": 1 } ] }
Both tiers are wallet-only - every call burns real upstream transcription credit. See Security Model.
Supported audio formats
mp3, mp4, mpeg, mpga, m4a, wav, ogg, flac, webm (max 25 MB)
Request / Response
// Request
{ "url": "https://example.com/audio.mp3", "language": "en" }
// Response
{ "model": "gpt-transcribe", "provider": "openai",
"text": "Hello, this is a sample transcription.",
"language": "en", "duration": 3.5 }
Only url is required. language (ISO-639-1 code) is optional but improves accuracy.
See also
- Text-to-Speech - the reverse: text to audio
- LLM Proxy Gateway - text inference
- Paying with x402 - the USDC payment flow