Text-to-Speech
Payment wires: every paid endpoint accepts x402 and MPP (Machine Payments Protocol) on the same 402 - see Paying with x402 and Paying with MPP. Agent402 is the applied layer of Agentic Finance: agents that pay and get paid on their own.
Acceptable use. Hosted-instance traffic is governed by the Terms of Service - including a generative-content acceptable-use policy - and by the upstream model providers' usage policies. Wallets used for prohibited content are blocked before settlement. Outputs are generated by third-party models from your inputs; you are responsible for how you use them.
Three tiers of text-to-speech, paywalled via x402. Send text, get back base64-encoded audio. The same request shape and the same 10 voice names on every tier. All three run on the operator's OPENROUTER_API_KEY; until OpenAI retires tts-1 and tts-1-hd on 2027-01-06, an OPENAI_API_KEY lets the two ElevenLabs tiers fall back to them during an ElevenLabs outage.
Tiers
| Endpoint | Price | Model | Quality | Formats | Text cap |
|---|---|---|---|---|---|
POST /api/tts-lite |
$0.005 | hexgrad/kokoro-82m |
Synthetic-sounding, the cheapest | mp3, pcm | 800 chars |
POST /api/tts |
$0.12 | elevenlabs/eleven-v4-turbo |
Natural, low latency | all six | 2,000 chars |
POST /api/tts-hd |
$0.24 | elevenlabs/eleven-v4 |
Most expressive; reads audio tags such as [whispering] |
all six | 2,000 chars |
Each ElevenLabs tier tries same-rate ElevenLabs models in turn if the first is unavailable (tts: Eleven v4 Turbo, Turbo v2.5, Flash v2.5; tts-hd: Eleven v4, v3, Multilingual v2). If ElevenLabs is busy, throttling or returning errors, the call is served by a backup instead of failing: OpenAI tts-1 / tts-1-hd until OpenAI retires them on 2027-01-06, then Grok voice or MAI-Voice-2 (fewer voices and languages). A call that timed out or came back empty is not retried, since it may already have been generated; it fails uncharged. /api/gateway-status reports speechUpstream as ok, cooling or probing. The model field in the answer always names the one that spoke.
Which to use. The lite tier is for high-volume narration, notifications and agent speech, where the cost per call matters more than the timbre; the voice is audibly synthetic next to the ElevenLabs tiers, which is the trade. It takes the same ten voice names and maps each to the nearest Kokoro voice, naming the one that spoke in the response. Ask for a format it cannot serve and you get a 400 pointing at /api/tts, never a silent downgrade.
All three tiers are wallet-only - every call burns real upstream TTS credit. See Security Model.
Voices
alloy (default), ash, ballad, coral, echo, fable, nova, onyx, sage, shimmer. On tts and tts-hd each maps to its own ElevenLabs voice:
| Name | ElevenLabs voice |
|---|---|
| alloy | river |
| ash | chris |
| ballad | callum |
| coral | jessica |
| echo | eric |
| fable | george |
| nova | sarah |
| onyx | brian |
| sage | matilda |
| shimmer | lily |
Those two tiers also take any ElevenLabs voice by name: george, sarah, adam, alice, bella, bill, brian, callum, charlie, chris, daniel, eric, harry, jessica, laura, liam, lily, matilda, river, roger, will.
Output formats
mp3 (default), opus, aac, flac, wav, pcm
Request / Response
// Request
{ "text": "Hello from Agent402!", "voice": "alloy", "format": "mp3" }
// Response
{ "model": "elevenlabs/eleven-v4-turbo", "provider": "openrouter", "voice": "river", "format": "mp3",
"audio": "<base64-encoded audio>", "chars": 20 }
Only text is required. voice defaults to alloy, format defaults to mp3.
See also
- Speech-to-Text - the reverse: audio to text
- LLM Proxy Gateway - text inference
- Paying with x402 - the USDC payment flow