Subtitle pipeline
Audio URL → finished subtitles in one call: transcribe the audio, emit the transcript as SRT/WebVTT/JSON cues, and report the text statistics - length, reading time, word count.
3 tools run server-side in one request. You pay once, settle once, and get a single response - no orchestration, no per-step payments, and a partial-success envelope if any step fails. USDC over x402 on any supported chain.
When to use this pack
An agent processing podcasts, voice notes, or video audio needs shippable subtitle files plus the stats to budget downstream steps (summarization, translation, chapters) - without stitching three tools by hand.
Tools in this pack
All 3 run inside the single $0.03 call above. Each is also callable on its own if you only need one part.
- Speech-to-text POST /api/transcribe Transcribe audio to text using OpenAI (gpt-transcribe). Provide a URL to an audio file (mp3, wav, m4a, etc.) and get back the transcript. Add diarize:true for speaker labels and word timestamps (ElevenLabs Scribe v2, same price). No API key needed; pay per call via x402. Max 4 minutes of audio, 25 MB file size; /api/transcribe-pro takes the same models to 10 minutes.
- Subtitle convert (SRT/VTT) POST /api/srt-convert Convert subtitles between SRT, WebVTT, plain text, and JSON cues. Send SRT or VTT text (auto-detected) - or a JSON cues array [{start,end,text}] with times in ms - and the target format. Deterministic, pure CPU.
- Text statistics POST /api/text-stats Counts for any text in one call: characters (Unicode code points), words, sentences, paragraphs, avgWordLength, readingTimeMinutes (200 words a minute) and estimatedTokens, a rough LLM token estimate (about 4 characters a token, one per Chinese or Japanese character). Chinese and Japanese text, which has no spaces, is counted a character per word, and 。!? end sentences. Use it to budget a prompt or size a document before sending it; for an exact count on one model's tokenizer use a tokenizer tool.
Bought one at a time, these 3 tools cost $0.033 together; the pack is that sum less a 10% bundle discount, rounded up to the $0.001 settlement floor, which is $0.03.
Workflow
- Transcribe the audio with transcribe - OpenAI speech-to-text with language detection and duration.
- Convert the transcript into subtitle cues with srt-convert in your chosen format (SRT, WebVTT, plain text, or JSON cues).
- Run text-stats over the transcript - word count, sentence count, and estimated reading time for downstream budgeting.
Arguments
| Name | Required | Description | Example |
|---|---|---|---|
url | yes | Public URL of the audio file to transcribe | https://agent402.tools/fixtures/sample-speech.wav |
format | no | Subtitle output format: srt | vtt | text | json (default vtt) | vtt |
What one call returns
A JSON object with pack, args, steps, summary; steps holds one entry per tool (transcribe, srt-convert, text-stats), each with its own result or error. Full example on the API page.
Call it directly
Any x402 client pays the 402 and gets the whole workflow back in one response. With the agent402-client SDK (npm i agent402-client, an ES module):
import { Agent402 } from "agent402-client";
// payFetch: an x402-wrapped fetch your wallet signs (@x402/fetch).
// Tools on the free tier need no options: new Agent402() pays them by proof-of-work.
// an existing prepaid credits key also works: new Agent402({ creditsKey })
const client = new Agent402({ fetch: payFetch });
const result = await client.call("skill-subtitle-pipeline", {"url":"https://agent402.tools/fixtures/sample-speech.wav","format":"vtt"});
Run it in Claude
claude mcp add agent402 -s user -- npx -y agent402-mcp@latest
Then paste this prompt into Claude:
Turn the audio at https://agent402.tools/fixtures/sample-speech.wav into subtitles using Agent402's subtitle-pipeline skill pack. (1) Transcribe the audio, (2) convert the transcript to vtt subtitles, (3) get the text statistics. Return the subtitle file content, the detected language and duration, and the word count.