LLM inference
POST /api/llmLLM inference proxy - send an OpenAI-format chat/completions request and get a response from GPT-4o-mini. Send POST /api/llm with the required fields model and messages and pay $0.010 per call over x402 or MPP (there is no free tier). It returns a JSON object with model, provider, usage and choices.
Supports vision (up to 2 image URLs, low detail) and structured output (response_format: json_object or json_schema). No API key needed; pay per call via x402. Input capped at 16k chars, output at 4096 tokens.
Parameters
| Name | Type | Required | Description |
|---|---|---|---|
model | string | yes | Model ID - gpt-4o-mini |
messages | array | yes | Array of {role, content} objects. content can be a string or array of {type:'text',text} and {type:'image_url',image_url:{url,detail}} blocks |
max_tokens | number | no | Max output tokens (default 1024, cap 4096) |
response_format | object | no | Optional: {type:"json_object"} or {type:"json_schema",json_schema:{name,schema}} |
temperature | number | no | Sampling temperature (0-2) |
top_p | number | no | Nucleus sampling (0-1) |
stop | string | no | Stop sequence(s) |
Example request
curl -i -X POST https://agent402.tools/api/llm \
-H "Content-Type: application/json" \
-d '{"model":"gpt-4o-mini","messages":[{"role":"user","content":"Say hello in one sentence."}],"max_tokens":64}'
Without payment this returns HTTP 402 Payment Required with the exact price for llm; any x402 v2 or MPP client pays it and retries.
Example response
{
"model": "gpt-4o-mini",
"provider": "openai",
"usage": {
"prompt_tokens": 12,
"completion_tokens": 8,
"total_tokens": 20
},
"choices": [
{
"message": {
"role": "assistant",
"content": "Hello! How can I help you today?"
},
"finish_reason": "stop"
}
]
}
| Field | Type | Always present | In the example |
|---|---|---|---|
model | string | yes | gpt-4o-mini |
provider | string | yes | openai |
usage | object | yes | 3 fields: prompt_tokens, completion_tokens, total_tokens |
choices | array of objects | yes | 1 item in the example |
From an MCP client
catalog.call {
"slug": "llm",
"params": {
"model": "gpt-4o-mini",
"messages": [
{
"role": "user",
"content": "Say hello in one sentence."
}
],
"max_tokens": 64
}
}
The hosted connector at https://agent402.tools/mcp needs a payment for llm; the stdio package pays it from a wallet or from AGENT402_CREDITS_KEY. Local install: npx -y agent402-mcp.
Errors and behavior
modelandmessagesare required. An input the tool rejects returns an HTTP 4xx whose body carrieserror,tool,expected,requiredandexample, so the caller can correct it.- A paid call that ends in any status of 400 or above is not charged over x402, MPP or a prepaid credits key: settlement is cancelled when the tool fails. The exception is a Tempo push credential, a transfer the buyer sent before the call: it settles before the tool runs, so if the tool then fails the payment is recorded as a refund owed to the paying wallet.
- Wallet-only: this tool runs a model, so it has no proof-of-work tier. A prepaid card-credits key issued earlier (
Authorization: Bearer a402_...) also pays it. - Model-backed: the answer is generated by a model, so the same input can produce different wording.
- A
GETorHEADto /api/llm returns the same 402 quote, so the price can be read without a body. - An
Idempotency-Keyheader makes a retried paid call replay the first 200 instead of charging again (an answer larger than 1 MB is not replayed).
Paid call (JavaScript agent)
import { wrapFetchWithPayment } from "@x402/fetch";
import { x402Client } from "@x402/core/client";
import { registerExactEvmScheme } from "@x402/evm/exact/client";
import { privateKeyToAccount } from "viem/accounts";
const client = new x402Client();
client.setSpendControls?.(false); // keep your own spending ceiling in code
registerExactEvmScheme(client, { signer: privateKeyToAccount(KEY) });
const payFetch = wrapFetchWithPayment(fetch, client);
const res = await payFetch("https://agent402.tools/api/llm", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({
"model": "gpt-4o-mini",
"messages": [
{
"role": "user",
"content": "Say hello in one sentence."
}
],
"max_tokens": 64
}),
});
Related tools
LLM inference (Premium)
POST /api/llm-premiumLLM inference proxy (Premium tier) - o3 or o3-mini reasoning models. Supports vision (up to 2 image URLs) and structured…
LLM inference (Pro)
POST /api/llm-proLLM inference proxy (Pro tier) - GPT-4o or GPT-4.1. Supports vision (up to 2 image URLs) and structured output (response…
Text-to-speech (OpenAI-compatible)
POST /v1/audio/speechOpenAI-compatible text-to-speech over x402 - point any OpenAI SDK's audio.speech.create() at base_url https://agent402.t…
Chat completions (OpenAI-compatible)
POST /v1/chat/completionsOpenAI-compatible chat completions, base tier: point any OpenAI SDK at base_url https://agent402.tools/v1 and pay per ca…
Chat completions - auto tier (eval-ranked routing)
POST /v1/auto/chat/completionsOpenAI-compatible chat completions with server-side model choice: omit "model" (or send "auto") and the gateway routes t…
Grounded chat (web search, OpenAI-compatible)
POST /v1/grounded/chat/completionsOpenAI-compatible chat completions GROUNDED in a live web search on every call: the gateway runs an Exa search (up to 5 …