LLM inference (Premium)

$0.500 per call · USDC via x402 · POST /api/llm-premium

LLM inference proxy (Premium tier) - o3 or o3-mini reasoning models. Send POST /api/llm-premium with the required fields model and messages and pay $0.500 per call over x402 or MPP (there is no free tier). It returns a JSON object with model, provider, usage and choices.

Supports vision (up to 2 image URLs) and structured output (response_format: json_object or json_schema). No API key needed; pay per call via x402. Input capped at 32k chars, output at 2048 tokens.

Category: AI & compute · Tags: llm ai inference chat proxy openai o3 o3-mini

TRY IN PLAYGROUND →

Parameters

NameTypeRequiredDescription
modelstringyesModel ID - o3 or o3-mini
messagesarrayyesArray of {role, content} objects. content can be a string or array of {type:'text',text} and {type:'image_url',image_url:{url,detail}} blocks
max_tokensnumbernoMax output tokens (default 1024, cap 2048)
response_formatobjectnoOptional: {type:"json_object"} or {type:"json_schema",json_schema:{name,schema}}
temperaturenumbernoSampling temperature (0-2)
top_pnumbernoNucleus sampling (0-1)
stopstringnoStop sequence(s)

Example request

curl -i -X POST https://agent402.tools/api/llm-premium \
  -H "Content-Type: application/json" \
  -d '{"model":"o3-mini","messages":[{"role":"user","content":"Say hello in one sentence."}],"max_tokens":64}'

Without payment this returns HTTP 402 Payment Required with the exact price for llm-premium; any x402 v2 or MPP client pays it and retries.

Example response

{
  "model": "o3-mini",
  "provider": "openai",
  "usage": {
    "prompt_tokens": 12,
    "completion_tokens": 8,
    "total_tokens": 20
  },
  "choices": [
    {
      "message": {
        "role": "assistant",
        "content": "Hello! How can I help you today?"
      },
      "finish_reason": "stop"
    }
  ]
}
FieldTypeAlways presentIn the example
modelstringyeso3-mini
providerstringyesopenai
usageobjectyes3 fields: prompt_tokens, completion_tokens, total_tokens
choicesarray of objectsyes1 item in the example

From an MCP client

catalog.call {
  "slug": "llm-premium",
  "params": {
    "model": "o3-mini",
    "messages": [
      {
        "role": "user",
        "content": "Say hello in one sentence."
      }
    ],
    "max_tokens": 64
  }
}

The hosted connector at https://agent402.tools/mcp needs a payment for llm-premium; the stdio package pays it from a wallet or from AGENT402_CREDITS_KEY. Local install: npx -y agent402-mcp.

Errors and behavior

Paid call (JavaScript agent)

import { wrapFetchWithPayment } from "@x402/fetch";
import { x402Client } from "@x402/core/client";
import { registerExactEvmScheme } from "@x402/evm/exact/client";
import { privateKeyToAccount } from "viem/accounts";

const client = new x402Client();
client.setSpendControls?.(false); // keep your own spending ceiling in code
registerExactEvmScheme(client, { signer: privateKeyToAccount(KEY) });
const payFetch = wrapFetchWithPayment(fetch, client);

const res = await payFetch("https://agent402.tools/api/llm-premium", {
  method: "POST",
  headers: { "Content-Type": "application/json" },
  body: JSON.stringify({
    "model": "o3-mini",
    "messages": [
      {
        "role": "user",
        "content": "Say hello in one sentence."
      }
    ],
    "max_tokens": 64
  }),
});

Related tools

LLM inference

$0.010 · POST /api/llm

LLM inference proxy - send an OpenAI-format chat/completions request and get a response from GPT-4o-mini. Supports visio…

LLM inference (Pro)

$0.100 · POST /api/llm-pro

LLM inference proxy (Pro tier) - GPT-4o or GPT-4.1. Supports vision (up to 2 image URLs) and structured output (response…

Text-to-speech (OpenAI-compatible)

$0.060 · POST /v1/audio/speech

OpenAI-compatible text-to-speech over x402 - point any OpenAI SDK's audio.speech.create() at base_url https://agent402.t…

Chat completions (OpenAI-compatible)

$0.02 · POST /v1/chat/completions

OpenAI-compatible chat completions, base tier: point any OpenAI SDK at base_url https://agent402.tools/v1 and pay per ca…

Chat completions - auto tier (eval-ranked routing)

$0.01 · POST /v1/auto/chat/completions

OpenAI-compatible chat completions with server-side model choice: omit "model" (or send "auto") and the gateway routes t…

Grounded chat (web search, OpenAI-compatible)

$0.03 · POST /v1/grounded/chat/completions

OpenAI-compatible chat completions GROUNDED in a live web search on every call: the gateway runs an Exa search (up to 5 …