One API for chat, image and speech modelsOpenAI-compatibleNo per-request markupCredit never expiresRead the docs

Quick start

Create a key in your account, then use an OpenAI-compatible client for chat, images and speech. Copy the API base URL from Account → API keys; it includes /v1. Requests charge the balance of the account that owns the key. All keys on that account share its balance, with separate key spending caps.

import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.RPROUTER_API_KEY,
  baseURL: process.env.RPROUTER_BASE_URL // ends with /v1
});

const { data: models } = await client.models.list();
if (!models.length) throw new Error("No routes currently available");
const response = await client.chat.completions.create({
  model: models.find(m => m.architecture.output_modalities.includes("text")).id,
  messages: [{ role: "user", content: "Tell me a short story." }],
  max_tokens: 500, stream: true
});
for await (const chunk of response) {
  process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}

Authentication

Send Authorization: Bearer rp_…. API keys are shown once, stored as hashes, and can be revoked immediately. Keep keys on your server. The account website uses separate email authentication.

Each key has a lifetime spending cap, and accounts have a daily UTC cap. Pending reservations count against both. Provider keys cannot be connected in this release.

Keys carry permissions (scopes). Keys you create yourself have all of them. Keys issued to a developer app carry only what the user approved; a missing permission returns 403 insufficient_scope.

ScopeAllows
inferenceMake chat, image and speech requests billed to your RProuter credit
balanceSee your available and reserved credit
subscriptionSee whether you have an active RProuter membership
profileSee your RProuter account ID and email address
billingSee when you add credit and how much (not your balance or usage)
usageSee requests made with this key

Developer apps

Let your users bring their own RProuter account. Your site sends the user to RProuter; they sign in, choose a spending limit and approve; RProuter sends them back with a one-time code; your server exchanges the code for a key. From then on, requests you make with that key are billed to the user’s credit and listed in their usage under your app’s name. This is the standard OAuth 2.0 authorization-code flow, with PKCE for apps that can’t keep a secret.

  1. Register the app in Account → Apps with its exact redirect URIs. Keep the client secret on your server.
  2. Send the user to /oauth/authorize with client_id, redirect_uri, scope, a random state and, optionally, limit_usd to suggest a spending limit.
  3. RProuter redirects to your redirect_uri with code and your state, or with error=access_denied. Check state first.
  4. Within 10 minutes, POST the code to /oauth/token. The response’s access_token is an RProuter API key (rp_…); store it encrypted, per user.
  5. Call /v1 with it. Check membership with GET /v1/me/subscription; the user can disconnect at any time, after which the key returns 401.
// 1. Redirect the user (server-side route)
const state = crypto.randomUUID();           // store in the user's session
const url = new URL("https://rprouter.example/oauth/authorize");
url.search = new URLSearchParams({
  client_id: process.env.RPROUTER_CLIENT_ID,
  redirect_uri: "https://your-app.example/rprouter/callback",
  scope: "inference balance subscription",
  state,
  limit_usd: "20",
}).toString();
return Response.redirect(url, 302);

// 2. Callback: verify state, then exchange the code
const res = await fetch("https://rprouter.example/oauth/token", {
  method: "POST",
  headers: { "content-type": "application/x-www-form-urlencoded" },
  body: new URLSearchParams({
    grant_type: "authorization_code",
    code: callbackUrl.searchParams.get("code"),
    redirect_uri: "https://your-app.example/rprouter/callback",
    client_id: process.env.RPROUTER_CLIENT_ID,
    client_secret: process.env.RPROUTER_CLIENT_SECRET,
  }),
});
const { access_token, user_id, scope, spending_limit_usd } = await res.json();

// 3. Use the key on the user's behalf
const sub = await fetch("https://rprouter.example/v1/me/subscription", {
  headers: { Authorization: `Bearer ${access_token}` },
}).then((r) => r.json());
if (sub.active) unlockPremiumFeatures();   // plan: "premium"

Public clients (single-page and mobile apps) omit the secret and use PKCE: send code_challenge=BASE64URL(SHA256(verifier)) and code_challenge_method=S256 to /oauth/authorize, then code_verifier to /oauth/token. The token endpoint accepts cross-origin requests.

GET /v1/me/subscription
Authorization: Bearer rp_…

{
  "object": "subscription",
  "plan": "premium",          // null when not a member
  "active": true,
  "status": "active",
  "paid_until": "2026-10-28T00:00:00.000Z",
  "current_period_end": "2026-10-28T00:00:00.000Z",
  "cancel_at_period_end": false,
  "topup_fee_bps": 500
}

GET /v1/me returns the stable RProuter user ID for the key (use it to link accounts on your side), the app, the granted scopes and the key’s spending limit; the email address is included only with the profile scope. Connecting again replaces the user’s previous key for your app. Revoke a key yourself with POST /oauth/revoke and token=rp_…. Codes are single use; redirect URIs must match a registered URI exactly.

Top-ups from your app

To let users add credit without leaving your flow, link them to /topup?client_id=…&return_to=…. They sign in if needed, pay by card (minimum $10), and once the credit has landed RProuter sends them back to return_to, which must be on the same site as one of your redirect URIs.

With the billing scope, GET /v1/me/funding?days=30 reports the credit the user added in that window (net of refunds) and their latest top-up. Use it for your own rules, such as unlocking features after a recent top-up; RProuter only reports the facts. Check it when the user returns from /topup and when they sign in to your app.

GET /v1/me/funding?days=30
Authorization: Bearer rp_…

{
  "object": "funding",
  "currency": "USD",
  "window_days": 30,
  "topped_up_usd": 25,
  "topups": 1,
  "last_topup_at": "2026-09-28T10:00:00.000Z",
  "last_topup_usd": 25
}

API endpoints

MethodPathResult
GET/v1/modelsAvailable model IDs and capabilities
POST/v1/chat/completionsText completion or SSE stream
POST/v1/images/generationsSynchronous image results
POST/v1/images/editsImage editing (multipart or owned uploads)
POST/v1/images/estimateConfiguration-based maximum charge
POST/v1/jobsDurable asynchronous image job
GETDELETE/v1/jobs/:idRetrieve results / request cancellation
POST/v1/uploadsAccount-owned image, mask or voice sample
GET/v1/media/:tokenSigned link streaming an image from the provider
POST/v1/audio/speechAudio bytes, streaming by default
GETPOST/v1/voicesList public/private voices or create a private clone
GETDELETE/v1/voices/:idInspect / delete an owned voice
POST/v1/voices/searchDiscover public Fish voices
GET/v1/balanceAvailable and reserved usage credit
GET/v1/meThe key’s user ID, app, scopes and limit
GET/v1/me/subscriptionWhether the user has an active membership
GET/v1/me/fundingCredit the user added, net of refunds
GET/v1/usageRequests made with this key
GET/oauth/authorizeStart connecting a user’s account
POST/oauth/tokenExchange a code for the user’s key
POST/oauth/revokeRevoke an app key
GET/topupSend a user to add credit and back

Full request and response schemas are in the API reference (OpenAPI JSON). Video and transcription are not supported.

Images and speech

Use the same key and wallet for every modality. Only models listed by GET /v1/models are available. Set RPROUTER_BASE_URL to your API base URL ending in /v1.

Generate an image

OpenAI-compatible fields. RProuter’s max_cost and provider_options are extensions. Send a stable Idempotency-Key when retrying.

curl "$RPROUTER_BASE_URL/images/generations" \
  -H "Authorization: Bearer $RPROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: $REQUEST_ID" \
  -d '{"model":"tensorart/wai-illustrious-v16","prompt":"A quiet mountain lake","size":"1024x1024","n":1,"response_format":"url"}'

Estimate and submit an asynchronous job

RProuter extension. The estimate does not reserve funds. max_cost prevents a later quote from exceeding your limit. Poll the original job, including after a timeout; pending means reconciliation, not permission to resubmit.

# Save the exact settings you intend to submit.
printf '%s' '{"model":"tensorart/wai-illustrious-v16","prompt":"A quiet mountain lake","size":"1024x1024"}' > image.json
curl "$RPROUTER_BASE_URL/images/estimate" \
  -H "Authorization: Bearer $RPROUTER_API_KEY" -H "Content-Type: application/json" -d @image.json
# Add max_cost from the quote to image.json, then submit once.
curl "$RPROUTER_BASE_URL/jobs" -H "Authorization: Bearer $RPROUTER_API_KEY" \
  -H "Idempotency-Key: $REQUEST_ID" -H "Content-Type: application/json" -d @image.json
curl "$RPROUTER_BASE_URL/jobs/$JOB_ID" -H "Authorization: Bearer $RPROUTER_API_KEY"
# Request cancellation (dispatched work can still be billable):
curl -X DELETE "$RPROUTER_BASE_URL/jobs/$JOB_ID" -H "Authorization: Bearer $RPROUTER_API_KEY"

Edit an image

Multipart image and mask fields are compatible with image-edit clients. One source image is supported: a PNG/JPEG file, or an upload ID from /uploads. TensorArt routes support masks; fal routes take the reference image only. provider_options is JSON in multipart requests. Masks use TensorArt’s white-to-repaint convention.

curl "$RPROUTER_BASE_URL/images/edits" \
  -H "Authorization: Bearer $RPROUTER_API_KEY" -H "Idempotency-Key: $REQUEST_ID" \
  -F model=tensorart/wai-illustrious-v16 -F 'prompt=Change the sky to sunset' \
  -F image=@source.png -F mask=@mask.png \
  -F 'provider_options={"mode":"inpaint","strength":0.5}'

Stream speech

OpenAI-compatible speech fields. voice is a stable RProuter voice UUID from /voices, not a Fish provider ID. Expressive tags are forwarded unchanged. Successful usage is billed by submitted UTF-8 bytes. Audio is not retained.

curl "$RPROUTER_BASE_URL/voices?scope=public&search=" -H "Authorization: Bearer $RPROUTER_API_KEY"
curl "$RPROUTER_BASE_URL/audio/speech" \
  -H "Authorization: Bearer $RPROUTER_API_KEY" -H "Content-Type: application/json" \
  -d "{\"model\":\"fish/s2-pro\",\"voice\":\"$VOICE_ID\",\"input\":\"Hello there.\",\"response_format\":\"mp3\"}" \
  --output speech.mp3

Create and delete a private voice

RProuter extension. Use your own voice or obtain the speaker’s permission. Upload WAV, MP3 or FLAC; wait until ready before synthesizing. Deletion is asynchronous and retried upstream. The sample is deleted once the voice is ready.

curl "$RPROUTER_BASE_URL/uploads" -H "Authorization: Bearer $RPROUTER_API_KEY" -F file=@sample.wav
# Use the returned upload id:
curl "$RPROUTER_BASE_URL/voices" -H "Authorization: Bearer $RPROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d "{\"name\":\"My voice\",\"uploadId\":\"$UPLOAD_ID\",\"consent\":true,\"consentVersion\":\"2026-09-09\",\"rights\":\"my_voice\"}"
curl "$RPROUTER_BASE_URL/voices/$VOICE_ID" -H "Authorization: Bearer $RPROUTER_API_KEY"
curl -X DELETE "$RPROUTER_BASE_URL/voices/$VOICE_ID" -H "Authorization: Bearer $RPROUTER_API_KEY"

Image options and retention

Sizes: square 1024×1024, portrait 896×1336 or 1024×1536, landscape 1336×896 or 1536×1024, where supported by the route. Standard quality uses 20 steps; high uses 30. provider_options accepts a public checkpoint for tensorart/custom, negative_prompt, loras: [{id, strength}], detail_refinement, and editing strength from 0 to 1. Reference mode uses image-to-image conditioning. Detail refinement is not available with masks.

Synchronous image calls wait up to 120 seconds. A 504 image_wait_timeout includes job_id and request_id; retrieve that job instead of generating again. Generated images are not stored by RProuter. Result URLs are signed /v1/media/… links that stream the image from the provider; they work without an API key (for example in an <img> tag) and expire after 24 hours, or sooner if the provider removes the file. Save images you want to keep. Unused uploads expire after one hour; accepted jobs keep their inputs until processing finishes.

The accepted image quote is the maximum charge. Unexpected supplier overages are absorbed and flagged for review. Speech uses the pinned price per million UTF-8 bytes; cancelled or incomplete speech stays pending until reconciled. Standard and Premium fees apply only when purchasing credit.

Streaming and routing

Text uses Server-Sent Events. x-request-id identifies the billable request. Disconnecting cancels upstream work where supported, but it does not guarantee zero supplier usage.

Default routing favors first-answer latency among compatible routes within your price ceiling. Use provider.sort with latency, throughput, or price, and optional provider.max_price input/output unit ceilings. Unknown measurements do not become zero latency.

Send an Idempotency-Key to avoid duplicate dispatch. A repeated text request returns 409 and the original request ID; it does not replay stored conversation content. Unsupported settings return 400, inactive keys 401, insufficient credit 402, ownership failures 404, capacity exhaustion 429 with Retry-After, and unavailable routes 503.

Retries are limited to one confirmed unbilled transient refusal before output. Models and private voices are never silently substituted. Missing provider usage remains reserved for reconciliation.

Billing and limits

Before dispatch, RProuter reserves a conservative upper bound. Text estimates use a UTF-8 byte bound plus message framing and the requested output cap. Large requests may reserve more than their final charge. Final settlement uses provider-reported tokens, including cached and billed reasoning tokens.

Missing usage stays pending; unused reservations are released. Cancellations and partial results are reconciled against actual supplier usage.

Purchased credit never expires. Top-ups carry a 12% platform fee, or 5% with Premium. There is no additional platform markup on requests. Premium is a $15 monthly membership with no included usage. Refunds revoke corresponding credit; refunded credit already used becomes a balance adjustment that must be cleared before further spending. Taxes do not create usage credit.

Premium priority

Standard requests run immediately when capacity is free. When both queues are waiting, up to three Premium admissions precede the next Standard admission. Running streams are never interrupted. A request waits up to five seconds on Standard or ten seconds on Premium, then receives a 429 response if capacity is still unavailable. Per-customer concurrency limits apply to both plans.

Premium priority requires a paid, unexpired membership. Canceling renewal preserves priority through the paid period. A full membership refund removes the refunded period’s priority. Status updates alone cannot grant a new unpaid period.

Priority only affects RProuter’s admission queue. It cannot change a supplier’s internal scheduling, rate limits, or token-generation speed.

Performance

First-answer latency starts with the incoming request and ends at the first visible answer content. Reasoning-only chunks do not end this timer. Generation speed uses output tokens and the measured generation interval; completion time covers the full request.

Summaries group provider route, region, workload, and settings. Medians require 30 comparable samples; p95 requires 100. Summaries refresh every five minutes and display sample count, window, method, and freshness. Supplier claims must be labeled separately.

External benchmark figures are separate from RProuter route observations. Customer conversations are not benchmark material without consent.

Benchmark sources

The September 15, 2026 snapshot matches 22 catalog models to EQ-Bench Creative Writing v3 and 11 to EQ-Bench 4. Writing measures story generation; EQ measures social and emotional intelligence in multi-turn persona conversations. Neither is a complete roleplay evaluation. The EQ4 export was generated July 26, 2026; a retrieval date is not a new test date.

Six exact model versions have speed and first-token figures from linked Artificial Analysis model summaries. These are dated external observations at the stated provider and reasoning setting, not live RProuter performance. Prompt lengths and sample windows were not specified in the cited summaries. Different setups limit direct latency comparisons.

We map exact versions and preserve provisional labels. Confidence intervals and sample counts appear when published. Missing values stay unavailable; they do not become zero.

Source records include URLs, tested model IDs, settings, and retrieval dates. These snapshots are refreshed separately from the five-minute summaries of RProuter traffic. No external benchmark is presented as our own test.