Quick start
Create a key in your account, then use an OpenAI-compatible client for chat, images and speech. Copy the API base URL from Account → API keys; it includes /v1. Requests charge the balance of the account that owns the key. All keys on that account share its balance, with separate key spending caps.
import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.RPROUTER_API_KEY,
baseURL: process.env.RPROUTER_BASE_URL // ends with /v1
});
const { data: models } = await client.models.list();
if (!models.length) throw new Error("No routes currently available");
const response = await client.chat.completions.create({
model: models.find(m => m.architecture.output_modalities.includes("text")).id,
messages: [{ role: "user", content: "Tell me a short story." }],
max_tokens: 500, stream: true
});
for await (const chunk of response) {
process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}Authentication
Send Authorization: Bearer rp_…. API keys are shown once, stored as hashes, and can be revoked immediately. Keep keys on your server. The account website uses separate email authentication.
Each key has a lifetime spending cap, and accounts have a daily UTC cap. Pending reservations count against both. Provider keys cannot be connected in this release.
Keys carry permissions (scopes). Keys you create yourself have all of them. Keys issued to a developer app carry only what the user approved; a missing permission returns 403 insufficient_scope.
| Scope | Allows |
|---|---|
inference | Make chat, image and speech requests billed to your RProuter credit |
balance | See your available and reserved credit |
subscription | See whether you have an active RProuter membership |
profile | See your RProuter account ID and email address |
billing | See when you add credit and how much (not your balance or usage) |
usage | See requests made with this key |
Developer apps
Let your users bring their own RProuter account. Your site sends the user to RProuter; they sign in, choose a spending limit and approve; RProuter sends them back with a one-time code; your server exchanges the code for a key. From then on, requests you make with that key are billed to the user’s credit and listed in their usage under your app’s name. This is the standard OAuth 2.0 authorization-code flow, with PKCE for apps that can’t keep a secret.
- Register the app in Account → Apps with its exact redirect URIs. Keep the client secret on your server.
- Send the user to
/oauth/authorizewithclient_id,redirect_uri,scope, a randomstateand, optionally,limit_usdto suggest a spending limit. - RProuter redirects to your
redirect_uriwithcodeand yourstate, or witherror=access_denied. Checkstatefirst. - Within 10 minutes, POST the code to
/oauth/token. The response’saccess_tokenis an RProuter API key (rp_…); store it encrypted, per user. - Call
/v1with it. Check membership withGET /v1/me/subscription; the user can disconnect at any time, after which the key returns 401.
// 1. Redirect the user (server-side route)
const state = crypto.randomUUID(); // store in the user's session
const url = new URL("https://rprouter.example/oauth/authorize");
url.search = new URLSearchParams({
client_id: process.env.RPROUTER_CLIENT_ID,
redirect_uri: "https://your-app.example/rprouter/callback",
scope: "inference balance subscription",
state,
limit_usd: "20",
}).toString();
return Response.redirect(url, 302);
// 2. Callback: verify state, then exchange the code
const res = await fetch("https://rprouter.example/oauth/token", {
method: "POST",
headers: { "content-type": "application/x-www-form-urlencoded" },
body: new URLSearchParams({
grant_type: "authorization_code",
code: callbackUrl.searchParams.get("code"),
redirect_uri: "https://your-app.example/rprouter/callback",
client_id: process.env.RPROUTER_CLIENT_ID,
client_secret: process.env.RPROUTER_CLIENT_SECRET,
}),
});
const { access_token, user_id, scope, spending_limit_usd } = await res.json();
// 3. Use the key on the user's behalf
const sub = await fetch("https://rprouter.example/v1/me/subscription", {
headers: { Authorization: `Bearer ${access_token}` },
}).then((r) => r.json());
if (sub.active) unlockPremiumFeatures(); // plan: "premium"Public clients (single-page and mobile apps) omit the secret and use PKCE: send code_challenge=BASE64URL(SHA256(verifier)) and code_challenge_method=S256 to /oauth/authorize, then code_verifier to /oauth/token. The token endpoint accepts cross-origin requests.
GET /v1/me/subscription
Authorization: Bearer rp_…
{
"object": "subscription",
"plan": "premium", // null when not a member
"active": true,
"status": "active",
"paid_until": "2026-10-28T00:00:00.000Z",
"current_period_end": "2026-10-28T00:00:00.000Z",
"cancel_at_period_end": false,
"topup_fee_bps": 500
}GET /v1/me returns the stable RProuter user ID for the key (use it to link accounts on your side), the app, the granted scopes and the key’s spending limit; the email address is included only with the profile scope. Connecting again replaces the user’s previous key for your app. Revoke a key yourself with POST /oauth/revoke and token=rp_…. Codes are single use; redirect URIs must match a registered URI exactly.
Top-ups from your app
To let users add credit without leaving your flow, link them to /topup?client_id=…&return_to=…. They sign in if needed, pay by card (minimum $10), and once the credit has landed RProuter sends them back to return_to, which must be on the same site as one of your redirect URIs.
With the billing scope, GET /v1/me/funding?days=30 reports the credit the user added in that window (net of refunds) and their latest top-up. Use it for your own rules, such as unlocking features after a recent top-up; RProuter only reports the facts. Check it when the user returns from /topup and when they sign in to your app.
GET /v1/me/funding?days=30
Authorization: Bearer rp_…
{
"object": "funding",
"currency": "USD",
"window_days": 30,
"topped_up_usd": 25,
"topups": 1,
"last_topup_at": "2026-09-28T10:00:00.000Z",
"last_topup_usd": 25
}API endpoints
| Method | Path | Result |
|---|---|---|
| GET | /v1/models | Available model IDs and capabilities |
| POST | /v1/chat/completions | Text completion or SSE stream |
| POST | /v1/images/generations | Synchronous image results |
| POST | /v1/images/edits | Image editing (multipart or owned uploads) |
| POST | /v1/images/estimate | Configuration-based maximum charge |
| POST | /v1/jobs | Durable asynchronous image job |
| GETDELETE | /v1/jobs/:id | Retrieve results / request cancellation |
| POST | /v1/uploads | Account-owned image, mask or voice sample |
| GET | /v1/media/:token | Signed link streaming an image from the provider |
| POST | /v1/audio/speech | Audio bytes, streaming by default |
| GETPOST | /v1/voices | List public/private voices or create a private clone |
| GETDELETE | /v1/voices/:id | Inspect / delete an owned voice |
| POST | /v1/voices/search | Discover public Fish voices |
| GET | /v1/balance | Available and reserved usage credit |
| GET | /v1/me | The key’s user ID, app, scopes and limit |
| GET | /v1/me/subscription | Whether the user has an active membership |
| GET | /v1/me/funding | Credit the user added, net of refunds |
| GET | /v1/usage | Requests made with this key |
| GET | /oauth/authorize | Start connecting a user’s account |
| POST | /oauth/token | Exchange a code for the user’s key |
| POST | /oauth/revoke | Revoke an app key |
| GET | /topup | Send a user to add credit and back |
Full request and response schemas are in the API reference (OpenAPI JSON). Video and transcription are not supported.
Images and speech
Use the same key and wallet for every modality. Only models listed by GET /v1/models are available. Set RPROUTER_BASE_URL to your API base URL ending in /v1.
Generate an image
OpenAI-compatible fields. RProuter’s max_cost and provider_options are extensions. Send a stable Idempotency-Key when retrying.
curl "$RPROUTER_BASE_URL/images/generations" \
-H "Authorization: Bearer $RPROUTER_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: $REQUEST_ID" \
-d '{"model":"tensorart/wai-illustrious-v16","prompt":"A quiet mountain lake","size":"1024x1024","n":1,"response_format":"url"}'Estimate and submit an asynchronous job
RProuter extension. The estimate does not reserve funds. max_cost prevents a later quote from exceeding your limit. Poll the original job, including after a timeout; pending means reconciliation, not permission to resubmit.
# Save the exact settings you intend to submit.
printf '%s' '{"model":"tensorart/wai-illustrious-v16","prompt":"A quiet mountain lake","size":"1024x1024"}' > image.json
curl "$RPROUTER_BASE_URL/images/estimate" \
-H "Authorization: Bearer $RPROUTER_API_KEY" -H "Content-Type: application/json" -d @image.json
# Add max_cost from the quote to image.json, then submit once.
curl "$RPROUTER_BASE_URL/jobs" -H "Authorization: Bearer $RPROUTER_API_KEY" \
-H "Idempotency-Key: $REQUEST_ID" -H "Content-Type: application/json" -d @image.json
curl "$RPROUTER_BASE_URL/jobs/$JOB_ID" -H "Authorization: Bearer $RPROUTER_API_KEY"
# Request cancellation (dispatched work can still be billable):
curl -X DELETE "$RPROUTER_BASE_URL/jobs/$JOB_ID" -H "Authorization: Bearer $RPROUTER_API_KEY"Edit an image
Multipart image and mask fields are compatible with image-edit clients. One source image is supported: a PNG/JPEG file, or an upload ID from /uploads. TensorArt routes support masks; fal routes take the reference image only. provider_options is JSON in multipart requests. Masks use TensorArt’s white-to-repaint convention.
curl "$RPROUTER_BASE_URL/images/edits" \
-H "Authorization: Bearer $RPROUTER_API_KEY" -H "Idempotency-Key: $REQUEST_ID" \
-F model=tensorart/wai-illustrious-v16 -F 'prompt=Change the sky to sunset' \
-F image=@source.png -F mask=@mask.png \
-F 'provider_options={"mode":"inpaint","strength":0.5}'Stream speech
OpenAI-compatible speech fields. voice is a stable RProuter voice UUID from /voices, not a Fish provider ID. Expressive tags are forwarded unchanged. Successful usage is billed by submitted UTF-8 bytes. Audio is not retained.
curl "$RPROUTER_BASE_URL/voices?scope=public&search=" -H "Authorization: Bearer $RPROUTER_API_KEY"
curl "$RPROUTER_BASE_URL/audio/speech" \
-H "Authorization: Bearer $RPROUTER_API_KEY" -H "Content-Type: application/json" \
-d "{\"model\":\"fish/s2-pro\",\"voice\":\"$VOICE_ID\",\"input\":\"Hello there.\",\"response_format\":\"mp3\"}" \
--output speech.mp3Create and delete a private voice
RProuter extension. Use your own voice or obtain the speaker’s permission. Upload WAV, MP3 or FLAC; wait until ready before synthesizing. Deletion is asynchronous and retried upstream. The sample is deleted once the voice is ready.
curl "$RPROUTER_BASE_URL/uploads" -H "Authorization: Bearer $RPROUTER_API_KEY" -F file=@sample.wav
# Use the returned upload id:
curl "$RPROUTER_BASE_URL/voices" -H "Authorization: Bearer $RPROUTER_API_KEY" \
-H "Content-Type: application/json" \
-d "{\"name\":\"My voice\",\"uploadId\":\"$UPLOAD_ID\",\"consent\":true,\"consentVersion\":\"2026-09-09\",\"rights\":\"my_voice\"}"
curl "$RPROUTER_BASE_URL/voices/$VOICE_ID" -H "Authorization: Bearer $RPROUTER_API_KEY"
curl -X DELETE "$RPROUTER_BASE_URL/voices/$VOICE_ID" -H "Authorization: Bearer $RPROUTER_API_KEY"Image options and retention
Sizes: square 1024×1024, portrait 896×1336 or 1024×1536, landscape 1336×896 or 1536×1024, where supported by the route. Standard quality uses 20 steps; high uses 30. provider_options accepts a public checkpoint for tensorart/custom, negative_prompt, loras: [{id, strength}], detail_refinement, and editing strength from 0 to 1. Reference mode uses image-to-image conditioning. Detail refinement is not available with masks.
Synchronous image calls wait up to 120 seconds. A 504 image_wait_timeout includes job_id and request_id; retrieve that job instead of generating again. Generated images are not stored by RProuter. Result URLs are signed /v1/media/… links that stream the image from the provider; they work without an API key (for example in an <img> tag) and expire after 24 hours, or sooner if the provider removes the file. Save images you want to keep. Unused uploads expire after one hour; accepted jobs keep their inputs until processing finishes.
The accepted image quote is the maximum charge. Unexpected supplier overages are absorbed and flagged for review. Speech uses the pinned price per million UTF-8 bytes; cancelled or incomplete speech stays pending until reconciled. Standard and Premium fees apply only when purchasing credit.
Streaming and routing
Text uses Server-Sent Events. x-request-id identifies the billable request. Disconnecting cancels upstream work where supported, but it does not guarantee zero supplier usage.
Default routing favors first-answer latency among compatible routes within your price ceiling. Use provider.sort with latency, throughput, or price, and optional provider.max_price input/output unit ceilings. Unknown measurements do not become zero latency.
Send an Idempotency-Key to avoid duplicate dispatch. A repeated text request returns 409 and the original request ID; it does not replay stored conversation content. Unsupported settings return 400, inactive keys 401, insufficient credit 402, ownership failures 404, capacity exhaustion 429 with Retry-After, and unavailable routes 503.
Retries are limited to one confirmed unbilled transient refusal before output. Models and private voices are never silently substituted. Missing provider usage remains reserved for reconciliation.
Billing and limits
Before dispatch, RProuter reserves a conservative upper bound. Text estimates use a UTF-8 byte bound plus message framing and the requested output cap. Large requests may reserve more than their final charge. Final settlement uses provider-reported tokens, including cached and billed reasoning tokens.
Missing usage stays pending; unused reservations are released. Cancellations and partial results are reconciled against actual supplier usage.
Purchased credit never expires. Top-ups carry a 12% platform fee, or 5% with Premium. There is no additional platform markup on requests. Premium is a $15 monthly membership with no included usage. Refunds revoke corresponding credit; refunded credit already used becomes a balance adjustment that must be cleared before further spending. Taxes do not create usage credit.
Premium priority
Standard requests run immediately when capacity is free. When both queues are waiting, up to three Premium admissions precede the next Standard admission. Running streams are never interrupted. A request waits up to five seconds on Standard or ten seconds on Premium, then receives a 429 response if capacity is still unavailable. Per-customer concurrency limits apply to both plans.
Premium priority requires a paid, unexpired membership. Canceling renewal preserves priority through the paid period. A full membership refund removes the refunded period’s priority. Status updates alone cannot grant a new unpaid period.
Priority only affects RProuter’s admission queue. It cannot change a supplier’s internal scheduling, rate limits, or token-generation speed.
Performance
First-answer latency starts with the incoming request and ends at the first visible answer content. Reasoning-only chunks do not end this timer. Generation speed uses output tokens and the measured generation interval; completion time covers the full request.
Summaries group provider route, region, workload, and settings. Medians require 30 comparable samples; p95 requires 100. Summaries refresh every five minutes and display sample count, window, method, and freshness. Supplier claims must be labeled separately.
External benchmark figures are separate from RProuter route observations. Customer conversations are not benchmark material without consent.
Benchmark sources
The September 15, 2026 snapshot matches 22 catalog models to EQ-Bench Creative Writing v3 and 11 to EQ-Bench 4. Writing measures story generation; EQ measures social and emotional intelligence in multi-turn persona conversations. Neither is a complete roleplay evaluation. The EQ4 export was generated July 26, 2026; a retrieval date is not a new test date.
Six exact model versions have speed and first-token figures from linked Artificial Analysis model summaries. These are dated external observations at the stated provider and reasoning setting, not live RProuter performance. Prompt lengths and sample windows were not specified in the cited summaries. Different setups limit direct latency comparisons.
We map exact versions and preserve provisional labels. Confidence intervals and sample counts appear when published. Missing values stay unavailable; they do not become zero.
Source records include URLs, tested model IDs, settings, and retrieval dates. These snapshots are refreshed separately from the five-minute summaries of RProuter traffic. No external benchmark is presented as our own test.