Pay for what you use
One plan. You top up a balance, and every request draws from it at the model's published per-token price. No subscription, no seats, no minimum commitment.
What's included
- Every model — one endpoint, one key, one balance across the whole catalog
- OpenAI-compatible API — point an existing SDK at us and change the model name
- Streaming — server-sent events with token usage and cost
- Tool calling — up to 128 function tools per request
- Image, audio & video input — on the models that read them
- Presets — save a model, prompt and parameters under one name
- Per-request logs — tokens, price and latency for every call
- Activity charts — spend by day, model and key
- Unlimited API keys — each with its own spend cap and expiry
- Per-key rate limits — so one integration can't starve another
- Workspace budgets — daily, weekly, monthly or lifetime caps
- Guardrails — spend limits plus a list of models a key may use
How the fee works
The 4% is charged once, when you buy credits, and covers card processing. It is shown before you pay. Requests themselves carry no fee — you are billed the model's token price and nothing else.
| Buying $25 of credits | Amount |
|---|---|
| Credits added to your balance | $25.00 |
| Platform fee (4%) | $1.00 |
| Charged to your card | $26.00 |
Top up between $5 and $5,000 at a time. Credits never expire, and an unused balance stays yours.
Model rates
Prices are USD per million tokens, billed separately for what you send and what the model writes back. They span more than a hundredfold across the catalog, so the model you pick matters far more than any plan would.
| Model | Input | Output |
|---|---|---|
| Nova Microamazon/nova-micro-v1:0 | $0.04/M | $0.161/M |
| Qwen3 32B (dense)qwen/qwen3-32b-v1:0 | $0.172/M | $0.69/M |
| Qwen3 Coder Nextqwen/qwen3-coder-next | $0.575/M | $1.38/M |
| Mistral Large (24.02)mistral/mistral-large-2402-v1:0 | $4.6/M | $13.8/M |
Not available yet
Things people expect from a platform like this that Sothe genuinely does not do today. They are listed here so you can tell before you build on them.
- Structured outputs — response_format is accepted but not enforced
- Automatic routing and fallbacks between models
- Embeddings, image generation, speech and transcription endpoints
- Bring-your-own provider keys
- A management API, and automatic top-ups
Start with $5
Create an account, add credits and send your first request in a few minutes.