Pricing

All prices per million tokens. No minimums, no seats, no monthly platform fee. You pay for tokens.

ModelContextInput (list)Input (yours)Output (list)Output (yours)Discount
Claude Opus 4.1Anthropic200K$15.00$7.50$75.00$37.5050%
Claude Sonnet 4.5Anthropic200K$3.00$1.50$15.00$7.5050%
Claude Haiku 4.5Anthropic200K$1.00$0.600$5.00$3.0040%
GPT-5OpenAI400K$1.25$0.688$10.00$5.5045%
GPT-5 miniOpenAI400K$0.250$0.163$2.00$1.3035%
GPT-4.1OpenAI1M$2.00$1.10$8.00$4.4045%
Claude Haiku 3.5Anthropic200K$0.800$0.520$4.00$2.6035%
o4-miniOpenAI200K$1.10$0.660$4.40$2.6440%

All prices USD per 1M tokens.

Cached input, batch, and long context

These are the line items that quietly decide your bill. Where our discount does not apply, the table says so.

ModelCached input (list)Cached input (yours)Batch APILong-context tier
Claude Opus 4.1$1.50$0.750Passed through at listSame discount
Claude Sonnet 4.5$0.300$0.150Passed through at listSame discount
Claude Haiku 4.5$0.100$0.060Passed through at listSame discount
GPT-5$0.125$0.069Passed through at listSame discount
GPT-5 mini$0.025$0.016Passed through at listSame discount
GPT-4.1$0.500$0.275Passed through at listSame discount
Claude Haiku 3.5$0.080$0.052Passed through at listSame discount
o4-mini$0.275$0.165Passed through at listSame discount

Why batch isn’t discounted. Batch jobs already carry the provider’s own 50% reduction, and our contracted rate is negotiated against standard throughput. We pass batch through at the provider’s list price with no markup and no discount. If your workload is mostly batch, you will not save money with us, and we would rather tell you that here than after you migrate.

Throughput

Rate limits

Tiers are per-key and unlock automatically as you add credit. There is no application and no sales call to move up until the Dedicated tier.

TierRequests / minTokens / minConcurrentHow to upgrade
Free2050K2Add $20 in credit
Tier 15001M25Add $200 in credit
Tier 22,0005M100Add $1,000 in credit
Tier 310,00025M500Talk to us
DedicatedCustomCustomCustomReserved capacity contract

Shared or dedicated? Free through Tier 3 draw from a shared pool of capacity we have contracted in advance. Your tier limit is a guaranteed ceiling on what you can request, not a guaranteed reservation of upstream throughput — during an industry-wide spike, a shared-pool request may queue. The Dedicated tier is reserved capacity that is contractually yours and is not pooled with anyone else.

Billing

How you pay

Prepaid credits

You add credit, we draw it down per request at the rates in the table above. There is no monthly invoice, no subscription, and no minimum spend.

We chose prepaid over postpaid deliberately: it caps what you can lose if we fail, and it means we never have to chase anyone for money.

Hitting zero mid-request

An in-flight request always completes — we never truncate a response you are already paying for, even if it takes the balance negative.

The next request returns 402 with the shortfall in the body. Set a low-balance alert and auto-reload in the dashboard and you will not see it.

Refunds and invoicing

Unused credit is refundable to the original payment method within 12 months, minus payment processing fees. Email support and a human replies.

Companies that need a PO, a W-9, net-30 terms, or a countersigned order form should talk to sales — that path exists and is not a fight.

Our pricing commitment

Rates are fixed through [YYYY-MM-DD] under our current provider agreement. If our costs change after that, we will give you 60 days’ written notice before any price change, and your keys keep working at the old rate through the entire notice period.

A price increase never applies to credit you have already purchased. Credit bought at today’s rates is spent at today’s rates, however long it takes you to use it.

Do the arithmetic

What this actually saves you

36M
1M1B

Assumes a 70/30 input-to-output split and no prompt caching. Cached reads make the gap wider, not narrower.

You’d save

$120/mo

Direct from provider
$239.63
Through lowcostllm
$119.82
Discount
50% off

Three worked examples

Priced on Claude Sonnet 4.5 at a 70/30 input-to-output split, with no prompt caching applied.

Hobby

5M tokens / mo

A side project with a few hundred users

Direct
$33
With us
$16.50

$16.50/mo

Startup

150M tokens / mo

A funded team shipping an AI feature

Direct
$990
With us
$495

$495/mo

Scale

2B tokens / mo

Agentic workloads running continuously

Direct
$13200
With us
$6600

$6600/mo