Key takeaways
- The Claude API is priced per million tokens, input and output billed separately. Cost scales with use, so the number moves every month.
- You lower it three ways: prompt caching (cache reads run roughly $0.10-$0.25 per MTok), 50% off with batch processing, and matching the model to the job.
- Anthropic API pricing is usage-based, USD-billed, and spiky. Pay it from a dedicated USD card with a hard limit, so one busy week cannot decline your other tools.
Building on the Claude API feels predictable until the first heavy month lands and finance sees a number nobody forecast. A demo that cost $30 in testing becomes $3,000 in production, because real users send more tokens than your test suite ever did. That is not a billing error. Anthropic API pricing is metered per token, so the cost tracks usage directly: more calls, longer context, and chattier outputs each push the bill up. Understand the pricing model and the number stops being a surprise. If you are also weighing Claude against GPT, see how Anthropic and OpenAI API pricing compare.
How does Anthropic API pricing work?
The Claude API charges per million tokens (MTok), with separate rates for input (what you send) and output (what the model writes back). A token is roughly three-quarters of a word. There is no seat license and no monthly minimum: you pay for the tokens you move, metered to the request.
Two facts shape every bill under Anthropic API pricing. First, output almost always costs more than input, typically about 5x on these models, so a model that writes long answers spends faster than one that reads long prompts. Second, context counts as input on every call. If you resend a 40,000-token document with each request, you pay to read it every time unless you cache it.
How much does the Claude API cost?
Here are the headline rates for the four current Claude models, the core of Anthropic API pricing. The cheapest and the priciest sit a wide gap apart, which is the whole point: you match the model to the job.
| Model | Input per 1M | Output per 1M | Best for |
|---|---|---|---|
| Haiku 4.5 | $1 | $5 | Fastest, most cost-efficient. High-volume simple tasks. |
| Sonnet 5 | $2 | $10 | Balanced daily workloads and production apps. |
| Opus 5.5 | $4 | $20 | Agentic coding and enterprise reasoning. |
| Fable 5.1 | $10 | $50 | Long-running agents and extended tasks. |
Rates as of September 2026; see claude.com/pricing. Prices change, so re-verify before you budget.
A worked illustration of Anthropic API pricing, so the table means something. Say your app runs on Sonnet 5 and handles 25 million input tokens and 5 million output tokens in a month.
- Input: 25 MTok × $2 = $50
- Output: 5 MTok × $10 = $50
- Monthly total: $100
This is an illustration, not a quote. Your real split of input to output, and your model mix, decide the actual number. But the shape holds: output punches above its token count, and the model you choose sets the multiplier on everything.
How can you reduce Claude API costs?
Three levers cut Anthropic API pricing, in order of effort.
Prompt caching. If you send the same context repeatedly, a long system prompt, a knowledge base, a codebase, you cache it once and pay a cache-read rate on later calls instead of the full input rate. Cache reads run roughly $0.10-$0.25 per MTok depending on the model, well below the $1-$10 input rates above. For agents that reload the same context on every step, this is the single biggest saving.
Batch processing. Work that does not need an answer this second, overnight enrichment, evals, bulk classification, goes through the Batch API for 50% off input and output. You trade immediacy for half the bill.
Model choice. Route the simple 80% of your traffic to Haiku 4.5 at $1 / $5 and reserve Opus 5.5 at $4 / $20 for the hard 20%. A tiered setup can cut a bill by more than half without users noticing. If you need data residency, US-only inference is available at 1.1x the standard rate.
How do teams pay for a usage-based API bill?
Here is the part most Anthropic API pricing breakdowns skip. You have modeled the cost. Now something has to actually pay it, every month, in US dollars, at whatever number the meter lands on.
That is harder than it sounds for a global team. The charge is usage-based, so it is different every cycle and you cannot set it and forget it. It is billed in USD, so a card denominated in euros, pesos, or rupees takes an FX hit on every payment. And it is spiky: a product launch or a viral week can triple your token volume overnight, which is exactly when a shared company card hits its limit and declines, taking your other subscriptions down with it.
The clean fix, the same one behind paying for the OpenAI API cleanly, is a dedicated card that spends USD without a currency penalty and carries a hard limit sized to your forecast plus headroom.
Claude API prices at a glance
Haiku 4.5
$1 / $5Fastest, cheapest. Per 1M tokens, in / out.
Sonnet 5
$2 / $10High-performance default for coding and agents.
Opus 5.5
$4 / $20Daily driver for agentic and enterprise work.
Fable 5.1
$10 / $50Long-running agents.
Per million tokens, input / output, as of September 2026. Batch saves 50%; cached reads cost far less. See claude.com/pricing.
How Endl fits
An Endl virtual card is built for exactly the bill Anthropic API pricing produces. It spends USD at $0, so a US-dollar Anthropic charge carries no FX surcharge. Spend in another currency and it is the Visa rate plus a flat 1%, with no hidden markup layered on top.
Give the Claude API its own virtual card with a per-card limit set to your monthly forecast plus a buffer. If usage spikes past the limit, that one card declines and the rest of your stack keeps running, no cascade. If a key leaks, you freeze the card instantly and issue a new one, without touching the payment method behind payroll or your other tools. Virtual and physical cards are both available, each with its own limit and instant freeze.
A note on what Endl is. The balance is self-custodial and funded on stablecoin rails at a flat 0.5%. Endl is not a bank and balances are not insured. It operates as a registered VASP in the EU and an MSB in Canada. The cards are debit and spend, not credit, so you are funding real dollars you hold, not drawing a line you repay later. For a metered API bill, that is the point: you decide the ceiling before the meter runs.
See the pricing page for the full fee list, or start free and issue a card for your API spend today.




