Usage and metering
Every request draws down credits at the model catalog rates.
Quick path
- Tokens are counted per request: prompt plus completion.
- Cost = input tokens × input rate + output tokens × output rate. Cache hits halve the input rate.
- Reconcile in usage: per-request rows with model, tokens, and cost.
Details
| Topic | Decision |
|---|---|
| Counting | Tokenizer of the serving model. Prompt tokens include messages, tools, and media tokens where the model accepts them. |
| Cache hits | Matched input tokens bill at 50% of the input rate. See Caching. |
| Partial delivery | Only tokens actually delivered are billed — including revoked-key cutoffs and mid-stream failures. |
| Refusals | A refused request bills nothing. See Content safety. |
| Validation errors | A 400 before inference bills nothing. |
| Currency | USD. Balances and rows are in dollars to four decimals. See Pricing. |
Checklist
- [ ] You can map every charge to a request row.
- [ ] Cache-hit savings are visible in the rows.
- [ ] Failed-stream retries are not double-billed for undelivered tokens.
Next step
Control repeat-work cost: Caching.