Request-level pricing modifiers for Anthropic (fast mode, batch, priority, data residency)
A coverage audit of Anthropic's pricing surfaced several request-level modifiers that change the effective rate but aren't currently modeled for Anthropic. Filing one consolidated issue (separate from the cache-TTL #295 and web-search #202 threads).
Each is signaled by a field already present in Anthropic's usage object:
| Modifier |
Effect |
usage signal |
Values |
| Fast mode |
premium flat rates (Opus 4.6/4.7 = $30 in / $150 out; Opus 4.8 = $10/$50), across full context window |
usage.speed |
standard | fast |
| Batch API |
0.5× input & output |
usage.service_tier |
standard | batch |
| Priority tier |
commitment pricing |
usage.service_tier |
standard | priority |
| Data residency (US) |
1.1× on all token categories (Opus 4.6 / Sonnet 4.6+) |
usage.inference_geo |
global(default) | us |
Stacking, per Anthropic's pricing page: fast mode applies across the full window and stacks with prompt-caching multipliers and data residency, but not with batch; batch + caching combine; data residency 1.1× applies to input / output / cache-write / cache-read.
The batch/priority pair is the Anthropic side of the same "request tier changes the rate" mechanism already requested for OpenAI in #115 — so a shared, condition-keyed price-variant mechanism (matching on a usage field rather than only date/time) would likely cover both providers at once.
Questions for maintainers:
- Is modeling these as request-condition-keyed price variants the direction you'd want, or is there a preferred shape?
- Would you accept a PR adding these for Anthropic once the shape is agreed?
Happy to contribute — we have verified rate tables and real usage samples for all four.
Request-level pricing modifiers for Anthropic (fast mode, batch, priority, data residency)
A coverage audit of Anthropic's pricing surfaced several request-level modifiers that change the effective rate but aren't currently modeled for Anthropic. Filing one consolidated issue (separate from the cache-TTL #295 and web-search #202 threads).
Each is signaled by a field already present in Anthropic's
usageobject:usagesignalusage.speedstandard|fastusage.service_tierstandard|batchusage.service_tierstandard|priorityusage.inference_geoglobal(default) |usStacking, per Anthropic's pricing page: fast mode applies across the full window and stacks with prompt-caching multipliers and data residency, but not with batch; batch + caching combine; data residency 1.1× applies to input / output / cache-write / cache-read.
The batch/priority pair is the Anthropic side of the same "request tier changes the rate" mechanism already requested for OpenAI in #115 — so a shared, condition-keyed price-variant mechanism (matching on a
usagefield rather than only date/time) would likely cover both providers at once.Questions for maintainers:
Happy to contribute — we have verified rate tables and real
usagesamples for all four.