pay per task, not per seat
the bernstein router picks the cheapest model that still passes your tests on each task. this page lets you put your own bills in and see what the math suggests for your situation. heuristic, not a promise.
your bill, audited
what cheapest-passing-test routing would shift.
enter your last month's spend on three of the bills bernstein users typically pay. the calculation below uses a documented heuristic, hardcoded model prices, and shows every step so you can audit it.
estimated band [?]
$180-$360 /mo
show the math
- total monthly llm spend$400 + $200 + $0 = $600
- fraction of tasks routable to a cheaper model that still passes tests40%-80% (heuristic)
- cost ratio: cheapest-passing model vs current premium model~25% of original (see model-prices table below)
- saving = total × routable% × (1 − cost ratio)$180-$360 /mo
the routable% is a heuristic, not a measured number. real saving depends on whether the cheaper models pass your project's tests on each task. on a repo where tests are flaky or coverage is low, routing falls back to the premium model and the band shrinks toward zero. on a repo with tight tests and a lot of mechanical work, the band shifts higher than 80%.
your last month's bill was $600.
sponsoring at $25/mo is 4% of $600.
bernstein keeps routing the cheapest model that passes tests.
model prices
| model | input / 1m | cached input / 1m | output / 1m |
|---|---|---|---|
| deepseek | |||
deepseek-v4.1-flash | $0.02 | $0.01 | $0.60 |
deepseek-v4-pro | $0.9553 | $0.07961 | $1.911 |
| openai | |||
gpt-oss-safeguard-20b | $0.075 | $0.0375 | $0.30 |
gpt-6-astra | $10 | $1 | $50 |
gemma-4-26b-a4b-it | $0.09 | $0.05 | $0.30 |
gemini-3.5-flash | $1.5 | $0.15 | $9 |
| anthropic | |||
claude-haiku-4.5 | $1 | $0.10 | $5 |
claude-fable-5.1 | $10 | $0.25 | $50 |
| xai | |||
grok-build-0.1 | $1 | $0.20 | $2 |
grok-4.6 | $2 | $0.50 | $6 |
sources: claude.com/pricing, platform.openai.com/docs/pricing, ai.google.dev/gemini-api/docs/pricing, openrouter.ai/models.
the table is generated, not typed: the catalogue is filtered down to the current flagship and the current cheap tier per provider, so a new release replaces the row it supersedes on the next refresh. cached input is the provider's own cache-read rate where one is published, and n/a where it is not. absolute numbers can shift by 20-40% between revisions even when relative ordering survives, and the catalogue rate is what this listing charges rather than what you would pay on a direct contract, so check the source links before quoting these in a procurement conversation.
how the band is computed
the calculator multiplies your total monthly llm spend by the fraction of tasks bernstein could route to a cheaper model (40-80%, a heuristic) and by the cost gap between the premium and cheap models (about 75% saving on the routable tasks). the spread is visible in the table above: the flagship rows and the cheap-tier rows are typically one to two orders of magnitude apart on input, so swapping one for the other on a routable task saves nearly all of the input cost; the 75% blended figure is the band-weighted average across mixed task types. the result is a band, not a point. the math is shown step by step in the calculator block so you can substitute your own assumptions if the heuristic does not fit your repo.
the routable fraction is the load-bearing assumption. on a codebase with flaky tests it skews toward zero - the cheaper models do not pass, the bandit falls back to the premium model, and the saving collapses. on a codebase with tight tests and a lot of mechanical work (typed refactors, test scaffolding, lint fixes) it skews higher than the upper bound. neither extreme is a promise.
if you want to verify any of this on your own repo before you sponsor, install bernstein with pipx install bernstein and check the cost column in the run report after one parallel run. the numbers there are real, not heuristic.
frequently asked
How much does Bernstein save on LLM bills?
It depends on how much of your work is routable to a cheaper model that still passes your tests. The calculator on this page uses a heuristic: 40-80% of tasks are routable, and the cheapest passing model costs about a quarter of the premium model. On a $600/month combined Claude + Codex + Cursor bill, that suggests a band of roughly $180-360/month shifted. Real saving will be lower if your tests are flaky and higher if you have a lot of mechanical work in the repo.
How does Bernstein decide which model to route a task to?
Bernstein runs an epsilon-greedy contextual bandit over a per-task pass-rate history. Each task type (lint fix, test generation, refactor, architecture, tests-and-boilerplate) has its own arm. The bandit prefers the cheapest model whose recent pass rate on that task type is above a configurable threshold, and explores a more expensive model with probability epsilon.
Why is the calculator output a band and not a single number?
A single number would be marketing, not honest. The actual saving depends on how many tasks route to a cheaper model (varies with task mix), how often the cheaper model passes your tests (varies with test quality), and how aggressively you tune the bandit explore rate. The band is the lower and upper bounds of a heuristic that assumes routing kicks in 40-80% of the time.
Does sponsoring Bernstein affect what it routes to?
No. Bernstein is on-prem only. It runs on your machine, calls the model APIs you configure with your own keys, and writes state to disk you own. Sponsorship funds the operator, not the routing logic. Routing decisions are deterministic Python in src/bernstein/core/routing/bandit_router.py, scheduled from src/bernstein/core/orchestration/ - what model wins is a function of the bandit history, the cost table, and your test results.