Skip to main content

Benchmarks

Bernstein benchmarks

Two things, each reproducible and each honest about its limits: a simulated parallel-scheduling speedup model, and a community cost benchmark you can add your own run to. Every number below is either computed from a published data file or quoted from the repo, and the page tells you which.

No leaderboard sheet where Bernstein wins everything. The scheduling figures are a labelled simulation, not measured runs.

Parallel scheduling speedup (simulation)

This is a simulation, not measured agent runs. It models a greedy list scheduler over 10 tasks with realistic dependency graphs, so treat it as a capacity-planning estimate rather than a leaderboard claim.

ConfigurationAvg speedup vs single agentCost effect
3 agents1.78xmodel mixing reduces cost ~23%
5 agents2.18xsame mix, more parallel slots

Full task table, scheduling model, and cost model are in docs/benchmarks/BENCHMARKS.md. Reproduce locally with uv run python benchmarks/run_benchmark.py.

Community cost benchmark

What does a Bernstein session actually cost on your setup? This benchmark collects real bernstein cost --json runs so you can sanity-check the bill before you install, across different models and hardware rather than from one machine.

To add a row, run bernstein cost --json after a session and submit the task count, model breakdown, total cost, cost-per-task, wall-clock time, and hardware. Two paths:

Submitted runs are published with attribution to the setup they came from. The schema is whatever bernstein cost --json emits, so there is nothing to fill in by hand.