Understand the scores
Equal eval weights. Visible unknowns. One final percentage per model.
On this page
Three averages, in order
- Repetitions → metric
- Metrics → eval
- Evals → benchmark %
aggregation
metric = mean(repetitions)
eval = mean(metrics in this eval)
benchmark = 100 × mean(evals)
An eval with more metrics does not get more benchmark weight. Each eval contributes equally to the final percentage.
Change a decision. See the score.
Score explorerIllustration · one repetition per metric
SQL queryAsks for dialect
SQL queryBound parameters
Issue summaryActionable summary
Benchmark score75%
Equal eval weights
Select a decision to cycle through pass → fail → unknown. The calculation uses OpenEval's score projection.
Read each number in context
| Viewer value | Interpretation |
|---|---|
| Final percentage | All required metric checks have known values |
| Checks scored | Metric decisions across evals and repetitions |
| Eval drilldown | Individual metrics for one selected eval |
| Runtime | Union of active execution intervals; overlapping work counts once |
| ETA | Estimate for scheduled work using candidate and judge stage timings |
| Cost | Reported candidate and judge spend |
