OpenEvalDocumentation
GitHub
RUN & INSPECT

Understand the scores

Equal eval weights. Visible unknowns. One final percentage per model.

On this page

Three averages, in order

  1. Repetitions → metric
  2. Metrics → eval
  3. Evals → benchmark %
aggregation
metric = mean(repetitions)
eval = mean(metrics in this eval)
benchmark = 100 × mean(evals)

An eval with more metrics does not get more benchmark weight. Each eval contributes equally to the final percentage.

Change a decision. See the score.

Score explorerIllustration · one repetition per metric
SQL queryAsks for dialect
SQL queryBound parameters
Issue summaryActionable summary
Benchmark score75%
Equal eval weights

Select a decision to cycle through pass → fail → unknown. The calculation uses OpenEval's score projection.

Read each number in context

Viewer valueInterpretation
Final percentageAll required metric checks have known values
Checks scoredMetric decisions across evals and repetitions
Eval drilldownIndividual metrics for one selected eval
RuntimeUnion of active execution intervals; overlapping work counts once
ETAEstimate for scheduled work using candidate and judge stage timings
CostReported candidate and judge spend
OpenEval model score chart with three illustrative models
Illustrative viewer data. Every model has one benchmark score.Open full size
Find in documentation