Eval Regressions Between Versions
Before a new model version ships, an AI lab runs it through a suite of evaluation tasks and compares the scores with the version before it. A release that gets better on average can still get worse at something specific, and the release review wants every such regression listed.
Tables
models
model_version: the version's name, such asv9orv10.released_on: the day it was released; no two versions share a day.
eval_results
model_version: the version that was evaluated.task: the evaluation task, such asmathorcoding.score: the score on that task, from 0 to 100 with two decimals.
Not every version is evaluated on every task.
Task
For each task, compare every version's score with the score of the previous version evaluated on that same task, where "previous" means the most recent earlier release. List the regressions: the cases where the score dropped by more than 2 points.
Example
In the sample, v8 was not evaluated on translation, so v9's 61.00 is compared with v7's 66.00: a drop of 5.00, listed first. On summarization, v8 dropped exactly 2.00 from v7, which is not more than 2, so it is not listed.
| task | model_version | previous_version | previous_score | score | score_drop |
|---|---|---|---|---|---|
| translation | v9 | v7 | 66 | 61 | 5 |
| math | v9 | v8 | 74 | 70.25 | 3.75 |
| coding | v8 | v7 | 60 | 57.5 | 2.5 |
Submitting also runs your answer against 3 hidden datasets, each built around an edge case: NULLs, ties, empty tables. A failure names the case without showing its data.
Follow-up: Scores on small tasks are noisy. How would you change the report to flag a regression only when it is larger than the run-to-run variation you measured by evaluating the same version several times?
- Return the columns `task`, `model_version`, `previous_version`, `previous_score`, `score` and `score_drop` (`previous_score - score`), in that order. - Versions are ordered by `released_on`, never by name: `v10` was released after `v9`. - Only drops strictly greater than 2.00 are listed. A task's first evaluated version has nothing to compare with. - Sort by `score_drop` from largest to smallest, then by `task`, then by `model_version`.
- Views
- 3