Evaluation record / catalog simulationRun 024
Candidate 024 / evaluation record
A simulated AI lab workspace with evaluation history, exact benchmark values, failure sets, and reproducible run context.
Specimen data
This route demonstrates product composition. It does not report a live model, deployed capability, or external benchmark result.
- Checkpoints
- 07
- Candidate score
- 73.6
- final
- Baseline delta
- +9.3
- points
- Median latency
- 42
- milliseconds
Evaluation trajectory
Candidate and fixed baseline across seven evaluation checkpoints.
Composite score / higher is better
- Candidate
- Baseline
View data table
| Checkpoint | Candidate | Baseline | Difference |
|---|---|---|---|
| CP-01 | 61.2 | 60.4 | +0.8 |
| CP-02 | 64.8 | 61.1 | +3.7 |
| CP-03 | 66.1 | 62.2 | +3.9 |
| CP-04 | 68.9 | 62.8 | +6.1 |
| CP-05 | 70.4 | 63.4 | +7.0 |
| CP-06 | 72.8 | 64.1 | +8.7 |
| CP-07 | 73.6 | 64.3 | +9.3 |
Benchmark families
Score / 100
- Candidate
- Baseline
| Measure | Accuracy | Stability | Task fit |
|---|---|---|---|
| Candidate | 0.91 | 0.84 | 0.88 |
| Baseline | 0.82 | 0.79 | 0.80 |
| Ablation | 0.76 | 0.71 | — |
Run record
- Dataset
- Catalog evaluation set / revision 03
- Candidate
- 024 / fixed checkpoint
- Baseline
- Reference 011 / frozen
- Evaluator
- Deterministic specimen protocol
- Status
- Review ready / not deployed