Data Driver
PUBLIC SCORING

Track Record

Every prediction we publish is scored against reality. This page is our permanent, public accountability record. No retroactive adjustments. No cherry-picking.

Every prediction is scored against the actual outcome and surfaced on the chart below. We share both hits and misses — the regulation reset taught the model new patterns and our rookie predictions are still calibrating, so each scored race shifts the picture.

scoring LIVE API

Brier score

Lower is better, 0 is perfect

Skill score

vs grid baseline

11

Live races scored

+4 reconstructed · 15 in the history

BRIER EVOLUTION

Brier score evolution

Per-race score over the season — hover for details

TDD
Grid baseline
0.0000.0120.0240.0350.0470.059AustraliJapaneseCanadianBarcelonBritishHungariaItalianAzerbaij
RACE RESULTS MATRIX

Race results matrix

Hover any cell for race details

Correct
Miss
Not published
—Aust
—Chin
—Japa
—Miam
—Cana
—*Mona
—*Barc
—*Aust
—*Brit
—Belg
—Hung
—Dutc
—Ital
—Span
—Azer
Top 1 (Win)
Top 3 (Podium)
Beats Grid
Brier score
0.040
0.035
0.037
0.033
0.036
0.040
0.038
0.039
0.040
0.038
0.041
0.051
0.032
0.028
0.043

* Reconstructed by walk-forward backfill, not a live pre-race prediction. Headline figures count live races only.

CALIBRATION

Live calibration

Hover any point for details

perfect0%10%20%30%40%50%0%10%20%30%40%50%Predicted probabilityObserved frequency
Observed vs predicted
Perfect calibration
95% confidence interval

Drag to read the curve

The model was well calibrated here — prediction and reality agree. Based on 202 scored win probabilities in this probability band.

When we predict a 30% chance, it happens ~30% of the time. Points on the diagonal indicate perfect calibration.

BRIER DECOMPOSITION

Reliability

calibration error — lower is better

Resolution

how much the model separates outcomes

Uncertainty

inherent unpredictability of F1

BASELINES
ModelBrierSkill Score
The Data Driver0.038reference
Grid baseline0.047+20.0%
Championship baseline0.060+37.3%
Random uniform0.054+30.4%

Skill Score = improvement over baseline. BSS = 1 - (model / reference).

SHARPNESS

Mean favourite probability

Higher sharpness = model is more decisive

30.3%

11%
Australian
36%
Chinese
41%
Japanese
42%
Miami
40%
Canadian
25%
Belgian
20%
Hungarian
34%
Dutch
26%
Italian
29%
Spanish
30%
Azerbaijan

Trend: stable — model confidence is stable across races

LOG LOSS

Log loss penalises confident wrong predictions more heavily than Brier score

—Australian Grand Prix
0.153

log loss

0.040

brier

—Chinese Grand Prix
0.106

log loss

0.035

brier

—Japanese Grand Prix
0.110

log loss

0.037

brier

—Miami Grand Prix
0.098

log loss

0.033

brier

—Canadian Grand Prix
0.108

log loss

0.036

brier

—Monaco Grand Prix
0.152

log loss

0.040

brier

—Barcelona-Catalunya Grand Prix
0.139

log loss

0.038

brier

—Austrian Grand Prix
0.141

log loss

0.039

brier

—British Grand Prix
0.150

log loss

0.040

brier

—Belgian Grand Prix
0.125

log loss

0.038

brier

—Hungarian Grand Prix
0.138

log loss

0.041

brier

—Dutch Grand Prix
0.190

log loss

0.051

brier

—Italian Grand Prix
0.104

log loss

0.032

brier

—Spanish Grand Prix
0.091

log loss

0.028

brier

—Azerbaijan Grand Prix
0.140

log loss

0.043

brier

HISTORICAL VALIDATION

Walk-forward backtest

Validated on 114 races (2021-2025)

Model retrained before each race using only past data — no future leakage

0.041

Brier score

0.558

Spearman rank

Transparency

Predictions published before each race. Track record computed automatically after each race using Brier score. We never modify predictions retroactively. Every probability is timestamped and immutable.

Scoring methodology: Brier score = mean squared error of probabilistic predictions. Skill score = 1 - (model Brier / baseline Brier). Grid baseline uses qualifying order.