RaoscaffResearch
Prediction Series · Lock · Issue F-03
Prediction Series · F-03

74 resolved. 70 right. And the four we got wrong.

The RAOSCAFF Prediction Series is 194 public forecasts, each committed to a git timestamp before its event. 74 have now resolved: 70 hits, 4 misses — a 94.6% resolved hit rate at a Brier score of 0.0496. This report publishes the wins, the losses, the calibration curve, and the three caveats that qualify the headline. A forecast record you cannot check is not a record.

Type · Flagship · consolidated scorecard As of · 2026-08-20 · recomputed from the public scorecard Scope · 194 locked forecasts · 74 resolved · 120 still open Method · every lock timestamped before its event, scored on its literal published criterion
Resolved forecasts · hits vs misses
70 / 74
94.6% resolved hit rate · Brier 0.0496

70 hits and 4 misses across 74 resolved forecasts. Brier score of 0.0496 across the 57 resolved forecasts that carry a single binary probability, where 0 is perfect, 0.25 is a coin flip and 1.0 is maximally wrong. 120 of the 194 locks have not yet reached their resolution date.

— 1 · The Record

194 locked forecasts. 74 resolved. 70 right.

Every forecast in this series is committed to a public git history before its event occurs — the commit timestamp is the lock, and it cannot be backdated. Each one names, in advance, the exact metric it will be scored against and the source that will settle it. As of 20 August 2026, 74 of the 194 locks have reached their resolution date. Seventy resolved as hits and four as misses.

Series scorecard as of 2026-08-20
MeasureValue
Total locked forecasts194
Resolved74
Hits70
Misses4
Resolved hit rate94.6%
Still open (not yet due)120
Brier score (binary subset, n=57)0.0496
— 2 · The Misses

All four, in full. This is the part most track records omit.

A scorecard that publishes only the wins is advertising. Here are all four losses, with what we said and what happened. The most expensive was P-02: we stated 100% probability on a document being signed as drafted, and it was signed a day later as a modified text. A 100% claim that fails carries the maximum possible Brier penalty of 1.0 — a single forecast that costs more than sixty correct ones earn. We have not stated 100% on anything since.

Every resolved miss
#The callWhat happened
P-02US–Iran 14-point MoU signed as drafted — stated 100%Signed as a MODIFIED text. Headline miss, Brier 1.0. The clause we ranked most likely to slip was the one that slipped.
P-03India June CPI headline at 4.6%Printed 4.38%. A point-estimate forecast, missed on the literal number.
P-05Tesla Q2 deliveries in the 400–425k bucket480,126 — blew through the top of the bucket.
P-69Coinbase Q2 subscription & services revenue ≥$565M$555M. Missed the floor by $10M, or 1.8%.
— 3 · Calibration

We are systematically under-confident by 9.3 points — and that is a design choice, not an edge.

Across the 57 resolved forecasts carrying a single binary probability, the average stated confidence was 87.2% and the actual hit rate was 96.5% — a gap of +9.3 points. That is not evidence of superior judgement. It is the arithmetic consequence of how these forecasts are built: floors are deliberately set BELOW company guidance and ceilings ABOVE the latest print, so that the call clears unless something genuinely breaks. A series designed that way should beat its stated confidence. The honest reading is that our probabilities are conservative, not that our foresight is exceptional.

Calibration by stated-confidence band (n=57)
Stated bandnActualStatedGap
60–75%4100.0%69.5%+30.5pt
75–85%887.5%80.5%+7.0pt
85–92%31100.0%88.5%+11.5pt
92%+1492.9%93.4%−0.5pt
— 4 · What Qualifies All Of The Above

Three caveats. Including one that makes our own Brier score look better than it should.

First, the record is partial: 120 of 194 locks have not yet resolved, and a hit rate computed on the 74 that have is not a claim about the 120 that have not. Second, the Brier score of 0.0496 is computed on 57 forecasts — the ones carrying a single binary probability. Seventeen resolved forecasts are point estimates, buckets, distributions or Monte-Carlo ensembles with no single number to score, and TWO OF OUR FOUR MISSES sit in that excluded group. The Brier subset is therefore structurally kinder than the full record, and 94.6% is the honest topline rather than the 96.5% the subset shows. Third, the highest-confidence band is the only one where we are not under-confident: at 92%+ stated we delivered 92.9%, which is calibration, not luck. We publish all three because a number you cannot interrogate is not evidence.

— 5 · How A Forecast Is Scored

The rule that makes the score mean something: the criterion is fixed before the event.

Each lock names its metric, its threshold, its resolution date and its settling source in advance. At resolution we score that criterion and no other. This matters more than it sounds. Walmart's Q2 FY2027 (P-76) cleared our locked total-net-sales floor comfortably while missing Street expectations on US comparable sales in the same release — we recorded the hit and disclosed the miss on the metric we had not locked. Alibaba (P-139) renamed its cloud segment between our lock and the print, so we published the composition change alongside the result. And P-01 turned on the difference between the CPI energy-commodities sub-index and the broader energy index, which moved seventeen points apart in the same month. Choosing which number counts, in public, before you know the answer, is the whole discipline.

Scorecard as of 2026-08-20 — recomputed from the public series data.

Every figure in this report is derived from the published scorecard that drives the live track record, recomputed on the date of publication. The series and its results can be audited independently.

As of
2026-08-20
Locked forecasts
194 (P-01 through P-194)
Resolved
74 — 70 hits, 4 misses
Resolved hit rate
94.6%
Brier score
0.0496 across 57 binary-probability forecasts
Calibration
stated 87.2% vs actual 96.5% — under-confident by 9.3pt
Live track record
raoscaff.com/intel/track-record

120 forecasts remain open. This scorecard will move — in both directions — as they resolve.