The RAOSCAFF Prediction Series is 194 public forecasts, each committed to a git timestamp before its event. 74 have now resolved: 70 hits, 4 misses — a 94.6% resolved hit rate at a Brier score of 0.0496. This report publishes the wins, the losses, the calibration curve, and the three caveats that qualify the headline. A forecast record you cannot check is not a record.
70 hits and 4 misses across 74 resolved forecasts. Brier score of 0.0496 across the 57 resolved forecasts that carry a single binary probability, where 0 is perfect, 0.25 is a coin flip and 1.0 is maximally wrong. 120 of the 194 locks have not yet reached their resolution date.
Every forecast in this series is committed to a public git history before its event occurs — the commit timestamp is the lock, and it cannot be backdated. Each one names, in advance, the exact metric it will be scored against and the source that will settle it. As of 20 August 2026, 74 of the 194 locks have reached their resolution date. Seventy resolved as hits and four as misses.
| Measure | Value |
|---|---|
| Total locked forecasts | 194 |
| Resolved | 74 |
| Hits | 70 |
| Misses | 4 |
| Resolved hit rate | 94.6% |
| Still open (not yet due) | 120 |
| Brier score (binary subset, n=57) | 0.0496 |
A scorecard that publishes only the wins is advertising. Here are all four losses, with what we said and what happened. The most expensive was P-02: we stated 100% probability on a document being signed as drafted, and it was signed a day later as a modified text. A 100% claim that fails carries the maximum possible Brier penalty of 1.0 — a single forecast that costs more than sixty correct ones earn. We have not stated 100% on anything since.
| # | The call | What happened |
|---|---|---|
| P-02 | US–Iran 14-point MoU signed as drafted — stated 100% | Signed as a MODIFIED text. Headline miss, Brier 1.0. The clause we ranked most likely to slip was the one that slipped. |
| P-03 | India June CPI headline at 4.6% | Printed 4.38%. A point-estimate forecast, missed on the literal number. |
| P-05 | Tesla Q2 deliveries in the 400–425k bucket | 480,126 — blew through the top of the bucket. |
| P-69 | Coinbase Q2 subscription & services revenue ≥$565M | $555M. Missed the floor by $10M, or 1.8%. |
Across the 57 resolved forecasts carrying a single binary probability, the average stated confidence was 87.2% and the actual hit rate was 96.5% — a gap of +9.3 points. That is not evidence of superior judgement. It is the arithmetic consequence of how these forecasts are built: floors are deliberately set BELOW company guidance and ceilings ABOVE the latest print, so that the call clears unless something genuinely breaks. A series designed that way should beat its stated confidence. The honest reading is that our probabilities are conservative, not that our foresight is exceptional.
| Stated band | n | Actual | Stated | Gap |
|---|---|---|---|---|
| 60–75% | 4 | 100.0% | 69.5% | +30.5pt |
| 75–85% | 8 | 87.5% | 80.5% | +7.0pt |
| 85–92% | 31 | 100.0% | 88.5% | +11.5pt |
| 92%+ | 14 | 92.9% | 93.4% | −0.5pt |
First, the record is partial: 120 of 194 locks have not yet resolved, and a hit rate computed on the 74 that have is not a claim about the 120 that have not. Second, the Brier score of 0.0496 is computed on 57 forecasts — the ones carrying a single binary probability. Seventeen resolved forecasts are point estimates, buckets, distributions or Monte-Carlo ensembles with no single number to score, and TWO OF OUR FOUR MISSES sit in that excluded group. The Brier subset is therefore structurally kinder than the full record, and 94.6% is the honest topline rather than the 96.5% the subset shows. Third, the highest-confidence band is the only one where we are not under-confident: at 92%+ stated we delivered 92.9%, which is calibration, not luck. We publish all three because a number you cannot interrogate is not evidence.
Each lock names its metric, its threshold, its resolution date and its settling source in advance. At resolution we score that criterion and no other. This matters more than it sounds. Walmart's Q2 FY2027 (P-76) cleared our locked total-net-sales floor comfortably while missing Street expectations on US comparable sales in the same release — we recorded the hit and disclosed the miss on the metric we had not locked. Alibaba (P-139) renamed its cloud segment between our lock and the print, so we published the composition change alongside the result. And P-01 turned on the difference between the CPI energy-commodities sub-index and the broader energy index, which moved seventeen points apart in the same month. Choosing which number counts, in public, before you know the answer, is the whole discipline.
Every figure in this report is derived from the published scorecard that drives the live track record, recomputed on the date of publication. The series and its results can be audited independently.
120 forecasts remain open. This scorecard will move — in both directions — as they resolve.