Official warnings and forecasts come from the National Hurricane Center. This page independently scores SHIPS-RII's published probabilities against what happened; it supplements NHC guidance and never replaces it.

RII Verification

Did SHIPS-RII's rapid-intensification probabilities actually predict what happened? Scored against best-track outcomes, 2018–2025 backfill plus the live 2026 season — a live number, not a fixed retrospective study, so it grows as pairs accrue. As of Jul 27, 2026, 6:46 AM UTC.

These numbers refresh automatically — every scheduled pipeline cycle re-scores the pool against that cycle's newly-elapsed SHIPS windows, not just once at launch. A refresh only lands if it passes an internal sanity check first; if a cycle's result looks implausible, this page keeps showing the last good numbers rather than a broken update, so "As of" above can lag by more than one cycle on occasion.

Headline pair: 30 kt / 24 h, SHIPS-RII row, pre-registered before backfill ran. Given SHIPS issued a forecast, how skillful was it — not skill across all storm-hours: cycles where SHIPS structurally couldn't run (sub-threshold intensity, center inland) create no pairs and are excluded from this number, not scored as zero. n = 5290 pairs from 312 storms.

Brier Skill Score (vs. sample base rate)
0.133
base rate 7.3%, n=5290
Brier Skill Score (vs. SHIPS climatology)
0.108
SHIPS' own embedded climatology reference
AUC (discrimination)
0.870
95% CI [0.846, 0.891]

SHIPS-RII ranks RI cases well while being systematically over-confident — these answer different questions, not one. Discrimination (AUC 0.870) asks whether RI cases got higher probabilities than non-RI cases; on that question SHIPS-RII does strongly. Calibration (the reliability table below, and the MHW-stratum calibration further down) asks whether a 30% forecast verifies about 30% of the time; on that question SHIPS-RII runs systematically high — real skill in ranking is not the same claim as well-calibrated probabilities, and this page reports both because a forecast can pass one and fail the other.

Reliability

Forecast binnObserved freq.Wilson 95%Block-bootstrap 95%
0.0%–10.0%24970.7%[0.4%, 1.1%][0.3%, 1.1%]
10.0%–20.0%17265.0%[4.1%, 6.2%][3.7%, 6.5%]
20.0%–30.0%57118.0%[15.1%, 21.4%][13.9%, 22.2%]
30.0%–40.0%25123.9%[19.0%, 29.5%][17.6%, 30.6%]
40.0%–50.0%11843.2%[34.6%, 52.2%][31.2%, 55.3%]
50.0%–60.0%4753.2%[39.2%, 66.7%][36.8%, 68.9%]
60.0%–70.0%3342.4%[27.2%, 59.2%][21.1%, 63.3%]
70.0%–80.0%1838.9%[20.3%, 61.4%][15.8%, 66.7%]
80.0%–90.0%1369.2%[42.4%, 87.3%][40.0%, 92.9%]
90.0%–100.0%1681.3%[57.0%, 93.4%][54.5%, 100.0%]

Bins with n < 5 are shown hollow, not dropped — a bin's absence should never be silently invisible.

Marine heatwave stratified analysis

Storm-level split across the full dataset: 208 storms reached MHW category ≥1 somewhere along their track, 101 never did. 7 storms (6 Central Pacific, structurally outside the tracked OISST domain; 1 whose track never resolved to a position on it) are excluded from every MHW comparison below, not silently folded into either side.

Within-season comparison (required, not a sensitivity row)

MHW-active and inactive storms are compared separately within each season, then pooled (Mantel-Haenszel) — a naive pool of all seasons together can't be trusted here, because 2023–2024 were both far more MHW-active and far more RI-active than every other season in this dataset, entangling an MHW effect with a season effect. The within-season method is what makes it possible to ask whether an effect survives once that confound is controlled for.

Risk difference (MHW-active − inactive), pooled
+12.2 pts
95% CI [+1.9 pts, +22.6 pts] · bootstrap cross-check [+1.3 pts, +23.2 pts]
Odds ratio, pooled (secondary)
1.729×
95% CI [0.990, 3.020]

Suggestive, not established. The odds-ratio confidence interval's lower bound sits at 0.990 — barely above 1 — and the positive direction holds in 6 of 9 seasons shown below, not all of them. Two data points is not a trend and eight seasons is not a settled result; this is evidence worth taking seriously, not a confirmed effect. The per-season table below is shown in full every time this result is, on purpose — a pooled number without the breakdown would hide exactly the season-to-season variation that matters for judging how strong this evidence actually is.

Seasonn activeRate activen inactiveRate inactiveΔ
20182147.6%1428.6%+19.0 pts
20192927.6%90.0%+27.6 pts
20203330.3%1811.1%+19.2 pts
20212035.0%1822.2%+12.8 pts
20221926.3%1637.5%-11.2 pts
20233243.8%837.5%+6.3 pts
20242839.3%40.0%+39.3 pts
20251931.6%1233.3%-1.8 pts
2026520.0%0

7 storm(s) excluded for unmeasurable MHW, 4 for having no resolved 30kt/24h pair —305 storms enter the comparison above.

No pair or storm in this dataset (2018-2025 backfill plus 2026 live) has ever had a tracked position coincide with MHW category 3 ("Severe") or 4 ("Extreme") water, even though both levels occur elsewhere in the tracked ocean domain. Every dose-response or MHW-stratum result on this page therefore covers categories 0-2 only and must not be read as extending to extreme marine-heatwave conditions (spec Decision 36).

Calibration by MHW stratum (secondary)

Not a BSS/AUC comparison here on purpose — BSS and AUC are base-rate- sensitive and the two strata differ in base rate by construction, so comparing them directly would partly just be re-measuring the split itself. This instead asks the more direct operational question: does SHIPS' own forecast probability track what actually happened, within each stratum?

MHW-active storms
-4.9 pts
observed 7.6% vs. forecast 12.4%, n=3913 · gap CI [-6.3 pts, -3.3 pts]
MHW-inactive storms
-6.0 pts
observed 6.1% vs. forecast 12.1%, n=1271 · gap CI [-8.2 pts, -3.7 pts]

SHIPS over-forecasts in both strata (negative gap, neither confidence interval crosses zero) — the same over-confidence the headline calibration finding above describes, and it shows up whether or not the storm ever crossed a marine heatwave.

No pair or storm in this dataset (2018-2025 backfill plus 2026 live) has ever had a tracked position coincide with MHW category 3 ("Severe") or 4 ("Extreme") water, even though both levels occur elsewhere in the tracked ocean domain. Every dose-response or MHW-stratum result on this page therefore covers categories 0-2 only and must not be read as extending to extreme marine-heatwave conditions (spec Decision 36).

Dose-response: forecast-track intersection with active MHW cells (secondary)

Category itself never exceeds 2 anywhere in this dataset (see the caveat below), so the fraction of a storm's forecast track intersecting an active MHW cell is the only real continuous exposure axis available here — a secondary, supporting look, not a second confirmation of the within-season result above.

Track intersectionnObserved RI freq.Wilson 95%Block-bootstrap 95%
0.0 (none)31757.0%[6.2%, 8.0%][5.5%, 8.6%]
partial10198.0%[6.5%, 9.9%][5.5%, 10.9%]
1.0 (full)59610.7%[8.5%, 13.5%][5.9%, 16.2%]

The trend runs in one direction across all three buckets — more track intersecting active MHW cells lines up with a higher observed RI frequency. Reported as supporting the within-season finding above, not as independent confirmation of it: both are drawn from the same underlying storms and the same MHW classification.

500 pair(s) excluded for unmeasurable MHW position.

No pair or storm in this dataset (2018-2025 backfill plus 2026 live) has ever had a tracked position coincide with MHW category 3 ("Severe") or 4 ("Extreme") water, even though both levels occur elsewhere in the tracked ocean domain. Every dose-response or MHW-stratum result on this page therefore covers categories 0-2 only and must not be read as extending to extreme marine-heatwave conditions (spec Decision 36).

Sensitivity analysis

Every row below is a reported delta from the primary number above, not an alternate view — the headline is always the primary pool.

Variantn (pairs/storms)BSS (sample)AUC
dissipated excluded4684 / 2860.1420.860
et included6397 / 3160.1360.889
landfall excluded4698 / 3060.1840.885
era 2018 20255216 / 3060.1360.870
era 202674 / 6-0.1130.979
basin al2819 / 1520.0750.844
basin ep2372 / 1540.1510.876
model version 2018 20191380 / 760.1590.865
model version 2020 20253836 / 2300.1270.872

The model-version rows (2018–2019 vs. 2020–2025) test a documented candidate SHIPS-RII predictor revision — sourced from secondary literature, not a primary read. A large delta there is suggestive of version-mixing, not proof.

MHW stratum, BSS/AUC (sensitivity only — see the calibration section above for the primary MHW comparison)

Stratumn (pairs/storms)Base rateBSS (sample)AUC
mhw active storms3913 / 2067.6%0.1290.863
mhw inactive storms1271 / 996.1%0.0920.883

Base rate shown deliberately next to BSS/AUC here — the two strata differ in base rate by construction, so a naive BSS/AUC comparison between them would partly just remeasure that difference. Read alongside the base rate, not instead of it.

What these numbers can and can't support: this is a student research tool. Reliability and skill scores here reflect 312 distinct storms, not 5290 independent observations — SHIPS cycles every 6 hours, so pairs from one storm are correlated, and the block-bootstrap intervals above account for that. The MHW within-season result is suggestive, not established — treat it as evidence worth taking seriously, not a confirmed effect. Where a sensitivity row's n is small, treat its delta as suggestive, not conclusive.