Official warnings and forecasts come from the National Hurricane Center. This page independently scores SHIPS-RII's published probabilities against what happened; it supplements NHC guidance and never replaces it.
RII Verification
Did SHIPS-RII's rapid-intensification probabilities actually predict what happened? Scored against best-track outcomes, 2018–2025 backfill plus the live 2026 season — a live number, not a fixed retrospective study, so it grows as pairs accrue. As of Jul 27, 2026, 6:46 AM UTC.
These numbers refresh automatically — every scheduled pipeline cycle re-scores the pool against that cycle's newly-elapsed SHIPS windows, not just once at launch. A refresh only lands if it passes an internal sanity check first; if a cycle's result looks implausible, this page keeps showing the last good numbers rather than a broken update, so "As of" above can lag by more than one cycle on occasion.
Headline pair: 30 kt / 24 h, SHIPS-RII row, pre-registered before backfill ran. Given SHIPS issued a forecast, how skillful was it — not skill across all storm-hours: cycles where SHIPS structurally couldn't run (sub-threshold intensity, center inland) create no pairs and are excluded from this number, not scored as zero. n = 5290 pairs from 312 storms.
SHIPS-RII ranks RI cases well while being systematically over-confident — these answer different questions, not one. Discrimination (AUC 0.870) asks whether RI cases got higher probabilities than non-RI cases; on that question SHIPS-RII does strongly. Calibration (the reliability table below, and the MHW-stratum calibration further down) asks whether a 30% forecast verifies about 30% of the time; on that question SHIPS-RII runs systematically high — real skill in ranking is not the same claim as well-calibrated probabilities, and this page reports both because a forecast can pass one and fail the other.
Reliability
| Forecast bin | n | Observed freq. | Wilson 95% | Block-bootstrap 95% |
|---|---|---|---|---|
| 0.0%–10.0% | 2497 | 0.7% | [0.4%, 1.1%] | [0.3%, 1.1%] |
| 10.0%–20.0% | 1726 | 5.0% | [4.1%, 6.2%] | [3.7%, 6.5%] |
| 20.0%–30.0% | 571 | 18.0% | [15.1%, 21.4%] | [13.9%, 22.2%] |
| 30.0%–40.0% | 251 | 23.9% | [19.0%, 29.5%] | [17.6%, 30.6%] |
| 40.0%–50.0% | 118 | 43.2% | [34.6%, 52.2%] | [31.2%, 55.3%] |
| 50.0%–60.0% | 47 | 53.2% | [39.2%, 66.7%] | [36.8%, 68.9%] |
| 60.0%–70.0% | 33 | 42.4% | [27.2%, 59.2%] | [21.1%, 63.3%] |
| 70.0%–80.0% | 18 | 38.9% | [20.3%, 61.4%] | [15.8%, 66.7%] |
| 80.0%–90.0% | 13 | 69.2% | [42.4%, 87.3%] | [40.0%, 92.9%] |
| 90.0%–100.0% | 16 | 81.3% | [57.0%, 93.4%] | [54.5%, 100.0%] |
Bins with n < 5 are shown hollow, not dropped — a bin's absence should never be silently invisible.
Marine heatwave stratified analysis
Storm-level split across the full dataset: 208 storms reached MHW category ≥1 somewhere along their track, 101 never did. 7 storms (6 Central Pacific, structurally outside the tracked OISST domain; 1 whose track never resolved to a position on it) are excluded from every MHW comparison below, not silently folded into either side.
Within-season comparison (required, not a sensitivity row)
MHW-active and inactive storms are compared separately within each season, then pooled (Mantel-Haenszel) — a naive pool of all seasons together can't be trusted here, because 2023–2024 were both far more MHW-active and far more RI-active than every other season in this dataset, entangling an MHW effect with a season effect. The within-season method is what makes it possible to ask whether an effect survives once that confound is controlled for.
Suggestive, not established. The odds-ratio confidence interval's lower bound sits at 0.990 — barely above 1 — and the positive direction holds in 6 of 9 seasons shown below, not all of them. Two data points is not a trend and eight seasons is not a settled result; this is evidence worth taking seriously, not a confirmed effect. The per-season table below is shown in full every time this result is, on purpose — a pooled number without the breakdown would hide exactly the season-to-season variation that matters for judging how strong this evidence actually is.
| Season | n active | Rate active | n inactive | Rate inactive | Δ |
|---|---|---|---|---|---|
| 2018 | 21 | 47.6% | 14 | 28.6% | +19.0 pts |
| 2019 | 29 | 27.6% | 9 | 0.0% | +27.6 pts |
| 2020 | 33 | 30.3% | 18 | 11.1% | +19.2 pts |
| 2021 | 20 | 35.0% | 18 | 22.2% | +12.8 pts |
| 2022 | 19 | 26.3% | 16 | 37.5% | -11.2 pts |
| 2023 | 32 | 43.8% | 8 | 37.5% | +6.3 pts |
| 2024 | 28 | 39.3% | 4 | 0.0% | +39.3 pts |
| 2025 | 19 | 31.6% | 12 | 33.3% | -1.8 pts |
| 2026 | 5 | 20.0% | 0 | — | — |
7 storm(s) excluded for unmeasurable MHW, 4 for having no resolved 30kt/24h pair —305 storms enter the comparison above.
No pair or storm in this dataset (2018-2025 backfill plus 2026 live) has ever had a tracked position coincide with MHW category 3 ("Severe") or 4 ("Extreme") water, even though both levels occur elsewhere in the tracked ocean domain. Every dose-response or MHW-stratum result on this page therefore covers categories 0-2 only and must not be read as extending to extreme marine-heatwave conditions (spec Decision 36).
Calibration by MHW stratum (secondary)
Not a BSS/AUC comparison here on purpose — BSS and AUC are base-rate- sensitive and the two strata differ in base rate by construction, so comparing them directly would partly just be re-measuring the split itself. This instead asks the more direct operational question: does SHIPS' own forecast probability track what actually happened, within each stratum?
SHIPS over-forecasts in both strata (negative gap, neither confidence interval crosses zero) — the same over-confidence the headline calibration finding above describes, and it shows up whether or not the storm ever crossed a marine heatwave.
No pair or storm in this dataset (2018-2025 backfill plus 2026 live) has ever had a tracked position coincide with MHW category 3 ("Severe") or 4 ("Extreme") water, even though both levels occur elsewhere in the tracked ocean domain. Every dose-response or MHW-stratum result on this page therefore covers categories 0-2 only and must not be read as extending to extreme marine-heatwave conditions (spec Decision 36).
Dose-response: forecast-track intersection with active MHW cells (secondary)
Category itself never exceeds 2 anywhere in this dataset (see the caveat below), so the fraction of a storm's forecast track intersecting an active MHW cell is the only real continuous exposure axis available here — a secondary, supporting look, not a second confirmation of the within-season result above.
| Track intersection | n | Observed RI freq. | Wilson 95% | Block-bootstrap 95% |
|---|---|---|---|---|
| 0.0 (none) | 3175 | 7.0% | [6.2%, 8.0%] | [5.5%, 8.6%] |
| partial | 1019 | 8.0% | [6.5%, 9.9%] | [5.5%, 10.9%] |
| 1.0 (full) | 596 | 10.7% | [8.5%, 13.5%] | [5.9%, 16.2%] |
The trend runs in one direction across all three buckets — more track intersecting active MHW cells lines up with a higher observed RI frequency. Reported as supporting the within-season finding above, not as independent confirmation of it: both are drawn from the same underlying storms and the same MHW classification.
500 pair(s) excluded for unmeasurable MHW position.
No pair or storm in this dataset (2018-2025 backfill plus 2026 live) has ever had a tracked position coincide with MHW category 3 ("Severe") or 4 ("Extreme") water, even though both levels occur elsewhere in the tracked ocean domain. Every dose-response or MHW-stratum result on this page therefore covers categories 0-2 only and must not be read as extending to extreme marine-heatwave conditions (spec Decision 36).
Sensitivity analysis
Every row below is a reported delta from the primary number above, not an alternate view — the headline is always the primary pool.
| Variant | n (pairs/storms) | BSS (sample) | AUC |
|---|---|---|---|
| dissipated excluded | 4684 / 286 | 0.142 | 0.860 |
| et included | 6397 / 316 | 0.136 | 0.889 |
| landfall excluded | 4698 / 306 | 0.184 | 0.885 |
| era 2018 2025 | 5216 / 306 | 0.136 | 0.870 |
| era 2026 | 74 / 6 | -0.113 | 0.979 |
| basin al | 2819 / 152 | 0.075 | 0.844 |
| basin ep | 2372 / 154 | 0.151 | 0.876 |
| model version 2018 2019 | 1380 / 76 | 0.159 | 0.865 |
| model version 2020 2025 | 3836 / 230 | 0.127 | 0.872 |
The model-version rows (2018–2019 vs. 2020–2025) test a documented candidate SHIPS-RII predictor revision — sourced from secondary literature, not a primary read. A large delta there is suggestive of version-mixing, not proof.
MHW stratum, BSS/AUC (sensitivity only — see the calibration section above for the primary MHW comparison)
| Stratum | n (pairs/storms) | Base rate | BSS (sample) | AUC |
|---|---|---|---|---|
| mhw active storms | 3913 / 206 | 7.6% | 0.129 | 0.863 |
| mhw inactive storms | 1271 / 99 | 6.1% | 0.092 | 0.883 |
Base rate shown deliberately next to BSS/AUC here — the two strata differ in base rate by construction, so a naive BSS/AUC comparison between them would partly just remeasure that difference. Read alongside the base rate, not instead of it.
What these numbers can and can't support: this is a student research tool. Reliability and skill scores here reflect 312 distinct storms, not 5290 independent observations — SHIPS cycles every 6 hours, so pairs from one storm are correlated, and the block-bootstrap intervals above account for that. The MHW within-season result is suggestive, not established — treat it as evidence worth taking seriously, not a confirmed effect. Where a sensitivity row's n is small, treat its delta as suggestive, not conclusive.