Can AI predict a rare natural phenomenon, or did I just spend thousands of dollars to watch a volcano play hard to get?
The plan for Hawaiʻi was a bit of everything: beaches, waterfalls, fresh pineapple and a lava fountain. Since December 2024, Kīlauea has been erupting in episodes, with a fountaining burst roughly every two to three weeks. The pattern is regular enough to tempt a forecast, but not so regular that you can simply book it.
How do you structure an itinerary rigid enough to lock in high-end bookings, yet fluid enough to pivot at a moment's notice the second USGS signals an episodic eruption?
The planning problemUSGS Hawaiian Volcano Observatory (HVO) publishes a near-term forecast window, but it usually firms up only a few days before an episode. That isn't enough notice when you are flying across the Pacific. The question became: can historical episode data, analyzed with Claude, give a window weeks ahead that is reliable enough to book around?
Each episode is one turn of a pressure cycle in the shallow magma reservoir beneath Halemaʻumaʻu. Summit tiltmeters see it as a sawtooth: fast deflation during fountaining, then a slow, steady inflation until the system reaches its threshold again.
Fountaining drains magma and pressure from the reservoir. The ground sinks within hours.
Magma keeps flowing in from below and the reservoir repressurizes. The ground swells again for days to weeks.
Near the threshold, pathways open: small overflows, gas venting and spattering. This phase lasts hours to days.
Past the point of no return, precursors turn into full fountaining, and the cycle repeats.
The repose interval, the pause from the end of one episode to the start of the next, is the quantity every model below tries to predict. USGS has noted that since episode 5, deflation during an episode has been almost exactly matched by inflation during the following pause, which is what makes the cycle forecastable at all.
Every episode since the eruption began on December 23, 2024 was compiled from the USGS HVO episode log. For each one: start and end time (HST), duration, maximum fountain height and erupted volume. The repose interval is computed exactly, from one episode's end to the next episode's start.
| Category | Variables | How it was used |
|---|---|---|
| Time-based | Start/end timestamps, episode duration, repose interval, start-to-start interval | Target variable (repose) and its history |
| Intensity-based | Maximum fountain height (m), erupted volume (million m³) | Predictor for volume regression; context |
| Derived | Regime flag (immature: episodes 1–28, mature: 29+), rolling means, trend | Filters which history each model trusts |
| Check markers | Inflation/deflation tilt (µrad), count of precursory overflows | Reported unevenly, so used only as plausibility checks, not model inputs |
Episodes 1–28 were short, erratic pauses of under a day to 13 days while the new vents settled. From episode 29 (July 2025) the system moved to a mature regime with longer and more stable pauses of 9 to 30 days, which is the regime the models are trained on.
An episode that empties more of the reservoir takes longer to refill. Across the mature regime, each additional million m³ of lava added about days to the pause that followed. This physical link is the basis of the volume regression model.
Before any modeling, the simplest useful answer is the distribution itself. In the mature regime the most common pause is about 16 days, with a full range of 9 to 30. That gives two practical planning bands:
The next pause will equal the last one. It is a surprisingly strong baseline for a system with inertia.
Averages the last five pauses to smooth single-episode noise while still following the recent trend.
Assumes the mature regime is stationary and uses its long-run average.
Fits a straight line through mature-regime pauses by episode number to catch a slow drift.
An exponentially weighted average in which recent pauses count more, but history is never fully dropped.
Predicts the pause from how much lava the previous episode erupted, based on the physics of refilling.
A point estimate alone is useless for booking, so every model gets empirical windows. Each model was run walk-forward over episodes 35–52, predicting each pause using only data available at the time. The 25th–75th percentiles of those real errors, added to the point forecast, give the 50% window. The 10th–90th percentiles give the 80% window. A model that has been wrong by a lot in the past gets a wide window, and a consistent one gets a tight window.
Trained on episodes 1–53 and anchored to the end of episode 53 (Aug 13, 2026, 1:23 a.m. HST), every model was asked when episode 54 would start. It actually began 12.4 days later, on Aug 25 at 10:30 a.m. HST.
The smoothing models (moving average, EWMA) follow the general level but trail sudden changes by an episode. Naive persistence overreacts to them. Volume regression anticipates long pauses after big eruptions, such as episode 43's 30-day pause after 11.9 million m³.
All six models put episode 54 inside their 50% window. The three simple time-series models used for episode 55 (moving average, EWMA and naive persistence) each missed by about two days.
Hold-out verdictA note on the rebuild: for this page the full analysis was recomputed from the public USGS episode log. In the recomputation, volume regression came out as the sharpest model, both on episode 54 and over the walk-forward test. The 5-episode moving average, EWMA and naive persistence ranked next, the same top three used for the original episode 55 forecast.
Episode 54 ended on Aug 25, 2026 at 7:33 p.m. HST after about 9 hours of fountaining up to ~150 m. Adding every episode through 54 to the training data, the three trip-planning models agreed closely:
Every model's 50% window overlapped the USGS near-term forecast of Sep 7–10. The two approaches, statistics on 54 past episodes and USGS's real-time tilt modeling, independently pointed to the same days, and that agreement is what made the trip feel like a calculated risk rather than a gamble.
So what happened when the forecast met the ground? Yes, and no. Kīlauea decided to play hard to get.
Why 0.9 and not 0? The model got the physics, the timing of the threshold and the USGS agreement right. It also passed the hold-out test and put the trip in the park during a dozen precursory episodes. What it could not see was a regime change. Every model here assumes the future looks like the last 25 pauses, and when the plumbing itself changes, with new vents and stalled inflation, no amount of history can forecast that.
A model built on history is a bet on probability, not a guarantee. The 50% window is a coin flip by definition, so build the itinerary around the 80% window.
A five-episode average came within about two days of a volcano. The physically grounded volume model did even better. Six models did not need machine learning to be useful.
Statistical windows and USGS's live tilt modeling agreed on Sep 7–10. That agreement was the real confidence signal, and it justified the booking.
No fountain, but a dozen precursory episodes, real lava flows and a park full of anticipation. That is still a great story to tell.
Read the full story on Medium: Nature: 1 — Claude: 0.9.