In brief
Correcting sub-daily satellite precipitation extremes (GPM IMERG V07) with machine learning becomes challenging when the ground observational reference itself carries spatial and density uncertainties. Benchmarking tabular models (LightGBM) and spatial deep learning (CNN) across strict temporal, spatial, and event-based holdouts demonstrates that continuous post-processing reliably reduces bulk error, but deterministic recovery of heavy rainfall tails hits fundamental limits. Direct probabilistic exceedance modeling provides the most dependable operational ranking for early warning and hazard screening.
Contribution
A systematic evaluation of satellite precipitation extreme correctability limits under operational reference uncertainty, benchmarking tabular, deep learning, and probabilistic exceedance models across out-of-sample validation splits.
Key finding
Continuous correction with LightGBM reduces global RMSE significantly across sub-daily timescales (e.g. from 0.461 to 0.360 at 1 h), but fixed extreme events remain difficult to reconstruct deterministically. Direct probabilistic exceedance modeling provides the most robust risk ranking for operational hazard screening.
Models and data
Distinctive research elements
Models and methods
- LightGBM (M3/M3b)
- CNN spatial specialist (M4)
- Probabilistic exceedance model
Data and evaluation
- GPM IMERG Final V07
- AVAMET rain-gauge network (556 gauges, 2019-2025)
Author-written overview
Research overview
We present a statistical assessment of out-of-sample correctability limits under observational uncertainty, using GPM IMERG Final V07 sub-daily areal precipitation extremes over the Comunitat Valenciana (eastern Spain) and the dense AVAMET network (556 gauges, 2019–2025). Rather than treating gauges as point truth, we construct a gauge-derived multi-station operational areal reference proxy at 30 min over 126 IMERG cells and build a closed benchmark at 1, 3, and 6 h under year-holdout, province-holdout, and event-fold validation. The benchmark compares raw IMERG (M0), two simple statistical corrections (M1–M2), a continuous LightGBM corrector (M3), a two-stage tail-oriented LightGBM model (M3b), and a conservative patch-based CNN specialist (M4). We further evaluate reference sensitivity under alternative proxy definitions and pseudo leave-one-gauge-out perturbations, and we treat direct probabilistic exceedance modelling as a core benchmark component alongside deterministic correction. Under the main operational-median proxy, M3 is the best global continuous corrector at all scales, reducing event-fold RMSE from 0.461 to 0.360 at 1 h, from 0.801 to 0.578 at 3 h, and from 1.445 to 1.259 at 6 h. However, these gains do not translate into clean recovery of severe and extreme events under fixed operational thresholds. Tail-oriented models improve severe-event skill relative to M3 at 3 h and 6 h, but only modestly, while fixed extreme recovery remains weak across reference perturbations and hard holdouts. The conservative local deep-learning specialist does not materially improve the tail-oriented tabular baseline. About half of the severe and extreme cases occur in cells with the minimum observational support of two gauges, and local jackknife diagnostics show increasing proxy sensitivity in the upper tail. Peak and alignment-oracle diagnostics indicate that residual error is only partly explained by local spatiotemporal misalignment; substantial amplitude underestimation persists even after temporal and neighborhood tolerance, and IMERG often enters far below the operational threshold in observed extreme cases. Importantly, direct probabilistic exceedance modelling emerges as the main positive result: it provides robust out-of-sample risk ranking for severe and extreme exceedance across split families, although operational precision remains limited by event rarity. Overall, the results support partial out-of-sample correctability of IMERG sub-daily areal extremes under an operational areal proxy, while probabilistic risk ranking emerges as the central operational output when clean deterministic recovery of the fixed extreme tail remains unattained. For operational use, the most cautious interpretation is to pair continuous correction with direct probabilistic risk ranking rather than rely on corrected satellite estimates alone for threshold-based event detection in early-warning or hazard-screening applications.
Cite this work
BibTeX
@article{semper2026out,
title = {Out-of-sample statistical correctability limits under an uncertain operational reference: the case of IMERG sub-daily areal precipitation extremes},
author = {Marc Semper and Manuel Curado and Jose F. Vicent and Leandro Tortosa},
journal = {Stochastic Environmental Research and Risk Assessment},
year = {2026},
volume = {40},
number = {8},
eid = {202},
doi = {10.1007/s00477-026-03336-6}
}Always check the publisher record for final volume, issue and page information before citing.