Aldo Tapia Araya

Universidad de La Serena;University Of Arizona

Subject Areas: Catchment hydrology,Model calibration

 Recent Activity

ABSTRACT:

Deep learning rainfall-runoff models have shown the ability to reproduce the observed hydrograph exceeding the performance of physics-based models, but their reliability outside the training distribution is still an open question. Most robustness evaluation have focused on a single architecture (LSTM), and have used controlled forcing perturbations (temperature offset or precipitation scaling) to probe the non-stationarity condition. In this study, we evaluate the robustness of five deep learning architectures (LSTM, minLSTM, minGRU, MLP-Mixer, Transformer) against a re-calibrated physics-based benchmark (SAC-SMA + Snow-17) across 554 CAMELS catchments, driven by a six-Global Change Model (GCM), four-Shared Socioeconomic Pathway (SSP) NEX-GDDP-CMIP6 and three-future periods ensemble. The results show that the deep learning models outperform the physics-based benchmark in terms of accuracy, but their robustness is limited under non-stationary conditions. The LSTM and minGRU architectures show the best performance among the deep learning models, while the MLP-Mixer and minLSTM architectures show the worst performance. We found that the predictive skill does not imply robustness: the LSTM, the most accurate architecture, is the least robust. Transformer, the most robust on annual volumes, fails to reproduce the intra-annual redistribution of streamflow, diverging from the physics-based benchmark in both directions, compensating that difference at annual basis. Architecture, not GCMs, SSPs nor period, is the dominant factor of divergence (45 \% of the evaluated catchments). The results highlight the need for a more comprehensive evaluation of deep learning models under non-stationary conditions, and the importance of considering both accuracy and robustness in model selection.

Show More

ABSTRACT:

Deep learning rainfall-runoff models have shown the ability to reproduce the observed hydrograph exceeding the performance of physics-based models, but their reliability outside the training distribution is still an open question. Most robustness evaluation have focused on a single architecture (LSTM), and have used controlled forcing perturbations (temperature offset or precipitation scaling) to probe the non-stationarity condition. In this study, we evaluate the robustness of five deep learning architectures (LSTM, minLSTM, minGRU, MLP-Mixer, Transformer) against a re-calibrated physics-based benchmark (SAC-SMA + Snow-17) across 554 CAMELS catchments, driven by a six-Global Change Model (GCM), four-Shared Socioeconomic Pathway (SSP) NEX-GDDP-CMIP6 and three-future periods ensemble. The results show that the deep learning models outperform the physics-based benchmark in terms of accuracy, but their robustness is limited under non-stationary conditions. The LSTM and minGRU architectures show the best performance among the deep learning models, while the MLP-Mixer and minLSTM architectures show the worst performance. We found that the predictive skill does not imply robustness: the LSTM, the most accurate architecture, is the least robust. Transformer, the most robust on annual volumes, fails to reproduce the intra-annual redistribution of streamflow, diverging from the physics-based benchmark in both directions, compensating that difference at annual basis. Architecture, not GCMs, SSPs nor period, is the dominant factor of divergence (45 \% of the evaluated catchments). The results highlight the need for a more comprehensive evaluation of deep learning models under non-stationary conditions, and the importance of considering both accuracy and robustness in model selection.

Show More

ABSTRACT:

Accurate short-term streamflow forecasting is crucial for effective water resource management and mitigating hydrological extremes, such as floods and droughts. HidroCL DB is a comprehensive lumped hydrometeorological database designed to support short-term streamflow forecasting across 432 catchments in Continental Chile, with daily records spanning 2000 to 2025. The database integrates diverse data sources, including GIS layers, climatic reanalysis, satellite observations, and numerical weather predictions. Data are organized by catchment and grouped into static variables (e.g., catchment characteristics, water rights, soil properties), observed variables (e.g., satellite-derived indicators, meteorological data, reservoir surface area), forecasted variables (e.g., numerical weather prediction outputs), and streamflow measurements from gauge stations. HidroCL DB provides a valuable resource for researchers and practitioners to develop, validate, and benchmark hydrological models for operational streamflow forecasting in Chile.

Show More

 Contact

Resources
All 0
Collection 0
Resource 0
App Connector 0
Resource Resource
HidroCL Database: A Comprehensive Hydrometeorological Database for Streamflow Forecasting in Continental Chile
Created: July 17, 2025, 7:55 p.m.
Authors: Tapia Araya, Aldo · Jorge Arevalo · Jorge Saavedra-Garrido · Luis De LaFuente · ChristopherParedes-Arroyo

ABSTRACT:

Accurate short-term streamflow forecasting is crucial for effective water resource management and mitigating hydrological extremes, such as floods and droughts. HidroCL DB is a comprehensive lumped hydrometeorological database designed to support short-term streamflow forecasting across 432 catchments in Continental Chile, with daily records spanning 2000 to 2025. The database integrates diverse data sources, including GIS layers, climatic reanalysis, satellite observations, and numerical weather predictions. Data are organized by catchment and grouped into static variables (e.g., catchment characteristics, water rights, soil properties), observed variables (e.g., satellite-derived indicators, meteorological data, reservoir surface area), forecasted variables (e.g., numerical weather prediction outputs), and streamflow measurements from gauge stations. HidroCL DB provides a valuable resource for researchers and practitioners to develop, validate, and benchmark hydrological models for operational streamflow forecasting in Chile.

Show More
Resource Resource

ABSTRACT:

Deep learning rainfall-runoff models have shown the ability to reproduce the observed hydrograph exceeding the performance of physics-based models, but their reliability outside the training distribution is still an open question. Most robustness evaluation have focused on a single architecture (LSTM), and have used controlled forcing perturbations (temperature offset or precipitation scaling) to probe the non-stationarity condition. In this study, we evaluate the robustness of five deep learning architectures (LSTM, minLSTM, minGRU, MLP-Mixer, Transformer) against a re-calibrated physics-based benchmark (SAC-SMA + Snow-17) across 554 CAMELS catchments, driven by a six-Global Change Model (GCM), four-Shared Socioeconomic Pathway (SSP) NEX-GDDP-CMIP6 and three-future periods ensemble. The results show that the deep learning models outperform the physics-based benchmark in terms of accuracy, but their robustness is limited under non-stationary conditions. The LSTM and minGRU architectures show the best performance among the deep learning models, while the MLP-Mixer and minLSTM architectures show the worst performance. We found that the predictive skill does not imply robustness: the LSTM, the most accurate architecture, is the least robust. Transformer, the most robust on annual volumes, fails to reproduce the intra-annual redistribution of streamflow, diverging from the physics-based benchmark in both directions, compensating that difference at annual basis. Architecture, not GCMs, SSPs nor period, is the dominant factor of divergence (45 \% of the evaluated catchments). The results highlight the need for a more comprehensive evaluation of deep learning models under non-stationary conditions, and the importance of considering both accuracy and robustness in model selection.

Show More
Resource Resource

ABSTRACT:

Deep learning rainfall-runoff models have shown the ability to reproduce the observed hydrograph exceeding the performance of physics-based models, but their reliability outside the training distribution is still an open question. Most robustness evaluation have focused on a single architecture (LSTM), and have used controlled forcing perturbations (temperature offset or precipitation scaling) to probe the non-stationarity condition. In this study, we evaluate the robustness of five deep learning architectures (LSTM, minLSTM, minGRU, MLP-Mixer, Transformer) against a re-calibrated physics-based benchmark (SAC-SMA + Snow-17) across 554 CAMELS catchments, driven by a six-Global Change Model (GCM), four-Shared Socioeconomic Pathway (SSP) NEX-GDDP-CMIP6 and three-future periods ensemble. The results show that the deep learning models outperform the physics-based benchmark in terms of accuracy, but their robustness is limited under non-stationary conditions. The LSTM and minGRU architectures show the best performance among the deep learning models, while the MLP-Mixer and minLSTM architectures show the worst performance. We found that the predictive skill does not imply robustness: the LSTM, the most accurate architecture, is the least robust. Transformer, the most robust on annual volumes, fails to reproduce the intra-annual redistribution of streamflow, diverging from the physics-based benchmark in both directions, compensating that difference at annual basis. Architecture, not GCMs, SSPs nor period, is the dominant factor of divergence (45 \% of the evaluated catchments). The results highlight the need for a more comprehensive evaluation of deep learning models under non-stationary conditions, and the importance of considering both accuracy and robustness in model selection.

Show More