Héctor Leopoldo Venegas Quiñones
Colorado School of Mines
| Subject Areas: | Hydrogeology,Artificial neural networks,Remote sensing applications |
Recent Activity
ABSTRACT:
A regional groundwater database for Arequipa and surroundings, southern Peru, with 4775 records. 1979 were recovered from 301 source files (university theses, government and consultancy reports) through an optical character recognition and large language model workflow with manual verification of the thesis corpus; 1865 were compiled from two groundwater-monitoring reports for the Chili aquifer; and 931 are depth-to-water records from Peru's national water information system (SNIRH-ANA). Records carry depth to water, well depth, saturated thickness, hydraulic conductivity, transmissivity, storage coefficient, test type, lithology and location, each traced to its source document. arequipa_final_gw_database.csv is the authoritative file; arequipa_final_gw_database_per_basin.xlsx holds the same records split by hydrographic unit; database.7z.001/.002 is an archive of harvested source documents. The README describes every file and four companion tables: a source catalogue, a unit-conversion table, a test-type vocabulary and a list of flagged records. This version replaces 10.4211/hs.1dfbdbe97e6d48d7b307591868ca70c4 and changes only the documentation. Data description paper: manuscript ESSD-2026-667, Earth System Science Data.
ABSTRACT:
We present the first comprehensive, open-access groundwater database for this region, developed by applying a novel workflow that integrates Optical Character Recognition (OCR) with Large Language Models (LLMs). This methodology systematically extracted and unified data from 3,675 previously unstructured 'gray literature' sources, including government reports and academic theses, spanning the period 1966–2024. The resulting database consolidates 4,775 multi-parameter data points—encompassing variables such as groundwater level, well elevation, water level (m a.s.l.), hydraulic conductivity, aquifer thickness, transmissivity, storage coefficient, specific yield, methodology used, and lithology, among others—and is augmented with 931 depth-to-water records from Peru's National Water Resources Information System (SNIRH). A rigorous three-phase quality assurance protocol, comprising 100% manual verification, statistical outlier screening, and spatial validation, was applied to ensure database accuracy. Designed to overcome data fragmentation, it enables trend analysis, groundwater modeling, and sustainable yield assessments. Publicly available via HydroShare with code on GitHub.
------------------------------------------
📌 Note on Dataset Versions:
- The file 'arequipa_final_gw_database.cvs' includes all the above data PLUS additional records sourced from the public database at [snirh.ana.gob.pe](https://snirh.ana.gob.pe), specifically for the following districts:
Apurímac, Arequipa, Ayacucho, Cusco, Ica, Moquegua, Puno, and Tacna.
- The file 'arequipa_final_gw_database_per_basin.xlsx' includes all of the aforementioned data, along with additional records sourced from the public database at snirh.ana.gob.pe. To facilitate easier classification and analysis, the workbook is organized so that each sheet corresponds to a specific basin.
➤ This extended version is intended for broader analysis and cross-validation with official national hydrological records.
ABSTRACT:
We present the first comprehensive, open-access groundwater database for this region, developed by applying a novel workflow that integrates Optical Character Recognition (OCR) with Large Language Models (LLMs). This methodology systematically extracted and unified data from 3,675 previously unstructured 'gray literature' sources, including government reports and academic theses, spanning the period 1966–2024. The resulting database consolidates 4,775 multi-parameter data points—encompassing variables such as groundwater level, well elevation, water level (m a.s.l.), hydraulic conductivity, aquifer thickness, transmissivity, storage coefficient, specific yield, methodology used, and lithology, among others—and is augmented with 931 depth-to-water records from Peru's National Water Resources Information System (SNIRH). A rigorous three-phase quality assurance protocol, comprising 100% manual verification, statistical outlier screening, and spatial validation, was applied to ensure database accuracy. Designed to overcome data fragmentation, it enables trend analysis, groundwater modeling, and sustainable yield assessments. Publicly available via HydroShare with code on GitHub.
------------------------------------------
📌 Note on Dataset Versions:
- The file 'arequipa_final_gw_database.cvs' includes all the above data PLUS additional records sourced from the public database at [snirh.ana.gob.pe](https://snirh.ana.gob.pe), specifically for the following districts:
Apurímac, Arequipa, Ayacucho, Cusco, Ica, Moquegua, Puno, and Tacna.
- The file 'arequipa_final_gw_database_per_basin.xlsx' includes all of the aforementioned data, along with additional records sourced from the public database at snirh.ana.gob.pe. To facilitate easier classification and analysis, the workbook is organized so that each sheet corresponds to a specific basin.
➤ This extended version is intended for broader analysis and cross-validation with official national hydrological records.
Contact
| (Log in to send email) |
| All | 0 |
| Collection | 0 |
| Resource | 0 |
| App Connector | 0 |
Created: Aug. 11, 2025, 10:11 p.m.
Authors: Venegas-Quiñones, Héctor Leopoldo · Madeleine Guillen · Pablo A. Garcia-Chevesich · Brett Uhle · Edgard González · Javier Ticona · José Díaz · Julia Zea · Francisco Alejo · McCray, John
ABSTRACT:
We present the first comprehensive, open-access groundwater database for this region, developed by applying a novel workflow that integrates Optical Character Recognition (OCR) with Large Language Models (LLMs). This methodology systematically extracted and unified data from 3,675 previously unstructured 'gray literature' sources, including government reports and academic theses, spanning the period 1966–2024. The resulting database consolidates 4,775 multi-parameter data points—encompassing variables such as groundwater level, well elevation, water level (m a.s.l.), hydraulic conductivity, aquifer thickness, transmissivity, storage coefficient, specific yield, methodology used, and lithology, among others—and is augmented with 931 depth-to-water records from Peru's National Water Resources Information System (SNIRH). A rigorous three-phase quality assurance protocol, comprising 100% manual verification, statistical outlier screening, and spatial validation, was applied to ensure database accuracy. Designed to overcome data fragmentation, it enables trend analysis, groundwater modeling, and sustainable yield assessments. Publicly available via HydroShare with code on GitHub.
------------------------------------------
📌 Note on Dataset Versions:
- The file 'arequipa_final_gw_database.cvs' includes all the above data PLUS additional records sourced from the public database at [snirh.ana.gob.pe](https://snirh.ana.gob.pe), specifically for the following districts:
Apurímac, Arequipa, Ayacucho, Cusco, Ica, Moquegua, Puno, and Tacna.
- The file 'arequipa_final_gw_database_per_basin.xlsx' includes all of the aforementioned data, along with additional records sourced from the public database at snirh.ana.gob.pe. To facilitate easier classification and analysis, the workbook is organized so that each sheet corresponds to a specific basin.
➤ This extended version is intended for broader analysis and cross-validation with official national hydrological records.
Created: June 9, 2026, 6:52 p.m.
Authors: Venegas-Quiñones, Héctor Leopoldo · Madeleine Guillen · Pablo A. Garcia-Chevesich · Brett Uhle · Edgard González · Javier Ticona · José Díaz · Julia Zea · Francisco Alejo · McCray, John
ABSTRACT:
We present the first comprehensive, open-access groundwater database for this region, developed by applying a novel workflow that integrates Optical Character Recognition (OCR) with Large Language Models (LLMs). This methodology systematically extracted and unified data from 3,675 previously unstructured 'gray literature' sources, including government reports and academic theses, spanning the period 1966–2024. The resulting database consolidates 4,775 multi-parameter data points—encompassing variables such as groundwater level, well elevation, water level (m a.s.l.), hydraulic conductivity, aquifer thickness, transmissivity, storage coefficient, specific yield, methodology used, and lithology, among others—and is augmented with 931 depth-to-water records from Peru's National Water Resources Information System (SNIRH). A rigorous three-phase quality assurance protocol, comprising 100% manual verification, statistical outlier screening, and spatial validation, was applied to ensure database accuracy. Designed to overcome data fragmentation, it enables trend analysis, groundwater modeling, and sustainable yield assessments. Publicly available via HydroShare with code on GitHub.
------------------------------------------
📌 Note on Dataset Versions:
- The file 'arequipa_final_gw_database.cvs' includes all the above data PLUS additional records sourced from the public database at [snirh.ana.gob.pe](https://snirh.ana.gob.pe), specifically for the following districts:
Apurímac, Arequipa, Ayacucho, Cusco, Ica, Moquegua, Puno, and Tacna.
- The file 'arequipa_final_gw_database_per_basin.xlsx' includes all of the aforementioned data, along with additional records sourced from the public database at snirh.ana.gob.pe. To facilitate easier classification and analysis, the workbook is organized so that each sheet corresponds to a specific basin.
➤ This extended version is intended for broader analysis and cross-validation with official national hydrological records.
Created: Sept. 28, 2026, 2:24 p.m.
Authors: Venegas-Quiñones, Héctor Leopoldo · Madeleine Guillen · Pablo A. Garcia-Chevesich · Brett Uhle · Edgard González · Javier Ticona · José Díaz · Julia Zea · Francisco Alejo · McCray, John
ABSTRACT:
A regional groundwater database for Arequipa and surroundings, southern Peru, with 4775 records. 1979 were recovered from 301 source files (university theses, government and consultancy reports) through an optical character recognition and large language model workflow with manual verification of the thesis corpus; 1865 were compiled from two groundwater-monitoring reports for the Chili aquifer; and 931 are depth-to-water records from Peru's national water information system (SNIRH-ANA). Records carry depth to water, well depth, saturated thickness, hydraulic conductivity, transmissivity, storage coefficient, test type, lithology and location, each traced to its source document. arequipa_final_gw_database.csv is the authoritative file; arequipa_final_gw_database_per_basin.xlsx holds the same records split by hydrographic unit; database.7z.001/.002 is an archive of harvested source documents. The README describes every file and four companion tables: a source catalogue, a unit-conversion table, a test-type vocabulary and a list of flagged records. This version replaces 10.4211/hs.1dfbdbe97e6d48d7b307591868ca70c4 and changes only the documentation. Data description paper: manuscript ESSD-2026-667, Earth System Science Data.