Checking for missing content type metadata ...
This resource contains content types with missing metadata required to make it public or discoverable. Show missing content type metadata.
Click on the edit button ( ) below to edit this resource.
Checking for non-preferred file/folder path names (may take a long time depending on the number of files/folders) ...
This resource contains some files/folders that have non-preferred characters in their name. Show non-conforming files/folders.
This resource contains content types with files that need to be updated to match with metadata changes. Show content type files that need updating.
Regional Groundwater Database from Unstructured Sources Based on an Integrated Optical Character Recognition and Large Language Model Workflow for Arequipa and Surroundings in Southern Peru
| Authors: |
|
|
|---|---|---|
| Owners: |
|
This resource does not have an owner who is an active HydroShare user. Contact CUAHSI (help@cuahsi.org) for information on this resource. |
| Type: | Resource | |
| Storage: | The size of this resource is 6.3 GB | |
| Created: | Sep 28, 2026 at 2:24 p.m. (UTC) | |
| Last updated: | Sep 28, 2026 at 2:59 p.m. (UTC) | |
| Citation: | See how to cite this resource |
| Sharing Status: | Public |
|---|---|
| Views: | 37 |
| Downloads: | 9 |
| +1 Votes: | Be the first one to this. |
| Comments: | No comments (yet) |
Abstract
A regional groundwater database for Arequipa and surroundings, southern Peru, with 4775 records. 1979 were recovered from 301 source files (university theses, government and consultancy reports) through an optical character recognition and large language model workflow with manual verification of the thesis corpus; 1865 were compiled from two groundwater-monitoring reports for the Chili aquifer; and 931 are depth-to-water records from Peru's national water information system (SNIRH-ANA). Records carry depth to water, well depth, saturated thickness, hydraulic conductivity, transmissivity, storage coefficient, test type, lithology and location, each traced to its source document. arequipa_final_gw_database.csv is the authoritative file; arequipa_final_gw_database_per_basin.xlsx holds the same records split by hydrographic unit; database.7z.001/.002 is an archive of harvested source documents. The README describes every file and four companion tables: a source catalogue, a unit-conversion table, a test-type vocabulary and a list of flagged records. This version replaces 10.4211/hs.1dfbdbe97e6d48d7b307591868ca70c4 and changes only the documentation. Data description paper: manuscript ESSD-2026-667, Earth System Science Data.
Subject Keywords
Coverage
Spatial
Temporal
| Start Date: | |
|---|---|
| End Date: |
Content
README.txt
README
Regional Groundwater Database for Arequipa and Surroundings, Southern Peru
====================================================================================================
Documentation version 2 (September 2026). Replaces the README of November 2025.
Resource title: Regional Groundwater Database from Unstructured Sources Based on an Integrated
Optical Character Recognition and Large Language Model Workflow for Arequipa and Surroundings in
Southern Peru
Authors: Venegas-Quiñones, H. L., Guillen, M., Garcia-Chevesich, P. A., Uhle, B., González, E.,
Ticona, J., Díaz, J., Zea, J., Alejo, F., McCray, J.
This version: https://www.hydroshare.org/resource/0f9223cb27584d00ba46b87a5af471ef/ (a DOI is
assigned when this version is published). It is a new version of
https://doi.org/10.4211/hs.1dfbdbe97e6d48d7b307591868ca70c4 (published 3 August 2026); the data
files are unchanged, and the documentation was corrected and extended.
Data description paper: Venegas-Quiñones et al., manuscript ESSD-2026-667 submitted to Earth System
Science Data; preprint DOI https://doi.org/10.5194/essd-2026-667 (active once the preprint is
posted).
Licence: CC BY 4.0 applies to the database files, the companion tables and this documentation. The
source documents in database.7z remain under the terms of their original publishers.
Contact: Héctor L. Venegas-Quiñones, Colorado School of Mines (hector.venegasquinones@mines.edu)
READ THIS FIRST
----------------------------------------------------------------------------------------------------
* arequipa_final_gw_database.csv is the authoritative data file (4775 records). The per-basin
workbook is a convenience copy that differs in formatting (Section 3). The 6.8 GB archive holds
source documents only and is not needed to use the data.
* Line 2 of the CSV (and row 2 of every basin worksheet) repeats the column names in Spanish. It is
a second header row, not a record. The CSV is encoded in Windows-1252 (cp1252), not UTF-8.
* Measurement_Date is written M/D/YYYY, month first, and its precision varies: 1024 dates are
1/1/YYYY year-only placeholders, the 981 dates of the 2014-2021 TGC compilation are campaign months
written as day 1, and 221 more are year-level or estimated. Read Measurement_Date_Justification
before using dates.
* Hydraulic conductivity and transmissivity are stored with the unit printed in the source. Only 42
of 224 K values are in metres per day (written 'm/d', 'm/día' or 'm/day'), and 36 of 58 T values are
per second. Convert with unit_conversion_table.csv before any calculation. For depth to water, well
depth and thickness, values whose unit string says cm or mm are already in metres.
* The K column mixes measurement supports. Filter on Test_Type (test_type_vocabulary.csv): 57 K
records come from pumping or recovery tests, which are 48 distinct tests once repeats are removed
(Section 8); their median is 16.4 m/d, about 34 times the median of the other 167 values (0.48 m/d).
* flagged_records.csv lists 221 records with values that are not field measurements (literature
values, estimates, geophysical interpretations, surface-water points, design values) or that carry a
unit problem. Check it before any analysis.
1. FILE INVENTORY
----------------------------------------------------------------------------------------------------
Part A. Data and documentation (small files; most users need only these)
arequipa_final_gw_database.csv [3.1 MB]
The database: 4775 records, 56 columns. Authoritative (Section 2).
arequipa_final_gw_database_per_basin.xlsx [1.0 MB]
The same records in 23 worksheets, one per hydrographic unit, plus a summary sheet (Section
3).
source_catalog.csv [new]
One row per source file (303 files plus the SNIRH network): title, authors, years, records
contributed, whether and where the document is in database.7z, and how to obtain it otherwise
(Section 6).
unit_conversion_table.csv [new]
Every unit string used for K, T, depth to water, well depth and thickness, with its factor to
the reference unit (Section 7).
test_type_vocabulary.csv [new]
Every Test_Type value attached to K, T or S, with its measurement support and suitability for
aquifer K (Section 8).
flagged_records.csv [new]
Records whose values are not field measurements or carry a unit problem (Section 9).
README.txt [this file]
Description of every file.
DATA_SOURCES_AND_TRACEABILITY.txt [small]
How to trace a record to its source document (Section 5).
Part B. Collection basis (large; needed only to check values against source documents)
database.7z.001 [5243 MB] and database.7z.002 [1555 MB]
Two volumes of one 7-Zip archive of source documents: 647 PDF files in 604 folders, 8.06 GB
uncompressed. It does not contain the database (Section 4).
2. arequipa_final_gw_database.csv
----------------------------------------------------------------------------------------------------
4775 records and 56 columns, one row per observation point and date as reported in a source. 3844
records come from 303 source files of theses and technical reports (the documentary stream); 931
come from the national monitoring network SNIRH-ANA (Full_Source = 'snirh.ana.gob.pe'). All K, T and
S values belong to the documentary stream.
Format: comma-separated values, Windows-1252 encoding, CRLF line endings. Line 1 holds the English
column names, line 2 the original Spanish names, lines 3 to 4777 the records. Missing values are
empty cells. Numbers use a full stop as decimal separator and are stored as text as printed; convert
after reading.
Python: df = pd.read_csv('arequipa_final_gw_database.csv', encoding='cp1252', skiprows=[1], dtype=str)
df['date'] = pd.to_datetime(df['Measurement_Date'], format='%m/%d/%Y')
R: df <- read.csv('arequipa_final_gw_database.csv', fileEncoding = 'CP1252',
colClasses = 'character')[-1, ]
df$date <- as.Date(df$Measurement_Date, format = '%m/%d/%Y')
Column dictionary. 'Spanish' is the name on line 2; 'n' is the number of records with a value.
# Column (line 1) Spanish (line 2) n Description
1 Record_ID ID_Registro 4775 Unique record identifier,
integers 1 to 4790; 15
numbers (2881 to 2895)
are unused, so Record_ID
is not the row number.
2 File_Code Codigo_Archivo 4775 Internal code of the
source file assigned
during extraction.
3 Full_Source Fuente_Completa 4775 File name or harvest
identifier of the source
document, or
'snirh.ana.gob.pe' for
national-network records.
Join key to
source_catalog.csv.
4 Document_Title Titulo_Documento 4775 Title of the source
document.
5 Authors Autor_es 4744 Author(s) or issuing
institution of the source
document.
6 Publication_Year Año_Publicacion 4744 Year of the source
document; for multi-year
compilations it varies
between records of the
same file.
7 Measurement_Date Fecha_Medicion 4732 Date of the measurement,
format M/D/YYYY (month
first). Precision varies
(Section 11): 1024 values
are 1/1/YYYY year-only
placeholders, the 981
dates of 'Informe
final_Victor
Camacho_ok.pdf' are
campaign months written
as day 1, and 221 more
are year-level or
estimated according to
Measurement_Date_Justification;
43 are empty.
8 Measurement_Date_Justification Justificacion_Fecha_Medicion 3801 How the date was
assigned; read it to
judge date precision.
Values starting 'Año',
'Estimad' or 'Periodo'
mean the day and often
the month are not known.
9 Original_Point_ID ID_Punto_Original 4767 Point identifier as
printed in the source.
10 Base_Point_ID ID_Punto_Base 3844 Harmonised point
identifier linking
records of the same
point.
11 Test_Depth_Start Profundidad_Ensayo_Inicio 175 Top of the tested
interval, as reported.
12 Test_Depth_End Profundidad_Ensayo_Fin 156 Bottom of the tested
interval, as reported.
13 Point_Type Tipo_Punto 4769 Observation point type as
reported, in Spanish (342
distinct values, e.g.
'Pozo', 'Pozo a tajo
abierto', 'Manantial',
'Piezómetro',
'Calicata').
14 Location_Description Descripcion_Ubicacion 3844 Location text as printed
in the source (some
values carry leading or
trailing spaces).
15 UTM_East_Coordinate Coordenada_Este_UTM 4758 Easting as stated in, or
inferred from, the source
(see Coordinate_Source).
16 UTM_North_Coordinate Coordenada_Norte_UTM 4758 Northing as stated in, or
inferred from, the source
(see Coordinate_Source).
17 UTM_Zone Zona_UTM 4749 UTM zone of the source
coordinates; letters are
latitude bands or
hemisphere (e.g. '18',
'18 K', '18 Sur', '18K',
'18L', '18S').
18 Coordinate_System Sistema_Coordenadas 4763 Coordinate reference
system stated in the
source or an EPSG code.
See Section 12 for the
SNIRH records.
19 Decimal_Latitude Latitud_Decimal 503 Latitude as stated in, or
inferred from, the
source; 453 of the 503
values are inferred
(Coordinate_Source
'Inferida').
20 Decimal_Longitude Longitud_Decimal 503 Longitude as stated in,
or inferred from, the
source.
21 Coordinate_Source Fuente_Coordenadas 4775 How the location was
obtained: 'Explícita'
(stated), 'Extraída de
reporte' (from report
text), 'Inferida' (from a
map or site description),
'No habia informacion'
(none).
22 Coordinate_Precision Precision_Coordenadas 4763 Qualitative positional
confidence: 'Alta',
'Media', 'Baja'.
23 Elevation_masl Cota_msnm 2690 Ground elevation (m
a.s.l.) as reported.
24 Total_Depth Profundidad_Total 716 Total depth of the point:
well or borehole depth,
but also pit, sampling or
investigation depth; see
Point_Type and
Total_Depth_Unit.
25 Total_Depth_Unit Profundidad_Total_Unidad 751 Unit of Total_Depth as
printed.
26 Water_Table Nivel_Freatico 3305 Depth to water below
ground (negative = above
ground, 1 record); see
Water_Table_Unit. Some
values are not measured
groundwater levels
(Section 9).
27 Water_Table_Unit Nivel_Freatico_Unidad 3331 Unit of Water_Table as
printed.
28 Water_Table_masl Nivel_Freatico_msnm 324 Water-table elevation as
reported.
29 Water_Table_masl_Unit Nivel_Freatico_msnm_Unidad 442 Unit of Water_Table_masl.
30 Level_Status Estado_Nivel 3513 Condition of the level as
reported, e.g.
'Estático', 'Surgente',
'Seco'. 1805 values carry
a leading space ('
Estático'); trim before
grouping.
31 Saturated_Thickness Espesor_Saturado 335 Saturated thickness b;
see
Saturated_Thickness_Unit.
32 Saturated_Thickness_Unit Espesor_Saturado_Unidad 355 Unit of
Saturated_Thickness as
printed.
33 Hydraulic_Conductivity_K Conductividad_Hidraulica_K 224 Hydraulic conductivity K
as printed. Convert with
K_Unit and
unit_conversion_table.csv;
filter on Test_Type.
34 K_Unit Unidad_K 224 Unit of K as printed (15
different strings).
35 Transmissivity_T Transmisividad_T 58 Transmissivity T as
printed. Convert with
T_Unit and
unit_conversion_table.csv.
36 T_Unit Unidad_T 58 Unit of T as printed (9
different strings).
37 Storage_Coefficient_S Coeficiente_Almacenamiento_S 27 Storage coefficient S
(dimensionless).
38 Specific_Yield_Sy Rendimiento_Especifico_Sy 0 Specific yield Sy. Empty
in this release.
39 Test_Type Tipo_Ensayo 1131 Measurement method as
reported, mostly in
Spanish (227 distinct
values).
test_type_vocabulary.csv
classifies the values
attached to K, T or S.
40 Calculation_Method Metodo_Calculo 368 Analytical solution or
method used by the source
(e.g. Theis, Jacob,
Lefranc).
41 Lithology Litologia 1280 Lithological description
as printed, in Spanish.
42 Aquifer_Type Tipo_Acuifero 1018 Aquifer condition as
reported, e.g. 'Libre',
'Fisurado', 'Fracturado'.
43 Water_Quality_Parameters Parametros_Calidad_Agua 1095 Water-quality values as
text. Of 1095 values, 506
are valid JSON, 117
become JSON after
replacing doubled quotes
("") with single ones,
and the rest are plain
'key: value' text. Keys
and units are not
harmonised.
44 Page_Reference Pagina_Referencia 2960 Page(s) of the source on
which the values appear,
as recorded by the
extraction ('Página 10',
'54, 115', or a file name
with a page). Empty for
1815 records (Section
10).
45 Observations Observaciones 2894 Free-text notes carried
over from the source (in
Spanish).
46 Processing_Date Fecha_Procesamiento 1979 Date the record was
processed by the
extraction workflow.
47 Source_Folder Carpeta_Origen 1979 Processing provenance:
harvest folder of the
intermediate file on the
authors' workstation
(e.g. thesis_UNSA,
repositorio_ana). Not a
file of this resource.
48 Source_File Archivo_Origen 1979 Processing provenance:
intermediate extraction
file on the authors'
workstation. Not a file
of this resource.
49 Consolidation_Date Fecha_Consolidacion 1979 Date the record entered
the consolidated
database.
50 Document_Part Parte_Documento 982 Part of a multi-part
source document the value
came from.
51 Latitude_WGS84 Latitud_WGS84 4763 Harmonised latitude, WGS
84 decimal degrees. Use
these coordinates (see
Section 12 for SNIRH).
52 Longitude_WGS84 Longitud_WGS84 4763 Harmonised longitude, WGS
84 decimal degrees.
53 Conversion_Source Fuente_Conversion 4775 Basis of the
harmonisation, e.g.
'Coordenadas decimales
existentes', 'UTM WGS84
convertido'.
54 Conversion_Quality Calidad_Conversion 4775 Confidence in the
harmonisation: 'Alta',
'Media'.
55 BASIN_NAME NOMBRE_CUENCA 4775 Hydrographic unit of the
ANA Pfafstetter
delineation; also the
worksheet name in the
per-basin workbook.
'Cuenca Honda' groups two
codes; see BASIN_CODE.
56 BASIN_CODE CODIGO_CUENCA 4775 Pfafstetter code of the
unit (24 codes; 'Cuenca
Honda' holds 137158 (15)
and 13178 (1)).
3. arequipa_final_gw_database_per_basin.xlsx
----------------------------------------------------------------------------------------------------
Excel workbook with 24 worksheets. 23 are named after a hydrographic unit (the BASIN_NAME value) and
hold that unit's records in the 56 columns of the CSV, with English names in row 1 and Spanish names
in row 2. Every CSV record appears exactly once. The sheet 'Resumen' gives the number and share of
records per unit (columns Cuenca, Registros, % del total, and a TOTAL row).
Python: sheets = pd.read_excel('arequipa_final_gw_database_per_basin.xlsx', sheet_name=None,
skiprows=[1], dtype=str)
df = pd.concat([d for name, d in sheets.items() if name != 'Resumen'], ignore_index=True)
The workbook was produced separately from the CSV and differs from it in formatting, not in content:
Measurement_Date is written YYYY-MM-DD in 3751 records and M/D/YYYY in 981; 384 text cells contain
control characters where the CSV has an en dash or another typographic character (for example 'P- BU
\x96 1' for 'P- BU – 1'); and numbers are stored as numbers, so leading or trailing zeros as printed
are not kept; Processing_Date and Consolidation_Date are written YYYY-MM-DD hh:mm:ss instead of
M/D/YYYY h:mm. Use the CSV for analysis and the workbook for browsing.
Worksheet (hydrographic unit) Records
Cuenca Quilca - Vitor - Chili 3398
Cuenca Camaná 289
Cuenca Acarí 208
Intercuenca 13719 189
Cuenca Tambo 150
Cuenca Yauca 111
Cuenca Ocoña 106
Intercuenca 1319 70
Intercuenca 135 63
Cuenca Chala 54
Intercuenca 133 25
Intercuenca 13179 20
Intercuenca Alto Apurímac 19
Cuenca Pescadores - Caraveli 17
Cuenca Honda 16
Intercuenca 137157 12
Intercuenca 13717 7
Cuenca Cháparra 6
Intercuenca 137155 6
Intercuenca 137159 3
Cuenca Choclón 3
Intercuenca 13713 2
Intercuenca 13177 1
Total 4775
4. database.7z.001 and database.7z.002
----------------------------------------------------------------------------------------------------
How to open: download both volumes into one folder and extract database.7z.001 with a program that
reads split 7-Zip archives (7-Zip on Windows; 7-Zip, Keka or p7zip on macOS and Linux). The archive
is solid, so any extraction reads all of it.
Content: a folder 'database' holding a collection log and 604 subfolders, one per harvested item.
355 are named after the title and the ANA repository handle (20.500.12543/NNNN); 116 after a
university-repository identifier (UUID; 88 of these folder names end in '.pdf' although they are
folders); 118 after a short identifier; and 15 after other names, including INGEMMET file names. The
folders hold 647 PDF files, some of them duplicate copies, and 477 harvest-metadata files named
_info.txt, plus four ArcGIS map packages (.mpk), two zipped ArcGIS maps, two images and two files
without extension. 112 folders contain only _info.txt: the document itself was not retrieved. Most
archived documents contribute no records; they were screened during harvesting and held no usable
measurements.
Coverage: the archive holds the source document for 64 of the 303 source files that contribute
records, i.e. for 584 of the 3844 documentary records. The other 239 files (3260 records) are not
archived; source_catalog.csv (columns availability and how_to_obtain) says where each can be
obtained:
not in a public repository 12 files 1911 records
university repository 182 files 931 records
ANA repository 42 files 408 records
journal article (open access) 2 files 9 records
INGEMMET repository 1 file 1 record
The two consultancy files are groundwater-monitoring compilations for the Chili aquifer by TGC
Ingeniería y Servicios Generales EIRL: 'Informe final_Victor Camacho_ok.pdf' (981 records, reports
of 2014-2021) and 'Consolidado datos monitoreo Chili 2025.pdf' (884 records, 2021-2025). They are
not in a public repository and were compiled directly rather than through the text-extraction
workflow (their Processing_Date and Source_File are empty).
5. DATA_SOURCES_AND_TRACEABILITY.txt
----------------------------------------------------------------------------------------------------
Explains the traceability columns, gives two worked examples (one archived, one not) and lists the
repositories for documents that are not archived. Corrected in this version (Section 13).
6. source_catalog.csv (UTF-8 with byte-order mark; in R use fileEncoding = 'UTF-8-BOM')
----------------------------------------------------------------------------------------------------
One row per value of Full_Source. Columns: Full_Source (join key), stream, origin_folder,
Document_Title (most frequent title; n_distinct_titles counts variants, which occur for 2 files),
Authors, Publication_Year (a range for multi-year compilations), n_records, n_depth_to_water, n_K,
n_T, n_S, n_without_page_reference, in_database_7z (yes / no / not applicable), availability,
match_rule, path_in_database_7z (a PDF file), other_pdfs_in_same_folder, how_to_obtain. A few
documents appear under more than one file name (for example the ANA 2018 Chili aquifer study), so
the number of distinct documents is slightly below the number of rows.
7. unit_conversion_table.csv (UTF-8 with byte-order mark)
----------------------------------------------------------------------------------------------------
One row per unit string and variable. Columns: variable, value_column, unit_column, unit_as_printed,
n_records, multiply_by_to_get, reference_unit (m/d for K, m2/d for T, m for lengths), note. Multiply
the stored value by multiply_by_to_get. The note column explains five special cases: Lugeon values
describe rock-mass permeability and the factor is nominal; one T value is in l/s, a discharge, with
no valid conversion; length values labelled cm or mm are already in metres (factor 1); three
depth-to-water values with no unit are qualitative zeros; and two are measured below the top of the
casing (mbtoc). Negative depths to water are levels above ground.
8. test_type_vocabulary.csv (UTF-8 with byte-order mark)
----------------------------------------------------------------------------------------------------
One row per Test_Type value attached to K, T or S: 47 of the 227 distinct Test_Type values, plus one
row for records with no Test_Type (48 rows). Columns: Test_Type_as_printed, n_records_with_K_T_or_S,
n_records_with_K, measurement_support, use_for_aquifer_K, note. Classes: 'pumping or recovery test
(saturated zone)', 'near-surface infiltration test', 'Lefranc borehole test', 'packer (Lugeon) test
in rock', 'auger-hole (Van Beers) test', 'laboratory test', 'not a measurement', 'unspecified'.
Values were classified by name and, where the name was ambiguous, by Calculation_Method; the note
column records these cases.
K values by class: near-surface infiltration test 81; pumping or recovery test (saturated zone) 57;
packer (Lugeon) test in rock 26; Lefranc borehole test 20; laboratory test 14; unspecified 12;
auger-hole (Van Beers) test 9; not a measurement 5. The data description paper (Sect. 3.5) used
seven classes built from keywords (infiltrometer 83, pumping 55, packer 25, laboratory 16, Lefranc
13, auger-hole 9, unspecified 23). The vocabulary reclassifies ambiguous names by their
Calculation_Method, which moves a recovery test and a constant-rate test into the pumping class and
several borehole tests into the Lefranc class, and it separates non-measurements.
The 57 pumping or recovery K records are not 57 tests. The ANA (2018) Chili aquifer study is stored
under several file names, so 8 records repeat tests already present (records 1935, 1937, 1942, 1943,
1944, 1945, 1962, 1964), often with different coordinates for the same well, and record 1961 is a
miscalculated copy (flagged). Counting each test once leaves 48 tests with a median of 16.4 m/d.
9. flagged_records.csv (UTF-8 with byte-order mark)
----------------------------------------------------------------------------------------------------
One row per flagged value (Record_ID and variable), with the reason and the rule that raised it.
Flags cover K values that are not measurements of aquifer material (29: literature values,
estimates, range limits, a design criterion, a model output, a laboratory criterion, a test on
concrete, infiltration rates of permeable pavements cited from other studies, a miscalculated copy
and two tests of uncertain method); a T value in l/s and a literature T; storage coefficients that
are estimated or physically implausible (about 1e-10); depths to water that are not measured
groundwater levels (surface-water points, geophysical interpretations, estimated, interpreted or
design levels, and pits where the water table was not reached); thicknesses interpreted from
geophysics; depths at points that are not groundwater points; and length values whose unit string
says cm or mm although they are in metres. Values by variable: K 29, S 6, T 2, depth to water 131,
saturated thickness 52, total depth 74. The rules are explicit; the list is not exhaustive.
10. HOW TO TRACE A RECORD TO ITS SOURCE
----------------------------------------------------------------------------------------------------
1. Find the record in the CSV and note Full_Source and Page_Reference.
2. Look up Full_Source in source_catalog.csv.
3. If in_database_7z is 'yes', extract the archive and open path_in_database_7z at the cited page.
4. Otherwise follow how_to_obtain (an ANA handle, a repository to search by Document_Title, or a
contact for documents not in a public repository).
5. For SNIRH-ANA records, query the station (Original_Point_ID) in https://snirh.ana.gob.pe.
Page_Reference is the page recorded by the extraction; it may refer to the page of the PDF file or
to the printed page number. It is empty for 1815 records: all 931 SNIRH records and the 884 records
of 'Consolidado datos monitoreo Chili 2025.pdf', which can be traced to their document but not to a
page.
11. COVERAGE
----------------------------------------------------------------------------------------------------
Space: 4763 records have WGS 84 coordinates, spanning 14.55 to 17.22 degrees S and 70.07 to 74.85
degrees W across 23 hydrographic units centred on the Arequipa region. 12 records have no
coordinates (all from 'Consolidado datos monitoreo Chili 2025.pdf'); they are assigned to 'Cuenca
Quilca - Vitor - Chili' from their document. The bounding box of the first published version (14.88
to 16.50 S, 71.30 to 73.30 W) covered the core area only; 1246 records lie outside it.
Time: measurement dates run from 1966 to 14 December 2025 (193 records carry that last date), except
5 records: 2 dated 1913, the date of a historical analysis cited in a 1968 document, and 3 dated
1964, the publication year of their source. SNIRH-ANA records span 11 May 1998 to 24 May 2023.
Variables (records with a value): depth to water 3305; ground elevation 2690; total well depth 716;
water-table elevation 324; saturated thickness 335; hydraulic conductivity 224; transmissivity 58;
storage coefficient 27; specific yield 0.
12. KNOWN LIMITATIONS
----------------------------------------------------------------------------------------------------
* Units are not harmonised in the stored values (Section 7), and K mixes measurement supports
(Section 8).
* 29 K values and a few other values are not field measurements (Section 9).
* Positional uncertainty is uneven: coordinates are explicit for 1668 records, taken from report
text for 1853, inferred from maps or site descriptions for 1242, and absent for 12. Use
Coordinate_Source and Coordinate_Precision.
* The 931 SNIRH records carry Coordinate_System '24879' (EPSG:24879, PSAD56 / UTM zone 19S), but
their WGS 84 coordinates were computed as if the UTM values were WGS 84, and 309 of them are in zone
18. If the source values are PSAD56, the true positions differ by about 400 m.
* Water_Quality_Parameters is only partly JSON, with keys and units that are not harmonised (Section
2).
* 656 records hold no groundwater or water-quality value at all: they are locations such as survey
control points, rock and soil samples, weather and surface-water stations, quarries and some wells
without data. Filter on Point_Type or on the value columns.
* In the 1981 Acarí inventory (ANA0000209), 22 of 38 depths to water are exactly 0 while the
saturated thickness equals the well depth; these may be missing values coded as 0. Check the source
before using them.
* The ANA (2018) Chili aquifer study appears under several file names, and some of its wells carry
different coordinates in different records (Section 8).
* Only a fraction of source documents is in the archive (Section 4), and some records have no page
reference (Section 10).
* Repeat measurements at the same point are rare and concentrated in the Quilca-Vitor-Chili basin,
so trend analysis is not supported elsewhere. Spatial density reflects where studies were done, not
hydrogeological importance.
13. WHAT CHANGED FROM THE NOVEMBER 2025 README
----------------------------------------------------------------------------------------------------
* All files are now listed and described: the six files of the first published version and four new
companion tables. The previous README listed only the two archive volumes and itself, did not
describe the per-basin workbook or the traceability guide, and stated wrongly that
arequipa_final_gw_database.csv is inside the archive.
* The archive is described correctly as a collection of source documents; it contains no database
file.
* The column dictionary uses the column names actually in the CSV (English on line 1, Spanish on
line 2). The previous dictionary listed Spanish names only and a column 'CUENCA' that does not exist
(the unit is BASIN_NAME), and gave the date format as YYYY-MM-DD, whereas the CSV uses M/D/YYYY.
* Record counts are stated consistently: 4775 records in total, of which 931 are SNIRH-ANA records.
* The CSV encoding, the second header row, the date placeholders and the differences of the workbook
are documented.
* The worked example of the traceability guide now reports Record 1 correctly: its depth to water is
0.95 m; the 8.85 m given before is its saturated thickness.
* No data value was changed. The companion tables were built from the CSV, the archive listing and a
review of the Observations column.
14. CITATION
----------------------------------------------------------------------------------------------------
Venegas-Quiñones, H. L., Guillen, M., Garcia-Chevesich, P. A., Uhle, B., González, E., Ticona, J.,
Díaz, J., Zea, J., Alejo, F., and McCray, J.: Regional Groundwater Database from Unstructured
Sources Based on an Integrated Optical Character Recognition and Large Language Model Workflow for
Arequipa and Surroundings in Southern Peru, HydroShare [data set],
https://www.hydroshare.org/resource/0f9223cb27584d00ba46b87a5af471ef/, 2026.
Related Resources
| This resource updates and replaces a previous version | Venegas-Quiñones, H. L., Guillen, M., Garcia-Chevesich, P. A., Uhle, B., González, E., Ticona, J., Díaz, J., Zea, J., Alejo, F., McCray, J. (2026). Regional Groundwater Database from Unstructured Sources Based on an Integrated Optical Character Recognition and Large Language Model Workflow for Arequipa and Surroundings in Southern Peru, HydroShare, http://www.hydroshare.org/resource/1dfbdbe97e6d48d7b307591868ca70c4 |
Credits
Funding Agencies
This resource was created using funding from the following sources:
| Agency Name | Award Title | Award Number |
|---|---|---|
| The Center for Mining Sustainability | None | 470266 |
How to Cite
This resource is shared under the Creative Commons Attribution CC BY.
http://creativecommons.org/licenses/by/4.0/
Comments
There are currently no comments
New Comment