A Corpus with Its Reasons Attached: The Qur'anic Structural Research Corpus — Architecture, Coordinate System, and Provenance Model of a Multi-Layer Reproducibility Resource

Claims about the structure of the Qur'an are abundant and the evidence for them is rarely inspectable. A reader is told that a passage opens a movement, that two surahs are paired, that a group of surahs forms a unit, that a phrase returns at a structurally significant point — and is given a conclusion without the record that would allow the conclusion to be checked, qualified, or refuted. The claim travels; the basis for it does not. This paper documents a corpus built to the opposite specification. The Qur'anic Structural Research Corpus comprises ten independently deposited reproducibility datasets covering 113 CSV files and 35,932 data records, annotating the structure of the received order at six research scales: surah, passage, ayah, adjacency, boundary, and corpus. Its distinguishing feature is not its size but what each record carries. Every substantive row is accompanied by a confidence grade and an evidence note; competing readings are recorded alongside adopted ones rather than discarded; and structural claims are tested against negative controls, including alternative segmentations and control pairs, so that a pattern appearing where the hypothesis says it should not can be identified as such. The paper describes the corpus as a single resource for the first time. It specifies the received-order coordinate system that makes the ten layers interoperable, consisting of eleven canonical coordinates expressed across 94 distinct field names; it reports a validated join map of 90 dataset-pair joins covering all 45 possible pairs, so that the corpus is a fully connected graph; and it documents a provenance model of 46 fields serving eight functions. Of 495 distinct column names in the corpus, 140 — more than one in four — carry coordinates or provenance rather than substantive content. Three findings emerge from treating the ten deposits as one object. The coordinate system is heavily redundant in naming, with a single coordinate expressed by up to 25 different field names, so that interoperability depends on an alias layer rather than on uniform naming. Connectivity is uneven: surah-level joins connect every pair at a resolution of one part in 114, while the finest coordinates connect only a handful of layers, which sets a present ceiling on what the corpus can answer. And only a quarter of coordinate-bearing fields identify a row's own position, the remainder referring to positions relative to a research object, which makes naive joining a predictable source of error — demonstrated in a worked join that shows how a complete, null-free result table can silently answer a different question. The paper reports the deposit chronology that explains the corpus's naming drift, states what the resource does not support, and labels its most contested component — a provisional seven-group macro-structural model carried by seven of the ten datasets — as a hypothesis offered for refutation, with the author's own boundary-testing and alternative-segmentation data provided for that purpose. Limitations are stated plainly, including that all coding is by a single coder and that no inter-coder reliability figure exists for any field. 38 pages, 10 figures, 21 tables, five appendices. Every statistic reported is computed from the deposited files and reproducible from them; Appendix E documents the method for each. The corpus is released under CC BY 4.0 as Figshare Collection 10.6084/m9.figshare.c.8755176, with a documentation layer at 10.6084/m9.figshare.34075545 providing the coordinate specification, field alias registry, file-level crosswalk, validated join map, and provenance model.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-06
DOI
https://doi.org/10.5281/zenodo.23179646
Primary Topic
Digital Humanities and Scholarship
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

A Corpus with Its Reasons Attached: The Qur'anic Structural Research Corpus — Architecture, Coordinate System, and Provenance Model of a Multi-Layer Reproducibility Resource

Syed Raheel Shahzad
Zenodo (CERN European Organization for Nuclear Research)
Digital Humanities and Scholarship
preprint

A Corpus with Its Reasons Attached: The Qur'anic Structural Research Corpus — Architecture, Coordinate System, and Provenance Model of a Multi-Layer Reproducibility Resource

Syed Raheel Shahzad
preprint en

Abstract

Claims about the structure of the Qur'an are abundant and the evidence for them is rarely inspectable. A reader is told that a passage opens a movement, that two surahs are paired, that a group of surahs forms a unit, that a phrase returns at a structurally significant point — and is given a conclusion without the record that would allow the conclusion to be checked, qualified, or refuted. The claim travels; the basis for it does not. This paper documents a corpus built to the opposite specification. The Qur'anic Structural Research Corpus comprises ten independently deposited reproducibility datasets covering 113 CSV files and 35,932 data records, annotating the structure of the received order at six research scales: surah, passage, ayah, adjacency, boundary, and corpus. Its distinguishing feature is not its size but what each record carries. Every substantive row is accompanied by a confidence grade and an evidence note; competing readings are recorded alongside adopted ones rather than discarded; and structural claims are tested against negative controls, including alternative segmentations and control pairs, so that a pattern appearing where the hypothesis says it should not can be identified as such. The paper describes the corpus as a single resource for the first time. It specifies the received-order coordinate system that makes the ten layers interoperable, consisting of eleven canonical coordinates expressed across 94 distinct field names; it reports a validated join map of 90 dataset-pair joins covering all 45 possible pairs, so that the corpus is a fully connected graph; and it documents a provenance model of 46 fields serving eight functions. Of 495 distinct column names in the corpus, 140 — more than one in four — carry coordinates or provenance rather than substantive content. Three findings emerge from treating the ten deposits as one object. The coordinate system is heavily redundant in naming, with a single coordinate expressed by up to 25 different field names, so that interoperability depends on an alias layer rather than on uniform naming. Connectivity is uneven: surah-level joins connect every pair at a resolution of one part in 114, while the finest coordinates connect only a handful of layers, which sets a present ceiling on what the corpus can answer. And only a quarter of coordinate-bearing fields identify a row's own position, the remainder referring to positions relative to a research object, which makes naive joining a predictable source of error — demonstrated in a worked join that shows how a complete, null-free result table can silently answer a different question. The paper reports the deposit chronology that explains the corpus's naming drift, states what the resource does not support, and labels its most contested component — a provisional seven-group macro-structural model carried by seven of the ten datasets — as a hypothesis offered for refutation, with the author's own boundary-testing and alternative-segmentation data provided for that purpose. Limitations are stated plainly, including that all coding is by a single coder and that no inter-coder reliability figure exists for any field. 38 pages, 10 figures, 21 tables, five appendices. Every statistic reported is computed from the deposited files and reproducible from them; Appendix E documents the method for each. The corpus is released under CC BY 4.0 as Figshare Collection 10.6084/m9.figshare.c.8755176, with a documentation layer at 10.6084/m9.figshare.34075545 providing the coordinate specification, field alias registry, file-level crosswalk, validated join map, and provenance model.

Zenodo (CERN European Organization for Nuclear Research)
Digital Humanities and Scholarship
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.