What a Retrieval Index Loses Before Anyone Notices: Three Silent Defects in Three Published Wikipedia Embedding Indexes

We re-read three Wikipedia retrieval indexes we had published ourselves and had been using in production — English, Japanese and Chinese, FAISS over intfloat/multilingual-e5-large — and found three defects, none of which had ever produced an error message. Chunks were cut by character count in front of a token window, so 13.7 % of the published Japanese chunks and 18.5 % of the Chinese ones ran past the encoder's 512-token limit and lost 21.7 % and 27.0 % of their tokens at embed time. The Kiwix licence footer, 184 characters of English legal text, sat inside 97.1 % of the published Japanese chunks and 92.9 % of the Chinese ones. The index type, IndexIVFPQ(m=64), returned 45.4–47.1 % of the exact top-10 and cost 7.8 MRR in Japanese and 10.7 in English against exact search. Rebuilding the whole pipeline moved retrieval against the previously published artifact from 26.6 to 43.3 MRR in Japanese and 26.2 to 47.5 in Chinese, on a control question set generated from the old extractor's own output. The components are not separated, and the rebuild is not uniformly better: indexed-article coverage falls in two of three languages. Two results may generalise. Our own first benchmark ranked the defective extractor higher, because it queried by article title, a field the index also stores. And clipping the scalar quantiser's range to the 0.1–99.9 percentile of the residuals rather than their min/max raises top-10 agreement with exact search from 81.6 % to 85.2 % at the same 4 bits per dimension; the first attempt computed the quantiles on raw vectors and made the index substantially worse, because an IVF index encodes residuals. Three rounds of adversarial review are reported at equal prominence with the results. Round one rejected the first rebuild. Round two found sixteen unsupported statements in the draft. Round three found a defect in the release that the first two had both passed — the acceptance script had not been re-run on two of the three shipped indexes after they were re-encoded — and three measurements that round two had itself introduced while correcting other numbers. Scope. This reports what happened to three specific artifacts. It does not claim these defects are common elsewhere. The retrieval numbers come from synthetic questions generated by one local model with no human relevance judgements and no answerability check; the reconstructed standard error of one MRR figure is 2.2–2.9 points. Index-type comparisons were run on pools of 150,000–206,000 vectors, not at the shipped scale of 3–15.6 million. Nothing here is peer reviewed, certified, or independently replicated; every reviewer was a model instance commissioned by the author. AI co-observer disclosure. The forensic measurements, the rebuild, the three review rounds and the drafting were carried out with Claude Opus 5 (Anthropic) as a working instrument under human direction. The registered author is the human author alone.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-16
DOI
https://doi.org/10.5281/zenodo.22782085
Primary Topic
Wikis in Education and Collaboration
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

What a Retrieval Index Loses Before Anyone Notices: Three Silent Defects in Three Published Wikipedia Embedding Indexes

Toeda Taiko
Zenodo (CERN European Organization for Nuclear Research)
Wikis in Education and Collaboration
article

What a Retrieval Index Loses Before Anyone Notices: Three Silent Defects in Three Published Wikipedia Embedding Indexes

Toeda Taiko
article en

Abstract

We re-read three Wikipedia retrieval indexes we had published ourselves and had been using in production — English, Japanese and Chinese, FAISS over intfloat/multilingual-e5-large — and found three defects, none of which had ever produced an error message. Chunks were cut by character count in front of a token window, so 13.7 % of the published Japanese chunks and 18.5 % of the Chinese ones ran past the encoder's 512-token limit and lost 21.7 % and 27.0 % of their tokens at embed time. The Kiwix licence footer, 184 characters of English legal text, sat inside 97.1 % of the published Japanese chunks and 92.9 % of the Chinese ones. The index type, IndexIVFPQ(m=64), returned 45.4–47.1 % of the exact top-10 and cost 7.8 MRR in Japanese and 10.7 in English against exact search. Rebuilding the whole pipeline moved retrieval against the previously published artifact from 26.6 to 43.3 MRR in Japanese and 26.2 to 47.5 in Chinese, on a control question set generated from the old extractor's own output. The components are not separated, and the rebuild is not uniformly better: indexed-article coverage falls in two of three languages. Two results may generalise. Our own first benchmark ranked the defective extractor higher, because it queried by article title, a field the index also stores. And clipping the scalar quantiser's range to the 0.1–99.9 percentile of the residuals rather than their min/max raises top-10 agreement with exact search from 81.6 % to 85.2 % at the same 4 bits per dimension; the first attempt computed the quantiles on raw vectors and made the index substantially worse, because an IVF index encodes residuals. Three rounds of adversarial review are reported at equal prominence with the results. Round one rejected the first rebuild. Round two found sixteen unsupported statements in the draft. Round three found a defect in the release that the first two had both passed — the acceptance script had not been re-run on two of the three shipped indexes after they were re-encoded — and three measurements that round two had itself introduced while correcting other numbers. Scope. This reports what happened to three specific artifacts. It does not claim these defects are common elsewhere. The retrieval numbers come from synthetic questions generated by one local model with no human relevance judgements and no answerability check; the reconstructed standard error of one MRR figure is 2.2–2.9 points. Index-type comparisons were run on pools of 150,000–206,000 vectors, not at the shipped scale of 3–15.6 million. Nothing here is peer reviewed, certified, or independently replicated; every reviewer was a model instance commissioned by the author. AI co-observer disclosure. The forensic measurements, the rebuild, the three review rounds and the drafting were carried out with Claude Opus 5 (Anthropic) as a working instrument under human direction. The registered author is the human author alone.

Zenodo (CERN European Organization for Nuclear Research)
Yulius (NL)
Openalex Percentile: Top 4%
Wikis in Education and Collaboration
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.