Synthetic Worldview Reconstruction (SWR): A Methodological Framework for LLM-Assisted Group-Level Belief System Analysis
Abstract (English) This paper presents Synthetic Worldview Reconstruction (SWR)—a method for reconstructing the collective belief systems of sociologically defined groups using large language models. SWR formulates worldview analysis as an inverse problem: given the aggregated public statements and documented actions of a group, it reconstructs the latent belief system that most coherently accounts for those outputs. The method constructs a fictive synthetic person—a Weberian ideal type—whose worldview represents the group's collective beliefs in distilled form. This paper abstracts SWR from its original application (a study of 100 influential AI actors; Geiger, 2026) into a transferable methodological framework. We describe the five-step protocol, the dimensional rating procedure as one operationalization example, and the focus control infrastructure. The pilot records five group profiles with two repetitions each: the historical mean of one-way coefficients and a separate two-way absolute-agreement calculation both round to 0.902; the latter is a descriptive profile index without a bound population uncertainty model. A separate N=15 protocol reports ICC(3,1)=0.847 for single-measure consistency and ICC(2,k)=0.988 for agreement of its 15-run mean. Cross-modal prediction records 72% confirmed and 0% contradicted in that pilot; the tabulated placeholder-versus-fictional-name profiles yield Pearson r=0.9585 and MAE=0.33. These local observations do not establish external construct validity or general equivalence of blinding strategies. Tiered acceptance criteria distinguish between minimum requirements (Tier 1) and desirable extensions (Tier 2). A step-by-step application guide enables researchers to apply SWR to their own populations. The method is designed for group-level analysis, where ready-made collective worldviews of sociological groups are less likely to exist in the training corpus than public knowledge about prominent individuals, thereby reducing a contamination risk that is typically more acute in individual-level applications. Zusammenfassung (Deutsch) Dieser Beitrag präsentiert Synthetic Worldview Reconstruction (SWR) – eine Methode zur Rekonstruktion der kollektiven Glaubenssysteme soziologisch definierter Gruppen mittels großer Sprachmodelle. SWR formuliert Weltanschauungsanalyse als inverses Problem: Aus den aggregierten öffentlichen Äußerungen und dokumentierten Handlungen einer Gruppe wird das latente Glaubenssystem rekonstruiert, das diese beobachtbaren Outputs am kohärentesten erzeugt. Die Methode konstruiert eine fiktive Syntheseperson – einen Weberschen Idealtypus – deren Weltanschauung die kollektiven Überzeugungen der Gruppe in destillierter Form repräsentiert. Der Beitrag abstrahiert SWR von seiner ursprünglichen Anwendung (einer Studie über 100 einflussreiche KI-Akteure; Geiger, 2026) zu einem übertragbaren methodologischen Rahmenwerk. Wir beschreiben das Fünfschritte-Protokoll, das Verfahren zur dimensionalen Bewertung als ein Operationalisierungsbeispiel und die Infrastruktur der Fokuskontrolle. Die Pilotdaten umfassen fünf Gruppenprofile mit je zwei Wiederholungen: Das historische Mittel von One-way-Koeffizienten und eine separate Zweiweg-Rechnung für absolute Übereinstimmung ergeben beide gerundet 0,902; Letztere ist ein deskriptiver Profilindex ohne gebundenes Unsicherheitsmodell für eine Objektpopulation. Ein getrenntes N=15-Protokoll berichtet ICC(3,1)=0,847 für Einzelmessungs-Konsistenz und ICC(2,k)=0,988 für die Übereinstimmung des Mittels aus 15 Läufen. Die modalitätsübergreifende Vorhersage verzeichnet in dieser Pilotstudie 72% bestätigt und 0% widerlegt; die tabellierten Profile mit Platzhaltern und fiktiven Namen ergeben Pearson r=0,9585 und MAE=0,33. Diese lokalen Beobachtungen belegen weder externe Konstruktvalidität noch eine allgemeine Äquivalenz der Verblindungsstrategien. Abgestufte Akzeptanzkriterien unterscheiden zwischen Mindestanforderungen (Stufe 1) und wünschenswerten Erweiterungen (Stufe 2). Ein schrittweiser Leitfaden ermöglicht es Forschenden, SWR auf eigene Populationen anzuwenden. Die Methode ist für die Analyse auf Gruppenebene konzipiert, wobei vorgefertigte kollektive Weltanschauungen soziologischer Gruppen im Trainingskorpus weniger wahrscheinlich sind als öffentlich verfügbares Wissen über prominente Einzelpersonen. Dadurch wird ein Vorwissensrisiko reduziert, das bei Ansätzen auf Individualebene typischerweise ausgeprägter ist. CHANGELOG Changes in Version 6.2 (3 October 2026) Pilot reliability clarified: Five fixed group profiles with two repetitions each yield two distinct descriptive calculations that both round to 0.902: the historical mean of one-way coefficients and a separately calculated two-way absolute-agreement index. Neither establishes population uncertainty or external construct validity. Blinding and say–do limits corrected: The tabulated placeholder-versus-fictional-name pair yields Pearson r = 0.9585 and MAE = 0.33; the older r = 0.987 text figure is not reproducible from that pair. Separately synthesized group and aggregate profiles do not prove a Simpson effect or cancellation. Method scope and files: The current English and German manuscripts distinguish the five-step protocol from the dimensional rating example and state the remaining validity limits. The public file set for this version contains only the English (26 pages) and German (28 pages) PDFs; the bilingual combination remains internal. Historical entries: Earlier changelog descriptions of validation and blinding are retained as version history; the corrected estimates and limitations in this version govern the current interpretation. Changes in 6.1 (Source and metadata maintenance) Source-check maintenance: Bibliographic metadata were checked against Crossref/DOI, arXiv, PMLR, and the Zenodo API. Corrections include Ornstein et al. 2025, Hornby 2025, Santurkar et al. 2023, Tai et al. 2024, Wang et al. 2025, and DOI/URL completions for the remaining checked entries. Method framing clarified: The public description now explicitly separates the SWR synthesis pipeline (Steps 1--4) from the dimensional rating procedure (Step 5), which is one operationalization example rather than a mandatory component. Build refresh: English, German, and combined bilingual PDFs were rebuilt after the source check. LaTeX logs contain no overfull boxes, undefined citations or references, rerun warnings, or LaTeX errors. Changes in 6.0 (Complete Rewrite) Complete rewrite: Paper B was rewritten from the ground up to resolve inconsistencies that had developed between the method paper and the companion application study (Paper A v7.0), and to address methodological divergences that had accumulated across earlier revision cycles. Streamlined from 32 to 21 pages (EN) / 22 pages (DE): Removed overengineered variable architecture and unexecuted operationalizations. Retained only what was actually implemented and validated. New Application Guide (Section 4): Step-by-step instructions including transfer sketch (climate scientists example), cost estimates, and common pitfalls. Tiered Acceptance Criteria: Tier 1 (instrument calibration -- required) vs. Tier 2 (model independence -- desirable for replication). Coherence bias expanded: RLHF mechanism, autoregressive tendency, premature closure, metacognitive insufficiency -- with control mechanisms and honest framing as instrument property. Circularity risk section: Same-model synthesis and validation discussed with instance separation as key mitigation. Reliability vs. validity distinction: IMIIRR explicitly framed as reliability, not validity. External expert validation remains an outstanding desideratum. Belief system terminology clarified: Three domains (self-image, view of humanity, worldview), anchored in Construal Level Theory (Trope and Liberman, 2010). Structural validation (PCA + HDBSCAN): Added to the quality assurance table. Additional references: Bourdieu (doxa), Gadamer (hermeneutics), Berger and Luckmann (social construction), Kommers et al. (computational hermeneutics), Hornby (Gadamer after ChatGPT), Wang et al. (LLM thematic analysis), and Barros et al. (LLM qualitative mapping). Changes in 5.0 (Consistency Patch) Aligned with Paper A v7.0. IMIIRR introduced. Instance separation established. Cross-modal prediction corrected from 88% to 72%. Narrative corrected. Changes in 4.0 Paradigm shift to group-level analysis. Self-contained methodology. Weber as primary anchor.
Authors
- Lukas Geiger (ORCID: https://orcid.org/0009-0005-7296-1534)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-03
- DOI
- https://doi.org/10.5281/zenodo.23121886
- Primary Topic
- Social Power and Status Dynamics
- Type
- article
- Field-Weighted Citation Impact
- 0.00