A Latent Map of Authorial Style: Training, Anatomy, and Applications to Dating, Idiolect Evolution, and Translation
We present a method for building a compact latent space of authorial style (z ∈ ℝ⁶⁴), trained on a corpus of literary prose in five languages (664 authors, 8,923 books). Unlike existing style embeddings, which are trained on short texts and fail to separate literary authors (top-1 accuracy 21–33%), the proposed space reaches 91.7% on leave-one-book-out author identification (Δ-separation 0.34 versus 0.02–0.10 for the competitors). Anatomy of the space reveals interpretable factors in the leading principal components — language, epoch, dialogism, gender direction, lexical richness — while the tail of components 8–64 carries an individual author signal (77% identification accuracy) that remains largely unidentified. Genre invariance is confirmed (η² = 0.001). Three applications demonstrate the practical value of the space: (1) style-based dating of texts with a median error of ±7 years for authors held out from the dating regressor's training; (2) idiolect evolution — each author drifts along an individual set of axes (58 of 64 engaged), not reducible to epoch or genre; (3) translation triangulation — the source author remains the dominant signal, but the authorial trace barely survives a change of language (1 of 43 translations keeps the original author in the top-20).
Authors
- Alexey Kravtsov (ORCID: https://orcid.org/0009-0004-4651-9948)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-28
- DOI
- https://doi.org/10.5281/zenodo.23019103
- Primary Topic
- Authorship Attribution and Profiling
- Type
- preprint