An LLM-Associated Register Shift in Korean Journal Abstracts: A Morphology-Aware Excess-Vocabulary Study, 2018-2026

{"Excess":[0],"vocabulary,":[1],"a":[2,70,96,119,206,259],"word's":[3],"frequency":[4],"above":[5],"its":[6,172,274],"pre-2023":[7],"trend,":[8],"is":[9,111,342],"how":[10],"the":[11,47,55,103,132,150,163,186,194,198,203,210,215,224,235,277,287,299,303,313,317,321,334,348,359,363,370,377,380,384,397],"change":[12],"in":[13,64,67,75,80,337],"scholarly":[14],"English":[15,201,211,295,356],"after":[16],"2022":[17],"has":[18],"been":[19],"measured.":[20],"We":[21],"adapt":[22],"it":[23,227],"to":[24,95,152,221],"Korean":[25,60,191,216,314,390],"with":[26,35,171,238,242,280,320],"morphological":[27],"units":[28],"on":[29,108,312],"398,296":[30],"KCI":[31],"abstracts":[32,38,61,84,110,202,230,315],"(2018–August":[33],"2026),":[34],"47,165":[36],"Vietnamese":[37],"for":[39,46,54,116,139,354],"comparison.":[40],"Placebo":[41],"floors":[42],"are":[43,247,400],"0.1–2.2":[44],"points":[45],"single-word":[48,104],"statistic":[49],"and":[50,114,118,125,137,166,273,282,285,294,309,324,332,357,367,373,383,393,403],"at":[51,219,401],"most":[52],"2.9":[53],"re-selected":[56],"split-half":[57,120,318],"set":[58,121,151],"statistic.":[59],"show":[62],"nothing":[63],"2023,":[65],"onset":[66],"late":[68],"2024,":[69],"rise":[71],"through":[72],"2025":[73],"flattening":[74],"mid-2026:":[76],"시사하다":[77],"\\"suggest\\"":[78],"appears":[79,205],"21.4%":[81],"of":[82,98,162,189,223,261,289,306],"2026":[83,169],"against":[85,376],"5.3%":[86],"expected;":[87],"plain":[88],"verbs":[89],"like":[90],"알아보다":[91],"\\"look":[92],"into\\"":[93],"fall":[94,192],"quarter":[97],"trend.":[99],"Under":[100],"stated":[101],"assumptions":[102],"conditional":[105],"lower":[106],"bound":[107,122],"LLM-processed":[109],"3.5%,":[112],"10.5%":[113],"16.1%":[115],"2024–2026":[117],"7.8%,":[123],"20.6%":[124],"33.0%.":[126],"Holzwarth":[127],"et":[128],"al.'s":[129],"estimator":[130,305],"under":[131,316],"same":[133,199,398],"discipline":[134],"gives":[135,276],"41.9%":[136],"72.1%":[138],"2025–2026.":[140],"Subject-matter":[141],"controls":[142],"reduce":[143],"but":[144],"do":[145,182],"not":[146,183],"remove":[147],"it:":[148,185],"restricting":[149],"lemmas":[153],"three":[154,232],"language-model":[155,262],"annotators":[156],"all":[157],"call":[158],"style":[159],"leaves":[160,177],"14.7":[161],"33.0":[164],"points,":[165],"pairing":[167],"each":[168],"abstract":[170,176],"journal's":[173],"closest":[174],"base-period":[175],"34.1.":[178],"Tested":[179],"translation":[180],"routes":[181],"explain":[184],"surface":[187],"marks":[188],"translated":[190],"as":[193],"markers":[195],"rise.":[196],"In":[197],"articles'":[200],"excess":[204],"year":[207],"earlier;":[208],"where":[209,226],"side":[212,327,329],"carries":[213,333],"none,":[214],"shift":[217],"persists":[218],"30":[220],"66%":[222],"rate":[225],"does.":[228],"Control":[229],"from":[231,396],"providers":[233],"reproduce":[234],"rising":[236],"words,":[237],"marker":[239],"turnover":[240],"consistent":[241],"model":[243,269],"generations;":[244],"implied":[245],"prevalences":[246],"scenario-dependent.":[248],"Working":[249],"paper,":[250],"version":[251,335],"8":[252,257],"(4":[253],"September":[254],"2026).":[255],"Version":[256,340],"adds":[258],"declaration":[260],"use":[263],"(Section":[264,330],"9)":[265],"that":[266,270],"names":[267],"every":[268],"took":[271],"part":[272],"role,":[275],"agent":[278],"pipeline":[279],"session":[281],"instruction":[283],"counts,":[284],"states":[286],"fractions":[288],"analysis":[290,349],"code,":[291,350],"build":[292],"scripts":[293,372],"text":[296],"produced":[297],"by":[298,328],"coding":[300],"agent;":[301],"reimplements":[302],"mixture":[304],"Holzwarth,":[307],"González-Márquez":[308],"Kobak":[310],"(2026)":[311],"discipline,":[319],"in-sample,":[322],"cross-fitted":[323],"placebo":[325],"values":[326],"5.15);":[331],"history":[336],"Appendix":[338],"J.":[339],"7":[341],"10.5281/zenodo.22110398.":[343],"The":[344],"reproducibility":[345],"package":[346],"contains":[347],"per-year":[351],"document-frequency":[352],"tables":[353],"Korean,":[355],"Vietnamese,":[358],"generated":[360],"control":[361],"abstracts,":[362],"annotation,":[364],"matching,":[365],"translation-route":[366],"robustness":[368],"results,":[369],"Holzwarth-estimator":[371],"their":[374],"validation":[375],"released":[378],"data,":[379],"results":[381],"manifest":[382],"manuscript":[385],"consistency":[386],"gate.":[387],"A":[388],"public":[389],"AI-style":[391],"dictionary":[392],"checker":[394],"built":[395],"measurements":[399],"https://os.intframe.com/report/ai-style-dictionary-ko":[402],"https://os.intframe.com/report/ai-style-check-ko.":[404]}

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-04
DOI
https://doi.org/10.5281/zenodo.22303588
Primary Topic
Authorship Attribution and Profiling
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

An LLM-Associated Register Shift in Korean Journal Abstracts: A Morphology-Aware Excess-Vocabulary Study, 2018-2026

Ahron Lee
Zenodo (CERN European Organization for Nuclear Research)
Authorship Attribution and Profiling
article

An LLM-Associated Register Shift in Korean Journal Abstracts: A Morphology-Aware Excess-Vocabulary Study, 2018-2026

Ahron Lee
article en

Abstract

Excess vocabulary, a word's frequency above its pre-2023 trend, is how the change in scholarly English after 2022 has been measured. We adapt it to Korean with morphological units on 398,296 KCI abstracts (2018–August 2026), with 47,165 Vietnamese abstracts for comparison. Placebo floors are 0.1–2.2 points for the single-word statistic and at most 2.9 for the re-selected split-half set statistic. Korean abstracts show nothing in 2023, onset in late 2024, a rise through 2025 flattening in mid-2026: 시사하다 "suggest" appears in 21.4% of 2026 abstracts against 5.3% expected; plain verbs like 알아보다 "look into" fall to a quarter of trend. Under stated assumptions the single-word conditional lower bound on LLM-processed abstracts is 3.5%, 10.5% and 16.1% for 2024–2026 and a split-half set bound 7.8%, 20.6% and 33.0%. Holzwarth et al.'s estimator under the same discipline gives 41.9% and 72.1% for 2025–2026. Subject-matter controls reduce but do not remove it: restricting the set to lemmas three language-model annotators all call style leaves 14.7 of the 33.0 points, and pairing each 2026 abstract with its journal's closest base-period abstract leaves 34.1. Tested translation routes do not explain it: the surface marks of translated Korean fall as the markers rise. In the same articles' English abstracts the excess appears a year earlier; where the English side carries none, the Korean shift persists at 30 to 66% of the rate where it does. Control abstracts from three providers reproduce the rising words, with marker turnover consistent with model generations; implied prevalences are scenario-dependent. Working paper, version 8 (4 September 2026). Version 8 adds a declaration of language-model use (Section 9) that names every model that took part and its role, gives the agent pipeline with session and instruction counts, and states the fractions of analysis code, build scripts and English text produced by the coding agent; reimplements the mixture estimator of Holzwarth, González-Márquez and Kobak (2026) on the Korean abstracts under the split-half discipline, with the in-sample, cross-fitted and placebo values side by side (Section 5.15); and carries the version history in Appendix J. Version 7 is 10.5281/zenodo.22110398. The reproducibility package contains the analysis code, per-year document-frequency tables for Korean, English and Vietnamese, the generated control abstracts, the annotation, matching, translation-route and robustness results, the Holzwarth-estimator scripts and their validation against the released data, the results manifest and the manuscript consistency gate. A public Korean AI-style dictionary and checker built from the same measurements are at https://os.intframe.com/report/ai-style-dictionary-ko and https://os.intframe.com/report/ai-style-check-ko.

Zenodo (CERN European Organization for Nuclear Research)
CellaMedic Biotechnology (South Korea) (KR)
Quality Education
Openalex Percentile: Top 8%
Authorship Attribution and Profiling
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.