What a nano-GPT can (not) tell us about spoken language

Abstract After the launch of ChatGPT in autumn 2022, a lot of research has focussed on the quality and near-naturalness Large Language Model-based tools present in the texts they produce. While one area of research has focussed on the similarities and differences between machine-produced and human-produced output (e.g. Berber Sardinha, 2024 ), others have explored how far such tools could process more complex tasks (e.g. Curry et al., 2024 ; Valmeekam, et al., 2023 ). While it can be assumed that ChatGPT makes use of written-to-be-spoken training material, there has been no investigation, as yet, into how far a Generative Pre-trained Transformer (GPT) algorithm is able to process (transcribed) natural, colloquial language. This research will investigate whether spoken language transcripts lead to processing difficulties; whether such generated language can be seen as a suitable reflection of natural speech; and whether machine produced texts offer new insights into the workings of language.

Authors

Institutions

Publication Details

Journal
International Journal of Corpus Linguistics
Published
2026-09-25
DOI
https://doi.org/10.1075/ijcl.24172.pac
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

What a nano-GPT can (not) tell us about spoken language

Michael Thomas Pace-Sigge, Evangs Mailoa
International Journal of Corpus Linguistics
Artificial Intelligence in Healthcare and Education
article

What a nano-GPT can (not) tell us about spoken language

Michael Thomas Pace-Sigge, Evangs Mailoa
article en

Abstract

Abstract After the launch of ChatGPT in autumn 2022, a lot of research has focussed on the quality and near-naturalness Large Language Model-based tools present in the texts they produce. While one area of research has focussed on the similarities and differences between machine-produced and human-produced output (e.g. Berber Sardinha, 2024 ), others have explored how far such tools could process more complex tasks (e.g. Curry et al., 2024 ; Valmeekam, et al., 2023 ). While it can be assumed that ChatGPT makes use of written-to-be-spoken training material, there has been no investigation, as yet, into how far a Generative Pre-trained Transformer (GPT) algorithm is able to process (transcribed) natural, colloquial language. This research will investigate whether spoken language transcripts lead to processing difficulties; whether such generated language can be seen as a suitable reflection of natural speech; and whether machine produced texts offer new insights into the workings of language.

International Journal of Corpus Linguistics
Satya Wacana Christian University (ID), University of Eastern Finland (FI), Finland University (FI)
Quality Education
Openalex Percentile: Top 15%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.