Characterizing LLM scientific concept generation: A multi-dimensional measurement study

Large Language Models (LLMs) can generate text describing scientific concepts, but the characteristics of these outputs remain poorly understood. We present a multi-dimensional characterization framework analyzing 9,666 outputs from seven models (GPT-4.1, GPT-5.2, GPT-5.5, o4-mini, Claude Sonnet 4.5, Claude Opus 4.5, and the open-weight Gemma-3-27B) across five scientific domains. Rather than making claims about creativity or novelty, we measure five independent dimensions: coherence, domain relevance, lexical profile, structural properties, and semantic position using four sentence embedding models spanning 2020-2024. After quality filtering (99.2% coherence, 99.9% domain relevance pass rates), we find that outputs exhibit graduate-level readability (median Flesch-Kincaid grade 16.3) and occupy semantic positions at the 83rd percentile of calibration distributions. All 39 metrics differ significantly across models (Kruskal-Wallis, p < 0.05), with structural properties showing the largest effects (ϵ2 = 0.35-0.54) and semantic position showing small-to-medium effects (ϵ2 ≈ 0.01-0.14). Dunn's post-hoc tests with Bonferroni correction confirm that all model pairs differ on the top structural metrics (21/21 pairs for paragraph count, 20/21 for the next three). Centroid positions are robust to calibration sampling (bootstrap cosine similarity ≥ 0.99) and exceed a shuffled-domain baseline in 96.8-98.7% of cases; cross-model embedding consistency is moderate (Spearman ρ = 0.38-0.67) and is not explained by embedding dimensionality. Temperature effects replicate across two independent full-range models (GPT-4.1 and Gemma). This work provides calibrated measurements and validated methodology for future research without making interpretive claims about novelty.

Authors

Publication Details

Journal
PLoS ONE
Published
2026-09-15
DOI
https://doi.org/10.1371/journal.pone.0357892
Primary Topic
Text Readability and Simplification
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Characterizing LLM scientific concept generation: A multi-dimensional measurement study

Jalil Ahmadpour
PLoS ONE
Text Readability and Simplification
article

Characterizing LLM scientific concept generation: A multi-dimensional measurement study

Jalil Ahmadpour
article en

Abstract

Large Language Models (LLMs) can generate text describing scientific concepts, but the characteristics of these outputs remain poorly understood. We present a multi-dimensional characterization framework analyzing 9,666 outputs from seven models (GPT-4.1, GPT-5.2, GPT-5.5, o4-mini, Claude Sonnet 4.5, Claude Opus 4.5, and the open-weight Gemma-3-27B) across five scientific domains. Rather than making claims about creativity or novelty, we measure five independent dimensions: coherence, domain relevance, lexical profile, structural properties, and semantic position using four sentence embedding models spanning 2020-2024. After quality filtering (99.2% coherence, 99.9% domain relevance pass rates), we find that outputs exhibit graduate-level readability (median Flesch-Kincaid grade 16.3) and occupy semantic positions at the 83rd percentile of calibration distributions. All 39 metrics differ significantly across models (Kruskal-Wallis, p < 0.05), with structural properties showing the largest effects (ϵ2 = 0.35-0.54) and semantic position showing small-to-medium effects (ϵ2 ≈ 0.01-0.14). Dunn's post-hoc tests with Bonferroni correction confirm that all model pairs differ on the top structural metrics (21/21 pairs for paragraph count, 20/21 for the next three). Centroid positions are robust to calibration sampling (bootstrap cosine similarity ≥ 0.99) and exceed a shuffled-domain baseline in 96.8-98.7% of cases; cross-model embedding consistency is moderate (Spearman ρ = 0.38-0.67) and is not explained by embedding dimensionality. Temperature effects replicate across two independent full-range models (GPT-4.1 and Gemma). This work provides calibrated measurements and validated methodology for future research without making interpretive claims about novelty.

PLoS ONEVol. 21(9)
Quality Education
Openalex Percentile: Top 47%
Text Readability and Simplification
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.