The Critical Path of Reasoning: Measuring, and Moving, the Depth of Learned Computation

Chain-of-thought reasoning in language models is serial, and its cost scales with the length of the chain; latent and looped alternatives promise to pay for reasoning with depth instead of size. How much of reasoning's seriality is actually necessary? We answer in three parts. (1) Measurement: we introduce a validated instrument for the dependency structure of reasoning artifacts (critical-path depth, parallelization headroom R = size/depth, antichain width, and a free-exponent scaling fit), governed by three propositions that convert its known failure modes into signed, computable error brackets. Across three real corpora with three independent extraction methods (GSM8K arithmetic, R = 1.44; EntailmentBank proofs, R = 1.14; MBPP programs, R = 1.33), real reasoning artifacts are mostly serial, but headroom grows with size and the share of purely serial instances collapses (78% to 4%, 100% to 0%, 87% to 7%). Exact execution DAGs show the same function can measure R = 1.0 or R = 31.9 depending only on how the solution is written: depth is a property of the expression, not the task. (2) Explanation: on an NC1-complete task with known size and depth (the S5 word problem), weight-tied looped transformers trained with a free iteration budget learn exactly serial composition; the minimum iteration count is K* = n-1 in 22 of 24 cells across seeds and widths, with both exceptions serial-consistent at the largest rung, and never the available logarithmic scan. (3) Manipulation: constraining the training budget to Theta(log n) iterations flips K* onto the trained log budget (2,2,2 at n=4; 3,3,3 at n=6; 4 and 5 at n=8,10), dose-dependently in the fraction of serially infeasible training batches, at 9 to 32 times the optimization cost, with brittleness outside the trained range and a fold ceiling; the same law reappears in a plain autoregressive transformer when scratchpad tokens are rationed. Our budget findings independently confirm, on an adjacent protocol, the concurrent budget law of Zhang et al. (2026), and extend it to the token substrate. Together the parts support a single claim: the seriality of learned reasoning is a default, not a necessity, and the price of parallelism is paid in optimization, not representation.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-17
DOI
https://doi.org/10.5281/zenodo.22803907
Primary Topic
Natural Language Processing Techniques
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

The Critical Path of Reasoning: Measuring, and Moving, the Depth of Learned Computation

Syed Shaaz
Zenodo (CERN European Organization for Nuclear Research)
Natural Language Processing Techniques
preprint

The Critical Path of Reasoning: Measuring, and Moving, the Depth of Learned Computation

Syed Shaaz
preprint en

Abstract

Chain-of-thought reasoning in language models is serial, and its cost scales with the length of the chain; latent and looped alternatives promise to pay for reasoning with depth instead of size. How much of reasoning's seriality is actually necessary? We answer in three parts. (1) Measurement: we introduce a validated instrument for the dependency structure of reasoning artifacts (critical-path depth, parallelization headroom R = size/depth, antichain width, and a free-exponent scaling fit), governed by three propositions that convert its known failure modes into signed, computable error brackets. Across three real corpora with three independent extraction methods (GSM8K arithmetic, R = 1.44; EntailmentBank proofs, R = 1.14; MBPP programs, R = 1.33), real reasoning artifacts are mostly serial, but headroom grows with size and the share of purely serial instances collapses (78% to 4%, 100% to 0%, 87% to 7%). Exact execution DAGs show the same function can measure R = 1.0 or R = 31.9 depending only on how the solution is written: depth is a property of the expression, not the task. (2) Explanation: on an NC1-complete task with known size and depth (the S5 word problem), weight-tied looped transformers trained with a free iteration budget learn exactly serial composition; the minimum iteration count is K* = n-1 in 22 of 24 cells across seeds and widths, with both exceptions serial-consistent at the largest rung, and never the available logarithmic scan. (3) Manipulation: constraining the training budget to Theta(log n) iterations flips K* onto the trained log budget (2,2,2 at n=4; 3,3,3 at n=6; 4 and 5 at n=8,10), dose-dependently in the fraction of serially infeasible training batches, at 9 to 32 times the optimization cost, with brittleness outside the trained range and a fold ceiling; the same law reappears in a plain autoregressive transformer when scratchpad tokens are rationed. Our budget findings independently confirm, on an adjacent protocol, the concurrent budget law of Zhang et al. (2026), and extend it to the token substrate. Together the parts support a single claim: the seriality of learned reasoning is a default, not a necessity, and the price of parallelism is paid in optimization, not representation.

Zenodo (CERN European Organization for Nuclear Research)
Peace, Justice and strong institutions
Natural Language Processing Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

The Critical Path of Reasoning: Measuring, and Moving, the Depth of Learned Computation — Syed Shaaz · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS