The Critical Path of Reasoning: Measuring, and Moving, the Depth of Learned Computation
Chain-of-thought reasoning in language models is serial, and its cost scales with the length of the chain; latent and looped alternatives promise to pay for reasoning with depth instead of size. How much of reasoning's seriality is actually necessary? We answer in three parts. (1) Measurement: we introduce a validated instrument for the dependency structure of reasoning artifacts (critical-path depth, parallelization headroom R = size/depth, antichain width, and a free-exponent scaling fit), governed by three propositions that convert its known failure modes into signed, computable error brackets. Across three real corpora with three independent extraction methods (GSM8K arithmetic, R = 1.44; EntailmentBank proofs, R = 1.14; MBPP programs, R = 1.33), real reasoning artifacts are mostly serial, but headroom grows with size and the share of purely serial instances collapses (78% to 4%, 100% to 0%, 87% to 7%). Exact execution DAGs show the same function can measure R = 1.0 or R = 31.9 depending only on how the solution is written: depth is a property of the expression, not the task. (2) Explanation: on an NC1-complete task with known size and depth (the S5 word problem), weight-tied looped transformers trained with a free iteration budget learn exactly serial composition; the minimum iteration count is K* = n-1 in 22 of 24 cells across seeds and widths, with both exceptions serial-consistent at the largest rung, and never the available logarithmic scan. (3) Manipulation: constraining the training budget to Theta(log n) iterations flips K* onto the trained log budget (2,2,2 at n=4; 3,3,3 at n=6; 4 and 5 at n=8,10), dose-dependently in the fraction of serially infeasible training batches, at 9 to 32 times the optimization cost, with brittleness outside the trained range and a fold ceiling; the same law reappears in a plain autoregressive transformer when scratchpad tokens are rationed. Our budget findings independently confirm, on an adjacent protocol, the concurrent budget law of Zhang et al. (2026), and extend it to the token substrate. Together the parts support a single claim: the seriality of learned reasoning is a default, not a necessity, and the price of parallelism is paid in optimization, not representation.
Authors
- Syed Shaaz (ORCID: https://orcid.org/0009-0006-9810-0641)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-17
- DOI
- https://doi.org/10.5281/zenodo.22803906
- Primary Topic
- Natural Language Processing Techniques
- Type
- preprint