Steering Tree-of-Thought Reasoning via Deductive Verification
Large language models (LLMs) have demonstrated great potential in code reasoning tasks, but their reasoning processes lack reliable verification mechanisms, making it difficult to ensure logical correctness. The Tree of Thoughts (ToT) framework improves reasoning by exploring multiple paths and employing backtracking, yet its path selection relies entirely on LLM-based self-evaluation—a heuristic and error-prone mechanism—leading to frequent erroneous pruning and unproductive exploration. We identify a key insight: LLMs’ encoding capability is stronger than their reasoning capability—translating code semantics into formal constraints is a pattern-matching task that LLMs can perform reliably, while verification should be delegated to SMT(Satisfiability Modulo Theories) solvers. Based on this insight, we propose Deductive Steering, a mechanism that integrates SMT solver verification into Tree of Thoughts exploration. It consists of four core components: (1) Candidate Generator produces candidate reasoning steps, each comprising a natural language thought t and its SMT constraint encoding ϕ; (2) Deductive Evaluator verifies whether a candidate constraint ϕ is a logical consequence of the accumulated constraint Φ by checking the unsatisfiability of Φ ∧ ¬ ϕ; (3) Counterexample Refinement uses counterexample to guide the LLM in correcting its reasoning when verification fails; (4) Exploration and Backtracking Strategy manages path exploration and backtracks to alternative candidates when verification fails. Experiments on five benchmarks covering fault localization, program synthesis, and loop invariant generation show that, compared with ToT, Deductive Steering improves task-level effectiveness by 9.2–32.6 percentage points while reducing token consumption by 35.7–52.4%. The method generalizes across different LLMs and extends to mathematical reasoning, demonstrating broad applicability to domains where reasoning can be encoded as formal constraints.
Authors
- Xinyu Gao (ORCID: https://orcid.org/0009-0004-7135-1833)
- Enyi Tang (ORCID: https://orcid.org/0000-0001-9004-1292)
- Yuchuan Liu (ORCID: https://orcid.org/0000-0003-0928-9296)
- Yu Tian (ORCID: https://orcid.org/0000-0003-3847-4378)
- Shuoxiao Zhang (ORCID: https://orcid.org/0009-0004-3023-5027)
- Cheng Haoliang (ORCID: https://orcid.org/0009-0000-4370-9309)
- Jason Ma (ORCID: https://orcid.org/0009-0007-0006-1021)
- Haibin Wang (ORCID: https://orcid.org/0009-0004-7865-2874)
- Jiahe Mao (ORCID: https://orcid.org/0009-0007-7539-9684)
- Keyu Cui (ORCID: https://orcid.org/0009-0008-2368-267X)
- Yanling Fu (ORCID: https://orcid.org/0009-0009-6879-3899)
Institutions
- Nanjing University (CN)
Publication Details
- Journal
- Proceedings of the ACM on software engineering.
- Published
- 2026-10-01
- DOI
- https://doi.org/10.1145/3832200
- Primary Topic
- Software System Performance and Reliability
- Type
- article
- Field-Weighted Citation Impact
- 0.00