SPERT‐AR: A Unified Framework for Resolving User Story Ambiguities to Improve Story Point Estimation

ABSTRACT Objectives Accurate story point estimation remains a core challenge in Agile software development because it directly affects sprint planning, resource allocation, and delivery predictability. Learning‐based estimators have improved prediction accuracy, but their performance can degrade when user stories contain vague wording, missing conditions, inconsistent phrasing, or overlapping scope. This paper presents SPERT‐AR, a framework designed to evaluate whether ambiguity‐aware refinement of user stories can improve story point estimation while preserving the original repository‐assigned estimation labels. Methods SPERT‐AR detects lexical, syntactic, semantic, and pragmatic ambiguity, refines unclear stories using targeted GPT‐4 prompts, and applies expert validation to preserve functional intent. The clarified stories are then used as input to a reinforced transformer estimator. To avoid conflating textual clarification with label revision, the primary evaluation keeps the original repository‐assigned story points unchanged for both original and resolved stories. Results Evaluated on ten real‐world GitLab projects, SPERT‐AR reduces mean absolute error by 29.52% and improves standardized accuracy by 23.41% relative to the SPERT baseline under this fixed‐label setting. Beyond the headline accuracy results, the findings show that SPERT‐AR reduces ambiguity while preserving semantic intent, improves estimation stability across baselines, and provides ablation evidence that ambiguity‐aware refinement contributes to the observed improvement beyond model complexity alone. Conclusion Overall, the findings suggest that clearer user‐story representations can make story point estimation more reliable, while expert‐adjusted effort interpretations should be treated as secondary evidence rather than primary ground truth.

Authors

Institutions

Publication Details

Journal
Software Practice and Experience
Published
2026-09-21
DOI
https://doi.org/10.1002/spe.70108
Primary Topic
Software Engineering Research
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

SPERT‐AR: A Unified Framework for Resolving User Story Ambiguities to Improve Story Point Estimation

Muhammad Ali Lodhi, Tahreem Iqbal, Rui Chen, Waleed Younas et al.
Software Practice and Experience
Software Engineering Research
article

SPERT‐AR: A Unified Framework for Resolving User Story Ambiguities to Improve Story Point Estimation

Muhammad Ali Lodhi, Tahreem Iqbal, Rui Chen, Waleed Younas, Jing Zhao
article en

Abstract

ABSTRACT Objectives Accurate story point estimation remains a core challenge in Agile software development because it directly affects sprint planning, resource allocation, and delivery predictability. Learning‐based estimators have improved prediction accuracy, but their performance can degrade when user stories contain vague wording, missing conditions, inconsistent phrasing, or overlapping scope. This paper presents SPERT‐AR, a framework designed to evaluate whether ambiguity‐aware refinement of user stories can improve story point estimation while preserving the original repository‐assigned estimation labels. Methods SPERT‐AR detects lexical, syntactic, semantic, and pragmatic ambiguity, refines unclear stories using targeted GPT‐4 prompts, and applies expert validation to preserve functional intent. The clarified stories are then used as input to a reinforced transformer estimator. To avoid conflating textual clarification with label revision, the primary evaluation keeps the original repository‐assigned story points unchanged for both original and resolved stories. Results Evaluated on ten real‐world GitLab projects, SPERT‐AR reduces mean absolute error by 29.52% and improves standardized accuracy by 23.41% relative to the SPERT baseline under this fixed‐label setting. Beyond the headline accuracy results, the findings show that SPERT‐AR reduces ambiguity while preserving semantic intent, improves estimation stability across baselines, and provides ablation evidence that ambiguity‐aware refinement contributes to the observed improvement beyond model complexity alone. Conclusion Overall, the findings suggest that clearer user‐story representations can make story point estimation more reliable, while expert‐adjusted effort interpretations should be treated as secondary evidence rather than primary ground truth.

Software Practice and Experience
Dalian University of Technology (CN), Yangzhou University (CN)
Openalex Percentile: Top 4%
Software Engineering Research
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.