aadya-m1: An Open, On-Device Probabilistic Model for Calibrated Next-Period Forecasting
Period-tracking applications are used by hundreds of millions of people, yet most show a single predicted date, conceal how uncertain it is, and publish no way for others to check their accuracy. We study next-period forecasting as probabilistic prediction and release an open benchmark, an open model and an open leaderboard. AadyaBench scores full probability distributions with the continuous ranked probability score (CRPS), fixes seeds, splits and decision rules before results are seen, and reports paired per-user bootstrap intervals; it combines a calibrated simulator with six scenarios, two public real-world cohorts, a cold-start section and a missed-log stress test. aadya-m1 is a 256M-parameter transformer that maps past cycle lengths, the days since the last period and an optional age group to a probability for every cycle length from 1 to 120 days; a history-qualitydependent blend with a hierarchical Bayesian reference makes it robust on short and irregular histories. It runs on the user's own device and is trained only on synthetic users. On 539 simulated validation users aadya-m1 attains the lowest mean CRPS of eleven open methods (2.925, against 2.965 for an LSTM trained on the same data, 3.026 for the Bayesian reference, 3.109 for a reimplementation of the Poisson-with-skips model of Li et al. and 3.702 for the rolling-median heuristic used by most trackers), and every paired difference is clear. On two public real-world cohorts it again has the lowest mean error (1.580 on Marquette, 113 users; 2.290 on a Creighton hold-out, 37 users): slightly ahead of the Bayesian reference, a difference within statistical noise at these sample sizes, and 15 to 23% below the rolling-median heuristic. An optional ovulation-test input narrows the 80% window from about 12 to 7 days on 41 participants of the public mcPHASES dataset. We document the checks and corrections made along the way, including a cold-start defect, a rejected fine-tune and a personalization rule that we removed, and we make no comparison with closed commercial applications, for which no shared data exist.
Authors
- Manas Kumar Dutta (ORCID: https://orcid.org/0009-0001-4273-0453)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-06
- DOI
- https://doi.org/10.5281/zenodo.23183872
- Primary Topic
- Forecasting Techniques and Applications
- Type
- preprint