A Typed Model That Moves Its Own Decision Boundary
This controlled study examines whether a model designed for typed decisions can be perfectly repeatable while still changing its decision boundary under controlled semantic and configuration changes. A frozen corpus of 41 supplier payment texts was evaluated across two Laya checkpoints and two class semantics conditions, with three executions per text in every condition, producing 492 observations. Identical inputs generated fully repeatable discrete decisions across all four configurations. Materially equivalent reformulations nevertheless produced substantial disagreement, while checkpoint replacement and changes to class definitions also shifted operational outcomes. A separate competence reproduction matched the published 80.4% invoice processing accuracy of the specialist checkpoint. The findings distinguish schema conformance, execution repeatability, reformulation invariance, benchmark competence and decision boundary stability as separate reliability properties. This publication reports a controlled single domain study and does not establish a general ranking of models or providers.
Authors
- José López López
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-30
- DOI
- https://doi.org/10.5281/zenodo.23065531
- Primary Topic
- Topic Modeling
- Type
- preprint