Does swapping the Claude API model change which businesses it names from the same search results? A pre-registered same-day A/B of claude-sonnet-4-6 and claude-sonnet-5 on Kansas City AI discoverability questions
Pre-registration, deposited before any study data is collected. Lefty Media Co measures Pick Rate: whether an AI engine names a business when asked who is best. Its Claude measurement calls the Anthropic Messages API with claude-sonnet-4-6 and is scheduled to switch to claude-sonnet-5 on October 16, 2026. This study asks whether that model swap, on its own, changes the result. 52 Kansas City AI discoverability questions are each answered 5 times by each model (520 calls), with the question, a frozen Tavily search-result packet, the system prompt, the request settings and the scoring rule held fixed. Only the model string differs. The claim is about the API under a supplied grounding packet, not about the Claude.ai app, and the result is the operational effect of the swap inside this exact configuration, including its 500-token output cap. Outcomes: Lefty Media Co's strict-citation Pick Rate (paired per-question difference with a question-level bootstrap), how often each model names any business, and the overlap of the businesses named, benchmarked against claude-sonnet-4-6's agreement with itself. This deposit contains the pre-registration (PREREGISTRATION.md, also inside the zip), the frozen inputs with SHA-256 checksums, the run, extraction and analysis scripts, their tests on synthetic data, a disclosed five-call configuration probe, two external reviews and how each point was handled, and a manifest. Results will be added as a new version under the same concept DOI, with no change to the pre-registration. Author's research page: leftymediaco.com/research
Authors
- Nathan Hall
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-25
- DOI
- https://doi.org/10.5281/zenodo.22961448
- Primary Topic
- Data Analysis with R
- Type
- preprint