Out of Time: Calibration and Stability of Tabular Foundation Models on Credit Default Vintages
We evaluate TabPFN-3 and TabICLv2, at the defaults of tabpfn 8.5.0 and tabicl 2.1.1, out of time on two credit books, Lending Club and Freddie Mac, against a weight-of-evidence scorecard and a tuned gradient-boosting model. Each model is scored on every later cohort, under kill criteria written before any foundation model was fitted on each book and deposited after the Lending Club runs and the Freddie Mac fits. At the shipped softmax temperature of 0.9, named the primary setting after the first build had been read, both foundation models predict too few defaults on both books: on Lending Club, observed over expected averaged over cells is 1.845 for TabPFN and 1.594 for TabICL, against 1.283 for gradient boosting. On Lending Club the shortfall is already there on the rows each model was conditioned on, and temperature one removes it there. Under the registered age specification, neither model's AUC falls with age faster than that of a gradient-boosting model fitted on the same context rows. Every rejection on the criterion's matrices rests on a pooling chosen after its result, on the scorecard's larger training pool, on the stability reference or on the temperature; the level was reported, not tested. Without the one Lending Club column whose clipped bound dates a loan, four criterion rows change state, TabPFN's calibration row to the state that meets its kill criterion; the verdicts stay those of the registered matrix. TabPFN-3.5, which later releases of tabpfn load by default, was scored after these runs on one Lending Club arm and decides no verdict: there, under one bootstrap seed, it misses the realised default rate by less than TabPFN-3 at either temperature, and its AUC falls with age faster than that of the gradient-boosting model on the same context rows when each cohort has its own intercept. The age specification, the stability reference and the temperature each move a verdict, and the total Brier score, within a thousandth across models, misses the size of the level gap.
Authors
- Nikolai Khobotov (ORCID: https://orcid.org/0009-0005-3612-830X)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-08
- DOI
- https://doi.org/10.5281/zenodo.23249719
- Primary Topic
- Financial Distress and Bankruptcy Prediction
- Type
- preprint