Out of Time: Calibration and Stability of Tabular Foundation Models on Credit Default Vintages

We evaluate TabPFN-3 and TabICLv2, at the defaults of tabpfn 8.5.0 and tabicl 2.1.1, out of time on two credit books, Lending Club and Freddie Mac, against a weight-of-evidence scorecard and a tuned gradient-boosting model. Each model is scored on every later cohort, under kill criteria written before any foundation model was fitted on each book and deposited after the Lending Club runs and the Freddie Mac fits. At the shipped softmax temperature of 0.9, named the primary setting after the first build had been read, both foundation models predict too few defaults on both books: on Lending Club, observed over expected averaged over cells is 1.845 for TabPFN and 1.594 for TabICL, against 1.283 for gradient boosting. On Lending Club the shortfall is already there on the rows each model was conditioned on, and temperature one removes it there. Under the registered age specification, neither model's AUC falls with age faster than that of a gradient-boosting model fitted on the same context rows. Every rejection on the criterion's matrices rests on a pooling chosen after its result, on the scorecard's larger training pool, on the stability reference or on the temperature; the level was reported, not tested. Without the one Lending Club column whose clipped bound dates a loan, four criterion rows change state, TabPFN's calibration row to the state that meets its kill criterion; the verdicts stay those of the registered matrix. TabPFN-3.5, which later releases of tabpfn load by default, was scored after these runs on one Lending Club arm and decides no verdict: there, under one bootstrap seed, it misses the realised default rate by less than TabPFN-3 at either temperature, and its AUC falls with age faster than that of the gradient-boosting model on the same context rows when each cohort has its own intercept. The age specification, the stability reference and the temperature each move a verdict, and the total Brier score, within a thousandth across models, misses the size of the level gap.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-08
DOI
https://doi.org/10.5281/zenodo.23249719
Primary Topic
Financial Distress and Bankruptcy Prediction
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

Out of Time: Calibration and Stability of Tabular Foundation Models on Credit Default Vintages

Nikolai Khobotov
Zenodo (CERN European Organization for Nuclear Research)
Financial Distress and Bankruptcy Prediction
preprint

Out of Time: Calibration and Stability of Tabular Foundation Models on Credit Default Vintages

Nikolai Khobotov
preprint en

Abstract

We evaluate TabPFN-3 and TabICLv2, at the defaults of tabpfn 8.5.0 and tabicl 2.1.1, out of time on two credit books, Lending Club and Freddie Mac, against a weight-of-evidence scorecard and a tuned gradient-boosting model. Each model is scored on every later cohort, under kill criteria written before any foundation model was fitted on each book and deposited after the Lending Club runs and the Freddie Mac fits. At the shipped softmax temperature of 0.9, named the primary setting after the first build had been read, both foundation models predict too few defaults on both books: on Lending Club, observed over expected averaged over cells is 1.845 for TabPFN and 1.594 for TabICL, against 1.283 for gradient boosting. On Lending Club the shortfall is already there on the rows each model was conditioned on, and temperature one removes it there. Under the registered age specification, neither model's AUC falls with age faster than that of a gradient-boosting model fitted on the same context rows. Every rejection on the criterion's matrices rests on a pooling chosen after its result, on the scorecard's larger training pool, on the stability reference or on the temperature; the level was reported, not tested. Without the one Lending Club column whose clipped bound dates a loan, four criterion rows change state, TabPFN's calibration row to the state that meets its kill criterion; the verdicts stay those of the registered matrix. TabPFN-3.5, which later releases of tabpfn load by default, was scored after these runs on one Lending Club arm and decides no verdict: there, under one bootstrap seed, it misses the realised default rate by less than TabPFN-3 at either temperature, and its AUC falls with age faster than that of the gradient-boosting model on the same context rows when each cohort has its own intercept. The age specification, the stability reference and the temperature each move a verdict, and the total Brier score, within a thousandth across models, misses the size of the level gap.

Zenodo (CERN European Organization for Nuclear Research)
Financial Distress and Bankruptcy Prediction
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.