FT-Behavioural Router: Knowledge vs Behaviour Separation in Adaptive Hybrid Inference

Hybrid assistants combine retrieval-augmented generation (RAG), tools, and multiple model sizes, but calling a large LLM with retrieval on every query is costly. We study a fine-tuned (FT) behavioural router that selects among four actions: answer as a small model, retrieve and answer as the same small model (FT+RAG), invoke tools, or escalate to a large LLM, while keeping factual knowledge non-parametric. We do not claim inventing adaptive RAG or LLM cascading. On in-domain workplace eval_v2 (n=133), FT DistilBERT reaches accuracy 0.9699 and MiniLM 0.8872, above a prompt-rubric simulation (0.7744; not an LLM API). Exact McNemar and paired bootstrap tests on the same eval items find DistilBERT significantly above MiniLM and keyword (measured p<0.01). On a hard OOD split (n=104; never used for training) DistilBERT drops to 0.8750 and the rubric collapses to 0.5000. A remapped public eval (SQuAD 1.1, HotpotQA, CLINC-150; n=240) is a second domain: workplace cue rules fall to about 0.25–0.28; workplace-trained DistilBERT transfers at 0.5000 while remapped-train DistilBERT reaches 0.9250. A Phase 2 unit-cost simulation shows FT mean cost roughly 0.25× always-RAG+escalate; this is not a cloud bill and no live answer EM/F1 is reported. We deepen the related-work contrast, error analysis, limitations, and when-FT-is-worth-it criteria.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-25
DOI
https://doi.org/10.5281/zenodo.22968232
Primary Topic
Topic Modeling
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

FT-Behavioural Router: Knowledge vs Behaviour Separation in Adaptive Hybrid Inference

Vijay Kumar
Zenodo (CERN European Organization for Nuclear Research)
Topic Modeling
preprint

FT-Behavioural Router: Knowledge vs Behaviour Separation in Adaptive Hybrid Inference

Vijay Kumar
preprint en

Abstract

Hybrid assistants combine retrieval-augmented generation (RAG), tools, and multiple model sizes, but calling a large LLM with retrieval on every query is costly. We study a fine-tuned (FT) behavioural router that selects among four actions: answer as a small model, retrieve and answer as the same small model (FT+RAG), invoke tools, or escalate to a large LLM, while keeping factual knowledge non-parametric. We do not claim inventing adaptive RAG or LLM cascading. On in-domain workplace eval_v2 (n=133), FT DistilBERT reaches accuracy 0.9699 and MiniLM 0.8872, above a prompt-rubric simulation (0.7744; not an LLM API). Exact McNemar and paired bootstrap tests on the same eval items find DistilBERT significantly above MiniLM and keyword (measured p<0.01). On a hard OOD split (n=104; never used for training) DistilBERT drops to 0.8750 and the rubric collapses to 0.5000. A remapped public eval (SQuAD 1.1, HotpotQA, CLINC-150; n=240) is a second domain: workplace cue rules fall to about 0.25–0.28; workplace-trained DistilBERT transfers at 0.5000 while remapped-train DistilBERT reaches 0.9250. A Phase 2 unit-cost simulation shows FT mean cost roughly 0.25× always-RAG+escalate; this is not a cloud bill and no live answer EM/F1 is reported. We deepen the related-work contrast, error analysis, limitations, and when-FT-is-worth-it criteria.

Zenodo (CERN European Organization for Nuclear Research)
Topic Modeling
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

FT-Behavioural Router: Knowledge vs Behaviour Separation in Adaptive Hybrid Inference — Vijay Kumar · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS