Evaluating Sequence Foundation Models for Insect Olfactory Gene Discovery: Failure Modes, Evidence Standards, and a Benchmark Agenda

Genome-scale searches for insect olfactory genes are vulnerable to rapid receptor divergence, tandem duplication, fragmented gene models, circular labels, and evolutionary leakage between training and test data. Sequence foundation models trained on DNA or proteins may extend homology-based annotation, but available evidence does not support their use as autonomous annotators or direct predictors of sensory function. This critical review organizes the problem around annotation failures rather than model catalogues. We distinguish six inference levels, from candidate locus detection to organism-level function, and define task-specific high-confidence reference criteria. We also propose a benchmark using biological hard negatives, sequence-cluster and taxonomic holdouts, conventional baselines, calibration, abstention, and workload-aware retrieval metrics. Evidence integration is defined operationally as a pipeline that retains model scores alongside similarity, domains, topology, phylogeny, expression, and assay evidence, without promoting a weak signal to a stronger biological claim. A published ESM-2-supported analysis of remote insect chemoreceptor homology illustrates both the potential and the current evidence boundary. A convincing benefit will require an incremental gain over transparent baselines under frozen biologically realistic splits, with reliable uncertainty and a clear route from ambiguous predictions to review or experiment.

Authors

Institutions

Publication Details

Journal
Insects
Published
2026-09-04
DOI
https://doi.org/10.3390/insects17090928
Primary Topic
Neurobiology and Insect Physiology Research
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Evaluating Sequence Foundation Models for Insect Olfactory Gene Discovery: Failure Modes, Evidence Standards, and a Benchmark Agenda

Xiao Chun, Huiqin Li, Guoxing Wu, Junfu Yu et al.
Insects
Neurobiology and Insect Physiology Research
article

Evaluating Sequence Foundation Models for Insect Olfactory Gene Discovery: Failure Modes, Evidence Standards, and a Benchmark Agenda

Xiao Chun, Huiqin Li, Guoxing Wu, Junfu Yu, Dian Zhou, Wei Pu, Xi Gao
article en

Abstract

Genome-scale searches for insect olfactory genes are vulnerable to rapid receptor divergence, tandem duplication, fragmented gene models, circular labels, and evolutionary leakage between training and test data. Sequence foundation models trained on DNA or proteins may extend homology-based annotation, but available evidence does not support their use as autonomous annotators or direct predictors of sensory function. This critical review organizes the problem around annotation failures rather than model catalogues. We distinguish six inference levels, from candidate locus detection to organism-level function, and define task-specific high-confidence reference criteria. We also propose a benchmark using biological hard negatives, sequence-cluster and taxonomic holdouts, conventional baselines, calibration, abstention, and workload-aware retrieval metrics. Evidence integration is defined operationally as a pipeline that retains model scores alongside similarity, domains, topology, phylogeny, expression, and assay evidence, without promoting a weak signal to a stronger biological claim. A published ESM-2-supported analysis of remote insect chemoreceptor homology illustrates both the potential and the current evidence boundary. A convincing benefit will require an incremental gain over transparent baselines under frozen biologically realistic splits, with reliable uncertainty and a clear route from ambiguous predictions to review or experiment.

InsectsVol. 17(9)
Yunnan Agricultural University (CN), Honghe University (CN), Wenshan University (CN)
National Natural Science Foundation of China
Openalex Percentile: Top 16%
Neurobiology and Insect Physiology Research
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.