FeeCLAP: Feature-Enhanced Contrastive Language-Audio Pretraining for Descriptive Caption Generation of Chicken Respiratory Sounds

Respiratory diseases in chickens are a major concern in poultry production, as they adversely affect animal health and production efficiency. Chicken vocalizations often contain important acoustic cues associated with respiratory conditions. However, conventional sound classification methods typically rely on predefined categories, limiting their ability to capture complex acoustic patterns. To address this limitation, this study introduces an audio captioning framework for the analysis of chicken respiratory sounds and proposes a model termed FeeCLAP for generating descriptive text from chicken vocalizations. The proposed model maps acoustic signals into natural language descriptions of sound quality, rhythm, and related attributes, thereby enabling semantic-level representation of vocal patterns. Built upon the baseline EnCLAP framework, a dataset of chicken respiratory sound descriptions with medically relevant semantic styles was constructed. An Acoustic Token Adapter (ATA) module was then introduced between discrete acoustic tokens and the text encoder to enhance the temporal modeling of acoustic features. In addition, a CLAP (Contrastive Language-Audio Pretraining)-aware semantic alignment mechanism was incorporated to improve consistency between acoustic representations and the semantic space. During inference, a CLAP-based similarity re-ranking strategy was further applied to improve the quality of generated descriptions. Experiments showed that FeeCLAP outperformed existing methods, achieving Consensus-based Image Description Evaluation (CIDEr), Semantic Propositional Image Caption Evaluation (SPICE), and SPIDEr scores of 0.423, 0.202, and 0.313, respectively, with generated descriptions approaching human-annotated references. Overall, the proposed approach demonstrated its effectiveness in generating semantic descriptions of chicken respiratory sounds under the experimental conditions used in this study. These findings suggest its potential to support respiratory health assessment in poultry; however, further validation under commercial poultry production conditions is required before the method can be considered a reliable tool for practical intelligent monitoring or early detection of respiratory abnormalities.

Authors

Institutions

Publication Details

Journal
Agriculture
Published
2026-08-27
DOI
https://doi.org/10.3390/agriculture16171849
Primary Topic
Phonocardiography and Auscultation Techniques
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

FeeCLAP: Feature-Enhanced Contrastive Language-Audio Pretraining for Descriptive Caption Generation of Chicken Respiratory Sounds

Yanrong Zhuang, Peng Yin, Ligen Yu, Feng Qiu et al.
Agriculture
Phonocardiography and Auscultation Techniques
article

FeeCLAP: Feature-Enhanced Contrastive Language-Audio Pretraining for Descriptive Caption Generation of Chicken Respiratory Sounds

Yanrong Zhuang, Peng Yin, Ligen Yu, Feng Qiu, Qifeng Li, Yue Wu, Jin He, Binzhou Li, Gan Yang, Jiaxing Liu
article en

Abstract

Respiratory diseases in chickens are a major concern in poultry production, as they adversely affect animal health and production efficiency. Chicken vocalizations often contain important acoustic cues associated with respiratory conditions. However, conventional sound classification methods typically rely on predefined categories, limiting their ability to capture complex acoustic patterns. To address this limitation, this study introduces an audio captioning framework for the analysis of chicken respiratory sounds and proposes a model termed FeeCLAP for generating descriptive text from chicken vocalizations. The proposed model maps acoustic signals into natural language descriptions of sound quality, rhythm, and related attributes, thereby enabling semantic-level representation of vocal patterns. Built upon the baseline EnCLAP framework, a dataset of chicken respiratory sound descriptions with medically relevant semantic styles was constructed. An Acoustic Token Adapter (ATA) module was then introduced between discrete acoustic tokens and the text encoder to enhance the temporal modeling of acoustic features. In addition, a CLAP (Contrastive Language-Audio Pretraining)-aware semantic alignment mechanism was incorporated to improve consistency between acoustic representations and the semantic space. During inference, a CLAP-based similarity re-ranking strategy was further applied to improve the quality of generated descriptions. Experiments showed that FeeCLAP outperformed existing methods, achieving Consensus-based Image Description Evaluation (CIDEr), Semantic Propositional Image Caption Evaluation (SPICE), and SPIDEr scores of 0.423, 0.202, and 0.313, respectively, with generated descriptions approaching human-annotated references. Overall, the proposed approach demonstrated its effectiveness in generating semantic descriptions of chicken respiratory sounds under the experimental conditions used in this study. These findings suggest its potential to support respiratory health assessment in poultry; however, further validation under commercial poultry production conditions is required before the method can be considered a reliable tool for practical intelligent monitoring or early detection of respiratory abnormalities.

AgricultureVol. 16(17)
Tianjin Agricultural University (CN), National Engineering Research Center for Information Technology in Agriculture (CN), Chemical Industry Press (CN), China Agricultural University (CN)
National Natural Science Foundation of China
Quality Education
Openalex Percentile: Top 11%
Phonocardiography and Auscultation Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.