Cattle-CLIP: A Multimodal Framework for Dairy Cattle Behaviour Recognition from Video

Monitoring cattle behaviour provides valuable insights into animal health, productivity and welfare conditions. However, robust behaviour recognition in real-world farm environments remains challenging due to several data-related limitations, including the scarcity of well-annotated livestock video datasets and the substantial domain gap between large-scale pretraining corpora and agricultural surveillance footage. To address these challenges, Cattle-CLIP, a domain-adaptive vision–language framework based on Contrastive Language-Image Pretraining (CLIP), is proposed, in which cattle behaviour recognition is reformulated as a cross-modal semantic alignment task rather than a purely visual classification problem. Instead of directly fine-tuning visual backbones, Cattle-CLIP incorporates a temporal integration module to extend image-level contrastive pretraining to video-based behaviour understanding, enabling consistent semantic alignment across time. In order to minimise the domain discrepancy between web-scale pretraining data and real-world cattle surveillance scenarios, customised data augmentation and specialised behaviour-oriented prompts are employed. Furthermore, CattleBehaviours6, a curated and behaviour-consistent video dataset comprising 1905 annotated clips across six indoor behaviours was constructed: feeding, drinking, standing-self-grooming, standing-ruminating, lying-self-grooming, lying-ruminating, to support model training and evaluation. Beyond serving as a benchmark for our proposed method, the dataset provides a standardised ethogram definition, offering a practical resource for future research in livestock behaviour analysis. The performance of Cattle-CLIP was assessed under both supervised and few-shot learning paradigms, particularly targeting behaviour recognition in data-constrained environments, an important yet insufficiently studied task in livestock surveillance. Experimental evaluation demonstrated that the proposed model reached 96.1% accuracy across six behaviours in supervised tasks, while achieving both high precision and recall for feeding, drinking and standing-ruminating behaviours. Moreover, the model maintained promising few-shot performance under limited labelled data, highlighting the potential of multimodal learning for data-constrained cow behaviour recognition.

Authors

Institutions

Publication Details

Journal
AgriEngineering
Published
2026-09-30
DOI
https://doi.org/10.3390/agriengineering8100410
Primary Topic
Animal Behavior and Welfare Studies
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Cattle-CLIP: A Multimodal Framework for Dairy Cattle Behaviour Recognition from Video

Jing Gao, Andrew W. Dowsey, Neill Campbell, Daria Baran et al.
AgriEngineering
Animal Behavior and Welfare Studies
article

Cattle-CLIP: A Multimodal Framework for Dairy Cattle Behaviour Recognition from Video

Jing Gao, Andrew W. Dowsey, Neill Campbell, Daria Baran, Huimin Liu, Axel X. Montout
article en

Abstract

Monitoring cattle behaviour provides valuable insights into animal health, productivity and welfare conditions. However, robust behaviour recognition in real-world farm environments remains challenging due to several data-related limitations, including the scarcity of well-annotated livestock video datasets and the substantial domain gap between large-scale pretraining corpora and agricultural surveillance footage. To address these challenges, Cattle-CLIP, a domain-adaptive vision–language framework based on Contrastive Language-Image Pretraining (CLIP), is proposed, in which cattle behaviour recognition is reformulated as a cross-modal semantic alignment task rather than a purely visual classification problem. Instead of directly fine-tuning visual backbones, Cattle-CLIP incorporates a temporal integration module to extend image-level contrastive pretraining to video-based behaviour understanding, enabling consistent semantic alignment across time. In order to minimise the domain discrepancy between web-scale pretraining data and real-world cattle surveillance scenarios, customised data augmentation and specialised behaviour-oriented prompts are employed. Furthermore, CattleBehaviours6, a curated and behaviour-consistent video dataset comprising 1905 annotated clips across six indoor behaviours was constructed: feeding, drinking, standing-self-grooming, standing-ruminating, lying-self-grooming, lying-ruminating, to support model training and evaluation. Beyond serving as a benchmark for our proposed method, the dataset provides a standardised ethogram definition, offering a practical resource for future research in livestock behaviour analysis. The performance of Cattle-CLIP was assessed under both supervised and few-shot learning paradigms, particularly targeting behaviour recognition in data-constrained environments, an important yet insufficiently studied task in livestock surveillance. Experimental evaluation demonstrated that the proposed model reached 96.1% accuracy across six behaviours in supervised tasks, while achieving both high precision and recall for feeding, drinking and standing-ruminating behaviours. Moreover, the model maintained promising few-shot performance under limited labelled data, highlighting the potential of multimodal learning for data-constrained cow behaviour recognition.

AgriEngineeringVol. 8(10)
University of Bristol (GB)
Zero hunger
Openalex Percentile: Top 10%
Animal Behavior and Welfare Studies
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.