Cattle-CLIP: A Multimodal Framework for Dairy Cattle Behaviour Recognition from Video
Monitoring cattle behaviour provides valuable insights into animal health, productivity and welfare conditions. However, robust behaviour recognition in real-world farm environments remains challenging due to several data-related limitations, including the scarcity of well-annotated livestock video datasets and the substantial domain gap between large-scale pretraining corpora and agricultural surveillance footage. To address these challenges, Cattle-CLIP, a domain-adaptive vision–language framework based on Contrastive Language-Image Pretraining (CLIP), is proposed, in which cattle behaviour recognition is reformulated as a cross-modal semantic alignment task rather than a purely visual classification problem. Instead of directly fine-tuning visual backbones, Cattle-CLIP incorporates a temporal integration module to extend image-level contrastive pretraining to video-based behaviour understanding, enabling consistent semantic alignment across time. In order to minimise the domain discrepancy between web-scale pretraining data and real-world cattle surveillance scenarios, customised data augmentation and specialised behaviour-oriented prompts are employed. Furthermore, CattleBehaviours6, a curated and behaviour-consistent video dataset comprising 1905 annotated clips across six indoor behaviours was constructed: feeding, drinking, standing-self-grooming, standing-ruminating, lying-self-grooming, lying-ruminating, to support model training and evaluation. Beyond serving as a benchmark for our proposed method, the dataset provides a standardised ethogram definition, offering a practical resource for future research in livestock behaviour analysis. The performance of Cattle-CLIP was assessed under both supervised and few-shot learning paradigms, particularly targeting behaviour recognition in data-constrained environments, an important yet insufficiently studied task in livestock surveillance. Experimental evaluation demonstrated that the proposed model reached 96.1% accuracy across six behaviours in supervised tasks, while achieving both high precision and recall for feeding, drinking and standing-ruminating behaviours. Moreover, the model maintained promising few-shot performance under limited labelled data, highlighting the potential of multimodal learning for data-constrained cow behaviour recognition.
Authors
- Jing Gao (ORCID: https://orcid.org/0000-0001-5983-1114)
- Andrew W. Dowsey (ORCID: https://orcid.org/0000-0002-7404-9128)
- Neill Campbell
- Daria Baran (ORCID: https://orcid.org/0009-0004-8456-9102)
- Huimin Liu (ORCID: https://orcid.org/0009-0005-6269-8322)
- Axel X. Montout
Institutions
- University of Bristol (GB)
Publication Details
- Journal
- AgriEngineering
- Published
- 2026-09-30
- DOI
- https://doi.org/10.3390/agriengineering8100410
- Primary Topic
- Animal Behavior and Welfare Studies
- Type
- article
- Field-Weighted Citation Impact
- 0.00