Domain Adaptation of LLaMA-2-7B for Veterinary and Livestock Question Answering: A Lexical, Semantic, Statistical and Expert-Based Evaluation

Large Language Models (LLMs) have demonstrated substantial capabilities in natural language processing, yet their adaptation to specialized domains such as veterinary medicine and livestock production remains comparatively underexplored. This study adapted LLaMA-2-7B to the veterinary and livestock domain using a domain-specific corpus comprising 24,182 records and 19,904,874 LLaMA-2 tokens. Domain adaptation was performed using Low-Rank Adaptation (LoRA)-based parameter-efficient fine-tuning. Model performance was evaluated on 100 veterinary and livestock-related questions using veterinarian-prepared reference responses, lexical metrics (BLEU, ROUGE, and METEOR), BERTScore, veterinary expert assessment, question-level paired statistical analyses, and selected general capability benchmarks (ARC-Challenge, HellaSwag, WinoGrande, and MMLU). Question-level paired analyses showed that METEOR increased significantly from 0.1741 to 0.2156 (Holm-adjusted p < 0.001), whereas BLEU, ROUGE-1, and ROUGE-2 showed no significant differences after multiple-comparison correction. ROUGE-L decreased from 0.1236 to 0.1096 (adjusted p = 0.0403), while BERTScore showed a small but significant decrease from 0.8442 to 0.8373 (adjusted p = 0.0113). Veterinary expert assessment of the LoRA-adapted model yielded an overall question-level score of 7.1960 ± 1.4589 across five qualitative dimensions. Agreement between the two complete domain-expert assessment streams showed a moderate-to-good reliability profile across complementary agreement statistics. General capability evaluation showed modest improvements on ARC-Challenge and HellaSwag and modest decreases on WinoGrande and MMLU. Overall, LoRA-based domain adaptation produced measurable but metric-dependent effects rather than uniform improvements. These findings emphasize the importance of multidimensional evaluation and provide a methodological foundation for the development and assessment of specialized LLMs for veterinary medicine and livestock production. The findings characterize model behavior within a controlled veterinary and livestock question-answering setting and should not be interpreted as evidence of clinical safety or readiness for autonomous veterinary decision support.

Authors

Institutions

Publication Details

Journal
Applied Sciences
Published
2026-09-28
DOI
https://doi.org/10.3390/app16199625
Primary Topic
Topic Modeling
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Domain Adaptation of LLaMA-2-7B for Veterinary and Livestock Question Answering: A Lexical, Semantic, Statistical and Expert-Based Evaluation

Mehmet Erkan Yüksel, Selçuk Sevgen, Tolgahan Öztürk
Applied Sciences
Topic Modeling
article

Domain Adaptation of LLaMA-2-7B for Veterinary and Livestock Question Answering: A Lexical, Semantic, Statistical and Expert-Based Evaluation

Mehmet Erkan Yüksel, Selçuk Sevgen, Tolgahan Öztürk
article en

Abstract

Large Language Models (LLMs) have demonstrated substantial capabilities in natural language processing, yet their adaptation to specialized domains such as veterinary medicine and livestock production remains comparatively underexplored. This study adapted LLaMA-2-7B to the veterinary and livestock domain using a domain-specific corpus comprising 24,182 records and 19,904,874 LLaMA-2 tokens. Domain adaptation was performed using Low-Rank Adaptation (LoRA)-based parameter-efficient fine-tuning. Model performance was evaluated on 100 veterinary and livestock-related questions using veterinarian-prepared reference responses, lexical metrics (BLEU, ROUGE, and METEOR), BERTScore, veterinary expert assessment, question-level paired statistical analyses, and selected general capability benchmarks (ARC-Challenge, HellaSwag, WinoGrande, and MMLU). Question-level paired analyses showed that METEOR increased significantly from 0.1741 to 0.2156 (Holm-adjusted p < 0.001), whereas BLEU, ROUGE-1, and ROUGE-2 showed no significant differences after multiple-comparison correction. ROUGE-L decreased from 0.1236 to 0.1096 (adjusted p = 0.0403), while BERTScore showed a small but significant decrease from 0.8442 to 0.8373 (adjusted p = 0.0113). Veterinary expert assessment of the LoRA-adapted model yielded an overall question-level score of 7.1960 ± 1.4589 across five qualitative dimensions. Agreement between the two complete domain-expert assessment streams showed a moderate-to-good reliability profile across complementary agreement statistics. General capability evaluation showed modest improvements on ARC-Challenge and HellaSwag and modest decreases on WinoGrande and MMLU. Overall, LoRA-based domain adaptation produced measurable but metric-dependent effects rather than uniform improvements. These findings emphasize the importance of multidimensional evaluation and provide a methodological foundation for the development and assessment of specialized LLMs for veterinary medicine and livestock production. The findings characterize model behavior within a controlled veterinary and livestock question-answering setting and should not be interpreted as evidence of clinical safety or readiness for autonomous veterinary decision support.

Applied SciencesVol. 16(19)
Istanbul University-Cerrahpaşa (TR), Burdur Mehmet Akif Ersoy Üniversitesi (TR), Istanbul University (TR)
Openalex Percentile: Top 9%
Topic Modeling
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.