Enhancing FT-MIR Prediction of Short-Chain Free Fatty Acids in Milk Through High-Lipolysis Field Samples, Reference Method Evaluation and Machine Learning Approaches
Lipolysis is a longstanding and appealing topic within the dairy industry. Having been extensively researched throughout the twentieth century, lipolysis is often prematurely dismissed as a settled or outdated field of study. In recent years, the advent of robotic milking systems and increased milking frequencies have led to a rise in lipolysis levels, sparking renewed interest in the monitoring of milk quality. Predicting lipolysis more effectively relies on the analysis of free fatty acid (FFA) concentrations. Recent research has demonstrated the feasibility of quantifying individual FFA in raw milk. In particular, Fourier-transform mid-infrared (FT-MIR) spectroscopy has shown considerable potential for predicting these compounds. However, that study presented several limitations, including the use of non-standardized reference methods and the reliance on artificially lipolyzed samples to enhance variability within the calibration dataset. To overcome these limitations, FT-MIR FFA models have been used as a preliminary screening tool in order to select samples (n = 248) with high lipolysis, and the selected samples were analyzed using the ISO reference method. A total of 248 samples were selected, of which 208 were obtained through the Walloon DHI in Belgium and 40 were sampled at the CRA-W experimental farm. A comparison between the two reference methods was performed and highlighted the need to harmonize reference analysis as the two reference methods are only comparable for short-chain free fatty acids (SCFFA). A mixed model for SCFFA combining both references was built using different algorithms such as Partial Least Squares (PLS), Kernel Ridge Regression (KRR), Support Vector Machine Regression (SVR), Random Forest Regression (RFR) and Gaussian Process Regression (GPR). The model developed for SCFFA showed predictive performances comparable to those reported in the literature with an R2val of 0.80 and an RMSEval of 32.08 mg/L. Qualitative models for SCFFA were created with a threshold highlighted in the literature using the same algorithms as the quantitative model and gave an accuracy of correct classification of 85% applied on the test dataset. This work has highlighted the necessity to correctly select the reference method in order to build the calibration dataset. Non-linear algorithms provided analyte-dependent improvements over PLS, particularly for C6, C8 and SCFFA prediction.
Authors
- Julie Leblois (ORCID: https://orcid.org/0000-0001-5112-0262)
- Frédéric Dehareng (ORCID: https://orcid.org/0000-0002-6733-4334)
- Hélène Soyeurt (ORCID: https://orcid.org/0000-0001-9883-9047)
- Octave S. Christophe (ORCID: https://orcid.org/0000-0002-9037-8079)
- Didier Veselko (ORCID: https://orcid.org/0009-0009-2906-0078)
Institutions
- University of Liège (BE)
- Gembloux Agro-Bio Tech (BE)
- Walloon Agricultural Research Centre (BE)
Publication Details
- Journal
- Foods
- Published
- 2026-09-24
- DOI
- https://doi.org/10.3390/foods15193423
- Primary Topic
- Spectroscopy and Chemometric Analyses
- Type
- article
- Field-Weighted Citation Impact
- 0.00