Details

Large language models in food and nutrition science : opportunities, challenges, and the case of FoodyLLM
ID Gjorgjevikj, Ana (Author), ID Martinc, Matej (Author), ID Cenikj, Gjorgjina (Author), ID Drole, Jan (Author), ID Ogrinc, Nives (Author), ID Džeroski, Sašo (Author), ID Koroušić-Seljak, Barbara (Author), ID Eftimov, Tome (Author)

.pdfPDF - Presentation file, Download (6,36 MB)
MD5: 4162ED67E2719E47ECDD80803B24960C
URLURL - Source URL, Visit https://www.sciencedirect.com/science/article/pii/S2665927126000511 This link opens in a new window

Abstract
Background Reliable nutrient profiling and semantic interoperability are essential for scalable dietary assessment, food labeling (e.g., traffic-light schemes), and FAIR integration of food composition and consumption data. However, general-purpose large language models (LLMs) are not systematically exposed to structured recipe–nutrition mappings and food ontologies, limiting their accuracy and trustworthiness in food and nutrition tasks. Scope and approach We review recent LLM advances in life sciences and healthcare and analyze the gap in food and nutrition applications. To address this gap, we introduce FoodyLLM, a domain-specialized LLM fine-tuned on 225k task-aligned QA pairs for (i) recipe nutrient estimation, (ii) traffic-light classification, and (iii) ontology-based entity linking to support FAIR food data interoperability. We benchmark FoodyLLM against strong general-purpose baselines (e.g., Llama 3 8B, Gemini 2.0) under zero-/few-shot prompting across five evaluation folds. Key findings Across all tasks, FoodyLLM substantially outperforms general-purpose LLMs for nutrient estimation across all macronutrients (fat, protein, salt, saturates, sugar), accuracy increases from 0.43 to 0.63 to 0.91–0.97; for traffic-light classification across all nutrients and color categories, macro F1 improves from 0.46 to 0.80 to 0.86–0.97; and for ontology-based food entity linking across FoodOn, SNOMED-CT, and Hansard, macro F1 increases from 0.33 to 0.44 (best general-purpose baseline) to 0.93–0.98 on artificial NEL data, and from 0.24 to 0.51 to 0.67–0.84 on real corpora (CafeteriaSA and CafeteriaFCD). Overall, our results demonstrate the practical value of domain-specialized LLMs in food and nutrition research. They enable automated dietary assessment, large-scale nutritional monitoring, and FAIR data integration, while opening new pathways toward sustainable and personalized nutrition.

Language:English
Keywords:FoodyLLM, nutrient estimation, data interoperability
Work type:Article
Typology:1.01 - Original Scientific Article
Organization:FRI - Faculty of Computer and Information Science
MF - Faculty of Medicine
Publication status:Published
Publication version:Version of Record
Year:2026
Number of pages:26 str.
Numbering:Vol. 12, art. 101351
PID:20.500.12556/RUL-185271 This link opens in a new window
UDC:004.8
ISSN on article:2665-9271
DOI:10.1016/j.crfs.2026.101351 This link opens in a new window
COBISS.SI-ID:270414595 This link opens in a new window
Publication date in RUL:30.07.2026
Views:57
Downloads:16
Metadata:XML DC-XML DC-RDF
:
Copy citation
Share:Bookmark and Share

Record is a part of a journal

Title:Current research in food science
Publisher:Elsevier
ISSN:2665-9271
COBISS.SI-ID:18959875 This link opens in a new window

Licences

License:CC BY 4.0, Creative Commons Attribution 4.0 International
Link:http://creativecommons.org/licenses/by/4.0/
Description:This is the standard Creative Commons license that gives others maximum freedom to do what they want with the work as long as they credit the author.

Secondary language

Language:Slovenian
Keywords:FoodyLLM, interoperabilnost podatkov

Projects

Funder:ARRS - Slovenian Research Agency
Project number:P2-0098
Name:Računalniške strukture in sistemi

Funder:ARRS - Slovenian Research Agency
Project number:P2-0103
Name:Tehnologije znanja

Funder:ARRS - Slovenian Research Agency
Project number:P1-0143
Name:Kroženje snovi v okolju, snovna bilanca in modeliranje okoljskih procesov ter ocena tveganja

Funder:ARRS - Slovenian Research Agency
Project number:J7-70265
Name:AI4Food: Developing and Validating a Comprehensive Food Composition Database and a Knowledge Base Aligned with FAIR Principles, Artificial Intelligence Methods, and Large Language Models

Funder:ARRS - Slovenian Research Agency
Project number:GC-0001
Name:Umetna inteligenca za znanost

Funder:ARIS - Slovenian Research and Innovation Agency
Project number:PR-12393
Name:Young Researchers Grant

Funder:ARRS - Slovenian Research Agency
Project number:BI-US/24-26-081-2024
Name:Preiskovanje znanstvene literature za napovedovanje interakcij med hrano, boleznimi in zdravili z uporabo grafov znanja

Funder:EC - European Commission
Project number:101211695
Acronym:HE MSCA-PF AutoLLMSelect

Funder:EC - European Commission
Project number:101187010
Acronym:HE ERA Chair AutoLearn SI

Funder:EC - European Commission
Project number:101198470
Name:Large Language Models for the European Union
Acronym:LLMs4EU

Funder:EC - European Commission
Project number:101060712
Acronym:FishEUTrust

Funder:EC - European Commission
Project number:101254461
Name:Slovenian AI Factory
Acronym:SLAIF

Similar documents

Similar works from RUL:
Similar works from other Slovenian collections:

Back