Your browser does not allow JavaScript!
JavaScript is necessary for the proper functioning of this website. Please enable JavaScript or use a modern browser.
Repository of the University of Ljubljana
Open Science Slovenia
Open Science
DiKUL
slv
|
eng
Search
Advanced
New in RUL
About RUL
In numbers
Help
Sign in
Details
Environmentally grounded pseudo-absence sampling for species distribution models : a language guided framework
ID
Miok, Kristian
(
Author
),
ID
Laza, Antonio V.
(
Author
),
ID
Škrlj, Blaž
(
Author
),
ID
Robnik Šikonja, Marko
(
Author
),
ID
Pârvulescu, Lucian
(
Author
)
PDF - Presentation file,
Download
(1,74 MB)
MD5: 353EE713B22F5B866BF7203226D70812
URL - Source URL, Visit
https://onlinelibrary.wiley.com/doi/10.1111/ddi.70199
Image galllery
Abstract
Aim: Species Distribution Models (SDMs) are widely used in conservation planning, invasive species management and global change assessments. Their reliability depends on both presence and absence data, yet biodiversity databases are dominated by presence records, while true absences are rarely collected and geographically restricted. We introduce a framework that integrates large language models (LLMs) to sample ecologically realistic pseudo-absences under the constraints of dendritic river networks. Innovation: We present the Language-Grounded Multivariate Pseudo-Absence Sampling (LGMPAS) framework, which converts environmental predictor profiles into natural-language eco-narratives that an LLM uses to score and rank pre-filtered candidate locations drawn exclusively from the river network and restricted to unlabelled sites. Retrieval-Augmented Generation (RAG) further anchors eco-narratives in published ecological knowledge. We tested LGMPAS on two ecologically contrasting crayfish species in the Danube basin, the widespread invasive Faxonius limosus and the narrowly endemic Austropotamobius bihariensis, validating outputs against independent field-collected true-absence data using Random Forest predictive performance, spatial overlap of high-suitability areas and predictor-space distances. LLM-derived pseudo-absences closely reproduced true-absence model outputs and consistently outperformed random sampling across both species. Main Conclusions: LGMPAS demonstrates that LLMs can reliably sample pseudo-absences that reproduce the ecological signal of true-absence data, even under complex freshwater network constraints. By reducing dependence on costly absence surveys in contexts where true-absence data are unavailable or spatially restricted, and by avoiding the biases inherent to random sampling, the framework strengthens the robustness of SDMs for conservation applications. Its reproducibility and adaptability across taxa and ecosystems offer particular value for biodiversity monitoring, invasive species management and conservation planning under global change.
Language:
English
Keywords:
species distribution models
,
pseudo-absence sampling
,
large language models
,
ecological modelling
,
spatial ecology
,
freshwater ecosystems
,
crayfish
,
retrieval-augmented generation
Work type:
Article
Typology:
1.01 - Original Scientific Article
Organization:
FRI - Faculty of Computer and Information Science
Publication status:
Published
Publication version:
Version of Record
Year:
2026
Number of pages:
13 str.
Numbering:
Vol. 32, iss. 5, art. e70199
PID:
20.500.12556/RUL-183939
UDC:
004.8:574
ISSN on article:
1366-9516
DOI:
10.1111/ddi.70199
COBISS.SI-ID:
277720323
Publication date in RUL:
22.06.2026
Views:
201
Downloads:
207
Metadata:
Cite this work
Plain text
BibTeX
EndNote XML
EndNote/Refer
RIS
ABNT
ACM Ref
AMA
APA
Chicago 17th Author-Date
Harvard
IEEE
ISO 690
MLA
Vancouver
:
Copy citation
Share:
Record is a part of a journal
Title:
Diversity and distributions : a journal of conservation biogeography
Shortened title:
Divers. distrib.
Publisher:
Wiley
ISSN:
1366-9516
COBISS.SI-ID:
30709
Licences
License:
CC BY 4.0, Creative Commons Attribution 4.0 International
Link:
http://creativecommons.org/licenses/by/4.0/
Description:
This is the standard Creative Commons license that gives others maximum freedom to do what they want with the work as long as they credit the author.
Secondary language
Language:
Slovenian
Keywords:
modeliranje razširjenosti vrst
,
psevdo-odsotnosti
,
veliki jezikovni modeli
,
ekološko modeliranje
,
prostorska ekologija
,
sladkovodni ekosistemi
,
raki
,
priklicno obogatena generacija
Projects
Funder:
Other - Other funder or multiple funders
Project number:
PN-III-P4-ID-PCE-2020-1187
Funder:
EC - European Commission
Funding programme:
HE
Project number:
101081355
Name:
Machine learning for Sciences and Humanities
Acronym:
SMASH
Funder:
ARIS - Slovenian Research and Innovation Agency
Project number:
P2-0103
Name:
Tehnologije znanja
Funder:
ARIS - Slovenian Research and Innovation Agency
Project number:
P6-0411
Name:
Jezikovni viri in tehnologije za slovenski jezik
Funder:
ARIS - Slovenian Research and Innovation Agency
Project number:
L2-50070
Name:
Tehnike vektorskih vložitev za medijske aplikacije
Funder:
ARIS - Slovenian Research and Innovation Agency
Project number:
GC-0002
Name:
Veliki jezikovni modeli za digitalno humanistiko
Funder:
ARIS - Slovenian Research and Innovation Agency
Project number:
J4-4555
Name:
Napovedovanje patogenosti in perzistence bakterij Listeria monocytogenes na osnovi značilnosti njihovih biofilmov in surfaktoma s pomočjo strojnega učenja
Similar documents
Similar works from RUL:
Similar works from other Slovenian collections:
Back