Advanced search
1 file | 701.36 KB Add to list

LLMs as chainsaws : evaluating open-weights generative LLMs for extracting fauna and flora from multilingual travelogues

Tess Dejaeghere (UGent) , Els Lefever (UGent) and Julie M. Birkholz (UGent)
Author
Organization
Project
Abstract
Named Entity Recognition (NER) is crucial in literary-historical research for tasks such as semantic indexing and entity linking. However, historical texts pose challenges for implementing said tasks due to language variations, OCR errors, and poor performance of off-the-shelf annotation tools. Generative Large Language Models (LLMs) present both novel opportunities and challenges in humanities research. These models, while powerful, raise valid concerns regarding biases, hallucinations, and opacity - making their evaluation for the Digital Humanities (DH) community all the more urgent. In response, we propose our work on the evaluation of 3 quantized open-weights LLMs (mistral-7b-instruct-v0.1, nous-hermes-llama2-13b, Meta-Llama-3-8B-instruct) through GPT4ALL for NER on literary-historical travelogues from the 18th to 20th centuries in English, French, Dutch, and German. All models were assessed both quantitatively and qualitatively across 5 incrementally more complex prompts - revealing common error types such as bias, parsing issues, the addition of redundant information, entity adaptations and hallucinations. We analyse prevalent examples per language, century, prompt and model. Our contributions include a publicly accessible annotated dataset, pioneering insights into LLMs’ performance in literary-historical contexts, and the publication of reusable workflows for utilizing and evaluating LLMs in humanities research.
Keywords
LT3

Downloads

  • CLIN2024 dejaeghere.pdf
    • full text (Published version)
    • |
    • open access
    • |
    • PDF
    • |
    • 701.36 KB

Citation

Please use this url to cite or link to this publication:

MLA
Dejaeghere, Tess, et al. “LLMs as Chainsaws : Evaluating Open-Weights Generative LLMs for Extracting Fauna and Flora from Multilingual Travelogues.” COMPUTATIONAL LINGUISTICS IN THE NETHERLANDS JOURNAL, vol. 14, 2025, pp. 255–78.
APA
Dejaeghere, T., Lefever, E., & Birkholz, J. M. (2025). LLMs as chainsaws : evaluating open-weights generative LLMs for extracting fauna and flora from multilingual travelogues. COMPUTATIONAL LINGUISTICS IN THE NETHERLANDS JOURNAL, 14, 255–278.
Chicago author-date
Dejaeghere, Tess, Els Lefever, and Julie M. Birkholz. 2025. “LLMs as Chainsaws : Evaluating Open-Weights Generative LLMs for Extracting Fauna and Flora from Multilingual Travelogues.” COMPUTATIONAL LINGUISTICS IN THE NETHERLANDS JOURNAL 14: 255–78.
Chicago author-date (all authors)
Dejaeghere, Tess, Els Lefever, and Julie M. Birkholz. 2025. “LLMs as Chainsaws : Evaluating Open-Weights Generative LLMs for Extracting Fauna and Flora from Multilingual Travelogues.” COMPUTATIONAL LINGUISTICS IN THE NETHERLANDS JOURNAL 14: 255–278.
Vancouver
1.
Dejaeghere T, Lefever E, Birkholz JM. LLMs as chainsaws : evaluating open-weights generative LLMs for extracting fauna and flora from multilingual travelogues. COMPUTATIONAL LINGUISTICS IN THE NETHERLANDS JOURNAL. 2025;14:255–78.
IEEE
[1]
T. Dejaeghere, E. Lefever, and J. M. Birkholz, “LLMs as chainsaws : evaluating open-weights generative LLMs for extracting fauna and flora from multilingual travelogues,” COMPUTATIONAL LINGUISTICS IN THE NETHERLANDS JOURNAL, vol. 14, pp. 255–278, 2025.
@article{01K3B64TJWM5NSEBT3VVSD23FB,
  abstract     = {{Named Entity Recognition (NER) is crucial in literary-historical research for tasks such as semantic indexing and entity linking. However, historical texts pose challenges for implementing said tasks due to language variations, OCR errors, and poor performance of off-the-shelf annotation tools. Generative Large Language Models (LLMs) present both novel opportunities and challenges in humanities research. These models, while powerful, raise valid concerns regarding biases, hallucinations, and opacity - making their evaluation for the Digital Humanities (DH) community all the more urgent. In response, we propose our work on the evaluation of 3 quantized open-weights LLMs (mistral-7b-instruct-v0.1, nous-hermes-llama2-13b, Meta-Llama-3-8B-instruct) through GPT4ALL for NER on literary-historical travelogues from the 18th to 20th centuries in English, French, Dutch, and German. All models were assessed both quantitatively and qualitatively across 5 incrementally more complex prompts - revealing common error types such as bias, parsing issues, the addition of redundant information, entity adaptations and hallucinations. We analyse prevalent examples per language, century, prompt and model. Our contributions include a publicly accessible annotated dataset, pioneering insights into LLMs’ performance in literary-historical contexts, and the publication of reusable workflows for utilizing and evaluating LLMs in humanities research.}},
  author       = {{Dejaeghere, Tess and Lefever, Els and Birkholz, Julie M.}},
  issn         = {{2211-4009}},
  journal      = {{COMPUTATIONAL LINGUISTICS IN THE NETHERLANDS JOURNAL}},
  keywords     = {{LT3}},
  language     = {{eng}},
  pages        = {{255--278}},
  title        = {{LLMs as chainsaws : evaluating open-weights generative LLMs for extracting fauna and flora from multilingual travelogues}},
  url          = {{https://www.clinjournal.org/clinj/article/view/199}},
  volume       = {{14}},
  year         = {{2025}},
}