- Author
- Florian Debaene (UGent)
- Organization
- Project
- Abstract
- Raw text extraction of 466 early modern Dutch comedies and farces after nltk sentence tokenization, with author indications. Gold data comes from DBNL and CENETON, OCR data is the raw output of Google Books scans by Transkribus Print M1.
- Keywords
- early modern Dutch theatre, comedy, farce, corpus, OCR
- License
- CC-BY-SA-4.0
- Access
- open access
Citation
Please use this url to cite or link to this publication: http://hdl.handle.net/1854/LU-01JWAXV73EVSXDD1CZHY22RDPQ
@misc{01JWAXV73EVSXDD1CZHY22RDPQ,
abstract = {{Raw text extraction of 466 early modern Dutch comedies and farces after nltk sentence tokenization, with author indications. Gold data comes from DBNL and CENETON, OCR data is the raw output of Google Books scans by Transkribus Print M1.}},
author = {{Debaene, Florian}},
keywords = {{early modern Dutch theatre,comedy,farce,corpus,OCR}},
language = {{dut}},
publisher = {{Hugging Face}},
title = {{EmDComF_raw}},
url = {{http://doi.org/10.57967/HF/5640}},
year = {{2025}},
}
- Altmetric
- View in Altmetric