Data collection and IPR in multilingual parallel corpora : Dutch parallel corpus
- Author
- Orphée De Clercq (UGent) and Maribel Montero Perez (UGent)
- Organization
- Abstract
- After three years of work the Dutch Parallel Corpus (DPC) project has reached an end. The finalized corpus is a ten-million-word high-quality sentence-aligned bidirectional parallel corpus of Dutch, English and French, with Dutch as central language. In this paper we present the corpus and try to formulate some basic data collection principles, based on the work that was carried out for the project. Building a corpus is a difficult and time-consuming task, especially when every text sample included has to be cleared from copyrights. The DPC is balanced according to five text types (literature, journalistic texts, instructive texts, administrative texts and texts treating external communication) and four translation directions (Dutch-English, English-Dutch, Dutch-French and French-Dutch). All the text material was cleared from copyrights. The data collection process necessitated the involvement of different text providers, which resulted in drawing up four different licence agreements. Problems such as an unknown source language, copyright issues and changes to the corpus design are discussed in close detail and illustrated with examples so as to be of help to future corpus compilers.
- Keywords
- Corpus Creation, IPR, Copyrights, Parallel Corpus, Steunpunt Diversiteit & Leren, language learning
Downloads
-
(...).pdf
- full text
- |
- UGent only
- |
- |
- 424.26 KB
Citation
Please use this url to cite or link to this publication: http://hdl.handle.net/1854/LU-1067765
- MLA
- De Clercq, Orphée, and Maribel Montero Perez. “Data Collection and IPR in Multilingual Parallel Corpora : Dutch Parallel Corpus.” LREC 2010 : Seventh Conference on International Language Resources and Evaluation, edited by Nicoletta Calzolari et al., 2010, pp. 3383–88.
- APA
- De Clercq, O., & Montero Perez, M. (2010). Data collection and IPR in multilingual parallel corpora : Dutch parallel corpus. In N. Calzolari, K. Choukri, B. Maegaard, J. Mariani, J. Odijk, S. Piperidis, … D. Tapias (Eds.), LREC 2010 : seventh conference on international language resources and evaluation (pp. 3383–3388).
- Chicago author-date
- De Clercq, Orphée, and Maribel Montero Perez. 2010. “Data Collection and IPR in Multilingual Parallel Corpora : Dutch Parallel Corpus.” In LREC 2010 : Seventh Conference on International Language Resources and Evaluation, edited by Nicoletta Calzolari, Khalid Choukri, Bente Maegaard, Joseph Mariani, Jan Odijk, Stelios Piperidis, Mike Rosner, and Daniel Tapias, 3383–88.
- Chicago author-date (all authors)
- De Clercq, Orphée, and Maribel Montero Perez. 2010. “Data Collection and IPR in Multilingual Parallel Corpora : Dutch Parallel Corpus.” In LREC 2010 : Seventh Conference on International Language Resources and Evaluation, ed by. Nicoletta Calzolari, Khalid Choukri, Bente Maegaard, Joseph Mariani, Jan Odijk, Stelios Piperidis, Mike Rosner, and Daniel Tapias, 3383–3388.
- Vancouver
- 1.De Clercq O, Montero Perez M. Data collection and IPR in multilingual parallel corpora : Dutch parallel corpus. In: Calzolari N, Choukri K, Maegaard B, Mariani J, Odijk J, Piperidis S, et al., editors. LREC 2010 : seventh conference on international language resources and evaluation. 2010. p. 3383–8.
- IEEE
- [1]O. De Clercq and M. Montero Perez, “Data collection and IPR in multilingual parallel corpora : Dutch parallel corpus,” in LREC 2010 : seventh conference on international language resources and evaluation, Valletta, Malta, 2010, pp. 3383–3388.
@inproceedings{1067765,
abstract = {{After three years of work the Dutch Parallel Corpus (DPC) project has reached an end. The finalized corpus is a ten-million-word high-quality sentence-aligned bidirectional parallel corpus of Dutch, English and French, with Dutch as central language. In this paper we present the corpus and try to formulate some basic data collection principles, based on the work that was carried out for the project. Building a corpus is a difficult and time-consuming task, especially when every text sample included has to be cleared from copyrights. The DPC is balanced according to five text types (literature, journalistic texts, instructive texts, administrative texts and texts treating external communication) and four translation directions (Dutch-English, English-Dutch, Dutch-French and French-Dutch). All the text material was cleared from copyrights. The data collection process necessitated the involvement of different text providers, which resulted in drawing up four different licence agreements. Problems such as an unknown source language, copyright issues and changes to the corpus design are discussed in close detail and illustrated with examples so as to be of help to future corpus compilers.}},
author = {{De Clercq, Orphée and Montero Perez, Maribel}},
booktitle = {{LREC 2010 : seventh conference on international language resources and evaluation}},
editor = {{Calzolari, Nicoletta and Choukri, Khalid and Maegaard, Bente and Mariani, Joseph and Odijk, Jan and Piperidis, Stelios and Rosner, Mike and Tapias, Daniel}},
isbn = {{9782951740860}},
keywords = {{Corpus Creation,IPR,Copyrights,Parallel Corpus,Steunpunt Diversiteit & Leren,language learning}},
language = {{eng}},
location = {{Valletta, Malta}},
pages = {{3383--3388}},
title = {{Data collection and IPR in multilingual parallel corpora : Dutch parallel corpus}},
year = {{2010}},
}