Advanced search
1 file | 424.26 KB Add to list

Data collection and IPR in multilingual parallel corpora : Dutch parallel corpus

Author
Organization
Abstract
After three years of work the Dutch Parallel Corpus (DPC) project has reached an end. The finalized corpus is a ten-million-word high-quality sentence-aligned bidirectional parallel corpus of Dutch, English and French, with Dutch as central language. In this paper we present the corpus and try to formulate some basic data collection principles, based on the work that was carried out for the project. Building a corpus is a difficult and time-consuming task, especially when every text sample included has to be cleared from copyrights. The DPC is balanced according to five text types (literature, journalistic texts, instructive texts, administrative texts and texts treating external communication) and four translation directions (Dutch-English, English-Dutch, Dutch-French and French-Dutch). All the text material was cleared from copyrights. The data collection process necessitated the involvement of different text providers, which resulted in drawing up four different licence agreements. Problems such as an unknown source language, copyright issues and changes to the corpus design are discussed in close detail and illustrated with examples so as to be of help to future corpus compilers.
Keywords
Corpus Creation, IPR, Copyrights, Parallel Corpus, Steunpunt Diversiteit & Leren, language learning

Downloads

  • (...).pdf
    • full text
    • |
    • UGent only
    • |
    • PDF
    • |
    • 424.26 KB

Citation

Please use this url to cite or link to this publication:

MLA
De Clercq, Orphée, and Maribel Montero Perez. “Data Collection and IPR in Multilingual Parallel Corpora : Dutch Parallel Corpus.” LREC 2010 : Seventh Conference on International Language Resources and Evaluation, edited by Nicoletta Calzolari et al., 2010, pp. 3383–88.
APA
De Clercq, O., & Montero Perez, M. (2010). Data collection and IPR in multilingual parallel corpora : Dutch parallel corpus. In N. Calzolari, K. Choukri, B. Maegaard, J. Mariani, J. Odijk, S. Piperidis, … D. Tapias (Eds.), LREC 2010 : seventh conference on international language resources and evaluation (pp. 3383–3388).
Chicago author-date
De Clercq, Orphée, and Maribel Montero Perez. 2010. “Data Collection and IPR in Multilingual Parallel Corpora : Dutch Parallel Corpus.” In LREC 2010 : Seventh Conference on International Language Resources and Evaluation, edited by Nicoletta Calzolari, Khalid Choukri, Bente Maegaard, Joseph Mariani, Jan Odijk, Stelios Piperidis, Mike Rosner, and Daniel Tapias, 3383–88.
Chicago author-date (all authors)
De Clercq, Orphée, and Maribel Montero Perez. 2010. “Data Collection and IPR in Multilingual Parallel Corpora : Dutch Parallel Corpus.” In LREC 2010 : Seventh Conference on International Language Resources and Evaluation, ed by. Nicoletta Calzolari, Khalid Choukri, Bente Maegaard, Joseph Mariani, Jan Odijk, Stelios Piperidis, Mike Rosner, and Daniel Tapias, 3383–3388.
Vancouver
1.
De Clercq O, Montero Perez M. Data collection and IPR in multilingual parallel corpora : Dutch parallel corpus. In: Calzolari N, Choukri K, Maegaard B, Mariani J, Odijk J, Piperidis S, et al., editors. LREC 2010 : seventh conference on international language resources and evaluation. 2010. p. 3383–8.
IEEE
[1]
O. De Clercq and M. Montero Perez, “Data collection and IPR in multilingual parallel corpora : Dutch parallel corpus,” in LREC 2010 : seventh conference on international language resources and evaluation, Valletta, Malta, 2010, pp. 3383–3388.
@inproceedings{1067765,
  abstract     = {{After three years of work the Dutch Parallel Corpus (DPC) project has reached an end. The finalized corpus is a ten-million-word high-quality sentence-aligned bidirectional parallel corpus of Dutch, English and French, with Dutch as central language. In this paper we present the corpus and try to formulate some basic data collection principles, based on the work that was carried out for the project. Building a corpus is a difficult and time-consuming task, especially when every text sample included has to be cleared from copyrights. The DPC is balanced according to five text types (literature, journalistic texts, instructive texts, administrative texts and texts treating external communication) and four translation directions (Dutch-English, English-Dutch, Dutch-French and French-Dutch). All the text material was cleared from copyrights. The data collection process necessitated the involvement of different text providers, which resulted in drawing up four different licence agreements. Problems such as an unknown source language, copyright issues and changes to the corpus design are discussed in close detail and illustrated with examples so as to be of help to future corpus compilers.}},
  author       = {{De Clercq, Orphée and Montero Perez, Maribel}},
  booktitle    = {{LREC 2010 : seventh conference on international language resources and evaluation}},
  editor       = {{Calzolari, Nicoletta and Choukri, Khalid and Maegaard, Bente and Mariani, Joseph and Odijk, Jan and Piperidis, Stelios and Rosner, Mike and Tapias, Daniel}},
  isbn         = {{9782951740860}},
  keywords     = {{Corpus Creation,IPR,Copyrights,Parallel Corpus,Steunpunt Diversiteit & Leren,language learning}},
  language     = {{eng}},
  location     = {{Valletta, Malta}},
  pages        = {{3383--3388}},
  title        = {{Data collection and IPR in multilingual parallel corpora : Dutch parallel corpus}},
  year         = {{2010}},
}

Web of Science
Times cited: