Advanced search
Add to list

Replication Data for: The semantic structuring of minimizing constructions in present-day Netherlandic Dutch: a distribution-based cluster analysis

Margot Van den Heede (UGent) and Peter Lauwers (UGent)
(2025)
Author
Organization
Abstract
Dataset abstract: This dataset contains the data files that were used for the cluster analysis of the Dutch minimizing construction, as described in the publication cited below. In addition to a ReadMe file, it contains three files: A txt file is provided with the corpus queries that were used to find tokens of the minimizing constructions in the Dutch Web 2014 (nlTenTen14) corpus, available via Sketch Engine (more information about the TenTen corpora: Jakubíček, M., A. Kilgarriff, V. Kovář, P. Rychlý & V. Suchomel (2013). The TenTen corpus family. In: 7th International Corpus Linguistics Conference CL. Lancaster, 125–127). A csv file is provided that forms the input file for the cluster analysis. It contains a list of 5,863 minimizer-predicate combinations, more specifically a list of the predicates that are combined with the minimizers that have a token frequency of at least 10 in my dataset. An R-script is provided with the code to perform the cluster analysis in R.
Keywords
minimizing constructions, Netherlandic Dutch, cluster analysis, corpus data, Construction Grammar
License
CC0-1.0
Access
open access

Citation

Please use this url to cite or link to this publication:

@misc{01K476HW3TZX2HTGKV977BK08Z,
  abstract     = {{Dataset abstract: This dataset contains the data files that were used for the cluster analysis of the Dutch minimizing construction, as described in the publication cited below. In addition to a ReadMe file, it contains three files: A txt file is provided with the corpus queries that were used to find tokens of the minimizing constructions in the Dutch Web 2014 (nlTenTen14) corpus, available via Sketch Engine (more information about the TenTen corpora: Jakubíček, M., A. Kilgarriff, V. Kovář, P. Rychlý & V. Suchomel (2013). The TenTen corpus family. In: 7th International Corpus Linguistics Conference CL. Lancaster, 125–127). A csv file is provided that forms the input file for the cluster analysis. It contains a list of 5,863 minimizer-predicate combinations, more specifically a list of the predicates that are combined with the minimizers that have a token frequency of at least 10 in my dataset. An R-script is provided with the code to perform the cluster analysis in R.}},
  author       = {{Van den Heede, Margot and Lauwers, Peter}},
  keywords     = {{minimizing constructions,Netherlandic Dutch,cluster analysis,corpus data,Construction Grammar}},
  language     = {{eng}},
  publisher    = {{DataverseNO}},
  title        = {{Replication Data for: The semantic structuring of minimizing constructions in present-day Netherlandic Dutch: a distribution-based cluster analysis}},
  url          = {{http://doi.org/10.18710/GIKMKM}},
  year         = {{2025}},
}

Altmetric
View in Altmetric