A novel approach to context-aware and responsible short text clustering
(2026)
FLEXIBLE QUERY ANSWERING SYSTEMS, FQAS 2025.
In Lecture Notes in Computer Science
16119.
p.19-30
- Author
- Maxime Deforche (UGent) , Ilse De Vos (UGent) and Guy De Tré (UGent)
- Organization
- Project
- Abstract
- Context-aware clustering of texts, particularly short texts, is a challenging task. Although metadata can provide valuable contextual cues to enhance clustering quality, such information is often incomplete or inconsistently available in real-world datasets. In this paper, we propose a novel clustering strategy that integrates both textual and metadata-based similarities, even when metadata is partially missing. Our method employs the Ordered Weighted Averaging (OWA) aggregator to fuse multiple similarity scores into a single aggregated value for each pair of texts. To handle missing metadata, we adapt the OWA mechanism by renormalising weights based only on available information, thereby avoiding potentially unreliable imputation or complete exclusion of certain metadata. We further introduce a confidence score that quantifies the reliability of each aggregated similarity, reflecting the pro-portion of missing metadata. Clustering is then performed using the K-Medoids algorithm on the resulting dissimilarity matrix. We demonstrate this approach on a real-world dataset of short Byzantine poems, where orthographic similarity is complemented with sparse metadata. The final clusters, stored in a graph database along with their confidence scores, enable meaningful interpretation and visualisation of the results, including the identification of uncertain cluster assignments due to missing contextual information.
- Keywords
- context-aware clustering, text clustering, missing data aggregation, OWA method, cluster confidence
Downloads
-
(...).pdf
- full text (Published version)
- |
- UGent only
- |
- |
- 372.64 KB
Citation
Please use this url to cite or link to this publication: http://hdl.handle.net/1854/LU-01K50XCMZW3MXY20MWNNM5601T
- MLA
- Deforche, Maxime, et al. “A Novel Approach to Context-Aware and Responsible Short Text Clustering.” FLEXIBLE QUERY ANSWERING SYSTEMS, FQAS 2025, edited by Guy De Tré et al., vol. 16119, Springer Cham, 2026, pp. 19–30, doi:10.1007/978-3-032-05607-8_4.
- APA
- Deforche, M., De Vos, I., & De Tré, G. (2026). A novel approach to context-aware and responsible short text clustering. In G. De Tré, S. Sotirov, J. Kacprzyk, G. Psaila, G. Smits, T. Andreasen, … H. L. Larsen (Eds.), FLEXIBLE QUERY ANSWERING SYSTEMS, FQAS 2025 (Vol. 16119, pp. 19–30). https://doi.org/10.1007/978-3-032-05607-8_4
- Chicago author-date
- Deforche, Maxime, Ilse De Vos, and Guy De Tré. 2026. “A Novel Approach to Context-Aware and Responsible Short Text Clustering.” In FLEXIBLE QUERY ANSWERING SYSTEMS, FQAS 2025, edited by Guy De Tré, Sotir Sotirov, Janusz Kacprzyk, Giuseppe Psaila, Grégory Smits, Troels Andreasen, Gloria Bordogna, and Henrik Legind Larsen, 16119:19–30. Springer Cham. https://doi.org/10.1007/978-3-032-05607-8_4.
- Chicago author-date (all authors)
- Deforche, Maxime, Ilse De Vos, and Guy De Tré. 2026. “A Novel Approach to Context-Aware and Responsible Short Text Clustering.” In FLEXIBLE QUERY ANSWERING SYSTEMS, FQAS 2025, ed by. Guy De Tré, Sotir Sotirov, Janusz Kacprzyk, Giuseppe Psaila, Grégory Smits, Troels Andreasen, Gloria Bordogna, and Henrik Legind Larsen, 16119:19–30. Springer Cham. doi:10.1007/978-3-032-05607-8_4.
- Vancouver
- 1.Deforche M, De Vos I, De Tré G. A novel approach to context-aware and responsible short text clustering. In: De Tré G, Sotirov S, Kacprzyk J, Psaila G, Smits G, Andreasen T, et al., editors. FLEXIBLE QUERY ANSWERING SYSTEMS, FQAS 2025. Springer Cham; 2026. p. 19–30.
- IEEE
- [1]M. Deforche, I. De Vos, and G. De Tré, “A novel approach to context-aware and responsible short text clustering,” in FLEXIBLE QUERY ANSWERING SYSTEMS, FQAS 2025, Burgas, Bulgaria, 2026, vol. 16119, pp. 19–30.
@inproceedings{01K50XCMZW3MXY20MWNNM5601T,
abstract = {{Context-aware clustering of texts, particularly short texts, is a challenging task. Although metadata can provide valuable contextual cues to enhance clustering quality, such information is often incomplete or inconsistently available in real-world datasets. In this paper, we propose a novel clustering strategy that integrates both textual and metadata-based similarities, even when metadata is partially missing. Our method employs the Ordered Weighted Averaging (OWA) aggregator to fuse multiple similarity scores into a single aggregated value for each pair of texts. To handle missing metadata, we adapt the OWA mechanism by renormalising weights based only on available information, thereby avoiding potentially unreliable imputation or complete exclusion of certain metadata. We further introduce a confidence score that quantifies the reliability of each aggregated similarity, reflecting the pro-portion of missing metadata. Clustering is then performed using the K-Medoids algorithm on the resulting dissimilarity matrix. We demonstrate this approach on a real-world dataset of short Byzantine poems, where orthographic similarity is complemented with sparse metadata. The final clusters, stored in a graph database along with their confidence scores, enable meaningful interpretation and visualisation of the results, including the identification of uncertain cluster assignments due to missing contextual information.}},
author = {{Deforche, Maxime and De Vos, Ilse and De Tré, Guy}},
booktitle = {{FLEXIBLE QUERY ANSWERING SYSTEMS, FQAS 2025}},
editor = {{De Tré, Guy and Sotirov, Sotir and Kacprzyk, Janusz and Psaila, Giuseppe and Smits, Grégory and Andreasen, Troels and Bordogna, Gloria and Larsen, Henrik Legind}},
isbn = {{9783032056061}},
issn = {{0302-9743}},
keywords = {{context-aware clustering,text clustering,missing data aggregation,OWA method,cluster confidence}},
language = {{eng}},
location = {{Burgas, Bulgaria}},
pages = {{19--30}},
publisher = {{Springer Cham}},
title = {{A novel approach to context-aware and responsible short text clustering}},
url = {{http://doi.org/10.1007/978-3-032-05607-8_4}},
volume = {{16119}},
year = {{2026}},
}
- Altmetric
- View in Altmetric
- Web of Science
- Times cited: