Mimicking how humans interpret out-of-context sentences through controlled toxicity decoding
- Author
- Maria Mihaela Trusca and Liesbeth Allein (UGent)
- Organization
- Abstract
- Interpretations of a single sentence can vary, particularly when its context is lost. This paper aims to simulate how readers perceive content with varying toxicity levels by generating diverse interpretations of out-of-context sentences. By modeling toxicity we can anticipate misunderstandings and reveal hidden toxic meanings. Our proposed decoding strategy explicitly controls toxicity in the set of generated interpretations by (i) aligning interpretation toxicity with the input, (ii) relaxing toxicity constraints for more toxic input sentences, and (iii) promoting diversity in toxicity levels within the set of generated interpretations. Experimental results show that our method improves alignment with human-written interpretations in both syntax and semantics while reducing model prediction uncertainty.
Citation
Please use this url to cite or link to this publication: http://hdl.handle.net/1854/LU-01K951E4W4NZNWSF8VTK8RJ4QF
- MLA
- Trusca, Maria Mihaela, and Liesbeth Allein. “Mimicking How Humans Interpret Out-of-Context Sentences through Controlled Toxicity Decoding.” Proceedings of the 5th Workshop on Trustworthy NLP (TrustNLP 2025), Association for Computational Linguistics (ACL), 2025, pp. 291–97, doi:10.18653/v1/2025.trustnlp-main.19.
- APA
- Trusca, M. M., & Allein, L. (2025). Mimicking how humans interpret out-of-context sentences through controlled toxicity decoding. Proceedings of the 5th Workshop on Trustworthy NLP (TrustNLP 2025), 291–297. https://doi.org/10.18653/v1/2025.trustnlp-main.19
- Chicago author-date
- Trusca, Maria Mihaela, and Liesbeth Allein. 2025. “Mimicking How Humans Interpret Out-of-Context Sentences through Controlled Toxicity Decoding.” In Proceedings of the 5th Workshop on Trustworthy NLP (TrustNLP 2025), 291–97. Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2025.trustnlp-main.19.
- Chicago author-date (all authors)
- Trusca, Maria Mihaela, and Liesbeth Allein. 2025. “Mimicking How Humans Interpret Out-of-Context Sentences through Controlled Toxicity Decoding.” In Proceedings of the 5th Workshop on Trustworthy NLP (TrustNLP 2025), 291–297. Association for Computational Linguistics (ACL). doi:10.18653/v1/2025.trustnlp-main.19.
- Vancouver
- 1.Trusca MM, Allein L. Mimicking how humans interpret out-of-context sentences through controlled toxicity decoding. In: Proceedings of the 5th Workshop on Trustworthy NLP (TrustNLP 2025). Association for Computational Linguistics (ACL); 2025. p. 291–7.
- IEEE
- [1]M. M. Trusca and L. Allein, “Mimicking how humans interpret out-of-context sentences through controlled toxicity decoding,” in Proceedings of the 5th Workshop on Trustworthy NLP (TrustNLP 2025), Albuquerque, New Mexico, 2025, pp. 291–297.
@inproceedings{01K951E4W4NZNWSF8VTK8RJ4QF,
abstract = {{Interpretations of a single sentence can vary, particularly when its context is lost. This paper aims to simulate how readers perceive content with varying toxicity levels by generating diverse interpretations of out-of-context sentences. By modeling toxicity we can anticipate misunderstandings and reveal hidden toxic meanings. Our proposed decoding strategy explicitly controls toxicity in the set of generated interpretations by (i) aligning interpretation toxicity with the input, (ii) relaxing toxicity constraints for more toxic input sentences, and (iii) promoting diversity in toxicity levels within the set of generated interpretations. Experimental results show that our method improves alignment with human-written interpretations in both syntax and semantics while reducing model prediction uncertainty.}},
author = {{Trusca, Maria Mihaela and Allein, Liesbeth}},
booktitle = {{Proceedings of the 5th Workshop on Trustworthy NLP (TrustNLP 2025)}},
isbn = {{9798891762336}},
language = {{eng}},
location = {{Albuquerque, New Mexico}},
pages = {{291--297}},
publisher = {{Association for Computational Linguistics (ACL)}},
title = {{Mimicking how humans interpret out-of-context sentences through controlled toxicity decoding}},
url = {{http://doi.org/10.18653/v1/2025.trustnlp-main.19}},
year = {{2025}},
}
- Altmetric
- View in Altmetric