Adapting multimodal foundation models for text-to-3D retrieval with domain-specific vocabulary
- Author
- Jensen Wiedler, Jarne Van den Herrewegen (UGent) , Tom Tourwé, Thomas Demeester (UGent) and Francis wyffels (UGent)
- Organization
- Project
- Abstract
- This report addresses the challenge of applying multimodal foundation models such as Uni3D for text-to-shape retrieval in a domain-specific context. While the recent multimodal models show strong performance on zero-shot 3D classification tasks, we found that their out-of-the-box performance on our dental manufacturing dataset is underwhelming, achieving only 13.21% zero-shot test accuracy. With appropriate fine-tuning however, we managed to improve the text-to-3D classification performance to 86.02% accuracy – a fine demonstration for the transfer learning capabilities of recent multimodal foundation models.
- Keywords
- Transfer Learning, 3D Retrieval, Text-to-3D, Representation learning
Downloads
-
20.pdf
- full text
- |
- open access
- |
- |
- 388.08 KB
Citation
Please use this url to cite or link to this publication: http://hdl.handle.net/1854/LU-01J9NZHXK7B6PDRAT7Y4C0W4XR
- MLA
- Wiedler, Jensen, et al. “Adapting Multimodal Foundation Models for Text-to-3D Retrieval with Domain-Specific Vocabulary.” CV4Metaverse 2024, 3rd Computer Vision for Metaverse Workshop, Abstracts, 2024.
- APA
- Wiedler, J., Van den Herrewegen, J., Tourwé, T., Demeester, T., & wyffels, F. (2024). Adapting multimodal foundation models for text-to-3D retrieval with domain-specific vocabulary. CV4Metaverse 2024, 3rd Computer Vision for Metaverse Workshop, Abstracts. Presented at the CV4Metaverse 2024, 3rd Computer Vision for Metaverse Workshop, Milan, Italy.
- Chicago author-date
- Wiedler, Jensen, Jarne Van den Herrewegen, Tom Tourwé, Thomas Demeester, and Francis wyffels. 2024. “Adapting Multimodal Foundation Models for Text-to-3D Retrieval with Domain-Specific Vocabulary.” In CV4Metaverse 2024, 3rd Computer Vision for Metaverse Workshop, Abstracts.
- Chicago author-date (all authors)
- Wiedler, Jensen, Jarne Van den Herrewegen, Tom Tourwé, Thomas Demeester, and Francis wyffels. 2024. “Adapting Multimodal Foundation Models for Text-to-3D Retrieval with Domain-Specific Vocabulary.” In CV4Metaverse 2024, 3rd Computer Vision for Metaverse Workshop, Abstracts.
- Vancouver
- 1.Wiedler J, Van den Herrewegen J, Tourwé T, Demeester T, wyffels F. Adapting multimodal foundation models for text-to-3D retrieval with domain-specific vocabulary. In: CV4Metaverse 2024, 3rd Computer Vision for Metaverse Workshop, Abstracts. 2024.
- IEEE
- [1]J. Wiedler, J. Van den Herrewegen, T. Tourwé, T. Demeester, and F. wyffels, “Adapting multimodal foundation models for text-to-3D retrieval with domain-specific vocabulary,” in CV4Metaverse 2024, 3rd Computer Vision for Metaverse Workshop, Abstracts, Milan, Italy, 2024.
@inproceedings{01J9NZHXK7B6PDRAT7Y4C0W4XR,
abstract = {{This report addresses the challenge of applying multimodal foundation models such as Uni3D for text-to-shape retrieval in a domain-specific context. While the recent multimodal models show strong performance on zero-shot 3D classification tasks, we found that their out-of-the-box performance on our dental manufacturing dataset is underwhelming, achieving only 13.21% zero-shot test accuracy. With appropriate fine-tuning however, we managed to improve the text-to-3D classification performance to 86.02% accuracy – a fine demonstration for the transfer learning capabilities of recent multimodal foundation models.}},
author = {{Wiedler, Jensen and Van den Herrewegen, Jarne and Tourwé, Tom and Demeester, Thomas and wyffels, Francis}},
booktitle = {{CV4Metaverse 2024, 3rd Computer Vision for Metaverse Workshop, Abstracts}},
keywords = {{Transfer Learning,3D Retrieval,Text-to-3D,Representation learning}},
language = {{eng}},
location = {{Milan, Italy}},
pages = {{4}},
title = {{Adapting multimodal foundation models for text-to-3D retrieval with domain-specific vocabulary}},
url = {{https://sites.google.com/view/cv4metaverse-2024/}},
year = {{2024}},
}