Advanced search
1 file | 388.08 KB Add to list

Adapting multimodal foundation models for text-to-3D retrieval with domain-specific vocabulary

Author
Organization
Project
Abstract
This report addresses the challenge of applying multimodal foundation models such as Uni3D for text-to-shape retrieval in a domain-specific context. While the recent multimodal models show strong performance on zero-shot 3D classification tasks, we found that their out-of-the-box performance on our dental manufacturing dataset is underwhelming, achieving only 13.21% zero-shot test accuracy. With appropriate fine-tuning however, we managed to improve the text-to-3D classification performance to 86.02% accuracy – a fine demonstration for the transfer learning capabilities of recent multimodal foundation models.
Keywords
Transfer Learning, 3D Retrieval, Text-to-3D, Representation learning

Downloads

  • 20.pdf
    • full text
    • |
    • open access
    • |
    • PDF
    • |
    • 388.08 KB

Citation

Please use this url to cite or link to this publication:

MLA
Wiedler, Jensen, et al. “Adapting Multimodal Foundation Models for Text-to-3D Retrieval with Domain-Specific Vocabulary.” CV4Metaverse 2024, 3rd Computer Vision for Metaverse Workshop, Abstracts, 2024.
APA
Wiedler, J., Van den Herrewegen, J., Tourwé, T., Demeester, T., & wyffels, F. (2024). Adapting multimodal foundation models for text-to-3D retrieval with domain-specific vocabulary. CV4Metaverse 2024, 3rd Computer Vision for Metaverse Workshop, Abstracts. Presented at the CV4Metaverse 2024, 3rd Computer Vision for Metaverse Workshop, Milan, Italy.
Chicago author-date
Wiedler, Jensen, Jarne Van den Herrewegen, Tom Tourwé, Thomas Demeester, and Francis wyffels. 2024. “Adapting Multimodal Foundation Models for Text-to-3D Retrieval with Domain-Specific Vocabulary.” In CV4Metaverse 2024, 3rd Computer Vision for Metaverse Workshop, Abstracts.
Chicago author-date (all authors)
Wiedler, Jensen, Jarne Van den Herrewegen, Tom Tourwé, Thomas Demeester, and Francis wyffels. 2024. “Adapting Multimodal Foundation Models for Text-to-3D Retrieval with Domain-Specific Vocabulary.” In CV4Metaverse 2024, 3rd Computer Vision for Metaverse Workshop, Abstracts.
Vancouver
1.
Wiedler J, Van den Herrewegen J, Tourwé T, Demeester T, wyffels F. Adapting multimodal foundation models for text-to-3D retrieval with domain-specific vocabulary. In: CV4Metaverse 2024, 3rd Computer Vision for Metaverse Workshop, Abstracts. 2024.
IEEE
[1]
J. Wiedler, J. Van den Herrewegen, T. Tourwé, T. Demeester, and F. wyffels, “Adapting multimodal foundation models for text-to-3D retrieval with domain-specific vocabulary,” in CV4Metaverse 2024, 3rd Computer Vision for Metaverse Workshop, Abstracts, Milan, Italy, 2024.
@inproceedings{01J9NZHXK7B6PDRAT7Y4C0W4XR,
  abstract     = {{This report addresses the challenge of applying multimodal foundation models such as Uni3D for text-to-shape retrieval in a domain-specific context. While the recent multimodal models show strong performance on zero-shot 3D classification tasks, we found that their out-of-the-box performance on our dental manufacturing dataset is underwhelming, achieving only 13.21% zero-shot test accuracy. With appropriate fine-tuning however, we managed to improve the text-to-3D classification performance to 86.02% accuracy – a fine demonstration for the transfer learning capabilities of recent multimodal foundation models.}},
  author       = {{Wiedler, Jensen and Van den Herrewegen, Jarne and Tourwé, Tom and Demeester, Thomas and wyffels, Francis}},
  booktitle    = {{CV4Metaverse 2024, 3rd Computer Vision for Metaverse Workshop, Abstracts}},
  keywords     = {{Transfer Learning,3D Retrieval,Text-to-3D,Representation learning}},
  language     = {{eng}},
  location     = {{Milan, Italy}},
  pages        = {{4}},
  title        = {{Adapting multimodal foundation models for text-to-3D retrieval with domain-specific vocabulary}},
  url          = {{https://sites.google.com/view/cv4metaverse-2024/}},
  year         = {{2024}},
}