Hugging Face's Sentence Transformers library v5.4 adds multimodal embedding and reranker capabilities, as announced by Tom Aarsen. The update enables encoding and comparing text, images, audio, and video through a unified API. Multimodal embedding models map different input types into a shared embedding space, while reranker models score relevance across mixed-modality pairs. Supported use cases include visual document retrieval, cross-modal search, and multimodal RAG pipelines.
No score is assigned. Sources and their independence are shown in the citation chain below.