WeMM-Embedding-9B maps modalities into a shared space
Published on August 25, 2026, the post describes a model that represents text, images, videos, and visual documents in a shared space, and claims state-of-the-art results on MMEB-v2 and MMEB-v3.
A post published on August 25, 2026 describes Tencent's WeMM-Embedding-9B as a model that maps text, images, videos, and visual documents into a shared embedding space. The post says the model achieves state-of-the-art results on the MMEB-v2 and MMEB-v3 benchmarks.
This kind of representation may be useful in retrieval systems that search textual and visual content jointly. To assess the claims, consult the original post and check the results and evaluation conditions in the MMEB-v2 and MMEB-v3 benchmark materials; the summary does not provide metrics or testing details.