Tencent launches the multimodal WeMM-Embedding family
Tencent’s WeChat Vision team announced multimodal embedding models for understanding and retrieving text, images, and video. The post also says engineers can evaluate them on multimodal retrieval tasks.
On August 26, 2026, Tencent’s WeChat Vision team announced WeMM-Embedding, a family of multimodal embedding models. According to the post, the models support search and recommendations involving text, images, and video.
The announcement says engineers can evaluate the models, described as open source, on multimodal retrieval tasks. To check the scope and findings, consult the original publication and verify which models, methods, and evaluations it presents; the supplied material gives no metrics or comparative results.