Tencent's WeChat team open-sources multimodal embedding model
Tencent's WeChat Vision team has open-sourced WeMM-Embedding, a family of multimodal embedding models designed to represent and match text, images, and videos.

Tencent's WeChat Vision team has released WeMM-Embedding, a set of multimodal embedding models, as open source. These models are capable of representing and matching diverse content types including text, images, and videos.
The team stated that the models are already in use across several WeChat services, such as WeChat Channels, Official Accounts, Moments, and e-commerce platforms. The release includes three versions of the models: 2B, 4B, and 9B.
According to evaluation results published by Tencent, the 9B model achieved the top performance on both the MMEB-v2 and MMEB-v3 benchmarks compared to other listed models.
The model code, evaluation tools, and weights have been made available through a public repository and model pages, allowing developers to integrate these models into multimodal search, retrieval, and recommendation applications.