News
Comparison and generation: Douyin SOTA multi-modal representation model DME
1 min read
Source: aixq.cc
The Douyin search multimodal team teamed up with the Hillhouse School of Artificial Intelligence at Renmin University of China to release the Douyin multimodal representation model DME (Douyin Multimodal Embedding). On the multimodal representation authoritative evaluation list MMEB-v2 (78 data sets, covering the three major fields of images, videos, and visual documents), the DME model achieved the corresponding scale SOTA at the two parameter levels of 2B and 9B, and the overall score reached 74.8 (2B), 78.4 (9B), the advantages of video and visual document retrieval are particularly prominent. DME's technology has been deployed into Douyin's online system: the overall internal offline evaluation set has a relative improvement of 2.92%; the online A/B experiment verified that the core business indicators have an L of 0.1%