This chapter examines how multi-modal data like text, image, audio, and videos can enhance recommendation systems. It introduces core integration strategies (early, late, and hybrid fusion) and contrasts them with emerging multi-modal large language models (LLMs) in terms of architecture, training, and use cases.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Leveraging Multi-modal Data

  • Jianqiang Jay Wang

摘要

This chapter examines how multi-modal data like text, image, audio, and videos can enhance recommendation systems. It introduces core integration strategies (early, late, and hybrid fusion) and contrasts them with emerging multi-modal large language models (LLMs) in terms of architecture, training, and use cases.