Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

RCML conditions multimodal representation learning on semantic relations, letting the same sample take different embeddings under different relational contexts and consistently beating strong baselines on retrieval and classification in zero-shot, fine-tuned, and out-of-domain settings

The work proposes Relation-Conditioned Multimodal Learning (RCML), a framework that treats natural-language relation descriptions as explicit conditions of multimodal representation learning by constructing relation-aware training pairs, introducing a relation-conditioned module to adapt embeddings to relation semantics, and employing a unified contrastive objective that jointly models cross-modal alignment and relation-induced inter-sample structure, so the same sample can be represented differently under different relational contexts; experiments on multiple datasets show RCML consistently outperforms strong baselines on retrieval and classification tasks in zero-shot, fine-tuned, and out-of-domain settings.