01
Modal Fusion
Multimodal systems must map different data types into representations that can interact. Fusion can occur through shared embeddings, cross-attention, token-like representations, encoders, decoders, or combinations of these mechanisms.
02
Multimodal Generation
- Text-to-image.
- Image-to-text.
- Text-to-audio.
- Speech-to-speech.
- Text-to-video.
- Image-to-video.
- Multimodal reasoning and tool use.