siddhant

Knowledge / Generative AI

Multimodal Generative Models

Models that jointly process text, images, audio, video, and other modalities.

By Siddhant Krishna · Published 2026-10-06 · Updated 2026-10-06

01

Modal Fusion

Multimodal systems must map different data types into representations that can interact. Fusion can occur through shared embeddings, cross-attention, token-like representations, encoders, decoders, or combinations of these mechanisms.

02

Multimodal Generation

  • Text-to-image.
  • Image-to-text.
  • Text-to-audio.
  • Speech-to-speech.
  • Text-to-video.
  • Image-to-video.
  • Multimodal reasoning and tool use.

References

  1. Vaswani et al. (2017), Advances in Neural Information Processing Systems.
    https://arxiv.org/abs/1706.03762
  2. Ho, Jain & Abbeel (2020).
    https://arxiv.org/abs/2006.11239
  3. Stanford Institute for Human-Centered Artificial Intelligence (2026).
    https://hai.stanford.edu/ai-index/2026-ai-index-report

Related

Contact

Get in Touch

Want to chat? Just shoot me a dm with a direct question on twitter and I'll respond whenever I can. I will ignore all soliciting.