Module 16 · Frontier
Beyond Text: Multimodal AI
In this module, we will learn how AI works with images and other types of data, and the generative models that create images from noise.
By the end of this module, we will know how models see images and how they generate new ones.
Lessons
- 16.1 Multimodal AI: Perceiving Text, Images, and Audio Together: The Big Picture · What is a Modality? · Unimodal AI vs Multimodal AI · Why Multimodal AI? · How Multimodal AI Works · Three Common Types of Multimodal AI · Real Examples of Multimodal AI · Use Cases of Multimodal AI · Common Mistakes to Avoid · Quick Summary
- 16.2 Vision Transformers: Applying Self-Attention to Image Patches: The Big Picture · Decoding Step 1: Splitting the Image into Patches · Decoding Step 2: Patch Embedding · Decoding Step 3: The CLS Token · Decoding Step 4: Position Embeddings · Decoding Step 5: The Transformer Encoder · Decoding Step 6: The Classification Head · Putting It All Together · ViT vs CNN · Quick Summary
- 16.3 Image Embeddings: Encoding Visual Content as Vectors: What is an embedding? · What is an image embedding? · Why do we need image embeddings? · How does a computer see an image? · How are image embeddings created? · A simple numeric walkthrough · How do we measure similarity between two embeddings? · A code example · Where are image embeddings used? · Summary
- 16.4 Diffusion Models: Iterative Denoising to Generate Images: What is a Diffusion Model? · Why do we need Diffusion Models? · The two processes: Forward and Reverse · The Forward Process (adding noise) · The Reverse Process (removing noise) · A step-by-step example walk-through · How the model is trained · A simple code example · Conditional Diffusion (text to image) · Advantages of Diffusion Models · Where Diffusion Models are used
- 16.5 GANs: A Generator and Discriminator in Constant Competition: What is a Generative Adversarial Network (GAN)? · The two players: Generator vs Discriminator · The counterfeiter vs police analogy · The adversarial training loop · The loss function and the minimax game in simple words · A tiny PyTorch-style code sketch · The mode collapse problem · Training stability · Types of GANs (DCGAN, Conditional GAN, StyleGAN, CycleGAN) · Real-world applications of GANs
- 16.6 Variational Autoencoders: Learning a Compressed Latent Space: What is an Autoencoder? · The problem with a normal Autoencoder · What is a Variational Autoencoder? · The encoder, the latent space, and the decoder · The reparameterization trick · The loss function of a Variational Autoencoder · A simple example walk-through · A simple code example · Advantages of Variational Autoencoders · Where Variational Autoencoders are used
← Module 15: Securing AI Systems · Module 17: Production AI Infrastructure →