Modern AI Engineering

Lesson 16.1 · 23 min

Multimodal AI: Perceiving Text, Images, and Audio Together

A customer sends our support bot a photo of a cracked blender jar and a voice note saying 'it leaks here' — how can one model understand both at once?

In short: Multimodal AI is AI that takes in or produces more than one kind of data, such as text, images, audio and video. It works by turning each kind of data into vectors with its own encoder, mapping those vectors into a shared space, and letting one model reason over all of them together. It is what powers photo-aware chatbots, text-to-image tools, image search by description and video understanding.

The big picture

Humans never experience the world through one channel. When a friend shows us a photo and says 'look at this crack', we combine what we see with what we hear without thinking. For years, most AI systems could only do one of these things. A language model read text. An image classifier looked at pixels. A speech recognizer listened to audio. Each one lived in its own world.

Multimodal AI breaks down those walls. A multimodal model can accept a mix of inputs — say a photo plus a question — and answer in text, or take a sentence and draw a picture. Throughout this lesson we will follow one running example: a support chatbot for a kitchen-appliance shop. Customers send photos of broken products, screenshots of error codes, and voice notes. A text-only bot would be stuck; a multimodal bot can look at the photo and answer.

Think of it like a doctor's visit A good doctor does not diagnose from one clue. They listen to what we say (text/audio), look at the rash (image), and read the X-ray (another image type). Each clue alone is weak; together they point to one answer. Multimodal AI tries to do the same: combine several weak or partial signals into one strong understanding.

What is a modality?

A modality is a type of data, defined by how the information is captured and structured. The word comes from 'mode' — a mode of perceiving. The common modalities in AI are:

  • Text — a sequence of characters, split into tokens (word pieces).
  • Images — a grid of pixels, each pixel a few numbers (red, green, blue).
  • Audio — a long list of air-pressure samples over time, often turned into a spectrogram (a picture of which frequencies are loud at each moment).
  • Video — a sequence of images (frames), usually with an audio track.
  • Others — sensor readings, depth maps, tables, code, molecules, robot joint angles. Anything with its own structure can be a modality.

Modalities differ in shape and meaning. A sentence is a short 1-D sequence of discrete symbols. A photo is a dense 2-D grid of continuous values; a 1024×1024 colour image holds over three million numbers. Ten seconds of 16 kHz audio is 160,000 samples. No single raw format fits them all, and that is the core engineering problem multimodal AI has to solve.

Pause and think: Is a screenshot of an error message 'text' or 'image' as a modality?

It is an image: the model receives pixels, not characters. A multimodal model has to read the text out of the pixels (like OCR does). That is why a model may misread a blurry screenshot that it would have understood perfectly as typed text.

Unimodal AI vs multimodal AI

A unimodal model handles exactly one modality in and usually one out: a text classifier, an image classifier, a speech-to-text system. A multimodal model handles two or more, either on the input side, the output side, or both.

Why multimodal AI?

There are four strong reasons to combine modalities:

  • Richer context. A photo of a cracked jar says in one shot what would take a paragraph to describe — and customers are bad at describing. 'It is broken near the bottom' is vague; the picture is exact.
  • Disambiguation. One modality can resolve ambiguity in another. The word 'bank' is ambiguous; a photo of a river is not. In video, lip movement helps speech recognition in a noisy room.
  • Robustness. If one signal is missing or noisy (blurry photo, background noise), the others can still carry the answer.
  • New abilities. Some tasks are multimodal by definition: describing an image for a blind user, generating an image from a prompt, searching photos with a sentence, answering questions about a chart.

There is also a training reason. The internet holds billions of images with nearby text (captions, alt text, surrounding articles). That pairing is free supervision: the text tells the model what the picture shows, without anyone labelling it by hand. Much of modern multimodal AI was built by learning from such pairs.

How multimodal AI works

Almost every modern multimodal system follows the same recipe. The key insight: neural networks only understand vectors (lists of numbers). So the job is to turn every modality into vectors that live in one shared space, where 'a photo of a dog' and the words 'a dog' end up close together.

The combining step is called fusion, and there are three classic ways to do it:

Three ways to fuse modalities

  1. Early fusion: Combine the raw or lightly processed inputs first, then run one model over everything. Example: turn image patches and text into tokens and feed them all into one transformer from the first layer. Lets modalities interact deeply, but needs lots of paired training data.
  2. Late fusion: Run a separate model per modality all the way to a prediction or embedding, then combine only at the end (average the scores, or compare the two embeddings). Simple and modular, but the modalities never 'talk' in detail. CLIP-style image-text matching is a form of this.
  3. Intermediate (hybrid) fusion: Encode each modality separately, then mix them in the middle with cross-attention layers or by inserting projected image tokens into a language model. This is how most vision-language chat models are built today: a pre-trained vision encoder, a projector, and a pre-trained LLM.
  4. Train the pieces to agree: Whatever the fusion style, training must teach the model that matching inputs belong together. Contrastive training pulls matching image-text pairs close and pushes mismatched pairs apart; instruction tuning then teaches the model to answer questions about images.

Images cost tokens When an image is inserted into a language model, it becomes many vectors — one per patch. A single image can become hundreds to a few thousand tokens depending on the model and the image resolution. This is why multimodal requests cost more and why many APIs resize or tile large images.

Three common types of multimodal AI

Most systems we meet in practice fall into one of three families. They differ in which direction information flows.

The three families, using our support-bot scenario
TypeDirectionExample taskTypical models
Multimodal understandingMany modalities in → text outPhoto of the jar + question → 'The crack is on the base; that part is covered for 2 years.'Vision-language chat models (GPT-4o, Gemini, Claude, LLaVA)
Cross-modal generationOne modality in → a different one out'Draw a diagram showing how to seat the jar' → image; text → speechText-to-image diffusion models, text-to-speech, text-to-video
Cross-modal retrieval / matchingCompare modalities in a shared embedding spaceCustomer types 'jar with a cracked base' → find matching photos from past ticketsCLIP-style dual encoders, multimodal embedding models

Some recent models are called any-to-any or natively multimodal: one network both reads and writes several modalities (for example, it can listen to speech and reply with speech directly). Exactly which inputs and outputs a given product supports changes often, so always check the current documentation of the model we plan to use.

Pause and think: Our shop wants a search box where staff type 'scratched stainless steel kettle' and get matching photos from 50,000 old tickets. Which family is that, and does it need a chat model?

Cross-modal retrieval. We embed every photo once with an image encoder, embed the query with the matching text encoder, and find the nearest vectors. No chat model is needed, which makes it much cheaper and faster than asking a vision-language model to look at 50,000 images.

Code: teaching two modalities to share a space

Let us build the heart of cross-modal matching from scratch. We pretend we already have an image encoder (it outputs 6 numbers per image) and a text encoder (5 numbers per caption). Their outputs live in different spaces with different sizes, so we cannot compare them directly. We learn two small projection heads that map both into one 3-d shared space, using a contrastive loss: for each image, the matching caption should score highest among all captions in the batch.

shared_space.py

import numpy as np
rng = np.random.default_rng(0)
# 4 concepts. Each "image" is a 6-d feature vector, each "caption" a 5-d one.
# The two modalities have different sizes and different meanings per dimension.
concepts = ["dog", "cat", "car", "pizza"]
img = rng.normal(size=(4, 6))          # pretend output of an image encoder
txt = rng.normal(size=(4, 5))          # pretend output of a text encoder
# Two projection heads map both modalities into one shared 3-d space.
W_img = rng.normal(scale=0.1, size=(6, 3))
W_txt = rng.normal(scale=0.1, size=(5, 3))
def match_accuracy():
sims = (img @ W_img) @ (txt @ W_txt).T     # 4x4 image-text scores
return (sims.argmax(axis=1) == np.arange(4)).mean(), sims
print("before training: accuracy =", match_accuracy()[0])
lr = 0.1
for step in range(300):
I, T = img @ W_img, txt @ W_txt
logits = I @ T.T                          # row i: image i vs every caption
p = np.exp(logits - logits.max(1, keepdims=True))
p /= p.sum(1, keepdims=True)              # softmax over captions
loss = -np.log(p[np.arange(4), np.arange(4)]).mean()
g = (p - np.eye(4)) / 4                   # d loss / d logits
W_img -= lr * img.T @ (g @ T)             # chain rule into each head
W_txt -= lr * txt.T @ (g.T @ I)
if step in (0, 299):
print(f"step {step:3d}: contrastive loss = {loss:.3f}")
acc, sims = match_accuracy()
print("after training: accuracy =", acc)
for i, c in enumerate(concepts):
print(f"image of {c:5s} -> best caption: '{concepts[sims[i].argmax()]}'")

Output:

before training: accuracy = 0.25
step   0: contrastive loss = 1.379
step 299: contrastive loss = 0.007
after training: accuracy = 1.0
image of dog   -> best caption: 'dog'
image of cat   -> best caption: 'cat'
image of car   -> best caption: 'car'
image of pizza -> best caption: 'pizza'

Real systems do the same thing at enormous scale: hundreds of millions of image-text pairs, large batches (so each image must beat many wrong captions), normalised vectors, a learned temperature, and the loss computed in both directions (image→text and text→image).

Real examples of multimodal AI

Milestones in multimodal AI

  1. CLIP (OpenAI): Trained an image encoder and a text encoder together on image-text pairs from the web with a contrastive loss. Enabled zero-shot image classification by comparing an image to captions like 'a photo of a cat'.
  2. DALL·E (OpenAI): Generated images from text prompts, bringing text-to-image generation to wide attention.
  3. Flamingo (DeepMind): Connected a frozen vision encoder to a frozen language model with new cross-attention layers, so the LLM could answer questions about images with few examples.
  4. Stable Diffusion and Whisper: An open text-to-image diffusion model, and an open speech-recognition model trained on a large multilingual audio-text corpus.
  5. Vision-language chat models: GPT-4 with vision, open models like LLaVA (vision encoder + projector + LLM), and Google's Gemini, which was designed to be multimodal from the start.
  6. Native audio and any-to-any: Models such as GPT-4o handle text, images and audio in one network, enabling real-time voice conversations. Video understanding and generation have improved rapidly since.

What this looks like day to day Taking a photo of a restaurant menu in another language and asking for a translation; uploading a chart and asking what trend it shows; asking a phone assistant 'what am I looking at?' through the camera; generating a product mock-up from a sentence; automatic captions on videos. Each of these combines at least two modalities.

Use cases of multimodal AI

Where multimodal AI earns its keep
DomainModalitiesWhat it does
Customer supportPhoto + text (+ voice)Diagnose damage, read error screens, check receipts, route tickets
DocumentsScanned pages + textRead invoices, forms and charts that are images rather than clean text
AccessibilityImage/video → text/speechDescribe surroundings or photos to blind and low-vision users
Healthcare (research and assistive)Medical images + notesHelp draft radiology reports or flag findings, always with expert review
E-commerceImage ↔ textSearch by photo, 'find similar', auto-generate product descriptions
Robotics and drivingCamera + lidar + languageUnderstand scenes and follow spoken instructions
Content and mediaText → image/audio/videoDraft illustrations, voice-overs and video clips; moderate uploaded media

Common mistakes to avoid

Mistake 1: trusting fine details in images Vision-language models can confidently misread small text, count objects wrongly, confuse left and right, or 'see' details that are not there (visual hallucination). For anything that matters — a serial number, a dosage, a price — verify with a dedicated tool (OCR, a barcode reader) or ask the user to confirm.

  • Using a multimodal model when text is enough. If the input is already text, sending it as a screenshot wastes tokens and adds reading errors.
  • Ignoring image resolution. Images are often resized before the model sees them; tiny text may become unreadable. Crop to the relevant region or send a higher-detail version if the API allows it.
  • Assuming the modalities are balanced. Models often lean on the text and ignore the image (or vice versa). Test cases where the image contradicts the text to see which one wins.
  • Forgetting privacy. Photos and voice notes leak far more than intended: faces, addresses on envelopes, other people in the background. Strip metadata and get consent.
  • Evaluating only on clean data. Real user photos are blurry, dark and rotated. Build the test set from real traffic.

When should we not use multimodal AI? When the extra modality adds no information, when latency and cost budgets are tight, or when a simple, specialised tool (a barcode scanner, a classic OCR engine, a speech-to-text API followed by a text model) solves the problem more reliably.

Worked example, step by step

Let us follow one support request through the recipe and count what the reasoning model really receives. The customer sends one photo of the cracked jar and types 'Is this covered by warranty?'. All sizes below are illustrative; real models differ, but the arithmetic is the same.

One photo plus one question, counted

  1. Resize the photo: Suppose the vision encoder expects 336×336 pixels and uses 14×14-pixel patches. The phone photo is shrunk to that size first, whatever its original resolution.
  2. Count the patches: 336 / 14 = 24 patches per side, so 24 × 24 = 576 patch vectors come out of the encoder.
  3. Project them: Say each patch vector has 1,024 numbers and the language model works with 4,096. The projector maps 1,024 → 4,096 for each of the 576 vectors. The count stays 576; only the width changes.
  4. Tokenise the question: 'Is this covered by warranty?' becomes roughly 7 text tokens, each looked up as a 4,096-number vector.
  5. Fuse: We place the image vectors and the text vectors in one sequence: 576 + 7 = 583 tokens. About 99% of the sequence is the photo.
  6. Reason: Self-attention now lets each of the 7 word tokens look at all 576 patches, so 'this' can attach to the region with the crack.
Illustrative sequence lengths with the same encoder (576 tokens per photo, 7-token question)
RequestImage tokensText tokensTotal
Question only077
1 photo + question5767583
3 photos + question1,72871,735
1 photo, patches merged 2×2 → 11447151

Two lessons fall out of the numbers. First, the image dominates the bill, so sending three photos 'just in case' triples the cost. Second, any step that merges or drops patch vectors saves tokens but also throws away fine detail. If the answer depends on a hairline crack or a tiny serial number, heavy merging is exactly what makes the model miss it. When a multimodal bot gives a vague answer about a small detail, check how many vectors that detail was given before blaming the language model.

Practice: try it yourself

We will simulate late fusion for our support bot. A photo model and a voice-note model each score three fault classes for 2,000 tickets. We measure each one alone, then average them, then see what happens when one modality goes bad.

practice_late_fusion.py

import numpy as np
rng = np.random.default_rng(7)
# Late fusion: two separate models each score 3 fault classes for a ticket.
# Each score = a signal on the true class + that modality's own noise.
N, K = 2000, 3
truth = rng.integers(0, K, size=N)
signal = np.eye(K)[truth]                        # 1.0 on the true class
def scores(noise):
"""Pretend output of one unimodal model (photo model or voice model)."""
return signal + rng.normal(0, noise, size=(N, K))
def accuracy(s):
return (s.argmax(axis=1) == truth).mean()
photo = scores(noise=0.9)                        # photo model: fairly noisy
voice = scores(noise=0.9)                        # voice-note model: just as noisy
print(f"photo only  : {accuracy(photo):.2f}")
print(f"voice only  : {accuracy(voice):.2f}")
print(f"late fusion : {accuracy((photo + voice) / 2):.2f}")
# Now the microphone is bad: the voice scores are almost pure noise.
bad_voice = scores(noise=4.0)
print(f"bad voice only        : {accuracy(bad_voice):.2f}")
print(f"equal-weight fusion   : {accuracy((photo + bad_voice) / 2):.2f}")
# Weight each modality by how much we trust it (1 / noise variance).
w_photo, w_voice = 1 / 0.9**2, 1 / 4.0**2
fused = (w_photo * photo + w_voice * bad_voice) / (w_photo + w_voice)
print(f"trust-weighted fusion : {accuracy(fused):.2f}")
print(f"weight on voice       : {w_voice / (w_photo + w_voice):.2f}")

Output:

photo only  : 0.67
voice only  : 0.68
late fusion : 0.80
bad voice only        : 0.42
equal-weight fusion   : 0.49
trust-weighted fusion : 0.69
weight on voice       : 0.05

Now change it:

  • Set the first voice noise to 0.3 (a very clear voice note). Predict first: will fusion beat the voice model alone, or will the noisy photo drag it down?
  • Add a third modality, text = scores(noise=0.9), and average all three. Predict whether accuracy rises above 0.80, and by more or less than the jump from one model to two.
  • Change K from 3 to 10 classes. Predict which way every accuracy moves, and whether fusion still helps.

Pause and think: In the run above, equal-weight fusion with the bad microphone scored 0.49 while the photo alone scored 0.67. Why can adding a second signal make things worse?

Averaging adds the noise of both inputs as well as their signal. The bad voice scores have noise far larger than the signal, so the average is mostly noise. Fusion only helps when each modality is weighted by how reliable it is; a model has to learn, or be told, when to ignore a modality.

Pause and think: Using the worked example's numbers, a customer attaches 3 photos and a 7-token question. A teammate suggests shortening the question to save cost. Is that a good plan?

No. The photos are 1,728 of the 1,735 tokens, so trimming the text saves almost nothing. The real levers are sending fewer photos, cropping to the damaged region, or using a lower-detail image setting when fine detail is not needed.

Quick summary

  • A modality is a type of data: text, image, audio, video, sensors and more.
  • Multimodal AI accepts and/or produces several modalities; unimodal AI handles one.
  • The recipe: one encoder per modality → projection into a shared space → fusion → a reasoning model → an output decoder.
  • Fusion can be early, late or intermediate; most chat models insert projected image tokens into an LLM.
  • Three families: understanding (many in, text out), generation (text to image/audio/video), and retrieval (shared embedding space).
  • Watch out for visual hallucinations, token cost and privacy; use it only when the extra modality carries real information.

Key takeaways

  • A modality is a type of data; multimodal AI handles two or more of them.
  • Every modality is encoded into vectors and projected into one shared space so a single model can reason over all of them.
  • Fusion can happen early, late or in the middle; most chat models insert projected image tokens into an LLM.
  • Three families: multimodal understanding, cross-modal generation, and cross-modal retrieval.
  • Use it when the extra modality carries real information, and verify fine visual details.

Key terms

  • Modality: A type of data with its own structure, such as text, images, audio or video.
  • Multimodal model: A model that takes in and/or produces more than one modality.
  • Encoder: A network that turns raw data of one modality into vectors.
  • Projection (connector): A small learned layer that maps one encoder's vectors into another model's embedding space.
  • Fusion: The step where information from different modalities is combined: early, late or intermediate.
  • Contrastive learning: Training that pulls matching pairs (an image and its caption) together and pushes mismatched pairs apart.
  • Visual hallucination: When a model describes image content that is not actually there.

← 15.3 LLM Watermarking: Embedding Invisible Signatures in AI Text · 16.2 Vision Transformers: Applying Self-Attention to Image Patches →