Modern AI Engineering

Lesson 17.11 · 29 min

How do Voice And Video Call Work?

When you video-call a friend, your faces often travel directly between your two phones, so how do two devices hidden behind home routers even find each other?

In short: A call has two phases. First, signaling: the two apps exchange messages through a server to agree on codecs and share possible network addresses (session descriptions and ICE candidates). Then media flows, ideally peer-to-peer over UDP using WebRTC. Because most devices sit behind NAT routers, a STUN server tells each device its public address so a direct path can be punched through, and when that fails a TURN server relays the media. Codecs such as Opus, VP8, H.264 and AV1 compress the audio and video so it fits the network.

The big picture: two phases of a call

Our running example: Asha in Bengaluru starts a video call with Ben in London from a chat app. Within a second or two, they see and hear each other. Under the hood, the call has two phases with very different needs.

  • Signaling (setting up the call). Small control messages: "Asha is calling Ben", "Ben accepted", "here are the codecs I support", "here are addresses where you might reach me". These go through our servers, typically over WebSocket or HTTPS, because the two devices cannot yet reach each other directly.
  • Media (the call itself). A continuous stream of audio and video packets, about 50 audio packets a second and dozens of video frames a second. This must be fast, so it goes as directly as possible between the two devices, usually over UDP.

Think of it like arranging to meet a pen pal You cannot phone your pen pal because you do not know their number, so you exchange letters through a mutual friend (the signaling server) to agree on a language and swap phone numbers. Once both have the numbers, you talk directly and the friend is no longer involved. If one of you is in a building that blocks direct calls, you ask an operator (TURN) to connect you through their switchboard.

Most modern browser and app calling is built on WebRTC (Web Real-Time Communication): an open standard and set of APIs, built into all major browsers and available as libraries for mobile, that handle capturing the camera and microphone, encoding, network traversal, encryption and playback. Many native calling apps use their own stacks built on the same ideas (SDP-like negotiation, ICE, STUN, TURN, RTP).

Signaling: agreeing how to talk

WebRTC deliberately does not define how signaling messages travel; every app chooses its own channel (often a WebSocket to the app server, as in the previous lesson). What WebRTC does define is what must be exchanged, using the offer/answer model.

The exchanged documents are written in SDP (Session Description Protocol), a text format listing the media streams (one audio, one video), the codecs each side supports in order of preference, encryption fingerprints, and connection parameters. The caller creates an offer; the callee replies with an answer that picks compatible options. Separately, each side sends ICE candidates: possible network addresses at which it might be reachable. Sending candidates one by one as they are discovered, instead of waiting for all of them, is called trickle ICE and speeds up setup.

Setting up Asha's call to Ben

  1. Asha creates an offer: Her app asks WebRTC for an SDP offer: "audio with Opus; video with VP8, H.264 or AV1; here is my DTLS fingerprint". It sets this as its local description and sends it to the signaling server.
  2. Server relays it to Ben: The signaling server looks up Ben's online session and forwards the offer; Ben's phone rings.
  3. Ben answers: When Ben accepts, his app sets Asha's offer as the remote description, creates an SDP answer choosing codecs both support (say Opus and VP8), and sends it back through the server.
  4. Both trickle ICE candidates: Each side gathers candidate addresses (local, public via STUN, relay via TURN) and sends each one through the signaling server as it is found.
  5. Connectivity checks: ICE pairs up candidates and tests them with small STUN packets sent directly between the devices, picking the best pair that works.
  6. Secure media flows: A DTLS handshake on the chosen path derives keys; audio and video flow as encrypted SRTP packets. The signaling server is no longer in the media path.

Pause and think: After the call is connected, Asha's chat server crashes. Does the video call drop?

Usually not. The signaling server was only needed to set the call up. Media flows directly (or via TURN) between the devices, so the call continues; only new signaling, such as adding a participant or renegotiating, would fail until the server recovers.

Peer-to-peer connection and the NAT problem

A peer-to-peer (P2P) connection means media goes straight from one device to the other without passing through our servers. It gives the lowest latency (no detour), costs us no server bandwidth, and keeps media off our infrastructure. The obstacle is NAT.

NAT (Network Address Translation) lets many devices on a home or office network share one public IPv4 address. Asha's phone has a private address such as 192.168.1.20 that is meaningless on the internet. When it sends a packet out, her router rewrites the source to its own public address and a port, say 203.0.113.7:54321, and remembers the mapping so replies can come back in. But the router drops unsolicited packets from strangers. So Ben cannot simply send packets to "192.168.1.20", and he does not know "203.0.113.7:54321" exists.

NATs differ in how they create mappings. Many home routers reuse the same public port for a device's socket no matter which destination it talks to; these are friendly to P2P. A symmetric NAT (common in some corporate and mobile carrier networks) creates a different public port for every destination, so the address learned from one server is useless for reaching another peer. Firewalls that block all UDP are a further obstacle.

The framework that solves this is ICE (Interactive Connectivity Establishment). Each device gathers three kinds of candidates: host candidates (its own local addresses, which work on the same network), server-reflexive candidates (its public NAT address, learned from STUN), and relay candidates (an address on a TURN server). Both sides exchange candidates via signaling, then systematically test pairs and pick the best working one, preferring direct paths over relays.

STUN server and TURN server

A STUN server (Session Traversal Utilities for NAT) answers one question: "what IP address and port do you see my packet coming from?". The device sends a small Binding request over UDP; the server replies with the source address it observed, which is the device's public NAT mapping. STUN servers are cheap to run because they handle only tiny packets and never carry media. Once both peers know their public addresses and send checks to each other simultaneously, their NATs usually let the traffic through.

A TURN server (Traversal Using Relays around NAT) is the fallback for when a direct path cannot be made, for example with symmetric NATs on both sides or firewalls that block UDP. The device allocates a relay address on the TURN server, and all media goes device → TURN → other device. TURN can also run over TCP or TLS on port 443, which gets through most restrictive firewalls. It always works, but it adds latency (a detour through the server) and costs us real bandwidth, since every byte of every relayed call passes through our infrastructure.

stun_ice_sim.py

# A STUN Binding exchange (RFC 5389 format), simulated offline, then ICE-style choice.
import ipaddress, struct
COOKIE = 0x2112A442
txn = bytes.fromhex("a1b2c3d4e5f6a7b8c9d0e1f2")   # 12-byte transaction id (fixed for demo)
# 1) Client -> STUN server: "What address do you see me as?"
request = struct.pack("!HHI", 0x0001, 0, COOKIE) + txn   # type=Binding Request, length=0
print("request:", request.hex(), f"({len(request)} bytes)")
# 2) Server sees the packet arrive from the NAT's public side and answers.
public_ip, public_port = "203.0.113.7", 54321            # what the NAT mapped us to
x_port = public_port ^ (COOKIE >> 16)
x_addr = int(ipaddress.IPv4Address(public_ip)) ^ COOKIE
attr = struct.pack("!HHBBHI", 0x0020, 8, 0, 0x01, x_port, x_addr)  # XOR-MAPPED-ADDRESS
response = struct.pack("!HHI", 0x0101, len(attr), COOKIE) + txn + attr
# 3) Client parses the response to learn its public (server-reflexive) address.
mtype, mlen, cookie = struct.unpack("!HHI", response[:8])
assert mtype == 0x0101 and response[8:20] == txn         # Binding Success, our txn
atype, alen, _, fam, xp, xa = struct.unpack("!HHBBHI", response[20:32])
port = xp ^ (COOKIE >> 16)
ip = ipaddress.IPv4Address(xa ^ COOKIE)
print(f"STUN says our public address is {ip}:{port}")
# 4) ICE: gather candidates, try pairs, prefer direct paths over relays
candidates = [("host", "192.168.1.20:50000", 126),
("srflx", f"{ip}:{port}", 100),            # learned via STUN
("relay", "198.51.100.9:3478", 0)]         # allocated on a TURN server
def connect(peer_nat):
for kind, addr, pref in sorted(candidates, key=lambda c: -c[2]):
works = ((kind == "host" and peer_nat == "same LAN") or
(kind == "srflx" and peer_nat != "symmetric") or kind == "relay")
if works:
return f"{kind:5s} via {addr}"
for nat in ["same LAN", "home router", "symmetric"]:
print(f"peer behind {nat:11s} -> {connect(nat)}")

Output:

request: 000100002112a442a1b2c3d4e5f6a7b8c9d0e1f2 (20 bytes)
STUN says our public address is 203.0.113.7:54321
peer behind same LAN    -> host  via 192.168.1.20:50000
peer behind home router -> srflx via 203.0.113.7:54321
peer behind symmetric   -> relay via 198.51.100.9:3478

Carrying the media: RTP, codecs and bad networks

Once a path exists, media travels in RTP (Real-time Transport Protocol) packets, each with a sequence number and timestamp so the receiver can reorder them and play them at the right time. In WebRTC, RTP is always encrypted as SRTP, with keys negotiated by a DTLS handshake, so a STUN or TURN server cannot listen in on a 1:1 call. RTCP packets carry feedback such as packet loss and round-trip time.

Media goes over UDP, not TCP, on purpose. TCP retransmits every lost packet and holds back everything behind it until the gap is filled. For a live call, a frame that arrives half a second late is useless; it is better to skip it and keep going. UDP lets the application decide what to do about loss.

A codec (coder-decoder) compresses media. Raw 720p video at 30 frames per second is roughly 1280 × 720 × 1.5 bytes × 30 ≈ 41 MB/s (in common YUV 4:2:0 format), hopeless on a phone network; a video codec brings that to around a megabit or two per second by sending mostly differences between frames. In WebRTC, Opus is the mandatory audio codec (flexible from a few kbps up to high-quality music, robust to loss); VP8 and H.264 are the mandatory video codecs, and VP9 and AV1 are widely supported for better compression at the cost of more CPU.

  • Jitter buffer: packets arrive with uneven delay (jitter); the receiver holds a few tens of milliseconds of audio to play it smoothly, growing the buffer on bad networks.
  • Packet loss concealment and FEC: Opus can guess a missing 20 ms of audio, and forward error correction sends a little redundant data so a lost packet can be rebuilt.
  • NACK and keyframes: the receiver can ask for a lost video packet again, or request a fresh full frame (keyframe) if decoding broke.
  • Congestion control and simulcast: the sender estimates available bandwidth and lowers bitrate or resolution when the network struggles; with simulcast it sends several qualities so an SFU can forward the right one to each receiver.
  • Echo cancellation and noise suppression: stop the remote voice coming out of your speaker from being sent back, and remove keyboard and fan noise.

Real-world architecture and timeline

A production calling stack Mobile and web clients with WebRTC; a signaling service over WebSocket handling presence, ringing, offer/answer and ICE candidates; STUN and TURN servers in several regions (TURN on UDP plus TLS 443 for strict networks); SFUs for group calls; push notifications to wake up a phone for incoming calls; and call-quality metrics (round-trip time, loss, jitter, bitrate, freeze counts) collected from clients. Voice AI agents, from the earlier lesson, plug into the same stack: the agent joins the call as a WebRTC peer or SFU participant.

Milestones in internet calling

  1. RTP standardised: The Real-time Transport Protocol for audio and video over IP is published (later updated as RFC 3550 in 2003).
  2. STUN revised: RFC 5389 defines modern STUN, including the magic cookie and XOR-MAPPED-ADDRESS used in our code.
  3. TURN and ICE: RFC 5766 (TURN) and RFC 5245 (ICE) define relaying and the candidate-checking process; ICE was later revised as RFC 8445.
  4. WebRTC project announced: Google open-sources WebRTC technology and standardisation work begins at the W3C and IETF.
  5. WebRTC 1.0 a W3C Recommendation: Real-time audio and video become a formal web standard supported by all major browsers.

Common mistakes and when P2P is not enough

Mistake: shipping without a TURN server Calls work perfectly in testing on office Wi-Fi and home networks, then fail for some users on corporate networks or certain mobile carriers. STUN alone cannot get through symmetric NATs or UDP-blocking firewalls. Always deploy TURN (with TLS on port 443) as a fallback and monitor how many calls use it.

  • Confusing signaling with media. The signaling server does not carry audio or video; scaling it is about connections and messages, not bandwidth.
  • Leaving TURN open to anyone. Use short-lived credentials generated per user, or strangers will use your relay as free bandwidth.
  • Mesh for big groups. Beyond about 3 or 4 participants, upload and CPU explode; use an SFU.
  • Ignoring network changes. When a phone switches from Wi-Fi to mobile data, the path breaks; use ICE restarts to find a new one.
  • Testing only on perfect networks. Simulate loss, jitter and low bandwidth; tune jitter buffers and bitrate adaptation.

Pause and think: Ben is on a corporate network with a symmetric NAT and a firewall that blocks UDP. Which candidate type will the call most likely use, and over what transport?

A relay candidate on a TURN server, reached over TCP or TLS on port 443. The STUN-learned address is useless with a symmetric NAT, and UDP is blocked, so TURN over TLS is the path that gets through.

Worked example, step by step

The jitter buffer got one line earlier. It deserves a closer look, because it is where a call trades delay against gaps. Suppose Asha's phone sends one audio packet every 20 ms. The network does not deliver them evenly. Here are five packets with illustrative network delays.

Five audio packets (all times in ms)
PacketSent atNetwork delayArrives atPlay time, 20 ms bufferPlay time, 60 ms buffer
10404060 (on time)100 (on time)
220557580 (on time)120 (on time)
3404585100 (on time)140 (on time)
46090150120 (too late)160 (on time)
58050130140 (on time)180 (on time)

Reading the table

  1. Without a buffer, playback stutters: If Ben's phone played each packet the moment it arrived, the gaps between them would be 35, 10, 65 ms instead of a steady 20. Speech would speed up and slow down.
  2. A buffer fixes a schedule: The receiver decides: 'I play each packet at its send time, plus the usual 40 ms of network delay, plus my buffer.' With a 20 ms buffer, packet 1 plays at 60, packet 2 at 80, and so on, exactly 20 ms apart.
  3. A late packet is a lost packet: Packet 4 arrives at 150 but was due at 120. Its slot has already passed, so the receiver must hide the gap (packet loss concealment) even though the packet was never lost by the network.
  4. A bigger buffer saves it: With a 60 ms buffer, packet 4 is due at 160 and arrives at 150, in time. But now every packet is played 40 ms later than before.
  5. Packets can overtake each other: Packet 5 arrives at 130, before packet 4 at 150. The sequence numbers in RTP let the buffer put them back in order.
  6. So the buffer adapts: Real receivers measure recent jitter and grow the buffer on a rough network and shrink it on a calm one, looking for the smallest delay that keeps late packets rare.

Delay matters because conversation is two-way. Once the total mouth-to-ear delay grows to a few hundred milliseconds, people start talking over each other. So a call that 'sounds choppy' and a call where 'we keep interrupting each other' can be the same problem seen from opposite ends of this one trade-off.

Practice: try it yourself

We will send 1,000 audio packets through a simulated network that loses a few, delays most a little and delays some a lot. Then we play them through jitter buffers of four sizes and count how many packets arrive too late to be played. At the end we weigh one audio packet to see how much of it is headers.

practice_jitter_buffer.py

# A jitter buffer: trade a little delay for fewer late (unplayable) packets.
import random
random.seed(21)
PACKET_MS = 20                                   # one audio packet every 20 ms
N = 1000                                         # 20 seconds of speech
# Network delay per packet: a 40 ms base, usually small jitter, sometimes a spike.
arrivals = []
for seq in range(N):
if random.random() < 0.02:
continue                                 # 2% of packets are lost for good
jitter = random.expovariate(1 / 8)           # mean 8 ms of extra delay
if random.random() < 0.05:
jitter += random.uniform(30, 90)         # an occasional burst of delay
arrivals.append((seq, seq * PACKET_MS + 40 + jitter))
lost = N - len(arrivals)
print(f"sent {N}, lost in the network {lost} ({lost / N:.1%})")
print("buffer   delay before playing   late packets   gaps heard")
for buffer_ms in (0, 20, 60, 120):
# Packet seq must be played at: its send time + base delay + the buffer.
late = sum(1 for seq, t in arrivals
if t > seq * PACKET_MS + 40 + buffer_ms)
gaps = (late + lost) / N                     # late packets are as bad as lost
print(f"{buffer_ms:4d} ms {40 + buffer_ms:16d} ms {late:14d} {gaps:12.1%}")
# What one audio packet weighs on the wire (24 kbps Opus, 20 ms packets)
payload = 24_000 / 8 * PACKET_MS / 1000          # bytes of compressed audio
headers = 12 + 8 + 20                            # RTP + UDP + IPv4 headers
print(f"payload {payload:.0f} B + headers {headers} B "
f"-> {(payload + headers) * 8 * 50 / 1000:.0f} kbps on the wire")

Output:

sent 1000, lost in the network 18 (1.8%)
buffer   delay before playing   late packets   gaps heard
0 ms               40 ms            982       100.0%
20 ms               60 ms            127        14.5%
60 ms              100 ms             32         5.0%
120 ms              160 ms              0         1.8%
payload 60 B + headers 40 B -> 40 kbps on the wire

Now change it:

  • Change the burst probability from 0.05 to 0.20 (a crowded Wi-Fi network). Predict what happens to the 'gaps heard' column for the 60 ms buffer.
  • Set PACKET_MS = 40, so each packet carries twice as much audio. Predict the kbps on the wire. What is the downside when one packet is lost?
  • Add a buffer size of 300 to the list. Predict its 'gaps heard' value. Then say why nobody would choose it for a live call.

Pause and think: In the run above, the network lost only 1.8% of packets, yet with a 20 ms buffer the listener hears gaps for 14.5% of them. Where do the other gaps come from?

From packets that did arrive, but after their play time had passed. A real-time player cannot wait: when a packet's slot comes and the packet is not there, the gap must be concealed, and the packet is useless when it shows up later. That is why a call can sound bad on a network with little real loss, and why the receiver's buffer size matters as much as the loss rate.

Pause and think: A teammate wants to send call audio over TCP so that no packet is ever lost. Using the simulation, what would happen to the packets behind one that needs retransmitting?

TCP delivers in order, so everything behind the missing packet waits until the retransmission arrives. One delayed packet turns into a whole run of late packets, and in a jitter buffer late means unplayable. To hide that, the buffer would have to grow, adding delay to the entire call. Over UDP the receiver simply conceals the one missing 20 ms and carries on, which is the better trade for live speech.

Key takeaways

  • A call = signaling (offer/answer and ICE candidates through a server) + media (ideally direct between devices).
  • NAT hides devices behind shared public addresses; ICE gathers host, server-reflexive and relay candidates and tests pairs.
  • STUN tells a device its public address so UDP hole punching can create a direct path; it never carries media.
  • TURN relays media when direct paths fail; it is essential in production but adds latency and bandwidth cost.
  • Media uses encrypted RTP over UDP with codecs (Opus, VP8/H.264, VP9, AV1), jitter buffers and congestion control; group calls use SFUs.

Key terms

  • Signaling: Exchanging setup messages (session descriptions and network candidates) between call participants through a server.
  • SDP: Session Description Protocol: a text format describing media streams, codecs and connection details in an offer or answer.
  • NAT: Network Address Translation: a router maps many private addresses to one public address and port.
  • ICE: Interactive Connectivity Establishment: gathers candidate addresses and tests pairs to find a working path.
  • STUN server: A server that tells a device which public IP address and port its packets appear to come from.
  • TURN server: A server that relays media between peers when they cannot connect directly.
  • SFU: Selective Forwarding Unit: a media server that receives each participant's stream once and forwards it to others.
  • Codec: Software that compresses and decompresses audio or video, such as Opus or VP8.

← 17.10 Transport Protocols: HTTP, WebSockets, and SSE Compared · 18.1 JEPA: LeCun's Vision for World Model AI →