Module 15 · Production
Securing AI Systems
In this module, we will learn how to keep an LLM application safe, how attackers try to break it, and how AI-generated text can be identified.
By the end of this module, we will be able to defend an LLM application against the most common attacks.
Lessons
- 15.1 LLM Guardrails: Filtering Inputs and Outputs for Safety: What is an LLM · What are LLM guardrails · Why do we need guardrails · Where guardrails sit: input and output · Types of guardrails · A simple input guardrail with code · A simple output guardrail with code · Using another model as a guardrail · A step-by-step walkthrough of a request · Limitations of guardrails · Best practices for guardrails
- 15.2 Prompt Injection: Attacks Against LLM-Powered Systems: What is a Large Language Model · What is a prompt · The system prompt and the user prompt · What is Prompt Injection · The root cause of Prompt Injection · A simple example of Prompt Injection · Direct Prompt Injection · Indirect Prompt Injection · A step-by-step walkthrough of a real attack · A code example of how the attack sneaks in · Prompt Injection vs Jailbreaking · Why Prompt Injection is not like SQL Injection · What an attacker can achieve · The defenses, one approach at a time · A defense checklist · How to test our own application · Why this problem is still not solved
- 15.3 LLM Watermarking: Embedding Invisible Signatures in AI Text: What is a watermark? · Why do we need a watermark in LLM-generated text? · How does an LLM write text? · How does an LLM choose the next word? · The hidden freedom that makes watermarking possible · Here comes the secret key into the picture · Preferred tokens and other tokens · Slightly changing the probabilities · Why the preferred set keeps changing · One token vs thousands of tokens · How does the detection work? · How is this different from an AI text detector? · Why the quality of the text does not break · What happens when someone edits the text? · Where is LLM watermarking used in the real world? · Advantages and disadvantages of LLM watermarking
← Module 14: Measuring What Matters · Module 16: Beyond Text: Multimodal AI →