Lesson 19.1 · 40 min
Cracking the AI Engineering Interview
You understand the material. Can you explain KV caching in 60 seconds, estimate the cost of a chatbot on a whiteboard, and design a RAG system while someone asks "why?" after every box you draw?
In short: AI engineering interviews usually combine coding, fundamentals questions, an AI system design round, a deep dive on your past projects, and behavioural questions. Concept answers land best when they follow a short structure: definition, problem it solves, how it works, trade-offs, example. The system design round rewards a clear process: clarify requirements, estimate, sketch the pipeline, go deep on retrieval, evaluation, safety, latency and cost. This lesson gives model answers from every module, a fully worked design of a RAG support assistant, and a 30-day revision plan.
How AI engineering interviews are structured
An AI engineer builds products on top of models: prompting, retrieval, agents, fine-tuning, evaluation, and serving. Interviews try to answer three questions about you: Do you understand how these systems work? Can you build and debug them? Can you make good trade-offs under real constraints like cost, latency and safety?
Formats vary a lot between companies, levels and teams, so treat the table below as a typical shape, not a fixed rule. Always ask the recruiter what each round covers; that is a normal, expected question.
| Round | What it tests | How to prepare |
|---|---|---|
| Recruiter / hiring-manager screen | Motivation, background, rough fit | A crisp 2-minute story of your experience and one project you are proud of |
| Coding | Practical programming: data structures, sometimes writing a small LLM/RAG utility, parsing, async API calls | Practise Python fluently; be ready to write clean, tested code while talking |
| ML / LLM fundamentals | Concepts from this course: attention, sampling, fine-tuning, RAG, inference, evaluation | Use the 60-second answer structure and the question bank below |
| AI system design | Designing an end-to-end AI product under constraints | Practise the framework and the worked RAG design in this lesson |
| Project deep dive | Depth and honesty about something you actually built | Know your numbers: data size, metrics before and after, what failed and why |
| Behavioural | Collaboration, ownership, handling ambiguity and mistakes | Prepare 5–6 stories in Situation, Task, Action, Result form |
Think like a senior colleague, not a student A student tries to recite the right answer. A senior colleague thinks aloud, asks what matters, states assumptions, offers options with trade-offs, and picks one. Interviewers are asking themselves: "Would I want this person in our design review?"
Explaining a concept clearly in 60 seconds
Most concept questions ("What is X?", "Why do we need Y?", "X vs Y?") are best answered in about a minute, then expanded only if the interviewer asks. A reliable structure:
The five-part answer
- 1. Define it in one sentence: Plain words first. "RAG retrieves relevant documents and adds them to the prompt so the model answers from them."
- 2. The problem it solves: "LLMs do not know private or recent data and can hallucinate."
- 3. How it works: Three or four steps, with a key detail that shows depth. "Chunk and embed documents offline; at query time embed the question, retrieve top-k by similarity, often hybrid with keyword search and a reranker, then generate with citations."
- 4. Trade-offs and when not to use it: "Quality depends on retrieval; adds latency and tokens; not the right tool for changing tone or format, that is fine-tuning."
- 5. A concrete example or number: "In a support bot, updating a policy document fixes answers immediately, no retraining."
Handling "I don't know" Say what you do know and reason from first principles: "I have not used that method, but if it is a speculative decoding variant, I would expect it to trade extra draft compute for fewer target-model passes..." Honest reasoning scores far better than bluffing, which interviewers spot quickly.
Pause and think: Practise: in one sentence each, give parts 1 and 4 of the five-part answer for "What is LoRA?"
Definition: LoRA fine-tunes a model by freezing its weights and training small low-rank matrices B·A added to chosen weight matrices. Trade-off: it trains a tiny fraction of the parameters and adapters are small to store and swap, but with a small rank it may underperform full fine-tuning on tasks that need large changes.
Must-know questions from every module
Below are high-frequency questions grouped by course module. Try to answer each aloud in 60 seconds before opening the model answer. The model answers are deliberately compact: they are what a strong first answer sounds like, not everything there is to say.
More practice questions Revisit any module whose quiz challenged you. Quiz scores persist, so you can track your weakest areas in the Dashboard and target them before your interview.
The AI system design round
In this round you get an open problem, such as "Design a support assistant for our help centre" or "Design a code-review bot", and 45–60 minutes. There is no single right answer. The interviewer watches how you structure the problem, whether your choices fit the requirements, and whether you can explain trade-offs and failure modes.
A framework that works for most AI design questions
- Clarify requirements (5 min): Users and use cases, scale (users, requests/day), latency target, accuracy bar, languages, data sources and freshness, privacy and compliance, budget. Write them down.
- Estimate (3 min): Requests per second at peak, tokens per request, daily cost, memory. Rough numbers drive model choice and architecture.
- High-level design (10 min): Draw the main pipeline: input, guardrails, retrieval, model, tools, output, logging. Name the components.
- Deep dives (15 min): Go deep where the interviewer leans in, typically retrieval quality, model choice, agent/tool safety, or serving.
- Evaluation (5 min): Offline eval set and metrics, online metrics, LLM-as-judge with human calibration, regression tests before every change.
- Reliability, safety, cost (5 min): Fallbacks, timeouts, rate limits, caching, prompt-injection defences, PII handling, monitoring and alerting.
- Summarise and iterate: Recap the design, the main risks, and what you would build first versus later.
back_of_envelope.py
# Back-of-envelope numbers for a RAG support assistant (all inputs are assumptions)
daily_users, msgs_per_user = 20_000, 3
prompt_tokens = 300 + 5 * 400 + 200 # system prompt + 5 chunks of 400 + chat/question
output_tokens = 250
price_in, price_out = 0.50, 2.00 # illustrative $ per 1M tokens (varies by model)
requests = daily_users * msgs_per_user
avg_qps = requests / 86_400
peak_qps = avg_qps * 5 # assume peak traffic is 5x the average
cost = requests * (prompt_tokens * price_in + output_tokens * price_out) / 1e6
print(f"requests/day {requests:,} avg QPS {avg_qps:.2f} peak QPS {peak_qps:.1f}")
print(f"tokens/request {prompt_tokens} in + {output_tokens} out -> ~${cost:,.0f}/day")
# KV-cache memory per token for a Llama-3-8B-like config (fp16 = 2 bytes)
layers, kv_heads, head_dim, bytes_ = 32, 8, 128, 2
kv_per_token = 2 * layers * kv_heads * head_dim * bytes_ # 2 = keys and values
print(f"KV cache: {kv_per_token / 1024:.0f} KiB/token, "
f"{kv_per_token * (prompt_tokens + output_tokens) / 2**20:.0f} MiB per request")
# Retrieval evaluation: did the right chunk appear in the top k results?
gold = ["refund-02", "ship-07", "acct-01", "refund-05"]
retrieved = [["refund-02", "refund-01", "ship-03"],
["ship-01", "ship-07", "ship-02"],
["acct-04", "acct-02", "pay-01"],
["refund-05", "refund-02", "acct-01"]]
for k in (1, 3):
hits = sum(g in r[:k] for g, r in zip(gold, retrieved))
print(f"recall@{k} = {hits}/{len(gold)} = {hits / len(gold):.2f}")
mrr = sum(1 / (r.index(g) + 1) if g in r else 0 for g, r in zip(gold, retrieved)) / len(gold)
print(f"MRR = {mrr:.3f}")Output:
requests/day 60,000 avg QPS 0.69 peak QPS 3.5 tokens/request 2500 in + 250 out -> ~$105/day KV cache: 128 KiB/token, 344 MiB per request recall@1 = 2/4 = 0.50 recall@3 = 3/4 = 0.75 MRR = 0.625
Pause and think: Per request, the estimate uses 2,500 input tokens at $0.50 per million and 250 output tokens at $2.00 per million. Which is the bigger cost driver, and what is one design change that reduces it?
Input: 2,500 × $0.50 / 1M = $0.00125 per request, versus 250 × $2.00 / 1M = $0.0005 for output, so input is about 70% of the cost. Retrieve fewer or shorter chunks (a reranker helps keep only the best 3), and use prompt caching for the fixed system prompt.
A worked system design: a RAG support assistant
Prompt: "Design an AI assistant that answers customer questions using our help centre (about 5,000 articles) and can check order status. 20,000 daily users. Answers must be accurate and cite sources. P95 latency to first token under 2 seconds."
1. Clarify. We would ask: Which languages? How often do articles change (daily)? Can the bot take actions or only read (read-only order lookup for now)? What should happen when it is unsure (hand off to a human)? Any PII constraints (order data must not be logged in plain text)? We assume English, daily updates, read-only tools, human handoff available.
| Decision | Choice | Reason / trade-off |
|---|---|---|
| Knowledge approach | RAG, not fine-tuning | Articles change daily and answers need citations |
| Retrieval | Hybrid search + reranker | Order numbers and product codes need keywords; paraphrased questions need vectors; reranker boosts precision of the final 3–5 chunks |
| Model | A mid-sized hosted model; a small model for intent routing | Meets quality and latency; routing simple queries to a cheaper model cuts cost |
| Tools | Read-only order lookup via a tool or MCP server, user-scoped | Least privilege: the bot can only see the logged-in user's orders |
| Latency | Streaming, prompt caching for the fixed system prompt, a semantic cache for very common questions | Keeps time to first token low; cache must be invalidated when articles change |
| Unsure answers | Abstain and offer a human handoff | A wrong policy answer costs more than a handoff |
Evaluation. Build a golden set of a few hundred real (anonymised) questions with the correct source articles and reference answers, including tricky and out-of-scope ones. Offline metrics: retrieval recall@5 and MRR; answer faithfulness and correctness scored by an LLM judge calibrated against human ratings; abstention rate on out-of-scope questions. Online: resolution rate without handoff, thumbs up/down, escalations, latency percentiles and cost per conversation. Run the offline suite on every prompt, model or index change.
Failure modes to mention. Retrieval misses (fix chunking, add hybrid search, add synonyms); stale answers (change-feed re-indexing, show last-updated dates); hallucinated citations (verify that cited chunk IDs were actually retrieved); prompt injection via user text or documents (treat retrieved text as data, least-privilege tools); cost spikes (rate limits, caching, routing); model provider outage (timeouts and a fallback model or graceful "please contact support").
A 30-day revision plan
Assume about 1.5–2 hours per weekday and a little more at weekends. Every day: revise the lessons, then answer 5 questions aloud using the five-part structure, recording yourself once a week to check clarity.
Four weeks to interview-ready
- Foundations: Module 1–2: the six words, supervised vs unsupervised, regression, features, precision/recall, losses, regularisation, RL, contrastive learning.
- Deep learning: Module 3: neurons, gradient descent, backprop, cross-entropy, dropout, normalisation, RNNs. Derive backprop for a tiny network on paper.
- Transformers and generation: Modules 4–5: tokenisation, embeddings, attention maths, causal masks, multi-head, RoPE, sampling, streaming. Implement attention in numpy.
- Modern architectures and model types: Modules 6–7: MoE, GQA, sliding window, FlashAttention, SLMs, reasoning models.
- Training and alignment: Module 8: fine-tuning, LoRA, distillation, RLHF, PPO, DPO, GRPO. Be able to compare them in a table.
- Prompting, RAG and agents: Modules 9–12: context engineering, vector search, chunking, hybrid search, reranking, agents, function calling, MCP, frameworks. Build a tiny RAG app end to end.
- Inference, evaluation, safety: Modules 13–15: KV cache, batching, paged attention, speculative decoding, quantization, evals, LLM-as-judge, guardrails, prompt injection.
- System design and mock interviews: Modules 16–17 skim, then two timed system designs (RAG assistant, agent for internal tools), one full mock loop with a friend, polish your project stories, rest the day before.
| End of week | You should be able to |
|---|---|
| Week 1 | Explain overfitting, precision/recall and backprop in 60 seconds each, without notes |
| Week 2 | Write attention from scratch and explain √dₖ, masking, multi-head, RoPE, temperature and top-p |
| Week 3 | Compare LoRA vs full fine-tuning, RLHF vs DPO, and build and evaluate a small RAG pipeline |
| Week 4 | Run a 45-minute system design with estimates, evaluation, safety and cost, and tell three project stories with numbers |
Worked example, step by step
Besides "design X", many loops include a debugging question: "Our support bot's answer quality dropped last week. How would you find out why?" There is no diagram to draw. The interviewer wants to see a method. The strongest move is to refuse to guess, and to split the failures by pipeline stage first.
A method you can say out loud
- Pin down the symptom: "Which metric dropped, by how much, and since when? Thumbs-down rate, escalations, or the offline eval?" A vague complaint becomes a number and a date.
- Ask what changed: "What shipped around that date? A prompt edit, a model version, a re-index, a new document source, a traffic shift?" Most regressions follow a change.
- Collect failing examples: Pull a sample of bad answers, say 100, with their full traces: rewritten query, retrieved chunks, prompt, tool calls and final answer.
- Triage by stage: For each one ask in order: was the right chunk retrieved? If yes, did the answer use it correctly? If the question was out of scope, did the bot abstain?
- Fix the biggest bucket: Work on the stage that explains the most failures. Do not start with the stage that is most fun to fix.
- Prove it and guard it: Re-run the offline eval, check that nothing else got worse, and add the failing cases to the regression set.
With these numbers the conclusion writes itself: 55 of 100 failures are retrieval misses, so changing the generation prompt or the model cannot fix more than 45. We would look at the retrieval stage first. Did the re-index drop documents? Did a new chunk size split answers in half? Are the missed questions full of product codes that vector search handles badly?
| Bucket | First things to check | Typical fix |
|---|---|---|
| Right chunk not retrieved | Index freshness, chunking, query rewrite, filters | Hybrid search, better chunks, fix the ingest |
| Chunk retrieved, answer wrong | Too many distracting chunks, unclear instructions, conflicting documents | Rerank and keep fewer chunks, tighten the prompt, show document dates |
| Should have abstained | Score threshold, out-of-scope examples in the eval set | Abstain below a rerank score, add a handoff path |
| Tool failure | Timeouts, error handling, argument validation | Retries with limits, clear error messages back to the model |
A sentence worth memorising "Before I change anything, I would measure where the failures are, because a fix aimed at the wrong stage cannot help." Said early, it tells the interviewer you debug with evidence.
Practice: try it yourself
We will write the triage step as code, the kind of small utility a coding round might ask for. Ten test questions come with recorded results. For each one we decide which stage failed, and we repeat this for three settings of k, the number of retrieved chunks placed in the prompt. Watch how the mix of failures shifts as k grows.
practice_failure_triage.py
# Failure triage for a RAG bot: WHICH stage broke each wrong answer?
# Ten test questions with recorded results (illustrative data).
# gold_rank: position of the correct chunk in the search results (None = never found)
# max_k: the answer is right only if the model reads at most this many chunks
# (0 = it misreads the chunk even alone; 9 = extra chunks never distract it)
RESULTS = [(1, 9), (1, 9), (2, 9), (1, 3), (3, 9),
(4, 9), (None, 9), (2, 0), (5, 9), (1, 3)]
CHUNK_TOKENS = 400
def triage(gold_rank, max_k, k):
"""Classify one question when the top k chunks are put in the prompt."""
if gold_rank is None or gold_rank > k:
return "retrieval miss" # the right chunk never reached the model
if k > max_k:
return "generation error" # it was there, but the answer is still wrong
return "correct"
print(" k correct retrieval miss generation error recall@k context tokens")
for k in (1, 3, 5):
counts = {"correct": 0, "retrieval miss": 0, "generation error": 0}
for gold_rank, max_k in RESULTS:
counts[triage(gold_rank, max_k, k)] += 1
recall = 1 - counts["retrieval miss"] / len(RESULTS)
print(f"{k:2d} {counts['correct']:7d} {counts['retrieval miss']:14d} "
f"{counts['generation error']:16d} {recall:8.2f} {k * CHUNK_TOKENS:14,d}")Output:
k correct retrieval miss generation error recall@k context tokens 1 4 6 0 0.40 400 3 6 3 1 0.70 1,200 5 6 1 3 0.90 2,000
Now change it:
- Add
10to the list of k values. Predict the number of correct answers and the recall before you run it. Is the best recall also the best system? - Simulate adding a reranker that keeps distracting chunks out: change both
(1, 3)entries to(1, 9). Predict the new "correct" count at k = 5. - Simulate better retrieval instead: change
(None, 9)to(2, 9)and(5, 9)to(2, 9). Predict which k gains the most, and compare with the reranker change. Which fix would you ship first at k = 3?
Pause and think: k = 3 and k = 5 both get 6 of 10 correct. Are the two systems equally good, and would the same fix help both?
No. At k = 3 the main problem is retrieval: 3 misses and 1 generation error, so better search is the next step. At k = 5 retrieval is nearly solved (recall 0.90) but 3 answers go wrong because extra chunks distract the model, so reranking or a tighter prompt is the next step. k = 5 also costs 2,000 context tokens per request instead of 1,200. The same accuracy can hide very different failure mixes, which is why we triage before fixing.
Pause and think: At k = 1 the recall is 0.40 and there are zero generation errors. A teammate concludes "the LLM is perfect, only search is broken". What is wrong with that conclusion?
The model was only tested on the 4 questions where the right chunk arrived, and with a single chunk there was nothing to distract it. End-to-end accuracy can never be higher than retrieval recall, so the generator's weaknesses stay hidden until retrieval improves. At k = 5 the same model makes 3 generation errors. We can only judge a later stage on the cases the earlier stage passed through.
Common mistakes and final tips
Mistakes that sink otherwise strong candidates Jumping into a design without clarifying requirements. Naming tools ("we use a vector DB and LangChain") without explaining why. Never mentioning evaluation. Ignoring cost, latency and safety. Bluffing on a detail instead of reasoning aloud. Talking for five minutes without checking whether the interviewer wants depth or breadth.
- Lead with structure. "I'll cover requirements, a rough estimate, the pipeline, then go deep on retrieval and evaluation. Sound good?"
- Quantify. Even rough numbers (tokens, QPS, dollars, milliseconds) show engineering judgement.
- Always offer a trade-off. "We could also do X; it is cheaper but loses Y; given the 2-second target I would choose Z."
- Bring evaluation into every answer. "How would we know it worked?" is the question behind most follow-ups.
- Use your projects. A real story ("our recall@5 was 0.62 until we added hybrid search, then 0.81") beats a textbook definition.
- Keep learning current. This field moves fast; read release notes and papers for the tools you claim, and say when something varies by vendor or model.
Keep practising Revisit any lesson in this course whose quiz you found hard, track your progress in the Dashboard, and keep building projects that touch inference, RAG, and agents. Good luck!
Key takeaways
- Typical rounds: coding, fundamentals, AI system design, project deep dive and behavioural; formats vary, so ask.
- Answer concepts in about 60 seconds: definition, problem, mechanism, trade-offs, example.
- In system design: clarify, estimate, sketch the pipeline, deep dive, evaluate, then cover safety, latency and cost.
- For a RAG assistant: hybrid retrieval, reranking, grounded generation with citations, abstention and human handoff, and a golden eval set.
- Quantify, offer trade-offs, and bring evaluation into every answer.
- Follow a structured 30-day plan and practise aloud with mock interviews.
Key terms
- System design round: An interview where you design an end-to-end system under stated requirements and explain trade-offs.
- Back-of-envelope estimate: A quick rough calculation of traffic, tokens, memory or cost used to guide design decisions.
- Recall@k: The fraction of queries for which a correct item appears in the top k retrieved results.
- MRR: Mean Reciprocal Rank: the average of 1 / rank of the first correct result across queries.
- Golden set: A curated set of inputs with expected outputs used to evaluate a system repeatedly.
- Abstention: When a system declines to answer, or hands off to a human, because it is not confident.
← 18.3 Recursive Self-Improvement: Can AI Improve Itself Indefinitely?