Modern AI Engineering

Lesson 19.1 · 40 min

Cracking the AI Engineering Interview

You understand the material. Can you explain KV caching in 60 seconds, estimate the cost of a chatbot on a whiteboard, and design a RAG system while someone asks "why?" after every box you draw?

In short: AI engineering interviews usually combine coding, fundamentals questions, an AI system design round, a deep dive on your past projects, and behavioural questions. Concept answers land best when they follow a short structure: definition, problem it solves, how it works, trade-offs, example. The system design round rewards a clear process: clarify requirements, estimate, sketch the pipeline, go deep on retrieval, evaluation, safety, latency and cost. This lesson gives model answers from every module, a fully worked design of a RAG support assistant, and a 30-day revision plan.

How AI engineering interviews are structured

An AI engineer builds products on top of models: prompting, retrieval, agents, fine-tuning, evaluation, and serving. Interviews try to answer three questions about you: Do you understand how these systems work? Can you build and debug them? Can you make good trade-offs under real constraints like cost, latency and safety?

Formats vary a lot between companies, levels and teams, so treat the table below as a typical shape, not a fixed rule. Always ask the recruiter what each round covers; that is a normal, expected question.

RoundWhat it testsHow to prepare
Recruiter / hiring-manager screenMotivation, background, rough fitA crisp 2-minute story of your experience and one project you are proud of
CodingPractical programming: data structures, sometimes writing a small LLM/RAG utility, parsing, async API callsPractise Python fluently; be ready to write clean, tested code while talking
ML / LLM fundamentalsConcepts from this course: attention, sampling, fine-tuning, RAG, inference, evaluationUse the 60-second answer structure and the question bank below
AI system designDesigning an end-to-end AI product under constraintsPractise the framework and the worked RAG design in this lesson
Project deep diveDepth and honesty about something you actually builtKnow your numbers: data size, metrics before and after, what failed and why
BehaviouralCollaboration, ownership, handling ambiguity and mistakesPrepare 5–6 stories in Situation, Task, Action, Result form

Think like a senior colleague, not a student A student tries to recite the right answer. A senior colleague thinks aloud, asks what matters, states assumptions, offers options with trade-offs, and picks one. Interviewers are asking themselves: "Would I want this person in our design review?"

Explaining a concept clearly in 60 seconds

Most concept questions ("What is X?", "Why do we need Y?", "X vs Y?") are best answered in about a minute, then expanded only if the interviewer asks. A reliable structure:

The five-part answer

  1. 1. Define it in one sentence: Plain words first. "RAG retrieves relevant documents and adds them to the prompt so the model answers from them."
  2. 2. The problem it solves: "LLMs do not know private or recent data and can hallucinate."
  3. 3. How it works: Three or four steps, with a key detail that shows depth. "Chunk and embed documents offline; at query time embed the question, retrieve top-k by similarity, often hybrid with keyword search and a reranker, then generate with citations."
  4. 4. Trade-offs and when not to use it: "Quality depends on retrieval; adds latency and tokens; not the right tool for changing tone or format, that is fine-tuning."
  5. 5. A concrete example or number: "In a support bot, updating a policy document fixes answers immediately, no retraining."

Handling "I don't know" Say what you do know and reason from first principles: "I have not used that method, but if it is a speculative decoding variant, I would expect it to trade extra draft compute for fewer target-model passes..." Honest reasoning scores far better than bluffing, which interviewers spot quickly.

Pause and think: Practise: in one sentence each, give parts 1 and 4 of the five-part answer for "What is LoRA?"

Definition: LoRA fine-tunes a model by freezing its weights and training small low-rank matrices B·A added to chosen weight matrices. Trade-off: it trains a tiny fraction of the parameters and adapters are small to store and swap, but with a small rank it may underperform full fine-tuning on tasks that need large changes.

Must-know questions from every module

Below are high-frequency questions grouped by course module. Try to answer each aloud in 60 seconds before opening the model answer. The model answers are deliberately compact: they are what a strong first answer sounds like, not everything there is to say.

More practice questions Revisit any module whose quiz challenged you. Quiz scores persist, so you can track your weakest areas in the Dashboard and target them before your interview.

The AI system design round

In this round you get an open problem, such as "Design a support assistant for our help centre" or "Design a code-review bot", and 45–60 minutes. There is no single right answer. The interviewer watches how you structure the problem, whether your choices fit the requirements, and whether you can explain trade-offs and failure modes.

A framework that works for most AI design questions

  1. Clarify requirements (5 min): Users and use cases, scale (users, requests/day), latency target, accuracy bar, languages, data sources and freshness, privacy and compliance, budget. Write them down.
  2. Estimate (3 min): Requests per second at peak, tokens per request, daily cost, memory. Rough numbers drive model choice and architecture.
  3. High-level design (10 min): Draw the main pipeline: input, guardrails, retrieval, model, tools, output, logging. Name the components.
  4. Deep dives (15 min): Go deep where the interviewer leans in, typically retrieval quality, model choice, agent/tool safety, or serving.
  5. Evaluation (5 min): Offline eval set and metrics, online metrics, LLM-as-judge with human calibration, regression tests before every change.
  6. Reliability, safety, cost (5 min): Fallbacks, timeouts, rate limits, caching, prompt-injection defences, PII handling, monitoring and alerting.
  7. Summarise and iterate: Recap the design, the main risks, and what you would build first versus later.

back_of_envelope.py

# Back-of-envelope numbers for a RAG support assistant (all inputs are assumptions)
daily_users, msgs_per_user = 20_000, 3
prompt_tokens = 300 + 5 * 400 + 200        # system prompt + 5 chunks of 400 + chat/question
output_tokens = 250
price_in, price_out = 0.50, 2.00           # illustrative $ per 1M tokens (varies by model)
requests = daily_users * msgs_per_user
avg_qps = requests / 86_400
peak_qps = avg_qps * 5                     # assume peak traffic is 5x the average
cost = requests * (prompt_tokens * price_in + output_tokens * price_out) / 1e6
print(f"requests/day {requests:,}  avg QPS {avg_qps:.2f}  peak QPS {peak_qps:.1f}")
print(f"tokens/request {prompt_tokens} in + {output_tokens} out  ->  ~${cost:,.0f}/day")
# KV-cache memory per token for a Llama-3-8B-like config (fp16 = 2 bytes)
layers, kv_heads, head_dim, bytes_ = 32, 8, 128, 2
kv_per_token = 2 * layers * kv_heads * head_dim * bytes_    # 2 = keys and values
print(f"KV cache: {kv_per_token / 1024:.0f} KiB/token, "
f"{kv_per_token * (prompt_tokens + output_tokens) / 2**20:.0f} MiB per request")
# Retrieval evaluation: did the right chunk appear in the top k results?
gold = ["refund-02", "ship-07", "acct-01", "refund-05"]
retrieved = [["refund-02", "refund-01", "ship-03"],
["ship-01", "ship-07", "ship-02"],
["acct-04", "acct-02", "pay-01"],
["refund-05", "refund-02", "acct-01"]]
for k in (1, 3):
hits = sum(g in r[:k] for g, r in zip(gold, retrieved))
print(f"recall@{k} = {hits}/{len(gold)} = {hits / len(gold):.2f}")
mrr = sum(1 / (r.index(g) + 1) if g in r else 0 for g, r in zip(gold, retrieved)) / len(gold)
print(f"MRR = {mrr:.3f}")

Output:

requests/day 60,000  avg QPS 0.69  peak QPS 3.5
tokens/request 2500 in + 250 out  ->  ~$105/day
KV cache: 128 KiB/token, 344 MiB per request
recall@1 = 2/4 = 0.50
recall@3 = 3/4 = 0.75
MRR = 0.625

Pause and think: Per request, the estimate uses 2,500 input tokens at $0.50 per million and 250 output tokens at $2.00 per million. Which is the bigger cost driver, and what is one design change that reduces it?

Input: 2,500 × $0.50 / 1M = $0.00125 per request, versus 250 × $2.00 / 1M = $0.0005 for output, so input is about 70% of the cost. Retrieve fewer or shorter chunks (a reranker helps keep only the best 3), and use prompt caching for the fixed system prompt.

A worked system design: a RAG support assistant

Prompt: "Design an AI assistant that answers customer questions using our help centre (about 5,000 articles) and can check order status. 20,000 daily users. Answers must be accurate and cite sources. P95 latency to first token under 2 seconds."

1. Clarify. We would ask: Which languages? How often do articles change (daily)? Can the bot take actions or only read (read-only order lookup for now)? What should happen when it is unsure (hand off to a human)? Any PII constraints (order data must not be logged in plain text)? We assume English, daily updates, read-only tools, human handoff available.

Key decisions and why
DecisionChoiceReason / trade-off
Knowledge approachRAG, not fine-tuningArticles change daily and answers need citations
RetrievalHybrid search + rerankerOrder numbers and product codes need keywords; paraphrased questions need vectors; reranker boosts precision of the final 3–5 chunks
ModelA mid-sized hosted model; a small model for intent routingMeets quality and latency; routing simple queries to a cheaper model cuts cost
ToolsRead-only order lookup via a tool or MCP server, user-scopedLeast privilege: the bot can only see the logged-in user's orders
LatencyStreaming, prompt caching for the fixed system prompt, a semantic cache for very common questionsKeeps time to first token low; cache must be invalidated when articles change
Unsure answersAbstain and offer a human handoffA wrong policy answer costs more than a handoff

Evaluation. Build a golden set of a few hundred real (anonymised) questions with the correct source articles and reference answers, including tricky and out-of-scope ones. Offline metrics: retrieval recall@5 and MRR; answer faithfulness and correctness scored by an LLM judge calibrated against human ratings; abstention rate on out-of-scope questions. Online: resolution rate without handoff, thumbs up/down, escalations, latency percentiles and cost per conversation. Run the offline suite on every prompt, model or index change.

Failure modes to mention. Retrieval misses (fix chunking, add hybrid search, add synonyms); stale answers (change-feed re-indexing, show last-updated dates); hallucinated citations (verify that cited chunk IDs were actually retrieved); prompt injection via user text or documents (treat retrieved text as data, least-privilege tools); cost spikes (rate limits, caching, routing); model provider outage (timeouts and a fallback model or graceful "please contact support").

A 30-day revision plan

Assume about 1.5–2 hours per weekday and a little more at weekends. Every day: revise the lessons, then answer 5 questions aloud using the five-part structure, recording yourself once a week to check clarity.

Four weeks to interview-ready

  1. Foundations: Module 1–2: the six words, supervised vs unsupervised, regression, features, precision/recall, losses, regularisation, RL, contrastive learning.
  2. Deep learning: Module 3: neurons, gradient descent, backprop, cross-entropy, dropout, normalisation, RNNs. Derive backprop for a tiny network on paper.
  3. Transformers and generation: Modules 4–5: tokenisation, embeddings, attention maths, causal masks, multi-head, RoPE, sampling, streaming. Implement attention in numpy.
  4. Modern architectures and model types: Modules 6–7: MoE, GQA, sliding window, FlashAttention, SLMs, reasoning models.
  5. Training and alignment: Module 8: fine-tuning, LoRA, distillation, RLHF, PPO, DPO, GRPO. Be able to compare them in a table.
  6. Prompting, RAG and agents: Modules 9–12: context engineering, vector search, chunking, hybrid search, reranking, agents, function calling, MCP, frameworks. Build a tiny RAG app end to end.
  7. Inference, evaluation, safety: Modules 13–15: KV cache, batching, paged attention, speculative decoding, quantization, evals, LLM-as-judge, guardrails, prompt injection.
  8. System design and mock interviews: Modules 16–17 skim, then two timed system designs (RAG assistant, agent for internal tools), one full mock loop with a friend, polish your project stories, rest the day before.
Weekly checkpoints
End of weekYou should be able to
Week 1Explain overfitting, precision/recall and backprop in 60 seconds each, without notes
Week 2Write attention from scratch and explain √dₖ, masking, multi-head, RoPE, temperature and top-p
Week 3Compare LoRA vs full fine-tuning, RLHF vs DPO, and build and evaluate a small RAG pipeline
Week 4Run a 45-minute system design with estimates, evaluation, safety and cost, and tell three project stories with numbers

Worked example, step by step

Besides "design X", many loops include a debugging question: "Our support bot's answer quality dropped last week. How would you find out why?" There is no diagram to draw. The interviewer wants to see a method. The strongest move is to refuse to guess, and to split the failures by pipeline stage first.

A method you can say out loud

  1. Pin down the symptom: "Which metric dropped, by how much, and since when? Thumbs-down rate, escalations, or the offline eval?" A vague complaint becomes a number and a date.
  2. Ask what changed: "What shipped around that date? A prompt edit, a model version, a re-index, a new document source, a traffic shift?" Most regressions follow a change.
  3. Collect failing examples: Pull a sample of bad answers, say 100, with their full traces: rewritten query, retrieved chunks, prompt, tool calls and final answer.
  4. Triage by stage: For each one ask in order: was the right chunk retrieved? If yes, did the answer use it correctly? If the question was out of scope, did the bot abstain?
  5. Fix the biggest bucket: Work on the stage that explains the most failures. Do not start with the stage that is most fun to fix.
  6. Prove it and guard it: Re-run the offline eval, check that nothing else got worse, and add the failing cases to the regression set.

With these numbers the conclusion writes itself: 55 of 100 failures are retrieval misses, so changing the generation prompt or the model cannot fix more than 45. We would look at the retrieval stage first. Did the re-index drop documents? Did a new chunk size split answers in half? Are the missed questions full of product codes that vector search handles badly?

Each bucket points to a different fix
BucketFirst things to checkTypical fix
Right chunk not retrievedIndex freshness, chunking, query rewrite, filtersHybrid search, better chunks, fix the ingest
Chunk retrieved, answer wrongToo many distracting chunks, unclear instructions, conflicting documentsRerank and keep fewer chunks, tighten the prompt, show document dates
Should have abstainedScore threshold, out-of-scope examples in the eval setAbstain below a rerank score, add a handoff path
Tool failureTimeouts, error handling, argument validationRetries with limits, clear error messages back to the model

A sentence worth memorising "Before I change anything, I would measure where the failures are, because a fix aimed at the wrong stage cannot help." Said early, it tells the interviewer you debug with evidence.

Practice: try it yourself

We will write the triage step as code, the kind of small utility a coding round might ask for. Ten test questions come with recorded results. For each one we decide which stage failed, and we repeat this for three settings of k, the number of retrieved chunks placed in the prompt. Watch how the mix of failures shifts as k grows.

practice_failure_triage.py

# Failure triage for a RAG bot: WHICH stage broke each wrong answer?
# Ten test questions with recorded results (illustrative data).
#   gold_rank: position of the correct chunk in the search results (None = never found)
#   max_k:     the answer is right only if the model reads at most this many chunks
#              (0 = it misreads the chunk even alone; 9 = extra chunks never distract it)
RESULTS = [(1, 9), (1, 9), (2, 9), (1, 3), (3, 9),
(4, 9), (None, 9), (2, 0), (5, 9), (1, 3)]
CHUNK_TOKENS = 400
def triage(gold_rank, max_k, k):
"""Classify one question when the top k chunks are put in the prompt."""
if gold_rank is None or gold_rank > k:
return "retrieval miss"        # the right chunk never reached the model
if k > max_k:
return "generation error"      # it was there, but the answer is still wrong
return "correct"
print(" k  correct  retrieval miss  generation error  recall@k  context tokens")
for k in (1, 3, 5):
counts = {"correct": 0, "retrieval miss": 0, "generation error": 0}
for gold_rank, max_k in RESULTS:
counts[triage(gold_rank, max_k, k)] += 1
recall = 1 - counts["retrieval miss"] / len(RESULTS)
print(f"{k:2d}  {counts['correct']:7d}  {counts['retrieval miss']:14d}  "
f"{counts['generation error']:16d}  {recall:8.2f}  {k * CHUNK_TOKENS:14,d}")

Output:

 k  correct  retrieval miss  generation error  recall@k  context tokens
1        4               6                 0      0.40             400
3        6               3                 1      0.70           1,200
5        6               1                 3      0.90           2,000

Now change it:

  • Add 10 to the list of k values. Predict the number of correct answers and the recall before you run it. Is the best recall also the best system?
  • Simulate adding a reranker that keeps distracting chunks out: change both (1, 3) entries to (1, 9). Predict the new "correct" count at k = 5.
  • Simulate better retrieval instead: change (None, 9) to (2, 9) and (5, 9) to (2, 9). Predict which k gains the most, and compare with the reranker change. Which fix would you ship first at k = 3?

Pause and think: k = 3 and k = 5 both get 6 of 10 correct. Are the two systems equally good, and would the same fix help both?

No. At k = 3 the main problem is retrieval: 3 misses and 1 generation error, so better search is the next step. At k = 5 retrieval is nearly solved (recall 0.90) but 3 answers go wrong because extra chunks distract the model, so reranking or a tighter prompt is the next step. k = 5 also costs 2,000 context tokens per request instead of 1,200. The same accuracy can hide very different failure mixes, which is why we triage before fixing.

Pause and think: At k = 1 the recall is 0.40 and there are zero generation errors. A teammate concludes "the LLM is perfect, only search is broken". What is wrong with that conclusion?

The model was only tested on the 4 questions where the right chunk arrived, and with a single chunk there was nothing to distract it. End-to-end accuracy can never be higher than retrieval recall, so the generator's weaknesses stay hidden until retrieval improves. At k = 5 the same model makes 3 generation errors. We can only judge a later stage on the cases the earlier stage passed through.

Common mistakes and final tips

Mistakes that sink otherwise strong candidates Jumping into a design without clarifying requirements. Naming tools ("we use a vector DB and LangChain") without explaining why. Never mentioning evaluation. Ignoring cost, latency and safety. Bluffing on a detail instead of reasoning aloud. Talking for five minutes without checking whether the interviewer wants depth or breadth.

  • Lead with structure. "I'll cover requirements, a rough estimate, the pipeline, then go deep on retrieval and evaluation. Sound good?"
  • Quantify. Even rough numbers (tokens, QPS, dollars, milliseconds) show engineering judgement.
  • Always offer a trade-off. "We could also do X; it is cheaper but loses Y; given the 2-second target I would choose Z."
  • Bring evaluation into every answer. "How would we know it worked?" is the question behind most follow-ups.
  • Use your projects. A real story ("our recall@5 was 0.62 until we added hybrid search, then 0.81") beats a textbook definition.
  • Keep learning current. This field moves fast; read release notes and papers for the tools you claim, and say when something varies by vendor or model.

Keep practising Revisit any lesson in this course whose quiz you found hard, track your progress in the Dashboard, and keep building projects that touch inference, RAG, and agents. Good luck!

Key takeaways

  • Typical rounds: coding, fundamentals, AI system design, project deep dive and behavioural; formats vary, so ask.
  • Answer concepts in about 60 seconds: definition, problem, mechanism, trade-offs, example.
  • In system design: clarify, estimate, sketch the pipeline, deep dive, evaluate, then cover safety, latency and cost.
  • For a RAG assistant: hybrid retrieval, reranking, grounded generation with citations, abstention and human handoff, and a golden eval set.
  • Quantify, offer trade-offs, and bring evaluation into every answer.
  • Follow a structured 30-day plan and practise aloud with mock interviews.

Key terms

  • System design round: An interview where you design an end-to-end system under stated requirements and explain trade-offs.
  • Back-of-envelope estimate: A quick rough calculation of traffic, tokens, memory or cost used to guide design decisions.
  • Recall@k: The fraction of queries for which a correct item appears in the top k retrieved results.
  • MRR: Mean Reciprocal Rank: the average of 1 / rank of the first correct result across queries.
  • Golden set: A curated set of inputs with expected outputs used to evaluate a system repeatedly.
  • Abstention: When a system declines to answer, or hands off to a human, because it is not confident.

← 18.3 Recursive Self-Improvement: Can AI Improve Itself Indefinitely?