Home › AI Interview Prep › AI Engineer (Prompt & Context Engineering)
Highest Demand

AI Engineer (Prompt & Context Engineering) Interview Questions and Answers

Thirty scenarios on prompt and context engineering, RAG, structured output, model selection and cost for production LLM features. Write your own answer first, by typing or speaking, then open the model answer to compare structure and reasoning.

30 scenariosWhat each question testsRed-flag answersNo sign-up, private
Written by Ayush Bisht · Reviewed by Sanjay Saini
Last updated 2026-09-30
AI Engineer (Prompt & Context Engineering) interview questions

Answer in your own words before opening a model answer, by typing or by pressing Speak your answer. Your text is saved in this browser only and is never uploaded to us. Voice input uses your browser's speech service to turn speech into text; in Chrome that audio is processed by Google.

Scenario 1Prompt EngineeringPractitioner

How do you structure a production prompt so it stays reliable as requirements change?

What the interviewer is testing: Whether prompts are structured, versioned and eval-gated.

Model answer, red flags and follow-up

A strong answer

I split the prompt into role, task, constraints, context and output format so each part can change without disturbing the others, and I keep it in version control with a changelog. Examples come from real inputs, not invented ones. Before any change I run a fixed eval set of common and edge cases and compare against the previous version. Small targeted edits beat rewrites because I can attribute a regression to one change. When a requirement shifts, I add the new cases to the eval set first, watch them fail, then adjust the prompt until they pass without breaking the older ones.

Answers that lose you the room

  • Writes one long unstructured prompt
  • Doesn't version prompts
  • Changes prompts without running evals

Expect this follow-up: A prompt fix helps one case and breaks another. How do you prevent that?

Scenario 2Context EngineeringPractitioner

What is context engineering and how does it differ from writing prompts?

What the interviewer is testing: Whether you curate what the model sees, not just what you say.

Model answer, red flags and follow-up

A strong answer

Prompting is the instruction; context engineering is deciding everything the model sees when it reads that instruction: retrieved documents, memory, tool results, conversation history, and the order and size of each. I rank, filter and compress context because irrelevant tokens cost money, add latency and pull attention away from what matters. I also decide what goes near the start or end of the window, since position affects recall. In practice most quality gains in production systems come from better context, not cleverer wording, and I measure that with evals rather than intuition.

Answers that lose you the room

  • Treats it as just better prompt wording
  • Sends all retrieved text unfiltered
  • Ignores context order and size

Expect this follow-up: What would you cut first when the context is too long?

Scenario 3RAGPractitioner

Your RAG answers are wrong even though the right document exists. How do you debug it?

What the interviewer is testing: Whether you separate retrieval failures from generation failures.

Model answer, red flags and follow-up

A strong answer

I split the problem in two: retrieval and generation. First I check whether the right chunk was retrieved and where it ranked. If it was missing, I look at chunking, embedding quality, hybrid search, metadata filters and query rewriting. If it was retrieved but ranked low, I add a reranker. If it was in the context and the answer is still wrong, the problem is generation: prompt clarity, context order, or too much noise. I keep retrieval metrics such as recall at k so I can tell which side broke before I touch the prompt.

Answers that lose you the room

  • Changes the prompt before checking retrieval
  • Blames the model by default
  • Has no retrieval metrics

Expect this follow-up: The right chunk is retrieved but ignored. What next?

Scenario 4HallucinationsPractitioner

How do you reduce hallucinations in a customer-facing assistant?

What the interviewer is testing: Whether you reduce hallucination with grounding and measurement.

Model answer, red flags and follow-up

A strong answer

I ground answers in retrieved sources and require citations, so every claim can be checked. I explicitly allow 'I don't know' and reward it in evals, because a model that must always answer will invent. I constrain output format, add a verification step for high-risk claims, and route low-confidence or high-stakes cases to a human. Then I measure: a labelled eval set gives me a hallucination rate I can track across releases. A stricter prompt helps a little, but only measurement tells me whether the assistant is actually safer.

Answers that lose you the room

  • Says a stricter prompt solves it
  • Never allows 'I don't know'
  • Tracks no hallucination rate on an eval set

Expect this follow-up: How do you decide when to route to a human?

Scenario 5Structured OutputFoundation

How do you get reliable structured output from an LLM?

What the interviewer is testing: Whether you enforce schemas instead of asking politely.

Model answer, red flags and follow-up

A strong answer

I stop asking politely and enforce a schema. Native structured output or schema-constrained decoding guarantees valid shape, and I still validate values against the schema in code, because valid JSON can hold wrong content. On failure I retry once with the validation error included so the model can correct itself, then fall back to a safe default or a human. I keep schemas simple, with enums where possible, and log every failure. Recurring failures usually point to an ambiguous field description or an over-complex schema, which I fix at the source.

Answers that lose you the room

  • Asks the model politely for JSON
  • Does no schema validation
  • Retries blindly without the error

Expect this follow-up: The schema is complex and fails often. What do you change?

Scenario 6Model SelectionFoundation

How do you choose a model for a new feature?

What the interviewer is testing: Whether you choose models by your own evals and constraints.

Model answer, red flags and follow-up

A strong answer

I start from requirements: quality bar, latency, cost per request, context length, data privacy and language coverage. I shortlist two or three models and compare them on my own eval set, not public benchmarks, because benchmarks rarely match my task. I pick the cheapest model that clears the quality bar with some margin, and I keep a second model as a fallback behind a thin abstraction so switching is cheap. I also plan to re-run the comparison periodically, since new models change the cost and quality picture every few months.

Answers that lose you the room

  • Picks the newest or biggest model
  • Relies on public benchmarks
  • Has no fallback model

Expect this follow-up: The cheaper model is 3% worse. How do you decide?

Scenario 7Long ContextPractitioner

When would you use a long context window instead of retrieval?

What the interviewer is testing: Whether you weigh long context against retrieval honestly.

Model answer, red flags and follow-up

A strong answer

Long context suits small, self-contained material and one-off analysis, like reading a single contract. Retrieval wins when the corpus is large, changes often, or needs citations, and it is cheaper per query because I send only what is relevant. I test both on the real task, because models can lose facts buried in the middle of a long prompt. Cost matters too: sending a full corpus on every request adds up quickly. Often the best design is hybrid: retrieve broadly, then place a generous but curated set of passages in a long window.

Answers that lose you the room

  • Says long context replaces retrieval
  • Ignores cost per query
  • Doesn't test middle-of-context recall

Expect this follow-up: At what corpus size would you switch to retrieval?

Scenario 8CostPractitioner

Your token bill doubled after a launch. What do you check first?

What the interviewer is testing: Whether you diagnose cost from tokens and traffic, not guesses.

Model answer, red flags and follow-up

A strong answer

I start with the numbers, not guesses. I break the bill down by feature: tokens per request, request volume, retries, and how much context grew. Common causes are ballooning conversation history, an over-retrieved context, retry storms and a new feature routed to an expensive model. Fixes follow the cause: trim prompts, cap history, cache repeated prefixes, route easy tasks to a smaller model and set budgets with alerts per feature. I avoid cutting quality first; I only accept a quality trade-off after the eval set shows it is small.

Answers that lose you the room

  • Blames traffic alone
  • Has no per-feature cost visibility
  • Cuts quality first

Expect this follow-up: Context per request grew 40%. Where would you look?

Scenario 9Few-shot PromptingFoundation

When do few-shot examples help, and when do they hurt?

What the interviewer is testing: Whether you know when examples help or bias output.

Model answer, red flags and follow-up

A strong answer

Few-shot examples help when the format or edge-case behaviour is hard to describe in words, such as tone, labelling rules or output structure. They hurt when examples are unrepresentative, when the model over-copies their surface pattern, or when they consume tokens that context could use better. I choose diverse examples from real data, include an edge case, and vary them so the model learns the rule, not the sample. Then I test with and without them on the eval set. If the gain is small, I drop them and save the tokens.

Answers that lose you the room

  • Adds many examples by default
  • Uses unrepresentative examples
  • Never tests with and without

Expect this follow-up: The model copies your example too literally. What do you do?

Scenario 10Prompt InjectionAdvanced

How do you defend an LLM app against prompt injection?

What the interviewer is testing: Whether you layer defences and treat inputs as untrusted.

Model answer, red flags and follow-up

A strong answer

I assume all retrieved and user-supplied text is untrusted. I separate instructions from data with clear delimiters, but I know that alone is weak. The real protections are architectural: least-privilege tool permissions, confirmation for sensitive actions, no secrets in the prompt, output filtering, and never letting retrieved text trigger actions unchecked. I add an input and output classifier as another layer and log suspicious attempts. Before launch I run adversarial tests, including indirect injection through documents. No single defence works, so I layer them and plan for the case where one fails.

Answers that lose you the room

  • Relies on 'ignore malicious instructions' in the prompt
  • Trusts retrieved text as instructions
  • Gives tools broad permissions

Expect this follow-up: How would you test your defences before launch?

Scenario 11ChunkingPractitioner

How do you decide on a chunking strategy for documents?

What the interviewer is testing: Whether chunking follows structure and is tuned on metrics.

Model answer, red flags and follow-up

A strong answer

I start with the document's own structure: headings, sections, lists and tables, so chunks hold complete ideas. Then I tune size and overlap against retrieval metrics on real user questions, since the best setting depends on the content and the questions. Each chunk carries metadata such as source, section title, date and access level, which supports filtering and citation. Tables and code need special handling so rows and columns stay together. I treat chunking as an experiment: change one parameter, re-run retrieval evals, and keep what measurably improves recall.

Answers that lose you the room

  • Uses one fixed chunk size everywhere
  • Ignores headings and tables
  • Has no retrieval metrics on real questions

Expect this follow-up: A table gets split across chunks. How do you handle it?

Scenario 12EmbeddingsPractitioner

How do you choose an embedding model and vector store?

What the interviewer is testing: Whether you test embeddings on your own data and languages.

Model answer, red flags and follow-up

A strong answer

I compare candidate embedding models on my own retrieval eval, using real queries and documents in the languages my users write in. I weigh quality against dimensions, storage, latency, cost and licensing. For the vector store I look at scale, metadata filtering, hybrid search support, update patterns, and whether it fits our operations, such as managed versus self-hosted. Leaderboards are a starting shortlist, not a decision. I also plan for re-embedding, since changing the model later means reprocessing the whole corpus.

Answers that lose you the room

  • Picks the top leaderboard model
  • Doesn't test on their own data
  • Ignores filtering and hybrid support

Expect this follow-up: Retrieval is good in English but poor in Hindi. What do you check?

Scenario 13Hybrid SearchPractitioner

Why combine keyword and vector search?

What the interviewer is testing: Whether you know where vector search alone fails.

Model answer, red flags and follow-up

A strong answer

Vector search captures meaning but often misses exact tokens such as order IDs, product codes, names and rare terms. Keyword search does the opposite: precise on exact terms, weak on paraphrase. Combining both, then reranking the merged list, gives better recall and precision than either alone. I confirm the benefit on a set of real queries rather than assuming it, and I tune the fusion weights. For structured identifiers I sometimes add a direct lookup path, because no retrieval trick beats an exact match.

Answers that lose you the room

  • Says vectors alone are enough
  • Doesn't know about exact-term failures
  • Adds no reranker

Expect this follow-up: Users search by order ID and get nothing. Why, and what's the fix?

Scenario 14Function CallingPractitioner

How do you make function calling reliable?

What the interviewer is testing: Whether tool calls are validated and evaluated.

Model answer, red flags and follow-up

A strong answer

Reliability starts with the tool definitions: clear names, precise descriptions and strict typed schemas with enums for closed choices. I validate every argument in code before executing, and return an informative error to the model so it can correct itself, with a hard cap on retries. Fewer, well-separated tools are chosen more accurately than many overlapping ones. I evaluate tool-selection and argument accuracy on realistic tasks, including ambiguous ones, and I log every call so failures can be traced and turned into new test cases.

Answers that lose you the room

  • Writes vague function descriptions
  • Does no argument validation
  • Allows unlimited retries

Expect this follow-up: The model calls a function with an invalid argument. What happens next?

Scenario 15Prompt EvaluationPractitioner

How do you evaluate whether a prompt change is actually better?

What the interviewer is testing: Whether you prove improvement statistically and by segment.

Model answer, red flags and follow-up

A strong answer

I run the old and new prompt on the same eval set and compare metrics overall and by segment, because an average can hide a regression in an important group. I read a sample of failures by hand rather than trusting a score alone, and I check whether the difference is bigger than run-to-run noise. If it holds up, I release to a small share of traffic first and watch real-world metrics before going wide. Judging by a handful of examples is how teams ship regressions with confidence.

Answers that lose you the room

  • Judges by a few examples
  • Reports only averages
  • Ships without a canary

Expect this follow-up: The new prompt is better overall but worse for one segment. What do you do?

Scenario 16CachingPractitioner

Where can caching reduce cost and latency in an LLM app?

What the interviewer is testing: Whether you cut cost without leaking or staling answers.

Model answer, red flags and follow-up

A strong answer

Prefix or prompt caching cuts cost and latency when a long system prompt or shared document is reused across requests. Response caching helps for identical queries. Embedding caching avoids recomputing vectors for unchanged text, and semantic caching can serve near-duplicate questions when the answers are safe to reuse. The risks are staleness and privacy: cached answers can go out of date, and a cache shared between users must never leak one person's data to another. I set expiry, scope caches by tenant, and measure hit rate to confirm the saving is real.

Answers that lose you the room

  • Caches everything
  • Ignores privacy and staleness
  • Doesn't know prefix caching

Expect this follow-up: What would you never put in a semantic cache?

Scenario 17Fine-tuningAdvanced

When is fine-tuning worth it over prompting?

What the interviewer is testing: Whether you justify fine-tuning with evidence of a gap.

Model answer, red flags and follow-up

A strong answer

Fine-tuning earns its place when prompting has a proven gap: consistent style or format at scale, lower latency and cost from a smaller model, or behaviour prompts cannot reach. It needs quality labelled data and an eval set to prove it worked. I always start with prompting, retrieval and evals, because they are faster to iterate and cheaper to undo. Fine-tuning also adds a maintenance burden, since I must retrain when the base model or the data changes. I fine-tune to close a measured gap, not to feel sophisticated.

Answers that lose you the room

  • Fine-tunes before trying prompting
  • Has no quality labelled data
  • Cannot state the gap fine-tuning would close

Expect this follow-up: What evidence would convince you fine-tuning is needed?

Scenario 18MultilingualPractitioner

How do you handle Indian languages in an LLM feature?

What the interviewer is testing: Whether you test quality per language, including code-mixing.

Model answer, red flags and follow-up

A strong answer

I test quality per language on real user text, not translated English, because performance often drops for lower-resource languages. Users mix languages and write Hindi in Roman script, so I test code-mixing and transliteration explicitly. I tune prompts and retrieval per language, check that the embedding model actually handles them, and involve native speakers in evaluation. I also watch token cost, since some scripts tokenise into many more tokens. If a language falls below the quality bar, I limit the feature or add human review rather than ship something unreliable.

Answers that lose you the room

  • Assumes English quality carries over
  • Ignores code-mixing and transliteration
  • Has no native-speaker evaluation

Expect this follow-up: Users type Hindi in Roman script. What breaks and how do you handle it?

Scenario 19Streaming UXFoundation

How does streaming change the user experience of an LLM feature?

What the interviewer is testing: Whether you understand streaming's effect on perceived latency.

Model answer, red flags and follow-up

A strong answer

Streaming shows the first words within a second or so, which sharply improves perceived latency even though total time is similar. The cost is complexity. Partial output can be wrong or later retracted, structured output cannot be parsed until it is complete, and safety checks may need the whole response. I handle cancellation so abandoned requests stop spending tokens, and I design the UI to handle a response that changes or stops midway. Where a check needs the full text, I may stream a draft and validate before enabling actions.

Answers that lose you the room

  • Says streaming makes it faster overall
  • Ignores validation of partial output
  • Doesn't handle cancellation

Expect this follow-up: Safety checks need the full response. How do you stream anyway?

Scenario 20DeterminismFoundation

How do you get more consistent outputs from an LLM?

What the interviewer is testing: Whether you accept variability and design validation around it.

Model answer, red flags and follow-up

A strong answer

I lower temperature, tighten the prompt, use structured output and fixed examples, and use a seed where the provider supports it. I also accept that full determinism is not guaranteed: batching, hardware and model updates can shift outputs even at temperature zero. So instead of assuming identical outputs, I design tolerance into the feature: validation, tests that check properties rather than exact strings, and monitoring for drift. If exact repeatability matters, such as for audit, I store the output rather than trying to regenerate it.

Answers that lose you the room

  • Says temperature zero makes it deterministic
  • Adds no validation of outputs
  • Ignores structured output

Expect this follow-up: The output still varies at temperature zero. Why?

Scenario 21PII HandlingAdvanced

How do you handle personal data in prompts?

What the interviewer is testing: Whether you minimise and control personal data in prompts.

Model answer, red flags and follow-up

A strong answer

I minimise what I send: only the fields the task needs. Identifiers are redacted or tokenised before the call and restored afterwards where required. I choose providers whose data terms match our obligations, such as no training on our data and suitable retention, and I restrict what is logged and for how long. I map the data flows end to end and review them with security and legal. I also make sure evals and debugging tools do not become a side door that copies personal data into less protected places.

Answers that lose you the room

  • Sends raw personal data to any provider
  • Logs full prompts indefinitely
  • Has no data-flow map

Expect this follow-up: Which fields would you redact before the model call?

Scenario 22Conversation MemoryPractitioner

How do you manage long conversations within context limits?

What the interviewer is testing: Whether summaries preserve the facts that matter.

Model answer, red flags and follow-up

A strong answer

I keep recent turns verbatim, summarise older ones, and store durable facts such as the user's name, preferences and open tasks in a separate memory that I retrieve when relevant. The risk is that summaries quietly drop something important, so I test that key facts survive summarisation across many turns. I also cap history to control cost and latency. For sensitive facts I make memory visible and correctable to the user. The goal is that the conversation feels continuous without paying to resend everything each time.

Answers that lose you the room

  • Sends the entire history every turn
  • Loses key facts when summarising
  • Never tests summarisation

Expect this follow-up: The user's name is forgotten after summarisation. How do you prevent it?

Scenario 23Ambiguous RequirementsFoundation

Product gives you a vague requirement for an AI feature. What do you do?

What the interviewer is testing: Whether you turn vague asks into testable examples.

Model answer, red flags and follow-up

A strong answer

I turn the vague ask into something testable. I talk to the requester about the user problem and what success would look like, then build a quick prototype and run it on real examples. Showing actual outputs to stakeholders is far more effective than debating a spec, because people recognise good and bad results when they see them. I collect those reactions into an eval set, which becomes the definition of done. That way the requirement becomes concrete through iteration instead of waiting for a perfect document.

Answers that lose you the room

  • Starts building without clarifying
  • Waits for a perfect spec
  • Doesn't show concrete outputs

Expect this follow-up: The stakeholder still can't say what 'good' means. What do you do?

Scenario 24DebuggingPractitioner

Outputs are inconsistent across users. How do you investigate?

What the interviewer is testing: Whether you debug by comparing traces and changing one variable.

Model answer, red flags and follow-up

A strong answer

I compare traces across affected and unaffected cases: the input, the retrieved context, the prompt version, the model and its settings, and the tool results. I look for what differs, segment failures by user type, language or input length, and form a hypothesis. Then I change one variable at a time and re-run the evals so I know what fixed it. Changing several things at once makes the result unexplainable. Good tracing from the start makes this fast, which is why I insist on logging the full chain of context per request.

Answers that lose you the room

  • Changes several variables at once
  • Doesn't compare traces
  • Has no hypothesis

Expect this follow-up: What is the first thing you record in a trace so you can compare users?

Scenario 25Model DeprecationPractitioner

A provider deprecates the model you depend on. How do you prepare?

What the interviewer is testing: Whether you plan migration before the deadline.

Model answer, red flags and follow-up

A strong answer

I treat deprecation as a scheduled event, not a surprise. I track provider notices, keep a model abstraction layer, and maintain an eval suite so I can test a replacement quickly on my own tasks. As soon as a successor is available, I compare it, adapt prompts where behaviour differs, and roll it out gradually with canary traffic while watching quality and cost. I keep the old model available until the new one has proven itself, and I never move all traffic at once. Starting early turns a deadline into a routine migration.

Answers that lose you the room

  • Notices only at the deadline
  • Has no abstraction or eval suite
  • Switches all traffic at once

Expect this follow-up: The replacement model behaves differently on your prompts. What do you do?

Scenario 26LatencyPractitioner

Your AI feature has good answers but a p95 latency of twelve seconds. How do you bring it down?

What the interviewer is testing: Whether you diagnose latency by stage and know the levers beyond a faster model.

Model answer, red flags and follow-up

A strong answer

I measure where the time goes: retrieval, prompt size, time to first token and output length. Output tokens usually dominate, so I shorten responses, stream them and cap length. I trim context, use prefix caching, and run independent calls such as retrieval and classification in parallel. Simple requests go to a smaller, faster model and only hard ones to the larger one. I set a latency budget per stage and watch p95, not the average, because users feel the slow tail.

Answers that lose you the room

  • Only proposes a faster model
  • Measures average latency, not p95
  • Has no per-stage timing

Expect this follow-up: Streaming is on and users still complain. What else could be slow?

Scenario 27RetrievalPractitioner

When is a reranker worth the extra latency and cost?

What the interviewer is testing: Whether you add components because measurement shows a gain.

Model answer, red flags and follow-up

A strong answer

A reranker helps when first-stage retrieval finds the right passage but ranks it too low, which is common with large corpora and ambiguous queries. I measure it: compare recall and answer quality at the top few results with and without reranking on real queries. If the right chunk is already first most of the time, the extra hop is waste. I keep the candidate list small to limit latency, and I consider a lighter reranker if speed matters. It earns its place only if the gain is visible in the eval.

Answers that lose you the room

  • Adds a reranker without measuring
  • Ignores the added latency
  • Reranks a huge candidate list

Expect this follow-up: Your reranker adds 400 ms. Product says it is too slow. What are your options?

Scenario 28Evaluation DataAdvanced

You are building a RAG assistant and have no labelled data. How do you create an eval set?

What the interviewer is testing: Whether you can bootstrap evaluation honestly from real material.

Model answer, red flags and follow-up

A strong answer

I start with real questions from support tickets, search logs or subject experts, since they reflect real usage. Where those are thin, I generate synthetic questions from the documents, then have an expert review and correct them, because unreviewed synthetic data flatters the system. Each item gets an expected answer or a rubric and the source passage. I include hard cases: ambiguous, multi-document and unanswerable questions. I keep a held-out part untouched for final checks and grow the set from production failures.

Answers that lose you the room

  • Uses only unreviewed synthetic questions
  • Tunes on the same set they report on
  • Includes no unanswerable questions

Expect this follow-up: How do you stop the synthetic questions from being too easy?

Scenario 29Model SelectionAdvanced

When would you use a reasoning model instead of a standard model, and what does it cost you?

What the interviewer is testing: Whether you match model type to task difficulty and account for the trade-offs.

Model answer, red flags and follow-up

A strong answer

Reasoning models help on multi-step problems such as planning, complex analysis or code with many constraints, where extra thinking improves correctness. They cost more, respond more slowly and consume extra tokens for the hidden reasoning. For simple extraction, classification or lookup they add cost with little benefit. I test both on my eval set by task type and route accordingly, sending only hard requests to the reasoning model. I also check that prompts written for standard models do not over-constrain the reasoning one.

Answers that lose you the room

  • Uses the reasoning model for everything
  • Ignores latency and token cost
  • Never compares against a standard model on their eval

Expect this follow-up: How would you decide which requests get routed to the reasoning model?

Scenario 30BehaviouralAdvanced

Tell me about an AI feature you built that did not work in production. What did you do?

What the interviewer is testing: Whether you own failures, learn from them and change your process.

Model answer, red flags and follow-up

A strong answer

A strong answer names a specific feature and what went wrong in concrete terms, such as answers that looked fine in demos but failed on real user phrasing. I would explain how I detected it, through user feedback, monitoring or an eval gap, and what I did first to limit harm. Then the root cause: usually an eval set that did not resemble real traffic. The lasting change is process: adding production samples to the eval set and gating releases on it. I would say plainly what I got wrong.

Answers that lose you the room

  • Blames the model or the users
  • Describes a success dressed up as a failure
  • Cannot say what changed afterwards

Expect this follow-up: What would you have needed to see before launch to catch it?

0 of 30 attempted