Thirty scenarios on deployment, monitoring, cost and latency control, reliability, self-hosting and incident response. Write your own answer first, by typing or speaking, then open the model answer to compare structure and reasoning.
30 scenariosWhat each question testsRed-flag answersNo sign-up, private
Answer in your own words before opening a model answer, by typing or by pressing Speak your answer. Your text is saved in this browser only and is never uploaded to us. Voice input uses your browser's speech service to turn speech into text; in Chrome that audio is processed by Google.
Voice input is not supported in this browser. Try Chrome, Edge or Safari, or use your keyboard's dictation (Windows + H on Windows, the microphone key on your phone keyboard).
No scenarios match that combination. Choose another topic or level.
Scenario 1DeploymentPractitioner
How do you deploy and version LLM applications safely?
What the interviewer is testing: Whether you version prompts, models and configs together.
Model answer, red flags and follow-up
A strong answer
I version prompts, model identifiers and configuration together as one release unit, stored in source control, and infrastructure is defined as code. Changes move through environments with eval gates, then reach production through a canary or shadow release with a small share of traffic. Rollback must be instant and tested, ideally a config switch, not a rebuild. Every release records which prompt, model and settings were live, so any behaviour can be traced to a version. Deploying straight to production is how silent quality regressions happen.
Answers that lose you the room
Versions code but not prompts and models
Deploys straight to production
Has no rollback path
Expect this follow-up: How do you deploy a prompt change without redeploying the app?
Scenario 2MonitoringPractitioner
What do you monitor in production LLM systems?
What the interviewer is testing: Whether you monitor quality and cost, not just uptime.
Model answer, red flags and follow-up
A strong answer
I monitor the operational basics: latency percentiles, error rates, throughput, token usage and cost. On top of that I watch quality signals: a sampled evaluation of outputs, guardrail hits, refusal rates, user feedback and drift in the types of inputs arriving. Each has agreed thresholds and alerts that reach the right on-call person. Dashboards show trends by feature and model version. The gap in many teams is quality monitoring, since a service can be perfectly healthy technically while producing poor answers.
Answers that lose you the room
Monitors uptime only
Doesn't sample quality
Sets alerts with no thresholds
Expect this follow-up: Everything is green but users complain. What is missing?
Scenario 3Cost ControlPractitioner
How do you control LLM costs at scale?
What the interviewer is testing: Whether you control cost without cutting quality.
Model answer, red flags and follow-up
A strong answer
I attack cost from several angles. Caching removes repeated work, prompt trimming and history caps reduce tokens, routing sends easy requests to smaller models, and batching handles non-urgent work at lower prices. I set budgets and rate limits per team and feature, and track cost per successful task, since a cheap failing request is waste. Regular reviews of the top spenders find surprises early. Every optimisation is checked against the eval set so savings do not silently cost quality.
Answers that lose you the room
Cuts quality to save cost
Tracks no cost per task
Sets no per-team budgets
Expect this follow-up: Which cost lever do you pull first, and why?
Scenario 4LatencyPractitioner
How do you reduce latency for a chat application?
What the interviewer is testing: Whether you reduce latency with measured techniques.
Model answer, red flags and follow-up
A strong answer
I measure where time goes first, then act on the biggest part. Streaming makes responses feel faster, and prompt caching and shorter context cut processing time. I run independent steps, such as retrieval and classification, in parallel, use smaller or faster models for simple requests, and deploy close to users. Output length often dominates, so I constrain it. I set a latency budget for each step and track p95 and p99, because the slow tail is what users remember.
Answers that lose you the room
Only switches to a smaller model
Ignores streaming and caching
Measures average latency
Expect this follow-up: p95 is bad but the average is fine. What do you check?
Scenario 5ReliabilityAdvanced
How do you handle provider outages and rate limits?
What the interviewer is testing: Whether you design for provider outages and limits.
Model answer, red flags and follow-up
A strong answer
I design for failure from the start. Every call has timeouts and retries with exponential backoff and jitter, and circuit breakers stop hammering a failing provider. Behind a model abstraction I keep a fallback provider or model, and queue non-urgent work. When everything fails, the system degrades gracefully, for example by showing a cached answer or a non-AI path. I respect rate limits with client-side throttling, and I rehearse failover regularly, because untested failover tends not to work when needed.
Answers that lose you the room
Uses a single provider with no fallback
Retries with no backoff
Never rehearses failover
Expect this follow-up: The fallback model behaves differently. How do you manage that?
Scenario 6Self-hostingPractitioner
When would you self-host an open model instead of using an API?
What the interviewer is testing: Whether you weigh self-hosting on total cost.
Model answer, red flags and follow-up
A strong answer
I consider self-hosting when data must stay within our control, when volume is high and steady enough to keep GPUs busy, when I need customisation such as fine-tuning, or when latency needs cannot be met by an API. I compare the full cost, including GPUs, engineering time, on-call, upgrades and idle capacity, against API pricing. At low or spiky volume, APIs are usually cheaper and simpler. I also weigh model quality, since the best hosted models may outperform open ones for my task.
Answers that lose you the room
Assumes self-hosting is always cheaper
Ignores GPU and on-call costs
Has no volume estimate
Expect this follow-up: At what volume does self-hosting break even?
Scenario 7Quality DriftAdvanced
Quality dropped and no code changed. What do you investigate?
What the interviewer is testing: Whether you diagnose drift when no code changed.
Model answer, red flags and follow-up
A strong answer
Since no code changed, I look outside our repository. The provider may have updated the model behind an alias, the mix of user inputs may have shifted, the retrieval index may be stale or a data pipeline may have broken. Seasonal or event-driven effects also matter. I compare current results with the eval set and traces from before the drop, segment by input type and time, and check provider change logs and index freshness. To prevent recurrence, I pin model versions and monitor quality continuously.
Answers that lose you the room
Says nothing changed so it can't be us
Ignores provider updates
Has no eval set to compare against
Expect this follow-up: How would you confirm the provider changed the model?
Scenario 8IncidentPractitioner
A prompt change caused a spike of bad outputs at 2am. How do you respond?
What the interviewer is testing: Whether you roll back first and learn afterwards.
Model answer, red flags and follow-up
A strong answer
First I stop the harm: roll back to the last known good version, then confirm recovery in the metrics. I communicate status to stakeholders and keep a timeline. Once stable, I capture the failing examples, find the cause and add those examples to the eval suite so the release gate would catch them next time. I tighten the release process, for example with canary stages or required eval passes for prompt changes. The follow-up is a blameless review focused on how the system allowed the change through.
Answers that lose you the room
Debugs live before rolling back
Skips the blameless review
Doesn't turn failures into tests
Expect this follow-up: What release gate would have caught this?
Scenario 9Prompt RegistryPractitioner
Why use a prompt registry, and what should it support?
What the interviewer is testing: Whether prompts get ownership, versions and audit.
Model answer, red flags and follow-up
A strong answer
A prompt registry treats prompts as managed assets. It gives each prompt versions, owners, environments, diffs and linked eval results, and supports instant rollback. Where appropriate it decouples prompt changes from application releases, so a fix does not require a full deployment, while access controls, review and audit history keep this safe. It also helps non-engineers such as product or domain experts propose changes through a controlled process. Without it, prompts get edited in place and nobody knows what was running when.
Answers that lose you the room
Keeps prompts scattered across code
Has no owners or audit history
Has no environment separation
Expect this follow-up: Who should be allowed to change a production prompt?
Scenario 10CI/CDPractitioner
What does CI/CD look like for an LLM application?
What the interviewer is testing: Whether evals gate releases in your pipeline.
Model answer, red flags and follow-up
A strong answer
It looks like conventional CI/CD with an extra layer. Code goes through linting and unit tests, then the eval suite runs on prompt, model and retrieval changes, with safety checks. Deployment is staged, with canary monitoring of quality and cost, and automatic rollback if thresholds are breached. The evals play the role of tests, and they gate each release. Because LLM outputs vary, I use statistical thresholds and confidence intervals instead of exact matches, and I keep the pipeline fast enough that developers do not bypass it.
Answers that lose you the room
Skips evals in the pipeline
Tests only the code
Has no automatic rollback
Expect this follow-up: What fails the build in your pipeline?
Scenario 11Model GatewayPractitioner
What does a model gateway provide?
What the interviewer is testing: Whether you centralise access, routing and policy.
Model answer, red flags and follow-up
A strong answer
A model gateway gives applications a single interface to multiple providers and models. It centralises authentication, rate limits, routing, retries, caching, logging, cost tracking and policy enforcement, so those features are not re-implemented in every service. It reduces lock-in because switching providers becomes a configuration change, and it gives security and finance a single point of visibility. The trade-off is another critical component to keep reliable and low-latency, so I treat it as production infrastructure with its own SLOs.
Answers that lose you the room
Calls providers directly from each app
Has no central logging
Enforces no policy
Expect this follow-up: What would you put in the gateway first?
Scenario 12TracingPractitioner
What should LLM tracing capture?
What the interviewer is testing: Whether traces make production bugs debuggable.
Model answer, red flags and follow-up
A strong answer
Each trace should capture the prompt and its version, the retrieved context, the model and parameters, tool calls and results, latency, token counts and cost, and any user feedback, all linked by a trace ID. This makes debugging possible, since I can reconstruct exactly what the model saw. It also lets me build evaluation sets from real production data and analyse cost by feature. I handle privacy through redaction and retention limits, and sample where volume is very high.
Answers that lose you the room
Logs only final responses
Uses no trace IDs
Stores raw personal data
Expect this follow-up: How do you debug one bad conversation from production?
Scenario 13Semantic CachingAdvanced
When is semantic caching a good idea?
What the interviewer is testing: Whether you cache safely and avoid wrong shared answers.
Model answer, red flags and follow-up
A strong answer
Semantic caching suits high-volume, repetitive questions where near-duplicates should get the same answer, such as FAQ-style support. I set the similarity threshold carefully and test it, because a threshold that is too loose returns wrong answers confidently. I avoid caching personalised, time-sensitive or sensitive responses, scope the cache by tenant and permissions, and expire entries. I measure the hit rate and the quality of cached answers to confirm the saving is worth the risk.
Answers that lose you the room
Caches every response
Sets loose similarity thresholds
Caches personalised answers
Expect this follow-up: A cached answer is wrong for a different user. How did that happen?
Scenario 14GPU AutoscalingAdvanced
How do you autoscale GPU inference?
What the interviewer is testing: Whether you scale GPUs on the right signals.
Model answer, red flags and follow-up
A strong answer
I scale on signals that reflect real load, such as queue depth, request latency and GPU utilisation, not just CPU. GPUs take minutes to start and models take time to load, so I keep some warm capacity and scale ahead of predictable peaks. Batching and mixed instance types improve utilisation, and I set cost ceilings to avoid runaway spend. I load-test with realistic prompt and output lengths to find true limits, since token counts vary widely and throughput depends on them.
Answers that lose you the room
Scales on CPU only
Ignores cold starts
Sets no cost ceiling
Expect this follow-up: Traffic spikes at 9am daily. How do you handle cold starts?
Scenario 15Inference OptimisationAdvanced
How can you make self-hosted inference cheaper and faster?
What the interviewer is testing: Whether you optimise inference while validating quality.
Model answer, red flags and follow-up
A strong answer
The main levers are quantisation to reduce memory and speed up inference, continuous batching to raise throughput, KV-cache reuse for shared prefixes, and an optimised serving engine such as vLLM. I right-size hardware, and consider smaller or distilled models where quality allows. Speculative decoding can cut latency for some workloads. Each change is validated against the eval set, because optimisations can degrade quality in subtle ways. I measure cost per thousand tokens and latency, not just raw speed.
Answers that lose you the room
Buys bigger GPUs first
Doesn't validate quality after quantisation
Ignores batching
Expect this follow-up: How do you check quality didn't drop after quantising?
Scenario 16Multi-tenancyAdvanced
How do you isolate tenants in a shared LLM platform?
What the interviewer is testing: Whether tenants are isolated and that is tested.
Model answer, red flags and follow-up
A strong answer
I isolate tenants at every layer: separate data stores or indexes, or strict filtering enforced in the retrieval layer, per-tenant keys, quotas and rate limits, and separate logs. Caches must not be shared across tenants, and prompts must never include another tenant's data. I test isolation explicitly with cross-tenant probes, including prompt-injection attempts to extract other tenants' data. A noisy-neighbour problem is also possible, so fair-use limits protect performance for everyone.
Answers that lose you the room
Shares indexes and caches across tenants
Sets no per-tenant quotas
Never tests isolation
Expect this follow-up: How do you prove one tenant can't see another's data?
Scenario 17SecretsFoundation
How do you manage API keys and secrets for LLM services?
What the interviewer is testing: Whether you handle keys and secrets safely.
Model answer, red flags and follow-up
A strong answer
API keys live in a secrets manager, never in code, prompts or logs. They are scoped by service and environment, rotated regularly, and issued with the minimum permissions. I monitor usage for anomalies such as sudden spikes, which can indicate a leak, and set spending limits on keys. Developers get separate keys from production. Where possible I use short-lived credentials or workload identity, and have a tested procedure for revoking and replacing a compromised key quickly.
Answers that lose you the room
Keeps keys in code or config files
Never rotates them
Puts secrets in prompts or logs
Expect this follow-up: A key leaks in a log. What do you do?
Scenario 18Logging & PrivacyPractitioner
How do you balance detailed logging with privacy?
What the interviewer is testing: Whether you balance debuggability with privacy.
Model answer, red flags and follow-up
A strong answer
I log what is needed for debugging and evaluation, and no more. Sensitive fields are redacted or hashed at the point of logging, retention is short, and access is restricted and audited. A separate, consented or sanitised sample set can be kept for review and evals. Privacy requirements from legal and the data protection team decide what may be stored and for how long. The trade-off is real: less logging makes debugging harder, so I invest in good redaction rather than turning logs off.
Answers that lose you the room
Logs everything forever
Logs nothing to protect privacy
Sets no access restrictions
Expect this follow-up: Debugging needs full prompts. How do you allow that safely?
Scenario 19Index RefreshPractitioner
How do you keep a RAG index fresh?
What the interviewer is testing: Whether you keep retrieval indexes fresh and clean.
Model answer, red flags and follow-up
A strong answer
I trigger incremental ingestion when sources change, and schedule periodic full checks for anything missed. Deleted or updated documents must be removed or re-embedded so stale answers disappear. I version indexes so I can roll back a bad refresh, and I monitor freshness, ingestion failures and index size. After each refresh I run a set of retrieval tests to confirm quality did not drop. Stale data is a quiet failure, so I alert on lag between the source and the index.
Answers that lose you the room
Re-indexes everything manually
Ignores deleted documents
Doesn't monitor freshness
Expect this follow-up: A document is deleted at the source. How does it leave the index?
Scenario 20Feature FlagsPractitioner
How do feature flags help with LLM releases?
What the interviewer is testing: Whether you decouple deployment from release.
Model answer, red flags and follow-up
A strong answer
Feature flags let me separate deployment from release. I can expose a new prompt or model to a small share of traffic or specific segments, compare metrics with the control, and switch it off instantly if something goes wrong without redeploying. They also support experiments and gradual rollout by tenant. I keep flags well managed, with owners and expiry dates, so they do not accumulate into untested combinations of settings.
Answers that lose you the room
Releases to everyone at once
Has no way to switch off instantly
Ties release to deployment
Expect this follow-up: How would you run a 5% rollout of a new model?
Scenario 21SLOsPractitioner
How do you define SLOs for an LLM service?
What the interviewer is testing: Whether SLOs include quality, not only availability.
Model answer, red flags and follow-up
A strong answer
I define SLOs that reflect the user's experience: availability, latency percentiles, error rate and quality indicators from sampled evals. Each has an objective and an error budget that guides how much risk we can take with releases. Because quality is harder to measure than uptime, I make the quality SLO explicit and track it. I review SLOs after incidents and adjust them when they do not match what users care about. Alerts are based on budget burn, not on every blip.
Answers that lose you the room
Sets availability targets only
Adds no quality indicators
Has no error budgets
Expect this follow-up: How do you turn a quality dip into an SLO breach?
Scenario 22Cost AttributionPractitioner
How do you attribute LLM costs to teams and features?
What the interviewer is testing: Whether costs are attributable by team and feature.
Model answer, red flags and follow-up
A strong answer
I tag every request with team, feature and environment through the gateway, then aggregate token and infrastructure costs in dashboards. Each team gets a budget with alerts, and anomalies are reviewed weekly. Shared costs, such as a common index or platform overhead, are allocated using a documented rule. Making cost visible changes behaviour: teams notice wasteful prompts when they can see the bill. I also report cost per successful outcome, so cheap-but-failing features do not look efficient.
Answers that lose you the room
Tracks total spend only
Uses no tags by team or feature
Reviews costs annually
Expect this follow-up: One feature's cost spikes. How fast can you find it?
Scenario 23Model RegistryAdvanced
How do you manage fine-tuned models across their lifecycle?
What the interviewer is testing: Whether models trace back to data, code and evals.
Model answer, red flags and follow-up
A strong answer
I use a model registry that records each model's lineage: training data version, code, hyperparameters and evaluation results. Models move through stages, such as staging and production, with approvals and automatic checks, and I can reproduce any model from its record. Rollback to a previous version is straightforward. Every deployed model traces to its evidence, which supports audit and debugging. I also track dependencies on base models, so a base-model deprecation triggers a plan.
Answers that lose you the room
Stores models with no lineage
Cannot reproduce training
Requires no approvals
Expect this follow-up: An auditor asks how model v3 was produced. What do you show?
Scenario 24RunbooksFoundation
What belongs in an on-call runbook for an LLM service?
What the interviewer is testing: Whether on-call runbooks are concrete and maintained.
Model answer, red flags and follow-up
A strong answer
It lists the symptoms and the dashboards to check, how to confirm provider status, and step-by-step actions for rollback and failover. It covers known failure modes and their fixes, escalation contacts, and templates for communicating with users and stakeholders. It is written to be usable at 3am by someone who did not build the system. I update it after every incident and test it in game days, because an out-of-date runbook is worse than none.
Answers that lose you the room
Writes runbooks with generic steps
Leaves out rollback and failover steps
Never updates them after incidents
Expect this follow-up: What is on the first page of your runbook?
Scenario 25StagingPractitioner
How do you make staging environments realistic for LLM apps?
What the interviewer is testing: Whether staging exposes real-world problems.
Model answer, red flags and follow-up
A strong answer
I make staging as close to production as practical: the same configuration, the same model versions, representative traffic patterns and real provider rate limits. Data is sanitised or synthetic but realistic in shape and difficulty. I run the eval suite there and load-test at expected volumes. Mocked model responses hide exactly the problems I need to find, such as latency, rate limiting and unexpected outputs, so I use real calls with cost controls. A staging environment that behaves differently gives false confidence.
Answers that lose you the room
Uses mocked responses only
Uses unrealistic traffic and data
Skips real provider limits
Expect this follow-up: What issue only shows up in realistic staging?
Scenario 26ObservabilityAdvanced
How would you detect that an LLM application has started producing lower-quality answers before customers complain?
What the interviewer is testing: Whether you monitor quality proactively, not only uptime.
Model answer, red flags and follow-up
A strong answer
I combine several signals. An LLM judge or rule-based checks score a sample of live traffic against rubrics, and a small share is reviewed by humans to keep the judge honest. I watch proxy signals such as edit rates, retries, abandonment and thumbs-down, and I track drift in input topics, output length and refusal rates. Alerts fire on shifts beyond normal variation. Detection is only useful if flagged cases feed the eval set and an owner investigates.
Answers that lose you the room
Monitors only uptime and latency
Waits for customer complaints
Uses a judge that was never validated
Expect this follow-up: The judge score is stable but complaints rise. What do you check?
Scenario 27Rate LimitsPractitioner
Your traffic doubles during a marketing campaign and you hit provider rate limits. What do you do in the moment and afterwards?
What the interviewer is testing: Whether you can manage capacity and degrade gracefully under load.
Model answer, red flags and follow-up
A strong answer
In the moment I shed or queue low-priority work, apply client-side throttling, turn on caching, route eligible traffic to a secondary provider or smaller model, and communicate status. Afterwards I ask what warning we missed: capacity was not requested in advance, there was no load test and there was no priority tiering. I request higher limits, set quotas per feature, add autoscaling queues and rehearse the scenario. Marketing should tell engineering about campaigns in advance.
Answers that lose you the room
Only asks the provider for more quota
Has no priority between traffic types
Never load-tests the peak
Expect this follow-up: Which requests would you drop first, and how would you decide?
Scenario 28DeploymentPractitioner
How do you roll out a new model version when you cannot fully predict how it will behave?
What the interviewer is testing: Whether you use staged exposure and comparison, not blind switches.
Model answer, red flags and follow-up
A strong answer
I run the eval suite first, including regression and safety tests, then shadow the new model on live traffic without showing its outputs so I can compare results, cost and latency. Next comes a canary at a few percent, watching quality, guardrail hits and user signals, and then a gradual ramp with flags for instant rollback. Prompts may need adjusting for behavioural differences. I keep the old model available until the new one has proven itself over a full traffic cycle.
Answers that lose you the room
Switches all traffic at once
Relies only on public benchmarks
Removes the old model immediately
Expect this follow-up: The canary looks fine on average but worse for one customer. What now?
Scenario 29Data GovernancePractitioner
How do you handle user data that ends up in prompts, logs and traces in your LLM platform?
What the interviewer is testing: Whether you treat observability data as sensitive data.
Model answer, red flags and follow-up
A strong answer
I classify what may appear in prompts and traces, then apply redaction at the point of capture, encryption, short retention and role-based access. Debug access is audited. Data that must be kept for evaluation is sampled, minimised and, where required, consented. I honour deletion requests across logs, caches and indexes. I make sure third-party observability tools are covered by the same data terms. Traces are enormously useful, but they are also a copy of user data that needs protecting.
Answers that lose you the room
Logs full prompts forever
Gives all engineers access to traces
Forgets caches when handling deletion
Expect this follow-up: A user requests deletion. Where might their data still exist in your platform?
Scenario 30BehaviouralAdvanced
Tell me about an outage or incident you handled on an ML or LLM system. What did you change afterwards?
What the interviewer is testing: Whether you respond calmly and turn incidents into lasting improvements.
Model answer, red flags and follow-up
A strong answer
A strong answer gives a concrete incident with its impact and timeline: how it was detected, what I did to stabilise it and how I communicated. Then the root cause, often a combination of factors instead of one mistake, and the fixes: better alerts, release gates, runbooks or capacity planning. I would describe the blameless review and one specific change that measurably reduced the risk of a repeat. I would be honest about what I would do differently.
Answers that lose you the room
Blames one person
Cannot describe any lasting change
Focuses only on the heroics
Expect this follow-up: How did you know your fix worked?
0 of 30 attempted
Keep going
Practise the neighbouring roles, or see every question in one place.