ML engineering interviews are unusual in that they test two different people. One round wants the statistics — why your model generalises, what your metric actually measures. Another round wants the engineer — how the model gets served, what happens when the data shifts, what you do at three in the morning when predictions go sideways. Candidates are usually strong at one and vague about the other, and the vague half is where offers are lost.
Every answer below is written as something you can say out loud in the time an interview actually gives you. Where a formula, a number or a real failure mode carries the point, it is there. Where it would only pad the answer, it is not.
Click any question to open it. Expand all when revising, closed when you want to test yourself.
How to use this properly: read it once, close the toggle, then say it back in your own words. If you cannot get through it without reopening, you do not know it yet. Recognising a correct answer and producing one under pressure are completely different skills, and only one of them is being assessed.
Tier 1 — Machine Learning and Data Fundamentals
The screening round. These look basic, and the follow-ups are where people come apart — particularly anything about leakage and splitting.
Q1 Explain the bias-variance trade-off.
Bias is error from the model being too simple to capture the real pattern — it makes the same mistakes regardless of which data you train it on. Variance is error from the model being too sensitive to the particular training set, so it fits noise and changes a lot if you resample the data. Total generalisation error decomposes into bias squared, variance, and irreducible noise you cannot do anything about.
The trade-off is that making a model more flexible reduces bias and increases variance. A linear model on a curved relationship is high bias. A deep tree grown until every leaf is pure is high variance.
How I read it in practice. Compare training and validation error. High error on both means high bias — the model is underfitting, so I need more capacity or better features. Low training error with a large gap to validation means high variance — so more data, regularisation, or a simpler model. That diagnostic is what interviewers actually want, more than the decomposition itself.
Q2 How do you detect and prevent overfitting?
I detect it by holding data out. If training performance keeps improving while validation performance flattens and then degrades, the model has started memorising. Plotting both curves against training epochs or model complexity makes it obvious.
Preventing it depends on where the capacity is coming from. More data is the most reliable fix when it is available, including augmentation for images and audio. Regularisation constrains the weights — L1 and L2, dropout in neural networks, or limiting tree depth and requiring a minimum number of samples per leaf in tree models. Early stopping halts training at the point validation loss turns. And simplifying the feature set helps more than people expect, because every extra weakly informative feature is another opportunity to fit noise.
The point I would add. Any hyperparameter I tune against the validation set leaks a little information into it, so validation performance is optimistically biased after a long search. That is why a genuinely untouched test set, looked at once at the end, matters.
Q3 Supervised, unsupervised, self-supervised, reinforcement learning — what is the difference?
Supervised learning has labelled examples and learns a mapping from inputs to outputs — classification and regression. It is most of what gets deployed commercially, because most business problems come with a label attached somewhere in the data.
Unsupervised learning has no labels and finds structure — clustering customers, reducing dimensionality, detecting anomalies. Useful for exploration, harder to evaluate precisely because there is no ground truth.
Self-supervised learning generates its own labels from the structure of unlabelled data — predicting the next token, or a masked patch of an image. This is the technique behind essentially every foundation model, and it matters because it turns an abundant resource, unlabelled data, into supervision.
Reinforcement learning learns from a reward signal by acting in an environment. It is the right frame for sequential decision making — control, recommendations with long-term value — and in practice it also underlies preference tuning for language models.
Q4 How do you split your data, and when is a random split wrong?
The default is a random train, validation and test split — train to fit, validation to tune, test untouched until the end. But random is wrong more often than people assume, and the exceptions are the interesting part of this question.
With time-dependent data, a random split lets the model train on the future and predict the past, which inflates every number and collapses in production. There you split chronologically and validate on a later window than you trained on.
With grouped data, a random split puts rows from the same entity on both sides — several visits from the same patient, several photos of the same person — so the model can recognise the entity instead of learning the pattern. There you split by group.
And with a rare positive class, I stratify so each split has a representative share, otherwise the validation estimate is dominated by sampling noise.
The one-line version I would offer. The split should mimic how the model will actually be used, which almost always means predicting something it has not seen, in a period it was not trained on.
Q5 What is data leakage, and how does it happen in practice?
Leakage is when information that would not be available at prediction time gets into training, so the model looks excellent offline and fails in production. It is the single most common reason a model with a 0.97 AUC becomes useless on deployment.
There are two main shapes. Target leakage, where a feature is a consequence of the outcome rather than a cause — a discount_applied_at_cancellation column when predicting cancellation, or an account status field that is only set after the event. And preprocessing leakage, where a transform is fit on the full dataset before splitting, so the scaler or the imputation mean or the target encoding has already seen the test set.
Example of the preprocessing case.
# leaks — the scaler has seen the test data
X = scaler.fit_transform(X_all)
X_train, X_test = train_test_split(X)
# correct — fit only on train, apply to test
pipe = Pipeline([("scale", StandardScaler()), ("model", LogisticRegression())])
pipe.fit(X_train, y_train) # cross_val_score on the pipeline is safe too
How I catch it. Suspiciously high scores are the first signal — I treat an unexpectedly good result as a bug until proven otherwise. Then I check feature importances, because leaked features usually dominate them, and I ask for each top feature when its value is actually written.
Q6 How do you handle missing data?
First I try to find out why it is missing, because that determines what is safe. Missing completely at random is the easy case. Missing because of something related to the outcome — an income field that people with high income skip — means the missingness is itself informative and dropping those rows biases the model.
Given that, the options are: drop the column if it is mostly empty and not important; drop rows only if there are few of them and the missingness looks random; or impute. Median for skewed numerics, mode or an explicit "unknown" category for categoricals, and for tree models a sentinel value works because trees can split on it. I usually add a boolean was_missing indicator alongside the imputed value, so if the missingness carries signal the model can use it.
The detail that matters. Imputation is fit on training data only and applied to validation and test, exactly like scaling — otherwise it is leakage. And some models, notably LightGBM and XGBoost, handle missing values natively by learning a default direction at each split, so imputing before them can actually lose information.
Q7 Your positive class is 1% of the data. How do you approach it?
First I change the metric, because accuracy is meaningless here — predicting the majority class always gives 99%. I would look at precision, recall and PR-AUC, and I would ask the business which error is more expensive, since that determines whether recall or precision matters more.
Then, in roughly increasing order of intervention: class weighting, which tells the loss function to penalise minority errors more heavily and is usually the first thing I try because it changes nothing about the data. Then resampling — undersampling the majority when there is plenty of data, oversampling or SMOTE for the minority, applied inside the cross-validation fold rather than before splitting, or the validation set gets contaminated with synthetic neighbours of its own rows.
The thing worth saying explicitly. Resampling changes the base rate, so the model outputs stop being calibrated probabilities. If a downstream system multiplies the score by a value to make a decision, that matters and you need to recalibrate. For heavy imbalance I would also ask whether this is really a classification problem or an anomaly detection one.
Q8 When does feature scaling matter, and when does it not?
It matters for anything that measures distance or is trained by gradient descent. K-nearest neighbours, K-means and SVMs compute distances, so a feature measured in rupees dominates one measured in years purely because of magnitude. Neural networks and linear models with gradient descent converge far faster and more stably on scaled inputs. Any regularised linear model needs it too, because L1 and L2 penalise coefficients, and coefficient size depends on feature scale — so unscaled features get penalised unevenly. PCA needs it because it maximises variance and variance is scale-dependent.
It does not matter for tree-based models. A decision tree splits on thresholds within a single feature at a time, so any monotonic rescaling produces exactly the same tree. That covers random forests, XGBoost and LightGBM.
Standardisation or min-max. Standardisation is my default and is more robust to outliers. Min-max when I need a bounded range, for example feeding image pixels into a network. And either way the scaler is fit on training data only.
Q9 How do you encode categorical features?
One-hot encoding for low cardinality, which is safe and interpretable but explodes the feature space as cardinality grows. Ordinal encoding only when the categories genuinely have an order, like small-medium-large — using it on unordered categories tells a linear model that category three is somehow between two and four, which is false. For tree models, ordinal encoding is often fine even without order, since a tree can carve out individual values with enough splits.
For high cardinality — postcodes, product IDs, user IDs — I use target encoding or learned embeddings. Target encoding replaces the category with a statistic of the target for that category, which is powerful and also the single easiest way to leak: it must be computed out-of-fold and smoothed towards the global mean for rare categories, or the model simply memorises the target.
The operational detail. Whatever I choose has to handle unseen categories at inference time, because production will produce a value the training set never contained. That is a mapping to an "unknown" bucket, decided up front rather than discovered by a stack trace.
Q10 What is the curse of dimensionality, and how does PCA help?
As the number of features grows, the volume of the space grows exponentially, so any fixed number of samples becomes increasingly sparse in it. Distances between points converge — everything becomes roughly equidistant from everything else — which breaks distance-based methods, and the data needed to cover the space grows exponentially, which makes overfitting almost inevitable.
PCA reduces dimensionality by finding orthogonal directions of maximum variance and projecting onto the top few. You keep most of the variance in far fewer dimensions, which speeds up training, reduces overfitting, and removes multicollinearity since the components are uncorrelated.
The caveats I would mention. Components are linear combinations of the originals, so interpretability is gone — a stakeholder cannot be told that component three drove the decision. It requires scaling first. It is unsupervised, so a direction of low variance can still be the one that predicts the target. And for visualisation specifically, UMAP or t-SNE usually reveal structure PCA misses, though neither should be used as a preprocessing step for a model.
Q11 What is cross-validation, and which variant do you use?
K-fold cross-validation splits the data into k parts, trains on k minus one and validates on the remaining one, rotating through all of them and averaging. It gives a more reliable estimate than a single split because every row is validated exactly once, and the spread across folds tells you how stable the model is — a large variance between folds is itself a finding.
Which variant depends on the data, and this is the part being tested. Stratified k-fold for classification, so class proportions are preserved in each fold. Group k-fold when rows share an entity, so the same entity never appears on both sides. Time-series split for temporal data, where each fold trains on the past and validates on the future rather than shuffling.
Cost and honesty. Five-fold means training five times, so for a large model I might use a single well-constructed validation split instead and be explicit that the estimate is noisier. And any preprocessing must live inside the fold, which in scikit-learn means putting it in a Pipeline and cross-validating the pipeline, not the model.
Tier 2 — Modelling, Metrics and Deep Learning
Where they find out whether you understand what your numbers mean.
Q12 Why is accuracy often the wrong metric, and what would you use instead?
Because accuracy treats all errors as equal and is dominated by the majority class. On a dataset with 1% fraud, a model that predicts "not fraud" every time scores 99% and is worthless.
I would use precision and recall, chosen by which error costs more. Precision is what fraction of the things I flagged were real, so it is the metric when a false positive is expensive — blocking a legitimate transaction, or sending a healthy patient for an invasive test. Recall is what fraction of real cases I caught, so it is the metric when a false negative is expensive — missing a fraud, missing a tumour. F1 is their harmonic mean, useful as a single number when both matter roughly equally, though I would rather report both.
The framing that lands well. I would say I do not choose the metric in isolation — I ask what happens downstream when the model is wrong in each direction, and the cost asymmetry picks the metric. For a regression problem the equivalent question is whether large errors are disproportionately bad, which is the difference between MAE and RMSE.
Q13 ROC-AUC or PR-AUC — when do you use each?
ROC-AUC plots true positive rate against false positive rate across all thresholds, and it has a clean interpretation: the probability that a randomly chosen positive scores above a randomly chosen negative. It is threshold-independent and good for comparing models on balanced data.
The problem is that under heavy imbalance it is misleadingly optimistic, because false positive rate has the large negative class in its denominator. Going from a hundred to a thousand false positives barely moves the false positive rate when there are a million negatives, but it destroys precision.
So for imbalanced problems I use PR-AUC, precision against recall, which ignores true negatives entirely and reflects what the user of the model actually experiences. The baseline for PR-AUC is the positive class rate rather than 0.5, which is worth stating so nobody reads 0.4 as a poor score on a 1% problem.
Practical answer. Report both if asked to compare models. Optimise the one that matches the deployment reality.
Q14 Your model outputs a probability. How do you pick the decision threshold?
Not at 0.5, which is only a default and almost never the right operating point. I pick it from the cost of each error type. If a false negative costs ten times what a false positive costs, the threshold that minimises expected cost is well below 0.5, and I can compute that directly from the precision-recall curve on the validation set.
In practice there is often a capacity constraint that decides it instead. If the fraud team can review two hundred cases a day, the threshold is whatever produces two hundred alerts, and the model's job is to make those two hundred as high-precision as possible. Framing it that way — as an operating point chosen with the business, not a hyperparameter — is usually what the interviewer is listening for.
The related point. Thresholding assumes the scores are calibrated. Boosted trees and SVMs produce scores that rank well but are not true probabilities, so if a downstream system treats the output as a probability I would calibrate with Platt scaling or isotonic regression on a held-out set, and check it with a reliability curve.
Q15 What does logistic regression actually output, and how does it differ from linear regression?
Linear regression predicts a continuous value as a weighted sum of the inputs, and fits by minimising squared error. Logistic regression takes that same linear combination and passes it through a sigmoid, which squashes it into zero to one, so the output is a calibrated probability of the positive class. It fits by maximising likelihood, equivalently minimising log loss, rather than squared error.
The reason you cannot just use linear regression for classification is that it produces values outside zero to one, it is highly sensitive to outliers in a way that shifts the decision boundary, and squared error is not the right loss for a probability.
The interpretability point that is worth having ready. The coefficients of a logistic regression are log-odds, so exponentiating one gives an odds ratio — a coefficient of 0.7 on a feature means roughly a doubling of the odds per unit increase. That interpretability is a large part of why logistic regression is still the default in regulated domains like credit and healthcare, where a model has to be explainable to a regulator, not just accurate.
Q16 Decision tree, random forest, gradient boosting — what is the difference?
A single decision tree splits the feature space recursively to reduce impurity. It is fully interpretable and it overfits badly if grown deep, because it will keep splitting until leaves are pure.
A random forest is bagging: many deep trees trained in parallel on bootstrap samples, each considering only a random subset of features at each split, and their predictions averaged. Individual trees are high variance, but averaging decorrelated trees cancels much of that out. It is robust, hard to overfit badly, and needs very little tuning.
Gradient boosting is sequential: each new shallow tree is fit to the residual errors of the ensemble so far, so the model corrects itself incrementally. It generally reaches higher accuracy than a forest but is more sensitive to hyperparameters and can overfit if you keep adding trees without early stopping.
My default. A forest as a strong fast baseline that will not embarrass me, then LightGBM or XGBoost when the accuracy matters and I have time to tune the learning rate, tree depth and number of estimators with early stopping on a validation set.
Q17 Why does gradient boosting usually beat deep learning on tabular data?
Because the inductive biases match the data. Tabular features are heterogeneous — different scales, different meanings, many of them categorical — and the relationships are often irregular step functions rather than smooth ones. Trees split on axis-aligned thresholds, which is exactly that shape, and they are invariant to feature scaling and monotonic transformations. Neural networks are biased toward smooth functions, which is a great fit for images and text where the input is homogeneous and has spatial or sequential structure, and a poor fit for a table of mixed columns.
The practical side reinforces it: gradient boosting handles missing values natively, trains in minutes on a CPU, needs far less data, and needs far less tuning to get most of its performance.
Where I would still choose deep learning on tabular data. Very high cardinality categoricals where learned embeddings genuinely help, multimodal problems where a table sits alongside text or images, or where I want one end-to-end model rather than a pipeline. But for a plain table with a target column, boosting is the honest default and saying so shows judgement rather than fashion.
Q18 What is regularisation, and what is the difference between L1 and L2?
Regularisation adds a penalty on model complexity to the loss, so the optimiser trades a little training fit for better generalisation. It is the direct lever on the variance side of the bias-variance trade-off.
L2, ridge, penalises the sum of squared coefficients. It shrinks all of them toward zero smoothly without reaching it, and it handles correlated features gracefully by spreading weight across them. L1, lasso, penalises the sum of absolute values, and because of the shape of that penalty it drives some coefficients exactly to zero — so it performs feature selection as part of fitting. Elastic net combines both, which is what I would use with many correlated features where I still want sparsity.
How I would explain the intuition if pushed. The L1 constraint region has corners on the axes, and the optimum tends to land on a corner, which means a coefficient of exactly zero. The L2 region is a circle with no corners, so it shrinks but does not zero out. That geometric answer is short and usually ends the follow-up.
Q19 Explain gradient descent, and how do you choose an optimiser and learning rate?
Gradient descent computes the gradient of the loss with respect to the parameters and steps in the opposite direction, repeating until it converges. Batch gradient descent uses the whole dataset per step, which is stable but slow and often does not fit in memory. Stochastic gradient descent uses one sample, which is fast and noisy. Mini-batch, typically thirty-two to five hundred and twelve samples, is what everyone actually uses — it balances gradient quality against hardware efficiency.
The learning rate is the most important hyperparameter. Too high and the loss diverges or oscillates; too low and training crawls or settles into a poor region. I find it with a short learning rate range test and then use a schedule — warmup followed by cosine decay is a reliable default — rather than a constant value.
Optimiser choice. Adam or AdamW as the default because they adapt per-parameter and are forgiving about the initial learning rate, which is why they dominate transformer training. Plain SGD with momentum still generalises slightly better on some vision tasks and is worth trying if you have the budget. AdamW specifically because it decouples weight decay from the adaptive step, which is the version you want.
Q20 Explain backpropagation.
Backpropagation is how you get the gradient of the loss with respect to every parameter in a network efficiently. The forward pass computes the output and the loss, caching intermediate activations. The backward pass then applies the chain rule from the loss backwards through each layer, reusing the gradient already computed downstream so each parameter's gradient costs roughly the same as the forward pass rather than being recomputed from scratch.
That reuse is the entire point. Computing each gradient independently would be prohibitively expensive for a network with millions of parameters; backpropagation makes the whole gradient roughly as cheap as one forward pass.
The consequences worth naming. Because gradients multiply through layers, repeatedly multiplying by small numbers makes them vanish and by large numbers makes them explode — which is the reason ReLU replaced sigmoid in hidden layers, why residual connections exist, and why gradient clipping is standard in sequence models. Also, the cached activations are why training memory scales with batch size and depth, and why gradient checkpointing trades compute for memory by recomputing them.
Q21 What are vanishing gradients, and what techniques address them?
In a deep network the gradient at an early layer is a product of many terms. If those terms are consistently less than one, the product shrinks toward zero as it propagates back, so early layers barely update and effectively stop learning. Sigmoid and tanh saturate — their derivatives approach zero for large inputs — which is what made deep networks nearly untrainable before the fixes arrived.
The techniques that address it: ReLU and its variants, whose derivative is exactly one for positive inputs so nothing shrinks. Careful initialisation, He or Xavier, so activations keep a sensible variance through the layers. Batch or layer normalisation, which keeps activations in a well-behaved range and smooths the loss surface. And residual connections, which give the gradient a direct path backwards that skips the layer entirely — that is the reason networks went from tens of layers to hundreds.
The mirror image. Exploding gradients are the same mechanism with terms above one, common in RNNs, and the standard fix is gradient clipping by norm.
Q22 Your model is not learning — training loss is flat. How do you debug it?
I start by proving the pipeline can learn at all: take a handful of samples and try to overfit them completely. If the model cannot drive the loss to near zero on ten examples, the problem is a bug, not a modelling choice, and I should stop tuning hyperparameters and go find it.
Then I work through the usual causes in order. The learning rate is the most common — too high and the loss is noisy or NaN, too low and it is flat, so I sweep it across orders of magnitude. Then the data: are the labels aligned with the inputs after shuffling, is normalisation applied, is the target what I think it is. Then the loss function and the final layer — using the wrong pairing, like softmax plus a loss that already applies it, silently produces nonsense. Then gradients: print their norms, and if they are zero or NaN, look for dead ReLUs, a detached tensor, or a missing zero_grad.
The framing that helps. Flat from the start is usually a bug or the learning rate. Learning and then plateauing is usually capacity or optimisation. Training fine but validation poor is not this question at all — that is overfitting or leakage.
Tier 3 — MLOps and Production Systems
This is the round most candidates are weakest on, and it is the round that decides whether you are hired as an engineer or a researcher.
Q23 How do you take a model from a notebook to production?
The notebook stops being the artefact. I move the code into a package with the preprocessing and the model as one serialised pipeline, so inference cannot possibly apply a different transformation than training did — that single decision prevents most production incidents.
Then: a training script that is reproducible from a pinned dataset version, a config and a seed, with parameters and metrics logged to something like MLflow. Tests, including a test that the pipeline produces the expected output shape on a fixed input, and a data validation step that fails loudly on schema or range violations rather than predicting on garbage. The model gets served behind an API — FastAPI in a container is the common shape — or as a batch job writing to a table, depending on how it is consumed. Then CI/CD, and monitoring on both the inputs and the outputs.
What I would emphasise. Version the data and the model together, not just the code. A model artefact you cannot trace back to the exact data and code that produced it is not reproducible, and you will need that trace on the day the predictions look wrong.
Q24 What is training-serving skew, and how do you prevent it?
It is when the features a model sees in production differ from the ones it was trained on, even though the model itself is unchanged. It is one of the most common causes of a model that validated beautifully and performs poorly live, and it is almost always an engineering problem rather than a modelling one.
The classic version is two implementations of the same feature: a pandas transformation in the training notebook and a hand-written version in the serving code, which round differently, handle nulls differently, or use a slightly different time window. Another version is temporal — training computed a feature over a full day of data that at inference time is only partially available.
Prevention. Share one code path: the preprocessing lives inside the serialised pipeline, or in a library imported by both training and serving. A feature store solves the harder version by computing features once and serving the same values to both offline training and online inference. And I would log the actual feature values sent at inference and periodically compare their distribution against training, because that comparison catches skew that nobody predicted.
Q25 What is model drift, and what do you do about it?
Drift is the world changing out from under a model that has not changed. Data drift is the input distribution shifting — a new customer segment, a new device type, a marketing campaign that changes who arrives. Concept drift is the relationship between inputs and target shifting, so the same inputs now imply a different outcome; fraud patterns adapting to your detection is the textbook case, and it is adversarial by nature.
I detect input drift by monitoring feature distributions against a training baseline, using a population stability index or a KS test per feature, with alerting on the ones that matter. Prediction drift — a shift in the distribution of scores — is a useful early warning that needs no labels. Actual performance needs ground truth, and the delay before labels arrive is the hard operational constraint: for fraud you may know in days, for credit default in months.
The response. Retrain on recent data, which is why I would have an automated retraining pipeline before I need one. Scheduled retraining is simpler and usually sufficient; triggered retraining on a drift alert is better when drift is irregular. And any retrained model goes through the same validation gate as the original, because automated retraining on corrupted data is a fast way to ship a worse model.
Q26 How do you make an ML project reproducible?
Four things have to be pinned, and code is only one of them. Code in git, obviously. Data versioned — DVC, or an immutable snapshot in object storage referenced by hash, because "the customers table" changes every day and a model trained on Tuesday cannot be reproduced on Friday without it. Environment pinned, which in practice means a lockfile and a container image, since a minor version bump in a library can shift results. And configuration plus random seeds captured with the run.
Then every training run logs its parameters, metrics, and the resulting artefact against those four, so any model in production can be traced back to exactly what produced it.
The honest caveat. Full bit-for-bit reproducibility is hard on GPUs, because some kernels are non-deterministic and floating point addition is not associative. You can force deterministic algorithms at a performance cost. What I actually aim for is that a rerun lands within noise of the original and the lineage is complete — which is what matters when someone asks in six months why the model made a particular decision.
Q27 Batch or real-time inference — how do you decide?
It comes down to when the prediction is needed relative to when the input exists. If the inputs are known ahead of time and the prediction can be precomputed, batch is simpler, cheaper and more robust — run nightly, write to a table or a key-value store, and the serving path is a lookup with no model in the request path at all. Churn scores, next-day recommendations and lead scoring all fit this.
Real-time is necessary when the input only exists at request time — a fraud decision on the transaction in front of you, a search ranking for a query just typed, a price for a specific basket. Then you are running a service with a latency budget, and everything follows from that budget.
The middle option worth naming. Streaming, or near-real-time, where features update continuously and predictions are refreshed within seconds. And a common hybrid is precomputing an expensive candidate set in batch, then ranking a small number of candidates in real time — that is roughly how most large recommender systems are actually built, and mentioning it shows you have thought past the binary.
Q28 How do you safely deploy a new model version?
Never straight to all traffic. The sequence I would use is shadow, then canary, then a proper experiment.
Shadow deployment runs the new model alongside the old on real traffic but discards its predictions, so I can compare outputs and latency under production conditions with zero user risk. That catches skew, unexpected inputs and performance problems before anyone is affected.
Then a canary — a small slice of real traffic, watching operational metrics closely, with an automated rollback if error rate or latency degrades.
Then an A/B test, because the question that actually matters is not whether the new model has better offline AUC. Offline metrics and business metrics diverge regularly — a recommender with better click prediction can reduce long-run retention. So I would randomise users, measure the business metric the model exists to move, and run it long enough for significance rather than stopping when it looks good.
The one detail that impresses. Randomise at the user level, not the request level, or the same user sees both models and the experiment measures nothing.
Q29 What do you monitor for a model in production?
Three layers, and I would say them in this order because most candidates only mention the third.
Operational: latency percentiles, throughput, error rate, memory, and the failure rate of the feature-fetching path. A model that returns nothing in time is broken regardless of its accuracy.
Data: schema conformance, null rates, range violations, and distribution drift per feature against the training baseline. This is where problems appear first and it needs no labels, so it is the highest-value monitoring you can have. An upstream team silently changing a unit from seconds to milliseconds is the kind of thing this catches and nothing else does.
Model: the distribution of predictions, the rate above the decision threshold, and once labels arrive, the actual performance metric — sliced by segment, not just in aggregate, because overall performance can hold steady while the model degrades badly for one group.
The last piece. Alerts have to be actionable and go to someone who owns the model. A drift dashboard nobody looks at is not monitoring.
Q30 Inference is too slow and too expensive. What do you do?
I would measure where the time actually goes first, because it is frequently not the model — feature fetching, network hops and serialisation often dominate, and no amount of model optimisation fixes that.
If it is the model, the options in roughly increasing order of effort. Batching requests, which improves GPU utilisation enormously at the cost of a little latency. Quantisation, moving weights from float32 to int8 or 4-bit, which typically gives a large speedup and memory reduction for a small accuracy cost. Exporting to a compiled runtime like ONNX Runtime or TensorRT, which fuses operations and often gives a solid speedup for no accuracy change. Then distillation — training a smaller model to imitate the large one — which is more work but can give an order of magnitude.
And caching, which is undervalued: if inputs repeat, cache the prediction. For an LLM service, KV caching and prompt caching change the economics substantially.
The framing. I would set a latency and cost budget with the product owner and stop when it is met, because each of these steps costs accuracy or engineering time and there is no reason to spend either past the target.
Tier 4 — LLM Systems and Behavioural
Increasingly the round that decides ML engineering offers. Answer these with the same engineering discipline as the rest, not with hype.
Q31 How does a RAG system work, and where does it break?
Retrieval-augmented generation grounds a language model in your own data. Offline, documents are chunked, embedded and stored in a vector index. At query time the question is embedded, the most similar chunks are retrieved, and they are placed into the prompt as context so the model answers from them rather than from parametric memory.
Where it breaks is almost always retrieval, not generation, and saying that is the answer they are listening for. Chunking that splits a table or a definition across a boundary means the right answer is never retrievable. Pure vector search misses exact identifiers, error codes and names, which is why hybrid search — dense embeddings plus BM25 keyword matching — outperforms either alone. Top-k retrieval returns k chunks whether or not any of them are relevant, so a reranker and a relevance threshold matter. And the model will still confidently answer from a bad context unless you instruct it to say when the context is insufficient.
How I would debug one. Evaluate retrieval separately from generation. Measure whether the correct chunk is in the retrieved set at all — if recall at k is poor, no prompt engineering will save the answer.
Q32 Prompt engineering, RAG, or fine-tuning — how do you choose?
By what is actually missing, and I would work up the ladder rather than starting at the top.
If the model has the knowledge and just needs direction, that is prompting — clear instructions, a few examples, structured output. Cheapest to build and to change, so it is always the first attempt.
If the model lacks the knowledge, that is RAG. Facts that are private, that change, or that are too numerous to fit in a prompt belong in retrieval, because updating a document is far cheaper than retraining and you get citations for free.
If the model knows the facts but cannot produce the form — a specific output format, a domain tone, a specialised task where prompting plateaus — that is fine-tuning. It teaches behaviour, not knowledge, and expecting it to inject facts is the most common misconception here.
The practical answer. They compose: a fine-tuned smaller model with RAG is often cheaper and better than a large model prompted. But I would start with prompting plus RAG, measure, and only fine-tune when I have the evaluation set to prove it helped — which you need anyway.
Q33 How do you evaluate an LLM application?
The hard part is that there is no single correct output, so accuracy does not apply. What I would build first is an evaluation set — a few hundred real queries with expected behaviour, drawn from actual usage rather than invented — because without it every change is a guess and you cannot tell an improvement from a regression.
Then layered evaluation. Deterministic checks where possible: does the output parse as valid JSON, does it match the schema, does it contain a required citation, is it within length. These are cheap, run in CI, and catch most regressions. Then reference-based scoring where a ground truth exists. Then an LLM-as-judge for the subjective dimensions — faithfulness to the retrieved context, relevance, tone — which works but needs its own validation against human ratings, because a judge with an unexamined rubric drifts.
For RAG specifically I would evaluate retrieval and generation separately: recall at k for retrieval, and groundedness for generation, meaning whether every claim is supported by the context.
Production. Offline evaluation is a gate, not the truth. I would log real interactions, collect thumbs-up and thumbs-down, and sample outputs for human review, because the failures that matter are the ones nobody anticipated.
Q34 Tell me about an ML project you owned end to end, and what you would do differently.
Structure it as problem, approach, result, and reflection — and give the result in business terms, not just a metric, because a model with a 0.89 F1 means nothing to an interviewer without knowing what it changed.
A strong shape: the problem was framed by the business as one thing and I reframed it, which is a good signal in itself — for instance, "predict churn" becoming "rank the accounts a small retention team should call this week", which changes the metric from AUC to precision in the top hundred. Then the honest engineering detail: what the data actually looked like, the leakage I found and removed, the baseline I compared against. Then the outcome, including deployment, because a model that never shipped is a much weaker story.
The reflection is what they are actually asking for. Pick something real. Good answers: I built the model before validating that the labels meant what I assumed, and lost two weeks. I optimised offline metrics for too long before putting anything in front of a user. I did not set up monitoring at launch, so we found out about drift from a stakeholder rather than an alert. Each of those shows you learned an engineering lesson, which is more valuable to them than the model was.
A note on delivery
The content above gets you through the technical rounds. What separates candidates at that point is rarely knowledge:
- Answer the question, then stop. ML questions invite tangents more than most, and a rambling answer buries the correct one inside it.
- Say "I do not know" cleanly, then say how you would find out. They are probing for the edge of your knowledge. Bluffing a derivation is the single fastest way to lose a technical round.
- Bring one real number or one real failure per answer. A leakage bug you actually found, a latency figure you measured, a model that underperformed in production. That is what separates someone who has deployed from someone who has completed a course.
- Ask what the model is for. For any open-ended modelling question, clarifying the business objective, the cost of each error type and the latency constraint before proposing an architecture is itself the strongest signal you can send.