Loading...
Loading...
The most common machine learning interview mistakes in 2026 are using accuracy as the primary metric for an imbalanced classification problem without checking what a naive baseline achieves, responding to model underperformance with hyperparameter tuning before running a single diagnostic, applying random train-test splits to data with temporal or group structure, treating offline validation performance as equivalent to production performance, describing model deployment as the endpoint of an ML pipeline with no mention of monitoring or retraining, designing a RAG system without specifying a chunking strategy or retrieval evaluation methodology, and presenting every past ML project as a clean success with no unexpected results or model failures. Every mistake below has a specific, rehearsable fix built around the exact mechanism that causes the failure, not a general suggestion to study harder.
Machine learning interviews have a specific failure pattern that does not appear in other engineering interviews. The candidate gives a technically correct surface answer and the interviewer follows up with a single mechanism-level question. "You said you would tune hyperparameters. What diagnostic would you run first to decide which hyperparameters to tune and in which direction?" The candidate who has only trained models in Jupyter notebooks using default settings and measured the result freezes. The candidate who has debugged real training failures in production systems answers immediately, because they have lived the process of moving from a symptom to a diagnosis to a targeted fix.
This guide covers thirteen ML-specific interview mistakes at the mechanism level. Each mistake is named precisely, its real cost explained, and the fix is specific enough to rehearse before your next interview. The fifth category covers the mistakes that emerged from AI coding tools becoming part of the ML engineering workflow, introducing a category of subtle bugs that interviewers are actively screening for and that candidates who use these tools without review discipline consistently miss.

Real Interviews. Real Pressure. Practice until it feels easy.
ML interviews test two things that most candidates only prepare for one of: whether you can train a model, and whether you understand why it behaves the way it does. The second half is what gets tested in follow-up questions, and it is the half that predicts whether a candidate's models will perform as expected in production or will degrade unexpectedly.
ML systems fail in ways that are not visible during development. A model that achieved 94% validation accuracy and shipped to production can silently degrade to 71% effective performance over three months as the input data distribution shifts away from what the model saw during training. A recommendation system that tested beautifully offline can underperform in production because the test set used row-level random splitting that allowed the model to effectively memorize user behavior from training rather than generalize. These failure modes are completely invisible to a candidate who has only trained models on clean datasets and stopped at the validation metric.
Every mistake below is a reflection of this same root problem: knowing how to run the code without understanding what the code is actually doing and how it will behave when the assumptions it was built on are violated.
What This Looks Like
A candidate is given a binary classification problem during a case study round. After building and evaluating a model, they report "I achieved 97% accuracy, which is strong performance." The interviewer asks what accuracy a model that predicts the majority class for every single input would achieve. The candidate has not checked this and does not know the class distribution, so they cannot answer. The distribution turns out to be 3% positive class and 97% negative class, meaning the 97% accuracy achieved by the trained model is identical to what a completely useless naive classifier achieves.
Why Recruiters Flag This
Reporting accuracy without checking the naive baseline is not a minor oversight. It reveals that the candidate does not have a systematic mental model of what a metric actually measures and what it cannot tell you. Any metric that can be matched by always predicting the majority class is telling you nothing about your model's ability to distinguish between classes. An experienced ML interviewer treats this mistake as a significant signal about how a candidate would evaluate models in production, where the same confusion would produce dangerous false confidence.
The Real Cost
A fraud detection model reported as 99.2% accurate and shipped to production might be catching zero fraudulent transactions. A medical diagnosis classifier reported as 95% accurate might be completely ignoring the positive class because the prevalence of the condition is only 5%. Both scenarios have shipped in real organizations, and both resulted from accepting an accuracy report without checking what a trivial baseline would achieve.
The Complete Fix
Build a mandatory two-step pre-evaluation habit for every classification problem: before evaluating any model, check the class distribution and calculate the naive majority-class accuracy. Then choose a primary evaluation metric that is meaningful even under that distribution. For imbalanced problems where the positive class is rare and important, precision, recall, F1 score on the positive class, and AUC-PR are almost always more informative than accuracy. The choice between them is determined by the relative cost of false positives versus false negatives for the specific use case, which is a business question that must be answered before any model is evaluated.
In an interview, state this reasoning explicitly before proposing any metric: "Before I choose an evaluation metric, I want to check the class distribution, because if the dataset is imbalanced, accuracy will be misleading. Given a 3% positive rate, I would focus on precision and recall for the positive class and use F1 or AUC-PR as my primary metric, and I would report what accuracy a majority-class baseline achieves as a sanity check for any model I build."
Practice Method
Find any Kaggle binary classification dataset with a significant class imbalance, train a gradient boosting model with default settings, and write out five different metrics for the result including accuracy, precision, recall, F1, and AUC-PR. Then calculate what each metric would be for a classifier that always predicts the majority class. Observe which metrics reveal the model's actual discrimination ability and which ones are contaminated by the class distribution. Do this exercise on three different imbalance ratios: 70/30, 90/10, and 97/3. The progression builds the instinct for how severely class imbalance distorts different metrics.
What This Looks Like
An interviewer presents a scenario: a model achieves good training performance but poor validation performance. What do you do? The candidate says they would try adjusting the learning rate, increasing the model capacity, or trying different architectures. The interviewer asks what diagnostic they would run before making any of those changes. The candidate describes more hyperparameter experiments rather than a diagnostic that would tell them which direction to move in.
Why Recruiters Flag This
Random hyperparameter search without a preceding diagnostic is the ML equivalent of prescribing medication before examining the patient. High training performance and poor validation performance is a specific pattern with a specific name, high variance or overfitting, and it has a specific set of causes and interventions that follow a logical sequence. A candidate who responds to this pattern by trying different hyperparameters without running a diagnostic first will spend weeks making changes that may move in the wrong direction, and more importantly, will do the same thing in production when a model fails under real conditions.
The Real Cost
Unstructured hyperparameter experiments on a production model failure cost engineering time and delay resolution. Worse, they can accidentally improve a metric through a change that makes the underlying problem worse, creating the illusion of a fix before the real problem resurfaces.
The Complete Fix
Internalize a four-step diagnostic sequence and run through it out loud in any interview that presents a model performance problem. The first step is to plot the learning curves: training loss and validation loss over training iterations, which reveals whether the model has high bias, meaning both losses are high, high variance, meaning training loss is low and validation loss is high and the gap is large, or is in a good regime but simply needs more training. The second step is to examine the gap between training and validation performance as a function of training set size, which reveals whether adding more data would help close the gap. The third step is to inspect the error distribution by examining which examples the model most commonly gets wrong, which often reveals a data quality problem, a labeling inconsistency, or a subpopulation that is systematically underrepresented in training. Only after these three diagnostics produce a clear hypothesis does a specific intervention become justified.
Stating this sequence out loud before proposing any solution is the interview behavior that signals diagnostic thinking: "Before I suggest any hyperparameter changes, I would look at the learning curves. High training performance and poor validation performance with a large gap suggests high variance, which has a specific set of interventions. I would check whether the gap closes with more training data, because if it does, the issue is data volume rather than model capacity. If the gap does not close, I would increase regularization before increasing model complexity, because increasing complexity when you already have high variance typically makes the problem worse."
Practice Method
Deliberately induce four different model failure modes in any training pipeline: underfitting by using a too-small model or too little training, overfitting by training too long without regularization, training-validation distribution mismatch by creating a biased data split, and data leakage by including a target-correlated feature. For each failure mode, plot the learning curves and practice identifying the failure mode from the curve shape alone. Then practice naming the correct intervention for each pattern before seeing any validation accuracy number.

What This Looks Like
A candidate lists gradient boosting, random forests, and XGBoost on their resume. The interviewer asks: "You mentioned gradient boosting. Can you explain the specific insight that makes gradient boosting different from a standard boosting approach, and why it works?" The candidate says "it combines weak learners into a strong learner and it is very accurate for tabular data." This describes ensemble methods in general, not the specific mechanism of gradient boosting. When pushed further, the candidate cannot explain what a gradient in gradient boosting refers to or why fitting residuals enables the ensemble to improve with each added tree.
Why Recruiters Flag This
Algorithm name recognition without mechanism understanding is the most reliable signal available in an ML interview that a candidate's experience is primarily tutorial-based rather than production-based. On the job, a developer who does not understand why gradient boosting works cannot debug why a gradient boosting model is underperforming, cannot choose between gradient boosting and an alternative for a specific dataset regime, and cannot explain to a non-technical stakeholder why the model behaves the way it does. This is not a theoretical concern; it predicts specific production failures.
The Real Cost
A developer who does not understand gradient boosting's mechanism cannot identify that a gradient boosting model is underfitting because the learning rate is too high and the trees are not correcting residuals gradually enough, versus overfitting because the trees are too deep and the model has memorized training examples rather than learned patterns.
The Complete Fix
For every algorithm on your resume, practice a two-part explanation: the specific insight that makes the algorithm work, and a concrete failure mode that follows from that mechanism. For gradient boosting specifically: the insight is that each tree is trained to predict the residuals of all previous trees combined, meaning each subsequent tree corrects the errors of the ensemble so far, and the gradient in the name refers to the fact that fitting residuals of a mean squared error loss is equivalent to following the negative gradient of that loss function. The failure mode that follows from this mechanism is that if the learning rate is too high, each tree makes too large a correction and the ensemble oscillates rather than converging.
Apply this two-part structure to every algorithm you claim expertise in, and practice delivering it in under sixty seconds for each one. The speed matters because interviewers form their impression of depth in the first thirty seconds of the explanation.
Practice Method
Write a one-paragraph explanation of the mechanism behind every algorithm on your resume: not what the algorithm achieves but specifically how it achieves it at a mathematical or computational level. Then find someone technical, share your explanation, and ask them to identify the first sentence where the explanation becomes too vague to derive any implication from. That sentence marks the boundary of your current mechanistic understanding.

What This Looks Like
A candidate describes a project where they built a churn prediction model for a subscription product, trained on twelve months of customer behavior data. When asked how they split the data for evaluation, they say they used an 80/20 random split. The interviewer asks whether a given customer's data appears in both the training and test sets. The candidate confirms that it does, because the split was row-level. The interviewer asks what problem this creates. The candidate either cannot articulate the problem or gives a vague answer about data leakage without naming the specific mechanism by which the random split inflates the model's apparent performance.
Why Recruiters Flag This
Row-level random splitting on data where multiple rows come from the same entity, whether a customer, a patient, a session, or a document, allows the model to implicitly learn the entity's characteristics during training and be evaluated on the same entity's other records during testing. This is not test set contamination in the traditional sense; it is structural leakage that makes the model appear to generalize across entities when it is actually generalizing across time points within the same entity. The offline performance will be significantly higher than the production performance against unseen entities.
The Real Cost
A churn prediction model trained and evaluated with row-level random splitting might show 87% AUC-ROC offline and 71% AUC-ROC when evaluated on customers who joined after the training period ended. The 16-point gap is not random; it is a systematic consequence of the evaluation methodology. Acting on the 87% estimate wastes engineering resources and creates missed expectations for business stakeholders.
The Complete Fix
Before writing any code for a new ML project, identify the unit of independence: the entity such that two records from the same entity are not independently drawn from the same distribution. For customer data, the entity is the customer. For time-series data, the entity is time, meaning records before a cutoff date are not independent of records after it. For geographic data, the entity might be a region where local effects create correlations within regions.
Once the unit of independence is identified, the correct splitting strategy follows directly: group-based splitting that places all records from a given entity entirely in either training or test, never both; temporal splitting that uses all records before a cutoff date for training and all records after for testing; or a combination for entity-level time-series where all records for the same entity in the training time window are in training and all records for that entity in the test time window are in test.
State this reasoning explicitly in any interview that involves an ML project description: "Before deciding on a split strategy, I need to identify the unit of independence in this dataset. For a churn prediction model where I have multiple rows per customer, using row-level random splitting would allow the model to effectively memorize each customer during training and be evaluated on the same customers in the test set. The correct approach is a customer-level split where all of a given customer's records go to either training or test, combined with a temporal cutoff so the test set reflects future behavior rather than simultaneous behavior."
Practice Method
Take any public dataset that has a natural grouping structure, whether by user, by date, or by geographic region, and implement all three splitting approaches: random row-level, group-level, and temporal. Train the same model using each split strategy and record the offline evaluation metric. Then verify which metric best predicts performance by holding out a genuinely unseen cohort. The discrepancy between the random split metric and the group or temporal split metric is the quantified cost of the leakage mistake.
What This Looks Like
A candidate describes an ML project and presents the validation metric as evidence that the model is ready for deployment, or describes a past deployment where the validation metric was considered the primary signal of production readiness. When asked whether the validation metric predicted production performance, the candidate says yes. When asked what conditions would need to hold for that to be true, or what specific risks could cause the offline metric to overestimate production performance, the candidate cannot name any specific mechanisms.
Why Recruiters Flag This
The offline-to-online gap is one of the most documented and consequential phenomena in production ML. It arises from several distinct mechanisms, each of which a mid-level ML engineer should be able to name and discuss. Training-serving skew occurs when the feature preprocessing pipeline used during training produces slightly different outputs than the serving pipeline, either due to implementation differences or due to subtle differences in how real-time data is handled versus historical data. Distribution shift occurs when the incoming data in production has a different statistical distribution than the training data, which happens continuously as user behavior evolves, as product changes are made, and as external conditions change. Feedback loop effects occur in recommendation and ranking systems where deploying a model changes the data it collects, which then changes what the next training round learns.
The Real Cost
The standard pattern is a model that achieves a given offline metric, gets deployed, and underperforms relative to expectations. The investigation takes weeks because the team assumed the offline metric was reliable and did not build the monitoring infrastructure that would have revealed the offline-to-online discrepancy immediately.
The Complete Fix
Prepare a three-part answer for any production ML discussion that addresses the offline-to-online gap explicitly. First, name the specific mechanisms that could cause offline performance to overestimate production performance for this type of model and this type of data. Second, describe the pre-deployment check that would catch the most likely mechanism: for training-serving skew, a shadow deployment where the same input is processed through both the training and serving pipelines and the outputs are compared; for distribution shift, a comparison of the feature distribution in recent production data against the training data distribution. Third, describe the post-deployment monitoring that would detect the gap early: a business metric that reflects the model's actual purpose alongside a technical metric that reflects its predictive performance.
Practice Method
For any past ML project you have worked on or studied, write a one-page "what could go wrong post-deployment" analysis. For each risk, name the specific mechanism, the magnitude of impact you would expect, and the monitoring signal that would detect it within one week of deployment. This exercise builds the production ML thinking that interviewers are testing and that most candidates who have not deployed models to production have never explicitly practiced.
What This Looks Like
A candidate claims production ML deployment experience. When asked to describe the deployment, they explain that they trained the model, saved it as a pickle file, and the engineering team loads it in the application. When the interviewer asks how the model is served at inference time, what latency it achieves, how model updates are pushed without application downtime, and how the model's predictions are monitored in production, the candidate cannot answer any of these questions because the "deployment" in their description was artifact creation rather than production serving.
Why Recruiters Flag This
The word "deployment" in ML has a specific technical meaning that many candidates use loosely to mean "trained a model and handed off the artifact." Production model deployment involves a serving infrastructure that handles inference requests at the required latency, a model registry that tracks which model version is serving traffic, a mechanism for updating the serving model without downtime, and a monitoring layer that tracks prediction behavior in production. None of these components are present in "saved a pickle file and handed it to engineering," which means the candidate claiming deployment experience is describing something that does not address the hardest and most consequential parts of the ML engineering job.
The Real Cost
A candidate hired based on deployment experience they do not have will need significant ramp-up to contribute to the production ML lifecycle, and will be missing the instincts that production experience provides for anticipating deployment failures before they occur.
The Complete Fix
Be precise about what you have and have not done in production ML. If your experience is with model training and evaluation but not with production serving infrastructure, describe it accurately and add that you understand the components of a production serving setup even if you have not built one yourself. If you have built serving infrastructure, describe it specifically: the framework used for serving, the latency achieved, how model updates are pushed, and what monitoring is in place.
The distinction between components to learn is important: model training and model serving are separate engineering concerns. A production serving infrastructure includes at minimum an endpoint that accepts requests and returns predictions at the required latency, a mechanism for loading new model versions without restarting the service, health checking for the serving endpoint, and logging of inputs and predictions that enables downstream monitoring.
Practice Method
Deploy any trained model to a REST API using FastAPI, add a /predict endpoint that returns predictions, add health check endpoints, add request and response logging, and containerize it with Docker. Deploy it to any free cloud tier. This takes a day and converts "I have trained models" into "I have deployed a model to an API with logging," which is a significantly more credible production ML credential.

What This Looks Like
An interviewer asks a candidate to describe their ML pipeline from data to production. The candidate walks through data collection, preprocessing, feature engineering, model training, evaluation, and deployment. They stop there. When the interviewer asks "what happens next," the candidate is caught off guard. When specifically asked about model monitoring, drift detection, and retraining, the candidate either has never thought about these components or describes them as someone else's responsibility.
Why Recruiters Flag This
The deployment of a model is not the completion of an ML pipeline; it is the beginning of a lifecycle that is more demanding and more consequential than training. A deployed model begins its life representing the data distribution at training time, and that distribution changes continuously. The monitoring, drift detection, and retraining components of the pipeline are what determine whether the model remains useful over time or silently degrades. An ML engineer who considers deployment the endpoint has a fundamental misunderstanding of what production ML requires.
The Real Cost
A model with no monitoring and no retraining strategy will degrade. The question is not whether it degrades but how quickly and whether anyone notices before the degradation becomes a business problem. In recommendation systems, degradation typically becomes visible in three to six months. In fraud detection, it can happen in weeks as fraudsters adapt their behavior to avoid the current model's patterns.
The Complete Fix
Treat the ML pipeline description as having three distinct phases: the training lifecycle, the deployment phase, and the production lifecycle. When asked to describe an ML pipeline in any interview, always include all three phases and name specific components of each. The production lifecycle includes at minimum: input feature monitoring that detects when the distribution of incoming features shifts significantly from the training distribution using a statistical test such as the Kolmogorov-Smirnov test for continuous features; prediction monitoring that tracks the distribution of model outputs and alerts when it shifts significantly; business metric monitoring that tracks the downstream outcome the model is intended to improve; a retraining trigger that fires when any of these signals crosses a threshold; and a retraining pipeline that can produce and evaluate a new model candidate without manual intervention.
Practice describing this three-phase pipeline in under three minutes, because the ability to describe a complete ML lifecycle concisely demonstrates that the full lifecycle is how you think about ML rather than something you can enumerate under prompting.
Practice Method
Design the complete monitoring infrastructure for any model you have previously trained, as if you were now responsible for that model in production. Write the specific statistical tests you would use to detect drift in each feature type, the threshold that would trigger a retraining, and the evaluation criteria a new model candidate would need to meet before replacing the current serving model. The specificity required by this exercise reveals any gaps in your production ML mental model.
What This Looks Like
A candidate describes production ML work and mentions that they monitored the model in production. When the interviewer asks how they detected drift, the candidate says they tracked model accuracy over time. When the interviewer asks how they measured accuracy without real-time labels, the candidate pauses. Most production models do not have immediate ground truth labels available at inference time: a recommendation model does not know until days later whether a recommended item was purchased, a loan approval model does not know for months whether a loan was repaid. The monitoring approach that depends on real-time labels is infeasible in most production settings.
Why Recruiters Flag This
Drift detection without real-time labels is one of the most practically important and most commonly misunderstood aspects of production ML. Input drift detection, which compares the distribution of incoming features against the training distribution using statistical tests that require no labels, is the primary mechanism for early warning of model degradation. Candidates who have not implemented label-free drift detection do not have the production ML experience they claim, even if they have deployed models.
The Real Cost
A model monitored only by downstream business metrics will not show degradation signals until the degradation is severe enough to affect business performance. By that point, the retraining, evaluation, and redeployment cycle may take days or weeks, during which the degraded model continues to serve production traffic.
The Complete Fix
Memorize a four-signal monitoring framework and be able to describe the specific implementation of each signal. The first signal is input feature drift: for each important feature in the model, compute the Kolmogorov-Smirnov statistic between the current week's distribution and the reference distribution from training, and alert when the p-value drops below a threshold. For categorical features, use a chi-squared test or the Population Stability Index. The second signal is prediction drift: track the distribution of model output scores and alert when it shifts significantly, because prediction drift can reveal behavior change even when input distributions appear stable. The third signal is data quality monitoring: track the rate of missing values, out-of-range values, and unexpected categorical values in incoming features, since data quality problems are often the root cause of apparent drift. The fourth signal is delayed label monitoring: for problems where labels eventually become available, compare the model's predictions against the delayed labels on a regular cadence to estimate actual accuracy without waiting for enough labels to compute a statistically meaningful real-time estimate.
Practice Method
Add drift monitoring to any model you have already deployed or built locally. Implement KS-test-based monitoring for at least two continuous features, chi-squared monitoring for at least one categorical feature, and prediction distribution monitoring. Then deliberately shift the input distribution by modifying a few rows in a test input file and verify that your monitoring system detects the shift. This hands-on implementation is the experience that interview questions on production monitoring are designed to surface.
Real Conversations. Real Scenarios. Speak until it feels natural.

What This Looks Like
A candidate is asked to design a question-answering system over a large document collection. They describe a retrieval-augmented generation approach: split the documents into chunks, embed them using an embedding model, store them in a vector database, retrieve the top k most similar chunks for each query, and pass them to an LLM. When the interviewer asks what chunking strategy they would use and why, the candidate says "fixed-size chunks of about 500 tokens." When asked what they would do to evaluate whether the retrieval step is working correctly, the candidate has no systematic answer.
Why Recruiters Flag This
Chunking strategy is one of the most consequential decisions in a RAG pipeline, and "fixed-size chunks" is the lowest-effort approach with the most failure modes. Fixed-size chunking splits documents at arbitrary token boundaries that frequently fall within a sentence, separating context from its continuation and creating chunks whose embeddings do not meaningfully represent a complete thought. The retrieval evaluation gap is equally concerning: a RAG system where the retrieval is poor will produce responses that are confidently wrong regardless of the LLM's quality, and a candidate who has not designed retrieval evaluation cannot distinguish between a retrieval problem and an LLM generation problem when the system fails.
The Real Cost
A RAG system with poor chunking and no retrieval evaluation will underperform consistently on multi-sentence queries and queries that require context from adjacent document sections. The team will spend time adjusting prompts and changing LLMs looking for improvements, without realizing that the retrieval step is the actual bottleneck because no evaluation was designed to detect it.
The Complete Fix
Prepare a specific answer for chunking strategy that names the tradeoffs of at least two approaches. Semantic chunking, which splits at natural semantic boundaries like paragraph ends and section breaks, preserves the coherence of complete thoughts and produces higher-quality embeddings at the cost of variable chunk sizes that require more careful vector index configuration. Recursive character text splitting with overlap, which is the approach used in most production RAG implementations, splits at a hierarchy of separators starting with double newlines and falling back to single newlines and then sentence boundaries, with an overlap of fifty to one hundred tokens between adjacent chunks to prevent context loss at boundaries. The overlap is critical and often omitted: it ensures that a sentence split across a chunk boundary can still be retrieved when either half is queried.
For retrieval evaluation, prepare a specific methodology: create a held-out evaluation set of question-answer pairs where the correct answer exists in the document collection and the specific passage is known. Measure recall at k, meaning what fraction of questions have the correct passage in the top k retrieved chunks, separately from generation quality, meaning whether the LLM produces the correct answer given the retrieved chunks. This separation identifies whether a failure is a retrieval problem or a generation problem.
Practice Method
Build a RAG system for any document collection and implement both chunking strategies: fixed-size and recursive with overlap. Create twenty evaluation questions where you know the exact document passage that contains the correct answer. Measure recall at five for each chunking strategy. The recall difference between the two strategies makes the chunking decision concrete and defensible rather than a matter of convention.
What This Looks Like
An interviewer asks a candidate to choose between RAG and fine-tuning for a customer service chatbot that needs to answer questions about a company's product documentation. The candidate says "it depends on the use case." The interviewer says "yes, it does depend. What does it depend on specifically?" The candidate gives a vague answer about the amount of data available or the cost of fine-tuning without being able to name the three or four specific factors that should drive the decision and apply them to this scenario.
Why Recruiters Flag This
"It depends" without named decision criteria is a non-answer that reveals the candidate has not built a mental model for this decision. LLM engineering requires systematic judgment about build and design decisions, and the RAG versus fine-tuning decision is one of the most common design choices in production LLM systems. A candidate who cannot name the decision criteria cannot make this decision confidently in a team environment, which means the decision gets made by default rather than by design.
The Real Cost
Teams that choose the wrong approach for a specific use case either build a fine-tuned model that becomes outdated every time the product documentation changes, requiring retraining at significant cost, or build a RAG system for a task that actually requires behavior change rather than knowledge injection, producing poor response quality that no amount of retrieval improvement can fix.
The Complete Fix
Memorize four specific decision factors and apply each to the scenario given. The first factor is knowledge update frequency: if the information the system needs to convey changes frequently, such as product documentation, pricing, policies, or news, RAG is almost always the right choice because updating an embedding store is far cheaper and faster than retraining a model. If the knowledge is stable over time, fine-tuning becomes more attractive. The second factor is knowledge source type: if the knowledge exists in discrete documents that can be retrieved, RAG works well. If the knowledge is diffuse and cannot be easily segmented into retrievable chunks, such as a general reasoning style or a specific communication tone, fine-tuning is more appropriate. The third factor is behavior versus knowledge: RAG injects knowledge into context at inference time but does not change the model's underlying behavior. Fine-tuning changes how the model generates responses, making it appropriate for tasks where you need the model to respond in a particular style, format, or reasoning pattern regardless of the input. The fourth factor is cost and latency: RAG adds retrieval latency and infrastructure cost; fine-tuning adds training compute cost and may reduce inference cost by eliminating retrieval.
For the customer service chatbot over product documentation, the correct answer is RAG for the knowledge retrieval component and potentially a lightweight fine-tune for the response style component, with a clear explanation of why each decision follows from the specific factors.
Practice Method
Take five different LLM application scenarios and apply all four decision factors to each one, writing out the conclusion for each factor and the overall recommendation with justification. Then discuss your analysis with someone who disagrees with one of your conclusions and practice defending your reasoning under challenge, since the follow-up question in this interview topic is almost always an attempt to change a specific variable and see whether the candidate updates their recommendation accordingly.
What This Looks Like
A candidate builds a content moderation classifier using an LLM to classify user-generated content as safe or unsafe. They report that the classifier achieves 93% accuracy on a held-out test set and recommend deployment. The interviewer asks what the false negative rate is, meaning the fraction of genuinely unsafe content that the classifier labels as safe. The candidate has not measured this separately. When measured, it is 22%, meaning one in five pieces of unsafe content passes through the classifier undetected. The candidate did not evaluate the metric that actually matters for the safety use case.
Why Recruiters Flag This
The cost of a false positive and the cost of a false negative are almost never equal in real ML applications, and the choice of evaluation metric should reflect the cost asymmetry before any model is built. In content moderation, a false negative, allowing unsafe content through, is far more costly than a false positive, blocking safe content. In medical diagnosis, a false negative, missing a disease, is far more costly than a false positive, ordering a follow-up test. In fraud detection, the cost depends on the specific economics of the fraud type and the friction of the false positive experience for legitimate users. A candidate who defaults to accuracy without examining the cost function is producing an evaluation that may be completely uninformative for the actual decision at hand.
The Real Cost
A content moderation system with a 22% false negative rate will allow significant volumes of harmful content through, creating real platform safety failures. The 93% accuracy figure that was used to justify deployment masked this completely because the dataset was dominated by safe content, making accurate safe-content classification dominate the accuracy calculation.
The Complete Fix
Before defining any evaluation metric for an LLM feature, explicitly define the cost function: what is the consequence of a false positive, and what is the consequence of a false negative, and what is their relative magnitude. For content moderation: the false positive cost is user friction from blocked content, which is recoverable through an appeal process. The false negative cost is harm to users exposed to unsafe content, which may not be recoverable and creates regulatory and reputational consequences. Given this asymmetry, the evaluation should optimize for recall on the unsafe class, meaning the fraction of truly unsafe content that is correctly identified, with precision on the unsafe class as a secondary constraint that prevents the recall from being achieved by flagging everything.
State this cost function analysis explicitly before proposing any metric in an LLM evaluation interview: "Before I choose a metric, I need to think about what it costs to make each type of error. For a content moderation use case, a missed harmful content item has a much higher cost than a false alarm. That means I should optimize for recall on the harmful class, with precision as a secondary constraint, and report both separately rather than combining them into accuracy."
Practice Method
Take three different LLM application scenarios with different error cost asymmetries, write out the cost function explicitly for each, derive the primary evaluation metric that follows from the cost function, and calculate what the chosen metric would be for an LLM classifier that always predicts the majority class. This exercise builds the cost-function-first reasoning habit that experienced ML interviewers are testing.

What This Looks Like
A candidate uses an AI coding tool to generate a PyTorch training loop for a take-home assignment. The loop looks complete: it iterates over batches, computes a forward pass, calculates loss, calls backward, and updates the optimizer. The candidate submits it. The interviewer runs the code and training loss becomes NaN after three epochs. When asked to debug it, the candidate has not run the code themselves and cannot identify the issue quickly. The root cause is that the AI generated a training loop that computes the forward pass before moving the input tensors to the same device as the model, causing a RuntimeError in some configurations that the AI's default example silently avoided.
Why Recruiters Flag This
AI tools generate syntactically plausible ML code that frequently contains subtle bugs specific to the ML framework's semantics: device mismatches between tensors and models, validation loops that forget to disable gradient computation with torch.no_grad(), training augmentation that is accidentally also applied to the test set because both use the same transform object, and loss accumulation that builds a growing computation graph across batches because the loss value is retained rather than being detached. These are bugs that look correct in isolation, pass visual inspection, and only reveal themselves when the code runs under specific conditions. A candidate who does not catch them has revealed that their code review process ends at "it looks right."
The Real Cost
A take-home assignment with a NaN loss that the candidate cannot debug on the spot reveals to the interviewer that the candidate either did not run the code before submitting, ran it but did not notice the NaN, or noticed the NaN but could not diagnose it. Any of these outcomes is significantly more damaging to the evaluation than a correct but slower training loop.
The Complete Fix
Build a five-item ML-specific code review checklist that you run on every AI-generated training or evaluation code before submitting or merging it.
The first item is device placement: confirm that both the model and all input tensors are moved to the same device before any forward pass computation. A common AI tool error is generating model.to(device) but forgetting to move batch tensors inside the training loop.
The second item is gradient context correctness: confirm that torch.no_grad() is used consistently during validation and inference, and that training loops call optimizer.zero_grad() before every backward pass, not after, and not occasionally.
The third item is data augmentation scope: confirm that any transformation pipeline that includes random operations is applied only to the training data, and that evaluation data uses only deterministic preprocessing. A common AI error is sharing a single transform object between training and validation DataLoader calls.
The fourth item is loss detachment in logging: confirm that loss values accumulated for logging are detached from the computation graph using .item() rather than being stored as tensors, since storing loss tensors across iterations builds a computation graph that grows with training time and eventually causes OOM errors.
The fifth item is epoch versus batch metric accumulation: confirm that epoch-level metrics are correctly accumulated across batches and averaged rather than being overwritten on each batch or summed without normalization.
Narrating this checklist out loud during a live round that involves code review transforms AI tool use into a visible demonstration of ML engineering discipline.
Practice Method
Ask an AI coding tool to generate a complete PyTorch training loop for a simple image classification task. Before running it, apply the five-item checklist and write down every issue found. Then run the code and compare the runtime errors to the issues you found. Any runtime error you did not catch in the checklist review reveals a gap in your review process. Repeat this exercise until the checklist catches every class of error the code is likely to contain.
What This Looks Like
A candidate describes three ML projects during a behavioral round. In every project, the model achieved the intended performance improvement, the stakeholders were satisfied, and the project shipped on time. There are no unexpected results, no models that underperformed during testing, no data quality surprises discovered after training began, and no post-deployment behavior changes that required investigation. When the interviewer asks about a project where something did not go as planned, the candidate cannot identify one.
Why Recruiters Flag This
Real ML projects fail regularly in specific and instructive ways. Training data turns out to have mislabeled examples that inflate training accuracy. A feature that seemed predictive is later discovered to be a proxy for a label leakage artifact. A model that performed well on the evaluation dataset degrades in production because the evaluation distribution did not represent the deployment distribution. A retraining run produces a significantly worse model than the currently serving version for no immediately obvious reason. Every experienced ML engineer has lived through several of these scenarios, and the ability to describe them specifically, name the root cause, and describe what changed as a result is strong evidence of genuine production experience. A candidate with only clean successes has either worked on very simple problems or has not reflected honestly on projects that had difficulties.
The Real Cost
A candidate who cannot describe a project failure is evaluated as either lacking the production experience that produces such failures, or lacking the intellectual honesty to acknowledge them. Both are concerns that experienced interviewers will note.
The Complete Fix
Prepare one specific ML project failure story in advance using a structure that demonstrates both technical understanding and professional maturity. The story should name a specific unexpected result rather than a vague difficulty, explain the specific mechanism by which the unexpected result occurred, describe the investigation process that identified the root cause, and state what was changed in your process or your evaluation methodology as a result of the experience.
A strong example structure: "I built a churn prediction model that achieved excellent offline AUC but showed no improvement in production A/B test results. I spent a week investigating and discovered that the evaluation split was using a random row-level split on data where each customer had multiple rows. The model was effectively learning customer identity from training rows and predicting correctly on the same customers' test rows, rather than generalizing to new customers. Switching to a customer-level split dropped the offline AUC significantly but the production A/B test then showed a real improvement that matched the corrected offline estimate. I now specify the independence unit before writing any data splitting code as a mandatory first step."
This story demonstrates a specific technical error, the diagnostic investigation that found it, and a concrete process change that prevents recurrence. It is a far stronger interview answer than any description of a successful project.
Practice Method
Review every ML project you have worked on and write the honest version of each project story, specifically identifying the evaluation mistake, the data problem, the underperforming model, or the unexpected production behavior that occurred. If you genuinely cannot identify any, you likely worked on very constrained problems or have not reflected deeply enough. Even tutorial projects contain instructive failures if you examine the gap between your expected results and your actual results on each experiment.
Before reporting any classification metric, do you check what a naive majority-class classifier would achieve and select a metric that is informative under the actual class distribution?
[ ] When a model underperforms, do you plot and analyze learning curves before proposing any hyperparameter changes?
[ ] For every algorithm you list on your resume, can you explain the specific insight that makes it work and name a failure mode that follows from that mechanism?
[ ] Before writing any data splitting code, do you identify the unit of independence in the dataset and select a splitting strategy that respects it?
[ ] When describing offline evaluation results, can you name three specific mechanisms by which offline performance could overestimate production performance for the specific problem type?
[ ] When describing a deployed model, can you describe the serving infrastructure, the monitoring signals, and the retraining trigger separately from the training pipeline?
[ ] Can you describe the ML pipeline lifecycle in full, including the production phase with specific monitoring signals, without stopping at deployment?
[ ] Can you name and implement a label-free drift detection approach using a statistical test appropriate for each feature type in a model you have built?
[ ] When designing a RAG pipeline, do you specify a chunking strategy with a justification and a retrieval evaluation methodology before describing the LLM component?
[ ] Can you name four specific factors that drive the RAG versus fine-tuning decision and apply them to a given use case without using "it depends" as a complete answer?
[ ] Before choosing any evaluation metric for an LLM application, do you explicitly define the cost of each error type and select the metric that reflects the cost asymmetry?
[ ] Do you apply a five-item ML-specific code review checklist to any AI-generated training code before running or submitting it?
[ ] Do you have a specific ML project failure story prepared that names the root cause, the investigation process, and the process change that resulted?
If more than three of these are unchecked, this list is your concrete preparation plan for the next two weeks, not supplemental reading.
Thirteen ML-specific interview mistakes, each with a mechanism-level explanation and a rehearsable fix. The pattern connecting them is the same one that connects every production ML failure worth analyzing: the difference between knowing the algorithm and understanding what it does; the difference between achieving a metric and knowing what the metric actually measures; the difference between deploying a model and maintaining a model in production.
Every fix in this guide is a habit, not a reminder. Checking the naive baseline before any metric is a habit. Running learning curves before tuning hyperparameters is a habit. Identifying the unit of independence before writing a split is a habit. The challenge is not learning that these habits exist but building them strongly enough that they hold under the specific pressure of a live interview follow-up question that you did not expect.
That is the exact gap that structured, realistic practice under pressure closes. Understanding that you should plot learning curves before tuning and actually doing it when an interviewer is watching and waiting are meaningfully different skills. Platforms like Mocklingo's AI mock interview practice give you the environment where the habit gets tested under realistic follow-up pressure before it has to hold in an interview that decides your offer, which is the only place where the habit actually matters.
