In this guide
Machine Learning Interview Questions and Answers
Forty questions on the foundations underneath modern AI — how a model trains, why it fails, and how you'd actually know if it's any good. This page stays on traditional, foundational Machine Learning rather than LLMs, agents, or prompting; formulas appear only where they genuinely help, not as the explanation itself.
Machine Learning Fundamentals
1. What is Machine Learning?
A way of building software that improves at a task by learning from examples, instead of following rules a person wrote by hand. You give it data where the right answer is already known, and training adjusts the system until it can produce the right answer for new examples it hasn't seen.
2. How is Machine Learning different from Artificial Intelligence?
AI is the broad goal: software that performs tasks needing human-like judgment. Machine learning is the main approach used to achieve that today — training a system on examples rather than hand-writing rules. Every ML system is AI; not every AI approach uses machine learning.
3. What is the difference between Machine Learning and Deep Learning?
Deep learning is a specific kind of machine learning that uses neural networks with many stacked layers. It's the approach behind most recent breakthroughs, but machine learning also includes simpler methods — decision trees, linear regression — that don't use deep neural networks at all.
4. What is supervised learning?
Training on examples that already have the right answer labeled — show it inputs paired with known correct outputs, and it learns to predict the output for new, unseen inputs.
5. What is unsupervised learning?
Training on data with no labeled right answer at all — the model finds structure on its own, like grouping similar items together, without being told what the groups should be.
6. What is reinforcement learning?
Training a system by letting it take actions in an environment and rewarding or penalizing the outcomes, so it learns which actions lead to better results over repeated trials, rather than learning from a fixed set of labeled examples.
7. What is the difference between classification and regression?
Classification predicts a category — spam or not spam. Regression predicts a continuous numerical value — an estimated house price of €420,000. Same general approach, different kind of output.
8. What are features and labels?
Features are the pieces of information a model uses to make a prediction. A label is the value it's trying to predict, in supervised learning. For a house-price model, features might be the number of bedrooms, floor area, location, and age of the property; the label is the sale price.
9. What is a Machine Learning model?
The output of training — a set of learned parameters that, given new input matching the features it was trained on, produces a prediction. It's the thing that actually gets used at inference time — when the trained model is applied to a new, real input — after training is done.
10. What does it mean to train a model?
Repeatedly showing it examples with known correct answers and adjusting its internal parameters, over many passes, until its predictions get closer to those correct answers.
Data and Model Training
11. What is training data?
The set of examples, each with a known correct answer, that a model learns from during training. The model's eventual quality depends heavily on how well this data represents the real cases it'll actually see.
12. Why do we split data into training, validation, and test sets?
Training data teaches the model. Validation data is used to tune decisions during development — like comparing model variants — without touching the data used for the final check. Test data measures how well the finished model performs on examples it has never influenced in any way, which is the only honest estimate of real-world performance.
13. What is data preprocessing?
Cleaning and transforming raw data into a form a model can actually learn from — handling missing values, converting categories into numbers, scaling values onto comparable ranges — before training begins.
14. What is feature engineering?
Creating or transforming the input features a model uses, based on judgment about what's actually predictive for the task, rather than only using raw data exactly as it was collected. Combining a start and end date into a single "duration" feature is a simple example.
15. What is feature selection?
Choosing which features to actually use for training, and dropping the ones that add noise or redundancy rather than useful signal.
16. Why is data quality important in Machine Learning?
Because a model can only learn the patterns present in its training data — mislabeled examples, missing values, or data that doesn't represent real cases well all teach the model something wrong, no matter how good the algorithm is.
17. What is data leakage?
When information that wouldn't actually be available at real prediction time accidentally makes it into training — for example, a feature that's only known after the outcome already happened. It makes a model look far more accurate during testing than it will ever be in real use.
18. What is class imbalance?
When one category in the training data is far more common than another — thousands of normal transactions for every fraudulent one, for instance. It can cause a model to become very good at predicting the common class and effectively ignore the rare one, since doing so still scores well on simple measures like accuracy.
19. How would you handle missing data?
Depends on how much is missing and why. Options include removing rows or features with too much missing data, filling gaps with a reasonable estimate, or explicitly flagging that a value was missing so the model can account for it rather than treating a filled-in guess as if it were real.
20. Why can more training data improve a model?
More examples generally give a model a better chance of learning the real underlying pattern instead of quirks specific to a small sample. It's not unlimited — beyond a point, more data helps less, and data quality and relevance matter as much as sheer quantity.
Model Performance
21. What is overfitting?
A model that has learned the training data too specifically, including its noise and quirks, rather than the general pattern. Imagine a student who memorizes the exact answers to 100 practice questions and scores perfectly on them, but does badly when the real exam has new questions — that's overfitting. It performs well on data it's already seen and poorly on data it hasn't.
22. What is underfitting?
A model that hasn't learned the useful pattern in the data well enough, and performs poorly even on the training data itself — the opposite failure from overfitting.
23. What is bias in Machine Learning?
Error that comes from a model being too simple to capture the real pattern in the data — consistently missing in the same direction regardless of which training examples it saw. High bias looks like underfitting.
24. What is variance?
How much a model's predictions change if it were trained on a different sample of the same kind of data. High variance means the model is very sensitive to the specific examples it happened to train on — a sign of overfitting.
25. What is the bias-variance trade-off?
Reducing bias, by making a model more flexible so it can capture more complex patterns, tends to increase variance. Reducing variance tends to increase bias. The goal is finding a balance where the model captures the real pattern without also learning noise specific to its training set.
26. What is cross-validation?
Splitting the training data into several parts, training on some and validating on the rest, then rotating which part is held out and repeating — giving a more reliable estimate of real performance than a single split would.
27. What is regularization?
A technique that discourages a model from fitting the training data too closely, by penalizing overly complex patterns during training. It's a direct tool against overfitting — trading a small amount of training-data accuracy for a model that generalizes better to new data.
28. What are hyperparameters?
Settings chosen before or around the training process, rather than learned by the model during training — how complex to allow the model to get, how many training passes to run, how strongly to apply regularization. Parameters are what the model learns; hyperparameters are what a person or a search process decides beforehand.
29. What is hyperparameter tuning?
Systematically trying different hyperparameter settings and comparing the resulting model's performance on validation data, to find the combination that generalizes best rather than guessing at reasonable-sounding values.
30. How can you tell whether a model is performing well?
Check its performance on data it never saw during training or tuning, not on the training data itself. And check it using a metric that actually matches what the application needs, since a model can look excellent on the wrong metric and still be the wrong model for the job.
Evaluation Metrics
31. What is accuracy?
The share of all predictions a model got right. Simple to understand, but it can be a misleading measure on its own, especially when one outcome is much rarer than the other.
32. What are precision and recall?
Precision asks: of everything the model predicted as positive, how many actually were? Recall asks: of all the actual positive cases, how many did the model find? A model can score high on one and poorly on the other, which is why they're reported separately rather than as a single number.
33. What is an F1 score?
A single number that combines precision and recall, useful when you want to compare models without weighing two separate numbers against each other by hand. It's a summary, not a replacement for looking at precision and recall individually when they disagree sharply.
34. What is a confusion matrix?
A table breaking predictions down into correct and incorrect outcomes by category — how many true positives, false positives, true negatives, and false negatives a model produced. It's what precision, recall, and related metrics are actually calculated from.
35. When can accuracy be misleading?
Whenever one outcome is much rarer than the other. A fraud-detection dataset with 9,900 normal transactions and 100 fraudulent ones lets a model that predicts "not fraud" for everything score 99% accurate while catching zero fraud — completely useless despite the impressive-looking number. This is exactly why precision, recall, F1, and the confusion matrix matter: they reveal what a single accuracy figure hides.
Practical Machine Learning Questions
36. How would you choose a Machine Learning algorithm?
Start from the problem shape — classification or regression, how much labeled data exists, how important it is that the model's reasoning be explainable, and how much time and compute is available for training and serving. Simpler algorithms like linear or logistic regression are often a reasonable starting point before reaching for something more complex like gradient boosting or a neural network.
37. What would you do if a model performs well on training data but poorly on new data?
That's overfitting. Try simplifying the model, adding regularization, gathering more or more representative training data, and checking that the train/validation/test split was done properly in the first place — a leak between them can also produce this exact symptom.
38. How would you improve a Machine Learning model?
Depends on the diagnosis. Improve data quality or quantity if the model is underfitting from insufficient signal, simplify or regularize it if it's overfitting, engineer better features if the raw inputs aren't capturing what actually matters, and tune hyperparameters systematically rather than by guesswork.
39. What should you monitor after deploying a Machine Learning model?
Its real-world performance on the metric that actually matters for the application, whether the incoming data still resembles what it was trained on, and how its predictions are actually being used downstream. A model can quietly degrade as real-world data drifts away from its training data, well after it looked fine at launch.
40. When should you use Machine Learning instead of a simple rule-based system?
When the pattern is genuinely difficult to express as a small set of reliable explicit rules, and enough representative data exists to learn it from. "Customers under 18 cannot create this account type" needs a simple rule, not machine learning. "Determine whether a transaction looks fraudulent based on hundreds of behavioral signals" is a much better fit, because that pattern is far harder to hand-write as fixed rules. Don't reach for machine learning simply because it's available — use it when the problem actually calls for it.
Distinctions worth remembering
Classification predicts a category; regression predicts a number. Spam or not spam is classification. An estimated price of €420,000 is regression.
Training data teaches; test data measures. Training data helps the model learn. Test data measures how well the trained model performs on examples it never saw.
Parameters are learned; hyperparameters are chosen. Parameters come from training. Hyperparameters are settings decided before or around it.
Overfitting learns too specifically; underfitting doesn't learn enough. One memorizes noise; the other misses the real pattern.
Going deeper on one area
For modern LLM-based AI instead of classical ML: AI Interview Questions and Answers, LLM Interview Questions, or Generative AI Interview Questions.