Machine Learning with Scikit-Learn, Step by Step
Built from a real working notebook, refined into a teaching document.
📖 How to Use This Document
This is a reference you can come back to, not something you need to memorize in one sitting.
- New to ML? Start from the beginning and follow the examples in order.
- Already comfortable? Jump directly to the section you need, such as evaluation metrics, hyperparameter tuning, pipelines, or saving models.
- Following the code? Every example uses real, runnable Scikit-Learn code.
- Learning the concepts? Each important idea is explained before the code that uses it.
The goal is simple:
You should understand why you're writing the code, not just know what to type.
📦 Datasets Used in This Tutorial
We'll use four datasets throughout the tutorial.
| Dataset | Download link |
|---|---|
| Heart Disease (CSV) | 🔗 Link |
| California Housing | ✅ Built into Scikit-Learn (fetch_california_housing()) — no download needed |
| Car Sales (CSV) | 🔗 Link |
| Car Sales Extended (with missing data) (CSV) | 🔗 Link |
💡 Tip: If you're using Google Colab, upload the CSV files to your Colab session or mount Google Drive. Then make sure the file path in your code points to the correct location.
📚 Table of Contents
- What is Machine Learning?
- The Scikit-Learn Workflow (the big picture)
- Setup
- Getting Your Data Ready
- Choosing the Right Estimator
- Fitting a Model and Making Predictions
- Evaluating a Model
- Improving a Model
- Saving and Loading a Trained Model
- Putting It All Together (Pipelines)
- Common Errors & How to Fix Them
- Quick-Reference Cheat Sheet
1. What is Machine Learning?
Machine Learning (ML) is a way of building computer programs that learn patterns from examples. In traditional programming, you write the rules. In machine learning, you give the computer examples and let the algorithm learn the rules.
Traditional programming vs. machine learning
| Traditional Programming | Machine Learning | |
|---|---|---|
| You provide | Explicit rules | Examples |
| The computer | Follows your rules | Finds patterns in the examples |
| Output | Result based on your rules | A learned model that can make predictions |
| Good for | Problems with clear rules | Problems where rules are difficult to write manually |
🧠 A simple example
Imagine you want a computer to recognize dogs in pictures. With traditional programming, you might try to write rules such as:
"If the picture has four legs, fur, two ears, and a tail, call it a dog."
But this quickly becomes difficult. Some dogs are sitting, some are lying down, some are partially hidden, and other animals can have similar characteristics. With machine learning, you can instead give the computer many examples:
Picture → Dog
Picture → Dog
Picture → Not Dog
Picture → Dog
Picture → Not Dog
...
The algorithm looks at these examples and learns patterns that help it distinguish dogs from other things. That's the basic idea behind machine learning.
The same idea works for many problems:
- Spam detection → learn from emails labelled spam/not spam.
- House price prediction → learn from houses and their known prices.
- Medical prediction → learn from patient information and known outcomes.
In every case, we have:
Inputs → Model → Prediction
The inputs are usually called features, and the answer we're trying to predict is called the target or label.
Where Scikit-Learn fits
Scikit-Learn is a Python library that provides many tools for machine learning. It is especially useful for structured data — data organized into rows and columns.
For example:
Age Salary Experience Purchased
25 40000 2 0
32 60000 5 1
41 90000 10 1
This looks like a spreadsheet, which makes it a good fit for Scikit-Learn.
| Data type | Examples | Common tools |
|---|---|---|
| Structured | CSVs, SQL tables, spreadsheets | Scikit-Learn |
| Unstructured | Images, audio, video, raw text | Deep learning frameworks such as PyTorch/TensorFlow |
🎯 Rule of thumb: If your data naturally looks like rows and columns, Scikit-Learn is a very good place to start.
2. The Scikit-Learn Workflow (the big picture)
Most basic Scikit-Learn projects follow the same general process. Learn this flow because you'll see it again and again:
1. Get data ready
↓
2. Pick a model
↓
3. Fit the model to data
↓
4. Make predictions
↓
5. Evaluate the model
↓
Good enough? ── no ──▶ 6. Improve via tuning ──┐
│ yes │
│ │
▼ │
7. Save and reload the model ◀────────────────────┘
Let's translate that into normal language.
1. Get your data ready
Load your data and prepare it so the model can use it. This might involve:
- separating features and targets
- handling missing values
- converting text categories into numbers
- scaling features when necessary
2. Pick a model
Choose an algorithm that matches your problem. For example:
- Predicting Yes/No → classification
- Predicting a price → regression
3. Fit the model
Give the model training examples.
model.fit(X_train, y_train)This is where the model learns patterns from your data.
4. Make predictions
Give the trained model data it hasn't seen before.
model.predict(X_test)5. Evaluate
Check how good those predictions are. For example:
Classification metrics
- Accuracy
- Precision
- Recall
- F1
Regression metrics
- R²
- MAE
- MSE
The metric depends on the problem.
6. Improve
If the model isn't good enough, experiment. You might:
- improve the data
- try another model
- change model settings
- tune hyperparameters
- add or remove useful features
7. Save
Once you have a model you're happy with, save it to disk.
Then you can load it later without training it again.
🚀 A first end-to-end example
Before learning each individual piece, let's look at the whole process. Don't worry if some lines don't make sense yet. We'll explain each part later.
# 1. Get the data ready
import pandas as pd
heart_disease = pd.read_csv("sklearn-heart-disease.csv")
X = heart_disease.drop("target", axis=1) # features
y = heart_disease["target"] # labels
# 2. Choose the right model/estimator
from sklearn.ensemble import RandomForestClassifier
clf = RandomForestClassifier()
# 3. Fit the model to the training data
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=1
)
clf.fit(X_train, y_train)
# 4. Make a prediction
y_preds = clf.predict(X_test)
# 5. Evaluate the model
clf.score(X_train, y_train) # accuracy on data it has already seen
clf.score(X_test, y_test) # accuracy on data it has NOT seen — this is what matters
from sklearn.metrics import classification_report, confusion_matrix, accuracy_score
print(classification_report(y_test, y_preds))
confusion_matrix(y_test, y_preds)
accuracy_score(y_test, y_preds)
# 6. Improve the model (a simple manual sweep)
for i in range(10, 100, 10):
clf = RandomForestClassifier(n_estimators=i).fit(X_train, y_train)
print(f"n_estimators={i} -> test accuracy: {clf.score(X_test, y_test) * 100:.2f}%")
# 7. Save the model and load it back
import pickle
pickle.dump(clf, open("random_forest_model_1.pkl", "wb"))
loaded_model = pickle.load(open("random_forest_model_1.pkl", "rb"))
loaded_model.score(X_test, y_test)This is the basic machine learning loop:
Data → Split → Train → Predict → Evaluate → Improve → SaveThat's the entire universe of a basic ML project. Everything else in this handbook goes deeper into each of these seven steps.
3. Setup
Let's get the environment ready. Only a handful of lines.
📂 Mount Google Drive (if using Colab)
If you're using Google Colab and your datasets are stored in Google Drive, you can connect your Drive like this:
from google.colab import drive
drive.mount('/content/drive')After running this, Colab will ask you for permission. Once connected, your Drive becomes accessible through /content/drive/. If you're running the notebook locally, you don't need this.
💡 If running locally, skip this and use direct file paths.
📊 Standard Imports
# Import NumPy for numerical operations and working with arrays
import numpy as np
# Import Pandas for working with data and DataFrames
import pandas as pd
# Import Matplotlib for creating visualizations and plots
import matplotlib.pyplot as plt
# Display Matplotlib plots directly inside the Jupyter Notebook
%matplotlib inline📝 A note on GPUs: Scikit-Learn runs on the CPU. Unlike deep learning, you don't need a GPU — checking
!nvidia-smiin Colab is harmless but irrelevant here.
4. Getting Your Data Ready
This is where a lot of the real work in machine learning happens. Training a model is often just a few lines of code. The harder part is preparing the data so the model can actually learn from it.
There are three important things we usually need to do:
- Split the data into features and labels — usually called
Xandy. - Handle missing values — fill them with a suitable value (imputation) or remove the rows/columns containing them.
- Convert non-numeric data into numbers — this is called feature encoding.
We'll go through each of these steps one by one. Before we can do any of that, we need data loaded into the notebook.
Load the Heart Disease dataset:
heart_disease = pd.read_csv("sklearn-heart-disease.csv")pd.read_csv() reads the CSV file and creates a Pandas DataFrame.
Place the CSV in the same folder as your notebook (or update the path inside read_csv to match where you saved it).
Let's take a quick look at what we loaded:
heart_disease.head()You should see a table with columns like age, sex, chol, target, and so on. target is the column we'll try to predict — 1 means "has heart disease", 0 means "does not".
Check the size of the data:
Let's also check how many rows and columns we have:
heart_disease.shapeThis returns something like:
(rows, columns)For example:
(303, 14)would mean:
- 303 rows
- 14 columns
Check for missing values:
heart_disease.isna().sum()isna() checks which values are missing. sum() counts them.
So:
heart_disease.isna().sum()answers:
"How many missing values does each column contain?"
4.1 Splitting into Features (X) and Labels (y)
X= the features / "the questions" — everything you use to make a prediction.y= the label / "the answer" — the thing you're trying to predict.
X = heart_disease.drop("target", axis=1)This creates X by dropping the target column. X now holds every other column — the features the model will learn from.
y = heart_disease["target"]This creates y as just the target column — the labels the model will try to predict.
That's the first step done. Next, we split both into training and test sets.
A concrete example
Suppose our data looks like this:
| Age | Cholesterol | Exercise | Disease |
|---|---|---|---|
| 50 | 220 | 2 | 1 |
| 35 | 180 | 5 | 0 |
| 62 | 250 | 1 | 1 |
We want to predict Disease. So:
X = data.drop("Disease", axis=1) # Age, Cholesterol, Exercise
y = data["Disease"] # DiseaseThe result is:
X:
Age Cholesterol Exercise
50 220 2
35 180 5
62 250 1
y:
1
0
1A useful way to remember this:
X tells the model what it can look at.
y tells the model what answer it should learn to predict.
🧠 Vocabulary check
You'll see different names for these throughout ML documentation.
| Symbol | Also called |
|---|---|
X |
Features, feature matrix, input data, independent variables |
y |
Labels, targets, target variable, dependent variable, answer |
These terms refer to the same basic concepts. Get comfortable with them.
4.2 Splitting into training and test sets
You never want to evaluate a model on data it has already memorized — that tells you nothing about how it performs on new, real-world data. So you hold back a chunk of data purely for testing.
First, import the splitter:
from sklearn.model_selection import train_test_splitThis imports Scikit-Learn's tool for splitting data.
Now use it:
X_train, X_test, y_train, y_test = train_test_split(
X,
y,
test_size=0.2,
random_state=1
)Here's what just happened:
- We gave
train_test_splitour featuresXand our labelsy. - It returned four things, in this exact order:
X_train→ features used for trainingX_test→ features used for testingy_train→ correct answers for trainingy_test→ correct answers for testing
test_size=0.2→ means 20% of the data is held back for testing.random_state=1→ makes the split reproducible — same shuffle every run.
Let's confirm the split:
X_train.shape, X_test.shape, y_train.shape, y_test.shapeYou should see the training set is about 80% of the rows, and the test set is about 20%.
🧠 Mental model: the exam
Imagine studying for an exam.
X_train/y_trainare the practice questions with answers you study from.X_test/y_testare the real exam questions — you must not have seen these answers beforehand, or your "score" is meaningless.
If you trained on the test set, you'd be grading your own exam with the answer key open — you'd get a great score and learn nothing about how you'd actually perform.
What does random_state=1 mean?
train_test_split() randomly chooses which rows go into each set. If those rows change, your model's score can also change slightly. random_state=1 makes the random split repeatable.
So every time you run the code, you get the same split. The actual number doesn't matter.
These are all valid:
random_state=1
random_state=42
random_state=123The important thing is using a fixed value when you want reproducible results.
Common test sizes
test_size |
When to use |
|---|---|
0.2 (20%) |
Default for most datasets |
0.1 (10%) |
Very large datasets (millions of rows) |
0.3 (30%) |
Very small datasets (hundreds of rows) — you want more evaluation confidence |
These aren't strict rules. The right split depends on the amount and type of data you have.
4.3 Making sure everything is numerical (Feature Encoding)
Most traditional Scikit-Learn models expect their input data to be numbers.
Let's first check the data types of our columns:
heart_disease.dtypesIf every column shows int64 or float64, you’re okay. If you see object, that column probably contains text and needs to be converted into numbers.
Our heart disease dataset is already numerical, so for this section we'll use a car sales dataset that contains some actual text columns.
Real-world datasets often contain categorical values such as: "Make": "Toyota" or "Colour": "Blue".
Scikit-Learn models generally cannot use these text values directly. We need to convert them into numbers first.
This process is called feature encoding.
🔹 Why not simply assign numbers?
You might think:
"Toyota = 1, Honda = 2, BMW = 3. Done!"
The problem is that these numbers can make the model think there is an order or distance between the categories.
For example:
BMW = 3
Honda = 2
Toyota = 1
This could make the model think:
- BMW > Honda > Toyota
- BMW is "one unit" away from Honda
- Honda is "one unit" away from Toyota
But these relationships don't actually exist. BMW isn't mathematically greater than Honda, for example.
This is especially important for models such as linear models, SVMs, KNN, and neural networks, because they treat these numbers as real numerical values.
For categories without a natural order, called nominal categories, we usually use one-hot encoding instead.
📝 Two exceptions worth knowing:
- Ordinal categories have a natural order, such as
Small < Medium < LargeorPoor < Good < Excellent. In this case, integer values can represent the order. We can useOrdinalEncoderand explicitly define the order when needed.- Tree-based models such as decision trees, random forests and gradient boosting split values using thresholds, so they can handle integer-coded categories better than many other models. However, one-hot encoding is still a safe general-purpose choice.
🔹 One-Hot Encoding
One-hot encoding creates a separate column for each category.
Each column contains either:
0 = category does not apply
1 = category applies
For example:
Colour → Blue Green Red
Red 0 0 1
Blue 1 0 0
Green 0 1 0
Each category gets its own position, so there is no artificial ordering between categories.
📖 Worked example
One-hot encoding can also be used for words.
Suppose we have:
Corpus:
["cats like milk", "dogs like milk", "cats eat fish", "dogs eat fish"]
The unique words are:
Vocabulary: [cats, dogs, like, milk, eat, fish]
There are 6 unique words, so each word gets a vector with 6 positions:
cats → [1, 0, 0, 0, 0, 0]
dogs → [0, 1, 0, 0, 0, 0]
like → [0, 0, 1, 0, 0, 0]
milk → [0, 0, 0, 1, 0, 0]
eat → [0, 0, 0, 0, 1, 0]
fish → [0, 0, 0, 0, 0, 1]
Notice that "cats" and "dogs" are simply represented by different positions.
One-hot encoding does not understand that cats and dogs are both animals or that they are semantically related. It only tells the model that they are different categories.
This is one reason techniques such as Word2Vec and BERT embeddings are useful for text: embeddings can represent relationships and meaning, while one-hot encoding cannot.
If we add the one-hot vectors for all the words in:
"cats like milk"
we get:
[1, 0, 1, 1, 0, 0]
Tabular example
For a car dataset:
Colour (input) → Colour_Blue Colour_Green Colour_Red
---------------- ----------- ------------ ----------
Red 0 0 1
Blue 1 0 0
Green 0 1 0
Each category gets its own feature. There is no artificial ordering between Red, Blue and Green.
By default, Scikit-Learn sorts the categories, which is why the columns appear as:
Blue, Green, Red
🔹 In practice, with tabular data
For tabular data, Scikit-Learn provides OneHotEncoder.
First, load the data and separate the features (X) from the target (y):
import pandas as pd
car_sales = pd.read_csv("sklearn-car-sales.csv")
X = car_sales.drop("Price", axis=1)
y = car_sales["Price"]Split before encoding
We should split our data before fitting the encoder.
Why?
Because the encoder learns which categories exist in the data when we call fit().
If we fit it using the entire dataset, it can learn information from the test set. This is a form of data leakage.
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=1
)Now we have:
Training data → used to learn
Test data → used only for evaluation
Now let's create our encoder:
from sklearn.preprocessing import OneHotEncoder
from sklearn.compose import ColumnTransformer
# Pick which columns to encode
categorical_features = ["Make", "Colour", "Doors"]
one_hot = OneHotEncoder(handle_unknown="ignore")
transformer = ColumnTransformer(
[("one_hot", one_hot, categorical_features)],
remainder="passthrough"
)📝 Note:
Doorsis stored as a number (3, 4, 5), but we treat it as a category here.A 3-door and a 5-door car are different kinds of car, and treating it as categorical stops the model assuming the effect grows in a straight line with the number of doors.
This is a judgement call: the values do have an order, so numeric is also defensible.
Here’s what each piece does:
OneHotEncoder(handle_unknown="ignore")creates the encoder (more onhandle_unknownbelow).ColumnTransformerapplies a transformation to specific columns and leaves everything else alone.[("one_hot", one_hot, categorical_features)]is a list of transformations. Each one is a tuple:(name, transformer, columns). Here it means: “applyone_hottocategorical_features, and call this stepone_hot.”remainder="passthrough"keeps the untouched columns as they are.
Now we fit the encoder using only the training data:
X_train_enc = transformer.fit_transform(X_train)fit_transform() does two things:
- Learns the categories from the training data.
- Converts the training data into encoded numerical data.
For the test data, we only use transform():
X_test_enc = transformer.transform(X_test)This uses the same categories learned from the training data.
We don't fit the encoder again on the test data.
The result is a numerical array where:
- categorical columns have been converted into one-hot columns
- numerical columns remain unchanged
If the encoded data contains many zeros, Scikit-Learn may store it as a sparse matrix to save memory. Many Scikit-Learn models can work with sparse matrices directly.
| Method | What it does |
|---|---|
fit() |
Learns information from the data |
transform() |
Converts data using what was learned |
fit_transform() |
Learns and converts in one step |
A simple way to remember this is:
fit() → Learn
transform() → Use what was learned
fit_transform() → Learn + transform
🔹 A pandas shortcut
For quick exploration, pandas provides:
dummies = pd.get_dummies(
car_sales[["Make", "Colour", "Doors"]],
dtype=float
)pd.get_dummies does the same one-hot expansion but returns a pandas DataFrame.
A simple rule:
pd.get_dummies()is convenient for exploration.
OneHotEncoderis usually a better choice when building Scikit-Learn pipelines.
The reason: get_dummies doesn’t remember which categories it saw.
If the training and test sets contain different categories, you end up with different columns in each.
OneHotEncoder remembers its categories after fit() and applies them consistently.
🔹 Fit a model on the encoded data
Once our categorical values have been converted into numbers, we can give the data to a model.
from sklearn.ensemble import RandomForestRegressor
model = RandomForestRegressor()
model.fit(X_train_enc, y_train)
model.score(X_test_enc, y_test)Here’s what each piece does:
RandomForestRegressor()creates the Random Forest model.train_test_split()(used earlier) splits the data into training and testing sets, withtest_size=0.2putting 20% in the test set andrandom_state=1making the split reproducible.model.fit(X_train_enc, y_train)trains the model.model.score(X_test_enc, y_test)evaluates the model on the test data and returns its R² score.
🔹 The handle_unknown gotcha
Because we fit the encoder on the training data only, the test set (or future data) can contain a category the encoder never saw.
By default, OneHotEncoder crashes in that case:
ValueError: Found unknown categories ['Ferrari'] in column 0 during transform
Fix it with:
one_hot = OneHotEncoder(handle_unknown="ignore")Now unseen categories produce a row of zeros instead of crashing. This is critical for production systems where new categories appear over time.
Key idea: Machine learning models generally need numerical input. For categorical features without a natural order, one-hot encoding converts each category into its own numerical feature. With Scikit-Learn, fit the encoder on the training data and use the learned encoder to transform validation or test data.
4.4 Handling missing values
Real-world data often contains missing values. Many ML models cannot work directly with missing values, so we need to decide what to do with them.
There are two common choices:
- Fill the missing values (called imputation).
- Drop the rows/columns with missing values.
To see what's missing in the car sales dataset, load it first:
import pandas as pd
car_sales_missing = pd.read_csv("sklearn-car-sales-missing-data.csv")Now check how much is missing:
car_sales_missing.isna().sum()This tells us how many missing values each column contains.
🔹 Option 1 — Fill missing data with pandas
Start with the categorical columns:
car_sales_missing["Make"] = car_sales_missing["Make"].fillna("missing")fillna("missing") replaces every missing value in the Make column with the string "missing". We use a string because this column holds categories.
Do the same for Colour:
car_sales_missing["Colour"] = car_sales_missing["Colour"].fillna("missing")For the numeric odometer column, we can use the mean:
car_sales_missing["Odometer (KM)"] = car_sales_missing["Odometer (KM)"].fillna(
car_sales_missing["Odometer (KM)"].mean()
)This means:
"If the odometer value is missing, use the average odometer value."
For Doors, we can use a sensible default:
car_sales_missing["Doors"] = car_sales_missing["Doors"].fillna(4)Finally, remove rows where the target is missing:
car_sales_missing = car_sales_missing.dropna(subset=["Price"])Why remove these rows?
Because Price is the answer we're trying to predict.
If a row doesn't contain the answer:
Features → ?
we can't use that row as a normal training example.
Why fill features but drop a missing target?
Imagine:
Age: 35
Salary: missing
Purchased: 1This row still contains a useful answer:
Purchased = 1We can estimate the missing salary. But consider:
Age: 35
Salary: 50000
Purchased: missingWe don't know the answer we're supposed to train against. So a common approach is:
Missing feature → fill it
Missing target → remove the row🔹 Option 2 — Fill missing values with Scikit-Learn's SimpleImputer
This is a more production-ready approach because it works well with pipelines and can be reused on new data. It remembers the value calculated from the training data (such as the mean) and uses the same value for future data.
from sklearn.impute import SimpleImputer
from sklearn.compose import ColumnTransformer
cat_imputer = SimpleImputer(strategy="constant", fill_value="missing")
door_imputer = SimpleImputer(strategy="constant", fill_value=4)
num_imputer = SimpleImputer(strategy="mean")Three imputers, each with a different strategy:
constantwithfill_value="missing"→ replaces NA in categorical columns with the string"missing".constantwithfill_value=4→ replaces NA inDoorswith4.mean→ replaces NA in the numeric column with the column mean.
Now tell ColumnTransformer which imputer goes with which columns:
cat_features = ["Make", "Colour"]
door_feature = ["Doors"]
num_features = ["Odometer (KM)"]
# Each tuple is `(name, transformer, columns)`
imputer = ColumnTransformer([
("cat_imputer", cat_imputer, cat_features),
("door_imputer", door_imputer, door_feature),
("num_imputer", num_imputer, num_features)
])Then fit and transform:
filled_X = imputer.fit_transform(X)filled_X is now a NumPy array with no missing values. Convert it back to a DataFrame for readability:
car_sales_filled = pd.DataFrame(
filled_X,
columns=["Make", "Colour", "Doors", "Odometer (KM)"]
)
car_sales_filled.isna().sum() # should be zeroThen encode the complete data as usual with OneHotEncoder, and finally fit your model.
🔹 Common imputation strategies
SimpleImputer(strategy="mean") # numeric: replace with mean
SimpleImputer(strategy="median") # numeric: replace with median (outlier-robust)
SimpleImputer(strategy="most_frequent") # categorical or numeric: replace with mode
SimpleImputer(strategy="constant", fill_value=0) # replace with a fixed value🧭 Order matters
When preparing data, the steps often look like:
Missing values → Imputaion → Encoding → Split → Fit
And when splitting the data, we need to make sure preprocessing is learned only from training data.
💡 Put all of these steps inside one
Pipelineso you never have to remember the order. Pipelines also prevent data leakage.
⚠️ The quantity vs. cleanliness trade-off
If you drop rows with missing values before imputing, you may end up with a noticeably smaller dataset (e.g., 950 rows after dropping vs. 1000 originally). A model trained on less data may score slightly worse, even though it's "cleaner." There's no universally correct choice here — it's a genuine trade-off:
- Imputation keeps more data but introduces estimated values.
- Dropping keeps data pure but reduces quantity.
For missing features, imputation is often useful. For a missing target, dropping the row is usually the straightforward choice.
5. Choosing the Right Estimator
In Scikit-Learn terminology, a model/algorithm is called an estimator. Which one to use depends on your problem type:
| Problem type | What you're predicting | Example | Common variable name |
|---|---|---|---|
| Classification | A category / class | Heart disease or not? | clf (classifier) |
| Regression | A continuous number | House price | model |
🗺️ **Scikit-Learn's official flowchart for picking a model is genuinely useful — bookmark it:**https://scikit-learn.org/stable/machine_learning_map.html
5.1 Picking a model for regression
We'll use the California Housing dataset. It's built into Scikit-Learn, so no download is needed.
First, import it:
from sklearn.datasets import fetch_california_housingfetch_california_housing downloads (or loads from a local cache) the dataset and returns a dictionary-like object.
Load it:
housing = fetch_california_housing()housing now holds the data, feature names, and target values.
Turn it into a pandas DataFrame, so it's easier to work with:
housing_df = pd.DataFrame(
housing["data"],
columns=housing["feature_names"]
)housing["data"]is the feature matrix.housing["feature_names"]gives us the column names.- We build a DataFrame from those two.
Then add the target column:
housing_df["target"] = housing["target"]Now housing_df has all features plus a target column.
Quick look:
housing_df.head()You should see columns like MedInc, HouseAge, AveRooms, target, and so on.
The target represents the median house value in units of $100,000.
🔹 Start with a simple model: Ridge
Ridge is a type of linear regression model. It tries to find a relationship between the input features and the target. It's a useful baseline because it's relatively simple and fast.
from sklearn.linear_model import Ridge
from sklearn.model_selection import train_test_split
X = housing_df.drop("target", axis=1)
y = housing_df["target"]
X_train, X_test, y_train, y_test = train_test_split(
X,
y,
test_size=0.2,
random_state=1
)
model = Ridge()
model.fit(X_train, y_train)
model.score(X_test, y_test) # returns an R² scoreFor a regression model, .score() returns R². We'll explain R² properly later.
🔹 Step up to an ensemble: RandomForestRegressor
A Random Forest uses many decision trees. Instead of relying on one tree, it builds many trees and combines their predictions.
from sklearn.ensemble import RandomForestRegressor
model = RandomForestRegressor()
model.fit(X_train, y_train)
model.score(X_test, y_test)Random Forest can capture relationships that a simple linear model may not be able to capture. For many structured/tabular datasets, tree-based models are useful starting points.
5.2 Picking a model for classification
Now let's return to the Heart Disease dataset.
heart_disease = pd.read_csv("sklearn-heart-disease.csv")Following Scikit-Learn's map might suggest LinearSVC first. LinearSVC is a linear Support Vector Classifier — fast and often a good baseline for classification.
from sklearn.svm import LinearSVC
from sklearn.model_selection import train_test_split
X = heart_disease.drop("target", axis=1)
y = heart_disease["target"]
X_train, X_test, y_train, y_test = train_test_split(
X,
y,
test_size=0.2,
random_state=1
)
clf = LinearSVC(max_iter=10000)
clf.fit(X_train, y_train)
clf.score(X_test, y_test)Now compare against an ensemble method:
from sklearn.ensemble import RandomForestClassifier
clf = RandomForestClassifier(random_state=1)
clf.fit(X_train, y_train)
clf.score(X_test, y_test)Usually performs better on this kind of data.
The important lesson is:
Start with a reasonable baseline, measure it, and then experiment.
6. Fitting a Model and Making Predictions
6.1 Fitting ("training") the model
"Fitting" is the process of the model looking at X_train and y_train and learning the statistical patterns that connect them.
Import the model:
from sklearn.ensemble import RandomForestClassifierCreate it:
clf = RandomForestClassifier(random_state=1)At this point, clf is an untrained classifier. It knows nothing about your dataset.
Now fit it:
clf.fit(X_train, y_train)This is where the learning happens. The model looks at every training row and adjusts its internal state to make good predictions.
Evaluate it:
clf.score(X_test, y_test)This reports accuracy on data the model has never seen — the number that actually matters.
Before fitting:
Model
↓
Knows nothing about this datasetAfter fitting:
X_train + y_train
↓
model.fit()
↓
Trained Model
The model learns its internal parameters during fitting.
6.2 Making predictions
There are two important prediction methods, and knowing when to use which matters a lot in real applications.
🔹 predict() → What class should I choose?
y_preds = clf.predict(X_test)This gives one predicted class for each test example.
For example:
[1, 0, 1, 1, 0]The model has made a final decision for each example.
Check the accuracy manually:
np.mean(y_preds == y_test)This compares each prediction with the actual answer.
Predicted: [1, 0, 1, 1]
Actual: [1, 0, 0, 1]Three out of four are correct:
3 / 4 = 0.75So the accuracy is 75%. You can also use the metric function:
from sklearn.metrics import accuracy_score
accuracy_score(y_test, y_preds)🔹 predict_proba() → How likely is each class?
Some classifiers can provide probabilities for each class.
clf.predict_proba(X_test[:5])You might see:
[[0.89, 0.11],
[0.49, 0.51],
[0.43, 0.57],
[0.84, 0.16],
[0.18, 0.82]]
Each row sums to 1 and represents [P(class 0), P(class 1)]. predict() is really just predict_proba() under the hood, picking whichever class has the higher probability.
🧠 Why does this matter in real life?
A patient with
[0.49, 0.51](barely leaning toward "has heart disease") is a very different clinical situation than one with[0.02, 0.98](highly confident) — even thoughpredict()returns the same class label ("1") for both.If you only ever look at
predict(), you lose this crucial confidence information.predict_probaalso lets you adjust the decision threshold (e.g., only flag "positive" above 70% confidence) instead of the default 50/50 cutoff.
Get only the positive-class probabilities:
y_probs = clf.predict_proba(X_test)
y_probs_positive = y_probs[:, 1][:, 1] means:
"Take every row, but only column 1."
So we keep the probability of class 1.
🔹 predict() for regression
predict() works the same way for regression models. The difference is that a regressor returns numbers instead of class labels.
from sklearn.ensemble import RandomForestRegressor
model = RandomForestRegressor()
model.fit(X_train, y_train)
y_preds = model.predict(X_test)You might get:
[2.41, 1.83, 3.05, 1.74, ...]These are the model's predicted numeric values.
To see how far the predictions are from the actual values, we can measure the error:
from sklearn.metrics import mean_absolute_error
mean_absolute_error(y_test, y_preds)For example, if the result is:
0.52it means that, on average, each prediction differs from the actual value by about 0.52.
For example, if you are predicting house prices in millions, an MAE of 0.52 means the predictions are off by about 0.52 million on average.
📝
predict_proba()is generally not used for regression because regression predicts a continuous number, not a class probability.
7. Evaluating a Model
Training a model isn't enough. We need to answer:
How well does it actually work?
Scikit-Learn gives us several ways to do this.
There are three ways to evaluate a Scikit-Learn model, from least to most detailed:
- The estimator's built-in
.score()method - The
scoringparameter (used with cross-validation) - Problem-specific metric functions from
sklearn.metrics
📖 Full reference: https://scikit-learn.org/stable/modules/model_evaluation.html
7.1 The .score() method
The simplest option. For classifiers, .score() returns accuracy by default: 0.0 means no correct predictions, while 1.0 means all predictions are correct. For regressors, it returns R².
clf.score(X_train, y_train)This measures how well the model performs on the training data — data it has already seen. The score can be overly high because the model has already learned from these examples.
clf.score(X_test, y_test)This measures how well the model performs on new, unseen data. This is usually the more important score because it tells us how well the model generalizes.
⚠️ Overfitting check: If the training score is much higher than the test score, such as 99% vs 65%, the model may be overfitting. This means it learned the training examples very well but doesn't perform as well on new data.
📝 Important:
.score()does not always mean the same thing for every model. For example,RandomForestClassifieruses accuracy, whileRidgeuses R². Always check the model's documentation to know what.score()returns.
7.2 The scoring parameter (and cross-validation)
A normal train/test split gives us one test score. But there is a problem. What if our test set happens to contain mostly easy examples? Our model might get a high score simply because it got an easy test set.
Cross-validation helps us get a more reliable idea of how well our model performs. Instead of testing the model only once, we split the data into several parts and test the model multiple times.
This gives us multiple scores instead of just one.
🔹 5-Fold Cross-Validation
In 5-fold cross-validation, we split our data into 5 roughly equal parts, called folds.
Then we train and test the model 5 times.
Each time:
- 1 fold is used for testing.
- The other 4 folds are used for training.
┌─────────┬─────────┬─────────┬─────────┬─────────┐
│ Fold 1 │ Fold 2 │ Fold 3 │ Fold 4 │ Fold 5 │
└─────────┴─────────┴─────────┴─────────┴─────────┘
TEST TRAIN TRAIN TRAIN TRAIN → Score 1
TRAIN TEST TRAIN TRAIN TRAIN → Score 2
TRAIN TRAIN TEST TRAIN TRAIN → Score 3
TRAIN TRAIN TRAIN TEST TRAIN → Score 4
TRAIN TRAIN TRAIN TRAIN TEST → Score 5
So every data point gets a chance to be part of the test set. At the end, we have 5 scores. We can then calculate their average to get an overall estimate of the model’s performance.
🔹 Use cross_val_score
Scikit-Learn provides cross_val_score() for this:
from sklearn.model_selection import cross_val_scoreCreate the model:
clf = RandomForestClassifier(
random_state=1,
n_estimators=100
)n_estimators=100 means the Random Forest will build 100 decision trees.
Now run cross-validation:
scores = cross_val_score(
clf,
X,
y,
cv=5
)Here’s what each argument means:
| Argument | Meaning |
|---|---|
clf |
The model we want to evaluate |
X, y |
The features and target |
cv=5 |
Use 5-fold cross-validation |
scores will contain 5 numbers, one for each fold.
scoresmight produce:
[0.82, 0.79, 0.84, 0.81, 0.80]Each number represents the model’s score on one test fold.
To calculate the average:
np.mean(scores)This gives us the model’s average cross-validation score.
Why use cross-validation?
A single train/test split can give us a score that is slightly lucky or unlucky depending on which examples ended up in the test set.
Cross-validation gives us several measurements.
For example:
Single test split:
0.84versus:
Cross-validation:
0.82
0.79
0.84
0.81
0.80
Average:
0.812The cross-validation result gives us more information about how consistently the model performs across different parts of the dataset.
The trade-off is that the model must be trained multiple times, so cross-validation takes more time.
🔹 Scoring options
By default, cross_val_score() uses the model’s default scoring method.
For example:
- Classifiers normally use accuracy.
- Regressors normally use R².
You can also tell cross_val_score() exactly which metric you want:
cross_val_score(clf, X, y, cv=5, scoring="accuracy")
cross_val_score(clf, X, y, cv=5, scoring="precision")
cross_val_score(clf, X, y, cv=5, scoring="recall")
cross_val_score(clf, X, y, cv=5, scoring="f1")
cross_val_score(clf, X, y, cv=5, scoring="roc_auc")Each call returns an array containing one score for each fold.
For example:
[0.82, 0.79, 0.84, 0.81, 0.80]The important thing is that the numbers represent the specific metric you requested.
🔹 Regression scoring
For regression models, .score() normally returns R².
We can calculate cross-validated R² like this:
cv_r2 = cross_val_score(
model,
X,
y,
cv=3
)
np.mean(cv_r2)Here, cv=3 means we use 3-fold cross-validation instead of 5-fold.
We can also choose a different metric.
For example, suppose we want Mean Absolute Error (MAE):
cv_mae = cross_val_score(
model,
X,
y,
cv=3,
scoring="neg_mean_absolute_error"
)
np.mean(cv_mae)You may wonder:
Why is the scoring name
neg_mean_absolute_errorinstead of justmean_absolute_error?
Scikit-Learn’s scoring system follows one simple rule:
Higher score = better
But MAE works the opposite way:
Lower error = better
So Scikit-Learn returns the negative value of the error.
For example:
Actual MAE:
10
Scikit-Learn score:
-10A smaller MAE is better:
MAE = 5 → score = -5
MAE = 10 → score = -10Since -5 is higher than -10, Scikit-Learn can still follow its rule that higher scores are better.
When we want to report the actual MAE, we convert it back to a positive number:
actual_mae = -np.mean(cv_mae)
print(
f"Cross-validated MAE: {actual_mae:.2f}"
)The same idea applies to Mean Squared Error (MSE):
cv_mse = cross_val_score(
model,
X,
y,
cv=3,
scoring="neg_mean_squared_error"
)
actual_mse = -np.mean(cv_mse)So remember:
Higher-is-better metrics
↓
accuracy, precision, recall, F1, R²
Lower-is-better metrics
↓
MAE, MSE
Scikit-Learn represents the second group
as negative values when used as scoring metrics.
Key idea: Cross-validation tests the model on different parts of the dataset multiple times. The
scoringparameter lets us choose exactly which metric we want to measure. For error metrics such as MAE and MSE, Scikit-Learn uses negative values so that higher scores always represent better performance.
7.3 Classification Evaluation Metrics
There are several classification metrics you should understand well:
- Accuracy
- ROC-AUC
- Confusion Matrix
- Precision
- Recall
- F1
Each metric answers a different question about how well our classifier is performing.
📌 Accuracy
Accuracy asks:
"Out of all predictions, how many were correct?"
For example, suppose our model makes 100 predictions:
100 predictions
80 correctThen:
Accuracy = 80%We can calculate accuracy with:
from sklearn.metrics import accuracy_score
accuracy_score(y_test, y_preds)We can also calculate it using cross-validation:
scores = cross_val_score(clf, X, y, cv=5)
print(
f"Cross-validated accuracy: "
f"{np.mean(scores) * 100:.2f}%"
)⚠️ The problem with accuracy
Accuracy can be misleading when one class is much more common than another.
Imagine we have 10,000 patients:
9,999 → no disease
1 → diseaseNow imagine a model that always predicts no disease.
It would get:
9,999 correct out of 10,000So its accuracy would be:
99.99%That sounds very good. But the model completely missed the only patient who had the disease. This is why accuracy alone is not always enough. For imbalanced datasets, we may also need metrics such as precision, recall, and F1-score.
📌 ROC Curve and AUC
A classifier can give us more than just a predicted class, such as 0 or 1. It can also give us a probability (or score) showing how strongly it thinks an example belongs to the positive class.
For example:
P(positive) = 0.72This means the model gives this example a positive-class probability of 0.72. But how do we turn 0.72 into an actual prediction?
We use a threshold.
Step 1: Understanding the Threshold
A threshold is simply the line we use to decide whether a prediction is positive or negative.
For example, with a threshold of 0.50:
0.72 >= 0.50 → predict positive
0.30 < 0.50 → predict negative0.50 is a common default threshold, but it is not the only possible choice.
We could use:
0.30
0.50
0.70
0.80or any other value.
The important idea is:
Changing the threshold changes which examples are classified as positive.
For example:
Probability = 0.72With a threshold of 0.50:
0.72 >= 0.50 → PositiveBut with a threshold of 0.80:
0.72 < 0.80 → NegativeThe model's probability did not change. Only the threshold changed. This idea is the foundation of the ROC curve.
Step 2: The Four Possible Outcomes
After choosing a threshold, every prediction falls into one of four categories.
| Term | Meaning |
|---|---|
| True Positive (TP) | Predicted positive and actually positive |
| False Positive (FP) | Predicted positive but actually negative |
| True Negative (TN) | Predicted negative and actually negative |
| False Negative (FN) | Predicted negative but actually positive |
For example:
Actual: 1 Predicted: 1 → True Positive
Actual: 1 Predicted: 0 → False Negative
Actual: 0 Predicted: 1 → False Positive
Actual: 0 Predicted: 0 → True NegativeA simple way to remember them is:
True → the prediction was correct
False → the prediction was wrong
Positive → model predicted positive
Negative → model predicted negative
So:
True Positive → correctly predicted positive
False Positive → incorrectly predicted positive
True Negative → correctly predicted negative
False Negative → incorrectly predicted negative
From these four values, we calculate two important rates.
Step 3: True Positive Rate (TPR)
TPR asks:
Of all the actual positive examples, how many did the model correctly find?
The formula is:
TPR = TP / (TP + FN)TPR is also called:
- Recall
- Sensitivity
For example, suppose we have:
10 actual positive examples
Model correctly finds 8
Model misses 2Then:
TPR = 8 / (8 + 2)
= 0.80So the model found 80% of the actual positives.
Think of TPR as:
"How good is the model at finding the positives?"
Step 4: False Positive Rate (FPR)
FPR asks:
Of all the actual negative examples, how many did the model incorrectly label as positive?
The formula is:
FPR = FP / (FP + TN)For example, suppose we have:
10 actual negative examples
Model incorrectly labels 2 as positive
Model correctly labels 8 as negativeThen:
FPR = 2 / (2 + 8)
= 0.20So the model incorrectly flagged 20% of the actual negatives.
Think of FPR as:
"How many false alarms does the model create?"
Step 5: Why Does the Threshold Matter?
Now we can understand why the threshold is important.
Imagine our model gives:
P(positive) = 0.72With a threshold of 0.50:
0.72 >= 0.50 → PositiveBut with a threshold of 0.80:
0.72 < 0.80 → NegativeAgain:
The probability didn't change. Only the threshold changed.
And when we change the threshold, the number of TP, FP, TN, and FN predictions can also change.
This means TPR and FPR can change too.
Generally:
- Lower threshold → more examples are predicted positive → usually TPR increases, but FPR also increases.
- Higher threshold → fewer examples are predicted positive → usually TPR decreases, but FPR also decreases.
So there is a trade-off:
Lower threshold
↓
Predict more things as positive
↓
Catch more actual positives
↓
TPR increases
BUT
More negative examples may also be called positive
↓
FPR increases
And:
Higher threshold
↓
Predict fewer things as positive
↓
Fewer false alarms
↓
FPR decreases
BUT
We may miss more actual positives
↓
TPR decreases
Step 6: What Is the ROC Curve?
Now we are ready to understand the ROC curve. Instead of choosing just one threshold, we try many different thresholds.
For example:
Threshold = 0.90
Threshold = 0.80
Threshold = 0.70
Threshold = 0.60
Threshold = 0.50
Threshold = 0.40
...For each threshold, we calculate:
FPR
TPRSo every threshold gives us one pair:
(FPR, TPR)For example:
Threshold 0.80 → (FPR = 0.10, TPR = 0.60)
Threshold 0.60 → (FPR = 0.20, TPR = 0.80)
Threshold 0.40 → (FPR = 0.40, TPR = 0.95)We then plot all these points. Connecting these points gives us the ROC curve.
So the easiest definition is:
The ROC curve shows how TPR and FPR change as we change the classification threshold.
Step 7: Calculating the ROC Curve with Scikit-Learn
Scikit-Learn can calculate these values for us:
from sklearn.metrics import roc_curve, roc_auc_score
y_probs = clf.predict_proba(X_test)
y_probs_positive = y_probs[:, 1]
fpr, tpr, thresholds = roc_curve(
y_test,
y_probs_positive
)Let's understand what is happening.
First:
y_probs = clf.predict_proba(X_test)The model gives us probabilities for each class.
Then:
y_probs_positive = y_probs[:, 1]We take the probability of the positive class.
Finally:
fpr, tpr, thresholds = roc_curve(
y_test,
y_probs_positive
)Scikit-Learn tries different thresholds and calculates the corresponding FPR and TPR values.
This gives us three arrays:
| Array | Meaning |
|---|---|
fpr |
False Positive Rate at each threshold |
tpr |
True Positive Rate at each threshold |
thresholds |
The probability thresholds used |
Each threshold gives us one point on the ROC curve.
Step 8: Plotting the ROC Curve
We can plot the curve like this:
def plot_roc_curve(fpr, tpr):
plt.plot(
fpr,
tpr,
color="orange",
label="ROC Curve"
)
plt.plot(
[0, 1],
[0, 1],
color="darkblue",
linestyle="--",
label="Guessing"
)
plt.xlabel("False Positive Rate (FPR)")
plt.ylabel("True Positive Rate (TPR)")
plt.title("Receiver Operating Characteristic (ROC) Curve")
plt.legend()
plt.show()
plot_roc_curve(fpr, tpr)The graph has:
Y-axis → TPR
X-axis → FPR
The diagonal dashed line represents roughly random guessing.
A good classifier will generally have a curve that moves toward the top-left corner.
Why?
Because we want:
High TPR → correctly find many actual positives
Low FPR → create few false alarmsSo the ideal situation is:
TPR = 1.0
FPR = 0.0
That means:
Find all the positives while making no false-positive mistakes.
Step 9: Summarizing the ROC Curve with AUC
The ROC curve gives us a whole graph. Sometimes we want to summarize that graph with one number. That's where AUC comes in.
AUC means:
Area Under the ROC Curve
We can calculate it with:
roc_auc_score(
y_test,
y_probs_positive
)A simple interpretation is:
1.0 → perfect ranking
0.5 → roughly random ranking
For example:
AUC = 0.90
means the model is generally very good at ranking positive examples above negative examples.
One useful way to understand AUC is:
If we randomly choose one positive example and one negative example, the model will give the positive example a higher score about 90% of the time.
So AUC is measuring how well the model separates and ranks the two classes.
Step 10: Why Is AUC Called Threshold-Independent?
Remember that the ROC curve uses many different thresholds.
Instead of asking:
"How good is my model at threshold
0.50?"
AUC considers the model's performance across the different thresholds used to create the ROC curve. That's why we often describe AUC as threshold-independent. It evaluates how well the model ranks positive examples above negative examples across different thresholds.
⚠️ One Important Limitation
AUC is useful for evaluating ranking quality, but it isn't always enough, especially when the positive class is very rare.
For example, imagine we are detecting a rare disease.
There may be:
Thousands of healthy patients
+
Only a few patients with the diseaseIn this situation, a model can have a good ROC-AUC while still producing many false positives.
That's why, for heavily imbalanced datasets, we may also look at a Precision-Recall curve (PR curve) and PR-AUC.
These focus more directly on the quality of positive predictions.
📌 Confusion Matrix
A confusion matrix helps us understand exactly where our model's predictions were correct and where they were wrong.
Accuracy tells us only the overall percentage of correct predictions. A confusion matrix gives us more detail by showing what type of mistakes the model made.
Import the Required Tools
from sklearn.metrics import confusion_matrix, ConfusionMatrixDisplayCreate a Confusion Matrix
If we already have the actual labels (y_test) and our model's predictions (y_preds), we can create a confusion matrix with:
confusion_matrix(y_test, y_preds)For example, we might get:
[[22, 8],
[ 7, 24]]At first, this might look confusing.
We can understand it by putting the numbers into a table.
Predicted
0 1
Actual 0 TN FP
Actual 1 FN TPThe rows represent the actual labels. The columns represent the predicted labels.
So we can read the matrix like this:
| Cell | Meaning |
|---|---|
[0, 0] |
Actual class 0, predicted class 0 → correct |
[0, 1] |
Actual class 0, predicted class 1 → wrong |
[1, 0] |
Actual class 1, predicted class 0 → wrong |
[1, 1] |
Actual class 1, predicted class 1 → correct |
For our example:
[[22, 8],
[ 7, 24]]We can interpret it as:
Predicted
0 1
Actual 0 22 8
Actual 1 7 24
Therefore:
22 → Actual 0, predicted 0 → Correct
8 → Actual 0, predicted 1 → Wrong
7 → Actual 1, predicted 0 → Wrong
24 → Actual 1, predicted 1 → Correct
If we think of class 1 as the positive class, these become:
22 → True Negative
8 → False Positive
7 → False Negative
24 → True Positive
A More Readable Version with pandas
The basic confusion_matrix() output gives us just the numbers. We can make it easier to read using pandas:
pd.crosstab(
y_test,
y_preds,
rownames=["Actual Labels"],
colnames=["Predicted Labels"]
)This contains the same numbers as confusion_matrix(). The difference is that the row and column labels make it easier to understand which numbers represent the actual labels and which represent the predictions.
Instead of simply seeing:
[[22, 8],
[ 7, 24]]we get a table that clearly tells us:
Predicted Labels
0 1
Actual Labels
0 22 8
1 7 24This is often easier to read when we are exploring our model's performance.
Visualize the Confusion Matrix
We can also display the confusion matrix as a visual table. If we want Scikit-Learn to make predictions and create the confusion matrix for us, we can use:
ConfusionMatrixDisplay.from_estimator(
clf,
X=X,
y=y
)Here, clf is our trained classifier.
Alternatively, if we already have the predictions, we can use:
ConfusionMatrixDisplay.from_predictions(
y_true=y_test,
y_pred=y_preds
)The second version uses:
y_test → actual labels
y_preds → our model's predictions
So both approaches help us see the same type of information, just in a more visual way.
🧠 Why Is a Confusion Matrix Useful?
Imagine two models both have:
Accuracy = 90%Accuracy alone tells us that both models got 90% of the predictions correct. But we don't know what mistakes they made.
For example, one model might produce:
Very few False Positives
Many False Negativeswhile another might produce:
Many False Positives
Very few False NegativesBoth could have similar accuracy, but they behave very differently. The confusion matrix lets us see these mistakes directly.
So remember:
A confusion matrix doesn't just tell us how many predictions were correct. It shows us exactly which predictions were correct, which were wrong, and what type of mistake was made.
📌 Classification Report
Scikit-Learn can calculate several classification metrics at once using classification_report():
from sklearn.metrics import classification_report
print(
classification_report(
y_test,
y_preds
)
)You might see:
precision recall f1-score support
0 0.88 0.70 0.78 30
1 0.76 0.90 0.82 31
accuracy 0.80 61
macro avg 0.82 0.80 0.80 61
weighted avg 0.82 0.80 0.80 61Let's understand each part.
Precision
Precision asks:
When the model says "YES", how often is it correct?
Formula:
Precision = TP / (TP + FP)High precision means the model makes fewer false alarms.
For example, if a model says "spam" 100 times and 80 of those emails are actually spam:
Precision = 80 / 100 = 80%Recall
Recall asks:
Of all the actual positive cases, how many did the model find?
Formula:
Recall = TP / (TP + FN)High recall means the model misses fewer positive cases.
For example, if there are actually 200 spam emails and the model finds 80:
Recall = 80 / 200 = 40%So precision and recall answer different questions:
Precision → "When I say YES, am I usually right?"
Recall → "Did I find most of the actual YES cases?"F1-score
F1 combines precision and recall into one number. It is useful when we care about both precision and recall.
It uses the harmonic mean:
2 × precision × recall
F1 = -----------------------------------------
precision + recallA high F1-score means the model has a good balance between precision and recall.
Support
support is not a performance metric. It simply tells us how many actual examples belong to each class.
For example:
Class 0 → 30 actual examples
Class 1 → 31 actual examplesSo in the report:
0 0.88 0.70 0.78 30
1 0.76 0.90 0.82 31The 30 and 31 are the support values.
Macro Average
macro avg calculates the metric for each class and then gives every class equal importance.
For example, suppose we have three classes:
Class 0 → 90%
Class 1 → 70%
Class 2 → 50%The macro average is:
(90 + 70 + 50) / 3 = 70%This is useful when we want to treat every class equally, regardless of how many examples each class contains.
Weighted Average
weighted avg also calculates the metric for each class, but gives more weight to classes with more examples.
For example:
Class 0 → 1,000 examples
Class 1 → 100 examplesClass 0 will have a much larger influence on the weighted average because it has more examples.
This makes weighted averages useful when we want the class distribution to be reflected in the final number.
Precision vs. Recall
An easy way to remember them:
Precision → "The model said YES." How many were actually YES?
Recall → "These are all the actual YES cases." How many did we find?
Another simple way to think about it:
Precision → Avoid false alarms
Recall → Avoid missing positives
Precision vs. Recall in Real Problems
Which metric matters more depends on which type of mistake is more costly.
| Metric | More important when... | Example |
|---|---|---|
| Recall | Missing a positive case is costly | Disease screening |
| Precision | False positives are costly | Spam filtering |
| F1 | You care about both | General classification |
These are not universal rules. The right metric depends on the actual problem and the cost of different mistakes.
7.4 Regression evaluation metrics
Classification models predict categories, such as:
Disease / No Disease
Cat / Dog
Spam / Not SpamRegression models predict numbers, such as:
House price → $300,000
Temperature → 28°C
Sales → 1,500 unitsBecause regression predicts numbers, we need different metrics to measure how close our predictions are to the actual values.
The three important regression metrics here are:
- R²
- MAE
- MSE
📌 R² (Coefficient of Determination)
R² stands for coefficient of determination.
The easiest way to think about R² is:
How much better is my model than simply predicting the average value every time?
Let's use house prices as an example.
Imagine the average house price in our dataset is:
$300,000A very simple model could predict:
$300,000for every house.
R² compares our actual model against this simple average-value baseline.
You can calculate R² using:
model.score(X_test, y_test)or:
from sklearn.metrics import r2_score
r2_score(y_test, y_preds)Typical interpretation:
| R² | Meaning |
|---|---|
1.0 |
Perfect predictions |
0.8 |
Explains about 80% of the variation in the target |
0.0 |
About as good as always predicting the mean |
< 0 |
Worse than the mean baseline |
For example:
R² = 0.80
roughly means:
The model explains about 80% of the variation in the target values.
A higher R² is generally better, but don't treat one particular value as universally "good" or "bad." What counts as a useful R² depends on the problem and the data.
📌 Mean Absolute Error (MAE)
MAE stands for Mean Absolute Error.
The easiest way to understand MAE is:
On average, how far are my predictions from the actual values?
For example, suppose we have:
Actual: [100, 200, 300]
Predicted: [110, 180, 320]First, calculate the error for each prediction:
Prediction - Actual
110 - 100 = 10
180 - 200 = -20
320 - 300 = 20Some errors are positive, and some are negative. If we simply added them together, they could cancel each other out. So MAE uses the absolute value of each error.
10
20
20Now calculate the average:
(10 + 20 + 20) / 3
= 16.67Therefore:
MAE = 16.67This means:
Our predictions are off by about 16.67 units on average.
Using Scikit-Learn
We can calculate MAE using:
from sklearn.metrics import mean_absolute_error
mean_absolute_error(
y_test,
y_preds
)MAE returns one number representing the average size of our prediction errors.
Why is MAE useful?
One of the biggest advantages of MAE is that it uses the same units as the target.
For example, suppose we are predicting house prices:
MAE = $12,000We can easily understand this as:
"My predictions are off by about $12,000 on average."
That's why MAE is easy to interpret. A lower MAE means our predictions are, on average, closer to the actual values.
Calculate MAE Manually
We can also calculate MAE ourselves using pandas and NumPy:
df = pd.DataFrame({
"Actual": y_test,
"Predicted": y_preds
})
df["Differences"] = (
df["Predicted"] - df["Actual"]
)
np.abs(
df["Differences"]
).mean()This does three simple things:
1. Calculate the difference
Predicted - Actual2. Remove the negative sign
We use:
np.abs()so that an error such as -20 becomes 20.
3. Calculate the average
We use:
.mean()The result should be the same as:
mean_absolute_error(y_test, y_preds)📌 Mean Squared Error (MSE)
MSE stands for Mean Squared Error. MSE is similar to MAE because both measure prediction errors. The main difference is what they do with those errors:
MAE → takes the absolute value
MSE → squares the errorWe can calculate MSE using:
from sklearn.metrics import mean_squared_error
mean_squared_error(
y_test,
y_preds
)Why does MSE square the errors?
The important reason is:
Squaring makes large errors much more important than small errors.
For example:
Error = 2
Squared = 2² = 4But:
Error = 20
Squared = 20² = 400The original error is:
20 / 2 = 10So the error of 20 is 10 times larger than the error of 2.
But after squaring:
400 / 4 = 100The squared error is 100 times larger. This is why MSE strongly penalizes big mistakes.
MAE vs. MSE
Suppose our prediction errors are:
Errors = [1, 1, 1, 20]Let's calculate MAE first.
MAE
MAE takes the absolute value of each error and calculates the average:
(1 + 1 + 1 + 20) / 4
= 5.75So:
MAE = 5.75Now let's calculate MSE.
MSE
MSE squares each error first:
(1² + 1² + 1² + 20²) / 4
= 100.75So:
MSE = 100.75Notice the difference:
MAE → 5.75
MSE → 100.75The large error of 20 has a much bigger influence on MSE because it gets squared.
🧠 Easy way to remember MAE vs. MSE
Think of it this way:
MAE
↓
Treats errors more evenly
↓
Easy to understand
MSE
↓
Squares errors
↓
Large errors become much more important
So:
- Use MAE when you want an error measure that is easy to understand and less affected by very large errors.
- Use MSE when you want large errors to be penalized much more heavily.
One More Difference: The Units
MAE is easier to interpret because it stays in the same units as the thing we are predicting.
For example, suppose we are predicting house prices:
Actual price → $200,000
Predicted → $190,000The error is:
$200,000 - $190,000 = $10,000MAE is also measured in dollars.
For example:
MAE = $10,000We can say:
"The model is off by about $10,000 on average."
MSE is different.
Because MSE squares the errors, its result is in squared units.
For example:
Error = $10,000
MSE contribution = $10,000²
= $100,000,000So the MSE number is much harder to interpret directly.
🧠 Easy way to remember
R²
→ How well does my model explain the target compared
with simply predicting the average?
MAE
→ How far off are my predictions on average?
MSE
→ How large are my errors, especially the large ones?Or even more simply:
R² → How well does the model explain the data?
MAE → How far off are my predictions?
MSE → How much should I punish large mistakes?
For understanding prediction errors, MAE is usually easier to explain. MSE is useful when we specifically want large errors to have a much bigger impact.
7.5 Combining scoring with cross-validation
You can use different scoring metrics with cross-validation.
from sklearn.model_selection import cross_val_score
# Classification
cv_acc = cross_val_score(clf, X, y, cv=5, scoring="accuracy")
cv_precision = cross_val_score(clf, X, y, cv=5, scoring="precision")
cv_recall = cross_val_score(clf, X, y, cv=5, scoring="recall")
# Regression (note the "neg_" prefix — Scikit-Learn's convention is that
# higher scores are always better internally, so error metrics are negated)
cv_mae = cross_val_score(model, X, y, cv=3, scoring="neg_mean_absolute_error")
cv_mse = cross_val_score(model, X, y, cv=3, scoring="neg_mean_squared_error")7.6 Using sklearn.metrics functions directly
The third evaluation approach: call metric functions directly on your predictions, without cross-validation.
Classification
from sklearn.metrics import accuracy_score, precision_score, recall_score, f1_score
y_preds = clf.predict(X_test)
print(f"Accuracy: {accuracy_score(y_test, y_preds) * 100:.2f}%")
print(f"Precision: {precision_score(y_test, y_preds)}")
print(f"Recall: {recall_score(y_test, y_preds)}")
print(f"F1: {f1_score(y_test, y_preds)}")Regression
from sklearn.metrics import r2_score, mean_absolute_error, mean_squared_error
y_preds = model.predict(X_test)
print(f"R2 score: {r2_score(y_test, y_preds)}")
print(f"MAE: {mean_absolute_error(y_test, y_preds)}")
print(f"MSE: {mean_squared_error(y_test, y_preds)}")7.7 Bonus: Feature Importance
After training a model, you may naturally wonder:
"Which features were most useful to the model?"
For many tree-based models, Scikit-Learn provides feature importance scores directly.
Extract and sort feature importances
feature_importances = pd.Series(
model.feature_importances_,
index=X.columns
).sort_values(ascending=False)
feature_importances.head(10).plot.bar(figsize=(10, 5))What's happening here?
model.feature_importances_→ gives an importance score for each feature.index=X.columns→ gives those scores their feature names..sort_values(ascending=False)→ puts the most important features first..head(10)→ selects the top 10 features..plot.bar()→ displays them as a bar chart.figsize=(10, 5)→ makes the chart 10 inches wide and 5 inches tall.
You might get something like:
Feature A ████████████████ 0.35
Feature B ███████████ 0.24
Feature C ████████ 0.18
Feature D █████ 0.12
Feature E ███ 0.06
This gives you a quick idea of which features the model relied on most.
Why is this useful?
Feature importance can help you:
- Understand what the model is using.
- Identify potentially useful features.
- Decide where collecting better data might be valuable.
- Get a general idea of which features have less influence on the model.
⚠️ Important limitations
Feature importance is useful, but don't treat it as a perfect explanation of the model.
1. Correlated features can be misleading
Suppose you have:
house_size
number_of_roomsThese features may be strongly related. A tree-based model may give more importance to one and less to the other, even though both contain useful information.
2. It doesn't tell you the direction
An importance score tells you how useful a feature was, but not whether increasing that feature makes the prediction go up or down.
For example:
Feature importance = 0.30
doesn't tell us whether:
More feature → higher prediction
or:
More feature → lower prediction
For more detailed explanations, techniques such as SHAP can help.
🧠 Easy way to remember
Feature importance tells you which features the model relied on most, but it doesn't fully explain how those features affect the prediction.
8. Improving a Model
Your first model is usually called a baseline. A baseline is your starting point—the model's initial performance that you use as a reference to measure future improvements against.
For example:
Baseline accuracy → 78%Now, whenever you change something, you can compare the new result with 78%.
Ways to improve, from two different angles
| From a data perspective | From a model perspective |
|---|---|
| Collect more data | Try a different, more powerful model |
| Improve/clean the existing data | Tune the current model's hyperparameters |
ML development is usually an experiment loop:
Baseline
↓
Measure performance
↓
Change one thing
↓
Train the model
↓
Measure performance again
↓
Compare with the baseline
↓
Keep the change if it helps
↓
RepeatThe important idea is:
Change something, measure the result, and compare it with your previous result.
8.1 Parameters vs. Hyperparameters
This distinction trips up a lot of beginners:
| Parameters | Hyperparameters | |
|---|---|---|
| What they are | Values the model learns automatically from the data during .fit() (e.g., the splits/thresholds inside a decision tree) |
Settings you choose before training that control how the model learns (e.g., how many trees in a forest, how deep each tree can grow) |
| Set by | The learning algorithm | You |
| Tuning? | You never set these by hand | These are what you "tune" |
A simple way to remember it:
Parameters are learned by the model. Hyperparameters are chosen by you.
Create a model and inspect its settings:
clf = RandomForestClassifier()clf.get_params()This prints the model's tunable hyperparameters and their current values — n_estimators, max_depth, max_features, and many more.
🔹 Three ways to search for good hyperparameters
- By hand → change one value, retrain, compare.
- RandomizedSearchCV → randomly try a fixed number of combinations.
- GridSearchCV → exhaustively try every combination in a grid.
🔹 Common hyperparameters to tune on a Random Forest
| Hyperparameter | What it controls |
|---|---|
n_estimators |
Number of trees |
max_depth |
Maximum depth of each tree |
max_features |
Number of features considered when splitting a node |
min_samples_split |
Minimum samples required to split a node |
min_samples_leaf |
Minimum samples allowed in a leaf node |
8.2 Tuning by hand (using train / validation / test splits)
When you're comparing several hand-picked models or tuning their hyperparameters, it's best practice to use three splits instead of two:
Full Dataset
┌────────────┼────────────┐
▼ ▼ ▼
Train Validation Test
(70%) (15%) (15%)
│ │ │
│ │ └─ ▶ score ONCE, at the very end
│ └───────────────▶ tune here, repeatedly
└────────────────────────────▶ model learns from this
- Training set → the model uses this data to learn patterns and relationships.
- Validation set → you use this data to compare models and tune hyperparameters.
- Test set → you use this data only once at the very end to measure the model's final performance.
❓ Why not tune using the test set?
If you keep checking the test score while changing your model or hyperparameters, you're indirectly using the test data to make decisions.
Over time, you may end up choosing the model that performs best on that specific test set, instead of the model that actually generalizes best to new, unseen data.
The validation set gives you a separate dataset where you can experiment, compare, and tune your model while keeping the test set completely untouched.
Think of the test set as your final exam. You can use the validation set to practice, but don't study from the final exam.
First, define a helper function to keep things tidy:
def evaluate_preds(y_true, y_preds):
"""Compare predictions to truth for a classifier and print/report metrics."""
accuracy = accuracy_score(y_true, y_preds)
precision = precision_score(y_true, y_preds)
recall = recall_score(y_true, y_preds)
f1 = f1_score(y_true, y_preds)
metric_dict = {
"accuracy": round(accuracy, 2),
"precision": round(precision, 2),
"recall": round(recall, 2),
"f1": round(f1, 2)
}
print(f"Accuracy: {accuracy * 100:.2f}%")
print(f"Precision: {precision:.2f}")
print(f"Recall: {recall:.2f}")
print(f"F1 score: {f1:.2f}")
return metric_dictThis function computes four metrics, prints them nicely, and returns them in a dictionary so we can compare models side by side later.
Now shuffle the dataset so the split isn't biased by row order:
np.random.seed(42)
heart_disease_shuffled = heart_disease.sample(frac=1)np.random.seed(42) makes the shuffle reproducible. sample(frac=1) returns the whole DataFrame in a random order.
Define X and y:
X = heart_disease_shuffled.drop("target", axis=1)
y = heart_disease_shuffled["target"]Compute the split points:
train_split = round(0.7 * len(heart_disease_shuffled))
valid_split = round(train_split + 0.15 * len(heart_disease_shuffled))train_split— index where the training set ends (70%).valid_split— index where the validation set ends (70% + 15% = 85%).
Slice into three sets:
X_train, y_train = X[:train_split], y[:train_split]
X_valid, y_valid = X[train_split:valid_split], y[train_split:valid_split]
X_test, y_test = X[valid_split:], y[valid_split:]Now train a baseline:
clf = RandomForestClassifier(n_estimators=100)
clf.fit(X_train, y_train)100 trees, trained on the training set.
Evaluate on validation:
baseline_metrics = evaluate_preds(y_valid, clf.predict(X_valid))Now try a tweaked version — limit tree depth:
clf_2 = RandomForestClassifier(n_estimators=100, max_depth=10)
clf_2.fit(X_train, y_train)max_depth=10 stops trees from growing too deep — a common overfitting guard.
Evaluate it:
clf_2_metrics = evaluate_preds(y_valid, clf_2.predict(X_valid))Compare the two metric dictionaries. Whichever performs better on validation is the one you'd move forward with.
📝 Note: When you use
GridSearchCVorRandomizedSearchCV, the validation process is handled internally through cross-validation, so you don't need a separate validation set. The train/validation/test split above is for when you're comparing hand-picked models.
8.3 RandomizedSearchCV
Instead of hand-picking values, define a range of options and let Scikit-Learn randomly sample a fixed number of combinations, testing each with cross-validation.
Import the tools:
from sklearn.model_selection import RandomizedSearchCV, train_test_splitDefine the search space:
grid = {
"n_estimators": [10, 100, 200, 500, 1000, 1200],
"max_depth": [None, 5, 10, 20, 30],
"max_features": ["sqrt", "log2"],
"min_samples_split": [2, 4, 6],
"min_samples_leaf": [1, 2, 4]
}That's 6 × 5 × 2 × 3 × 3 = 540 possible combinations. RandomizedSearchCV won't try them all — it'll sample.
Prepare the data (same as before):
heart_disease_shuffled = heart_disease.sample(frac=1, random_state=42)
X = heart_disease_shuffled.drop("target", axis=1)
y = heart_disease_shuffled["target"]
X_train, X_test, y_train, y_test = train_test_split(
X,
y,
test_size=0.2,
random_state=42
)Create the model:
clf = RandomForestClassifier(n_jobs=1, random_state=42)Set up the search:
rs_clf = RandomizedSearchCV(
estimator=clf,
param_distributions=grid,
n_iter=10, # try 10 random combinations
cv=5, # 5-fold cross-validation for each combination
verbose=2,
random_state=42,
refit=True # automatically retrain best combo on the FULL training set
)Here's what the important settings mean:
estimator=clf→ the model we want to tuneparam_distributions=grid→ the values we want to tryn_iter=10→ try 10 random combinationscv=5→ 5-fold cross-validation for each combination.verbose=2→ print progress while it runs.random_state=42→ make the random search reproduciblerefit=True→ after finding the best params, retrain on the full training set.
Run the search:
rs_clf.fit(X_train, y_train)Find the best settings:
rs_clf.best_params_Predict and evaluate:
rs_y_preds = rs_clf.predict(X_test)
rs_metrics = evaluate_preds(y_test, rs_y_preds)💡
n_iter=10means: randomly try 10 parameter combinations from the distributions you supplied. If there are 500 possible combinations, you only evaluate 10 — which can save substantial computation.
💡
n_jobs: number of CPU cores to use.-1means "use every available core" — much faster on multi-core machines, but leaves your computer less responsive for other tasks while it runs.
8.4 GridSearchCV
Once RandomizedSearchCV has pointed you toward a promising region of hyperparameters, narrow your grid and search it exhaustively — every single combination, guaranteed.
Import:
from sklearn.model_selection import GridSearchCVDefine a smaller, focused grid:
grid_2 = {
"n_estimators": [100, 200, 500],
"max_depth": [None],
"max_features": ["sqrt", "log2"],
"min_samples_split": [2],
"min_samples_leaf": [2]
}3 × 1 × 2 × 1 × 1 = 6 combinations. Small enough to try them all.
Create the model:
clf = RandomForestClassifier(n_jobs=1, random_state=42)Set up the search:
gs_clf = GridSearchCV(
estimator=clf,
param_grid=grid_2,
cv=5,
verbose=2
)Note there's no n_iter here — GridSearchCV tries every combination.
Fit:
gs_clf.fit(X_train, y_train)Get the best settings:
gs_clf.best_params_Predict and evaluate:
gs_y_preds = gs_clf.predict(X_test)
gs_metrics = evaluate_preds(y_test, gs_y_preds)🧠 RandomizedSearchCV vs. GridSearchCV — when to use which
| RandomizedSearchCV | GridSearchCV | |
|---|---|---|
| How it searches | Random combinations | Every combination |
| Number of attempts | Controlled by n_iter |
Determined by grid size |
| Best for | Large search spaces | Small focused searches |
| Guarantees | Finds a good combo | Finds the best combo in the grid |
Common workflow: RandomizedSearchCV broadly first → then GridSearchCV on a narrowed-down grid.
8.5 Comparing everything you've tried
Build a DataFrame from the metric dictionaries:
compare_metrics = pd.DataFrame({
"baseline": baseline_metrics,
"clf_2": clf_2_metrics,
"random_search": rs_metrics,
"grid_search": gs_metrics
})Each column is one experiment, each row is one metric.
Plot it:
compare_metrics.plot.bar(figsize=(10, 8))A grouped bar chart. It immediately shows which of your experiments actually helped — and by how much.
💡 A crucial reminder: Always evaluate all experiments on the same test set. If you evaluate different models on different splits, your comparison is meaningless.
9. Saving and Loading a Trained Model
Training a model can take minutes, hours, or even days. You don't want to train it again every time you restart your program. Instead, you can save ("persist") the trained model to disk and load it again whenever you need it.
9.1 pickle (Python's built-in serialization module)
import pickle
# Save
pickle.dump(gs_clf, open("gs_random_forest_model_1.pkl", "wb"))
# Load
loaded_pickle_model = pickle.load(open("gs_random_forest_model_1.pkl", "rb"))
# Use it exactly like the original
pickle_y_preds = loaded_pickle_model.predict(X_test)
evaluate_preds(y_test, pickle_y_preds)The "wb" and "rb" mean write-binary and read-binary. Since pickle files are stored in binary format, you need "wb" when saving and "rb" when loading.
9.2 joblib (usually preferred for Scikit-Learn models)
from joblib import dump, load
# Save
dump(gs_clf, filename="gs_random_forest_model_2.joblib")
# Load
loaded_joblib_model = load(filename="gs_random_forest_model_2.joblib")
joblib_y_preds = loaded_joblib_model.predict(X_test)
evaluate_preds(y_test, joblib_y_preds)🧠
picklevs.joblib: Both can save and load trained models. For Scikit-Learn models,joblibis usually preferred, especially for models containing large NumPy arrays.
🔒 Security note: Only load model files from trusted sources. Pickle and joblib files can execute code when loaded.
📝 Version compatibility: Saved models may depend on the versions of Python and Scikit-Learn used to create them. Keep track of your library versions so you can load the model reliably later.
10. Putting It All Together (Pipelines)
Up until now, we've done every preprocessing step by hand — imputing, encoding, splitting, and fitting as separate operations. This works for learning, but in a real project it becomes:
- Error-prone — you have to remember the right order every time.
- Repetitive — the same preprocessing has to run on training and test data.
- Dangerous — easy to accidentally leak information from test data into training.
Scikit-Learn's Pipeline solves all three problems by chaining every step — imputation, encoding, scaling, and the final model — into a single object that behaves like any other estimator.
10.1 The Problem Pipelines Solve
Suppose your data has:
- Missing values in categorical and numeric columns.
- Text categories that need one-hot encoding.
- A regression target (
Price) you want to predict.
Doing it by hand means:
- Fit the imputer on training data → transform both train and test.
- Fit the encoder on training data → transform both train and test.
- Then train the model.
- Repeat this whole sequence every time you tweak a parameter. 😩
With a Pipeline, you write it once, and it handles all of that automatically — including during cross-validation and hyperparameter tuning.
🧠 Pipeline mental model
Think of a pipeline like an assembly line:
Raw Data
↓
[Fill Missing Values]
↓
[Encode Categories]
↓
[Model]
↓
Prediction
You give the pipeline raw data. It handles the process.
10.2 Building the Pipeline — Step by Step
🔹 Step 1: Load the data
import pandas as pd
import numpy as np
data = pd.read_csv("sklearn-car-sales-missing-data.csv")
data
data.dtypes
data.isna().sum()At this point, we have three things we want to do:
- Fill missing data
- Convert data to numbers
- Build a model on the data
Rather than doing them one at a time, we'll chain them together with a Pipeline.
🔹 Step 2: Set up imports and reproducibility
# ============================================================
# Imports
# ============================================================
import pandas as pd
import numpy as np
from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline
from sklearn.impute import SimpleImputer
from sklearn.preprocessing import OneHotEncoder
# Modelling
from sklearn.ensemble import RandomForestRegressor
from sklearn.model_selection import train_test_split, GridSearchCV
# ============================================================
# Setup
# ============================================================
# Make random operations reproducible
np.random.seed(42)🔹 Step 3: Load data and define features
# ============================================================
# Load Data
# ============================================================
# Load the car sales dataset
data = pd.read_csv("sklearn-car-sales-missing-data.csv")
# Price is our target, so we cannot train on rows
# where the target value is missing
data.dropna(subset=["Price"], inplace=True)
# ============================================================
# Define Features
# ============================================================
# Categorical features contain text/category values
categorical_features = ["Make", "Colour"]
# Doors is numerical, but requires a specific
# strategy for handling missing values
door_feature = ["Doors"]
# Numerical features contain continuous numeric values
numeric_features = ["Odometer (KM)"]💡 Why does
Doorsget its own category? Because we want to fill missing door values with a specific constant (4) rather than the mean or mode. It behaves differently from both truly categorical and truly continuous columns, so it gets its own transformer.
🔹 Step 4: Build preprocessing sub-pipelines
Instead of applying each transformation manually, we nest mini-pipelines inside a ColumnTransformer. Each mini-pipeline says: "for these columns, do these steps in this order."
# ============================================================
# Preprocessing
# ============================================================
# For categorical data:
# 1. Fill missing values with "missing"
# 2. Convert categories into numerical columns
categorical_transformer = Pipeline(
steps=[
(
"imputer",
SimpleImputer(
strategy="constant",
fill_value="missing"
)
),
(
"onehot",
OneHotEncoder(handle_unknown="ignore")
)
]
)
# For Doors:
# Replace missing values with 4
door_transformer = Pipeline(
steps=[
(
"imputer",
SimpleImputer(
strategy="constant",
fill_value=4
)
)
]
)
# For numerical data:
# Replace missing values with the mean value
numeric_transformer = Pipeline(
steps=[
(
"imputer",
SimpleImputer(strategy="mean")
)
]
)🧠 What does
handle_unknown="ignore"do? If the test set contains a category (e.g., a car make) that the encoder never saw during training,OneHotEncoderwould normally throw an error.handle_unknown="ignore"tells it to just produce a row of zeros for that unseen category instead of crashing. This is critical for production systems where new categories appear over time.
🔹 Step 5: Combine the sub-pipelines with ColumnTransformer
# ============================================================
# Combine Preprocessing
# ============================================================
# Apply the appropriate preprocessing pipeline
# to each group of features
preprocessor = ColumnTransformer(
transformers=[
("cat", categorical_transformer, categorical_features),
("door", door_transformer, door_feature),
("num", numeric_transformer, numeric_features)
]
)At this point, preprocessor is a single object that knows how to:
- Fill missing
Make/Colourvalues with"missing". - One-hot encode them.
- Fill missing
Doorswith4. - Fill missing
Odometer (KM)with the mean.
And it applies the right transformation to the right columns — automatically.
🔹 Step 6: Attach the model to the pipeline
Now we chain the preprocessor and the model together:
# ============================================================
# Create Model Pipeline
# ============================================================
# The pipeline first preprocesses the data,
# then passes the processed data to the model
model = Pipeline(
steps=[
("preprocessor", preprocessor),
("model", RandomForestRegressor())
]
)🧠 This is the key insight.
modelis now a single object. When you call.fit(), it:
- Fits the imputers on the training data.
- Fits the encoder on the training data.
- Transforms the data.
- Trains the
RandomForestRegressoron the transformed data.And when you call
.predict(), it applies the same transformations — using the learned parameters from training — before making predictions.
🔹 Step 7: Split, fit, score
The pipeline behaves exactly like a normal estimator. The code looks identical to what you'd write for a bare model:
# ============================================================
# Split Features and Target
# ============================================================
# X contains the input features
X = data.drop("Price", axis=1)
# y contains the value we want to predict
y = data["Price"]
# Split the data into training and testing sets
# 80% for training and 20% for testing
X_train, X_test, y_train, y_test = train_test_split(
X,
y,
test_size=0.2
)
# ============================================================
# Train Model
# ============================================================
# The pipeline handles preprocessing and model training
model.fit(X_train, y_train)
# ============================================================
# Evaluate Model
# ============================================================
# Evaluate the model using the test data
# For regression, .score() returns the R² score
model.score(X_test, y_test)🎉 Notice how clean that is. No manual imputation, no manual encoding, no chance of getting the order wrong. The pipeline handles it all.
🔹 Step 8: Hyperparameter tuning on the pipeline
Here's where pipelines really shine: you can tune both the preprocessing and the model simultaneously — in a single GridSearchCV.
The notation uses double underscores __ to drill into the pipeline structure:
# ============================================================
# Hyperparameter Tuning with GridSearchCV
# ============================================================
from sklearn.model_selection import GridSearchCV
# Define the hyperparameters we want to test.
#
# The "__" notation allows us to access steps inside
# our Pipeline and ColumnTransformer.
pipe_grid = {
# Try different strategies for filling missing
# numerical values
"preprocessor__num__imputer__strategy": [
"mean",
"median"
],
# Number of trees in the Random Forest
"model__n_estimators": [
100,
1000
],
# Maximum depth of each decision tree
# None means the tree can grow until stopping criteria
# are reached
"model__max_depth": [
None,
5
],
# Number of features considered when splitting
# each node
"model__max_features": [
"sqrt"
],
# Minimum number of samples required to split a node
"model__min_samples_split": [
2,
4
]
}
# GridSearchCV tests every combination of the
# hyperparameters defined above.
#
# cv=5 means 5-fold cross-validation:
# The training data is divided into 5 parts and the model
# is trained/validated multiple times using different parts.
gs_model = GridSearchCV(
model,
pipe_grid,
cv=5,
# Show progress while GridSearchCV is running
verbose=2,
# Automatically train the best combination on
# the entire training dataset after the search
refit=True
)
# Run the hyperparameter search using the training data.
gs_model.fit(X_train, y_train)
gs_model.score(X_test, y_test)🧠 How to read
"preprocessor__num__imputer__strategy":preprocessor → the first step in our outer Pipeline __num → the "num" transformer in our ColumnTransformer __imputer → the "imputer" step inside the numeric sub-pipeline __strategy → the hyperparameter we want to tuneEach
__is one level down the nested pipeline. If you can read the pipeline structure, you can read the hyperparameter name.
💡 Why is this powerful? You're letting
GridSearchCVdecide whethermeanormedianimputation works better for this specific model on this specific data — without writing a single line of manual comparison code.
10.3 Why Use a Pipeline?
There are several important reasons.
| Benefit | What it means |
|---|---|
| Less data leakage | Preprocessing is learned within each training process |
| One object to save | Preprocessing + model stay together |
| Cleaner tuning | Search preprocessing and model settings together |
| Reusable | New raw data can go directly into .predict() |
| Less manual code | The pipeline handles the repeated steps |
The biggest idea is:
The preprocessing needed by the model becomes part of the model workflow.
Instead of saving:
imputer
encoder
modelseparately, you can save the whole pipeline.
raw data
↓
pipeline
↓
predictionThis is one of the most useful habits to develop when moving from notebook experiments to real ML applications.
🌟 This is genuinely how experienced practitioners structure real projects. Everything you learned in isolation in Sections 3–8 slots into a
Pipelineas soon as a project grows past a quick notebook experiment.
11. Common Errors & How to Fix Them
Beginners don't get stuck on the happy path — they get stuck on tracebacks. Here are the most common ones you'll hit, and what they actually mean.
| Error message | What it means | Fix |
|---|---|---|
ValueError: could not convert string to float: 'Nissan' |
You tried to fit a model on text data | Encode categorical values |
ValueError: Input contains NaN, infinity or a value too large |
Your data has missing values | Impute or drop NAs (SimpleImputer, dropna()) |
ValueError: Found unknown categories [...] in column 0 during transform |
Test data has a category the encoder never saw | Use OneHotEncoder(handle_unknown="ignore") |
AttributeError: 'RandomForestRegressor' object has no attribute 'predict_proba' |
You tried predict_proba on a regressor |
Regression returns a number — probabilities don't apply |
ValueError: Unknown label type: 'continuous' |
You tried a classifier on a regression target | Use a regressor (RandomForestRegressor, Ridge) |
ConvergenceWarning: lbfgs failed to converge |
LinearSVC/LogisticRegression didn't finish training |
Increase max_iter, and scale your features |
| Train score = 0.99, test score = 0.60 | Possible overfitting | Simplify/tune model or improve data |
| Train score = 0.55, test score = 0.53 | Possible underfitting | Try a more suitable/complex model or better features |
| Accuracy very high but model misses minority class | Class imbalance | Check precision, recall and F1 |
🛠️ Handling Class Imbalance
If you spot imbalance (one class dominates), you have three main options:
Option 1: tell the model to weight classes:
RandomForestClassifier(class_weight="balanced")This makes the model penalize mistakes on the minority class more heavily.
Option 2: stratified k-fold:
from sklearn.model_selection import StratifiedKFold
cross_val_score(
clf,
X,
y,
cv=StratifiedKFold(
n_splits=5
)
)StratifiedKFold keeps the class ratios intact in every fold — critical when one class is rare.
Option 3: oversample the minority class (requires the imbalanced-learn package):
Another approach is to increase the number of minority-class examples. One popular technique is SMOTE. It is available through the imbalanced-learn package.
# from imblearn.over_sampling import SMOTESMOTE creates synthetic minority-class examples based on existing data. When using techniques like SMOTE, apply them only to the training data, not the test data.
12. Quick-Reference Cheat Sheet
The workflow, one more time
1. Get data ready → train_test_split, SimpleImputer, OneHotEncoder, StandardScaler
2. Pick a model → RandomForestClassifier / RandomForestRegressor / Ridge / LinearSVC ...
3. Fit → model.fit(X_train, y_train)
4. Predict → model.predict(X_test) / model.predict_proba(X_test)
5. Evaluate → model.score(), cross_val_score(), sklearn.metrics functions
6. Improve → RandomizedSearchCV → GridSearchCV
7. Save / load → pickle.dump/load or joblib.dump/loadClassification vs. Regression
| Classification | Regression | |
|---|---|---|
| Output | Category | Number |
| Example | Spam / Not Spam | House Price |
| Example model | RandomForestClassifier |
RandomForestRegressor |
.predict() |
Class | Number |
| Common metrics | Accuracy, Precision, Recall, F1, ROC-AUC | R², MAE, MSE, RMSE |
| Confusion Matrix | Yes | No |
predict_proba() |
Often available | Usually not applicable |
Classification metrics at a glance
| Metric | Import | Question it answers |
|---|---|---|
| Accuracy | accuracy_score |
What % of predictions were correct overall? |
| Precision | precision_score |
Of predicted positives, how many were right? |
| Recall | recall_score |
Of actual positives, how many did we catch? |
| F1 | f1_score |
Balanced score of precision & recall |
| ROC AUC | roc_auc_score |
How well does the model rank positive examples above negative ones? |
| Confusion Matrix | confusion_matrix |
Where exactly are predictions going wrong? |
Regression metrics at a glance
| Metric | Import | Question it answers |
|---|---|---|
| R² | r2_score |
How does the model compare with predicting the mean? |
| MAE | mean_absolute_error |
On average, how far off are predictions (same units as target)? |
| MSE | mean_squared_error |
How large are the errors, with extra weight on large mistakes? |
The practical checklist
When starting a new ML project:
- Understand the problem → am I predicting a category or a number?
- Load the data →
pd.read_csv(...) - Inspect →
data.head(),data.info(),data.describe(),data.isna().sum() - Define X and y →
X = data.drop("target", axis=1),y = data["target"] - Split →
train_test_split(X, y, test_size=0.2, random_state=42) - Identify preprocessing needs → missing values? categories? scaling?
- Build preprocessing →
SimpleImputer,OneHotEncoder,StandardScaler,ColumnTransformer,Pipeline - Choose a baseline model →
RandomForestClassifier()orRandomForestRegressor() - Fit →
model.fit(X_train, y_train) - Predict →
y_preds = model.predict(X_test) - Evaluate → task-appropriate metrics
- Cross-validate →
cross_val_score(model, X, y, cv=5) - Improve → data / features / preprocessing/algorithms
- Tune →
RandomizedSearchCV→GridSearchCV - Final evaluation → held-out test set
- Save →
joblib.dump(model, "model.joblib")
🧠 Mental models to remember above everything else
- Never evaluate on data the model was trained on → that's like grading your own exam with the answer key open.
- Accuracy lies on imbalanced datasets → always check precision/recall too.
- Parameters are learned; hyperparameters are chosen by you and tuned.
- A validation set lets you tune freely; the test set gets touched exactly once, at the end.
Pipelineisn't just tidier code; it actively prevents data-leakage bugs.- If train >> test, you've overfit. If train ≈ test but both are bad, you've underfit.
predict_probatells you how confident the model is.predictthrows that away.- Use double underscores
__to reach hyperparameters inside aPipeline.
📚 Official documentation
- User guide: https://scikit-learn.org/stable/user_guide.html
- Model selection map: https://scikit-learn.org/stable/machine_learning_map.html
- Metrics reference: https://scikit-learn.org/stable/modules/model_evaluation.html
- Pipelines: https://scikit-learn.org/stable/modules/compose.html
🎉 You now have a complete mental model of the Scikit-Learn workflow. From raw data → to a trained, tuned, saved, and deployed model — every step is here. The pattern
fit → predict → evaluate → improveis the heartbeat of every machine learning project you'll ever build.And when you're ready to move from a quick notebook experiment to a real project, reach for a
Pipeline— it's the difference between "it worked on my machine" and "it works everywhere."