Intel Image Classification using PyTorch

Gajendra Saraswat, 28 August 2020

I will be trying to make this notebook very beginner friendly, if you find anything that should be changed or added, please leave a comment!

Now, Let's make this fun!

Importing Essential Libraries

In [1]:
import numpy as np
import pandas as pd
import os

from PIL import Image

import torch
import torch.nn as nn
import torch.utils.data as D
from torch.optim.lr_scheduler import ExponentialLR

from torchvision import models, transforms as T

from tqdm import tqdm_notebook

from glob import glob
from sklearn.preprocessing import LabelEncoder
from sklearn.model_selection import train_test_split

import matplotlib.pyplot as plt
import seaborn as sns
sns.set(style = "darkgrid")
%matplotlib inline

import warnings
warnings.filterwarnings("ignore")

from ignite.engine import Events, create_supervised_evaluator, create_supervised_trainer
from ignite.metrics import Loss, Accuracy
from ignite.contrib.handlers.tqdm_logger import ProgressBar
from ignite.handlers import  EarlyStopping, ModelCheckpoint

Setting parameters

In [2]:
# Set Seed

def set_seed(seed = 1234):
    '''Sets the seed of the entire notebook so results are the same every time we run.
    This is for REPRODUCIBILITY.'''
    np.random.seed(seed)
#     random.seed(seed)
    torch.manual_seed(seed)
    if torch.cuda.is_available(): 
        torch.cuda.manual_seed(seed)
        torch.cuda.manual_seed_all(seed) # gpu vars
    if torch.backends.cudnn.is_available:
        torch.backends.cudnn.deterministic = True
        torch.backends.cudnn.benchmark = False
    # Set a fixed value for the hash seed
    os.environ['PYTHONHASHSEED'] = str(seed)
    
set_seed(42)
In [3]:
BASE_PATH = "/kaggle/input/intel-image-classification"
device = torch.device("cuda:0" if torch.cuda.is_available() else "cpu")
BATCH_SIZE = 32

Getting and Preprocessing data

Defining necessary functions

In [4]:
def get_data(repo):
    image_path_dict = {os.path.basename(x).split(".")[0] : x for x in glob(os.path.join(BASE_PATH, repo, repo, "*", "*.jpg"))}
    df = pd.DataFrame()
    image_id = [os.path.basename(x).split(".")[0] for x in glob(os.path.join(BASE_PATH, repo, repo, "*", "*.jpg"))] 
    df["image_id"] = image_id
    df["path"] = df["image_id"].map(image_path_dict.get)
    df["target"] = df.path.apply(lambda x: x.split("/")[-2])
    return df
In [5]:
def label_encode():
    df = pd.concat([df_train, df_test])
    le = LabelEncoder()
    le.fit(df["target"])
    df_train["label"] = le.transform(df_train["target"])
    df_test["label"] = le.transform(df_test["target"])
    print(le.classes_)
    print("Encoding Successful")
    return le.classes_

Let's define our Image preparing pipeline which will take jpg image as input and after preprocessing and doing some augmentations on it, will provide us with augmented PIL image.

In [6]:
class ImagesDS(D.Dataset):
    def __init__(self, df, img_dir, mode = 'seg_train'):
        self.records = df.to_records(index = False)
        self.mode = mode
        self.img_dir = img_dir
        self.len = df.shape[0]
        
    @staticmethod
    def _load_img_as_tensor(file_name):
        with Image.open(file_name) as img:
            return T.Compose([T.RandomResizedCrop(size = 150, scale = (0.7, 1.0)),
                              T.RandomHorizontalFlip(),
                              T.ToTensor(), 
                              T.Normalize(mean = [0.485, 0.456, 0.406], std = [0.229, 0.224, 0.225])])(img)

    def _get_img_path(self, index):
        image_id, target = self.records[index].image_id, self.records[index].target
        return '/'.join([self.img_dir,self.mode, self.mode, target, f"{image_id}.jpg"])
        
    def __getitem__(self, index):
        path = self._get_img_path(index)
        img = self._load_img_as_tensor(path)
        if self.mode == 'seg_train':
            return img, int(self.records[index].label)
        else:
            return img, self.records[index].image_id

    def __len__(self):
        return self.len

Let's get the training and test data.

In [7]:
df_train = get_data("seg_train")
df_test = get_data("seg_test")

Encoding the target classes.

In [8]:
encodings = label_encode()
['buildings' 'forest' 'glacier' 'mountain' 'sea' 'street']
Encoding Successful

Let's see the distribution of data over all the 6 classes we have!

In [9]:
plt.rcParams["figure.figsize"] = [10, 8]
df_train.target.value_counts().plot(kind = "bar", color = ["B", "G", "R", "Y", "Orange", "Violet"])
Out[9]:
<matplotlib.axes._subplots.AxesSubplot at 0x7f7f67eaac10>

Okay, seems like the data is almost balanced.

Now, let's take a sneak peek into some of the images from the the data with their labels.

In [10]:
temp = df_train.sample(9)

fig=plt.figure(figsize=(10, 10))
columns = 3
rows = 3
for i in range(1, columns*rows +1):
    img = plt.imread(temp.path.iloc[i-1])
    fig.add_subplot(rows, columns, i)
    plt.imshow(img)
    plt.xticks([])
    plt.yticks([])
    label = temp.target.iloc[i-1]
    plt.xlabel(f"Label: {label}")
plt.tight_layout()
plt.show()

Training

Let's define our model!

We will be using Resnet50 for prediction with weights from imagenet!

In [11]:
classes = 6
model = models.resnet50(pretrained = True)
model.fc = nn.Linear(2048, 1000)
model._fc = nn.Linear(1000, classes)
Downloading: "https://download.pytorch.org/models/resnet50-19c8e357.pth" to /root/.cache/torch/checkpoints/resnet50-19c8e357.pth

Let's now split our data into two parts, i.e. training and validation, which will help us in comparing our model's accuracy. We will using 80% of data in training while 20% of data will be used for test.

In [12]:
train_data, val_data = train_test_split(df_train, random_state = 42, test_size = 0.2, stratify = df_train["label"])

Here comes the part we are all waiting for, in the next few cells, we will pass our data through image preparing pipeline we defined earlier and then creating batches of the data using Dataloader class from Pytorch!

In [13]:
ds = ImagesDS(train_data, BASE_PATH, mode = "seg_train")
ds_val = ImagesDS(val_data, BASE_PATH, mode = "seg_train")
ds_test = ImagesDS(df_test, BASE_PATH, mode = "seg_test")
In [14]:
loader = D.DataLoader(ds, batch_size = BATCH_SIZE, shuffle = True, num_workers = 4)
val_loader = D.DataLoader(ds_val, batch_size = BATCH_SIZE, shuffle = True, num_workers = 4)
test_loader = D.DataLoader(ds_test, batch_size = BATCH_SIZE, shuffle = False, num_workers = 4)

Defining Loss, optimizer and metrics!

We have set learning rate to be 0.0003 with a weight decay of 0.0001.

In [15]:
criterion = nn.CrossEntropyLoss()
optimizer = torch.optim.Adam(model.parameters(), lr = 3e-4, weight_decay = 0.0001)
In [16]:
metrics = { "loss" : Loss(criterion), "accuracy" : Accuracy()}

trainer = create_supervised_trainer(model, optimizer, criterion, device = device)
val_eval = create_supervised_evaluator(model, metrics = metrics, device = device)

Below cells use some of the very powerful inbuilt functions from Ignite!

In [17]:
@trainer.on(Events.EPOCH_COMPLETED)
def compute_and_display(engine):
    epoch = engine.state.epoch
    metrics = val_eval.run(val_loader).metrics
    print("Validation Results - Epoch: {}  Average Loss: {:.4f} | Accuracy: {:.4f} "
          .format(engine.state.epoch, 
                      metrics['loss'], 
                      metrics['accuracy']))

Early stopping in case our model stops learning and we want to stop it midway, in order to prevent overfitting!

In [18]:
handler = EarlyStopping(patience = 4, score_function = lambda engine : engine.state.metrics['accuracy'], trainer = trainer)
val_eval.add_event_handler(Events.COMPLETED, handler)
Out[18]:
<ignite.engine.engine.RemovableEventHandle at 0x7f7f6784aad0>
In [19]:
checkpoints = ModelCheckpoint("models", "Model", n_saved = 3, create_dir = True)
trainer.add_event_handler(Events.EPOCH_COMPLETED, checkpoints, {"Resnet50" : model})
Out[19]:
<ignite.engine.engine.RemovableEventHandle at 0x7f7f6013b2d0>
In [20]:
pbar = ProgressBar(bar_format = '')
pbar.attach(trainer, output_transform = lambda x : {"loss" : x})

Yayyy, Training time! :)

In [21]:
trainer.run(loader, max_epochs = 30)
Validation Results - Epoch: 1  Average Loss: 0.2579 | Accuracy: 0.9070 
Validation Results - Epoch: 2  Average Loss: 0.2471 | Accuracy: 0.9141 
Validation Results - Epoch: 3  Average Loss: 0.2818 | Accuracy: 0.9056 
Validation Results - Epoch: 4  Average Loss: 0.2473 | Accuracy: 0.9109 
Validation Results - Epoch: 5  Average Loss: 0.2432 | Accuracy: 0.9081 
Validation Results - Epoch: 6  Average Loss: 0.2728 | Accuracy: 0.8985 
Out[21]:
State:
	iteration: 2106
	epoch: 6
	epoch_length: 351
	max_epochs: 30
	output: 0.24409562349319458
	batch: <class 'list'>
	metrics: <class 'dict'>
	dataloader: <class 'torch.utils.data.dataloader.DataLoader'>
	seed: 12

Inference

Evaluation mode, here, we will predict the labels for the test data we had put aside in the starting!

In [22]:
model.eval()
with torch.no_grad():
    preds = np.empty(0)
    for x, _ in tqdm_notebook(test_loader): 
        x = x.to(device)
        output = model(x)
        idx = output.max(dim=-1)[1].cpu().numpy()
        preds = np.append(preds, idx, axis=0)

preds = pd.Series(preds).astype(np.int)

Post Prediction Analysis

This is now time to see some of the images with their predicted labels and original ones side by side!

In [23]:
random_index = df_test.sample(9).index
fig=plt.figure(figsize=(10, 10))
columns = 3
rows = 3
for i in range(1, columns*rows +1):
    img = plt.imread(df_test.path[random_index[i-1]])
    fig.add_subplot(rows, columns, i)
    plt.imshow(img)
    plt.xticks([])
    plt.yticks([])
    prediction = encodings[preds[random_index[i-1]]]
    original = df_test.target[random_index[i-1]]
    plt.xlabel(f"Prediction: {prediction} \n Original: {original}")
plt.tight_layout()
plt.show()

Let's now see the %age accuracy of our model on all of the test data!

In [24]:
count = 0
for i in range(len(preds)):
    if df_test.label[i] == preds[i]:
        count += 1
        
print("Test accuracy: {}%".format((count/len(preds))*100))
Test accuracy: 90.06666666666666%

I hope you found reading this notebook worthwhile!

Thanks for reading!

Until next time! :)