Credit Card Fraud Detection

An End-to-End Machine Learning Pipeline

Presented by: Jayan Gupta

Q1. Load Dataset

Imports the necessary libraries and loads the dataset from the CSV file.

Q2. Data Types & Missing Values

Checks the structural matrix of the dataset and confirms there are absolutely zero missing values.

Total Records
284,807
Features
31
Missing Nulls
0

Q3. Class Imbalance

Reveals extreme class imbalance. A model guessing 'Legitimate' every time would achieve 99.82% accuracy, making raw metrics highly deceptive.

Legitimate (Class 0) 284,315 (99.82%)
Fraudulent (Class 1) 492 (0.17%)
*Fraud scaling amplified visually

Q4. EDA: Amounts

Splits the dataset and calculates summary statistics. Fraudulent transactions have a higher average amount but are capped at lower maximums.

Legit Mean
$88.29
Fraud Mean
$122.21
Legit Max
$25,691
Fraud Max
$2,125

Q5. EDA: Time vs Amount

Visualizes transactions over time. Fraud occurs consistently across all hours, bypassing standard diurnal patterns observed in legitimate activity.

Q6. Feature Correlation

PCA features (V1-V28) show zero correlation with each other. Noticeable negative correlations exist between the Fraud Class and specific features like V14, V17.

Q7. Splitting the Data

Splits the dataset into 80% training and 20% testing sets. The 'stratify' parameter ensures the extreme class imbalance is preserved in both sets.

Q8. Preprocessing

Scales numerical features (like Amount) to have a mean of 0 and variance of 1, preventing features with larger scales from dominating the distance metrics.

Q9. KNN Training

Trains a K-Nearest Neighbors classifier on a subset of the data. It maps multidimensional spatial zones to cluster similar transactions.

Q10. KNN Evaluation

Evaluates the KNN model. While it achieves high accuracy, standard accuracy fails to expose hidden errors due to severe class imbalance.

Q11. XGBoost Training

Trains an XGBoost classifier. We balance class weights via the 'scale_pos_weight' parameter to heavily penalize missing fraud cases sequentially.

Q12. XGBoost Evaluation

Evaluates the XGBoost model. It significantly outperforms KNN, especially in Recall, correctly flagging more total fraud occurrences.

Q13. GridSearch Tuning

Exhaustively searches over a specified parameter grid to find the optimal model settings. It checks every combination iteratively to maximize recall.

Q14. RandomSearch Tuning

Randomly samples from parameter distributions, offering a computationally efficient alternative to Grid Search while still finding highly optimized parameters.

Q15. Model Comparison

Final Model Comparison. XGBoost outscores KNN across every operational checkpoint metric, notably boosting Recall parameters and tracking fewer false negatives.

1 / 16