Task 4 — Customer Churn Prediction System¶

This notebook implements a machine learning model to predict whether a customer will leave (churn) or continue using a service.

1. Data Collection¶

Load the customer churn dataset.

In [1]:
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import seaborn as sns

# Load the dataset
df = pd.read_csv('Telco-Customer-Churn.csv')
display(df.head())
customerID gender SeniorCitizen Partner Dependents tenure PhoneService MultipleLines InternetService OnlineSecurity ... DeviceProtection TechSupport StreamingTV StreamingMovies Contract PaperlessBilling PaymentMethod MonthlyCharges TotalCharges Churn
0 7590-VHVEG Female 0 Yes No 1 No No phone service DSL No ... No No No No Month-to-month Yes Electronic check 29.85 29.85 No
1 5575-GNVDE Male 0 No No 34 Yes No DSL Yes ... Yes No No No One year No Mailed check 56.95 1889.5 No
2 3668-QPYBK Male 0 No No 2 Yes No DSL Yes ... No No No No Month-to-month Yes Mailed check 53.85 108.15 Yes
3 7795-CFOCW Male 0 No No 45 No No phone service DSL Yes ... Yes Yes No No One year No Bank transfer (automatic) 42.30 1840.75 No
4 9237-HQITU Female 0 No No 2 Yes No Fiber optic No ... No No No No Month-to-month Yes Electronic check 70.70 151.65 Yes

5 rows × 21 columns

2. Data Preprocessing & 4. Feature Engineering¶

Handle missing values, encode categorical variables, create new features, and scale data.

In [2]:
# 4. Feature Engineering: Remove irrelevant columns
df.drop('customerID', axis=1, inplace=True)

# Convert TotalCharges to numeric, coerce errors to NaN
df['TotalCharges'] = pd.to_numeric(df['TotalCharges'], errors='coerce')

# 2. Data Preprocessing: Handle missing values
df['TotalCharges'].fillna(df['TotalCharges'].median(), inplace=True)

# 4. Feature Engineering: Create Average Monthly Spend (already have MonthlyCharges, but let's demonstrate)
df['Avg_Monthly_Spend'] = np.where(df['tenure'] > 0, df['TotalCharges'] / df['tenure'], df['MonthlyCharges'])

# 4. Feature Engineering: Contract type grouping (Month-to-month vs Long-term)
df['Is_Long_Term_Contract'] = df['Contract'].apply(lambda x: 1 if x in ['One year', 'Two year'] else 0)

# 2. Data Preprocessing: Encode categorical variables
from sklearn.preprocessing import LabelEncoder
le = LabelEncoder()

for col in df.select_dtypes(include=['object']).columns:
    df[col] = le.fit_transform(df[col])

display(df.head())
C:\Users\hari shanker s.t\AppData\Local\Temp\ipykernel_8316\2032627952.py:8: FutureWarning: A value is trying to be set on a copy of a DataFrame or Series through chained assignment using an inplace method.
The behavior will change in pandas 3.0. This inplace method will never work because the intermediate object on which we are setting values always behaves as a copy.

For example, when doing 'df[col].method(value, inplace=True)', try using 'df.method({col: value}, inplace=True)' or df[col] = df[col].method(value) instead, to perform the operation inplace on the original object.


  df['TotalCharges'].fillna(df['TotalCharges'].median(), inplace=True)
gender SeniorCitizen Partner Dependents tenure PhoneService MultipleLines InternetService OnlineSecurity OnlineBackup ... StreamingTV StreamingMovies Contract PaperlessBilling PaymentMethod MonthlyCharges TotalCharges Churn Avg_Monthly_Spend Is_Long_Term_Contract
0 0 0 1 0 1 0 1 0 0 2 ... 0 0 0 1 2 29.85 29.85 0 29.850000 0
1 1 0 0 0 34 1 0 0 2 0 ... 0 0 1 0 3 56.95 1889.50 0 55.573529 1
2 1 0 0 0 2 1 0 0 2 2 ... 0 0 0 1 3 53.85 108.15 1 54.075000 0
3 1 0 0 0 45 0 1 0 2 0 ... 0 0 1 0 0 42.30 1840.75 0 40.905556 1
4 0 0 0 0 2 1 0 1 0 0 ... 0 0 0 1 2 70.70 151.65 1 75.825000 0

5 rows × 22 columns

3. Exploratory Data Analysis (EDA)¶

Visualize churn distribution, charges vs churn, tenure vs churn, and correlation.

In [3]:
plt.figure(figsize=(6,4))
sns.countplot(data=df, x='Churn')
plt.title('Churn Distribution (0: No, 1: Yes)')
plt.show()

plt.figure(figsize=(6,4))
sns.boxplot(x='Churn', y='MonthlyCharges', data=df)
plt.title('Monthly Charges vs Churn')
plt.show()

plt.figure(figsize=(6,4))
sns.boxplot(x='Churn', y='tenure', data=df)
plt.title('Tenure vs Churn')
plt.show()

plt.figure(figsize=(12,8))
sns.heatmap(df.corr(), annot=False, cmap='coolwarm')
plt.title('Correlation Heatmap')
plt.show()
No description has been provided for this image
No description has been provided for this image
No description has been provided for this image
No description has been provided for this image

5. Model Building & 6. Model Evaluation¶

Train Logistic Regression, Decision Tree, Random Forest, and KNN. Evaluate using Accuracy, Precision, Recall, F1 Score, and Confusion Matrix.

In [4]:
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
from sklearn.metrics import accuracy_score, precision_score, recall_score, f1_score, confusion_matrix, ConfusionMatrixDisplay
from sklearn.linear_model import LogisticRegression
from sklearn.tree import DecisionTreeClassifier
from sklearn.ensemble import RandomForestClassifier
from sklearn.neighbors import KNeighborsClassifier

# Prepare data
X = df.drop('Churn', axis=1)
y = df['Churn']

# Train-test split
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

# Feature scaling
scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train)
X_test_scaled = scaler.transform(X_test)

# Initialize models
models = {
    'Logistic Regression': LogisticRegression(),
    'Decision Tree': DecisionTreeClassifier(random_state=42),
    'Random Forest': RandomForestClassifier(random_state=42),
    'KNN': KNeighborsClassifier()
}

# Train and evaluate models
results = []
for name, model in models.items():
    model.fit(X_train_scaled, y_train)
    y_pred = model.predict(X_test_scaled)
    
    acc = accuracy_score(y_test, y_pred)
    prec = precision_score(y_test, y_pred)
    rec = recall_score(y_test, y_pred)
    f1 = f1_score(y_test, y_pred)
    cm = confusion_matrix(y_test, y_pred)
    
    results.append({
        'Model': name,
        'Accuracy': acc,
        'Precision': prec,
        'Recall': rec,
        'F1 Score': f1
    })
    
    print(f"\n{name} Confusion Matrix:")
    disp = ConfusionMatrixDisplay(confusion_matrix=cm)
    disp.plot(cmap='Blues')
    plt.title(f'{name} Confusion Matrix')
    plt.show()

# Display evaluation metrics
results_df = pd.DataFrame(results)
display(results_df)
Logistic Regression Confusion Matrix:
No description has been provided for this image
Decision Tree Confusion Matrix:
No description has been provided for this image
Random Forest Confusion Matrix:
No description has been provided for this image
KNN Confusion Matrix:
No description has been provided for this image
Model Accuracy Precision Recall F1 Score
0 Logistic Regression 0.815472 0.677116 0.579088 0.624277
1 Decision Tree 0.731725 0.493473 0.506702 0.500000
2 Random Forest 0.804826 0.681481 0.493298 0.572317
3 KNN 0.765791 0.560907 0.530831 0.545455