Task 4 — Customer Churn Prediction System¶
This notebook implements a machine learning model to predict whether a customer will leave (churn) or continue using a service.
1. Data Collection¶
Load the customer churn dataset.
In [1]:
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import seaborn as sns
# Load the dataset
df = pd.read_csv('Telco-Customer-Churn.csv')
display(df.head())
| customerID | gender | SeniorCitizen | Partner | Dependents | tenure | PhoneService | MultipleLines | InternetService | OnlineSecurity | ... | DeviceProtection | TechSupport | StreamingTV | StreamingMovies | Contract | PaperlessBilling | PaymentMethod | MonthlyCharges | TotalCharges | Churn | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 0 | 7590-VHVEG | Female | 0 | Yes | No | 1 | No | No phone service | DSL | No | ... | No | No | No | No | Month-to-month | Yes | Electronic check | 29.85 | 29.85 | No |
| 1 | 5575-GNVDE | Male | 0 | No | No | 34 | Yes | No | DSL | Yes | ... | Yes | No | No | No | One year | No | Mailed check | 56.95 | 1889.5 | No |
| 2 | 3668-QPYBK | Male | 0 | No | No | 2 | Yes | No | DSL | Yes | ... | No | No | No | No | Month-to-month | Yes | Mailed check | 53.85 | 108.15 | Yes |
| 3 | 7795-CFOCW | Male | 0 | No | No | 45 | No | No phone service | DSL | Yes | ... | Yes | Yes | No | No | One year | No | Bank transfer (automatic) | 42.30 | 1840.75 | No |
| 4 | 9237-HQITU | Female | 0 | No | No | 2 | Yes | No | Fiber optic | No | ... | No | No | No | No | Month-to-month | Yes | Electronic check | 70.70 | 151.65 | Yes |
5 rows × 21 columns
2. Data Preprocessing & 4. Feature Engineering¶
Handle missing values, encode categorical variables, create new features, and scale data.
In [2]:
# 4. Feature Engineering: Remove irrelevant columns
df.drop('customerID', axis=1, inplace=True)
# Convert TotalCharges to numeric, coerce errors to NaN
df['TotalCharges'] = pd.to_numeric(df['TotalCharges'], errors='coerce')
# 2. Data Preprocessing: Handle missing values
df['TotalCharges'].fillna(df['TotalCharges'].median(), inplace=True)
# 4. Feature Engineering: Create Average Monthly Spend (already have MonthlyCharges, but let's demonstrate)
df['Avg_Monthly_Spend'] = np.where(df['tenure'] > 0, df['TotalCharges'] / df['tenure'], df['MonthlyCharges'])
# 4. Feature Engineering: Contract type grouping (Month-to-month vs Long-term)
df['Is_Long_Term_Contract'] = df['Contract'].apply(lambda x: 1 if x in ['One year', 'Two year'] else 0)
# 2. Data Preprocessing: Encode categorical variables
from sklearn.preprocessing import LabelEncoder
le = LabelEncoder()
for col in df.select_dtypes(include=['object']).columns:
df[col] = le.fit_transform(df[col])
display(df.head())
C:\Users\hari shanker s.t\AppData\Local\Temp\ipykernel_8316\2032627952.py:8: FutureWarning: A value is trying to be set on a copy of a DataFrame or Series through chained assignment using an inplace method.
The behavior will change in pandas 3.0. This inplace method will never work because the intermediate object on which we are setting values always behaves as a copy.
For example, when doing 'df[col].method(value, inplace=True)', try using 'df.method({col: value}, inplace=True)' or df[col] = df[col].method(value) instead, to perform the operation inplace on the original object.
df['TotalCharges'].fillna(df['TotalCharges'].median(), inplace=True)
| gender | SeniorCitizen | Partner | Dependents | tenure | PhoneService | MultipleLines | InternetService | OnlineSecurity | OnlineBackup | ... | StreamingTV | StreamingMovies | Contract | PaperlessBilling | PaymentMethod | MonthlyCharges | TotalCharges | Churn | Avg_Monthly_Spend | Is_Long_Term_Contract | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 0 | 0 | 0 | 1 | 0 | 1 | 0 | 1 | 0 | 0 | 2 | ... | 0 | 0 | 0 | 1 | 2 | 29.85 | 29.85 | 0 | 29.850000 | 0 |
| 1 | 1 | 0 | 0 | 0 | 34 | 1 | 0 | 0 | 2 | 0 | ... | 0 | 0 | 1 | 0 | 3 | 56.95 | 1889.50 | 0 | 55.573529 | 1 |
| 2 | 1 | 0 | 0 | 0 | 2 | 1 | 0 | 0 | 2 | 2 | ... | 0 | 0 | 0 | 1 | 3 | 53.85 | 108.15 | 1 | 54.075000 | 0 |
| 3 | 1 | 0 | 0 | 0 | 45 | 0 | 1 | 0 | 2 | 0 | ... | 0 | 0 | 1 | 0 | 0 | 42.30 | 1840.75 | 0 | 40.905556 | 1 |
| 4 | 0 | 0 | 0 | 0 | 2 | 1 | 0 | 1 | 0 | 0 | ... | 0 | 0 | 0 | 1 | 2 | 70.70 | 151.65 | 1 | 75.825000 | 0 |
5 rows × 22 columns
3. Exploratory Data Analysis (EDA)¶
Visualize churn distribution, charges vs churn, tenure vs churn, and correlation.
In [3]:
plt.figure(figsize=(6,4))
sns.countplot(data=df, x='Churn')
plt.title('Churn Distribution (0: No, 1: Yes)')
plt.show()
plt.figure(figsize=(6,4))
sns.boxplot(x='Churn', y='MonthlyCharges', data=df)
plt.title('Monthly Charges vs Churn')
plt.show()
plt.figure(figsize=(6,4))
sns.boxplot(x='Churn', y='tenure', data=df)
plt.title('Tenure vs Churn')
plt.show()
plt.figure(figsize=(12,8))
sns.heatmap(df.corr(), annot=False, cmap='coolwarm')
plt.title('Correlation Heatmap')
plt.show()
5. Model Building & 6. Model Evaluation¶
Train Logistic Regression, Decision Tree, Random Forest, and KNN. Evaluate using Accuracy, Precision, Recall, F1 Score, and Confusion Matrix.
In [4]:
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
from sklearn.metrics import accuracy_score, precision_score, recall_score, f1_score, confusion_matrix, ConfusionMatrixDisplay
from sklearn.linear_model import LogisticRegression
from sklearn.tree import DecisionTreeClassifier
from sklearn.ensemble import RandomForestClassifier
from sklearn.neighbors import KNeighborsClassifier
# Prepare data
X = df.drop('Churn', axis=1)
y = df['Churn']
# Train-test split
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
# Feature scaling
scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train)
X_test_scaled = scaler.transform(X_test)
# Initialize models
models = {
'Logistic Regression': LogisticRegression(),
'Decision Tree': DecisionTreeClassifier(random_state=42),
'Random Forest': RandomForestClassifier(random_state=42),
'KNN': KNeighborsClassifier()
}
# Train and evaluate models
results = []
for name, model in models.items():
model.fit(X_train_scaled, y_train)
y_pred = model.predict(X_test_scaled)
acc = accuracy_score(y_test, y_pred)
prec = precision_score(y_test, y_pred)
rec = recall_score(y_test, y_pred)
f1 = f1_score(y_test, y_pred)
cm = confusion_matrix(y_test, y_pred)
results.append({
'Model': name,
'Accuracy': acc,
'Precision': prec,
'Recall': rec,
'F1 Score': f1
})
print(f"\n{name} Confusion Matrix:")
disp = ConfusionMatrixDisplay(confusion_matrix=cm)
disp.plot(cmap='Blues')
plt.title(f'{name} Confusion Matrix')
plt.show()
# Display evaluation metrics
results_df = pd.DataFrame(results)
display(results_df)
Logistic Regression Confusion Matrix:
Decision Tree Confusion Matrix:
Random Forest Confusion Matrix:
KNN Confusion Matrix:
| Model | Accuracy | Precision | Recall | F1 Score | |
|---|---|---|---|---|---|
| 0 | Logistic Regression | 0.815472 | 0.677116 | 0.579088 | 0.624277 |
| 1 | Decision Tree | 0.731725 | 0.493473 | 0.506702 | 0.500000 |
| 2 | Random Forest | 0.804826 | 0.681481 | 0.493298 | 0.572317 |
| 3 | KNN | 0.765791 | 0.560907 | 0.530831 | 0.545455 |