top of page

Customer Segmentation retention Analytics

A data analytics project using machine learning and RFM segmentation to predict churn and profile customer segments for improved retention strategies in telco and retail datasets.

Overview:


This project builds an end-to-end customer analytics system that segments users and predicts churn using real-world datasets from the telecom and online retail domains. It simulates how subscription platforms, SaaS companies, e-commerce firms, and mobility apps (e.g., Netflix, Spotify, Amazon, Uber, Slack) analyse customer behavior to improve retention and maximize customer lifetime value.


The project combines data cleaning, feature engineering, machine learning, clustering, and business decision modeling into a single reproducible pipeline.


Objectives:

  1. Segment customers based on behavioral patterns using RFM (Recency, Frequency, Monetary) analysis

  2. Predict customer churn using supervised machine learning

  3. Identify high-value and at-risk customers

  4. Design data-driven retention and resource allocation strategies


Datasets

  1. Telco Customer Churn (Kaggle)

    Customer demographics, services, billing information, contract types

    Target variable: Churn (Yes / No)

    Used for churn prediction and churn driver analysis

2. Online Retail (Kaggle)

Transaction-level purchase history (invoices, quantities, prices, timestamps)

Used for RFM segmentation and customer value analysis

Techniques & Tools Tech Stack

Python (pandas, numpy)

Scikit-learn (Random Forest, K-Means, preprocessing) Matplotlib / Seaborn (visualization)

Jupyter Notebook


Methods

  1. Data Cleaning & Feature Engineering Label Encoding of categorical variables RFM (Recency, Frequency, Monetary) analysis

  2. K-Means clustering for customer segmentation

  3. Random Forest classifier for churn prediction Feature importance analysis for churn drivers


Project Workflow


Data Ingestion

↓

Data Cleaning & Feature Engineering

↓

Customer Segmentation (RFM + Clustering)

↓

Churn Prediction Model

↓

Segment Risk Analysis

↓

Retention Strategy & Business Decisions


Key Results & Insights Churn Prediction


  1. Built a Random Forest churn model with ROC-AUC around 0.84

  2. Key churn drivers: Short customer tenure Month-to-month contracts High monthly charges Lack of bundled services Customer Segmentation (RFM)

  3. Customers were grouped into four major behavioural segments:

    High Value Loyal -frequent and high-spending customers with recent activity

    Frequent Buyers -high frequency, moderate spending

    At Risk - inactive customers with low frequency and high recency

    Low Engagement - low spending and low activity


Retention Targeting

  1. Identified high-value but inactive customers as critical retention targets

  2. Flagged low-value churn-prone users where marketing spend should be minimized



Business Impact:

This project demonstrates how analytics can support real business decisions:

  1. Prioritize retention offers for high-value customers with rising churn risk

  2. Reward loyal customers with early access and loyalty benefits

  3. Avoid wasting marketing resources on customers with low lifetime value and high churn probability


The approach mirrors real-world customer lifecycle management used in streaming, SaaS, e-commerce, and mobility platforms.


Future Improvements:

  1. Add customer lifetime value (CLV) prediction

  2. Build an interactive dashboard using Power BI or Tableau

  3. Experiment with XGBoost and SHAP for model explainability

  4. Integrate churn + segmentation into a single scoring framework





Project Gallery

bottom of page