Real-world style machine learning application for EMI eligibility classification and maximum monthly EMI prediction.
EMIPredict AI helps evaluate loan or EMI applications using customer profile, income, expense, credit, and requested loan details.
The Streamlit app includes:
- Dashboard for portfolio overview and model readiness.
- Single applicant prediction with eligibility, safe EMI, risk signals, policy checks, and stress testing.
- Batch scoring from CSV upload.
- Saved application queue with approval/manual-review/decline filters.
- Model performance and EDA views.
- MLflow experiment summary.
- Clean the raw EMI dataset.
- Perform exploratory data analysis.
- Engineer financial risk features.
- Train classification models for
emi_eligibility. - Train regression models for
max_monthly_emi. - Track experiments with MLflow.
- Deploy the Streamlit app from GitHub.
For running only the Streamlit app:
python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
streamlit run app.pyFor the full ML pipeline, including training and MLflow:
pip install -r requirements-dev.txtOn macOS, XGBoost may also require:
brew install libompUse this repository setup:
- App entrypoint:
streamlit_app.py - Python dependencies:
requirements.txt - Linux dependency for XGBoost:
packages.txt - Streamlit theme:
.streamlit/config.toml - Required model files:
models/ - Required report files:
reports/tables/
In Streamlit Community Cloud:
- Push this project to GitHub.
- Open
https://share.streamlit.io. - Click
Create app. - Select the GitHub repository and branch.
- Set the main file path to
streamlit_app.py. - In Advanced settings, choose Python
3.12. - Deploy.
If this folder is not already a Git repository:
git init
git add .
git commit -m "Prepare EMIPredict AI for Streamlit deployment"
git branch -M main
git remote add origin https://github.com/YOUR_USERNAME/emipredict-ai.git
git push -u origin mainReplace YOUR_USERNAME and repository name with your actual GitHub details.
- Do not commit the large raw CSV dataset; it is ignored in
.gitignore. - The deployed app uses the already-trained models in
models/. - Local MLflow files such as
mlflow.db,mlruns/, andmlartifacts/are ignored because they are not needed by Streamlit Cloud. - Saved applications in the deployed app use the app container filesystem and may reset when the app restarts.
Preprocess data:
python src/preprocess_data.pyRun EDA:
python src/eda_analysis.pyTrain models:
python src/train_models.pyTrain on the full cleaned dataset:
python src/train_models.py --sample-size 0Open the local MLflow dashboard:
mlflow ui --backend-store-uri sqlite:///mlflow.db- Classification target:
emi_eligibility - Regression target:
max_monthly_emi