Production-grade loan prediction platform with MLflow model registry, DVC-backed datasets, FastAPI inference service, and monitoring/drift tooling.
- MLOps: MLflow experiment tracking and registry, Hyperopt HPO, SHAP explainability, Fairlearn fairness checks, DVC-backed datasets (S3 remote).
- API: FastAPI service on port 8005 with
/health,/prediction_api,/batch_prediction,/model_info,/list_models, and/metrics(Prometheus). - Monitoring: Prometheus + Grafana stack (
monitoring/docker-compose.yml) and Evidently-based Streamlit drift app (drift_monitoring/app_v1.py). - Deployment: Dockerfile that syncs MLflow artifacts from
s3://loanpred-mlops-20251118-120330/mlruns/, Kubernetes manifests ink8s/, and GitHub Actions CI/CD (.github/workflows/cicd.yml).
Loan-Prediction_MLOps/
- .dvc/ # DVC configuration (remote: s3://loanpred-mlops-20251118-120330)
- .github/workflows/ # CI/CD pipelines (cicd, docs-check, OIDC test)
- docs/ # Setup, API, monitoring, CI/CD, infra guides
- drift_monitoring/ # Evidently Streamlit drift app + Dockerfile
- grafana/ , prometheus/ # Provisioning for monitoring stack
- k8s/ # Kubernetes manifests for EKS deployment
- monitoring/ # docker-compose for Prometheus + Grafana
- prediction_model/ # Training, preprocessing, inference, config
- scripts/ # Helper scripts (DVC push, Docker build, EKS auth)
- tests/ # API and prediction tests
- Dockerfile # Production image (port 8005)
- datasets.dvc # DVC tracking for datasets/ (pulled from S3)
- main.py # FastAPI entrypoint
- requirements*.txt # Dependencies
- Python 3.11+
- pip
- Docker + Docker Compose (for monitoring)
- AWS CLI with access to the S3 bucket (
loanpred-mlops-20251118-120330by default) for data/artifacts - DVC with S3 support:
pip install dvc[s3]
- Clone and create a virtual environment
git clone https://github.com/your-org/Loan-Prediction_MLOps.git
cd Loan-Prediction_MLOps
# venv (Linux/Mac)
python -m venv venv && source venv/bin/activate
# venv (Windows)
python -m venv venv && venv\Scripts\activate- Install dependencies
pip install -r requirements.txt
# Optional: development tooling
pip install -r requirements-dev.txt- Configure environment variables
cp .env.example .env
# Key variables from .env.example / code defaults:
# MLFLOW_TRACKING_URI=./mlruns # File-based MLflow tracking by default
# MODEL_STAGE=Production # Which stage the API serves
# API_PORT=8005
# AWS_REGION=eu-west-2
# S3_BUCKET=loanpred-mlops-20251118-120330 # default in configs/DVC/Dockerfile
# DVC_REMOTE=s3://loanpred-mlops-20251118-120330
# DRIFT_S3_BUCKET=loanpred-mlops-20251118-120330- Pull datasets with DVC (requires AWS credentials)
aws configure # Ensure access to <your AWS_REGION>
dvc pull # Uses remote 'myremote' from .dvc/configThe default DVC remote points to s3://loanpred-mlops-20251118-120330.
- (Optional) Download existing MLflow runs/models
aws s3 sync s3://loanpred-mlops-20251118-120330/mlruns/ ./mlruns/Train models and log to MLflow (experiment loan_prediction_model):
python -m prediction_model.training_pipeline- Reads data from
datasets/train.csvanddatasets/test.csv(tracked by DVC). - Logs metrics, SHAP plots, and models to
./mlruns(or the URI inMLFLOW_TRACKING_URI). - S3 bucket used for artifacts by default:
loanpred-mlops-20251118-120330(seeprediction_model/config/config.py).
Start FastAPI (serves the latest model stage via MLflow):
python main.py
# or: uvicorn main:app --host 0.0.0.0 --port 8005Environment notes:
- Ensure
mlruns/contains a trained model (train locally or sync from S3). MODEL_STAGEcontrols which stage to load (Productiondefault).- Metrics are exposed at
/metricsfor Prometheus scraping. Note: Inputs are parsed directly into pandas without Pydantic validation; ensure the JSON matches expected feature names and types.
Single prediction:
curl -X POST http://localhost:8005/prediction_api \
-H "Content-Type: application/json" \
-d '{
"Gender": "Male",
"Married": "Yes",
"Dependents": "0",
"Education": "Graduate",
"Self_Employed": "No",
"ApplicantIncome": 5000,
"CoapplicantIncome": 2000,
"LoanAmount": 150,
"Loan_Amount_Term": 360,
"Credit_History": 1,
"Property_Area": "Urban"
}'Batch prediction (CSV upload):
curl -X POST http://localhost:8005/batch_prediction \
-F "file=@batch_data.csv" \
-o predictions.csvModel info:
curl "http://localhost:8005/model_info?stage=Production"- Prometheus/Grafana:
cd monitoring docker-compose up -d- Prometheus: http://localhost:9090
- Grafana: http://localhost:3000 (admin/admin by default)
- The API exposes
/metrics(usesprometheus_fastapi_instrumentator). - Data drift UI:
Set
cd drift_monitoring pip install -r requirements.txt streamlit run app_v1.pyDRIFT_S3_BUCKETif pulling batches from S3; defaults are indrift_monitoring/config.py. - Script
collect_k8s_predictions.pycan collect sample predictions from a cluster endpoint for drift analysis.
- Datasets are tracked via DVC (
datasets.dvc). Remotemyremoteis defined in.dvc/config:- URL:
s3://loanpred-mlops-20251118-120330 - Region:
eu-west-2
- URL:
- Common commands:
dvc pull # fetch datasets dvc add datasets/ # track updated data dvc push # push data to S3
- MLflow artifacts: default bucket
loanpred-mlops-20251118-120330(Docker startup script syncsmlruns/from there).
- Build locally:
scripts\build_docker_local.bat # Windows helper # or docker build -t loan-prediction:local .
- Run container:
The image syncs
docker run -p 8005:8005 \ -e MLFLOW_TRACKING_URI="http://host.docker.internal:5000" \ loan-prediction:localmlruns/from S3 on startup (see Dockerfilestart.sh). - Kubernetes manifests are in
k8s/(deployments, services, ServiceMonitor, namespace/auth helpers). - GitHub Actions CI/CD (
.github/workflows/cicd.yml) trains, validates metrics, pushes to ECR, and deploys to EKS using DVC data from the same S3 remote.
Tests live in tests/:
- Unit/API tests target the FastAPI app and prediction helpers.
- Integration tests are guarded by a
--run-integrationflag (defined intests/conftest.py) and require MLflow models inmlruns/. - Run:
pytest tests/ -v # fast checks pytest tests/ -v --run-integration # requires MLflow models present in mlruns/
If you see Client.__init__() got an unexpected keyword argument 'app', reinstall the pinned HTTPX version: pip install httpx==0.24.1. If collection fails on pytest.config, update the test skip logic to use request.config.getoption("--run-integration") or the RUN_INTEGRATION flag from tests/conftest.py.
docs/SETUP.md: Detailed environment setupdocs/API.md: Endpoint contract and schemasdocs/MONITORING.md: Metrics, Grafana, Fairlearn, MLflow instructionsdocs/CI_CD_WORKFLOW.md: CI/CD pipeline breakdowndocs/infra/*.md: ECR, EKS, Prometheus, Grafana, and S3 notes
MIT License. See LICENSE.