A next-generation road hazard analysis stack powered by YOLOv8, Depth-Anything-V2, DINOv2, and Shape-From-Shading.
Traditional pothole detection systems rely entirely on bounding boxes, treating 2D pixel anomalies as road damage. RoadLens bridges the gap between 2D segmentation and 3D physical reality. By synthesizing monocular depth estimation, photometric shadow analysis, and foundation semantic features, RoadLens accurately categorizes road damage severity, predicts hazard levels, and rejects visual illusions (like shadows or manhole covers) that fool traditional models.
RoadLens utilizes four advanced, concurrent pipelines to analyze a single frame:
Global depth estimation models (like Depth-Anything-V2) excel at scene structure but fail to capture the high-frequency micro-textures of potholes. RoadLens combines global depth gradients with Shape-From-Shading (SFS). By calculating surface normals from lighting gradients and integrating them via the Frankot-Chellappa algorithm, the system generates localized, high-resolution pseudo-depth maps of craters.
Tree shadows on flat roads often trigger false positive "deep" readings in traditional depth networks. To counter this, we integrated DINOv2 Foundation Features. The system extracts patch-level feature vectors from the bounding box. High internal variance in these vectors indicates jagged, chaotic textures (true potholes), while low variance indicates a flat surface (shadows/illusions), allowing the system to actively downgrade false severities.
Submerged potholes disguise their true depth and pose hydroplaning risks. The pipeline uses HSV color thresholding, reflection variance analysis, and contour area checking to determine a Water Presence Boolean. If a pothole is classified as filled with water, its severity is mathematically inflated due to the hidden danger.
For video feeds or sequential inspections, the temporal_analysis.py module tracks individual potholes across frames using Intersection over Union (IoU) and structural similarity (SSIM). It calculates degradation slopes over time, alerting authorities to exponentially worsening road conditions.
graph TD
A[Upload Road Image] --> B{YOLOv8 Segmentation}
%% Parallel Feature Extraction Paths
B -->|Mask & BBox| C[Multi-Modal Feature Extractor]
A -->|Monocular Frame| D[Depth Anything V2]
A -->|Photometric Grayscale| E[Shape-From-Shading]
A -->|RGB Tensor| F[DINOv2 Foundation Vision]
%% Merging Features
D -->|Scene Depth Map| C
E -->|Micro-texture Gradients| C
F -->|Semantic Patch Variance| C
%% Water Detection Branch
A -->|Color & Reflection| W[Water Hazard Detector]
W -->|Hidden Danger Boolean| C
%% Machine Learning
C -->|Feature Vector: Volume, Depth Var, Context| G[Ensemble ML Models]
C -->|Constraints| H[Rule-Based Fallback]
%% Output
G --> I((Final Packaged JSON))
H --> I
style A fill:#2d3436,stroke:#74b9ff,stroke-width:2px,color:#fff
style I fill:#00b894,stroke:#55efc4,stroke-width:4px,color:#fff
The system extracts a 31-dimensional feature vector for each segmented pothole. These features are scaled and fed into an ensemble of machine learning classifiers to determine severity (Low, Medium, High).
Models Evaluated: Random Forest, XGBoost, SVM, KNN, LightGBM, Logistic Regression, MLP, and Naive Bayes.
Detailed classification reports, SHAP feature importance graphs, and t-SNE projections are generated automatically and saved to the ml_results/ directory during training.
RoadLens is trained and validated on a merged dataset combining:
- RDD2022 (Road Damage Detector): Global bounding box dataset.
- Pothole-600: High-resolution semantic segmentation masks.
The backend exposes a highly optimized FastAPI service for integration with dashboards or edge devices.
| Method | Endpoint | Description |
|---|---|---|
GET |
/healthz |
Confirms the API and ML models are loaded into memory. |
POST |
/analyze |
Main inference endpoint. Accepts multipart/form-data image upload. Returns bounding boxes, severity labels, confidence scores, water hazard flags, and Base64 overlay images. |
GET |
/insights/summary |
Aggregates the latest metrics from the ml_results/ directory for dashboard charting. |
GET |
/insights/files/{file} |
Serves artifacts (like SHAP charts) to the frontend. |
RoadLens/
├── api.py # FastAPI endpoints and route handlers
├── main.py # CLI interface for standard processing
├── inference.py # Headless inference for batch scripts
├── src/
│ ├── core/ # YOLO handling, Geometry extraction, ML parsing
│ └── advanced/ # SFS math, DINOv2 tensors, Water reflections
├── tests/ # PyTest/Unittest scripts (API, Water, Flow)
├── ml_models/ # Serialized .pkl files (Scalers, RandomForests)
├── ml_results/ # Auto-generated PDF reports and PNG charts
├── docs/ # Vector assets, architectural plans, diagrams
├── scripts/ # Dataset ingestion (RDD2022/Pothole600 converters)
└── web-ui/ # React/Vite Frontend (Bento-Grid Dashboard)
1. Python Environment (Click to expand)
We highly recommend using a virtual environment to prevent dependency conflicts.
python -m venv .venv
.venv\Scripts\activate
python -m pip install --upgrade pip
python -m pip install -r requirements.txt
copy .env.example .env2. Downloading Foundation Weights
Due to GitHub file limits, large models must be downloaded locally:
- Depth-Anything-V2 Repository:
git clone https://github.com/DepthAnything/Depth-Anything-V2.git
- Depth Checkpoint: Download the
vitscheckpoint from the official Depth-Anything-V2 releases and place it at:Depth-Anything-V2/checkpoints/depth_anything_v2_vits.pth - YOLOv8 Custom Weights: Place your trained segmentation model at
yolo-segmentation/model/best.pt
3. Web UI Setup
cd web-ui
copy .env.example .env
npm installEnsure your web-ui/.env contains VITE_API_BASE_URL=http://localhost:8000.
Run the Backend Server:
python api.pyRun the React Dashboard:
cd web-ui
npm run devRun Automated Verification Tests:
# Runs the full end-to-end processing pipeline on test images
python tests/test_pipeline.py
# Runs specific sub-module checks
python tests/test_water_detection.py