An end-to-end computer-vision pipeline that detects the body keypoints of Caenorhabditis elegans
worms in microscopy video, then classifies their locomotion behavior from the extracted pose
dynamics. The project covers the full applied-ML lifecycle: dataset engineering (COCO to YOLO
conversion), transfer learning on a keypoint-detection model, biologically-motivated feature
engineering, model selection across classical and deep-learning classifiers, and a reproducible
inference pipeline that turns a raw .avi clip into a predicted behavior.
Keypoint detection (YOLOv8s-pose, validation set):
| Metric | Box | Pose |
|---|---|---|
| mAP@50 | 96.8 | 96.9 |
| mAP@50-95 | 87.8 | 89.1 |
Behavior classification (LabGym clips, held-out test set of 490 samples):
| Model | Accuracy | F1 (macro) | F1 (weighted) |
|---|---|---|---|
| Random Forest | 0.686 | 0.632 | 0.679 |
| Gradient Boosting | 0.688 | 0.624 | 0.682 |
| LSTM | 0.576 | 0.536 | 0.583 |
Random Forest was selected as the production classifier on the basis of macro-F1 and 5-fold cross-validation stability (0.692 +/- 0.014 accuracy).
Quantifying animal behavior from video is a core task in neuroscience and toxicology research. C. elegans is a model organism whose locomotion (crawling, reversing, coiling into an "omega" bend) reflects its neural and physiological state. Manual annotation of these behaviors is slow and subjective. This project automates the two steps that make automated behavioral phenotyping possible:
- Where is the worm and how is its body posed? A keypoint-detection model localizes five anatomical landmarks along the worm body.
- What is the worm doing? A classifier reads the temporal dynamics of those landmarks across a clip and predicts one of five locomotion behaviors.
Raw video (.avi)
|
v
YOLOv8s-pose ------> 5 keypoints per frame: head, front, middle, back, tail
|
v
Feature extraction --> per-clip biological descriptors
| (speed, body curvature, elongation, path linearity, detection rate)
v
Random Forest classifier
|
v
Predicted behavior: {forward crawling, reverse crawling, omega bend, twitching, immobile}
Pose estimation (SpaceAnimal). The SpaceAnimal dataset provides multi-species keypoint annotations (C. elegans, Drosophila, zebrafish). This project focuses on C. elegans:
| Property | Value |
|---|---|
| Images | 6,996 (5,622 train / 1,374 val) |
| Annotated instances | 15,378 (12,349 train / 3,029 val) |
| Keypoints per worm | 5 (head, front, middle, back, tail) |
| Skeleton | head-front-middle-back-tail chain |
| Source annotation format | COCO keypoints |
Annotations were converted from COCO to the YOLO-pose label format as part of the pipeline.
Behavior classification (LabGym). Locomotion behavior is trained on annotated .avi clips from
the LabGym C. elegans locomotion categorizer, spanning five classes: forward crawling,
reverse crawling, omega bend, twitching, and immobile.
- Model:
yolov8s-pose, pretrained on COCO (17 human keypoints). The pose head is reshaped for the 5-point worm skeleton and the full network — backbone, neck, and head — is fine-tuned end-to-end on the C. elegans data (transfer learning via full fine-tuning, not a frozen-backbone / linear-probe setup). - Training: 50 epochs on a single Tesla T4 GPU, with mosaic and horizontal-flip augmentation and a configuration-driven, reproducible training setup.
- Evaluation: standard box and pose mAP at IoU thresholds 0.50 and 0.50-0.95.
Each clip is summarized into a fixed-length vector of biologically interpretable features aggregated over all frames:
- Centroid speed (mean, std, max) captures forward/reverse motion intensity.
- Body curvature (mean, std) is the head-middle-tail angle, low for an omega bend.
- Elongation (mean, std) and body length describe posture and contraction.
- Path linearity separates directed crawling from local, non-directed movement.
- Detection rate encodes how reliably the worm was tracked across the clip.
Three classifiers were trained and compared under identical splits: Random Forest, Gradient Boosting, and an LSTM. Selection used macro-F1 (to account for class imbalance) plus 5-fold cross-validation. The classical tree ensembles outperformed the LSTM, which is consistent with the compact, hand-crafted feature representation and moderate dataset size.
| File | Description |
|---|---|
SpaceAnimal_YOLOv8_Pose_Estimation.ipynb |
Notebook 1: dataset analysis, COCO-to-YOLO conversion, pose training and evaluation |
SpaceAnimal_Behavior_Classification.ipynb |
Notebook 2: feature extraction and classifier training/selection |
SpaceAnimal_Behavior_Inference_Pipeline.ipynb |
Notebook 3: end-to-end inference from .avi to predicted behavior |
Behavior_Analysis-LabGym.ipynb |
Exploratory analysis of the LabGym behavior dataset |
best.pt |
Trained YOLOv8s-pose weights for C. elegans |
yolov8s_pose_celegans_results.html |
Rendered pose-estimation results report |
SpaceAnimal_Rapport final.pdf |
Full project report |
SpaceAnimal_BigData_Presentation.pdf |
Project presentation slides |
The notebooks are designed to run on Google Colab or Kaggle with GPU acceleration.
pip install ultralytics==8.3.0
pip install "numpy==1.26.4"Quick inference with the trained pose model:
from ultralytics import YOLO
model = YOLO("best.pt")
results = model.predict("path/to/worm_clip.avi", conf=0.15)
for r in results:
keypoints = r.keypoints.xy # 5 body landmarks per detected wormTo reproduce the full workflow, run the notebooks in order: pose estimation, then behavior classifier training, then the inference pipeline.
- Deep learning / CV: Ultralytics YOLOv8-pose, PyTorch
- Classical ML: scikit-learn (Random Forest, Gradient Boosting), Keras/TensorFlow (LSTM)
- Data and video: OpenCV, NumPy, pandas
- Tooling: Google Colab, Kaggle, Jupyter, joblib
- Building an end-to-end vision pipeline from raw data to a deployable model artifact.
- Dataset engineering across annotation formats (COCO to YOLO-pose) and reproducible training configuration.
- Transfer learning and evaluation of a keypoint-detection model with domain-appropriate metrics.
- Feature engineering grounded in domain knowledge and rigorous model selection with cross-validation.
- Comparative analysis of classical versus deep-learning approaches, and honest reporting of the trade-offs.
- Behavior accuracy is bounded by class overlap in the feature space; minority classes such as
twitchingandimmobileremain the hardest to separate. - Temporal models (LSTM/Temporal CNN) could improve given more data or per-frame sequence features rather than aggregated clip descriptors.
- The pose model is trained only on C. elegans; extending to Drosophila and zebrafish from the SpaceAnimal dataset is a natural next step.