Ask questions over long videos, lectures, interviews, demos, podcasts, and research talks without paying the full cost of naive frame-by-frame multimodal processing.
Started as a Google Summer of Code 2025 project with Google DeepMind. Now maintained as an independent MIT-licensed open-source project.
HALO is a Python package for long-form video understanding. It combines semantic chunking, transcript-aware retrieval, multimodal analysis, and multi-tier caching to make video question-answering cheaper, faster, and more scalable.
Long videos are painful for AI systems.
A naive multimodal video pipeline often does the expensive thing: sample too many frames, repeatedly send similar context to an LLM, lose cross-segment memory, and pay again every time the same video is queried.
HALO takes a different approach.
Instead of treating video as a flat sequence of frames, HALO builds a hierarchy of useful context:
video
-> audio transcript
-> semantic segments
-> visual keyframes
-> cached context
-> question-aware retrieval
-> grounded answer
The goal is simple:
Make long-form video AI usable for students, researchers, developers, educators, and builders who want practical video understanding without burning unnecessary API calls.
HALO began as my Google Summer of Code 2025 project with Google DeepMind and has now moved into its next phase as a community-facing open-source project.
| Signal | Status |
|---|---|
| Package | halo-video on PyPI |
| Downloads | 5.9k+ PyPI downloads |
| License | MIT |
| Python | >=3.8 |
| Current release | 1.0.8 |
| Maintainer | Jeet Dekivadia |
| Origin | Google Summer of Code 2025 with Google DeepMind |
| Community interest | 800,000+ project/demo views, 50+ contributor leads, 3 active secondary maintainers |
Note: HALO is an independent open-source project. It originated through Google Summer of Code with Google DeepMind, but it is not an official Google or Google DeepMind product.
- Analyze YouTube videos and local video files
- Extract transcripts and visual context
- Ask questions about video content
- Use Google Gemini for multimodal reasoning
- Cache processed context across sessions
- Reduce redundant API calls through batching and reuse
- Process long-form videos with lower memory pressure
- Provide a CLI-first workflow for fast experimentation
- Serve as a base for video agents, lecture QA, research talk search, demo understanding, and multimodal retrieval systems
pip install halo-videoexport GEMINI_API_KEY="your_api_key_here"For Windows PowerShell:
$env:GEMINI_API_KEY="your_api_key_here"Get a Gemini API key here:
https://makersuite.google.com/app/apikey
halo-videoThen paste a YouTube URL or provide a local video path and start asking questions.
Video: 90-minute machine learning lecture
Question: "What intuition did the professor give for attention mechanisms?"
Video: conference talk or seminar
Question: "What are the key limitations the speaker mentions?"
Video: product walkthrough
Question: "Which features were shown, and what user problems do they solve?"
Video: long-form interview
Question: "Where do they discuss model evaluation?"
Video: technical demo
Question: "Summarize the setup steps and list anything that might break."
flowchart TD
A[Video Input] --> B[Audio and Metadata Extraction]
B --> C[Transcript Generation]
A --> D[Visual Keyframe Selection]
C --> E[Semantic Chunking]
D --> E
E --> F[Hierarchical Context Builder]
F --> G[Multi-Tier Cache]
G --> H[Question-Aware Retrieval]
H --> I[Gemini Multimodal Reasoning]
I --> J[Grounded Answer]
HALO is built around four core ideas:
| Layer | Purpose |
|---|---|
| Semantic chunking | Split long videos around meaning, not arbitrary timestamps |
| Multimodal fusion | Combine transcript context with visual keyframes |
| Context caching | Avoid reprocessing the same video and similar segments |
| Question-aware retrieval | Retrieve only the most relevant context for each query |
halo_video/
├── cli.py # Interactive command-line interface
├── config_manager.py # API key and config handling
├── context_cache.py # Multi-tier caching system
├── gemini_batch_predictor.py # Gemini API integration and batching
├── transcript_utils.py # Transcript and video processing utilities
└── __init__.py
tests/
├── test_basic.py # Core functionality tests
├── test_imports.py # Import and dependency checks
└── test_vision.py # Vision/API integration tests
demos/
├── demo.ipynb # Interactive notebook demo
├── demo.py # Minimal usage demo
└── demo_optimized.py # Optimized processing demo
docs/
├── GSoC_PROJECT_DOCUMENTATION.md
└── CONTRIBUTING.md
Most video QA demos work on short clips. HALO was designed around the harder case: lectures, technical walkthroughs, research talks, podcasts, and long YouTube videos.
HALO is not just a wrapper around a vision model. It tries to avoid waste by reducing redundant processing, caching context, and batching requests intelligently.
Many video QA systems ignore the visual stream. HALO combines speech, transcript, metadata, and selected visual frames to preserve richer context.
The project is intentionally modular. You can replace the model provider, improve chunking, add new cache backends, build a web UI, add evals, or extend it into a full video-agent framework.
git clone https://github.com/jeet-dekivadia/google-deepmind.git
cd google-deepmind
python -m venv venv
source venv/bin/activate
pip install -e ".[dev]"
pytestOn Windows:
git clone https://github.com/jeet-dekivadia/google-deepmind.git
cd google-deepmind
python -m venv venv
venv\Scripts\activate
pip install -e ".[dev]"
pytestRun the CLI locally:
python -m halo_video.cliHALO is usable today, but it is still early. The next goal is to turn it from a successful GSoC deliverable into a serious open-source video AI toolkit.
- Add a cleaner Python SDK interface
- Add provider adapters for OpenAI, Anthropic, Gemini, and local models
- Improve long-video regression tests
- Add reproducible benchmarks for cost, latency, and answer quality
- Add better cache invalidation and cache inspection tools
- Add structured output support for summaries, chapters, and citations
- Add async processing for large batch jobs
- Harden API key handling
- Audit local file and path handling
- Add URL validation for remote video sources
- Add dependency scanning and supply-chain checks
- Add safer temporary file cleanup
- Add prompt-injection tests for untrusted transcripts and video content
- Improve onboarding docs
- Add more demos and example notebooks
- Create
good first issuetasks - Add a simple web UI
- Add plugin hooks for researchers and developers
- Publish a full technical architecture guide
HALO welcomes contributors.
Good starting points:
- Improve the quickstart and docs
- Add tests for edge cases
- Add examples for different video types
- Improve transcript handling
- Add model-provider adapters
- Build a web UI
- Improve security around file handling, secrets, and untrusted URLs
- Add benchmarks for long-video cost and latency
To contribute:
- Fork the repository
- Create a branch
- Make a focused change
- Run tests
- Open a pull request with a clear explanation
git checkout -b feat/my-improvement
pytestRead the contributing guide:
docs/CONTRIBUTING.md
Open issues:
https://github.com/jeet-dekivadia/google-deepmind/issues
HALO was originally built during Google Summer of Code 2025 with Google DeepMind.
| Field | Details |
|---|---|
| Program | Google Summer of Code 2025 |
| Organization | Google DeepMind |
| Contributor | Jeet Dekivadia |
| Mentor | Paige Bailey |
| Timeline | May 2025 to September 2025 |
| Progress tracker | GSoC Progress Tracker |
| Final package | halo-video |
The original GSoC goal was to explore hierarchical abstraction for efficient long-form video analysis. The project delivered a working Python package, documentation, demos, tests, and a PyPI release.
The next goal is broader: make HALO useful to the open-source community.
HALO has grown beyond a summer research deliverable.
| Metric | Current signal |
|---|---|
| PyPI package | halo-video |
| PyPI downloads | 5.9k+ |
| Project/demo reach | 800,000+ views across project posts and demos |
| Contributor interest | 50+ people interested in contributing |
| Maintainer group | 1 primary maintainer, 3 active secondary maintainers |
| License | MIT |
| Focus | Efficient long-form video QA and multimodal context systems |
If you use HALO in research, teaching, demos, or derivative open-source work, you can cite it as:
@software{dekivadia2025halo,
author = {Dekivadia, Jeet},
title = {HALO: Hierarchical Abstraction for Longform Optimization},
year = {2025},
url = {https://github.com/jeet-dekivadia/google-deepmind},
note = {Google Summer of Code 2025 project with Google DeepMind}
}- PyPI: https://pypi.org/project/halo-video/
- Downloads: https://pepy.tech/projects/halo-video
- Repository: https://github.com/jeet-dekivadia/google-deepmind
- GSoC: https://summerofcode.withgoogle.com/
- Google DeepMind: https://deepmind.google/
- Progress tracker: https://docs.google.com/document/d/1QOIEO70PyZwIOS5W2nZWcum9mdTPrMzWScX19IaovIE/edit?usp=sharing
- Issues: https://github.com/jeet-dekivadia/google-deepmind/issues
Built and maintained by Jeet Dekivadia.
HALO started as a GSoC project. The next chapter is open source.