Platform for deploying and routing GPU-accelerated inference, streaming, and batch workloads at scale.
-
Updated
Oct 11, 2026 - Go
Platform for deploying and routing GPU-accelerated inference, streaming, and batch workloads at scale.
Run AI models anywhere. https://muna.ai/explore
Run ComfyUI on Modal with auto-scaling, GPU snapshots, and easy model management. Try image, video generation via ComfyUI on Modal.
Own your AI video pipeline. LTX-2.3 (22B) self-hosted on your Modal GPU via a Claude Code skill — t2v, i2v, keyframes, v2v + synced audio. ~$0.02 per 5s clip, idle = $0.
Single-user training and generation platform on Modal. Train Krea 2 LoRAs, generate stills, and turn them into video with MiniMax-H3 (sound and picture in one pass) — one deploy, one URL, no infrastructure to keep alive.
State-aware hedged requests for serverless GPU inference — return the first valid result and cancel the losers with an audited receipt.
Automated daily AI research engine powered by CrewAI & serverless Nvidia T4 GPUs.
AI-powered audio transcription, voice cloning, and image generation on Modal serverless GPUs. Real-time streaming, speaker diarization, meeting minutes, saved voice profiles, and FLUX.1 image gen — all in one service.
Queue-driven, scale-from-zero GPU inference for any Kubernetes — bursts to cross-region VMs when GPUs run dry
An open-source, BYOK YouTube thumbnail and channel analyzer powered by the TRIBE v2 neuroscience model, Next.js, and serverless GPU inference via Modal.
🔎 Search your photos by meaning — semantic photo search powered by CLIP on Runpod Flash serverless GPUs
Compare two videos by the brain response Meta's TRIBE v2 predicts. Serverless GPU backend on Modal, 3D cortex viewer in React + three.js.
WhisperX + pyannote meeting transcription on a scale-to-zero RunPod GPU, with R2 storage and a Next.js UI. $0 idle.
High-Performance Serverless event and data processing platform
Run Metaflow steps on serverless GPUs — no infrastructure to manage
One-command PowerShell deployment of Ollama running gpt-oss:20b (or any Ollama model) on Azure Container Apps serverless GPU — nginx API-key auth proxy, scale-to-zero billing.
A cost-effective, serverless AI image generation pipeline using local n8n and Modal.com to run the uncensored FLUX.2-klein-9B model on cloud A100 GPUs.
AI-powered genetic variant analysis platform using Stanford's Evo 2 model to predict mutation pathogenicity. Built with Next.js, FastAPI, and Serverless GPU acceleration for real-time genomic research and clinical decision support.
Serverless GPU cold start latency and cost benchmark for LLM inference (Modal, RunPod, Replicate, Together AI).
AI-powered medical imaging platform for DICOM ingestion, MONAI multi-organ segmentation, and MPR visualization with PHI de-identification, real-time inference streaming, and clinical PDF reporting.
To associate your repository with the serverless-gpu topic, visit your repo's landing page and select "manage topics."