Run ComfyUI on Modal.com with auto-scaling, GPU snapshots, and easy model management.
Good for testing Wan2.2/LTX-2.5/MiniMax-H3 or other image/video/audio generation models.
Note: If you want to work (creating/editing) on your workflows locally while seamlessly runs the workflow on Cloud GPU, you can install ComfyUI-Proxy custom node on your local ComfyUI https://github.com/anr2me/comfyui-proxy
(ComfyGPU's URL for Remote GPU URL, and optionally ComfyCPU's URL for Remote CPU URL).
- Clone this repository:
git clone https://github.com/anr2me/modal-comfyui.git cd modal-comfyui - Install the Modal client:
uv sync
- Set up your modal account (if not done already):
modal setup
Copy models.example.py to models.py and edit it to manage your models. You can specify:
- Hugging Face models(
models) usingrepo_idandfilename. Set yourHF_TOKENinhuggingface-secretSecrets for a faster download speed and gated models. - External models(
models_ext, e.g. civitai) using a directurl. You can also set yourCIVITAI_TOKENincustom-secretSecrets to download gated models.
Models are downloaded to persistent volumes and symlinked to the specified model_dir.
model_dir accepts two styles:
- Relative path (recommended for standard ComfyUI folders): resolved under
/root/comfy/ComfyUI/models/. e.g."checkpoints"→/root/comfy/ComfyUI/models/checkpoints. - Absolute path: used as-is. Use this when the target lives outside
ComfyUI/models/(e.g. a custom node's own model directory).
See models.example.py for reference.
Copy plugins.example.py to plugins.py and edit it to add custom node IDs or titles to be installed via comfy-cli.
- Workflow Dependencies: If you have a
workflow_api.jsonin the root directory, the setup will automatically install the necessary custom nodes for that workflow.
Open ComfyUI manager on comfyui and click "Used in Workflow" to see which custom nodes are used in the workflow.
*Note: We are using Legacy ComfyUI Manager.
- Add these custom nodes to
comfy_pluginsinplugins.py(be careful of node id). You can find the node id at https://registry.comfy.org/ - You can also install custom nodes repository using
giturl. Add the url, branch, and their dependencies tocomfy_plugins_extinplugins.py(be careful of dependency conflicts).
Run the following command to start ComfyUI in development mode:
modal serve comfyui.pyThis will provide a temporary URL where you can access the ComfyUI interface.
To deploy ComfyUI as a persistent app using the default L4 GPU:
modal deploy comfyui.pyOr rebuild the image and deploy (ie. after removing models at persistent volume to redownload them):
MODAL_FORCE_BUILD=1 modal deploy comfyui.pyOr deploy with cleared shared_dict (ie. when the App forcefully stopped):
python comfyui.pyOr change the GPU with:
MODAL_GPU=RTX-PRO-6000 modal deploy comfyui.pyYou can find the GPU types available on modal.com at https://modal.com/docs/guide/gpu
Other Environment Variables you can use are:
COMFY_VER="latest" # or "nightly --commit <hash>" or "nightly --pr kijai:minimax_fun"
COMFYGPU_ARGS="--use-flash-attention --preview-method auto --front-end-version Comfy-Org/ComfyUI_frontend@1.45.21"
COMFYMIX_ARGS="--preview-method auto --front-end-version Comfy-Org/ComfyUI_frontend@1.45.21"
JOBS_CUTOFFTIME=172800
MODAL_MAXTIME=3600
MODAL_IDLETIME=38
MODAL_WAITTIME=20
MODAL_MAXSTARTTIME=300You can access ComfyUI from the provided persistent URL when successfully deployed.
- Auto-scaling: Scales down to zero when not in use to save costs (modal's serverless can also auto-scales vertically, where CPU cores and RAM size can grow automatically as needed, so you don't need to overprovision them).
- GPU Snapshots: Fast startup times using Modal's GPU snapshots (cold-start can be under 3 seconds).
- Model Caching: Uses Modal Volumes to cache models across runs (modal's persistent volume is free for the first 1 TiB).
- Custom Node Management: Integrated with
comfy-clifor easy plugin installation. - Mixed CPU and GPU instance: Create/Edit your workflows using CPU-only instance for cheaper rates, but runs workflows on GPU instance seamlessly. Also have persistent completed jobs across sessions with their output assets accessible from Media Assets panel.
- Pre-installed Wheels:
- PyTorch+CUDA 13.0
- FlashAttention 2.8.3, 3, and 4
- SageAttention 2.2 and 3
- SpargeAttention 0.1
- nunchaku 1.2.1
- llama-cpp-python
Please feel free to contribute to make this project better. Performance improvements/optimizations are very welcome.