Reference implementation for Textual Equilibrium Propagation for Deep Compound AI Systems
Accepted to ICLR 2026
TEP is a local learning framework for optimizing deep compound AI systems built from black-box LLM modules. Instead of relying on global textual backpropagation, TEP separates optimization into a free phase for local equilibration and a nudged phase for bounded task-driven updates, improving stability as workflow depth grows.
This repository accompanies the paper Textual Equilibrium Propagation for Deep Compound AI Systems by Minghui Chen, Wenlong Deng, James Zou, Han Yu, and Xiaoxiao Li, and includes runnable pipelines for HotpotQA, PubMedQA, STARK-PRIME, and BigCodeBench.
Overview of TEP: local free-phase optimization followed by nudged, bounded prompt updates for deep compound AI workflows.
- Implementation of Textual Equilibrium Propagation for deep compound AI systems.
- Supports four benchmark settings: HotpotQA, PubMedQA, STARK-PRIME, and BigCodeBench.
- Includes dedicated training, evaluation, and multi-dataset experiment runners.
- Vendors the TextGrad runtime under
src/tep/vendor/textgradfor reproducible prompt optimization workflows.
| Dataset | Task Type | Config | Notes |
|---|---|---|---|
| HotpotQA | Multi-hop question answering | configs/hotpotqa_tep.yaml |
Loaded via HuggingFace datasets |
| PubMedQA | Biomedical question answering | configs/pubmed_tep.yaml |
Expects local JSONL files under data/pubmed/ |
| STARK-PRIME | Retrieval-style QA over a knowledge base | configs/stark_tep.yaml |
Requires stark_qa; VSS mode also needs STaRK embeddings |
| BigCodeBench | Code generation and repair | configs/bigcodebench_tep.yaml |
Loaded via HuggingFace datasets with the fixed benchmark split |
pip install -r requirements.txt
pip install -e .Create a .env file (or pass --dotenv_path) for model access:
OPENROUTER_API_KEY=your_key_here
OPENROUTER_DEFAULT_MODEL=openrouter/openai/gpt-4o-miniOptional dataset-specific extras:
pip install stark_qa huggingface_hubNotes:
- STARK-PRIME support depends on
stark_qaand, for VSS retrieval mode, the corresponding embedding assets. - BigCodeBench evaluation may require additional tooling from the official BigCodeBench ecosystem depending on how you reproduce the benchmark.
- The CLI entrypoints read
.envby default, so environment-based model configuration works out of the box.
Train TEP on HotpotQA:
tep-train configs/hotpotqa_tep.yaml \
--dataset hotpotqa \
--workflow_name hotpotqa \
--optimization_mode tepEvaluate a configured workflow:
tep-evaluate configs/hotpotqa_tep.yaml \
--dataset hotpotqa \
--workflow_name hotpotqaIf you prefer direct Python invocation, the same entrypoints are available under src/tep/cli/.
configs/ Dataset-specific experiment configs
data/ Local data assets (for example PubMedQA files)
src/tep/cli/ Training, evaluation, and experiment entrypoints
src/tep/runtime/ Core runtime, logging, registry, and split utilities
src/tep/tasks/ Benchmark-specific datasets and workflows
src/tep/metrics/ Evaluation metrics
src/tep/vendor/textgrad/ Vendored TextGrad runtime
assets/readme/ README visual assets
If you find this repository useful, please cite:
@inproceedings{chen2026textual,
title = {Textual Equilibrium Propagation for Deep Compound AI Systems},
author = {Chen, Minghui and Deng, Wenlong and Zou, James and Yu, Han and Li, Xiaoxiao},
booktitle = {International Conference on Learning Representations (ICLR)},
year = {2026},
url = {https://arxiv.org/abs/2601.21064}
}This project is released under the MIT License. See LICENSE for details.
