A benchmark testing whether LLM SQL agents can answer causal 'why did this metric move?' questions — scored against a deterministic decomposition engine.
data-science benchmark analytics evaluation causal-inference lmdi root-cause-analysis llm llm-evaluation bi-agents ai-data-analyst sql-agents
-
Updated
Jun 21, 2026 - Python