This repo defines a sequential multi-agent workflow using Google ADK, following the multi-agent patterns described in the ADK docs: https://google.github.io/adk-docs/agents/multi-agents/
The workflow is sequential and uses three logical agents:
- Agent 1 (Clarifier) receives the user question and asks for any needed clarifications to start the workflow.
- Agent 2 (Document Reader) reads every file in a documentation directory and stores their contents in shared session state.
- Agent 1 (File Summarizer) summarizes each document with key points that help answer the question.
- Agent 3 (Synthesizer) creates:
- an executive summary,
- a longer summary,
- references to the documents used.
agent.pyexposesroot_agent/appfor ADK CLI usage.workflow.pywires the agents into a sequential workflow.config.pycentralizes environment configuration and path normalization.agents/contains reader, clarifier, summarizer, and synthesizer agents.adk_templates/documents the instrumented LLM factory used by clarifier/summarizer/synthesizer.observability/holds OTLP setup (observability/otel_sdk.py), header parsing, ADK defaults, and session logging (readme-logs.md).tests/contains workflow and OTLP export tests.init/bootstraps Databricks UC trace tables + MLflow experiment and wires OTEL env (see below).sql/mlflow_trace_tables/holds reference SQL for UC trace table columns and views (not used to provision tables; see folder README).terraform/contains Databricks Terraform modules (catalog/schema/SQL warehouse); seeterraform/README.md.bicep/holds optional Azure templates (resource group, managed identity, etc.); seebicep/README.md. Deploy Azure-side resources first when they supply workspace URL or identities used by Terraform.scripts/contains Databricks utility scripts (see scripts/README.md).input_files/holds project documentation to summarize.
The workflow expects these environment variables:
MODEL(default:gemini-2.0-flash)DOCUMENTS_DIR(default:./input_files)MAX_FILE_CHARS(default:12000)OPENAI_API_BASE(for LM Studio / Azure Foundry, e.g.http://localhost:1234/v1)OPENAI_API_KEY(for LM Studio / Azure Foundry)OTEL_SERVICE_NAME(default: SE_workflow_test)MLFLOW_TRACING_SQL_WAREHOUSE_ID(for UC setup; find in SQL warehouse URL)DATABRICKS_CATALOG,DATABRICKS_SCHEMA(optional; for UC trace storage)
- Text formats:
.md,.txt,.rst,.log,.csv,.json,.yaml,.yml - PDFs:
.pdf(text-based PDFs only; scanned PDFs may extract no text)
- Put your project documentation in
./input_files(or setDOCUMENTS_DIR). - Install dependencies with UV:
uv sync. For unit tests (and coverage helpers), useuv sync --all-groupsoruv sync --group dev/--group testas needed. - Configure OpenTelemetry for MLflow tracing (see MLflow + Google ADK):
OTEL_EXPORTER_OTLP_ENDPOINT– OTLP traces URL (e.g..../api/2.0/otel/v1/traces)OTEL_EXPORTER_OTLP_HEADERS– headers (e.g.x-mlflow-experiment-id=<id>)- Logs are sent to
.../api/2.0/otel/v1/logs(same base URL). Requiresopentelemetry-exporter-otlp-proto-http(declared inpyproject.toml).
- From the parent directory of this repo, run
adk run SE_workflow_test. - Provide the user question as the initial message.
- If the clarifier asks questions, pass your answers by setting
clarification_answersin session state before re-running the workflow.
For Databricks, set:
OTEL_EXPORTER_OTLP_ENDPOINT="https://<workspace>/api/2.0/otel/v1/traces"OTEL_EXPORTER_OTLP_HEADERS="Authorization=Bearer <token>,x-mlflow-experiment-id=<id>"
Create the experiment in Databricks if needed, then use its ID in the header.
The init package creates the MLflow experiment (if missing), calls set_experiment_trace_location so Unity Catalog OTEL tables are created when they do not exist, then updates OTEL_EXPORTER_OTLP_HEADERS with x-mlflow-experiment-id and X-Databricks-UC-Table-Name (and sets OTEL_EXPORTER_OTLP_ENDPOINT from DATABRICKS_HOST if the endpoint is unset).
Option A — one shot from the shell (loads .env from repo root):
uv run python -m initUse uv run python -m init --dry-run to run MLflow linking steps and print planned os.environ updates without applying them. Use --quiet for warnings/errors only.
Option B — legacy script (same implementation):
uv run python scripts/setup_uc_tracing.pyOption C — before adk run, set in .env:
AUTO_CONFIGURE_DATABRICKS_TRACING=truesoagent.pyruns initialization afterload_dotenv().
Required: MLFLOW_TRACING_SQL_WAREHOUSE_ID. Token: DATABRICKS_TOKEN and DATABRICKS_HOST (if OTEL_EXPORTER_OTLP_ENDPOINT is not already set), or an existing Authorization=Bearer ... in OTEL_EXPORTER_OTLP_HEADERS. Optional: DATABRICKS_CATALOG (default: main), DATABRICKS_SCHEMA (default: mlflow_traces), MLFLOW_EXPERIMENT_ID or MLFLOW_EXPERIMENT_NAME.
The workflow stores intermediate outputs in session state:
clarificationdocumentsdocuments_jsonfile_summariesfinal_answer
To run a local model via LM Studio (OpenAI-compatible API), set:
MODEL="openai/<model-name>"(e.g.,openai/llama-3.1-8b-instruct)OPENAI_API_BASE="http://localhost:1234/v1"OPENAI_API_KEY="lm-studio"
LM Studio must be running with the local server enabled.