Problem
What the agent can do only reaches a chat when the agent object is rebuilt. The gateway builds one agent per profile at start (Gateway._agent) plus one per LLM config (_model_agents) and keeps them until ProfileManager.reload. So a change to skills, MCP servers, tools or Cards applies on the next reload at best — and some never reach a running profile:
| Input |
Reaches a chat today |
Persona, per-turn guidance, the rich-views switch, A2UI middleware, permissions/HITL |
every turn |
| LLM config, tools, Google tools, MCP servers, Docker sandbox, memory store, skills catalog |
on reload |
A Card added, removed or switched on/off (the rich-views description, ADR 0039) |
never until reload — nothing triggers one |
A skill the agent installs itself via install_skill, a skill edited on disk |
never until reload |
Settings routes for voice, connections, profiles and ACP change inputs without reloading at all.
Root of it: AG2's SkillPlugin freezes both the <available_skills> catalog (a static prompt string) and the load_skill name set at construction, and AG2's dynamic prompts only run when a turn passes no prompt= — the gateway always passes one (it now prepends agent.system_prompt itself, see #117).
Direction (to be verified)
Build the agent per turn and keep only the long-lived handles in a per-profile cache.
Measured on a real profile (openai, local sandbox, no MCP): create_agent is 120–200 ms warm (~390 ms cold), and ~75% of that is re-parsing every Card file because resolve_a2ui_skill builds a fresh CardCatalog per build instead of sharing the gateway's fingerprinted one. With the catalog shared, a build is ~30–40 ms. AG2 already creates the LLM client per turn; MCP servers start lazily.
Per-turn construction must not recreate what outlives a turn:
- MCP toolkits hold one live connection (300 s idle close) — rebuilding would respawn the server every turn and leave the old one alive for up to 5 min.
DockerEnvironment caches its container until process exit — one leaked container per turn.
- ACP model configs (Claude Code / Codex) hold per-chat sessions in
ACPConfig._sessions — a new config per turn loses the session and spawns a new subprocess.
- Memory store — one SQLite connection per profile, not per turn.
Sketch:
Gateway.send_message builds the turn's agent (_make_agent(cfg_for_turn)); drop _agent, _model_agents, the rebuild in _ensure_subscription_fresh and the rebuild half of reload().
- A per-profile resource cache, created in
Gateway.start and passed into create_agent as a parameter (no globals, testable per AGENTS.md): the shared CardCatalog, the memory store, MCP toolkits keyed by server settings (removed/changed ones closed), DockerEnvironment keyed by image+network, model configs keyed by config id + token. close() closes all of it.
- Callers that hold an agent (
require_agent(): ACP listener, voice tool names, A2UI server actions) get a builder instead.
- Remove the
reload calls the routes make only to pick up a change.
Tests (no monkeypatch): a real Gateway on temp Paths with a fake model config that records its prompt — send, drop a Card / switch a skill off via SkillStateStore, send again, assert the second prompt's catalog changed; an MCP stub via write_stub counting spawns — two turns, one spawn; removing the server closes it.
Existing bugs found on the way (worth fixing regardless)
_aclose_agents closes only the model config, so every reload today leaks MCP servers (until idle timeout) and Docker containers (until exit).
- The ACP listener captures the agent once at start (
profile_manager.py) and never sees a reload.
Upstream (AG2) opportunities
Some of this may belong in AG2 rather than here:
Related
Problem
What the agent can do only reaches a chat when the agent object is rebuilt. The gateway builds one agent per profile at start (
Gateway._agent) plus one per LLM config (_model_agents) and keeps them untilProfileManager.reload. So a change to skills, MCP servers, tools or Cards applies on the next reload at best — and some never reach a running profile:rich-viewsswitch, A2UI middleware, permissions/HITLrich-viewsdescription, ADR 0039)install_skill, a skill edited on diskSettings routes for voice, connections, profiles and ACP change inputs without reloading at all.
Root of it: AG2's
SkillPluginfreezes both the<available_skills>catalog (a static prompt string) and theload_skillname set at construction, and AG2's dynamic prompts only run when a turn passes noprompt=— the gateway always passes one (it now prependsagent.system_promptitself, see #117).Direction (to be verified)
Build the agent per turn and keep only the long-lived handles in a per-profile cache.
Measured on a real profile (openai, local sandbox, no MCP):
create_agentis 120–200 ms warm (~390 ms cold), and ~75% of that is re-parsing every Card file becauseresolve_a2ui_skillbuilds a freshCardCatalogper build instead of sharing the gateway's fingerprinted one. With the catalog shared, a build is ~30–40 ms. AG2 already creates the LLM client per turn; MCP servers start lazily.Per-turn construction must not recreate what outlives a turn:
DockerEnvironmentcaches its container until process exit — one leaked container per turn.ACPConfig._sessions— a new config per turn loses the session and spawns a new subprocess.Sketch:
Gateway.send_messagebuilds the turn's agent (_make_agent(cfg_for_turn)); drop_agent,_model_agents, the rebuild in_ensure_subscription_freshand the rebuild half ofreload().Gateway.startand passed intocreate_agentas a parameter (no globals, testable per AGENTS.md): the sharedCardCatalog, the memory store, MCP toolkits keyed by server settings (removed/changed ones closed),DockerEnvironmentkeyed by image+network, model configs keyed by config id + token.close()closes all of it.require_agent(): ACP listener, voice tool names, A2UI server actions) get a builder instead.reloadcalls the routes make only to pick up a change.Tests (no monkeypatch): a real
Gatewayon tempPathswith a fake model config that records its prompt — send, drop a Card / switch a skill off viaSkillStateStore, send again, assert the second prompt's catalog changed; an MCP stub viawrite_stubcounting spawns — two turns, one spawn; removing the server closes it.Existing bugs found on the way (worth fixing regardless)
_aclose_agentscloses only the model config, so every reload today leaks MCP servers (until idle timeout) and Docker containers (until exit).profile_manager.py) and never sees a reload.Upstream (AG2) opportunities
Some of this may belong in AG2 rather than here:
SkillPluginrendering its catalog andload_skillnames per turn (a dynamic prompt/tool schema over the runtime), instead of freezing them at construction. — [Feature Request]: SkillPlugin renders its skills catalog per turn ag2#3325prompt=(or a way to extend rather than replace the agent's prompt per turn). — [Feature Request]: A turn's prompt= should extend the agent's prompt, not silently replace it ag2#3324Agentobject, or cheapAgentre-derivation that shares them.Related
rich-viewsSkill whose description is rendered from the Cards (ADR 0039); its description is a construction-time snapshot because of this.