Target release
Unsure
How were you running Hyperloom?
Local
Optimization domain
Inference
Issue type
Wrong result or regression
Which phase failed?
Optimization loop
Short Summary
Resuming a target_reached session with --force-resume and a higher --target-gain closes again within one phase transition, because target_reached_at is never cleared.
Command / prompt submitted
# Leg 1 — ran to completion, stopped with stop_reason=target_reached at 17.10%
python3 -m hyperloom.inference_optimizer.cli optimize \
--model /path/to/gpt-oss-120b --framework sglang --gpu-type mi355x \
--model-class moe_swa --precision mxfp4 \
--isl 8192 --osl 1024 --conc 64 --tp 1 --ep 1 \
--max-model-len 13312 --max-hours 24 --target-gain 10 \
--server-args "--moe-runner-backend aiter --page-size 64 --prefill-attention-backend triton --decode-attention-backend aiter --context-length 11264 --watchdog-timeout 1800" \
--gpu-specialist-capacity 1 --research-lane-capacity 2 \
--extra-env SGLANG_USE_AITER=1 --extra-env HF_HUB_TRUST_REMOTE_CODE=1
# Leg 2 — resume, asking it to keep going to 50% with 15h of budget
python3 -m hyperloom.inference_optimizer.cli optimize \
--resume-from "$SESSION_DIR" --force-resume \
--max-hours 15 --target-gain 50
Expected behavior
I had a session that stopped at 17.10% because it hit its --target-gain 10. I still had most of my budget left, so I wanted to keep optimizing toward a much higher target. --force-resume exists for exactly this ("you have changed the workload / search space / strategy and want to continue regardless"), so I expected the resumed run to pick up at 17.10% and keep hunting for the full 15 hours.
Actual behavior
It ran for 31 minutes and closed again with stop_reason: target_reached, still at 17.10%.
The banner said the new objective had been accepted:
Objective : kind=gain_pct target_gain_pct=50.0
Max ticks : unlimited (budget = 15.0h)
→ cleared stop_reason and reset crash_count (was 0) for fresh resume (--force-resume override)
→ reset start_ts to 2026-09-10T04:18:43.360611+00:00 (resume budget)
...but the phase machine went FRAMEWORK_AGENT → SWEEP → CLOSE more or less straight away.
The cause is that target_reached_at survives the resume. --force-resume clears stop_reason, re-anchors start_ts and clears deadline_unix, but nothing clears target_reached_at, and that field is the only thing the phase machine consults:
orchestrator/phases/machine_state.py:549 — return bool(str(getattr(state, "target_reached_at", "") or "").strip())
orchestrator/phases/machine_state.py:2942 — return PHASE_SWEEP, "target_reached", {...}
orchestrator/loop/coordinator.py:1818 — only ever sets it: if objective.reached(self.shared_state) and not self.shared_state.target_reached_at:
So on the first tick the machine reads the previous leg's timestamp and short-circuits to SWEEP, regardless of the new objective. In my state.json after the resume:
target_reached_at = "2026-09-09T20:25:10.517858+00:00" # from leg 1
cumulative_gain_validated = 17.09678314194902
target_gain_pct (banner) = 50.0
target_summary is also left stale — it still read "...drive cumulative_gain_validated to >= 10.0% within 24.0h." after resuming with --target-gain 50 --max-hours 15, so the raised target reaches the banner and the objective object but not the persisted session summary that agents and reports read.
Setting target_reached_at = "" in state.json by hand and re-running the exact same resume command fixed it — the run then stayed in FRAMEWORK_AGENT and kept optimizing, with the 17.10% and the winning recipe preserved.
Since coordinator.py:1818 re-sets the flag whenever the objective is genuinely met, clearing it on a re-anchoring resume looks safe: with target_gain_pct=50 and 17.10% banked, objective.reached() is false, so it simply wouldn't be re-set until the new target is actually hit.
Suggested fix: clear target_reached_at in _begin_resume_leg() on the reanchor_budget=True path, alongside stop_reason and deadline_unix, and refresh target_summary from the new objective/budget on resume.
Steps to reproduce
- Run
optimize with a modest --target-gain N and let it finish with stop_reason=target_reached.
- Resume it:
optimize --resume-from <session_dir> --force-resume --max-hours <lots> --target-gain <much larger than N>.
- Watch it transition
FRAMEWORK_AGENT → SWEEP → CLOSE and exit at the old gain, well inside the new budget.
Reproduced twice on the same session; clearing the field by hand made the same command behave correctly.
Logs
# leg 2 (resume) — starts fine...
Objective : kind=gain_pct target_gain_pct=50.0
Max ticks : unlimited (budget = 15.0h)
→ cleared stop_reason and reset crash_count (was 0) for fresh resume (--force-resume override)
→ reset start_ts to 2026-09-10T04:18:43.360611+00:00 (resume budget)
# ...and 31 minutes later:
================ Final summary ================
stop_reason : target_reached
baseline : 3688.8 tok/s/GPU
cumulative_gain_val : 17.10% (validated_at_stack_len=1, ts=2026-09-09T20:12:03.232157+00:00)
crash_count : 0
===============================================
Note the validated_at_stack_len=1, ts=...T20:12:03 timestamp on the final summary is from leg 1 — leg 2 never validated anything of its own.
Additional artifacts / links
- Session:
gpt-oss-120b_20260909T191958Z_484ecf65 (local run, sglang 0.5.18.dev20260825+g0c7ff19e3b, ROCm 7.2.4, MI355X, TP=1)
- Hyperloom at
074382e57
Data handling
Target release
Unsure
How were you running Hyperloom?
Local
Optimization domain
Inference
Issue type
Wrong result or regression
Which phase failed?
Optimization loop
Short Summary
Resuming a
target_reachedsession with--force-resumeand a higher--target-gaincloses again within one phase transition, becausetarget_reached_atis never cleared.Command / prompt submitted
Expected behavior
I had a session that stopped at 17.10% because it hit its
--target-gain 10. I still had most of my budget left, so I wanted to keep optimizing toward a much higher target.--force-resumeexists for exactly this ("you have changed the workload / search space / strategy and want to continue regardless"), so I expected the resumed run to pick up at 17.10% and keep hunting for the full 15 hours.Actual behavior
It ran for 31 minutes and closed again with
stop_reason: target_reached, still at 17.10%.The banner said the new objective had been accepted:
...but the phase machine went
FRAMEWORK_AGENT → SWEEP → CLOSEmore or less straight away.The cause is that
target_reached_atsurvives the resume.--force-resumeclearsstop_reason, re-anchorsstart_tsand clearsdeadline_unix, but nothing clearstarget_reached_at, and that field is the only thing the phase machine consults:orchestrator/phases/machine_state.py:549—return bool(str(getattr(state, "target_reached_at", "") or "").strip())orchestrator/phases/machine_state.py:2942—return PHASE_SWEEP, "target_reached", {...}orchestrator/loop/coordinator.py:1818— only ever sets it:if objective.reached(self.shared_state) and not self.shared_state.target_reached_at:So on the first tick the machine reads the previous leg's timestamp and short-circuits to SWEEP, regardless of the new objective. In my
state.jsonafter the resume:target_summaryis also left stale — it still read"...drive cumulative_gain_validated to >= 10.0% within 24.0h."after resuming with--target-gain 50 --max-hours 15, so the raised target reaches the banner and the objective object but not the persisted session summary that agents and reports read.Setting
target_reached_at = ""instate.jsonby hand and re-running the exact same resume command fixed it — the run then stayed inFRAMEWORK_AGENTand kept optimizing, with the 17.10% and the winning recipe preserved.Since
coordinator.py:1818re-sets the flag whenever the objective is genuinely met, clearing it on a re-anchoring resume looks safe: withtarget_gain_pct=50and 17.10% banked,objective.reached()is false, so it simply wouldn't be re-set until the new target is actually hit.Suggested fix: clear
target_reached_atin_begin_resume_leg()on thereanchor_budget=Truepath, alongsidestop_reasonanddeadline_unix, and refreshtarget_summaryfrom the new objective/budget on resume.Steps to reproduce
optimizewith a modest--target-gain Nand let it finish withstop_reason=target_reached.optimize --resume-from <session_dir> --force-resume --max-hours <lots> --target-gain <much larger than N>.FRAMEWORK_AGENT → SWEEP → CLOSEand exit at the old gain, well inside the new budget.Reproduced twice on the same session; clearing the field by hand made the same command behave correctly.
Logs
Note the
validated_at_stack_len=1, ts=...T20:12:03timestamp on the final summary is from leg 1 — leg 2 never validated anything of its own.Additional artifacts / links
gpt-oss-120b_20260909T191958Z_484ecf65(local run, sglang 0.5.18.dev20260825+g0c7ff19e3b, ROCm 7.2.4, MI355X, TP=1)074382e57Data handling