Skip to content

--force-resume with a raised --target-gain closes immediately: target_reached_at survives the resume #1469

Description

@rodosingh

Target release

Unsure

How were you running Hyperloom?

Local

Optimization domain

Inference

Issue type

Wrong result or regression

Which phase failed?

Optimization loop

Short Summary

Resuming a target_reached session with --force-resume and a higher --target-gain closes again within one phase transition, because target_reached_at is never cleared.

Command / prompt submitted

# Leg 1 — ran to completion, stopped with stop_reason=target_reached at 17.10%
python3 -m hyperloom.inference_optimizer.cli optimize \
  --model /path/to/gpt-oss-120b --framework sglang --gpu-type mi355x \
  --model-class moe_swa --precision mxfp4 \
  --isl 8192 --osl 1024 --conc 64 --tp 1 --ep 1 \
  --max-model-len 13312 --max-hours 24 --target-gain 10 \
  --server-args "--moe-runner-backend aiter --page-size 64 --prefill-attention-backend triton --decode-attention-backend aiter --context-length 11264 --watchdog-timeout 1800" \
  --gpu-specialist-capacity 1 --research-lane-capacity 2 \
  --extra-env SGLANG_USE_AITER=1 --extra-env HF_HUB_TRUST_REMOTE_CODE=1

# Leg 2 — resume, asking it to keep going to 50% with 15h of budget
python3 -m hyperloom.inference_optimizer.cli optimize \
  --resume-from "$SESSION_DIR" --force-resume \
  --max-hours 15 --target-gain 50

Expected behavior

I had a session that stopped at 17.10% because it hit its --target-gain 10. I still had most of my budget left, so I wanted to keep optimizing toward a much higher target. --force-resume exists for exactly this ("you have changed the workload / search space / strategy and want to continue regardless"), so I expected the resumed run to pick up at 17.10% and keep hunting for the full 15 hours.

Actual behavior

It ran for 31 minutes and closed again with stop_reason: target_reached, still at 17.10%.

The banner said the new objective had been accepted:

Objective       : kind=gain_pct target_gain_pct=50.0
Max ticks       : unlimited (budget = 15.0h)
  → cleared stop_reason and reset crash_count (was 0) for fresh resume (--force-resume override)
  → reset start_ts to 2026-09-10T04:18:43.360611+00:00 (resume budget)

...but the phase machine went FRAMEWORK_AGENT → SWEEP → CLOSE more or less straight away.

The cause is that target_reached_at survives the resume. --force-resume clears stop_reason, re-anchors start_ts and clears deadline_unix, but nothing clears target_reached_at, and that field is the only thing the phase machine consults:

  • orchestrator/phases/machine_state.py:549 — return bool(str(getattr(state, "target_reached_at", "") or "").strip())
  • orchestrator/phases/machine_state.py:2942 — return PHASE_SWEEP, "target_reached", {...}
  • orchestrator/loop/coordinator.py:1818 — only ever sets it: if objective.reached(self.shared_state) and not self.shared_state.target_reached_at:

So on the first tick the machine reads the previous leg's timestamp and short-circuits to SWEEP, regardless of the new objective. In my state.json after the resume:

target_reached_at         = "2026-09-09T20:25:10.517858+00:00"   # from leg 1
cumulative_gain_validated = 17.09678314194902
target_gain_pct (banner)  = 50.0

target_summary is also left stale — it still read "...drive cumulative_gain_validated to >= 10.0% within 24.0h." after resuming with --target-gain 50 --max-hours 15, so the raised target reaches the banner and the objective object but not the persisted session summary that agents and reports read.

Setting target_reached_at = "" in state.json by hand and re-running the exact same resume command fixed it — the run then stayed in FRAMEWORK_AGENT and kept optimizing, with the 17.10% and the winning recipe preserved.

Since coordinator.py:1818 re-sets the flag whenever the objective is genuinely met, clearing it on a re-anchoring resume looks safe: with target_gain_pct=50 and 17.10% banked, objective.reached() is false, so it simply wouldn't be re-set until the new target is actually hit.

Suggested fix: clear target_reached_at in _begin_resume_leg() on the reanchor_budget=True path, alongside stop_reason and deadline_unix, and refresh target_summary from the new objective/budget on resume.

Steps to reproduce

  1. Run optimize with a modest --target-gain N and let it finish with stop_reason=target_reached.
  2. Resume it: optimize --resume-from <session_dir> --force-resume --max-hours <lots> --target-gain <much larger than N>.
  3. Watch it transition FRAMEWORK_AGENT → SWEEP → CLOSE and exit at the old gain, well inside the new budget.

Reproduced twice on the same session; clearing the field by hand made the same command behave correctly.

Logs

# leg 2 (resume) — starts fine...
Objective       : kind=gain_pct target_gain_pct=50.0
Max ticks       : unlimited (budget = 15.0h)
  → cleared stop_reason and reset crash_count (was 0) for fresh resume (--force-resume override)
  → reset start_ts to 2026-09-10T04:18:43.360611+00:00 (resume budget)

# ...and 31 minutes later:
================ Final summary ================
  stop_reason          : target_reached
  baseline             : 3688.8 tok/s/GPU
  cumulative_gain_val  : 17.10% (validated_at_stack_len=1, ts=2026-09-09T20:12:03.232157+00:00)
  crash_count          : 0
===============================================

Note the validated_at_stack_len=1, ts=...T20:12:03 timestamp on the final summary is from leg 1 — leg 2 never validated anything of its own.

Additional artifacts / links

  • Session: gpt-oss-120b_20260909T191958Z_484ecf65 (local run, sglang 0.5.18.dev20260825+g0c7ff19e3b, ROCm 7.2.4, MI355X, TP=1)
  • Hyperloom at 074382e57

Data handling

  • I confirmed the attached logs/artifacts contain no sensitive information
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

domain:inferenceRelated to inference optimizationtype:bugSomething is broken or not working as expected

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions