Summary
The harness can classify a CLI process that exited nonzero as a successful agent session when the driver does not return a fatal message.
On the initial attempt, the generic fallback synthesizes a failure only when the exit code is nonzero and both reported token counts are zero. A failed process that produced tokens bypasses fatal and retry handling.
The retry path is less strict: it logs retry: session succeeded whenever agent_detect_fatal returns an empty string, without checking AGENT_RUN_EXIT. A retried process can therefore exit nonzero with either zero or nonzero usage and still end the backoff loop as a success.
This makes success depend on each driver recognizing every terminal failure shape. An uncaught CLI error, signal exit, truncated structured stream, or incomplete detector can turn a failed or partial run into a successful session. The concrete Kimi behavior and reproduction are documented in this PR #113 review comment.
Steps to reproduce
- Run an agent with
max_retry_wait enabled.
- Have the first CLI attempt exit nonzero with zero usage while
agent_detect_fatal returns an empty string. The generic zero-token fallback correctly starts the retry loop.
- Have the retry exit nonzero again while
agent_detect_fatal returns an empty string.
- Observe that the harness logs
retry: session succeeded and leaves the retry loop despite the nonzero exit.
The initial-attempt variant requires no retry: return nonzero with nonzero input or output usage while agent_detect_fatal returns an empty string. The generic fallback does not run, so the failed attempt proceeds as a normal session.
Expected vs actual behavior
Expected: a session is successful only when the CLI exits zero and the driver reports no semantic fatal error. A nonzero exit without a driver message should produce a generic failure regardless of token usage, then pass through the existing retry classifier. Driver detectors must remain able to reject exit-zero results that encode failures, such as quota exhaustion reported in a successful result. The same post-run outcome logic should run after the initial attempt and every retry.
Actual: the initial path accepts otherwise-undetected nonzero exits after token use, and the retry path accepts any otherwise-undetected nonzero exit. This can stop retrying, count an incomplete run as normal, and continue to commit shipping or idle-success handling.
The fix should centralize the post-run outcome decision and use it from both paths. Tests should cover initial and retried nonzero exits with both zero and nonzero usage, plus exit-zero semantic failures. The Claude Code detector should also be audited because it suppresses structured errors when a result reports nonzero usage.
Version
0.22.0
Summary
The harness can classify a CLI process that exited nonzero as a successful agent session when the driver does not return a fatal message.
On the initial attempt, the generic fallback synthesizes a failure only when the exit code is nonzero and both reported token counts are zero. A failed process that produced tokens bypasses fatal and retry handling.
The retry path is less strict: it logs
retry: session succeededwheneveragent_detect_fatalreturns an empty string, without checkingAGENT_RUN_EXIT. A retried process can therefore exit nonzero with either zero or nonzero usage and still end the backoff loop as a success.This makes success depend on each driver recognizing every terminal failure shape. An uncaught CLI error, signal exit, truncated structured stream, or incomplete detector can turn a failed or partial run into a successful session. The concrete Kimi behavior and reproduction are documented in this PR #113 review comment.
Steps to reproduce
max_retry_waitenabled.agent_detect_fatalreturns an empty string. The generic zero-token fallback correctly starts the retry loop.agent_detect_fatalreturns an empty string.retry: session succeededand leaves the retry loop despite the nonzero exit.The initial-attempt variant requires no retry: return nonzero with nonzero input or output usage while
agent_detect_fatalreturns an empty string. The generic fallback does not run, so the failed attempt proceeds as a normal session.Expected vs actual behavior
Expected: a session is successful only when the CLI exits zero and the driver reports no semantic fatal error. A nonzero exit without a driver message should produce a generic failure regardless of token usage, then pass through the existing retry classifier. Driver detectors must remain able to reject exit-zero results that encode failures, such as quota exhaustion reported in a successful result. The same post-run outcome logic should run after the initial attempt and every retry.
Actual: the initial path accepts otherwise-undetected nonzero exits after token use, and the retry path accepts any otherwise-undetected nonzero exit. This can stop retrying, count an incomplete run as normal, and continue to commit shipping or idle-success handling.
The fix should centralize the post-run outcome decision and use it from both paths. Tests should cover initial and retried nonzero exits with both zero and nonzero usage, plus exit-zero semantic failures. The Claude Code detector should also be audited because it suppresses structured errors when a result reports nonzero usage.
Version
0.22.0