You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
We are recruiting independent contributors for a falsification-first evaluation of one frozen software agent on outsider-created situations. The target is stricter than patch/test success: a case passes only through an independently verified real outcome or a truth-checked non-win naming the smallest recovery condition, with zero unauthorized actions, fixed budgets, and deterministic replay. Cases should include realistic traps such as passing tests over broken behavior, incorrect documentation, misleading tickets, incomplete observability, or actions that look successful without achieving the outcome.
Needed roles: software-situation author, game author, independent custodian, outcome/non-win adjudicator, or adversarial reviewer. Authors must not have seen the candidate before sealing their case; custodian and adjudicator must be distinct. Hidden case details, evaluator logic, winning actions, secrets, and private repository content must never be posted publicly. No production access or destructive authority is requested.
Current state is deliberately negative: zero outside participants, zero capsules, zero scored attempts, and zero results. Participation does not imply endorsement, coauthorship, or a favorable finding. The public intake and review issues are at https://github.com/Parslee-ai/errata-external-evaluation/issues .
The most useful first response is a mismatch-first protocol critique or a statement of which independent role you could take. Please reply here only with public, non-sealed information.
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
We are recruiting independent contributors for a falsification-first evaluation of one frozen software agent on outsider-created situations. The target is stricter than patch/test success: a case passes only through an independently verified real outcome or a truth-checked non-win naming the smallest recovery condition, with zero unauthorized actions, fixed budgets, and deterministic replay. Cases should include realistic traps such as passing tests over broken behavior, incorrect documentation, misleading tickets, incomplete observability, or actions that look successful without achieving the outcome.
The candidate, eight-arm protocol, 12-attempt ceiling, custody rules, and an opaque 22-role capsule tool are frozen in immutable release v0.3.0: https://github.com/Parslee-ai/errata-external-evaluation/releases/tag/v0.3.0
We separately seek outside authors/custodians for an independent causal-game replication frozen in immutable v0.2.0: https://github.com/Parslee-ai/errata-external-evaluation/releases/tag/v0.2.0
Needed roles: software-situation author, game author, independent custodian, outcome/non-win adjudicator, or adversarial reviewer. Authors must not have seen the candidate before sealing their case; custodian and adjudicator must be distinct. Hidden case details, evaluator logic, winning actions, secrets, and private repository content must never be posted publicly. No production access or destructive authority is requested.
Current state is deliberately negative: zero outside participants, zero capsules, zero scored attempts, and zero results. Participation does not imply endorsement, coauthorship, or a favorable finding. The public intake and review issues are at https://github.com/Parslee-ai/errata-external-evaluation/issues .
The most useful first response is a mismatch-first protocol critique or a statement of which independent role you could take. Please reply here only with public, non-sealed information.
All reactions