Ultrafuzz eval suites benchmark the fuzzing pipeline against targets with known ground-truth bugs. Run journals retain node lifecycle, heartbeat, and artifact evidence locally. Scoring, comparison, reports, and history work without an external reporting account.
The eval YAML defines the experiment and grading policy. Project TOML supplies the default suite location and machine-specific ground-truth directory.
[eval]
eval_config = ".ultrafuzz/evals/bug-finding.yml"
ground_truth_root = "/secure/eval-ground-truth"
provider = "none"- Only
provider = "none"is supported. No external reporter is installed. - Precedence: CLI flag (
--provider,--suite) > environment (ULTRAFUZZ_EVAL_PROVIDER,ULTRAFUZZ_EVAL_CONFIG) > project TOML. An unsupported final provider selection fails before credential access or workflow launch; an explicit--provider nonecan override old settings. - Ground truth must live outside the repository. Suite targets reference files
relative to
ground_truth_root; absolute entries, traversal, symlinks, non-regular files, and files larger than 1 MiB are rejected. - Generic connection metadata in saved configuration remains readable but cannot activate an external reporter. Fresh defaults contain no connection profiles. See the migration guidance.
See .ultrafuzz/evals/bug-finding.yml for the default suite. It defines model
profiles, targets (repo/ref/ground truth/sensitivity), variants, trial counts,
grading metrics, and the reporting: block. Nothing in the YAML names a
provider, an endpoint, or an env var.
reporting.node_telemetry decides whether ultrafuzz eval run watches
launched rows by default; --no-watch overrides it. It and
reporting.heartbeat_interval_seconds remain inputs to the execution-policy
fingerprint. reporting.experiment_prefix and reporting.artifacts are still
validated, but no eval behaviour depends on them.
The current suite version is exactly ultrafuzz.eval.v2, validated against
urn:ultrafuzz:schema:evals:suite:2. JSON Schema is canonical. The retained
Zod parser is strict and non-transforming, and a shared acceptance corpus keeps
it aligned with the JSON Schema. Defaults are applied only after both shape
validators accept the operator-authored document.
All eval-owned objects are closed. In particular, model profiles contain only
agent, optional model, optional reasoning, and optional
timeout_seconds; variants contain only id, optional topology, optional
workflow_input, and optional runner/judge profile overrides; metrics
contains only recall_threshold. Historical no-op fields such as model
config, variant prompts or model_profiles, root prompt_overlays, and
metrics.primary/metrics.secondary are rejected. ground_truth_root is
machine-specific TOML/CLI configuration and is not a suite-YAML field.
Ordinary operator-defined workflow_input remains an intentional extension
seam. It must be a JSON object with nonempty, non-whitespace keys; its values
may be nested JSON objects, arrays, strings, numbers, booleans, or null. It may
not use the eval-owned reserved keys benchmark_execution, benchmark_lane,
excluded_strategy_families, target_frameworks, or ultrafuzz_eval.
Scalars and arrays at the workflow_input root are rejected instead of being
silently discarded.
Benchmark variants use one of three exact shapes. Private benchmark controls contain only the execution budget:
workflow_input:
benchmark_execution:
strategy_loops: 2
excluded_node_ids: [optional-analysis]The public full lane contains all fields below and permits neither excluded families nor excluded nodes:
workflow_input:
benchmark_lane: full
target_frameworks: { target-a: foundry }
excluded_strategy_families: []
benchmark_execution:
strategy_loops: 1
excluded_node_ids: []The public smoke lane fixes the workflow profile, all three excluded strategy families, and all four selected strategies:
workflow_input:
benchmark_lane: smoke
target_frameworks: { target-a: foundry }
excluded_strategy_families: [stateful-invariant, differential, dynamic-strategy]
benchmark_execution:
workflow_profile: smoke-benchmark-v1
audit_profile: smoke
audit_profile_catalog_digest: <64-lowercase-hex catalog digest>
topology_digest: <64-lowercase-hex packaged smoke topology digest>
selected_strategy_ids:
- time-warp-sequences
- external-dependency-boundaries
- externalized-state-accounting
- lifecycle-view-boundaries
strategy_loops: 1
excluded_node_ids: []Missing fields, extra fields, duplicate array entries, wrong lane constants,
reserved operator keys, and historical version literals fail validation. There
is no alias conversion, compatibility fallback, or repair pass.
The checked-in benchmark adapter preserves that released full-lane shape. The
launcher derives the current packaged exhaustive policy from benchmark_lane: full;
it refuses to launch public variants that declare a topology override. Before
startRun, it resolves the target's effective policy, requires the topology to
originate from that audit profile, and refuses any audit-profile, catalog, or
topology digest mismatch. The run journal captures that actual origin and
policy, and public-history publication checks every launched record against the
candidate checkout's packaged policy, so a stale or target-local override
cannot silently replace the packaged graph.
A benchmark that ships a reference solution beside the code under test would otherwise hand the run its own answer key. A target may withhold those paths:
targets:
- id: target
repo: "https://github.com/scfuzzbench/aave-v4-scfuzzbench"
ref: "edd6c82721512540c8c90e7a36a4a8e19fd7bdf3"
ground_truth: aave-v4/findings.yml
held_out_paths: ["tests/recon"]Materializing a pinned target rewrites the checkout to a parentless revision that never contained the declared paths, then destroys the benchmark commit and the withheld blobs.
A commit is required rather than a dirty worktree, because every node workspace
is a Git worktree of the pinned branch and would otherwise restore the files.
Keeping the benchmark commit as a parent is equally unsafe — git show HEAD^:tests/recon/Properties.sol would hand the answer key straight back — so
the rewritten revision has no parents and the original objects are pruned along
with the reflog. The single-revision isolation invariants are unchanged:
revision_count and commit_object_count stay 1.
The source proof records what was withheld:
"held_out": {
"source_commit": "edd6c82721512540c8c90e7a36a4a8e19fd7bdf3",
"source_tree": "3b3910d0657e4a639888ccae45822af680de6ef6",
"commit": "ac548733fd9c36cf9a32284a653a6429c089abfa",
"tree": "5e9c604d9474da946403369f652425880babee78",
"paths": ["tests/recon"],
"entries": [{ "path": "tests/recon/Properties.sol", "blob": "...", "size": 13722 }]
}commit and tree are the revision the run actually sees. source_commit and
source_tree are provenance only — those objects are deliberately absent from
the repository, so a reviewer can see which bytes were withheld without the run
being able to read them back.
Verification proves absence rather than diffing, since there is no parent left to diff against: HEAD must be the recorded parentless revision, none of the declared paths may be tracked, none of the recorded blobs may exist in the object store, the declaration must no longer match anything, and the benchmark commit must be unreadable. A launch is refused outright if a held-out path is still present in the checkout.
Ground truth is written against the benchmark commit, so it keeps binding
source_commit rather than the hold-out revision.
Declaring a path that matches nothing fails closed rather than silently running without the hold-out. An agent that writes a similar file inside its own workspace is a legitimate result and stays allowed; only the materialized input is constrained.
Each suite may declare how much model work a recovered row may repeat:
recovery_equivalence:
max_repeated_model_executions: 1
aggregate_non_comparable: separate # include | exclude | separate
publication: comparable # comparable | cleanUltrafuzz derives the row classification from the append-only node-attempt ledger and durable controller submissions. Normal retries within one workflow execution do not count as recovery re-execution. Re-running the same model-backed strategy attempt in a later controller generation does. The persisted row records unique and repeated model executions, infrastructure-only, model-work, and no-progress recovery generations, plus a reason whenever the evidence is non-comparable.
aggregate_non_comparable controls the primary variant aggregates. include
keeps every row, exclude removes non-comparable rows while reporting the
excluded count, and separate additionally emits dedicated non-comparable
variant aggregates. Individual rows and classification totals always remain
visible. publication: comparable rejects non-comparable evidence;
publication: clean also rejects otherwise comparable recovered rows. Public
benchmark history requires clean rows.
The classification is captured once in runs.jsonl and reused by later
scoring, reporting, and publication. Re-scoring therefore cannot reinterpret
the execution exposure after the fact. Missing, malformed, oversized, or
graph-inconsistent execution ledgers fail closed as non-comparable.
The eval driver owns the local loop and writes matrix.json, runs.jsonl,
scores.jsonl, and summary.json; nothing is wired into the workflow runner.
ultrafuzz eval run launches each row as a detached Ultrafuzz run. When it
watches a row, it repeats three steps until the run's durable state.json is
terminal or the watch deadline passes: synchronize the run, read state.json,
and sleep for the poll interval. It reads state.json once more before
recording the row, so a run that another command synchronized to a terminal
state during the last sleep is not recorded as timed out. A failed
synchronization is counted and recorded on the row as EVAL_ROW_SYNC_FAILED.
Only the same failure ten times in a row ends the watch early, recorded as
EVAL_ROW_SYNC_ABANDONED; the run itself is not cancelled. A row whose run
evidence cannot be read is recorded as EVAL_ROW_WATCH_FAILED, and the other
rows and run-summary.json are still written.
There is no reporter or telemetry-export interface.
Every new eval run records a versioned provenance block in eval.json:
- Candidate identity is resolved from the candidate checkout's exact commit, release tag when present, dirty status, and immutable local execution identity when available.
- Benchmark identity includes resolved target commits and clean-checkout state, ground-truth digests, model controls, trial budget, and a normalized execution-policy fingerprint. These controls produce the deterministic cohort fingerprint; tracked target modifications make it incomplete.
- Candidate-owned prompts, topology, strategies, and runtime configuration do
not alter the cohort. Their
graph_fingerprintandconfig_fingerprintare instead recorded on eachruns.jsonlrow so product changes remain visible. summary.jsonrecords a separate scoring identity covering the scorer implementation revision, deterministic or optional-judge mode, judge prompt version, judge models, and ground-truth digests. Historical artifacts without lineage remain readable and are labeled as having unavailable provenance.
Local provenance records preserve the benchmark series, cohort fingerprint, candidate identity, execution policy, row graph/config fingerprints, and scoring identity for comparisons and history publication.
For release-over-release comparisons, pass the candidate run followed by the baseline run:
ultrafuzz eval compare <candidate-eval-run-id> --against <baseline-eval-run-id>The comparison runs only when cohort and scoring identities are complete, match,
and cover the same variant IDs. Use --allow-incompatible as an explicit
waiver; the result remains marked incompatible and lists the compatibility
differences that were waived.
ultrafuzz eval plan # validate config + suite, print the matrix
ultrafuzz eval run # launch rows and poll them to a terminal state
ultrafuzz eval status # observe every row's durable node progress and ETA
ultrafuzz eval score # grade reports against ground truth (optional --llm-judge)
ultrafuzz eval report # show the scored variant ranking
ultrafuzz eval compare # diff variants or release runs with compatible lineage
ultrafuzz eval bundle # export privacy-safe aggregate evidence for offline analysis
ultrafuzz eval history # validate/render or append to public longitudinal historyThe public cohort and lane manifests under benchmarks/ adapt EVMbench detect
and the canonical Ultrafuzz benchmark cohort into the same eval-suite types.
Each cohort family has its own registered whole-document JSON Schema. The lane
policy is current-only ultrafuzz.benchmark.lanes.v2; v1 is not read or
converted. All three documents are accepted by the canonical JSON Schema
validator before their retained, non-transforming Zod parsers run. Cohort
identity joins and the pinned lane policy remain explicit named semantic gates.
Lane trial counts are required authored fields—omitting
trials_per_variant is invalid and never supplies a default.
The expanded target × variant × trial matrix has an absolute 100,000-row
ceiling. Planning rejects a larger product with overflow-safe division before
target resolution, row allocation, iteration, or artifact writes; each
individual dimension is capped at the same value in both JSON Schema and the
runtime schema. Target-level concurrency is capped at 32 and concurrently
launched or scored matrix rows are capped at 80. These ceilings are four times
the built-in full lane (8 targets and 20 runs) while bounding process and
provider fan-out for repository-selected suites.
The bounded smoke lane selects the three Foundry, Hardhat, and Vyper
Ultrafuzz-bench targets and pins GPT-5.6 Luna high for bug-finding. It selects
the CLI-packaged smoke audit profile instead of filtering the production
topology: one context pass feeds four strategies in one parallel wave, followed
by dedupe and report passes. Every node uses the selected runner model and
reasoning level. The four
strategies cover time, external dependencies, externalized accounting, and
lifecycle views. Its
lane definition still records strategy_loops: 1 and the three disabled
strategy families. The full lane selects every checked-in EVMBench target, pins
GPT-5.6 Luna high, Claude Sonnet 5 high, Kimi K3 max, and DeepSeek V4 Pro
max, sets the same
one strategy loop, and explicitly leaves all three disable flags off so the
packaged exhaustive audit profile and its complete specialist topology are included.
Both currently declare one trial per variant and use
GPT-5.6 Sol xhigh as an independent judge. Public Modal pairs contain one
runner variant. Maintainers may explicitly override runner models through the
local public benchmark config generator's BENCHMARK_MODELS_JSON input. Smoke
selects one provider, model, and reasoning level; full retains its four-provider
matrix. These overrides retain the lane's fixed provider count, target
selection, and topology. There are no paid benchmark or history-publication
Actions workflows. Manual publication validates every pair as an exact projection
of the candidate commit's trusted lane policy before merging its observations.
Smoke publication additionally requires at least one normalized finding for
every successful target row; the single report-backed failed target may publish
an empty normalized finding list. The smoke workflow profile and selected strategy IDs are part
of the execution-policy fingerprint, so its charts cannot mix with full or
legacy smoke observations.
The README overview emits one point only when a candidate run contains every
target pinned by its cohort, exactly once, for the same benchmark, lane,
variant, model profile, cohort, execution policy, and scoring lineage.
Incomplete or duplicate target sets are retained in the append-only history but
omitted from the overview. The UltrafuzzBench Score is macro-F1: the
arithmetic mean of the per-target F1 values, so each target and framework has
equal weight. Macro precision and recall use the same equal-target weighting.
Precision and recall remain available in benchmarks/ultrafuzzbench/history.json but are omitted
from the overview charts.
The latest-result summary sums recorded target cost and uses the slowest
recorded target as the parallel run's wall clock. Complete values appear
without a qualifier. If only part of the target evidence is usable, the known
subtotal or maximum remains visible and is labeled partial with its target
coverage (for example, 2/3); it is not presented as a complete-run total. If
no usable value exists, the summary renders n/a with its recorded status.
When the latest run contains multiple model profiles, the summary shows every
profile instead of choosing one by identifier order. The quality overview
shows at most the 12 latest complete candidate runs, gives each model profile a
distinct marker color and shape, and breaks lines when the cohort or execution
policy changes. Solid line segments connect identical scoring identities.
Because an exact scoring identity records the candidate commit itself, dashed
segments provide a visual guide across scoring-identity changes without
claiming strict comparability. Exact scoring identities remain available in
point tooltips and continue to gate strict eval compare compatibility.
The performance × cost overview uses the same complete-run aggregation for the
smoke lane: its vertical value is target-macro F1 and its horizontal value is
the sum of target costs. It groups runs by model, places the dot at the marginal
median of each metric, and draws horizontal cost and vertical F1 bands from
Type-7 first and third quartiles. These bands describe observed run dispersion,
not confidence intervals. To avoid mixing Luna's old and current prices, its
comparison starts at the first run in the non-overlapping current-price regime,
2026-07-31T14:52:13.635Z; other models use all priced smoke runs. A complete
target cohort with a known partial cost remains in the bivariate summary, with
the model point and legend explicitly marked partial and target coverage
shown. Missing targets are never treated as zero. A run with no usable cost is
not plotted; models with no priced run retain their median F1 and sample count
in an explicit unplotted annotation. A one-run model has a dot and a collapsed
IQR.
eval history consumes complete scored generations, stores aggregate metrics
plus immutable candidate, cohort, execution-policy, and scoring lineage in
benchmarks/ultrafuzzbench/history.json, and renders the README SVGs without network or model
calls. A known partial efficiency value remains numeric and renders with a
partial marker and typed reasons. An unavailable value remains null and
renders with an n/a cross. Legacy partial observations whose numeric value was
not recorded remain null, but use a distinct dashed-ring partial n/a marker
so they cannot be mistaken for unavailable data.
The history's top-level supersessions ledger can name an immutable source run
and the source run that replaces it, with a bounded reason and repository issue
URL. A pending entry does not hide data. An identity-compatible subset of the
replacement may arrive without invalidating history; activation waits for exact
per-target observation cardinality parity. At that point renderers exclude the
old source run while retaining its observations in the append-only file. The
replacement must match the old run's benchmark identity and target revisions;
over-counted targets, duplicate ledger entries, self-references, and chains are
rejected. Cohort fingerprints must match by default. An exceptional migration
between provenance schemas can declare a cohort_transition with the exact
superseded and replacement fingerprints; the old pin is checked even while the
entry is pending, the new pin is checked when its source run arrives, and every
non-cohort identity field (including the execution-policy fingerprint) must
still match. This is an auditable migration guardrail, not a general
comparability waiver.
EVMBench and Ultrafuzz-bench reports are non-sensitive public benchmark output.
The Modal publication bundle therefore includes the scored generation and the
allowlisted report and normalized-finding files. It also carries the strict
post-eval diagnostic that certifies each scoreable terminal outcome. A complete
cohort may include one report-backed failed workflow, preserving that failure as
a scored datapoint; two failed rows or missing terminal evidence still block
publication. Bundle path, size, and SHA-256 checks are distinct from
the aggregate-only eval bundle privacy contract used for arbitrary targets.
Deterministic grading runs locally. Optional LLM judging uses the generic
FindingJudge interface and requires an explicitly configured endpoint and
dedicated judge credential. Results remain in local artifacts; no telemetry
publishing service is involved.
eval status <eval-run-id> is the read-only live view across a whole matrix.
It derives node counts, row lifecycle state, checkpoint age, and estimated
remaining time only from recorded eval links, durable run state, and bounded
linked-workflow evidence. It never resumes, retries, synchronizes, collects,
publishes, or otherwise changes a workflow. The compact table shows at most
three active or waiting node IDs per row, truncates each displayed ID after 128
Unicode characters while keeping the JSON value exact, preserves actionable
wait reason → next action pairs, and uses +N for the remainder. It also
distinguishes an admitted active linked workflow from a finished, stopped,
failed, paused, waiting, or cancelled one using the current durable workflow
binding rather than a superseded launch-time ID. Recognized Smithers lifecycle
aliases retain their recorded spelling in JSON for compatibility across runner
versions.
All rows use matrix-order opaque labels. The versioned
ultrafuzz.eval.status.v1 JSON keeps the full active_node_ids and
waiting_nodes lists; each waiting entry includes its node status,
wait_reason, and next_eligible_action. linked_workflow_status carries the
reconciled lifecycle value. It is null when no workflow is linked and
"unknown" when linked evidence is missing, unsupported, or ambiguous. Wait
reason and next action are likewise null when absent or newer than the known
telemetry vocabulary. The existing typed counts, percentages, timestamps,
checkpoint freshness, and ETA availability remain unchanged. Private target
metadata, repository locations, findings, and diagnostics are not
representable.
The compact table renders an unlinked workflow as none and ambiguous or
unavailable linked evidence as unknown.
Configure an independent judge panel at the root of the eval suite YAML selected
by --suite or [eval].eval_config:
judge_panel:
total: 4
quorum: 3The default is three independent members with quorum two. Explicit strict-majority
overrides remain supported, including one member with quorum one and four members
with quorum three. total and quorum must be positive integers, quorum cannot
exceed total, and the quorum must be a strict majority. Each member gets a fresh
model context containing the same
versioned adjudicator prompt; panel requests use bounded concurrency while
retaining each request's retry, timeout, privacy, and credential boundaries.
The evaluator applies the normal classification policy to every response and
requires quorum on an identical decision. A true-positive vote's identity
includes its canonical matched bug ID. Findings without quorum enter the human
review queue with the panel-disagreement reason.
Each judged finding in scores.jsonl records the panel total and quorum,
model, reasoning effort when configured, prompt version, deterministic vote
split, individual member decisions and rationales, and aggregate decision.
Scoring provenance includes the effective panel configuration and versioned
prompt/scorer revisions, so panel scores are not interchangeable with older
single-judge artifacts. Credentials are never part of score artifacts.
eval bundle <eval-run-id> --output <directory> is the explicit offline
analysis export. It derives fixed-schema aggregate files rather than copying
the eval or run directories. analysis-bundle.json records bundle-relative
paths, sizes, and SHA-256 checksums; omissions.json records typed reasons for
expected evidence that was unavailable. Collection validates every payload,
reference, checksum, and the privacy allowlist before replacing the output
directory. The resulting bundle contains no raw agent output, findings,
configuration, absolute execution paths, or deployment identifiers.
When Modal recovery lifecycle evidence is supplied, the bundle includes
data/recovery-summary.json with exactly reconciled generation, progress,
model-work, resumption, rotation, and genuine-failure counts. Historical
exports without that evidence record a typed source omission rather than
guessing from launcher attempts.
Scored row summaries take lifecycle timestamps and terminal status from the
durable run state rather than the detached launcher process. Their typed
efficiency block reports wall/active/wait time, total tokens, and cost together
with explicit completeness states, and summary.md renders those same
structured fields.
The optional LLM judge requires both ULTRAFUZZ_EVAL_JUDGE_URL and its own
ULTRAFUZZ_EVAL_JUDGE_API_KEY. There is no implicit gateway; reporting or
general model-provider credentials are never reused. The endpoint must be
HTTPS without embedded credentials, and redirects are rejected.
For a target marked sensitivity: private, the judge is disabled unless the
operator explicitly sets ULTRAFUZZ_EVAL_JUDGE_ALLOW_PRIVATE_DATA=true.
Ground-truth IDs are replaced with candidate aliases in the request, and an
LLM result cannot downgrade a deterministic true positive.
Full per-command flags are in the CLI reference. Local eval
artifacts (eval.json, matrix.json, runs.jsonl, scores.jsonl,
summary.json, summary.md) are documented in
Run Artifacts and Reports. For a
task-oriented walkthrough, see
Run Eval Suites.