Skip to content

Latest commit

 

History

History
505 lines (437 loc) · 26.7 KB

File metadata and controls

505 lines (437 loc) · 26.7 KB

Eval Suites

Ultrafuzz eval suites benchmark the fuzzing pipeline against targets with known ground-truth bugs. Run journals retain node lifecycle, heartbeat, and artifact evidence locally. Scoring, comparison, reports, and history work without an external reporting account.

Configuration split

The eval YAML defines the experiment and grading policy. Project TOML supplies the default suite location and machine-specific ground-truth directory.

ultrafuzz.toml — the [eval] section

[eval]
eval_config = ".ultrafuzz/evals/bug-finding.yml"
ground_truth_root = "/secure/eval-ground-truth"
provider = "none"
  • Only provider = "none" is supported. No external reporter is installed.
  • Precedence: CLI flag (--provider, --suite) > environment (ULTRAFUZZ_EVAL_PROVIDER, ULTRAFUZZ_EVAL_CONFIG) > project TOML. An unsupported final provider selection fails before credential access or workflow launch; an explicit --provider none can override old settings.
  • Ground truth must live outside the repository. Suite targets reference files relative to ground_truth_root; absolute entries, traversal, symlinks, non-regular files, and files larger than 1 MiB are rejected.
  • Generic connection metadata in saved configuration remains readable but cannot activate an external reporter. Fresh defaults contain no connection profiles. See the migration guidance.

Eval YAML — the experiment definition

See .ultrafuzz/evals/bug-finding.yml for the default suite. It defines model profiles, targets (repo/ref/ground truth/sensitivity), variants, trial counts, grading metrics, and the reporting: block. Nothing in the YAML names a provider, an endpoint, or an env var.

reporting.node_telemetry decides whether ultrafuzz eval run watches launched rows by default; --no-watch overrides it. It and reporting.heartbeat_interval_seconds remain inputs to the execution-policy fingerprint. reporting.experiment_prefix and reporting.artifacts are still validated, but no eval behaviour depends on them.

Suite contract and workflow input

The current suite version is exactly ultrafuzz.eval.v2, validated against urn:ultrafuzz:schema:evals:suite:2. JSON Schema is canonical. The retained Zod parser is strict and non-transforming, and a shared acceptance corpus keeps it aligned with the JSON Schema. Defaults are applied only after both shape validators accept the operator-authored document.

All eval-owned objects are closed. In particular, model profiles contain only agent, optional model, optional reasoning, and optional timeout_seconds; variants contain only id, optional topology, optional workflow_input, and optional runner/judge profile overrides; metrics contains only recall_threshold. Historical no-op fields such as model config, variant prompts or model_profiles, root prompt_overlays, and metrics.primary/metrics.secondary are rejected. ground_truth_root is machine-specific TOML/CLI configuration and is not a suite-YAML field.

Ordinary operator-defined workflow_input remains an intentional extension seam. It must be a JSON object with nonempty, non-whitespace keys; its values may be nested JSON objects, arrays, strings, numbers, booleans, or null. It may not use the eval-owned reserved keys benchmark_execution, benchmark_lane, excluded_strategy_families, target_frameworks, or ultrafuzz_eval. Scalars and arrays at the workflow_input root are rejected instead of being silently discarded.

Benchmark variants use one of three exact shapes. Private benchmark controls contain only the execution budget:

workflow_input:
  benchmark_execution:
    strategy_loops: 2
    excluded_node_ids: [optional-analysis]

The public full lane contains all fields below and permits neither excluded families nor excluded nodes:

workflow_input:
  benchmark_lane: full
  target_frameworks: { target-a: foundry }
  excluded_strategy_families: []
  benchmark_execution:
    strategy_loops: 1
    excluded_node_ids: []

The public smoke lane fixes the workflow profile, all three excluded strategy families, and all four selected strategies:

workflow_input:
  benchmark_lane: smoke
  target_frameworks: { target-a: foundry }
  excluded_strategy_families: [stateful-invariant, differential, dynamic-strategy]
  benchmark_execution:
    workflow_profile: smoke-benchmark-v1
    audit_profile: smoke
    audit_profile_catalog_digest: <64-lowercase-hex catalog digest>
    topology_digest: <64-lowercase-hex packaged smoke topology digest>
    selected_strategy_ids:
      - time-warp-sequences
      - external-dependency-boundaries
      - externalized-state-accounting
      - lifecycle-view-boundaries
    strategy_loops: 1
    excluded_node_ids: []

Missing fields, extra fields, duplicate array entries, wrong lane constants, reserved operator keys, and historical version literals fail validation. There is no alias conversion, compatibility fallback, or repair pass. The checked-in benchmark adapter preserves that released full-lane shape. The launcher derives the current packaged exhaustive policy from benchmark_lane: full; it refuses to launch public variants that declare a topology override. Before startRun, it resolves the target's effective policy, requires the topology to originate from that audit profile, and refuses any audit-profile, catalog, or topology digest mismatch. The run journal captures that actual origin and policy, and public-history publication checks every launched record against the candidate checkout's packaged policy, so a stale or target-local override cannot silently replace the packaged graph.

Held-out benchmark paths

A benchmark that ships a reference solution beside the code under test would otherwise hand the run its own answer key. A target may withhold those paths:

targets:
  - id: target
    repo: "https://github.com/scfuzzbench/aave-v4-scfuzzbench"
    ref: "edd6c82721512540c8c90e7a36a4a8e19fd7bdf3"
    ground_truth: aave-v4/findings.yml
    held_out_paths: ["tests/recon"]

Materializing a pinned target rewrites the checkout to a parentless revision that never contained the declared paths, then destroys the benchmark commit and the withheld blobs.

A commit is required rather than a dirty worktree, because every node workspace is a Git worktree of the pinned branch and would otherwise restore the files. Keeping the benchmark commit as a parent is equally unsafe — git show HEAD^:tests/recon/Properties.sol would hand the answer key straight back — so the rewritten revision has no parents and the original objects are pruned along with the reflog. The single-revision isolation invariants are unchanged: revision_count and commit_object_count stay 1.

The source proof records what was withheld:

"held_out": {
  "source_commit": "edd6c82721512540c8c90e7a36a4a8e19fd7bdf3",
  "source_tree": "3b3910d0657e4a639888ccae45822af680de6ef6",
  "commit": "ac548733fd9c36cf9a32284a653a6429c089abfa",
  "tree": "5e9c604d9474da946403369f652425880babee78",
  "paths": ["tests/recon"],
  "entries": [{ "path": "tests/recon/Properties.sol", "blob": "...", "size": 13722 }]
}

commit and tree are the revision the run actually sees. source_commit and source_tree are provenance only — those objects are deliberately absent from the repository, so a reviewer can see which bytes were withheld without the run being able to read them back.

Verification proves absence rather than diffing, since there is no parent left to diff against: HEAD must be the recorded parentless revision, none of the declared paths may be tracked, none of the recorded blobs may exist in the object store, the declaration must no longer match anything, and the benchmark commit must be unreadable. A launch is refused outright if a held-out path is still present in the checkout.

Ground truth is written against the benchmark commit, so it keeps binding source_commit rather than the hold-out revision.

Declaring a path that matches nothing fails closed rather than silently running without the hold-out. An agent that writes a similar file inside its own workspace is a legitimate result and stays allowed; only the materialized input is constrained.

Recovery equivalence

Each suite may declare how much model work a recovered row may repeat:

recovery_equivalence:
  max_repeated_model_executions: 1
  aggregate_non_comparable: separate # include | exclude | separate
  publication: comparable # comparable | clean

Ultrafuzz derives the row classification from the append-only node-attempt ledger and durable controller submissions. Normal retries within one workflow execution do not count as recovery re-execution. Re-running the same model-backed strategy attempt in a later controller generation does. The persisted row records unique and repeated model executions, infrastructure-only, model-work, and no-progress recovery generations, plus a reason whenever the evidence is non-comparable.

aggregate_non_comparable controls the primary variant aggregates. include keeps every row, exclude removes non-comparable rows while reporting the excluded count, and separate additionally emits dedicated non-comparable variant aggregates. Individual rows and classification totals always remain visible. publication: comparable rejects non-comparable evidence; publication: clean also rejects otherwise comparable recovered rows. Public benchmark history requires clean rows.

The classification is captured once in runs.jsonl and reused by later scoring, reporting, and publication. Re-scoring therefore cannot reinterpret the execution exposure after the fact. Missing, malformed, oversized, or graph-inconsistent execution ledgers fail closed as non-comparable.

Architecture

The eval driver owns the local loop and writes matrix.json, runs.jsonl, scores.jsonl, and summary.json; nothing is wired into the workflow runner. ultrafuzz eval run launches each row as a detached Ultrafuzz run. When it watches a row, it repeats three steps until the run's durable state.json is terminal or the watch deadline passes: synchronize the run, read state.json, and sleep for the poll interval. It reads state.json once more before recording the row, so a run that another command synchronized to a terminal state during the last sleep is not recorded as timed out. A failed synchronization is counted and recorded on the row as EVAL_ROW_SYNC_FAILED. Only the same failure ten times in a row ends the watch early, recorded as EVAL_ROW_SYNC_ABANDONED; the run itself is not cancelled. A row whose run evidence cannot be read is recorded as EVAL_ROW_WATCH_FAILED, and the other rows and run-summary.json are still written. There is no reporter or telemetry-export interface.

Versioned lineage

Every new eval run records a versioned provenance block in eval.json:

  • Candidate identity is resolved from the candidate checkout's exact commit, release tag when present, dirty status, and immutable local execution identity when available.
  • Benchmark identity includes resolved target commits and clean-checkout state, ground-truth digests, model controls, trial budget, and a normalized execution-policy fingerprint. These controls produce the deterministic cohort fingerprint; tracked target modifications make it incomplete.
  • Candidate-owned prompts, topology, strategies, and runtime configuration do not alter the cohort. Their graph_fingerprint and config_fingerprint are instead recorded on each runs.jsonl row so product changes remain visible.
  • summary.json records a separate scoring identity covering the scorer implementation revision, deterministic or optional-judge mode, judge prompt version, judge models, and ground-truth digests. Historical artifacts without lineage remain readable and are labeled as having unavailable provenance.

Local provenance records preserve the benchmark series, cohort fingerprint, candidate identity, execution policy, row graph/config fingerprints, and scoring identity for comparisons and history publication.

For release-over-release comparisons, pass the candidate run followed by the baseline run:

ultrafuzz eval compare <candidate-eval-run-id> --against <baseline-eval-run-id>

The comparison runs only when cohort and scoring identities are complete, match, and cover the same variant IDs. Use --allow-incompatible as an explicit waiver; the result remains marked incompatible and lists the compatibility differences that were waived.

CLI surface

ultrafuzz eval plan      # validate config + suite, print the matrix
ultrafuzz eval run       # launch rows and poll them to a terminal state
ultrafuzz eval status    # observe every row's durable node progress and ETA
ultrafuzz eval score     # grade reports against ground truth (optional --llm-judge)
ultrafuzz eval report    # show the scored variant ranking
ultrafuzz eval compare   # diff variants or release runs with compatible lineage
ultrafuzz eval bundle    # export privacy-safe aggregate evidence for offline analysis
ultrafuzz eval history   # validate/render or append to public longitudinal history

The public cohort and lane manifests under benchmarks/ adapt EVMbench detect and the canonical Ultrafuzz benchmark cohort into the same eval-suite types. Each cohort family has its own registered whole-document JSON Schema. The lane policy is current-only ultrafuzz.benchmark.lanes.v2; v1 is not read or converted. All three documents are accepted by the canonical JSON Schema validator before their retained, non-transforming Zod parsers run. Cohort identity joins and the pinned lane policy remain explicit named semantic gates. Lane trial counts are required authored fields—omitting trials_per_variant is invalid and never supplies a default. The expanded target × variant × trial matrix has an absolute 100,000-row ceiling. Planning rejects a larger product with overflow-safe division before target resolution, row allocation, iteration, or artifact writes; each individual dimension is capped at the same value in both JSON Schema and the runtime schema. Target-level concurrency is capped at 32 and concurrently launched or scored matrix rows are capped at 80. These ceilings are four times the built-in full lane (8 targets and 20 runs) while bounding process and provider fan-out for repository-selected suites. The bounded smoke lane selects the three Foundry, Hardhat, and Vyper Ultrafuzz-bench targets and pins GPT-5.6 Luna high for bug-finding. It selects the CLI-packaged smoke audit profile instead of filtering the production topology: one context pass feeds four strategies in one parallel wave, followed by dedupe and report passes. Every node uses the selected runner model and reasoning level. The four strategies cover time, external dependencies, externalized accounting, and lifecycle views. Its lane definition still records strategy_loops: 1 and the three disabled strategy families. The full lane selects every checked-in EVMBench target, pins GPT-5.6 Luna high, Claude Sonnet 5 high, Kimi K3 max, and DeepSeek V4 Pro max, sets the same one strategy loop, and explicitly leaves all three disable flags off so the packaged exhaustive audit profile and its complete specialist topology are included. Both currently declare one trial per variant and use GPT-5.6 Sol xhigh as an independent judge. Public Modal pairs contain one runner variant. Maintainers may explicitly override runner models through the local public benchmark config generator's BENCHMARK_MODELS_JSON input. Smoke selects one provider, model, and reasoning level; full retains its four-provider matrix. These overrides retain the lane's fixed provider count, target selection, and topology. There are no paid benchmark or history-publication Actions workflows. Manual publication validates every pair as an exact projection of the candidate commit's trusted lane policy before merging its observations. Smoke publication additionally requires at least one normalized finding for every successful target row; the single report-backed failed target may publish an empty normalized finding list. The smoke workflow profile and selected strategy IDs are part of the execution-policy fingerprint, so its charts cannot mix with full or legacy smoke observations.

Published history

The README overview emits one point only when a candidate run contains every target pinned by its cohort, exactly once, for the same benchmark, lane, variant, model profile, cohort, execution policy, and scoring lineage. Incomplete or duplicate target sets are retained in the append-only history but omitted from the overview. The UltrafuzzBench Score is macro-F1: the arithmetic mean of the per-target F1 values, so each target and framework has equal weight. Macro precision and recall use the same equal-target weighting. Precision and recall remain available in benchmarks/ultrafuzzbench/history.json but are omitted from the overview charts.

The latest-result summary sums recorded target cost and uses the slowest recorded target as the parallel run's wall clock. Complete values appear without a qualifier. If only part of the target evidence is usable, the known subtotal or maximum remains visible and is labeled partial with its target coverage (for example, 2/3); it is not presented as a complete-run total. If no usable value exists, the summary renders n/a with its recorded status. When the latest run contains multiple model profiles, the summary shows every profile instead of choosing one by identifier order. The quality overview shows at most the 12 latest complete candidate runs, gives each model profile a distinct marker color and shape, and breaks lines when the cohort or execution policy changes. Solid line segments connect identical scoring identities. Because an exact scoring identity records the candidate commit itself, dashed segments provide a visual guide across scoring-identity changes without claiming strict comparability. Exact scoring identities remain available in point tooltips and continue to gate strict eval compare compatibility.

The performance × cost overview uses the same complete-run aggregation for the smoke lane: its vertical value is target-macro F1 and its horizontal value is the sum of target costs. It groups runs by model, places the dot at the marginal median of each metric, and draws horizontal cost and vertical F1 bands from Type-7 first and third quartiles. These bands describe observed run dispersion, not confidence intervals. To avoid mixing Luna's old and current prices, its comparison starts at the first run in the non-overlapping current-price regime, 2026-07-31T14:52:13.635Z; other models use all priced smoke runs. A complete target cohort with a known partial cost remains in the bivariate summary, with the model point and legend explicitly marked partial and target coverage shown. Missing targets are never treated as zero. A run with no usable cost is not plotted; models with no priced run retain their median F1 and sample count in an explicit unplotted annotation. A one-run model has a dot and a collapsed IQR.

eval history consumes complete scored generations, stores aggregate metrics plus immutable candidate, cohort, execution-policy, and scoring lineage in benchmarks/ultrafuzzbench/history.json, and renders the README SVGs without network or model calls. A known partial efficiency value remains numeric and renders with a partial marker and typed reasons. An unavailable value remains null and renders with an n/a cross. Legacy partial observations whose numeric value was not recorded remain null, but use a distinct dashed-ring partial n/a marker so they cannot be mistaken for unavailable data.

The history's top-level supersessions ledger can name an immutable source run and the source run that replaces it, with a bounded reason and repository issue URL. A pending entry does not hide data. An identity-compatible subset of the replacement may arrive without invalidating history; activation waits for exact per-target observation cardinality parity. At that point renderers exclude the old source run while retaining its observations in the append-only file. The replacement must match the old run's benchmark identity and target revisions; over-counted targets, duplicate ledger entries, self-references, and chains are rejected. Cohort fingerprints must match by default. An exceptional migration between provenance schemas can declare a cohort_transition with the exact superseded and replacement fingerprints; the old pin is checked even while the entry is pending, the new pin is checked when its source run arrives, and every non-cohort identity field (including the execution-policy fingerprint) must still match. This is an auditable migration guardrail, not a general comparability waiver.

EVMBench and Ultrafuzz-bench reports are non-sensitive public benchmark output. The Modal publication bundle therefore includes the scored generation and the allowlisted report and normalized-finding files. It also carries the strict post-eval diagnostic that certifies each scoreable terminal outcome. A complete cohort may include one report-backed failed workflow, preserving that failure as a scored datapoint; two failed rows or missing terminal evidence still block publication. Bundle path, size, and SHA-256 checks are distinct from the aggregate-only eval bundle privacy contract used for arbitrary targets.

Deterministic grading runs locally. Optional LLM judging uses the generic FindingJudge interface and requires an explicitly configured endpoint and dedicated judge credential. Results remain in local artifacts; no telemetry publishing service is involved.

eval status <eval-run-id> is the read-only live view across a whole matrix. It derives node counts, row lifecycle state, checkpoint age, and estimated remaining time only from recorded eval links, durable run state, and bounded linked-workflow evidence. It never resumes, retries, synchronizes, collects, publishes, or otherwise changes a workflow. The compact table shows at most three active or waiting node IDs per row, truncates each displayed ID after 128 Unicode characters while keeping the JSON value exact, preserves actionable wait reason → next action pairs, and uses +N for the remainder. It also distinguishes an admitted active linked workflow from a finished, stopped, failed, paused, waiting, or cancelled one using the current durable workflow binding rather than a superseded launch-time ID. Recognized Smithers lifecycle aliases retain their recorded spelling in JSON for compatibility across runner versions.

All rows use matrix-order opaque labels. The versioned ultrafuzz.eval.status.v1 JSON keeps the full active_node_ids and waiting_nodes lists; each waiting entry includes its node status, wait_reason, and next_eligible_action. linked_workflow_status carries the reconciled lifecycle value. It is null when no workflow is linked and "unknown" when linked evidence is missing, unsupported, or ambiguous. Wait reason and next action are likewise null when absent or newer than the known telemetry vocabulary. The existing typed counts, percentages, timestamps, checkpoint freshness, and ETA availability remain unchanged. Private target metadata, repository locations, findings, and diagnostics are not representable.

The compact table renders an unlinked workflow as none and ambiguous or unavailable linked evidence as unknown.

Configure an independent judge panel at the root of the eval suite YAML selected by --suite or [eval].eval_config:

judge_panel:
  total: 4
  quorum: 3

The default is three independent members with quorum two. Explicit strict-majority overrides remain supported, including one member with quorum one and four members with quorum three. total and quorum must be positive integers, quorum cannot exceed total, and the quorum must be a strict majority. Each member gets a fresh model context containing the same versioned adjudicator prompt; panel requests use bounded concurrency while retaining each request's retry, timeout, privacy, and credential boundaries. The evaluator applies the normal classification policy to every response and requires quorum on an identical decision. A true-positive vote's identity includes its canonical matched bug ID. Findings without quorum enter the human review queue with the panel-disagreement reason.

Each judged finding in scores.jsonl records the panel total and quorum, model, reasoning effort when configured, prompt version, deterministic vote split, individual member decisions and rationales, and aggregate decision. Scoring provenance includes the effective panel configuration and versioned prompt/scorer revisions, so panel scores are not interchangeable with older single-judge artifacts. Credentials are never part of score artifacts.

eval bundle <eval-run-id> --output <directory> is the explicit offline analysis export. It derives fixed-schema aggregate files rather than copying the eval or run directories. analysis-bundle.json records bundle-relative paths, sizes, and SHA-256 checksums; omissions.json records typed reasons for expected evidence that was unavailable. Collection validates every payload, reference, checksum, and the privacy allowlist before replacing the output directory. The resulting bundle contains no raw agent output, findings, configuration, absolute execution paths, or deployment identifiers. When Modal recovery lifecycle evidence is supplied, the bundle includes data/recovery-summary.json with exactly reconciled generation, progress, model-work, resumption, rotation, and genuine-failure counts. Historical exports without that evidence record a typed source omission rather than guessing from launcher attempts.

Scored row summaries take lifecycle timestamps and terminal status from the durable run state rather than the detached launcher process. Their typed efficiency block reports wall/active/wait time, total tokens, and cost together with explicit completeness states, and summary.md renders those same structured fields.

The optional LLM judge requires both ULTRAFUZZ_EVAL_JUDGE_URL and its own ULTRAFUZZ_EVAL_JUDGE_API_KEY. There is no implicit gateway; reporting or general model-provider credentials are never reused. The endpoint must be HTTPS without embedded credentials, and redirects are rejected. For a target marked sensitivity: private, the judge is disabled unless the operator explicitly sets ULTRAFUZZ_EVAL_JUDGE_ALLOW_PRIVATE_DATA=true. Ground-truth IDs are replaced with candidate aliases in the request, and an LLM result cannot downgrade a deterministic true positive.

Full per-command flags are in the CLI reference. Local eval artifacts (eval.json, matrix.json, runs.jsonl, scores.jsonl, summary.json, summary.md) are documented in Run Artifacts and Reports. For a task-oriented walkthrough, see Run Eval Suites.