Skip to content

Repository files navigation

Jev Computer Use Bench

View the live report on GitHub Pages

Pass rate versus cost per success chart for the Jev Computer Use Bench

Light mode screenshot · Dark mode screenshot

Test prompts and verification: see TESTS.md.

This repository is a reproducibility snapshot for a benchmark comparing four approaches to browser and desktop computer use:

  • Jev Browser - fast structured browser decisions.
  • Jev Ultrafast - a faster browser variant with a separate text helper.
  • Hybrid - Jev first, a postcondition check, then Luna-style visual recovery when Jev misses.
  • Luna - visual computer use through the visible browser or desktop surface.

Headline results

The ranking rule is success rate first, then estimated cost per successful task. Time is shown separately.

Approach Overall Simple browser Real-world browser Computer use
Hybrid 55/57 - 96.5% 42/42 - 100% 8/10 - 80% 5/5 - 100%
Luna 55/57 - 96.5% 42/42 - 100% 8/10 - 80% 5/5 - 100%
Jev Ultrafast 19/51 - 37.3% 18/42 - 42.9% 1/9 - 11.1% not measured
Jev Browser 18/52 - 34.6% 18/42 - 42.9% 0/10 - 0% not measured

Hybrid's ten-task browser result is a staged recorded fallback comparison, not one uninterrupted live interleaved turn. Jev Ultrafast's real-world number is a separate public nine-flow discovery lane.

The complete interactive report is report.html. It has light/dark mode, a pass-rate-versus-cost chart, lane filters, approach filters for recorded runs, and links to the locally retained clips when those clips are present. data/benchmark-final-manifest.json is the generated machine-readable report snapshot.

When GitHub Pages is enabled for this repository, the same report and bundled videos are available at https://builderio.github.io/jev-computer-use-tests/.

Cost and speed

Approach Overall cost / success Overall average time Native computer-use cost / success Native average time
Hybrid $0.0090 29.68 s/task ~$0.0126 proxy ~36.1 s/task
Luna $0.0128 44.57 s/task ~$0.0091 proxy ~37.8 s/task
Jev Ultrafast $0.0012 1.07 s/task not measured not measured
Jev Browser $0.000625 3.62 s/task not measured not measured

The native Hybrid cost is higher than Luna because every native task pays for the Jev-first attempt plus the verification/fallback path. Jev directly passed only the clean TextEdit task; Keynote and Numbers still required Luna recovery. The Hybrid path was slightly faster in this small lane, but it did not reduce the expensive work enough to beat pure Luna on cost.

Costs are comparison estimates, not invoices. Jev uses the recorded API input-token estimate. Luna uses an action/token proxy because the computer-use surface did not expose a product invoice. Local browser and desktop tool fees are treated as $0.

What was actually tested

Simple browser

42 controlled browser tasks covering navigation, clicking, selecting, scrolling, forms, small edits, and state checks. This lane measures the mechanical ceiling of each browser-action approach.

Real-world browser

Ten longer workflows included flight search and fare selection, a Thai restaurant reservation up to a safe confirmation boundary, TodoMVC project tracking, Drive, Notion, Figma, mail, Calendar, Spotify, and SauceDemo. No purchase, reservation, payment, message send, or publish action was submitted.

The authenticated Figma rerun succeeded for both Luna and Hybrid. Its time and cost were not separately metered, so the published time and cost columns remain the known-run estimates.

Jev Browser reached partial pages and controls but passed 0/10 strict postconditions. A separate Jev Ultrafast public lane covered nine read-only discovery flows and is not the same denominator.

Native computer use

Five desktop tasks used Spotify, Pages, TextEdit, Keynote, and Numbers. TextEdit was the one direct Jev pass. Luna completed all five. Hybrid used Jev when the result was verifiable and Luna when it was not.

We also ran a separate mac-cua Accessibility-driver control on the four native failures. It passed Pages and failed strict Keynote, Numbers, and Spotify end states. Because it has no planning model, those results are capability evidence, not a fourth leaderboard system.

Browser harness and file upload

Luna drove the normal Chrome browser through the ChatGPT desktop bridge; an earlier in-app-browser pass scored 39/42 because that browser has no file-picker and is not the intended setup. No page JavaScript execution was used, matching normal ChatGPT Chrome use.

That is why the three repeated 42-task Luna misses were file-upload fixtures. It is a limitation of that harness configuration, not evidence that Luna or Codex can never upload files. A different browser adapter can expose a native chooser or bind directly to a file input.

Three model-controlled Luna turns in the normal Chrome bridge ran the upload fixture. All three opened the native macOS chooser, selected benchmark.txt, clicked Upload fixture, and reached the visible success state: 3/3 targeted uploads. Combined with the original 39/39 non-upload passes, this supports a connected-Chrome result of 42/42. Traces are data/luna-upload-trace.json, data/luna-upload-attempt-2.json, and data/luna-upload-attempt-3.json.

Hybrid was also rerun through the model-controlled normal-Chrome upload path: three Jev-first attempts stopped at the unsupported file input, and connected-Chrome native-chooser recovery reached the visible success state on all three (3/3). Combined with the original 39/39 non-upload passes, this supports a connected-Chrome result of 42/42. Trace is data/hybrid-upload-attempts.json.

The native Jev lane used the local arc-cua runner with a macOS accessibility/OCR backend and a TypeSafe Jev policy. The Hybrid result is a policy plus verification and fallback, not a separate model. The local report includes short window-scoped Jev browser/app stop clips, a window-scoped TextEdit success, and a Pages false-completion recording. Long frozen tails were excluded after frame review; the long Luna walkthrough and staged Hybrid app walkthrough remain supplemental evidence rather than new scores.

Repository contents

Reproduce the safe snapshot

python3 -m json.tool data/results.json
python3 -m json.tool data/evidence.json
python3 scripts/verify_results.py
python3 -m http.server 8000
# then open http://127.0.0.1:8000/report.html

The full authenticated run requires private test accounts, local browser profiles, a Jev API key supplied through an environment variable, and native macOS applications. Those credentials and profiles are intentionally not part of this repository. The long Luna five-app clip is a successful supplemental walkthrough. The bundled Hybrid companion is an honest four-app fallback marked Fail: it does not reach Slides and is not a Hybrid score. See REPRODUCE.md for the boundary between the safe snapshot and the private local rerun.

Recording policy

Only frame-reviewed recordings belong in the local report. Earlier captures that showed the wrong blank Chrome window or the wrong native app window were excluded. The local raw video directory is about 2.6 GB and includes authenticated material, so it is not copied into this Git repository. One five-app Chrome bridge walkthrough is bundled as a supplemental, unscored example.

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages