Skip to content

Support the GPT-6 model catalog and reasoning effort - #177

Merged
cpojer merged 2 commits into
nkzw-tech:mainfrom
pascalandy:feat/codex-model-catalog
Oct 6, 2026
Merged

cpojer merged 2 commits into
nkzw-tech:mainfrom
pascalandy:feat/codex-model-catalog

Conversation

@pascalandy

@pascalandy pascalandy commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Why

Codiff's Codex menu is limited to a predefined model list, and unrecognized model IDs are replaced with the default. This prevents selecting newer installed models such as GPT-6.1 Sol with high reasoning.

This change reads the installed Codex CLI's visible, paginated model catalog in the background and adds a model-specific Reasoning Effort menu. Custom model IDs remain usable when discovery is unavailable. The new settings.openAIReasoningEffort override reaches both execution paths and separates walkthrough cache entries by effort. Existing automatic defaults remain compatible.

Scope

  • Load the installed CLI's visible models and supported reasoning efforts without an inference turn
  • Accept custom model IDs and persist settings.openAIReasoningEffort through both execution paths
  • Keep the selected model for large walkthroughs when an effort is explicit
  • Include explicit effort in walkthrough cache keys and document compatibility behavior

Example configuration:

{
  "settings": {
    "openAIModel": "gpt-6.1-sol",
    "openAIReasoningEffort": "high"
  }
}

Verification

  • pnpm exec vp check --fix passed
  • pnpm exec vp test: 1,143 passed, 7 skipped across 110 files
  • pnpm exec vp run build passed
  • Config defaults validated against the JSON Schema
  • Claude Opus 5.5 verified the fixes and found no remaining source defects
  • Real Codex 0.160.1 discovery returned current models, and an app-server turn acknowledged gpt-6.1-sol with high
  • GitHub CI needs maintainer action before jobs run (action_required)
  • Native menu interaction has not been visually tested; this PR remains a draft

AI assistance: GPT-6.1 Sol through Codex performed research, design, code, tests, documentation, and review responses. Codex agents reviewed design and source. Claude Opus 5.5 independently reviewed the change and verified its routing, error-reporting, and discovery-retry fixes.

banana banana banana

by GPT-6.1-Sol via Codex

- Purpose: let users select current Codex models and supported reasoning efforts
- Impact: preserve custom IDs, forward explicit effort, and separate effort-specific cache entries

by GPT-6.1-Sol via Codex
- Purpose: prevent automatic large-diff routing from sending explicit effort to an incompatible model
- Impact: retain selected models, report fallback causes, and allow later discovery retries

by GPT-6.1-Sol via Codex
@pascalandy pascalandy changed the title Support the Codex model catalog and reasoning effort Support the GPT-6 model catalog and reasoning effort Oct 6, 2026
@pascalandy
pascalandy marked this pull request as ready for review October 6, 2026 18:41
@pascalandy

Copy link
Copy Markdown
Contributor Author

Plan

Codiff should let users select the models and reasoning efforts offered by their installed Codex CLI, without requiring a Codiff release whenever OpenAI updates its catalog.

This includes GPT‑6 Astra, GPT‑6 Sol, GPT‑6 Luna, and GPT‑6.1 Sol when available to the user. My specific use case is GPT‑6.1 Sol with high reasoning effort. OpenAI model documentation

The scope includes model selection, reasoning effort, saved preferences, and compatibility with older Codex clients. Other agent backends and account-access management are out of scope.

CMO (current Mode of operation)

Codiff maintains its own model list and reasoning defaults.

  • The Model menu offers GPT‑5.6 Terra, Sol, Luna, and GPT‑5.5
  • GPT‑6 Astra, Sol, and Luna are accepted through configuration but absent from the predefined menu
  • Model IDs outside the allowlist are silently replaced with GPT‑5.6 Terra, including gpt-6.1-sol
  • Reasoning effort cannot be configured by the user

See the predefined catalog and model normalization.

This creates a gap between what Codex supports and what Codiff exposes.

FMO (future Mode of operation)

Updating the static list would address today's models but require repeated maintenance. I recommend using Codex's own catalog.

Codex app-server provides model/list, including model names, supported reasoning efforts, and default effort.

  • Populate the model selector from model/list, including all pages of results
  • Offer the reasoning efforts advertised for the selected model
  • Persist explicit choices and pass them unchanged through both app-server and exec
  • Accept custom model IDs in configuration without a separate Codiff allowlist
  • Preserve saved choices and manual configuration when catalog discovery fails or an older CLI lacks support
  • Show clear errors for unavailable models or unsupported effort, and identify any fallback used

The resulting selector should expose new models as Codex adds them, without another Codiff catalog update.

How we'll know it works

  • Verify every returned model appears in the selector, including models beyond the first catalog page
  • Verify GPT‑6.1 Sol high reaches both execution transports unchanged
  • Verify saved choices survive startup, saving, and live reload
  • Verify catalog failures and older clients still permit manual configuration
  • Extend the existing Codex, agent-menu, and configuration tests
  • Run vp test, vp check --fix, and vpr build, then confirm the effective model and effort in a real session

Premortem

  1. A second allowlist still rejects discovered models. New models appear but execute as Terra. Remove restrictive normalization throughout the configuration and execution paths
  2. Catalog discovery fails on an older CLI. The selector becomes unusable. Keep manual configuration available and preserve saved choices
  3. Effort handling differs between transports. The same selection produces different settings. Verify both execution paths and app-server thread and turn requests

by GPT-6.1-Sol via Codex

@cpojer cpojer left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Awesome, this totally makes sense. I do think the defaults still require frequent eval runs before changing them though. If you look into the evals folder, for example GPT-6 was only better in one situation.

@cpojer
cpojer merged commit c2f270a into nkzw-tech:main Oct 6, 2026
@pascalandy

Copy link
Copy Markdown
Contributor Author

We have a little win here :


GPT-6.1 Sol walkthrough evaluation — 2026-10-06

GPT-6.1 Sol writes better walkthroughs than GPT-6 Sol for larger changes, and it is slower. It leads on the 67-hunk case at every effort and on the 145-hunk case at low and high effort. Its average quality is higher at low (+1.9) and high (+2.9) effort, and level at medium (+0.5). Except on the 2-hunk case, it takes 1.1–1.6x as long. On the two small cases, the models are effectively tied.

Method

  • Ran upstream main at b9b6e10 (v1.16.0). It includes Support the GPT-6 model catalog and reasoning effort #177, which passes an explicit model and effort through to Codex unchanged.
  • Ran GPT-6 Sol and GPT-6.1 Sol at low, medium, and high reasoning. Each run covered all four cases twice. Within each effort, GPT-6 Sol ran first. All 48 attempts returned a walkthrough through the app server, and Codiff reported no model fallback.
  • Scored each walkthrough with evals/judge.mjs, pinned locally with -m gpt-6-astra -c model_reasoning_effort="high", so neither contestant judged its own output. A fake model name passed through -m failed at the API, which shows the flag selects the model. Nothing recorded the judge's effort. The published GPT-6 report used the maintainer's Codex default.
  • Used Pascal's Codex CLI 0.160.1 with ChatGPT sign-in, service_tier = "priority", and his global Codex instructions. The judge's event log shows these instructions loading. The generator's input of about 23k tokens for a 5.6k-character prompt suggests they load there too. Every run shares this setup, so the pairs are fair, but the absolute scores and times are not directly comparable with the GPT-6 report.
  • Quality is the average of two 0–100 scores. Generation time is the median of two attempts.

Results

Model and effort 2-hunk small 9-hunk focused small 67-hunk medium 145-hunk large
GPT-6 Sol, low 100.0 / 8.0s 89.5 / 6.7s 78.5 / 13.0s 66.0 / 14.4s
GPT-6.1 Sol, low 100.0 / 6.2s 83.5 / 8.4s 87.0 / 20.6s 71.0 / 23.7s
GPT-6 Sol, medium 100.0 / 8.8s 91.0 / 8.1s 80.0 / 15.6s 67.5 / 20.7s
GPT-6.1 Sol, medium 100.0 / 6.2s 90.0 / 10.5s 84.0 / 20.3s 66.5 / 26.0s
GPT-6 Sol, high 100.0 / 7.1s 95.0 / 11.0s 75.5 / 23.3s 66.5 / 35.2s
GPT-6.1 Sol, high 99.5 / 8.8s 91.5 / 12.4s 85.0 / 29.0s 72.5 / 44.2s

Each cell is average judge quality / median generation time.

GPT-6.1 Sol minus GPT-6 Sol, in quality points and generation-time ratio:

Effort 2-hunk small 9-hunk focused small 67-hunk medium 145-hunk large
low +0.0 / 0.78x -6.0 / 1.25x +8.5 / 1.59x +5.0 / 1.64x
medium +0.0 / 0.70x -1.0 / 1.28x +4.0 / 1.30x -1.0 / 1.25x
high -0.5 / 1.25x -3.5 / 1.12x +9.5 / 1.24x +6.0 / 1.26x
Model and effort Average quality Median generation
GPT-6 Sol, low 83.5 10.5s
GPT-6.1 Sol, low 85.4 14.5s
GPT-6 Sol, medium 84.6 12.2s
GPT-6.1 Sol, medium 85.1 15.4s
GPT-6 Sol, high 84.3 17.1s
GPT-6.1 Sol, high 87.1 20.7s

Our GPT-6 Sol rerun lands within 6 points of the published GPT-6 Sol scores. It scores 0.5 to 6 points higher on the three smaller cases and 5 to 6 points lower on the 145-hunk case. It also runs faster than the published times. The published run's service tier is unknown, so the cause is unclear.

Findings

  • 2-hunk case: Both models reach the ceiling at 99.5 to 100. This case no longer separates them.
  • 9-hunk focused case: GPT-6 Sol scores 1 to 6 points higher, but the gap rests on an inconsistent judge. In 4 of 6 attempts, the judge penalized GPT-6.1 Sol for claiming that nested sections stay out of the parent's copied code. All 6 GPT-6 Sol walkthroughs made a similar claim, for example "excluding nested details", and drew no penalty. Treat this case as a tie.
  • 67-hunk case: GPT-6.1 Sol leads at every effort, by 8.5 at low, 4.0 at medium, and 9.5 at high. It brings regression tests into the main path, with 61–66% main-path hunk coverage against 37–44%. The judge flagged GPT-6 Sol in all 6 attempts for leaving regression tests out of the main path, and in 4 of them for filing every test under "Other changes". GPT-6.1 Sol drew that complaint for specific tests in 2 attempts. Both models still underexplain the backend fallback behavior.
  • 145-hunk case: GPT-6.1 Sol leads at low by 5.0 and at high by 6.0, and its attempt scores do not overlap with GPT-6 Sol's: 71 and 71 against 66 and 66, then 74 and 71 against 66 and 67. The models tie at medium. The judge flagged GPT-6 Sol in all 6 attempts and GPT-6.1 Sol in 4 for filing the MergeRequestReviewApp tests under support. It flagged GPT-6.1 Sol twice for describing provider-blocked actions as disabled, while Panels.tsx hides them.
  • Speed: GPT-6.1 Sol is faster on the 2-hunk case at low and medium effort, and 1.1–1.6x slower elsewhere. At low effort, it writes 44–63% more output tokens on the two larger cases. GPT-6.1 Sol had no cached input tokens in 20 of 24 attempts, against 6 of 24 for GPT-6 Sol, which may inflate its times. On the 2-hunk case, times vary up to 1.8x between two attempts of the same setting, so its ratios are noise. On the 9-hunk case, attempts agree within 1.12x.

Decision

  • Larger changes: GPT-6.1 Sol gives better quality. At low effort, it reaches 87.0 on the 67-hunk case in 20.6s. At high effort, it reaches 72.5 on the 145-hunk case in 44.2s.
  • Small changes: No difference in quality. GPT-6 Sol is 1.1–1.3x faster on the 9-hunk case.
  • Noise: Two attempts per case cannot establish a gap under 5 points, and scores for one setting vary by up to 8 points. The gaps that clear this bar are GPT-6.1 Sol's leads on the 67-hunk case at low and high effort and on the 145-hunk case at low and high effort.
  • Unanswered: This run does not show whether GPT-6.1 Sol beats the production defaults, GPT-5.6 Terra and GPT-5.5 at low reasoning, because this judge did not score them. GPT-6.1 Sol's best 145-hunk score of 72.5 sits below the 75.0 that GPT-5.5 scored in the published run, under a different judge.

Run labels

gpt6-sol-low-20261006, gpt61-sol-low-20261006, gpt6-sol-medium-20261006, gpt61-sol-medium-20261006, gpt6-sol-high-20261006, gpt61-sol-high-20261006. The smoke test is smoke-20261006.

@pascalandy

Copy link
Copy Markdown
Contributor Author

Awesome, this totally makes sense.

By the way, Is there something here that I could improve? I'm really trying hard here to to have the most important information. of sharing issues and PRs on my own projects and thank four projects I want to contribute to.

@cpojer

cpojer commented Oct 7, 2026

Copy link
Copy Markdown
Member

I think walkthroughs could be improved literally by 10x – this requires a bunch of interface exploration of what is the right information to surface to users.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants