Repository navigation
Support the GPT-6 model catalog and reasoning effort - #177
Conversation
- Purpose: let users select current Codex models and supported reasoning efforts - Impact: preserve custom IDs, forward explicit effort, and separate effort-specific cache entries by GPT-6.1-Sol via Codex
- Purpose: prevent automatic large-diff routing from sending explicit effort to an incompatible model - Impact: retain selected models, report fallback causes, and allow later discovery retries by GPT-6.1-Sol via Codex
PlanCodiff should let users select the models and reasoning efforts offered by their installed Codex CLI, without requiring a Codiff release whenever OpenAI updates its catalog. This includes GPT‑6 Astra, GPT‑6 Sol, GPT‑6 Luna, and GPT‑6.1 Sol when available to the user. My specific use case is GPT‑6.1 Sol with The scope includes model selection, reasoning effort, saved preferences, and compatibility with older Codex clients. Other agent backends and account-access management are out of scope. CMO (current Mode of operation)Codiff maintains its own model list and reasoning defaults.
See the predefined catalog and model normalization. This creates a gap between what Codex supports and what Codiff exposes. FMO (future Mode of operation)Updating the static list would address today's models but require repeated maintenance. I recommend using Codex's own catalog. Codex app-server provides
The resulting selector should expose new models as Codex adds them, without another Codiff catalog update. How we'll know it works
Premortem
by GPT-6.1-Sol via Codex |
cpojer
left a comment
There was a problem hiding this comment.
Awesome, this totally makes sense. I do think the defaults still require frequent eval runs before changing them though. If you look into the evals folder, for example GPT-6 was only better in one situation.
|
We have a little win here : GPT-6.1 Sol walkthrough evaluation — 2026-10-06GPT-6.1 Sol writes better walkthroughs than GPT-6 Sol for larger changes, and it is slower. It leads on the 67-hunk case at every effort and on the 145-hunk case at low and high effort. Its average quality is higher at low (+1.9) and high (+2.9) effort, and level at medium (+0.5). Except on the 2-hunk case, it takes 1.1–1.6x as long. On the two small cases, the models are effectively tied. Method
Results
Each cell is average judge quality / median generation time. GPT-6.1 Sol minus GPT-6 Sol, in quality points and generation-time ratio:
Our GPT-6 Sol rerun lands within 6 points of the published GPT-6 Sol scores. It scores 0.5 to 6 points higher on the three smaller cases and 5 to 6 points lower on the 145-hunk case. It also runs faster than the published times. The published run's service tier is unknown, so the cause is unclear. Findings
Decision
Run labels
|
By the way, Is there something here that I could improve? I'm really trying hard here to to have the most important information. of sharing issues and PRs on my own projects and thank four projects I want to contribute to. |
|
I think walkthroughs could be improved literally by 10x – this requires a bunch of interface exploration of what is the right information to surface to users. |
Why
Codiff's Codex menu is limited to a predefined model list, and unrecognized model IDs are replaced with the default. This prevents selecting newer installed models such as GPT-6.1 Sol with high reasoning.
This change reads the installed Codex CLI's visible, paginated model catalog in the background and adds a model-specific Reasoning Effort menu. Custom model IDs remain usable when discovery is unavailable. The new
settings.openAIReasoningEffortoverride reaches both execution paths and separates walkthrough cache entries by effort. Existing automatic defaults remain compatible.Scope
settings.openAIReasoningEffortthrough both execution pathsExample configuration:
{ "settings": { "openAIModel": "gpt-6.1-sol", "openAIReasoningEffort": "high" } }Verification
pnpm exec vp check --fixpassedpnpm exec vp test: 1,143 passed, 7 skipped across 110 filespnpm exec vp run buildpassedgpt-6.1-solwithhighaction_required)AI assistance: GPT-6.1 Sol through Codex performed research, design, code, tests, documentation, and review responses. Codex agents reviewed design and source. Claude Opus 5.5 independently reviewed the change and verified its routing, error-reporting, and discovery-retry fixes.
banana banana banana
by GPT-6.1-Sol via Codex