Skip to content

Latest commit

 

History

21 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Volume or Coupling? A Scale-Dependent Dissociation in Constraint Recovery of Language-Model Loops

In one sentence: a controlled experiment asked whether a language model recovers from a conflicting instruction because it generated more text, or because it was coupled to a second agent — and the answer flips with scale: at 1.5B, volume explains the whole effect; at 7B, volume fails completely (S = 0%, X = 0%) and only the coupled loop recovers (26.7%, p = .0078).

中文一句话:同一个受控实验,问"模型从冲突指令中恢复"的能力来自上下文变多还是来自被耦合到另一个 agent —— 答案随规模反转:1.5B 下体量完全解释得通,7B 下体量彻底失效、只有耦合能恢复。

v2 headline — at 7B, the volume account fails: only the coupled loop recovers (p = .0078).

DOI License: CC BY 4.0 GitHub stars

Author: Simin Yuan
Contact: yleven120@gmail.com
Zenodo (v2): 10.5281/zenodo.21200851 · concept DOI (always latest): 10.5281/zenodo.21157628
arXiv: pending (v1 under review; v2 replacement ready)


Why this is worth a closer look

  • Every exact p-value is recomputable with zero dependencies and no GPU. The raw outcomes are embedded in code/analysis_mcnemar.py; one stdlib-only command re-derives them all — including the headline p = .0078 at 7B and the deflationary p = 0.7744 at 1.5B.
  • The result that weakens the paper's own story is kept, not buried. The v1 finding was that the coupling advantage at 1.5B is fully explained by generation volume; v1 code, data, figures and PDF all remain in the repo alongside the v2 scale extension, so the research trajectory is auditable rather than overwritten.
  • Determinism is designed in, not claimed. Greedy decoding, 30 fixed paired openers, a no-perturbation control (violation 0.000 throughout), and a pre-registered STABLE/TRANSIENT rule for the 7B trajectories — replayed delay matches recorded delay in 8/8 runs.
  • File naming tells you which version a result belongs to (v1: no scale suffix; v2: _3B / _7B / _scale / _traj, or a v2/ subfolder), matching the Zenodo version record.

Quick start (no GPU, no install)

git clone https://github.com/simin-yuan/context-volume-not-coupling
cd context-volume-not-coupling
python code/analysis_mcnemar.py

That prints every exact McNemar p-value in the paper, from the embedded raw outcomes, using only the Python standard library. Expected headline numbers: 7B — C vs X p = 0.0078, C vs S p = 0.0078, C vs D4 p = 0.0391; 1.5B — C vs X p = 0.7744.

Regenerate the two paper figures (needs matplotlib; still no GPU):

pip install matplotlib && python code/make_figures.py   # writes fig_scale_rates.png, fig_7b_traj.png

How to verify the claims

  1. The statistics. python code/analysis_mcnemar.py — stdlib only, runs in under a second. If the printed p-values disagree with the abstract, the paper is wrong.
  2. Embedded data vs. the raw file. code/analysis_mcnemar.py and code/make_figures.py embed the outcomes (D / DATA dicts); data/raw_outcomes_v2.txt holds the same outcomes in readable form (sed -n '1,20p' data/raw_outcomes_v2.txt). Compare them — the dicts must match the file.
  3. The 7B trajectories. python code/exp7_replay_7B_trajectories.py (L4 GPU) replays all 8 recovered 7B coupled runs; fidelity is asserted by comparing replayed delay with recorded delay (8/8), and the STABLE/TRANSIENT classification rule was fixed before inspection.
  4. v1 vs. v2 boundary. ls paper paper-v2 code figures data — every v1 file is suffix-free; v2 additions carry _3B / _7B / _scale / _traj. Nothing was renamed to hide the v1 result.
  5. The record itself. Zenodo v2 DOI (versioned) and the concept DOI that always resolves to the latest version, both listed above.

Project Overview

This repository contains the full replication package for the paper "Volume or Coupling? A Scale-Dependent Dissociation in Constraint Recovery of Language-Model Loops" (v2, July 2026).

The study investigates whether coupling a language model to a second agent improves recovery from a conflicting instruction, or whether the effect is merely due to increased generation volume. We use a minimal perturbation–recovery protocol (uppercase constraint, polite lowercase perturbation, 8-exchange window) across four conditions (Coupled, Single, Context-matched single, Dose-4) and three model scales (1.5B, 3B, 7B) within the Qwen2.5-Instruct family.

Key findings:

  • At 1.5B (v1 base result): Coupling advantage is fully explained by volume (C: 43%, S: 17%, X: 37%; C vs. X p=0.774).
  • At 3B: All conditions at ceiling (90–100%) – perturbation fails to capture.
  • At 7B (v2 reversal): Volume completely fails (S=0%, X=0%, D4=3.3%), but coupling uniquely recovers (C=26.7%, C vs. X p=0.0078). Recovery is slow (delays 2–7) and driven by asymmetric capture depth: the directly perturbed agent is deeply captured, while its shielded partner (never directly receiving the instruction) remains only shallowly captured by mimicry and re-anchors, supplying constraint-consistent evidence.

The volume effect is therefore scale-bounded: it only operates at small scales; at larger scales, structural shielding (coupling) provides volume-irreducible recovery.


Repository Structure

How to distinguish v1 vs. v2 files:
All v1 files are named without scale suffixes (e.g., run_1.5B.py, raw_1.5B.txt, fig1_setup.png).
All v2 additions are clearly marked with _3B, _7B, _scale, or _traj in filenames, or placed in subfolders named v2/ where applicable. Check each folder for detailed file listings.

Why keep v1 files (paper/ and all v1 data/code/figures)?
To maintain a complete, transparent research trajectory – reviewers and readers can compare the original deflationary result (v1) with the extended scale-dependent findings (v2). This aligns with Zenodo versioning and arXiv replacement records.


Getting Started

Hardware Requirements

  • 1.5B tier: Runs on free Google Colab T4 (under 3 GPU-hours total).
  • 3B & 7B tiers: Require L4 GPU (approx. 5 GPU-hours combined).

Run the Experiments

All scripts are in code/, as plain Python (extracted from the notebooks as run; Colab setup: pip install -q -U transformers accelerate scipy).

  • 1.5B tier (T4): exp1_pilot_1p5B.py, exp2_1p5B_coupled_vs_single.py, exp3_1p5B_context_matched.py, exp4_1p5B_dose_response.py
  • 3B / 7B tiers: exp5_scale_3B_7B.py — set MODEL_NAME per tier and run once per tier in a fresh session. The script header records the constraint: 3B runs on a T4, 7B needs an L4 (~15 GB VRAM in fp16), and 4-bit quantization must not be used because it would confound the scale comparison.
  • 7B controls & trajectory replay: exp6_verify_7B_controls.py, exp7_replay_7B_trajectories.py
  • No-GPU analysis: analysis_mcnemar.py, make_figures.py

Greedy decoding ensures deterministic outputs, so re-running will reproduce exactly the reported data.

Reproduce Figures

The figures/ folder contains all .png outputs. You can regenerate the v2 main figures with code/make_figures.py (needs matplotlib, no GPU). The v1 figures are also included for reference.


Citation

If you use this work, please cite the v2 Zenodo record (versioned DOI below; the concept DOI 10.5281/zenodo.21157628 always resolves to the latest version):

@misc{yuan2026volume,
  author       = {Yuan, Simin},
  title        = {Volume or Coupling? A Scale-Dependent Dissociation in Constraint Recovery of Language-Model Loops},
  year         = {2026},
  doi          = {10.5281/zenodo.21200851},
  publisher    = {Zenodo},
  note         = {Version 2}
}

License

This work is licensed under a Creative Commons Attribution 4.0 International License.

Acknowledgments

Experiment code and manuscript drafting were assisted by an AI system (Claude, Anthropic). The author thanks Kai-Wei Chang for arXiv endorsement and feedback on figures.

About

A controlled, reproducible experiment showing that the coupled-loop recovery advantage is scale-dependent: generation volume fully explains it at 1.5B, but fails at 7B, where only the coupled loop recovers (p = .0078).

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages