Skip to content

lora - #2

Open
xrsrke wants to merge 118 commits into
mainfrom
phuc/multilora-native-nemorl
Open

lora#2
xrsrke wants to merge 118 commits into
mainfrom
phuc/multilora-native-nemorl

Conversation

@xrsrke

@xrsrke xrsrke commented Jul 24, 2026

Copy link
Copy Markdown

No description provided.

yfw and others added 30 commits February 17, 2026 23:22
Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
Co-authored-by: Peter Jin <pjin@nvidia.com>
Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
2f3dfb9fe621e887a943f62ba5cfd9f8212031cd
c6ceac80f58dac7928c6d1865a6c805584a7f002
bc22275aa2dbc734a413a82655964258d3618e10

https://github.com/NVIDIA-NeMo/RL/pull/1613/changes
NVIDIA-NeMo#1917

Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
Co-authored-by: Peter Jin <pjin@nvidia.com>
1a38b15aabc734746537cdf875b80ce757371e09
84610d93e21c44d4bb8fb860258376ab9db5d116

NVIDIA-NeMo#1918

Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
Co-authored-by: Peter Jin <pjin@nvidia.com>
Co-authored-by: Jiaqi Zeng <jiaqiz@nvidia.com>
b3f1a39a34fabf54d6551b058a9f97fc976bc2c3
e9b5d9f89cd0fd90d4a1f10a8874382db27c5b93

NVIDIA-NeMo#1955

Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
Co-Authored-By: Jiaqi Zeng <jiaqiz@nvidia.com>
Related to NVIDIA-NeMo#1913, but may need
update for final config

Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
Signed-off-by: Jiaqi Zeng <jiaqiz@nvidia.com>
Signed-off-by: Terry Kong <terryk@nvidia.com>
Co-authored-by: Jiaqi Zeng <jiaqiz@nvidia.com>
Co-authored-by: Terry Kong <terryk@nvidia.com>
d72308ca3bbb9a057a10be2e461ee84dea12c35e

part of NVIDIA-NeMo#1978

Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
Co-Authored-By: Gerald Shen <geshen@nvidia.com>
TODO: do we have to change sync path too?

e456a41066ad388c6744cbc8f656257526c37106
b3f1a39a34fabf54d6551b058a9f97fc976bc2c3

NVIDIA-NeMo#1960

Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
Co-Authored-By: Jiaqi Zeng <jiaqiz@nvidia.com>
e476235662b9537f32bce78da08917df1c64158e

part of NVIDIA-NeMo#1906

Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
Co-Authored-By: Parth Chadha <pchadha@nvidia.com>
e476235662b9537f32bce78da08917df1c64158e

part of NVIDIA-NeMo#1906

Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
Co-Authored-By: Jiaqi Zeng <jiaqiz@nvidia.com>
e476235662b9537f32bce78da08917df1c64158e
9fd46676352e065036b8f8f792492c8349798a6d
cc0d8813a786c16e02deb4e9fee014671f723e51

part of NVIDIA-NeMo#1906

Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
Co-Authored-By: Jiaqi Zeng <jiaqiz@nvidia.com>
c68baf7a2009b0dce297e312812a4a0bb8432884

NVIDIA-NeMo#1908

Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
012af2a51d22def1ff170ecb3594ffe538898ff9

NVIDIA-NeMo#1957

Related: NVIDIA-NeMo#1891

Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
Co-Authored-By: Jiaqi Zeng <jiaqiz@nvidia.com>
d2b9bf68b704a603a888953c496baa1fbfedf6bb

NVIDIA-NeMo#1956

Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
Co-Authored-By: Jiaqi Zeng <jiaqiz@nvidia.com>
ffcb0d176f6ecc5a6761ffb287fdc3692cb78fd4

NVIDIA-NeMo#1910
Related: NVIDIA-NeMo#1838

Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
Co-Authored-By: Jiaqi Zeng <jiaqiz@nvidia.com>
…ging.

c42464d8815babe9a04b2c461e68606b91475671

NVIDIA-NeMo#1958

Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
Co-Authored-By: Peter Jin <pjin@nvidia.com>
TODO: may need refactor, ask haifeng

967261cab1055fa8a8d93fa17d7c40bd03aa1241

NVIDIA-NeMo#1919

Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
Co-Authored-By: Gerald Shen <geshen@nvidia.com>
8e8cbc7939501e9b516120314d7e57a8c235cdf3

NVIDIA-NeMo#1954

Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
Co-Authored-By: Jiaqi Zeng <jiaqiz@nvidia.com>
db1886be891f8193ec80d04d61e49e60ce9e8b8f

NVIDIA-NeMo#1950

Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
Co-Authored-By: Jiaqi Zeng <jiaqiz@nvidia.com>
c1abc09b4b9fa1a75cbd10f5ba07c74442c5ff67

NVIDIA-NeMo#1949

Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
Co-Authored-By: Jiaqi Zeng <jiaqiz@nvidia.com>
…s generated and consumed

c7fb07e0ca37b1f6c3a362191fbf78e8db314cfa

part of NVIDIA-NeMo#1906

Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
Co-Authored-By: Jiaqi Zeng <jiaqiz@nvidia.com>
58caed9041f662a4a456bbf20c1d652479a0fab6

NVIDIA-NeMo#1947

Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
Co-Authored-By: Jiaqi Zeng <jiaqiz@nvidia.com>
e024f5bf4a90f1fd36040a336826ced764f3ef97

NVIDIA-NeMo#1946

Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
Co-Authored-By: Yashaswi Karnati <ykarnati@nvidia.com>
xrsrke added 26 commits July 17, 2026 16:17
…odels.multi_lora

- 13 modules vendored verbatim from nousnet src/nousnet/rl/lora/multi/ +
  core/multi_lora_config + debug helpers; only import paths rewritten
  (nousnet.rl.lora.multi -> nemo_rl.models.multi_lora)
- patched files (loss_functions, automodel/train, lm_policy,
  dtensor_policy_worker_v2) now import the vendored package; no nousnet
  dependency remains anywhere in nemo_rl/
- examples/run_sft_multi_lora.py: native entry point (stock run_sft.py data
  path + multi_lora pluggable wiring: MultiAdapterLoss + MultiAdapterDataLoader
  + n_adapters dispatch)
- patches/automodel/: the 1-hunk Automodel lora.py dispatch, import-rewritten
- tests: 8 suites / 176 tests ported, all passing (torchtitan env)
…location), restore campaign runtime env + mounts
…c_name/keep_top_k, sft.val_*, logger backends) that nousnet's registry auto-filled at runtime
…erge, campaign values win; dynamic_batching/seq_packing/megatron OFF) — replaces piecemeal key filling
…d_config — parity with run_sft.py; full config now resolve-validated offline
…atches/ (container Automodel predates attr_shard_placement_fn); driver PYTHONPATH -> native worktree (campaign parity, kills version-skew class)
…ampaign runtime configs, deep-merged + resolve-validated; submit script (exact-init, traces ON, campaign env parity); vendored pairwise-loss analyzer (stdlib+matplotlib only)
…ha 800/800 all adapters, deltas at platform floor
…gn used tokenizer default (uniform +14tok wrapper diff, sha 0/800 overlap vs campaign)
…-ports) into the recipe, drop nousnet slurm_helpers + external Automodel checkout references, add in-container Automodel-integration assertion and standalone audit script
…nfigs against NeMo-RL roots, battery scripts take CANON via env (fail closed with native-export instructions), chart script reference-dir optional, error strings point at nemo_rl.models.multi_lora
- delete nemo_rl/models/multi_lora/debug/ (dump_everything, forward_backward_fingerprint)
  and their env-gated call sites (NOUSNET_DUMP_EVERYTHING, NOUSNET_FORWARD_BACKWARD_DIAG).
  Both gates were OFF in every validated native100 run — no runtime change.
- drop unconditional step-0 [NOUSNET DEBUG] print blocks from sft_train
  (token_mask diagnosis leftovers from the Super-120B investigation; print-only).
- untrack results/* + charts_native100/* evidence artifacts (stay on disk;
  results/ was already gitignored, add charts_native100/ too).

Verified: py_compile all touched files; zero residual refs to deleted modules.
- base.yaml = the validated native100 V3 multi config with run identity
  factored into run_name/paths/single_dataset interpolations (values otherwise
  byte-identical). New runs are a 2-8 line overlay.
- parity100_{noclip,clip1}_{multi,single_a..d}.yaml replace the 20 v1/v3
  copy-paste configs; smoke_10step_multi.yaml for quick functional checks.
- Offline equivalence proof (OmegaConf-resolved, identity-normalized):
  all 10 old-vs-new pairs identical modulo run-name rename; fingerprints in
  golden_native100v3_251641_251650/config_equivalence_proof.txt.
- scripts: submit_parity_battery.sh replaces the v1/v3 battery submitters;
  analyze_pairwise_loss.py / plot_parity_loss_differences.py renamed and
  de-hardcoded (--prefix/--steps); drop plot_native100_curves.py (superseded).
- drop nemo_rl/models/multi_lora/nemo_env.py (vendoring vestige, zero refs).
- add recipe README: config scheme, battery, verification, env knob docs.
Overlay configs inherit via 'defaults: base.yaml', resolved relative to the
config file's directory; the run executes from the EXP_DIR copy, so the base
must be copied alongside it.
@github-actions github-actions Bot added the documentation Improvements or additions to documentation label Jul 24, 2026
xrsrke added 3 commits July 28, 2026 18:30
Multi-LoRA stacks all N adapters in the leading dimension of lora_A/lora_B, so
that axis carries adapter IDENTITY. If FSDP shards dim 0, a rank holds only a
subset of adapter slots while per-token routing still indexes global adapter
ids — silently wrong results, a hang, or a native crash. Nothing detected this.

Fail closed instead of warning:

* adapter.py: add assert_stacked_lora_fsdp_placement(). Requires every
  MultiLinearLoRA to keep all adapter slots LOCAL, rejects Shard(0), and
  requires Shard(1) in distributed setups. All-gathers its own result so a
  placement or implementation-source divergence between workers is caught
  rather than silently tolerated, then prints MULTILORA_FSDP_PLACEMENT_OK with
  the module count and implementation sha256.
* dtensor_policy_worker_v2.py: run that assertion on every worker after
  FSDP/PEFT setup and before exact-init import or the first forward, gated on
  n_adapters > 1.
* moe_routing.py: promote the Shard(0) case from logger.warning to RuntimeError.
  Warning-and-continuing walked straight into the corruption it detected.
* routing.py: drop the blanket try/except around install_moe_expert_routing.
  It swallowed real install failures and silently reverted to the legacy slot-0
  fallback — the exact behaviour per-token routing exists to replace.
* sft_8gpu_native.slurm: copy every file listed under `defaults:` from the
  submitted config's own directory (fail closed if one is missing) instead of
  only a hard-coded base.yaml, so multi-level config bundles survive the move
  to EXP_DIR.

Tests (+9): placement assertion accepts an unsharded unit model, rejects
Shard(0), rejects a missing dim-1 shard, rejects dropped local slots, rejects a
model with no MultiLinearLoRA; MoE install rejects adapter-axis sharding; and
the vendored parallelizer honours the _fsdp_shard_dim hint while preserving the
stock default.

Validation: focused suite tests/unit/models/multi_lora = 200 passed / 0 failed
in-container against this working tree; tests/unit collects 1853 tests with no
collection errors. The full unit suite was NOT executed and is not claimed.
Distributed evidence is a 10-lane parity battery (packed 4-adapter multi vs
four true singles): all 10 lanes execution-success under the strict tuple
(exit 0, COMMAND_DONE, anchored terminal marker after training-complete, exact
trace cardinality on 8 ranks); pairwise A<->A/B<->B/C<->C/D<->D with SHA-256 and
token-count joins gives A/B/C PASS and D FAIL at one isolated step of 100
(0.3380 @ step 57), while the same-node 10-step control passes 4/4.

PARITY_REPORT.md records the full method, per-adapter numbers, job ids and the
determinism caveat: a same-config-twice probe reproduces only to ~0.0898 per
step, so the single D excursion is real but NOT attributable to Multi-LoRA
until run-to-run determinism is restored.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants