lora - #2
Open
xrsrke wants to merge 118 commits into
Open
Conversation
Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
Co-authored-by: Peter Jin <pjin@nvidia.com> Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
2f3dfb9fe621e887a943f62ba5cfd9f8212031cd c6ceac80f58dac7928c6d1865a6c805584a7f002 bc22275aa2dbc734a413a82655964258d3618e10 https://github.com/NVIDIA-NeMo/RL/pull/1613/changes NVIDIA-NeMo#1917 Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com> Co-authored-by: Peter Jin <pjin@nvidia.com>
1a38b15aabc734746537cdf875b80ce757371e09 84610d93e21c44d4bb8fb860258376ab9db5d116 NVIDIA-NeMo#1918 Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com> Co-authored-by: Peter Jin <pjin@nvidia.com> Co-authored-by: Jiaqi Zeng <jiaqiz@nvidia.com>
b3f1a39a34fabf54d6551b058a9f97fc976bc2c3 e9b5d9f89cd0fd90d4a1f10a8874382db27c5b93 NVIDIA-NeMo#1955 Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com> Co-Authored-By: Jiaqi Zeng <jiaqiz@nvidia.com>
Related to NVIDIA-NeMo#1913, but may need update for final config Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
Signed-off-by: Jiaqi Zeng <jiaqiz@nvidia.com> Signed-off-by: Terry Kong <terryk@nvidia.com> Co-authored-by: Jiaqi Zeng <jiaqiz@nvidia.com> Co-authored-by: Terry Kong <terryk@nvidia.com>
d72308ca3bbb9a057a10be2e461ee84dea12c35e part of NVIDIA-NeMo#1978 Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com> Co-Authored-By: Gerald Shen <geshen@nvidia.com>
TODO: do we have to change sync path too? e456a41066ad388c6744cbc8f656257526c37106 b3f1a39a34fabf54d6551b058a9f97fc976bc2c3 NVIDIA-NeMo#1960 Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com> Co-Authored-By: Jiaqi Zeng <jiaqiz@nvidia.com>
e476235662b9537f32bce78da08917df1c64158e part of NVIDIA-NeMo#1906 Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com> Co-Authored-By: Parth Chadha <pchadha@nvidia.com>
e476235662b9537f32bce78da08917df1c64158e part of NVIDIA-NeMo#1906 Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com> Co-Authored-By: Jiaqi Zeng <jiaqiz@nvidia.com>
e476235662b9537f32bce78da08917df1c64158e 9fd46676352e065036b8f8f792492c8349798a6d cc0d8813a786c16e02deb4e9fee014671f723e51 part of NVIDIA-NeMo#1906 Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com> Co-Authored-By: Jiaqi Zeng <jiaqiz@nvidia.com>
c68baf7a2009b0dce297e312812a4a0bb8432884 NVIDIA-NeMo#1908 Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com>
012af2a51d22def1ff170ecb3594ffe538898ff9 NVIDIA-NeMo#1957 Related: NVIDIA-NeMo#1891 Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com> Co-Authored-By: Jiaqi Zeng <jiaqiz@nvidia.com>
d2b9bf68b704a603a888953c496baa1fbfedf6bb NVIDIA-NeMo#1956 Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com> Co-Authored-By: Jiaqi Zeng <jiaqiz@nvidia.com>
ffcb0d176f6ecc5a6761ffb287fdc3692cb78fd4 NVIDIA-NeMo#1910 Related: NVIDIA-NeMo#1838 Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com> Co-Authored-By: Jiaqi Zeng <jiaqiz@nvidia.com>
…ging. c42464d8815babe9a04b2c461e68606b91475671 NVIDIA-NeMo#1958 Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com> Co-Authored-By: Peter Jin <pjin@nvidia.com>
TODO: may need refactor, ask haifeng 967261cab1055fa8a8d93fa17d7c40bd03aa1241 NVIDIA-NeMo#1919 Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com> Co-Authored-By: Gerald Shen <geshen@nvidia.com>
8e8cbc7939501e9b516120314d7e57a8c235cdf3 NVIDIA-NeMo#1954 Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com> Co-Authored-By: Jiaqi Zeng <jiaqiz@nvidia.com>
db1886be891f8193ec80d04d61e49e60ce9e8b8f NVIDIA-NeMo#1950 Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com> Co-Authored-By: Jiaqi Zeng <jiaqiz@nvidia.com>
c1abc09b4b9fa1a75cbd10f5ba07c74442c5ff67 NVIDIA-NeMo#1949 Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com> Co-Authored-By: Jiaqi Zeng <jiaqiz@nvidia.com>
…s generated and consumed c7fb07e0ca37b1f6c3a362191fbf78e8db314cfa part of NVIDIA-NeMo#1906 Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com> Co-Authored-By: Jiaqi Zeng <jiaqiz@nvidia.com>
58caed9041f662a4a456bbf20c1d652479a0fab6 NVIDIA-NeMo#1947 Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com> Co-Authored-By: Jiaqi Zeng <jiaqiz@nvidia.com>
e024f5bf4a90f1fd36040a336826ced764f3ef97 NVIDIA-NeMo#1946 Signed-off-by: Yi-Fu Wu <yifu.wu@gmail.com> Co-Authored-By: Yashaswi Karnati <ykarnati@nvidia.com>
…odels.multi_lora - 13 modules vendored verbatim from nousnet src/nousnet/rl/lora/multi/ + core/multi_lora_config + debug helpers; only import paths rewritten (nousnet.rl.lora.multi -> nemo_rl.models.multi_lora) - patched files (loss_functions, automodel/train, lm_policy, dtensor_policy_worker_v2) now import the vendored package; no nousnet dependency remains anywhere in nemo_rl/ - examples/run_sft_multi_lora.py: native entry point (stock run_sft.py data path + multi_lora pluggable wiring: MultiAdapterLoss + MultiAdapterDataLoader + n_adapters dispatch) - patches/automodel/: the 1-hunk Automodel lora.py dispatch, import-rewritten - tests: 8 suites / 176 tests ported, all passing (torchtitan env)
… nousnet NOT importable in-container)
…location), restore campaign runtime env + mounts
…c_name/keep_top_k, sft.val_*, logger backends) that nousnet's registry auto-filled at runtime
…erge, campaign values win; dynamic_batching/seq_packing/megatron OFF) — replaces piecemeal key filling
…d_config — parity with run_sft.py; full config now resolve-validated offline
…atches/ (container Automodel predates attr_shard_placement_fn); driver PYTHONPATH -> native worktree (campaign parity, kills version-skew class)
…ampaign runtime configs, deep-merged + resolve-validated; submit script (exact-init, traces ON, campaign env parity); vendored pairwise-loss analyzer (stdlib+matplotlib only)
…ha 800/800 all adapters, deltas at platform floor
…gn used tokenizer default (uniform +14tok wrapper diff, sha 0/800 overlap vs campaign)
…-ports) into the recipe, drop nousnet slurm_helpers + external Automodel checkout references, add in-container Automodel-integration assertion and standalone audit script
…nfigs against NeMo-RL roots, battery scripts take CANON via env (fail closed with native-export instructions), chart script reference-dir optional, error strings point at nemo_rl.models.multi_lora
- delete nemo_rl/models/multi_lora/debug/ (dump_everything, forward_backward_fingerprint) and their env-gated call sites (NOUSNET_DUMP_EVERYTHING, NOUSNET_FORWARD_BACKWARD_DIAG). Both gates were OFF in every validated native100 run — no runtime change. - drop unconditional step-0 [NOUSNET DEBUG] print blocks from sft_train (token_mask diagnosis leftovers from the Super-120B investigation; print-only). - untrack results/* + charts_native100/* evidence artifacts (stay on disk; results/ was already gitignored, add charts_native100/ too). Verified: py_compile all touched files; zero residual refs to deleted modules.
- base.yaml = the validated native100 V3 multi config with run identity
factored into run_name/paths/single_dataset interpolations (values otherwise
byte-identical). New runs are a 2-8 line overlay.
- parity100_{noclip,clip1}_{multi,single_a..d}.yaml replace the 20 v1/v3
copy-paste configs; smoke_10step_multi.yaml for quick functional checks.
- Offline equivalence proof (OmegaConf-resolved, identity-normalized):
all 10 old-vs-new pairs identical modulo run-name rename; fingerprints in
golden_native100v3_251641_251650/config_equivalence_proof.txt.
- scripts: submit_parity_battery.sh replaces the v1/v3 battery submitters;
analyze_pairwise_loss.py / plot_parity_loss_differences.py renamed and
de-hardcoded (--prefix/--steps); drop plot_native100_curves.py (superseded).
- drop nemo_rl/models/multi_lora/nemo_env.py (vendoring vestige, zero refs).
- add recipe README: config scheme, battery, verification, env knob docs.
Overlay configs inherit via 'defaults: base.yaml', resolved relative to the config file's directory; the run executes from the EXP_DIR copy, so the base must be copied alongside it.
Multi-LoRA stacks all N adapters in the leading dimension of lora_A/lora_B, so that axis carries adapter IDENTITY. If FSDP shards dim 0, a rank holds only a subset of adapter slots while per-token routing still indexes global adapter ids — silently wrong results, a hang, or a native crash. Nothing detected this. Fail closed instead of warning: * adapter.py: add assert_stacked_lora_fsdp_placement(). Requires every MultiLinearLoRA to keep all adapter slots LOCAL, rejects Shard(0), and requires Shard(1) in distributed setups. All-gathers its own result so a placement or implementation-source divergence between workers is caught rather than silently tolerated, then prints MULTILORA_FSDP_PLACEMENT_OK with the module count and implementation sha256. * dtensor_policy_worker_v2.py: run that assertion on every worker after FSDP/PEFT setup and before exact-init import or the first forward, gated on n_adapters > 1. * moe_routing.py: promote the Shard(0) case from logger.warning to RuntimeError. Warning-and-continuing walked straight into the corruption it detected. * routing.py: drop the blanket try/except around install_moe_expert_routing. It swallowed real install failures and silently reverted to the legacy slot-0 fallback — the exact behaviour per-token routing exists to replace. * sft_8gpu_native.slurm: copy every file listed under `defaults:` from the submitted config's own directory (fail closed if one is missing) instead of only a hard-coded base.yaml, so multi-level config bundles survive the move to EXP_DIR. Tests (+9): placement assertion accepts an unsharded unit model, rejects Shard(0), rejects a missing dim-1 shard, rejects dropped local slots, rejects a model with no MultiLinearLoRA; MoE install rejects adapter-axis sharding; and the vendored parallelizer honours the _fsdp_shard_dim hint while preserving the stock default. Validation: focused suite tests/unit/models/multi_lora = 200 passed / 0 failed in-container against this working tree; tests/unit collects 1853 tests with no collection errors. The full unit suite was NOT executed and is not claimed. Distributed evidence is a 10-lane parity battery (packed 4-adapter multi vs four true singles): all 10 lanes execution-success under the strict tuple (exit 0, COMMAND_DONE, anchored terminal marker after training-complete, exact trace cardinality on 8 ranks); pairwise A<->A/B<->B/C<->C/D<->D with SHA-256 and token-count joins gives A/B/C PASS and D FAIL at one isolated step of 100 (0.3380 @ step 57), while the same-node 10-step control passes 4/4. PARITY_REPORT.md records the full method, per-adapter numbers, job ids and the determinism caveat: a same-config-twice probe reproduces only to ~0.0898 per step, so the single D excursion is real but NOT attributable to Multi-LoRA until run-to-run determinism is restored.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.