Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
60 commits
Select commit Hold shift + click to select a range
b79d6cb
Add large-scale GGA + GGA+U fp32 training config
forklady42 May 28, 2026
2e2ce78
Add Modal training + Globus data-load for large-scale GGA/GGA+U run
forklady42 May 28, 2026
0a1e0f6
Add prep_volume helper to wire the data Volume after Globus transfer
forklady42 May 28, 2026
c841b35
Point populate_volume at oa-electrai with dedicated read-only secret
forklady42 Jun 15, 2026
35160da
Parallelize populate_volume S3 download with ThreadPoolExecutor
forklady42 Jun 17, 2026
f279059
Use .spawn() so the populate_volume pull survives local CLI disconnect
forklady42 Jun 17, 2026
cf2e7fd
Pack zarr stores as .zarr.zip to fit under Modal Volume's inode cap
forklady42 Jun 18, 2026
1b65484
Use .spawn() in train.py so multi-day runs survive local CLI disconnects
forklady42 Jun 18, 2026
bee89bc
Add Lambda Cloud runbook (scripts/lambda/) for the long-run training …
forklady42 Jun 18, 2026
fea5263
Source ~/.bashrc in lambda run_training.sh for non-interactive callers
forklady42 Jun 19, 2026
6662ed6
Lambda: prep_data accepts both .zarr.zip and .zarr/ (S3 is unpacked)
forklady42 Jun 19, 2026
42ef893
Lambda run_training: grep WANDB_API_KEY from ~/.bashrc, not source it
forklady42 Jun 19, 2026
1924e0e
Lambda: default data + checkpoints to NFS filesystem; fix run_training
forklady42 Jun 19, 2026
3ffae7f
Lambda run_training: simplify /data/ -> DATA_ROOT sed (YAML list-dash…
forklady42 Jun 19, 2026
b4e16e0
Lambda: auto-detect NFS_ROOT from /lambda/nfs/*
forklady42 Jun 19, 2026
c094da0
Lambda run_training: WANDB_MODE_OVERRIDE env var for offline/disabled…
forklady42 Jun 20, 2026
fafb857
Lambda run_training: anchor path-rewrite sed on leading space
forklady42 Jun 20, 2026
81c6ea0
Lambda multinode: drop-in 2-node DDP launcher + runbook
forklady42 Jun 22, 2026
de3edf7
Add PORT_PLAN.md: migration plan from H100:4 -> 2-node H100:16
forklady42 Jun 22, 2026
5436fcd
Add MONITOR.md + monitor_status.sh: monitoring script + LLM-driven ho…
forklady42 Jun 22, 2026
483b306
Lambda monitor: EC2 resident Claude agent (active operator)
forklady42 Jun 22, 2026
af2b108
Lambda monitor: default to Opus (claude-opus-4-8) for monitoring
forklady42 Jun 22, 2026
f1aa848
Lambda monitor: escalation degrades to a journal banner when no Slack…
forklady42 Jun 22, 2026
8f44ba0
Lambda monitor: fix snapshot cmd allowlist match; wire Slack state vi…
forklady42 Jun 22, 2026
8a34d49
Add width-64 GGA+GGA+U config + w64 mode for Lambda training
forklady42 Jun 24, 2026
7c7fc95
W64 run: grid-size cap + DDP timeout + frequent resume checkpoint
forklady42 Jun 24, 2026
3c70889
W64: switch precision fp16-mixed -> bf16-mixed (fp16 diverged to NaN)
forklady42 Jul 13, 2026
2f79873
Add train_watchdog.sh: cron-driven NaN/stall/down alert + auto-stop
forklady42 Jul 13, 2026
bb249dc
Use PyTorch cu128 index for CUDA torch wheels on linux-aarch64
forklady42 Jul 17, 2026
fe8342b
Add width-96 config for the width ablation (single-variable vs W64)
forklady42 Jul 17, 2026
c6e9df1
Add CoreWeave stage-in script for charge-density data
forklady42 Jul 17, 2026
8afaa9f
Merge branch 'betsy/gga-ggau-w96' into betsy/coreweave-w96
forklady42 Jul 17, 2026
7e8b068
Point W96 config at the CoreWeave stage-in layout
forklady42 Jul 17, 2026
d7d4c69
Add CoreWeave training wrapper and dry-run script
forklady42 Jul 17, 2026
59051b0
Switch staging and checkpoint sync from s5cmd to rclone
forklady42 Jul 18, 2026
e04fc55
Stage into /uv/cache: the only host-persistent mount in Iris task pods
forklady42 Jul 18, 2026
4b85565
Run wandb offline with a sync sidecar on CoreWeave
forklady42 Jul 18, 2026
616991a
Use step-based interval for the resume checkpoint, not wall-clock
forklady42 Jul 19, 2026
179935a
Promote the newest last*.ckpt before resume
forklady42 Jul 20, 2026
8db696e
Add W128 config and generic Iris submit helper
forklady42 Jul 20, 2026
6966ae8
Add exploratory W256 config for dry-run evaluation
forklady42 Jul 21, 2026
40e4a4f
dry_run.py: grid ladder, multi-iter timing, cudnn-benchmark probe flags
forklady42 Jul 21, 2026
5d3bb7c
Add W192 config: next width rung if it clears the cuDNN kernel cliff
forklady42 Jul 21, 2026
333c1a2
Add voxel-band counter for width-specific kernel-cliff caps
forklady42 Jul 21, 2026
37ff4a0
dry_run.py: add --n-channels override for width-ceiling probes
forklady42 Jul 22, 2026
63325c0
Add W160 config: widest cliff-safe rung on the unmodified capped dataset
forklady42 Jul 22, 2026
fb6f51c
submit.sh: raise default retry budget for preemption-heavy periods
forklady42 Jul 28, 2026
07c80d7
Name wandb runs from cfg.run_name
forklady42 Jul 29, 2026
1222a42
Shorten wandb run names to w{width}_{MMDD-HHMM}
forklady42 Jul 29, 2026
bbb924b
Add LR/weight-decay-vs-width sweep: stage-A configs, split subsampler…
forklady42 Jul 30, 2026
a91a49f
Sweep stage A: W64 lr=8e-3 edge extension + W&B review script
forklady42 Jul 31, 2026
0a678da
Sweep stage B: WD grid at per-width lr* (32=2e-3, 64=4e-3, 96=2e-3)
forklady42 Jul 31, 2026
b10cf57
review_lrwd_sweep: keep _rep2 noise-bar reruns separate from stage-A …
forklady42 Jul 31, 2026
b2a9826
Sweep doc: record stage A+B results and stage-C recipes
forklady42 Aug 1, 2026
d69acba
Stage C confirmations (W32/64/96 tuned lr) + W128 at lr 2e-3
forklady42 Aug 1, 2026
a04d07c
Reconcile resume checkpoints against the bucket with --update
forklady42 Aug 1, 2026
9f0a387
Fix ckpt_path in stage-C/W128 tuned configs: match run_training.sh st…
forklady42 Aug 4, 2026
946c43d
Refuse resume when local last.ckpt is staler than the bucket's
forklady42 Aug 5, 2026
b142576
W128 withheld-test-set eval: 0.481% combined NMAE (GGA 0.480 / GGA+U …
forklady42 Aug 7, 2026
aafadb5
Add final width-ablation figures
forklady42 Aug 31, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions data/MP/sweep_splits/gga+u_split_sweep12k.json

Large diffs are not rendered by default.

1 change: 1 addition & 0 deletions data/MP/sweep_splits/gga_split_sweep12k.json

Large diffs are not rendered by default.

Binary file added docs/figures/w160_train_val.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file added docs/figures/width_ablation_nmae.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
234 changes: 234 additions & 0 deletions docs/lr_wd_width_sweep.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,234 @@
# LR / weight-decay sweep across width (W32 / W64 / W96)

Status: stages A+B COMPLETE (launched 2026-07-30/31, finished 2026-07-31 —
16 + 15 trials, all succeeded; results below). Stage C not launched.

## Results (see W&B `mp-gga-ggau-lrwd`; aggregate with scripts/review_lrwd_sweep.py)

Best val NMAE% on the 8-epoch/12K proxy. Noise bar from the three verbatim
anchor reruns: **±0.04–0.07 points** (1.914→1.968, 1.592→1.665, 1.524→1.566).

| lr | W32 | W64 | W96 |
|---------|-------|-------|-------|
| 2.5e-4 | 2.330 | 1.998 | 1.943 |
| 5e-4 | 2.133 | 1.841 | 1.836 |
| 1e-3 | 1.935 | 1.708 | 1.757 |
| 2e-3 | **1.914** | 1.759 | **1.524** |
| 4e-3 | 1.954 | **1.592** | 1.562 |
| 8e-3 | — | 1.580 | — |

| wd @ lr* (AdamW) | W32 | W64 | W96 |
|---------|-------|-------|-------|
| 0 (+rerun) | 1.914 / 1.968 | 1.592 / 1.665 | 1.524 / 1.566 |
| 1e-4 | 1.968 | 1.631 | 1.535 |
| 1e-3 | 1.995 | 1.610 | 1.640 |
| 1e-2 | 1.970 | 1.654 | 1.549 |
| cross (lr*/2, 1e-3) | 1.904 | 1.638 | 1.731 |

Conclusions:

1. **lr\* does not shrink with width** — every width sits in a flat 2–4e-3
basin (differences at/inside the noise bar). The production lr=1e-3 is
suboptimal at all widths, and the penalty grows with width: ~1% relative
at W32, ~7% at W64, ~13% at W96. Extrapolation to W128/W160: use 2e-3.
2. **Weight decay is neutral across 1e-4–1e-2 at every width** — even on
this subset, where overfitting bites ~9× earlier than at full scale, so
at 111K it has even less room to help. Recommend keeping **wd=0**
(continuity with all incumbents); wd up to 1e-2 is demonstrably safe if
ever wanted for other reasons. (Only outlier: W96 @ 1e-3 = 1.640,
marginally past the bar; with 1e-2 neutral on both sides it reads as an
unlucky draw, not a trend.)
3. W64's stage-A curiosities dissolved: the 8e-3 "win" (1.580 vs 1.592) and
the 2e-3 dip (1.759; cross term at same lr with wd landed 1.638) are
both within run-to-run noise.
4. Width ordering at tuned recipes is unchanged: W96 1.524 < W64 1.592 <
W32 1.914.

**Stage C recipes**: W32 (2e-3, 0), W64 (4e-3, 0), W96 (2e-3, 0) — with
2e-3 defensible everywhere given the flat basin.

## Why

Every rung of the width ladder (`mp-gga-ggau-width`: W64, W96, W128, W160)
trains with the same recipe — `lr: 0.001`, `weight_decay: 0.0` — inherited
from the June W64 run. For Adam in standard parametrization the optimal LR
typically *shrinks* as width grows (roughly ∝ 1/width in the μP limit), so
the wider rungs are plausibly mistuned and the width-scaling conclusions
conflate capacity with recipe mismatch. W96's 17.8% win over W64 survived
that handicap; the question is whether the gaps (and the W128/W160 verdicts)
change once each width gets its own recipe.

This sweep tunes LR and WD independently at widths 32, 64, 96, then uses the
per-width optima to (a) re-examine the width ranking and (b) extrapolate a
recipe for W128/W160 rather than sweeping at those (much more expensive)
widths.

## Trial design (proxy task)

Full production runs are ~5 days; sweep trials compress two axes and change
nothing else:

- **Data**: same staged capped dataset (111,257 ids). Train is subsampled to
~12K structures (frac 0.1112 of each functional's capped train split, so
the 76/24 gga/gga+u mix is preserved: 9,139 + 2,862). **Validation is the
untouched production val set** (847 + 262 = 1,109), so trial `val_loss` is
on the same yardstick as the production curves. The subset lives in two
small split JSONs checked into the repo (`data/MP/sweep_splits/`) and
referenced by repo-relative path — they ride along in the submit.sh bundle,
so no bucket upload and no interaction with the `.staged.ok` warm-node
skip.
- **Schedule**: `epochs: 8`, `warmup_length: 1` — each trial runs its own
complete warmup+cosine schedule, compressed. Comparing LRs mid-schedule is
misleading; comparing completed short schedules preserves ranking much
better. 8 epochs × 12K ≈ 24K optimizer steps ≈ 0.9 production epochs.
- **Hardware/batch**: 1 node, GB200x4, DDP, batch 1/GPU — the *same
effective batch (4) as production*, so the tuned LR transfers directly.
(A single-GPU trial would change the effective batch and invalidate it.)
- Everything else identical to `config_gga_gga+u_w96.yaml`: bf16-mixed, no
activation checkpointing, Adam betas (0.9, 0.99), offline W&B + sidecar.

Per-trial wall-clock (from measured production step times): W96 ≈ 4 h,
W64 ≈ 3 h, W32 ≈ 2–2.5 h.

## Stage A — LR grid at wd = 0

`lr ∈ {2.5e-4, 5e-4, 1e-3, 2e-3, 4e-3}` × `width ∈ {32, 64, 96}` = **15
trials, ≈ 50 node-hours**. Factor-2 spacing matches the flatness of Adam LR
basins; the 16× range is centered on the incumbent 1e-3 and wide enough for
a ~3× shift across 32→96.

Decision rule per width: `lr*` = argmin of final-epoch `val_loss` (sanity:
best-epoch val and curve shape agree; a diverged/NaN trial is a top-edge
signal, likely at 4e-3 for the wider models). **If a width's optimum lands
on a grid edge, extend one more factor-2 point before moving on.** The old
Optuna tier-1 result (lr ≈ 3.5e-3 best for W32, different data/precision)
hints W32 may press the top edge.

## Stage B — WD grid at lr*

Weight decay uses **AdamW** (`optimizer: adamw`, added to
`lightning.py`) — the incumbent plain Adam applies `weight_decay` as
coupled L2, which gets rescaled per-parameter by the adaptive denominator
and makes values neither interpretable nor comparable across widths. At
wd = 0 the two optimizers are identical, so nothing about existing runs or
Stage A changes.

Per width, 5 trials:

- `wd ∈ {1e-4, 1e-3, 1e-2}` at `lr*`
- one cross term `(lr*/2, wd=1e-3)` to catch LR–WD interaction (in AdamW the
effective per-step decay is `lr·wd`, so a strong-WD winner may prefer a
lower LR)
- a **verbatim rerun of the Stage-A anchor** `(lr*, wd=0)` — the run-to-run
noise bar that decides whether Stage B differences are real

**15 trials, ≈ 50 node-hours.**

Caveat to carry into analysis: on a 12K subset, overfitting appears ~9×
earlier than on the full set, so the subset will overstate how much WD
helps. Treat Stage B as establishing the *tolerance and trend* (does wd hurt
below some threshold? does λ_eff = lr·wd stay constant across widths?), and
confirm the absolute choice in Stage C. If a large wd wins on the subset,
prefer the largest wd that is *neutral-or-better* rather than the argmin.

## Stage C — full-scale confirmation

One run per width at the tuned `(lr*, wd*)`: full 111K dataset, production
config, **10–12 epoch budget** (~45 h W96, ~32 h W64 on GB200x4). Compare
val_loss at matched epochs against the incumbent curves, which already exist
at lr=1e-3/wd=0:

- W64: `mp-gga-ggau-width` run `zz3oecp7` (16-epoch best 0.008398)
- W96: run `z21di7sl` / `revived-energy-7` (ep10 0.008336 … ep27 0.006901)

W32 has no incumbent in this campaign; run it only if the trend fit needs
the third full-scale point.

## Analysis

- Fit `log lr*` vs `log width` → slope α; predict lr for W128/W160.
(Expected α between 0 and −1.) Same check on wd: is `lr*·wd*`
width-stable?
- Re-examine the width ranking at matched epochs with tuned recipes — this
feeds the width-ablation writeup.
- All sweep runs log to W&B project **`mp-gga-ggau-lrwd`**. Run names are
auto-generated `w{width}_{MMDD-HHMM}` (restart-segment convention), so
aggregate by `config.lr` / `config.weight_decay` /
`config.model.n_channels` — same pattern as
`scripts/review_width_ablation.py`.

## Runbook

```bash
# 1. Get the production split files locally (small; either source works)
rclone copy cw:mp/chg_datasets/functionals/gga/split_capped.json /tmp/sweep/gga/
rclone copy cw:mp/chg_datasets/functionals/gga+u/split_capped.json /tmp/sweep/gga+u/
# or: scp della:/scratch/gpfs/ROSENGROUP/common/globus_share_OA/mp/chg_datasets/functionals/gga/split_capped.json ...

# 2. Generate the sweep splits (checked into the branch for provenance)
uv run python scripts/coreweave/make_sweep_split.py \
/tmp/sweep/gga/split_capped.json data/MP/sweep_splits/gga_split_sweep12k.json --frac 0.1112
uv run python scripts/coreweave/make_sweep_split.py \
"/tmp/sweep/gga+u/split_capped.json" "data/MP/sweep_splits/gga+u_split_sweep12k.json" --frac 0.1112

# 3. Stage A configs (already generated in src/electrai/configs/MP/sweep_lrwd/)
uv run python scripts/coreweave/gen_lrwd_sweep.py --stage a
# ... submit the printed submit.sh lines from a clean worktree; jobs run at
# batch priority and won't starve the live W128 run.

# 4. After Stage A: regenerate for Stage B with the real per-width winners
uv run python scripts/coreweave/gen_lrwd_sweep.py --stage b \
--best-lr 32=<lr*> 64=<lr*> 96=<lr*>
```

Each config filename stem doubles as the checkpoint namespace
(`run_training.sh` derives ckpt dirs from the stem), so trials never
cross-resume; the Stage-B noise rerun gets a `_rep2` stem for the same
reason.

## Known limitations

- 24K steps ≈ 0.9 production epochs per trial: LR ranking at short horizon
with a completed cosine is a standard, usually-reliable proxy, but a
near-tie between adjacent LRs at a width should be broken toward the
*lower* LR (long-horizon optima drift down, not up).
- WD conclusions from the subset overstate regularization benefit (see
Stage B caveat); Stage C is the arbiter.
- The scheduler steps per epoch, so an 8-epoch cosine has coarse resolution;
all trials share it, so comparisons are fair.

## W128 withheld-test-set evaluation (Aug 5–7, 2026)

Checkpoint `w128_ckpt_epoch55_val0.005785.ckpt` (stage-C W128 at lr 2e-3,
epoch 55) evaluated on the withheld `test` split of `split_capped.json` —
never seen in training or validation. Della jobs 12057752 (GGA) + 12111587
(GGA+U), fp32, single A100-80GB; config
`src/electrai/configs/MP/config_gga_gga+u_w128_test_della.yaml`; per-sample
CSVs in `/scratch/gpfs/ROSENGROUP/bb9080/w128_test_eval/results{,_padsu}/`.

| Subset | n | mean NMAE | median | p90 | p99 | max | share <1% |
|---|---|---|---|---|---|---|---|
| GGA | 1700 | 0.480% | 0.373% | 0.900% | 1.842% | 2.95% | 92.0% |
| GGA+U (PADS) | 524 | 0.483% | 0.397% | 0.857% | 1.794% | 3.36% | 94.5% |
| **Combined** | **2224** | **0.481%** | 0.379% | 0.890% | 1.832% | 3.36% | 92.6% |

**Combined test NMAE 0.481% is under the 0.5% ChargE3Net threshold** and
consistent with the 0.579% val loss (val is bf16-on-GB200; test is fp32).

Provenance notes uncovered during this eval:

- The split files existed only in the buckets; they were staged from
`s3://oa-electrai` to `/scratch/gpfs/ROSENGROUP/bb9080/w128_test_eval/`
and their test/validation indices verified byte-identical to the
training-time splits (via `data/MP/sweep_splits/*_sweep12k.json`
pass-through).
- **The training data's `gga+u` inputs are PADS, not SAD.** Della's
`functionals/gga+u_sad` and `gga+u_pads` share identical filelists and
labels; only inputs differ. Evaluating with SAD inputs gives ~10% NMAE
across the whole GGA+U test set (the model degrades SAD inputs below
their own 7–9% input error); PADS inputs give 0.48%. Probe: Della job
12111541. Any "GGA+U" result from the gga-ggau width campaign is a
PADS-input result.
- bf16-mixed autocast OOMs this checkpoint on A100 (a ~75 GiB allocation
inside a decoder conv on a 108³ sample, Della job 12049722) — evaluate at
fp32 on Della, matching the throughput benchmark (job 12046200).
52 changes: 52 additions & 0 deletions job_w128_test_eval.slurm
Original file line number Diff line number Diff line change
@@ -0,0 +1,52 @@
#!/bin/bash
#SBATCH --job-name=w128-test-eval
#SBATCH --nodes=1
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=8
#SBATCH --mem=64G
#SBATCH --gres=gpu:1
#SBATCH --constraint=gpu80
#SBATCH --time=04:00:00
#SBATCH --output=/scratch/gpfs/ROSENGROUP/bb9080/logs/w128-test-eval-%j.out

# Evaluate the W128 lr2e-3 checkpoint (epoch 55, val 0.005785) on the withheld
# GGA + GGA+U test set (split_capped.json test keys: 1700 + 524 samples).
# Config: src/electrai/configs/MP/config_gga_gga+u_w128_test_della.yaml

module purge
module load anaconda3/2025.6
conda activate electrai
module load proxy/default
export PATH=$HOME/.local/bin:$PATH
# Variable grid sizes fragment the caching allocator; expandable segments
# lets large per-sample activations reuse freed reserved memory.
export PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True

cd /scratch/gpfs/ROSENGROUP/bb9080/electrai

EVAL=/scratch/gpfs/ROSENGROUP/bb9080/w128_test_eval
CKPT_SRC=/scratch/gpfs/ROSENGROUP/bb9080/checkpoints/w128_ckpt_epoch55_val0.005785.ckpt

mkdir -p ${EVAL}/ckpt ${EVAL}/results
# test entrypoint expects <ckpt_path>/last.ckpt
ln -sfn ${CKPT_SRC} ${EVAL}/ckpt/last.ckpt

echo "=== W128 test-set eval ==="
echo "Checkpoint: ${CKPT_SRC}"
echo "Results: ${EVAL}/results"
echo "Start: $(date)"
echo ""

srun uv run --no-sync python ./src/electrai/entrypoints/main.py test \
--config src/electrai/configs/MP/config_gga_gga+u_w128_test_della.yaml

echo ""
echo "End: $(date)"

METRICS=${EVAL}/results/metrics.csv
if [ ! -f ${METRICS} ]; then
echo "ERROR: metrics.csv not found at ${METRICS}"
exit 1
fi
N_ROWS=$(tail -n +2 ${METRICS} | wc -l)
echo "metrics.csv rows: ${N_ROWS} (expected 2224)"
Loading
Loading