Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 6 additions & 6 deletions .agents/workflows/nmr-reaction-kinetics.md
Original file line number Diff line number Diff line change
Expand Up @@ -18,9 +18,9 @@ Follow Steps 0–5 of the `reaction-to-nmr-quantification.md` workflow to identi
Perform Wasserstein deconvolution at each time point and assemble kinetics curves.

- *Skill Reference:* `chem-nmr-analysis` (`kinetics.py`)
- **Action:** Pass all reference spectra, all time-point crude spectra (in chronological order), corresponding time values, proton counts, and component names to `kinetics.py`.
- **Output:** `kinetics.csv` (mole fractions + Wasserstein distance per time point) and `kinetics_plot.png` (mole fraction vs time + fit quality vs time).
- **Decision:** Inspect the kinetics plot. If curves are non-monotonic or a Wasserstein distance spike occurs at a single time point, that spectrum likely has baseline or phasing issues — consider excluding it and re-running. If all WD values exceed 0.15, the fit is poor across the board; revisit reference alignment (Step 5) before trusting the kinetics.
- **Action:** Pass all reference spectra, all time-point crude spectra (in chronological order), corresponding time values, proton counts, and component names to `kinetics.py`, with `--baseline_window LO HI` set to a ppm stretch that is signal-free in every time point and reference. Do not use the legacy `--baseline_correct` on noisy spectra: each spectrum's minimum is a noise excursion, and subtracting it biases all fractions toward an even mixture.
- **Output:** `kinetics.csv` (mole fractions, Wasserstein distance, noise fraction and subtracted baseline per time point) and `kinetics_plot.png` (mole fraction vs time + fit quality vs time).
- **Decision:** Inspect the kinetics plot and the `noise_fraction` column. If curves are non-monotonic or a WD or noise spike occurs at a single time point, that spectrum likely has baseline or phasing issues — consider excluding it and re-running. If all WD values exceed 0.15, or noise fractions exceed about 0.2, revisit reference alignment and the baseline window (Step 5) before trusting the kinetics.

## Summary Checklist for the Agent
When tasked with "extract kinetics from NMR time series":
Expand All @@ -29,9 +29,9 @@ When tasked with "extract kinetics from NMR time series":
3. [ ] Resolve names to SMILES. (`drug-db-pubchem`)
4. [ ] Predict products via ReactionT5. (`chem-nmr-analysis` `predict_products.py`)
5. [ ] Generate reference spectra. (`chem-nmr-predict`)
6. [ ] Visual inspection of overlay. (`chem-nmr-analysis` `plot.py`)
7. [ ] Run kinetics deconvolution. (`chem-nmr-analysis` `kinetics.py`)
8. [ ] Inspect kinetics plot and verify chemical plausibility.
6. [ ] Visual inspection of overlay; choose a signal-free baseline window. (`chem-nmr-analysis` `plot.py`)
7. [ ] Run kinetics deconvolution with `--baseline_window`. (`chem-nmr-analysis` `kinetics.py`)
8. [ ] Inspect kinetics plot and noise fractions, and verify chemical plausibility.

## References
- Ciach, M. et al., "Masserstein: linear resampling of mass spectra by optimal transport", *Rapid Commun. Mass Spectrom.*, 2020.
Expand Down
8 changes: 4 additions & 4 deletions .agents/workflows/reaction-to-nmr-quantification.md
Original file line number Diff line number Diff line change
Expand Up @@ -50,16 +50,16 @@ Predict 1H NMR spectra for all identified species (reactants + products + solven
Overlay the crude spectrum with all references to verify alignment before deconvolution.

- *Skill Reference:* `chem-nmr-analysis` (`plot.py`)
- **Action:** Generate an overlay plot and inspect it. Check that reference peaks align with mixture peaks and that no major mixture peaks are unaccounted for.
- **Action:** Generate an overlay plot and inspect it. Check that reference peaks align with mixture peaks and that no major mixture peaks are unaccounted for. Pick a signal-free ppm stretch (no peaks in the mixture or any reference) for the baseline window of Step 6, and note any isolated diagnostic signals that could be fitted on their own with `--ppm-range`.
- **Decision:** If peaks don't align, ask the user about ppm referencing. If major peaks are unmatched, revisit Step 1 for missing components.

### Step 6: Deconvolution
Quantify component mole fractions via Wasserstein-distance deconvolution.

- *Skill Reference:* `chem-nmr-analysis` (`deconvolve.py`)
- **Action:** Run `deconvolve.py` with the crude spectrum, all reference spectra, proton counts, and component names. Always use `--plot` and `--json`.
- **Output:** Estimated mole fractions, Wasserstein distance (fit quality), and a multi-panel deconvolution plot.
- **Decision:** Review the Wasserstein distance. If WD > 0.15, the fit is poor — check for missing components or ppm offsets before trusting the proportions. If the proportions contradict known chemistry, flag this to the user.
- **Action:** Run `deconvolve.py` with the crude spectrum, all reference spectra, proton counts, and component names, and `--baseline-window LO HI` set to the signal-free stretch from Step 5 (add `--baseline-stat max` for digitized spectra, whose baseline is a one-sided floor). Do not use the legacy `--baseline-correct` (minimum subtraction) on noisy or digitized spectra. Always use `--plot` and `--json`.
- **Output:** Estimated mole fractions, Wasserstein distance (fit quality), unexplained (noise) signal fraction, subtracted baselines, and a multi-panel deconvolution plot.
- **Decision:** Review the Wasserstein distance *and* the noise fraction. If WD > 0.15, the fit is poor — check for missing components or ppm offsets before trusting the proportions. If the noise fraction exceeds about 0.2, revisit the baseline window/statistic and consider restricting `--ppm-range` to diagnostic signals (with `--range-protons` giving each component's protons in that range), even when WD looks acceptable. If the proportions contradict known chemistry, flag this to the user.

## References
- Ciach, M. et al., "Masserstein: linear resampling of mass spectra by optimal transport", *Rapid Commun. Mass Spectrom.*, 2020.
Expand Down
48 changes: 26 additions & 22 deletions skills/chem-db-mof/SKILL.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
---
name: chem-db-mof
description: Query multiple MOF databases (QMOF via MPContribs; ARC-MOF DB7/Majumdar et al. via Zenodo) and download CIF structures with optional element or identifier filters.
description: Query multiple MOF databases (QMOF via MPContribs; ARC-MOF DB7/Majumdar et al. via Materials Cloud) and download CIF structures with optional element or identifier filters.
metadata:
category: [chemistry]
venv: [cpu]
Expand All @@ -15,7 +15,7 @@ Provide a unified interface for retrieving Metal-Organic Framework (MOF) crystal
| Database | Alias | Size | Access | Structures |
|---|---|---|---|---|
| Quantum MOF (QMOF) | `qmof` | ~20,000 DFT-relaxed | MPContribs API | DFT-optimized CIFs + bandgaps |
| ARC-MOF DB7 (Majumdar et al.) | `arcmof-majumdar` | 12,316 hypothetical | Zenodo stream | CIFs with REPEAT partial charges |
| ARC-MOF DB7 (Majumdar et al.) | `arcmof-majumdar` | 23,891 hypothetical (`ddmof_1`–`ddmof_23891`) | Materials Cloud archive | CIFs with EQeq partial charges + `mof_data.csv` metadata |

## Prerequisites

Expand All @@ -34,9 +34,9 @@ Decide which database to query and which element/identifier filters to apply.
- Use `--identifier` for a specific CSD refcode (e.g., `KAXQIL`)

**For ARC-MOF DB7 (Majumdar et al.)** — best for diverse hypothetical MOFs with underrepresented inorganic SBUs:
- Use `--elements` for element filtering (e.g., `Zn,O,C`)
- Use `--identifier` for a specific structure ID (e.g., `DB7_00042`)
- **First run**: downloads `geometric_properties.csv` (~110 MB) to `~/.cache/arcmof/` — one-time only; subsequent runs are fast
- Use `--elements` for element filtering (e.g., `Zn`, or `Zn,F` for fluorinated Zn MOFs); by default every listed element must be present, while `--element-match any` keeps structures containing at least one of them
- Use `--identifier` for a specific structure name (e.g., `ddmof_559`); a full name matches exactly, anything else matches as a substring
- **First run**: downloads the Majumdar archive `mof_data.tar.gz` (~149 MB) from Materials Cloud to `~/.cache/arcmof/` — one-time only; subsequent runs are fast

### Step 2: Run the query

Expand All @@ -50,10 +50,10 @@ MP_API_KEY=<your_key> ${CLAUDE_SKILL_DIR}/../../venv/run cpu python ${CLAUDE_SKI
```

```bash
# ARC-MOF DB7 (Majumdar) — 20 Zn,O,C hypothetical MOFs
# ARC-MOF DB7 (Majumdar) — 20 Zn-containing hypothetical MOFs
${CLAUDE_SKILL_DIR}/../../venv/run cpu python ${CLAUDE_SKILL_DIR}/scripts/query_mof_db.py \
--database arcmof-majumdar \
--elements Zn,O,C \
--elements Zn \
--max-results 20 \
--output-dir ./research/<date>_<task>/structures/arcmof_db7
```
Expand All @@ -62,7 +62,7 @@ ${CLAUDE_SKILL_DIR}/../../venv/run cpu python ${CLAUDE_SKILL_DIR}/scripts/query_
# ARC-MOF DB7 — retrieve a specific structure by identifier
${CLAUDE_SKILL_DIR}/../../venv/run cpu python ${CLAUDE_SKILL_DIR}/scripts/query_mof_db.py \
--database arcmof-majumdar \
--identifier DB7_00042 \
--identifier ddmof_559 \
--output-dir ./research/<date>_<task>/structures/arcmof_db7
```

Expand All @@ -72,7 +72,8 @@ ${CLAUDE_SKILL_DIR}/../../venv/run cpu python ${CLAUDE_SKILL_DIR}/scripts/query_
|---|---|---|
| `--database` | both | `qmof` or `arcmof-majumdar` |
| `--formula` | qmof | Element/formula filter string (e.g., `Zn,O,C`) |
| `--elements` | arcmof-majumdar | Comma-separated required elements; ALL must be present |
| `--elements` | arcmof-majumdar | Comma-separated elements; by default a structure must contain ALL of them |
| `--element-match` | arcmof-majumdar | `all` (default) or `any` (keep structures containing at least one listed element, e.g. a mixed-metal screening set) |
| `--identifier` | both | Specific structure name or ID substring |
| `--max-results` | both | Max CIFs to download (default: 10) |
| `--output-dir` | both | Directory for output CIF files |
Expand All @@ -82,7 +83,7 @@ ${CLAUDE_SKILL_DIR}/../../venv/run cpu python ${CLAUDE_SKILL_DIR}/scripts/query_

The script saves:
- Individual `.cif` files named by structure identifier
- `arcmof_db7_metadata.csv` (ARC-MOF only) — geometric properties for the downloaded subset
- `majumdar_metadata.csv` (ARC-MOF only) — Majumdar `mof_data.csv` rows (building blocks, topology, volume, density, CO2 uptake) for the downloaded subset

Verify the download:
```bash
Expand All @@ -93,14 +94,16 @@ ls -lh <output-dir>/*.cif | head -20

The first call with `--database arcmof-majumdar` performs:

1. **Metadata download** (~110 MB, one-time): `geometric_properties.csv` cached at `~/.cache/arcmof/`
2. **DB7 filtering**: identifies the 12,316 Majumdar structures from the full 288k-entry CSV
3. **CIF streaming**: streams the ARC-MOF tarball (`ARCMOF_20241004.tar.gz`, ~670 MB) and extracts only the requested CIFs — the stream is read once but only matching files are written to disk
1. **Archive download** (~149 MB, one-time): the Majumdar et al. `mof_data.tar.gz` from Materials Cloud (DOI 10.24435/materialscloud:yn-de), cached at `~/.cache/arcmof/majumdar_mof_data.tar.gz`
2. **Metadata**: `mof_data.csv` inside the archive lists all 23,891 structures with their building blocks and properties
3. **CIF extraction**: the nested `mof_structures.tar` is read sequentially and only the CIFs that pass the identifier/metal filters are written to disk; an exact `--identifier` stops at the first hit

Subsequent runs with the same `--output-dir` skip already-downloaded CIFs.

## Examples

**Literature-validated example: ARC-MOF DB7 `ddmof_559`** (the top post-combustion CO₂ MOF in Figure 7 of Majumdar et al. 2021). See [examples/arcmof_db7_ddmof_559/README.md](examples/arcmof_db7_ddmof_559/README.md). It retrieves the structure by exact identifier and checks its composition, Ni₄(μ₃-OH)₂ node stoichiometry, volume and density against the paper and the ARC-MOF metadata.

**Example 1: Query Zn MOFs from QMOF for CO₂ screening pre-processing**
```bash
MP_API_KEY=<your_mp_api_key> \
Expand All @@ -116,38 +119,39 @@ ${CLAUDE_SKILL_DIR}/../../venv/run cpu python ${CLAUDE_SKILL_DIR}/scripts/query_
# Zn-based
${CLAUDE_SKILL_DIR}/../../venv/run cpu python ${CLAUDE_SKILL_DIR}/scripts/query_mof_db.py \
--database arcmof-majumdar \
--elements Zn,O,C \
--elements Zn \
--max-results 50 \
--output-dir ./research/2026-03-27_arcmof_zn

# Ni-based
${CLAUDE_SKILL_DIR}/../../venv/run cpu python ${CLAUDE_SKILL_DIR}/scripts/query_mof_db.py \
--database arcmof-majumdar \
--elements Ni,O,C \
--elements Ni \
--max-results 50 \
--output-dir ./research/2026-03-27_arcmof_ni

# Mg-based
${CLAUDE_SKILL_DIR}/../../venv/run cpu python ${CLAUDE_SKILL_DIR}/scripts/query_mof_db.py \
--database arcmof-majumdar \
--elements Mg,O,C \
--elements Mg \
--max-results 50 \
--output-dir ./research/2026-03-27_arcmof_mg
```

> **Tip:** You can expand diversity by adding more elements to `--elements` (e.g., `Zn,Ni,O,C,N` to retrieve MOFs containing all of those elements simultaneously), or run separate queries per metal node and combine the resulting CIF directories for a broader screening campaign.
> **Tip:** To build a mixed-metal screening set in one query, list several metals with `--element-match any` (e.g., `--elements Zn,Ni,Mg --element-match any`), or run separate queries per metal node and combine the resulting CIF directories. With the default `all`, adding elements narrows the search (e.g., `Zn,F` returns only fluorinated Zn MOFs).

## Constraints

- **API limits**: QMOF via MPContribs has rate limits; keep `--max-results` ≤ 100 per call.
- **ARC-MOF first-run time**: Downloading the metadata CSV (~110 MB) takes ~1–2 min; streaming the tarball for CIF extraction adds ~5–15 min depending on how many structures are requested and network speed.
- **ARC-MOF CIF fallback**: If some DB7 structures are not found in `ARCMOF_20241004.tar.gz`, they may reside in `all_structures_1.tar.gz` or `all_structures_2.tar.gz`. Update `ARCMOF_STRUCTURES_NAME` in the script if needed.
- **Element filtering (ARC-MOF)**: Requires a `formula` or `chemical_formula` column in `geometric_properties.csv`. If the column is absent, all DB7 entries are returned without element filtering.
- **Post-download**: Structures from ARC-MOF DB7 include REPEAT partial charges embedded in the CIF. These can be used directly for classical force-field simulations but should be relaxed with an MLIP before running Widom insertion (see [`chem-sorption-relax`](../chem-sorption-relax/SKILL.md)).
- **ARC-MOF first-run time**: Downloading the ~149 MB archive takes a few minutes depending on network speed; later runs read the cached copy.
- **ARC-MOF availability**: The Materials Cloud download server can be temporarily unavailable (HTTP 503); retry later if the first download fails.
- **Element filtering (ARC-MOF)**: Parsed from each CIF's `_atom_site_type_symbol`; all listed elements must be present unless `--element-match any` is given.
- **Existing outputs**: If `--output-dir` already holds `--max-results` CIFs, the query is skipped regardless of the filters; use a fresh directory per query.
- **Post-download**: Structures from ARC-MOF DB7 include EQeq partial charges embedded in the CIF. These can be used directly for classical force-field simulations but should be relaxed with an MLIP before running Widom insertion (see [`chem-sorption-relax`](../chem-sorption-relax/SKILL.md)).

## References

- Raza, A. et al., "ARC–MOF: A Diverse Database of Metal-Organic Frameworks with DFT-Derived Partial Atomic Charges and Descriptors for Machine Learning", *Chem. Mater.*, 2022. [DOI: 10.1021/acs.chemmater.2c02485](https://doi.org/10.1021/acs.chemmater.2c02485)
- Burner, J. et al., "ARC–MOF: A Diverse Database of Metal-Organic Frameworks with DFT-Derived Partial Atomic Charges and Descriptors for Machine Learning", *Chem. Mater.* 35, 900–916, 2023. [DOI: 10.1021/acs.chemmater.2c02485](https://doi.org/10.1021/acs.chemmater.2c02485)
- Majumdar, S., Moosavi, S.M., Jablonka, K.M., Ongari, D., Smit, B., "Diversifying Databases of Metal Organic Frameworks for High-Throughput Computational Screening", *ACS Appl. Mater. Interfaces*, 2021. [DOI: 10.1021/acsami.1c16220](https://doi.org/10.1021/acsami.1c16220); dataset: *Materials Cloud Archive* 2021.126, [DOI: 10.24435/materialscloud:yn-de](https://doi.org/10.24435/materialscloud:yn-de)
- Chung, Y.G. et al., "Computation-Ready, Experimental Metal-Organic Frameworks: A Tool To Enable High-Throughput Screening of Nanoporous Crystals", *Chem. Mater.*, 2014 (QMOF precursor). [DOI: 10.1021/cm502594j](https://doi.org/10.1021/cm502594j)
- Rosen, A.S. et al., "Machine learning the quantum-chemical properties of metal-organic frameworks for accelerated materials discovery", *Matter*, 2021 (QMOF). [DOI: 10.1016/j.matt.2021.02.015](https://doi.org/10.1016/j.matt.2021.02.015)
Expand Down
Loading
Loading