Spectroscopy with an audit trail — from raw FID to regulatory-ready.
SpectraCheck is the spectroscopy intelligence engine inside MolTrace. NMR, LC-MS, HRMS, and MS/MS arrive as one evidence stack — typed, traceable, multi-modal by default. Every numerical claim links back to the spectrum it came from.
Every stage emits a typed Pydantic model with stable JSON keys. Every emission is pinned to the immutable raw archive by SHA-256 + recipe hash — designed so you can replay a report from any prior date and reproduce bit-identical output.
01
Ingest
Raw FID archive (Bruker / Agilent-Varian) lands in the immutable vault. SHA-256 hashed. Vendor metadata extracted: SFO1/BF1 (Bruker) or sfrq/reffrq (Varian).
Outputs
Archive · vendor · field_mhz · acquisition params
02
Process
Fourier transform + apodization + phasing + baseline correction. Every parameter recipe-hash-linked to the unchanged raw archive — designed for byte-for-byte reproducible replay.
Legacy detector (default) or GSD-Prompt-3 (experimental opt-in). Both surface the same envelope: peaks, environments, category_counts, environment_counts.
Outputs
Peaks · environments · category mix · per-peak fit metrics
Structure-elucidation report composer designed for regulatory submission. Every numerical claim is hyperlinked to its source. Human signoff required before release.
A pharmaceutical R&D group operates across NMR + LC-MS + HRMS + MS/MS simultaneously. SpectraCheck fuses these as one evidence stack — not separate apps — and uses cross-modal contradictions (HRMS exact mass disagreeing with NMR-implied formula) as first-class warnings.
1D NMR (¹H, ¹³C)
Layers 22 · 24
Accepts
Raw FID · JCAMP-DX · CSV · vendor exports
Bruker + Agilent-Varian FID parsing via nmrglue
Solvent-aware shift windows + residual-peak masking
Voigt / Lorentzian fitting with per-peak QC residuals
Multiplet clustering for environment counting
2D NMR (COSY · HSQC · HMQC · HMBC)
Layers 23 · 25
Accepts
Processed 2D NMR · vendor archives
Guarded behind ENABLE_2D_NMR feature flag
Cross-peak detection + connectivity assignment
Symmetrisation + denoise pipelines
Evidence consumed by candidate-scoring layers
HRMS · MS/MS
Layers 29 · 30 · 31 · 32
Accepts
mzML · mzXML · processed peak lists
HRMS exact-mass candidate + bounded formula search
MS1 adduct + isotope pattern inference
MS/MS fragmentation tree + diagnostic neutral losses
Cross-modal corroboration with NMR-implied formula
Contradictions surface as first-class warnings
Auto-classification
Every peak gets a category.
Peaks aren't just numbers. SpectraCheck applies a 5-category taxonomy — backed by curated impurity-shift tables and solvent residual-peak windows — so reviewers can scan a peak list and immediately see what's signal, what's noise, and what needs follow-up.
Categories are exposed as structured fields ( peak.category, peak.solvent_hit, peak.impurity_match) — not buried in a confidence number. Each impurity match cites the reference shift, the observed shift, and the delta.
Compound
Target analyte signals — passed to candidate scoring.
Solvent
Residual CHCl₃ at 7.26, DMSO-d₆ at 2.50, CD₃OD at 3.31. Auto-masked from candidate scoring.
Phasing distortions, sidebands, t₁ noise. Excluded from candidate scoring; flagged for reviewer.
¹³C satellite
Spinning sidebands of strong ¹H peaks at natural-abundance ¹³C positions. Down-weighted.
Two detectors, one envelope
Legacy by default. GSD by opt-in. Both behind the same gate.
We ship a stable, evidence-pipeline legacy detector as the default. The experimental GSD-Prompt-3 backend (industry-standard peak detection with auto-classification) ships as opt-in until it clears our internal validation gate. Both return the same { peaks, environments, category_counts } envelope so consumer code is detector-agnostic.
Capability
Legacy (default)
GSD (experimental)
Peak detection on NMRShiftDB2 corpus
Median Δ = 14
Median Δ = 19 (Δ = 5 compound-only)
Solvent auto-detect (structured field)
Inferred via impurity-match table
94.4% on our 18-fixture internal benchmark
Per-peak fit QC (χ²ᵣ, RMSE, FWHM, S/N, baseline σ)
Now exposed (Phase 24, normalised pending)
Native — normalised to baseline σ
Multiplicity + J-values + integration
Native — multiplet notation per peak
Not exposed
Candidate structure matching · DP4 ranking
Full evidence pipeline
Peak detection only — no candidate matching
Category classification (5-set)
Open string (detection + chemical-region)
Closed enum, native
Promotion status
Default backend
Experimental · opt-in · gated
Full A/B methodology + per-fixture numbers documented in Field notes and the technical white paper §3.1.
Internal validation benchmarks
The numbers we benchmark.
Every figure below comes from our internal benchmark corpora. The regression gate runs in CI; any drift larger than 50% on any single fixture fails the build with the fixture_id called out by name.
94.4%
Solvent auto-detect (internal benchmark)
On our 18-fixture internal NMRShiftDB2-derived benchmark (17 of 18 fixtures with a reference). Strict gate target: 95%.
Δ ≤ 2
Compound-count median
Median absolute peak-count delta vs expert-curated references on the HMDB-style multiplet-line corpus. Strict gate cleared.
20
Fixture regression gate
Every detector change runs against a curated corpus before merge. CI fails by fixture_id when any single fixture drifts > 50%.
39
Evidence layers
Built additively Weeks 22 → 39. Every layer's output is a typed Pydantic model with stable JSON keys.
The closing loop
Spectroscopy isn't the end. It's the start of a loop.
Most analytical platforms stop at "the spectrum has been processed." Ours doesn't. A real worked example with acetic acid impurity:
1
Spectroscopy detects
Impurity peak at 2.10 ppm matched to acetic acid CH₃ within 0.001 ppm. Confidence 93%.
2
Regulatory routes
Acetic acid is ICH Q3C Class 3 — no need for action below 5000 ppm. Action item raised only if observed concentration crosses threshold.
3
Reaction optimization constrains
Next experiment's reaction recipe receives the impurity limit as a Bayesian prior. Solvent + workup updated automatically.
4
Loop closes
Re-acquired spectrum confirms impurity below threshold. Audit ledger records every step from FID hash to recipe update.
Designed against the regulations you'll be audited on.
SpectraCheck's controls are designed to support compliance — by design, not by accident. The audit ledger, the immutable raw vault, the recipe-hash provenance, and the human-signoff release gate were designed to support ICH Q2(R2) ALCOA+, the FDA's January 2025 AI framework, and the EMA reflection paper from day one.
Immutable raw vault
Every FID is SHA-256 hashed, vault path policy enforced, and never overwritten.
Recipe-hash provenance
Every processing run links a recipe hash to the unchanged raw archive. Designed for bit-identical replay.
Human signoff queue
No regulatory document is released without an explicit qualified-human attribution.
ALCOA+ audit ledger
Attributable · Legible · Contemporaneous · Original · Accurate · Complete · Consistent · Enduring · Available.
Cross-modal contradiction warnings
HRMS exact mass disagreeing with NMR-implied formula raises a first-class warning before signoff.
Tenant isolation by default
Controls designed to support SOC 2 Type II, data-residency controls designed to support GDPR, role-scoped audit-event ledger.
See it on your own spectra.
Open SpectraCheck on the platform or schedule a 30-minute walkthrough on a real analyte from your workflow.
What NMR and MS file formats and instrument vendors does SpectraCheck support?
SpectraCheck ingests raw FID archives from Bruker and Agilent-Varian (parsed via the open-source nmrglue library), plus JCAMP-DX, CSV, and vendor exports for NMR, and mzML, mzXML, and processed peak lists for LC-MS, HRMS, and MS/MS.
Can SpectraCheck do automated structure elucidation from NMR?
It ranks candidate SMILES across a multi-layer evidence stack (1H/13C, 2D NMR, HRMS, MS/MS, predicted shifts, fragmentation trees) and runs an automated structure-verification layer, but the platform frames every output as decision support, never a proof of identity.
How does SpectraCheck score confidence, and is it a calibrated DP4 probability?
It aggregates cross-modal evidence into a ranked candidate list and exposes a DP4/DP5 panel, but MolTrace states the result is decision support — not proof of identity, and not a calibrated DP4/DP5 probability. A human reviewer weighs it before signoff.
How does SpectraCheck keep spectroscopy results auditable for a regulatory submission?
Every numerical claim links back to its source spectrum; raw FIDs are SHA-256 hashed in an immutable vault, processing is recipe-hash-linked for reproducible replay, and no report releases without human signoff recorded in an ALCOA+-aligned ledger — controls designed to support 21 CFR Part 11 and ICH Q2(R2), not a compliance claim.
How does SpectraCheck handle solvent and impurity peaks?
Each peak is auto-classified as compound, solvent, impurity, artifact, or 13C satellite against curated Fulmer/Gottlieb residual-solvent and impurity-shift tables, so non-analyte signals are masked from candidate scoring and flagged for the reviewer.
Which modalities can SpectraCheck combine in one analysis?
NMR (1H/13C, with 2D COSY/HSQC/HMQC/HMBC behind a feature flag), HRMS, MS/MS, and LC-MS features fuse into one evidence stack, and cross-modal contradictions — such as an HRMS exact mass disagreeing with the NMR-implied formula — surface as first-class warnings.