feat(olink): data-driven per-disease effect-size catalog - #3
Closed
bschilder wants to merge 3 commits into
Closed
Conversation
Adds synthlab/olink.py, a small greenfield simulator for Olink-style
proteomics data (subject x protein -> NPX). As of 2026-04 there is no
widely-used open-source Olink simulator — the closest analogue,
MSstatsSampleSize, targets LC-MS/MS peptide intensities rather than
NPX/PEA, and OlinkAnalyze ships demo tables (npx_data1/npx_data2) but
no simulator.
Model: NPX[i,j] = mean[j] + plate_eff[i] + group_shift[i,j] + eps
- plate_eff ~ N(0, plate_effect_sd^2), 96 samples / plate
- group_shift from configurable per-protein {group -> delta} map
- eps ~ N(0, sd[j]^2)
Missingness
- mnar_lod: drops 80% of sub-LOD values (soft LOD, mirrors real Olink
behaviour where some sub-LOD values are still reported with
QC_WARN flags)
- mcar: uniform missing_rate drop on top
- mar: currently aliases mcar; full MAR deferred to follow-up
- none: skips missingness
Priors come from aggregate UKB-PPP statistics (Sun et al. 2023,
Nature — 2,923 proteins x ~54k participants) and OlinkAnalyze
npx_data1/npx_data2 demo tables: mean ~ 5 log2-units, sd ~ 0.6,
LOD ~ mean - 2*sd.
Ships a 50-protein preset (default_explore_3072_panel) subsetted from
UKB-PPP Explore 3072, plus parquet read/write helpers. Full NumPy-
style docstrings on every public function / class / method. 16 unit
tests cover determinism, group-effects mean shift, LOD drop,
MCAR overlay, plate variation, QC rate, parquet round-trip, schema
types, validation errors, and empty-frame edge cases.
Scope-deferred (future PR): full MAR missingness, multi-plate batch
effects beyond a flat per-plate intercept, realistic PEA dilution
noise model, and panel-version LOD bridging (Explore 3072 ↔ HT).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Adds notebooks/olink_demo.ipynb — a 24-cell walkthrough of synthlab.olink.simulate_olink_npx covering NPX distributions, LOD-driven missingness, plate batch effects, PCA/UMAP, a protein-protein correlation heatmap, a case-vs-control volcano plot, and a parquet round-trip. Runs end-to-end in ~20s on CPU with outputs embedded. Adds a notebooks/README.md index and a reproducible builder script (notebooks/_build_olink_demo.py). pyproject.toml picks up a new [viz] optional-dependency group (matplotlib, seaborn, scikit-learn, umap-learn) — not a hard dep; notebook users install via `pip install synthlab[viz]`.
Ships synthlab/data/olink_disease_effects.csv — a curated, source-cited per-disease protein effect-size catalog covering 7 diseases × 43 rows (Alzheimer, BRCA_hereditary, CAD, CKD, Cancer_broad, IBD, T2D). Every row cites a real DOI (Sun et al. 2023 UKB-PPP, Williams et al. 2022 Sci Transl Med, Eldjarn et al. 2023 deCODE, Cohen et al. 2018 CancerSEEK, Dubin et al. 2023 Nat Comm CRIC, Guo et al. 2024 Nat Aging, Hu et al. 2025 Nat Comm UKB-PPP, Ahn et al. 2021 Cancers). Public API in synthlab.olink: - DiseaseEffectCatalog (frozen dataclass wrapping the loaded frame with .diseases(), .proteins_for(), .effects_for(noise_sd, seed)) - load_disease_effect_catalog(path=None) — importlib.resources-based ship-in-wheel loader Both re-exported from synthlab/__init__.py. Tests (tests/test_olink_disease_catalog.py, 15 tests): - schema, minimum disease coverage, DOI shape, defensible magnitudes - effects_for shape, KeyError paths, noise determinism / perturbation - end-to-end round-trip: catalog → effects_for → simulate_olink_npx, assert per-disease mean NPX shift matches catalog within 3 SE at 500 samples/arm × 44 rows Demo notebook (notebooks/olink_demo.ipynb) gains Section 10 — disease-conditional generation — with loading, inspection, 3-group cohort simulation (T2D/CAD/baseline, 300 samples each), top-5 grouped-bar chart, and a catalog-vs-empirical scatter. Rebuilt via notebooks/_build_olink_demo.py; executed outputs embedded. Cell count goes from 24 → 34. Docs: - README.md gains a "Disease-conditional effect catalog" subsection with per-disease row-range citations. - docs/olink_disease_catalog.md — full schema, unit-conversion guidance, SE guidance, evidence-strength rubric, PR checklist for adding a new disease. Ship-in-wheel plumbing: - synthlab/data/__init__.py + data/*.csv in [tool.setuptools.package-data] - .gitignore carve-out so synthlab/data/*.csv are tracked while other data/ and *.csv remain ignored. Stacks on PR #2 (feat/olink-npx-simulator). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Owner
Author
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Ships a curated, source-cited per-disease Olink effect-size catalog so users can simulate realistic disease cohorts with
simulate_olink_npxwithout hand-coding effect sizes. Stacks on #2.Coverage (43 rows × 7 diseases in
synthlab/data/olink_disease_effects.csv)Every row cites a real DOI; effect sizes are log2 NPX units (Olink's native scale); SE defaults to 0.3 when only one source is used.
Sourcing methodology
docs/olink_disease_catalog.md(fold change → log2; SDs → multiplied by baseline NPX sigma).|delta_npx|in the catalog is 1.20 (well within the observed Olink Explore range; sepsis CRP would peak at +3 to +4). A test enforces|delta| <= 3.0.strong(replicated / MR-supported / clinical),moderate(single-N-large cohort),weak(small-N or null-hypothesis placeholder). Seedocs/olink_disease_catalog.md.metacolumn.What shipped
synthlab/data/olink_disease_effects.csvdisease,protein_uniprot,delta_npx,se_delta,source,doi,evidence_strength,meta). UTF-8, no trailing whitespace.synthlab/data/__init__.pysynthlab.dataas a package soimportlib.resources.files("synthlab.data")works.synthlab/olink.pyDiseaseEffectCatalog(frozen dataclass) +load_disease_effect_catalog(importlib-resources-based loader). Full NumPy-style docstrings with doctest-style examples.synthlab/__init__.pytests/test_olink_disease_catalog.pyeffects_forshape + errors,noise_sddeterminism, end-to-end round-trip vssimulate_olink_npx, user-supplied CSV path, error paths.notebooks/_build_olink_demo.py+notebooks/olink_demo.ipynbREADME.mddocs/olink_disease_catalog.mdpyproject.tomlpackage-data = ["py.typed", "data/*.csv"]— ships the CSV in the wheel..gitignoresynthlab/data/*.csv+synthlab/data/*.pyso they're tracked while otherdata/and*.csvremain ignored.Verification
Test plan
pytest tests/test_olink_disease_catalog.py -xvs— 15/15 passpytest tests/test_olink.py -q— 16/16 pass (no PR feat(olink): greenfield NPX-style proteomics simulator #2 regression)load_disease_effect_catalog()→effects_for(["T2D","CAD"])→OlinkSimConfig→simulate_olink_npx(500 per arm)→ per-row empirical mean shift within 3 SE of catalogdelta_npximportlib.resources.files("synthlab.data").joinpath("olink_disease_effects.csv")resolves the shipped CSVDeferred follow-ups
time_to_dx_years → delta_npxcurves for incidence-cohort simulations.docs/olink_disease_catalog.md.Stacks on #2 via
--base feat/olink-npx-simulator.🤖 Generated with Claude Code