Longitudinal Baseline
A set of molecular measurements taken on one person at known timepoints under controlled conditions, used as the personal reference against which every later measurement is compared.
A longitudinal baseline is your own molecular data measured at more than one timepoint, stored with enough metadata that a later measurement can be compared to the earlier ones rather than to a population reference interval. The genome is measured once. Everything else (transcript abundance, protein levels, blood chemistry, glucose) is a time series, and a single draw is one sample from a distribution you have not characterized.
How it works
The useful unit is the per-analyte within-person distribution. For an analyte measured k times under matched conditions, you estimate a personal mean and a within-person coefficient of variation (CV_i), then separately estimate analytical CV (CV_a) from the assay’s replicate data. The reference change value, the delta that exceeds noise, is roughly 2.77 × √(CV_a² + CV_i²) for a two-sided 95% criterion. For a protein with CV_a of 5% and CV_i of 8%, that is about a 26% change before you should treat a move as real. For hs-CRP, where CV_i often runs above 40%, a doubling is inside noise.
Population reference intervals answer a different question. They are 95% ranges over thousands of people, so the ratio CV_i / CV_g (between-person variation) determines whether they mean anything for you. When that ratio is below ~0.6, an individual can drift from their own setpoint to the far edge of their personal range while staying comfortably inside the population interval. That is the whole argument for baselines.
The clinical evidence for repeated measurement is strongest where it has been tested directly. In chronic heart failure, serially sampled panels of 92 circulating proteins produced dynamic, patient-specific risk estimates that a single baseline draw did not 1. In preclinical Alzheimer disease, CSF markers measured repeatedly in middle-aged, cognitively normal adults showed within-person trajectories that separated on a timescale of years, well before symptoms 2. Amyloid and tau measured in unimpaired older adults predicted subsequent cognitive and functional decline over a multi-year follow-up 3. The signal lives in slopes.
In your own data
Where each layer lands:
- Genome:
sample.vcf.gzplussample.cram. Measured once. Keep the CRAM, not just the VCF, so you can re-call against a newer reference or a new caller version. Re-interpretation of a fixed genome over time is the main way genomic data ages well, and also where most of its failure modes live 4. - RNA-seq: per-timepoint
quant.sf(salmon) or a counts matrix. Never compare raw TPMs across batches. Build oneDESeqDataSetwith timepoint and batch as columns, letDESeq2compute size factors across all samples together, and look atvst()output for drift. - Proteomics: Olink NPX values are log2 and plate-normalized. Extract the bridge-sample columns and check plate medians before you interpret any delta. Aptamer RFUs from SomaScan are not comparable to NPX for the same protein.
- Clinical chemistry: keep the units string and the lab’s instrument ID in the same row as the value. A lab switching platforms will move your ferritin by more than a year of real biology.
- CGM: 5-minute or 15-minute interval CSVs with epoch timestamps. Store raw interstitial values, not the app’s daily summary. Compute your own metrics (mean, CV, time in a self-chosen band) so the definition stays constant when the vendor changes theirs.
Four mistakes we see repeatedly. First, no condition metadata: fasting duration, time of day, days since last hard training session, and sleep the night before all move analytes, and sleep specifically tracks Alzheimer-relevant markers in cognitively unimpaired adults 5. Second, storing processed values only, so a reprocessed batch cannot be compared to the old one. Third, changing labs between timepoints, which confounds assay change with biology. Fourth, testing every analyte for change at each timepoint without correcting for multiple comparisons, then chasing the tail.
Store everything as one long-format table: timepoint, date, assay_platform, analyte, value, unit, lot, condition_flags. Version it. Keep raw files immutable.
Limitations
Two or three timepoints do not give you a slope you should trust. Estimating CV_i to ±20% precision typically needs on the order of 5 to 10 matched samples per analyte. Most molecular markers have no validated personal decision threshold, so a real change tells you something moved, not what it means or what to do. Genomic markers in particular have a long history of promising associations that did not survive prospective validation, including in areas like DNA repair deficiency and periodontal risk where the biology is well described 67. Interpretation of any deviation, and any decision that follows, belongs with a clinician who can see the full picture.
Woolf Software builds longitudinal molecular profiles of individuals: whole-genome sequencing, RNA sequencing, proteomics, blood biomarkers, and continuous glucose data, integrated into one model of you. Build your profile.
Once you hold the raw files, the baseline is a versioned dataset with per-analyte within-person variance, not a row of numbers next to a population range. Your analysis question changes from 'is this value normal?' to 'has this value moved more than my own noise?'
Related Terms
References
- Dominika Klimczak-Tomaniak, Marie de Bakker, Elke Bouwens, et al.. Dynamic personalized risk prediction in chronic heart failure patients: a longitudinal, clinical investigation of 92 biomarkers (Bio-SHiFT study) . Scientific Reports, 2022. DOI
- Courtney L. Sutphen, Mateusz S. Jasielec, Aarti R. Shah, et al.. Longitudinal Cerebrospinal Fluid Biomarker Changes in Preclinical Alzheimer Disease During Middle Age . JAMA Neurology, 2015. DOI
- Reisa A. Sperling, M.C. Donohue, R.A. Rissman, et al.. Amyloid and Tau Prediction of Cognitive and Functional Decline in Unimpaired Older Individuals: Longitudinal Data from the A4 and LEARN Studies . The Journal of Prevention of Alzheimer's Disease, 2024. DOI
- Jay Shendure, Gregory M. Findlay, Matthew W. Snyder. Genomic Medicine–Progress, Pitfalls, and Promise . Cell, 2019. DOI
- Jonathan Blackman, Laura Stankeviciute, Eider M Arenaza-Urquijo, et al.. Cross-sectional and longitudinal association of sleep and Alzheimer biomarkers in cognitively unimpaired adults . Brain Communications, 2022. DOI
- Ksenija Nesic, Matthew Wakefield, Olga Kondrashova, et al.. Targeting DNA repair: the genome as a potential biomarker . The Journal of Pathology, 2018. DOI
- Luigi Nibali, Kimon Divaris, Emily Ming‐Chieh Lu. The promise and challenges of genomics‐informed periodontal disease diagnoses . Periodontology 2000, 2024. DOI