PGS Catalog
An open EMBL-EBI database of published polygenic scores, distributing the per-variant weights and the metadata needed to apply and evaluate each score correctly.
The PGS Catalog is an open database of published polygenic scores, hosted at EMBL-EBI, that distributes each score as a downloadable table of variants and effect weights alongside the metadata needed to apply and evaluate it. Each entry gets a stable accession (PGS000001, PGS004696, and so on), a linked publication, the training GWAS, the development and evaluation cohorts with their ancestry composition, and reported performance metrics.
A polygenic score is a weighted sum. For each variant in the score, you count how many effect alleles you carry (0, 1, or 2), multiply by the published effect weight, and add it up across every variant in the file. The output is a single number. On its own it means nothing. It becomes interpretable only after you place it in a distribution of scores from a reference population, which is where most of the real work lives.
How it works
The Catalog serves two file types. The raw scoring file is what the authors published, on whatever genome build and with whatever variant identifiers they used. The harmonized files (_hmPOS_GRCh37 and _hmPOS_GRCh38) have been remapped to consistent chromosomal positions, with effect alleles checked against the reference genome. Use the harmonized files. Always.
Score sizes vary by three orders of magnitude. Some coronary artery disease scores carry six or seven variants chosen by a clinical panel. Others carry 1.1 million or 6.6 million variants from LDpred2 or a similar Bayesian shrinkage method applied to a full GWAS. Large scores usually predict better but demand imputed genotypes, because you will never observe six million variants directly on an array and even a 30x whole genome will have gaps.
The reference implementation is pgsc_calc, an nf-core Nextflow pipeline from the Catalog team. It pulls scores by accession, harmonizes them against your target variants, calls PLINK2 --score with cols=+scoresums,+denom, and optionally projects your samples onto 1000 Genomes principal components so it can report an ancestry-adjusted percentile rather than a raw sum. The typical invocation:
nextflow run pgscatalog/pgsc_calc \
--input samplesheet.csv --target_build GRCh38 \
--pgs_id PGS000018,PGS000337 \
--run_ancestry pgsc_HGDP+1kGP_v1.tar.zst
Older tools like PRSice do clumping and p-value thresholding from raw GWAS summary statistics, which is a different job 1. If you are applying a published score, do not re-derive it.
In your own data
Start with your joint-called VCF or the per-sample gVCF from whole-genome sequencing. Three checks before you trust any number.
Build. A GRCh37 score applied to GRCh38 coordinates will silently match a few percent of variants by chance and produce a plausible-looking wrong answer. Confirm your build from the VCF header contig lengths (chr1 is 249,250,621 in GRCh37 and 248,956,422 in GRCh38).
Strand and allele orientation. Ambiguous A/T and C/G variants cannot be resolved by allele matching alone. pgsc_calc drops them by default. In Samoan cohorts, harmonization choices alone changed lipid score performance materially, which tells you how much of the final number is downstream of plumbing rather than biology 2.
Match rate. The pipeline reports how many score variants were found in your data. Below roughly 90% overlap for a genome-wide score, treat the result as unreliable and check whether missing genotypes were imputed as the reference allele, which biases scores downward. A practical walkthrough of these steps in UK Biobank covers the same failure modes 3.
Reporting is its own pipeline stage: mapping a raw sum to a percentile, then to an interpretable statement, requires a reference distribution matched to your ancestry 4.
Limitations
Most scores were trained in European-ancestry cohorts and lose predictive power elsewhere, sometimes by half or more. This has been documented repeatedly and directly in Samoan and high-altitude Peruvian cohorts 56. The Catalog reports ancestry composition for each score’s evaluation samples. Read it before using the score.
A percentile is a population statement, not a personal one. The absolute risk difference between the 90th and 50th percentile depends on the baseline prevalence of the trait and on everything else about you. Scores also do not capture rare high-penetrance variants, which have entirely different implications 7. The ACMG position is that most polygenic scores are not currently ready to drive clinical decisions on their own 8, and the broader methodological cautions about calibration, portability, and overinterpretation are well documented 9. Research uses continue to be interesting, including score-by-treatment interaction work 10, but those are studies, not instructions. Anything you would act on belongs in front of a genetic counselor or physician.
Woolf Software builds longitudinal molecular profiles of individuals: whole-genome sequencing, RNA sequencing, proteomics, blood biomarkers, and continuous glucose data, integrated into one model of you. Build your profile.
With your own VCF you can compute any of thousands of published scores yourself using pgsc_calc; the hard part is harmonizing variants to your genome build and interpreting the result against a reference distribution that matches your ancestry.
Related Terms
References
- Jack Euesden, Cathryn M. Lewis, Paul F. O’Reilly. PRSice: Polygenic Risk Score software . Bioinformatics, 2014. DOI
- Toni-Ann J. Yapp, Mohanraj Krishnan, Shuwei Liu, et al.. Variant Harmonization Critically Determines Polygenic Score Transferability for Lipid Traits in Samoan Populations . Human Genetics and Genomics Advances, 2026. DOI
- Jennifer A. Collister, Xiaonan Liu, Lei Clifton. Calculating Polygenic Risk Scores (PRS) in UK Biobank: A Practical Guide for Epidemiologists . Frontiers in Genetics, 2022. DOI
- Pedro Victor Barbosa Araújo, Tayná da Silva Fiúza, José Eduardo Kroll, et al.. From Data Curation to Risk Reporting: A Pipeline for Polygenic Risk Scores . 2026. DOI
- Toni-Ann J. Yapp, Mohanraj Krishnan, Shuwei Liu, et al.. Evaluating Polygenic Score Transferability for Lipid Traits in Underrepresented Populations: Evidence from Samoan Cohorts . 2026. DOI
- Andy Castañeda, Natalie R. Hasbani, Adam S. Heath, et al.. Performance of cardiometabolic polygenic scores in a high-altitude Peruvian population: the CRONICAS cohort . Scientific Reports, 2026. DOI
- Ali Torkamani, Nathan E. Wineinger, Eric J. Topol. The personal and clinical utility of polygenic risk scores . Nature Reviews Genetics, 2018. DOI
- Aya Abu-El-Haija, Honey V. Reddi, Hannah Wand, et al.. The clinical application of polygenic risk scores: A points to consider statement of the American College of Medical Genetics and Genomics (ACMG) . Genetics in Medicine, 2023. DOI
- John Novembre, Catherine Stein, Samira Asgari, et al.. Addressing the challenges of polygenic scores in human genetic research . The American Journal of Human Genetics, 2022. DOI
- Peter D. Fransquet, Chenglong Yu, Cammie Tran, et al.. Triglyceride Polygenic Score Identifies Individuals Who May Respond Differently to Aspirin in Primary Prevention . Clinical Pharmacology & Therapeutics, 2026. DOI