Genome-Wide Association Studies
- NOTOC**
Genome-Wide Association Studies
Genome-wide association studies (GWAS) are large-scale investigations designed to identify statistical associations between genetic variants and human traits or diseases. By scanning hundreds of thousands or millions of variants across the genome, researchers can compare genetic differences among individuals and identify genomic regions associated with characteristics ranging from height and blood pressure to diabetes, cancer, psychiatric disorders, autoimmune disease, sleep patterns, reproductive traits, and molecular phenotypes.
The development of GWAS transformed human genetics by allowing researchers to investigate complex traits without first selecting specific candidate genes. Early studies demonstrated that common diseases and quantitative traits are frequently influenced by large numbers of genetic variants, most of which individually have relatively small effects. As sample sizes increased into the hundreds of thousands and eventually millions of participants, GWAS identified thousands of loci associated with human biological variation and disease susceptibility.
GWAS have also revealed major challenges. Statistical associations do not necessarily identify causal variants or biological mechanisms, many traits are extremely polygenic, environmental influences remain important, and genetic discoveries derived primarily from populations of European ancestry may not transfer accurately to other populations. Modern GWAS therefore increasingly incorporate diverse populations, functional genomics, fine-mapping, transcriptomics, proteomics, metabolomics, electronic health records, and advanced computational methods.
Development and Methodology
Genome-wide association studies became practical as high-throughput genotyping technologies made it possible to measure hundreds of thousands of single-nucleotide polymorphisms (SNPs) throughout the genome. Large reference datasets and genotype-imputation methods subsequently allowed researchers to infer millions of additional genetic variants that had not been directly measured.
A typical GWAS begins with careful definition of the phenotype being studied and collection of genetic data from a sufficiently large population. Researchers perform extensive quality-control procedures before testing each genetic variant for statistical association with the phenotype. Because millions of statistical comparisons may be conducted simultaneously, stringent significance thresholds are required to reduce false-positive findings.
Population structure is another important consideration. Genetic ancestry differences between study participants can create apparent associations unrelated to the trait itself. Principal-components analysis and mixed-model approaches are widely used to control population stratification and relatedness.
Genotype imputation uses reference haplotypes to predict variants that were not directly genotyped. Resources such as the 1000 Genomes Project greatly increased the number and diversity of variants available for imputation and cross-population analysis.
Results from multiple studies are frequently combined through meta-analysis. This approach increases statistical power and has enabled many of the largest discoveries in complex-trait genetics.
From Association to Biological Interpretation
Identifying a statistically significant genomic region is usually the beginning rather than the end of biological investigation. Variants located near one another are frequently inherited together because of linkage disequilibrium, making it difficult to determine which variant within an associated region is biologically responsible for the observed signal.
Statistical fine-mapping attempts to narrow association signals to smaller sets of plausible causal variants. Methods can incorporate linkage disequilibrium, functional genomic annotations, information from different ancestry groups, and other sources of biological evidence.
Many GWAS-associated variants are located outside protein-coding regions. Large functional-genomics projects have shown that disease-associated variants are frequently concentrated in regulatory DNA that influences when and where genes are expressed.
Researchers increasingly integrate GWAS with expression quantitative trait loci (eQTL) studies and transcriptomic data. Approaches such as transcriptome-wide association studies, PrediXcan, SMR, and related methods attempt to connect trait-associated variants with changes in gene expression.
Gene-based and pathway-based methods can aggregate evidence across many variants, helping researchers identify biological processes that may not be obvious from individual SNP associations.
Polygenicity and Genetic Architecture
One of the most important findings from GWAS is that many common human traits are highly polygenic. Rather than being controlled by a small number of genes with large effects, characteristics such as height, body mass index, blood pressure, educational attainment, psychiatric disorders, and many common diseases are influenced by large numbers of variants with individually small effects.
The early observation that genome-wide-significant variants explained only a fraction of estimated heritability became known as the missing heritability problem. Subsequent studies demonstrated that common variants collectively account for considerably more genetic variation than genome-wide-significant variants alone.
Methods such as genome-wide complex trait analysis and LD Score regression made it possible to estimate the aggregate contribution of common variants and distinguish genuine polygenic signals from statistical inflation caused by population structure or other confounding factors.
The concept of an omnigenic model extends polygenic thinking by proposing that highly interconnected regulatory networks may allow variants in many genes expressed within relevant cell types to influence complex traits indirectly.
GWAS have also demonstrated extensive pleiotropy, in which individual genomic regions influence multiple traits. Genetic-correlation analyses can quantify the degree to which diseases and characteristics share underlying genetic influences.
Population Diversity and Ancestry
The ancestry composition of GWAS has become an important scientific and equity issue. Historically, a substantial proportion of genomic discovery datasets were composed of participants of European ancestry. This imbalance limits understanding of human genetic diversity and can reduce the accuracy of genetic prediction in other populations.
Allele frequencies and patterns of linkage disequilibrium differ among populations as a result of human demographic history. Consequently, an association discovered in one population may have a different frequency, statistical strength, or predictive value in another.
Studies involving African, East Asian, Japanese, Latino, and other populations have demonstrated that ancestry-diverse GWAS can discover variants that are uncommon or poorly captured in European datasets. Multi-ancestry analyses can also improve fine-mapping because differences in linkage disequilibrium may help distinguish causal variants from neighboring correlated markers.
Large trans-ancestry studies of type 2 diabetes, cardiovascular disease, cancer, rheumatoid arthritis, lung function, psychiatric disorders, and other traits illustrate the scientific value of expanding genomic research across populations.
Polygenic Risk Scores
Polygenic risk scores combine information from many genetic variants into a single measure intended to summarize inherited predisposition toward a particular trait or disease. Because individual GWAS variants generally have very small effects, combining thousands or millions of variants can provide more information than considering genome-wide-significant loci alone.
Polygenic scores are being investigated for applications involving cardiovascular disease, cancer, diabetes, psychiatric conditions, and other complex disorders. Their clinical usefulness, however, depends on predictive accuracy, appropriate interpretation, and comparison with established environmental and clinical risk factors.
An important limitation is reduced portability between ancestry groups. Scores derived primarily from European-ancestry GWAS frequently perform less accurately in populations that were poorly represented in discovery datasets.
Clinical use therefore raises scientific and ethical questions concerning ancestry, informed consent, communication of uncertainty, interpretation of probabilistic risk, and the possibility that unequal genomic representation could reinforce existing health disparities.
Landmark Human Trait Studies
GWAS have mapped genetic influences on a wide range of human characteristics. Height became one of the best examples of extreme polygenicity. Increasingly large studies identified hundreds and eventually thousands of associated variants, culminating in analyses involving millions of individuals and more than 12,000 independent height-associated variants.
Body mass index studies identified numerous loci connected with metabolic regulation and central nervous system pathways. Lipid studies uncovered variants influencing cholesterol and triglyceride concentrations and helped identify biological pathways relevant to cardiovascular disease.
Researchers have also identified variants associated with facial morphology, pigmentation, puberty timing, birth weight, reproductive behavior, intelligence, educational attainment, blood pressure, lung function, bone mineral density, sleep patterns, chronotype, and many other measurable human characteristics.
These findings demonstrate both the broad influence of inherited variation and the importance of environmental and social factors. For highly environmentally influenced traits, genetic associations describe statistical contributions within particular populations rather than fixed biological outcomes.
Cardiovascular and Metabolic Disease
Cardiovascular genetics has been transformed by large GWAS consortia. Studies have identified numerous loci associated with coronary artery disease, myocardial infarction, atrial fibrillation, blood pressure, venous thromboembolism, heart failure, stroke, thoracic aortic disease, and other cardiovascular conditions.
Blood-pressure GWAS involving more than one million participants identified hundreds of loci and connected genetic variation with vascular, renal, adrenal, and metabolic pathways.
Type 2 diabetes has been another major focus. Studies across European, East Asian, African, and multi-ancestry populations have identified hundreds of susceptibility regions and demonstrated the value of population diversity for discovery and fine-mapping.
GWAS have also identified genetic influences on fasting glucose, insulin regulation, blood lipids, serum urate, liver enzymes, fatty liver disease, and numerous biochemical traits.
These studies increasingly combine association signals with functional genomic evidence to identify genes and pathways that may become candidates for biological investigation or therapeutic development.
Cancer Genetics
Genome-wide association studies have identified inherited susceptibility variants for numerous cancers, including breast, prostate, colorectal, pancreatic, ovarian, lung, and melanoma.
Early cancer GWAS demonstrated that common variants can contribute modestly to inherited cancer susceptibility. Subsequent consortium studies involving much larger populations identified dozens or hundreds of additional loci.
Some associations differ among cancer subtypes or populations. Multi-ancestry studies have therefore become increasingly important for discovering population-specific risk variants and improving understanding of shared genetic mechanisms.
GWAS generally identify inherited susceptibility rather than the acquired somatic mutations that develop inside tumors. The two approaches complement one another by addressing different aspects of cancer biology.
Autoimmune and Inflammatory Disease
Autoimmune diseases have produced some of the clearest examples of how GWAS can reveal biological pathways. Studies of rheumatoid arthritis, Crohn's disease, multiple sclerosis, systemic lupus erythematosus, psoriasis, celiac disease, ankylosing spondylitis, asthma, and atopic dermatitis have identified numerous immune-related loci.
Many associations involve pathways responsible for antigen presentation, cytokine signaling, T-cell function, inflammatory regulation, and innate immunity.
Cross-disease analyses reveal that autoimmune disorders frequently share genetic associations. These overlaps can provide insight into common mechanisms while also identifying variants that distinguish susceptibility to specific disorders.
GWAS of rheumatoid arthritis have additionally demonstrated how genetic association results can contribute to drug discovery by prioritizing genes and pathways with potential therapeutic relevance.
Psychiatric and Behavioral Genetics
Very large GWAS have substantially expanded knowledge of the genetic architecture of psychiatric disorders. Studies have identified associated loci for schizophrenia, major depression, bipolar disorder, autism spectrum disorder, attention-deficit/hyperactivity disorder, anorexia nervosa, post-traumatic stress symptoms, and other conditions.
Schizophrenia provides a prominent example of increasing statistical power. Large international studies expanded the number of associated genomic regions from a small number of loci to hundreds, while implicating synaptic signaling and neuronal biology.
Studies of depression and bipolar disorder similarly reveal extensive polygenicity and genetic overlap with other psychiatric, behavioral, cognitive, and metabolic characteristics.
Behavioral GWAS have also examined smoking, alcohol consumption, reproductive behavior, neuroticism, sleep, and related phenotypes. Because such characteristics are strongly shaped by environmental and social conditions, genetic associations should not be interpreted as deterministic explanations of individual behavior.
Sleep, Cognition, and Brain Imaging
Large biobank datasets have made it possible to investigate genetic contributions to sleep duration, insomnia, daytime sleepiness, chronotype, cognitive ability, and brain structure.
Studies involving hundreds of thousands or more than one million individuals have identified numerous loci associated with insomnia and circadian preference. These findings connect sleep-related variation with neurological, psychiatric, cardiovascular, and metabolic pathways.
Brain-imaging GWAS combine genetic data with magnetic resonance imaging and other quantitative measurements. Researchers have identified genetic influences on regional brain volumes, white-matter microstructure, connectivity, and other imaging-derived phenotypes.
These approaches demonstrate how GWAS can be extended from traditional diseases and traits to highly detailed quantitative measurements of biological structure and function.
Computational Tools and Statistical Methods
GWAS depend heavily on specialized computational and statistical tools. PLINK became one of the most widely used software packages for genotype quality control, association testing, population analysis, and genomic-data management.
METAL provides efficient meta-analysis of association results from multiple studies. GCTA enables estimation of SNP-based heritability and other aspects of complex-trait architecture.
MAGMA converts SNP-level associations into gene and gene-set analyses, while VEGAS provides gene-based testing that accounts for linkage disequilibrium.
Genotype-imputation tools and haplotype-phasing algorithms dramatically expanded the number of variants that could be investigated without directly genotyping every position.
Mixed-model methods such as EMMAX help control relatedness and population structure, while LD Score regression can estimate heritability, genetic correlations, and the degree to which statistical inflation results from polygenicity or confounding.
Molecular Phenotypes and Multi-Omics
GWAS methodology has expanded beyond diseases and visible traits into molecular phenotypes. Proteomic GWAS identify genetic variants that influence concentrations of circulating proteins. These protein quantitative-trait loci can connect disease-associated variants with specific biological molecules and potential therapeutic targets.
Metabolomic association studies examine inherited influences on circulating metabolites and biochemical pathways. Whole-genome sequencing has extended such analyses from common variants to less frequent and rare variants with potentially larger biological effects.
Microbiome GWAS investigate relationships between human genetic variation and the composition of microbial communities. Studies have identified associations involving loci such as LCT and ABO while also demonstrating that environmental influences remain major determinants of microbiome composition.
GWAS have additionally been applied to cytokines, gene expression, imaging phenotypes, and many other quantitative molecular measurements.
The increasing integration of genomics with transcriptomics, proteomics, metabolomics, epigenomics, and other forms of biological data represents an important direction in attempts to move from statistical association toward mechanistic understanding.
Limitations and Challenges
GWAS identify statistical association rather than proving causation. A significant SNP may simply be correlated with the biologically causal variant because nearby variants are inherited together.
Most individual associations have small effects, meaning that even highly significant variants often provide limited predictive information by themselves.
Phenotype definition can substantially affect results. Studies using proxy phenotypes, self-reported traits, electronic health records, or different diagnostic definitions may introduce measurement differences or systematic biases.
Population stratification, relatedness, environmental differences, and technical artifacts can create misleading associations if inadequately controlled.
Ancestry imbalance remains one of the most important limitations of existing genomic databases. Expanding participation across global populations is important both for equitable application and for improving scientific understanding of human genetic variation.
GWAS also require careful interpretation when studying behavioral, cognitive, or socially influenced traits. Genetic associations are population-level statistical relationships and do not imply genetic determinism or establish that environmental and social influences are unimportant.
Clinical Translation and Precision Medicine
One major goal of GWAS is to translate genetic discoveries into improved understanding, prevention, diagnosis, and treatment of disease. Association studies can reveal previously unknown biological pathways and identify genes that may become therapeutic targets.
Genetic findings may also support approaches such as Mendelian randomization, which uses genetic variants as instruments for investigating potential causal relationships between exposures and health outcomes.
Polygenic prediction is being studied as a possible addition to conventional clinical risk assessment. Its value varies substantially between diseases and populations, and reliable translation requires validation in the populations where scores will be used.
Rather than providing deterministic predictions, GWAS-derived information generally contributes probabilistic information that must be interpreted alongside age, environment, lifestyle, family history, clinical measurements, and other risk factors.
The long-term significance of GWAS may therefore extend beyond predicting individual disease. Their greatest contribution may be the systematic mapping of biological pathways underlying complex human traits and diseases.
Conclusion
Genome-wide association studies have become one of the central tools of modern human genetics. From the first large association scans of common diseases to contemporary studies involving millions of participants, GWAS have revealed that most common diseases and quantitative traits arise from highly complex genetic architectures involving many variants of individually small effect.
The field has progressed from simply identifying associated SNPs toward understanding biological mechanisms. Fine-mapping, functional genomics, transcriptomics, proteomics, metabolomics, brain imaging, multi-ancestry analysis, and increasingly sophisticated statistical methods are helping researchers connect association signals with genes, cells, tissues, and biological pathways.
At the same time, GWAS have demonstrated the limitations of genetic prediction and the importance of population diversity. Associations may vary among populations, polygenic scores may transfer poorly across ancestry groups, and environmental influences remain essential components of complex human traits.
The continuing expansion of diverse genomic datasets and integration of multiple forms of biological information may make GWAS increasingly useful for understanding disease mechanisms, identifying potential therapeutic targets, and developing more representative approaches to precision medicine. The fundamental lesson of GWAS is that human biological variation is complex, highly polygenic, and best understood through the combined study of genetics, environment, population history, and biological function.
- TOC**
Foundational Reviews and GWAS Methodology
| Laura Harris et al. | Nature Reviews Genetics | 2024
Genome-wide association testing beyond SNPs. Reviews expanding association analysis beyond standard single-nucleotide polymorphisms to additional forms of genomic variation that may explain biologically important trait differences.
| Emil Uffelmann et al. | Nature Reviews Methods Primers | 2021
Genome-wide association studies. Provides a modern methodological primer covering study design, genotyping, imputation, quality control, association testing, meta-analysis, fine-mapping, functional follow-up, and interpretation.
| Emmanuelle Genin | Human Genetics | 2020
Missing heritability of complex diseases: case solved? Reassesses the missing-heritability problem in light of increasingly large association studies, whole-genome approaches, polygenicity, rare variants, and improved heritability estimation.
| Vivian Tam et al. | Nature Reviews Genetics | 2019
Benefits and limitations of genome-wide association studies. Reviews what GWAS have revealed about complex-trait genetics while examining statistical power, biological interpretation, population representation, reproducibility, and limitations of association-based approaches.
| Daniel J. Schaid, Wenan Chen and Nicholas B. Larson | Nature Reviews Genetics | 2018
From genome-wide associations to candidate causal variants by statistical fine-mapping. Reviews methods for narrowing associated regions to credible sets of potential causal variants using linkage disequilibrium, annotations, and multi-population information.
| Evan A. Boyle, Yang I. Li and Jonathan K. Pritchard | Cell | 2017
An expanded view of complex traits: from polygenic to omnigenic. Proposes the omnigenic model, in which highly interconnected regulatory networks allow variants in many expressed genes to influence complex traits.
| Christian Benner et al. | Bioinformatics | 2016
FINEMAP: efficient variable selection using summary data from genome-wide association studies. Introduces a computationally efficient Bayesian method for identifying probable causal variants within genomic regions highlighted by GWAS.
| Alexander Gusev et al. | Nature Genetics | 2016
Integrative approaches for large-scale transcriptome-wide association studies. Develops approaches for combining GWAS and gene-expression prediction to identify genes whose genetically regulated expression is associated with complex traits.
| Zhihong Zhu et al. | Nature Genetics | 2016
Integration of summary data from GWAS and eQTL studies predicts complex trait gene targets. Introduces SMR, integrating GWAS and expression-QTL results to prioritize genes whose expression may help explain trait-associated loci.
| Brendan K. Bulik-Sullivan et al. | Nature Genetics | 2015
LD Score regression distinguishes confounding from polygenicity in genome-wide association studies. Introduces LD Score regression to separate inflation caused by widespread genuine polygenic effects from inflation caused by population stratification and other confounders.
| Hilary K. Finucane et al. | Nature Genetics | 2015
Partitioning heritability by functional annotation using genome-wide association summary statistics. Introduces stratified LD Score regression for identifying genomic annotations and cell types disproportionately contributing to complex-trait heritability.
| Brendan Bulik-Sullivan et al. | Nature Genetics | 2015
An atlas of genetic correlations across human diseases and traits. Uses GWAS summary statistics to quantify shared genetic influences among many diseases and traits, demonstrating widespread pleiotropy and correlated genetic architecture.
| Eric R. Gamazon et al. | Nature Genetics | 2015
A gene-based association method for mapping traits using reference transcriptome data. Introduces the PrediXcan framework for testing associations between genetically predicted gene expression and complex human phenotypes.
| Evangelos Evangelou and John P. A. Ioannidis | Nature Reviews Genetics | 2013
Meta-analysis methods for genome-wide association studies and beyond. Reviews techniques for combining association results across studies while addressing heterogeneity, population differences, publication bias, and increasingly complex genomic data.
| Peter M. Visscher et al. | American Journal of Human Genetics | 2012
Five Years of GWAS Discovery. Assesses the first five years of GWAS discoveries and explains what association studies revealed about polygenicity, heritability, disease biology, and quantitative traits.
| Geraldine M. Clarke et al. | Nature Protocols | 2011
Basic statistical analysis in genetic case-control studies. Presents practical procedures for quality control, association testing, population stratification assessment, and interpretation of genotype data in case-control association studies.
| Noah A. Rosenberg et al. | Nature Reviews Genetics | 2010
Genome-wide association studies in diverse populations. Explores population structure, allele-frequency variation, linkage disequilibrium, admixture, and the scientific value of extending GWAS beyond European-ancestry populations.
| Jonathan Marchini and Bryan Howie | Nature Reviews Genetics | 2010
Genotype imputation for genome-wide association studies. Describes statistical imputation of untyped variants using reference haplotypes and explains how imputation greatly expanded the genomic coverage and comparability of GWAS.
| Evan E. Eichler et al. | Nature Reviews Genetics | 2010
Missing heritability and strategies for finding the underlying causes of complex disease. Reviews technological and analytical strategies for identifying genetic influences not captured by early common-variant GWAS.
| John P. A. Ioannidis et al. | Nature Reviews Genetics | 2009
Validating, augmenting and refining genome-wide association signals. Explains why initial GWAS findings require replication, fine-mapping, functional investigation, larger samples, and careful assessment of statistical and biological significance.
| Teri A. Manolio et al. | Nature | 2009
Finding the missing heritability of complex diseases. Discusses why discovered loci explain only part of familial heritability and evaluates possible contributions from rare variants, structural variation, interactions, and imperfect measurements.
| Mark I. McCarthy et al. | Nature Reviews Genetics | 2008
Genome-wide association studies for complex traits: consensus, uncertainty and challenges. Reviews lessons from the first major wave of GWAS and discusses replication, effect sizes, causal interpretation, and unexplained heritability.
| Alkes L. Price et al. | Nature Genetics | 2006
Principal components analysis corrects for stratification in genome-wide association studies. Demonstrates a widely adopted principal-components approach for controlling ancestry-related population structure that could otherwise generate spurious genetic associations.
| Joel N. Hirschhorn and Mark J. Daly | Nature Reviews Genetics | 2005
Genome-wide association studies for common diseases and complex traits. An early overview explaining how dense genome-wide markers could systematically identify common variants contributing to complex human diseases and traits.
| William Y. S. Wang et al. | Nature Reviews Genetics | 2005
Genome-wide association studies: theoretical and practical concerns. Examines early GWAS challenges including statistical power, multiple testing, linkage disequilibrium, population structure, marker density, replication, and study design.
Functional Interpretation, Genetic Architecture, and Equity
| Boran Gao et al. | Nature Genetics | 2024
MESuSiE enables scalable and powerful multi-ancestry fine-mapping of causal variants in GWAS. Introduces a fine-mapping framework that combines information across ancestry groups while allowing ancestry-specific genetic effects and linkage disequilibrium.
| Yuchang Wu et al. | Nature Genetics | 2024
Pervasive biases in proxy genome-wide association studies based on parental history of Alzheimer's disease. Shows how proxy phenotypes based on parental disease history can introduce systematic biases into GWAS effect estimates and biological interpretation.
| Julia Sidorenko et al. | Nature Genetics | 2024
Genetic architecture reconciles linkage and association studies of complex traits. Examines how complex genetic architecture can explain apparent differences between findings from family linkage studies and population association studies.
| GTEx Consortium | Science | 2020
The GTEx Consortium atlas of genetic regulatory effects across human tissues. Maps genetic influences on gene expression across tissues, providing a key resource for interpreting how noncoding GWAS variants may affect biological function.
| Nicholas Mancuso et al. | Nature Genetics | 2019
Probabilistic fine-mapping of transcriptome-wide association studies. Develops FOCUS to distinguish potentially causal genes from correlated transcriptome-wide association signals generated by linkage disequilibrium and shared expression-prediction models.
| Kyoko Watanabe et al. | Nature Genetics | 2019
A global overview of pleiotropy and genetic architecture in complex traits. Systematically examines shared genetic associations across hundreds of traits, revealing extensive pleiotropy and recurring genomic regions influencing multiple phenotypes.
| Alicia R. Martin et al. | Nature Genetics | 2019
Clinical use of current polygenic risk scores may exacerbate health disparities. Demonstrates that polygenic scores generally transfer less accurately to non-European populations because of ancestry imbalance in GWAS discovery datasets.
| Latrice G. Landry et al. | Health Affairs | 2018
Lack of diversity in genomic databases is a barrier to translating precision medicine research into practice. Documents the ancestry imbalance of genomic research and explains how limited diversity can reduce the equitable clinical usefulness of genetic discoveries.
| Alicia R. Martin et al. | American Journal of Human Genetics | 2017
Human demographic history impacts genetic risk prediction across diverse populations. Shows how allele frequencies and linkage disequilibrium shaped by demographic history affect the portability of genetic risk prediction between populations.
| 1000 Genomes Project Consortium | Nature | 2015
A global reference for human genetic variation. Expands the 1000 Genomes reference across diverse populations and millions of variants, strengthening genotype imputation and cross-population association research.
| Matthew T. Maurano et al. | Science | 2012
Systematic localization of common disease-associated variation in regulatory DNA. Shows that many GWAS variants occur within regulatory DNA and often map to cell types biologically relevant to the associated disease.
| 1000 Genomes Project Consortium | Nature | 2012
An integrated map of genetic variation from 1,092 human genomes. Provides a global catalogue of common and lower-frequency genetic variants that became a major reference resource for GWAS imputation and population-genetic analysis.
| Daniel G. MacArthur et al. | Science | 2012
A systematic survey of loss-of-function variants in human protein-coding genes. Catalogues protein-disrupting variants and demonstrates how functional annotation can help distinguish biologically important variants from the large background of genomic variation.
| Jian Yang et al. | Nature Genetics | 2011
Genome partitioning of genetic variation for complex traits using common SNPs. Develops methods for estimating how different chromosomes and genomic regions collectively contribute to heritable variation in complex traits.
| Jian Yang et al. | Nature Genetics | 2010
Common SNPs explain a large proportion of the heritability for human height. Demonstrates that common SNPs collectively explain substantially more height variation than genome-wide-significant loci alone, helping clarify the missing-heritability problem.
Diversity, Polygenic Risk, and Cross-Ancestry GWAS
| Eleni Friligkou et al. | Nature Genetics | 2024
Multi-ancestry genome-wide association study of anxiety. Uses large, ancestrally diverse datasets to expand the genetic architecture of anxiety and investigate shared genetic relationships with psychiatric and behavioral traits.
| Yunfeng Ruan et al. | Nature Genetics | 2022
Improving polygenic prediction in ancestrally diverse populations. Develops methods for combining GWAS information across populations to improve polygenic prediction, particularly where non-European discovery datasets remain comparatively small.
| Abram B. Kamiza et al. | Nature Medicine | 2022
Transferability of genetic risk scores in African populations. Tests polygenic-score performance in African populations and highlights the substantial challenges created by differences in ancestry, linkage disequilibrium, and discovery-sample representation.
| Anna C. F. Lewis and Robert C. Green | Genome Medicine | 2021
Polygenic risk scores in the clinic: new perspectives needed on familiar ethical issues. Reviews clinical and ethical questions surrounding polygenic scores, including ancestry bias, informed consent, communication, uncertainty, and potential discrimination.
| Cassandra N. Spracklen et al. | Nature | 2020
Identification of type 2 diabetes loci in 433,540 East Asian individuals. Large East Asian GWAS identifies numerous type 2 diabetes associations, including loci and signals poorly captured by studies dominated by European ancestry.
| Masato Akiyama et al. | Nature Genetics | 2017
Genome-wide association study identifies 112 new loci for body mass index in the Japanese population. Uses Japanese and trans-ancestry data to substantially expand known BMI loci and investigate pathways involved in body-weight regulation.
| Anubha Mahajan et al. | Nature Genetics | 2014
Genome-wide trans-ancestry meta-analysis provides insight into the genetic architecture of type 2 diabetes susceptibility. Combines multiple ancestry groups to identify diabetes loci, compare effect directions, and refine association signals.
| Yoon Shin Cho et al. | Nature Genetics | 2012
Meta-analysis of genome-wide association studies identifies eight new loci for type 2 diabetes in East Asians. Demonstrates the discovery value of ancestry-specific meta-analysis and identifies previously unknown diabetes susceptibility loci.
| Yukinori Okada et al. | PLOS Genetics | 2012
Meta-analysis identifies an association between the AFF1 locus and systemic lupus erythematosus in Japanese populations. Expands autoimmune GWAS evidence using Japanese cohorts and illustrates population-specific discovery and replication.
| Ryo Takata et al. | Nature Genetics | 2010
Genome-wide association study identifies five new susceptibility loci for prostate cancer in the Japanese population. Demonstrates that GWAS in Japanese participants can uncover prostate-cancer risk variants not readily identified in European cohorts.
Landmark Trait and Cardiometabolic GWAS
| Sridharan Raghavan et al. | PLOS Genetics | 2022
A multi-population phenome-wide association study of genetically predicted height. Uses GWAS-based genetic prediction across populations to investigate relationships between height-associated variation and a broad range of health outcomes.
| James J. Lee et al. | Nature Genetics | 2018
Gene discovery and polygenic prediction from a genome-wide association study of educational attainment in 1.1 million individuals. Identifies more than a thousand associated variants while illustrating both the power and limitations of highly polygenic prediction.
| Felix R. Day et al. | Nature Genetics | 2017
Genomic analyses identify hundreds of variants associated with age at menarche and support a role for puberty timing in cancer risk. Expands puberty-timing loci and connects their genetic architecture with adult disease susceptibility.
| Suzanne Sniekers et al. | Nature Genetics | 2017
Genome-wide association meta-analysis of 78,308 individuals identifies new loci and genes influencing human intelligence. Identifies genome-wide-significant loci and implicates genes involved in brain development and neuronal function.
| Aysu Okbay et al. | Nature | 2016
Genome-wide association study identifies 74 loci associated with educational attainment. Demonstrates the increasingly large samples required for socially and environmentally influenced highly polygenic traits while highlighting neuronal and developmental pathways.
| Christian Fuchsberger et al. | Nature | 2016
The genetic architecture of type 2 diabetes. Combines large-scale genotyping and sequencing to examine common and rare variants and clarify their relative contributions to diabetes susceptibility.
| Adam E. Locke et al. | Nature | 2015
Genetic studies of body mass index yield new insights for obesity biology. Identifies 97 BMI-associated loci and emphasizes the importance of central nervous system pathways in genetic susceptibility to obesity.
| Andrew R. Wood et al. | Nature Genetics | 2014
Defining the role of common variation in the genomic and biological architecture of adult human height. Identifies hundreds of height variants and shows that common variation collectively accounts for a large proportion of height heritability.
| Cristen J. Willer et al. | Nature Genetics | 2013
Discovery and refinement of loci associated with lipid levels. Uses very large samples to identify and refine lipid-associated loci and connect genetic signals with biological pathways and potential therapeutic targets.
| Sonja I. Berndt et al. | Nature Genetics | 2013
Genome-wide meta-analysis identifies 11 new loci for anthropometric traits and provides insights into genetic architecture. Uses extremes and population-wide measures of height, obesity, and body composition to study shared genetic structure.
| Joshua C. Randall et al. | PLOS Genetics | 2013
Sex-stratified genome-wide association studies including 270,000 individuals show sexual dimorphism in genetic loci for anthropometric traits. Identifies variants whose effects on body size differ between women and men.
| Fan Liu et al. | PLOS Genetics | 2012
A genome-wide association study identifies five loci influencing facial morphology in Europeans. Demonstrates that normal variation in human facial features can be mapped through genome-wide association approaches.
| Yingleong Chan et al. | PLOS Genetics | 2011
Genome-wide analysis of body height extremes identifies common variants with unusually large effects. Tests individuals at the extremes of height and examines how polygenic and individual variants contribute to extreme stature.
| Heribert Schunkert et al. | Nature Genetics | 2011
Large-scale association analysis identifies 13 new susceptibility loci for coronary artery disease. Expands the number of CAD-associated regions and shows that many operate independently of traditional cardiovascular risk factors.
| Tanya M. Teslovich et al. | Nature | 2010
Biological, clinical and population relevance of 95 loci for blood lipids. Large meta-analysis greatly expands lipid-associated loci and links GWAS signals with pathways controlling cholesterol and triglyceride metabolism.
| Hana Lango Allen et al. | Nature | 2010
Hundreds of variants clustered in genomic loci and biological pathways affect human height. Demonstrates extreme polygenicity of height and shows that many common variants of small effect collectively shape normal human stature.
| Elizabeth K. Speliotes et al. | Nature Genetics | 2010
Association analyses of 249,796 individuals reveal 18 new loci associated with body mass index. Major early BMI meta-analysis demonstrates the highly polygenic architecture of obesity and implicates neural and metabolic pathways.
| Josee Dupuis et al. | Nature Genetics | 2010
New genetic loci implicated in fasting glucose homeostasis and their impact on type 2 diabetes risk. Identifies common variants influencing fasting glucose and provides insight into pancreatic beta-cell and glucose-regulatory biology.
| Chunyan He et al. | Nature Genetics | 2009
Genome-wide association study identifies variants at four loci associated with age at menarche. Identifies common genetic influences on pubertal timing and establishes a foundation for later large-scale reproductive-trait GWAS.
| Nicole Soranzo et al. | PLOS Genetics | 2009
Meta-analysis of genome-wide scans for human adult stature identifies novel loci and associations with measures of skeletal frame size. Expands early height GWAS and connects associated loci with skeletal growth.
| Daniel Levy et al. | Nature Genetics | 2009
Genome-wide association study of blood pressure and hypertension. Identifies loci influencing systolic and diastolic blood pressure and demonstrates the value of consortium-scale meta-analysis for cardiovascular traits.
| Jiali Han et al. | PLOS Genetics | 2008
A genome-wide association study identifies novel alleles associated with hair color and skin pigmentation. Identifies common variants contributing to pigmentation traits and illustrates GWAS application to readily measurable human phenotypes.
| Wellcome Trust Case Control Consortium | Nature | 2007
Genome-wide association study of 14,000 cases of seven common diseases and 3,000 shared controls. Landmark study demonstrating the feasibility and discovery power of large-scale GWAS across multiple common diseases using shared controls.
| Laura J. Scott et al. | Science | 2007
A genome-wide association study of type 2 diabetes in Finns detects multiple susceptibility variants. One of the landmark early diabetes GWAS, identifying reproducible common variants contributing to type 2 diabetes risk.
Disease GWAS and Clinical Translation
| Matthew R. Lincoln et al. | Nature Genetics | 2024
Genetic mapping across autoimmune diseases reveals shared associations and mechanisms. Cross-disease association analysis identifies genetic signals and biological mechanisms shared among multiple autoimmune disorders.
| Heidi Hautakangas et al. | Nature Genetics | 2022
Genome-wide analysis of 102,084 migraine cases identifies 123 risk loci and subtype-specific risk alleles. Greatly expands migraine loci and distinguishes genetic components shared or specific to migraine with and without aura.
| Mike A. Nalls et al. | The Lancet Neurology | 2019
Identification of novel risk loci, causal insights, and heritable risk for Parkinson's disease. Large GWAS meta-analysis expands Parkinson's disease loci and integrates genetic associations with expression, methylation, pathways, and risk prediction.
| Eli A. Stahl et al. | Nature Genetics | 2019
Genome-wide association study identifies 30 loci associated with bipolar disorder. Large psychiatric GWAS expands bipolar-disorder loci and implicates ion channels, neurotransmitter transport, and synaptic biological mechanisms.
| Matthias Wuttke et al. | Nature Genetics | 2019
A catalog of genetic loci associated with kidney function from analyses of a million individuals. Trans-ancestry analysis identifies hundreds of kidney-function loci and combines fine-mapping, gene expression, and functional evidence.
| Naomi R. Wray et al. | Nature Genetics | 2018
Genome-wide association analyses identify 44 risk variants and refine the genetic architecture of major depression. Large meta-analysis establishes dozens of depression loci and quantifies genetic overlap with psychiatric, behavioral, and metabolic traits.
| David M. Howard et al. | Nature Communications | 2018
Genome-wide association study of depression phenotypes in UK Biobank identifies variants in excitatory synaptic pathways. Uses multiple depression definitions to investigate genetic architecture and implicate neuronal and synaptic processes.
| Carolina Roselli et al. | Nature Genetics | 2018
Multi-ethnic genome-wide association study for atrial fibrillation. Expands atrial-fibrillation loci through multi-ethnic analysis and links association signals with genes and biological pathways relevant to cardiac electrophysiology.
| Rainer Malik et al. | Nature Genetics | 2018
Multiancestry genome-wide association study of 520,000 subjects identifies 32 loci associated with stroke and stroke subtypes. Expands stroke loci and identifies shared and subtype-specific genetic mechanisms across diverse ancestry groups.
| James C. Lee et al. | Nature Genetics | 2017
Genome-wide association study identifies distinct genetic contributions to prognosis and susceptibility in Crohn's disease. Shows that genetic variants influencing disease course can differ from variants affecting initial disease susceptibility.
| Fredrick R. Schumacher et al. | Nature Communications | 2015
Genome-wide association study of colorectal cancer identifies six new susceptibility loci. Uses large discovery and replication cohorts to expand known colorectal-cancer loci and illuminate biological mechanisms of inherited susceptibility.
| EAGLE Eczema Consortium | Nature Genetics | 2015
Multi-ancestry genome-wide association study of 21,000 cases and 95,000 controls identifies new risk loci for atopic dermatitis. Combines European, African, Japanese, and Latino data to identify additional immune-related eczema loci.
| Schizophrenia Working Group of the Psychiatric Genomics Consortium | Nature | 2014
Biological insights from 108 schizophrenia-associated genetic loci. Landmark psychiatric GWAS identifies 108 loci and implicates synaptic biology, neuronal signaling, calcium channels, and immune-related genomic regions.
| Brian M. Wolpin et al. | Nature Genetics | 2014
Genome-wide association study identifies multiple susceptibility loci for pancreatic cancer. Identifies inherited common variants contributing to pancreatic cancer risk and points to genomic regions for biological follow-up.
| Yukinori Okada et al. | Nature | 2014
Genetics of rheumatoid arthritis contributes to biology and drug discovery. Meta-analysis of more than 100,000 participants identifies 101 rheumatoid-arthritis loci and demonstrates how GWAS findings can prioritize biological pathways and therapeutic targets.
| Jean-Charles Lambert et al. | Nature Genetics | 2013
Meta-analysis of 74,046 individuals identifies 11 new susceptibility loci for Alzheimer's disease. Large consortium meta-analysis expands Alzheimer's risk loci and highlights immune, lipid-processing, and endocytic pathways.
| Schizophrenia Psychiatric Genome-Wide Association Study Consortium | Nature Genetics | 2011
Genome-wide association study identifies five new schizophrenia loci. Expands schizophrenia-associated regions and provides early large-scale evidence that immune and neuronal pathways contribute to its highly polygenic architecture.
Genetic risk and a primary role for cell-mediated immune mechanisms in multiple sclerosis. Large association study strongly implicates immune regulation and T-cell biology in inherited susceptibility to multiple sclerosis.
| Ellen L. Goode et al. | Nature Genetics | 2010
A genome-wide association study identifies susceptibility loci for ovarian cancer at 2q31 and 8q24. Expands the genetic architecture of ovarian cancer and demonstrates the value of large international cancer consortia.
| Miriam F. Moffatt et al. | New England Journal of Medicine | 2010
A large-scale, consortium-based genomewide association study of asthma. Large international GWAS identifies asthma susceptibility loci and demonstrates important genetic differences between childhood-onset and adult disease.
| Peter K. Gregersen et al. | Nature Genetics | 2009
REL, encoding a member of the NF-kappaB family of transcription factors, is a newly defined risk locus for rheumatoid arthritis. Identifies an immune-regulatory locus and strengthens evidence connecting NF-kappaB signaling with rheumatoid arthritis.
| Ian P. M. Tomlinson et al. | Nature Genetics | 2008
A genome-wide association scan of tag SNPs identifies a susceptibility variant for colorectal cancer at 8q24.21. Helps establish common inherited susceptibility loci for colorectal cancer through large-scale association analysis.
| Jeffrey C. Barrett et al. | Nature Genetics | 2008
Genome-wide association defines more than 30 distinct susceptibility loci for Crohn's disease. Large meta-analysis dramatically expands Crohn's disease loci and highlights innate immunity, autophagy, and inflammatory pathways.
| International Consortium for Systemic Lupus Erythematosus Genetics et al. | Nature Genetics | 2008
Genome-wide association scan in women with systemic lupus erythematosus identifies susceptibility variants in ITGAM, PXK, KIAA1542 and other loci. Major early autoimmune GWAS expands the genetic basis of lupus.
| Douglas F. Easton et al. | Nature | 2007
Genome-wide association study identifies novel breast cancer susceptibility loci. Landmark cancer GWAS identifies common breast-cancer risk loci including FGFR2 and demonstrates that common susceptibility alleles contribute to familial risk.
| Robert J. Klein et al. | Science | 2005
Complement factor H polymorphism in age-related macular degeneration. Landmark genome-wide association study identifies a strong complement-factor H association and helps establish GWAS as a powerful approach to common disease genetics.
Psychiatric and Behavioral GWAS
| Xiangrui Meng et al. | Nature Genetics | 2024
Multi-ancestry genome-wide association study of major depression aids locus discovery, fine mapping, gene prioritization and causal inference. A large multi-ancestry analysis identifies dozens of new depression loci and demonstrates how ancestry diversity can improve fine-mapping and gene discovery.
| Ditte Demontis et al. | Nature Genetics | 2023
Genome-wide analyses of ADHD identify 27 risk loci, refine the genetic architecture and implicate several cognitive domains. Identifies 27 ADHD loci and links genetic susceptibility with early brain development, neuronal biology, and cognitive traits.
| Trubetskoy et al. | Nature | 2022
Mapping genomic loci implicates genes and synaptic biology in schizophrenia. Analysis of more than 76,000 schizophrenia cases identifies 287 genomic loci and prioritizes genes involved in excitatory and inhibitory neuronal biology.
| Niamh Mullins et al. | Nature Genetics | 2021
Genome-wide association study of more than 40,000 bipolar disorder cases provides new insights into the underlying biology. Identifies 64 bipolar-disorder loci and implicates synaptic signaling, neuronal pathways, and genes targeted by several classes of psychiatric medications.
| Jonas Grove et al. | Nature Genetics | 2019
Identification of common genetic risk variants for autism spectrum disorder. Large-scale association analysis establishes genome-wide-significant common-variant loci for autism and demonstrates genetic overlap with several psychiatric and cognitive phenotypes.
| Hunna J. Watson et al. | Nature Genetics | 2019
Genome-wide association study identifies eight risk loci and implicates metabo-psychiatric origins for anorexia nervosa. Finds eight significant loci and demonstrates that anorexia nervosa shares genetic architecture with both psychiatric and metabolic traits.
| Joel Gelernter et al. | Nature Neuroscience | 2019
Genome-wide association study of post-traumatic stress disorder reexperiencing symptoms in over 165,000 US veterans. Uses a very large veteran cohort to identify genetic associations with PTSD reexperiencing symptoms and investigate their biological basis.
| Mengzhen Liu et al. | Nature Genetics | 2019
Association studies of up to 1.2 million individuals yield new insights into the genetic etiology of tobacco and alcohol use. Identifies hundreds of variants influencing smoking initiation, cessation, smoking intensity, and alcohol consumption and reveals extensive pleiotropy.
| Henry R. Kranzler et al. | Nature Communications | 2019
Genome-wide association study of alcohol consumption and use disorder in 274,424 individuals from multiple populations. Distinguishes genetic influences on alcohol consumption from those associated with alcohol-use disorder in a large ancestry-diverse sample.
| Mats Nagel et al. | Nature Genetics | 2018
Meta-analysis of genome-wide association studies for neuroticism in 449,484 individuals identifies novel genetic loci and pathways. Greatly expands the number of loci associated with neuroticism and highlights neuronal and developmental biological pathways.
Sleep, Cognition, and Brain Imaging
| Chun Chieh Fan et al. | Nature Communications | 2022
Multivariate genome-wide association study of brain diffusion measures identifies hundreds of genetic loci. Applies multivariate association analysis to diffusion MRI measurements and identifies genomic regions affecting white-matter microstructure.
| Stephen M. Smith et al. | Nature Neuroscience | 2021
An expanded set of genome-wide association studies of brain imaging phenotypes in UK Biobank. Uses thousands of MRI-derived phenotypes to map genetic influences on brain structure, connectivity, microstructure, and function.
| Philip R. Jansen et al. | Nature Genetics | 2019
Genome-wide analysis of insomnia in 1,331,010 individuals identifies new risk loci and functional pathways. One of the largest insomnia GWAS identifies hundreds of associated loci and demonstrates substantial genetic overlap with psychiatric and metabolic traits.
| Jacqueline M. Lane et al. | Nature Genetics | 2019
Biological and clinical insights from genetics of insomnia symptoms. Identifies dozens of loci influencing insomnia and connects associated variants with brain regions, neuronal pathways, cardiovascular disease, and psychiatric traits.
| Samuel E. Jones et al. | Nature Communications | 2019
Genome-wide association analyses of chronotype in 697,828 individuals provide insights into circadian rhythms. Identifies hundreds of variants associated with whether people naturally prefer earlier or later sleeping and waking times.
| Hassan S. Dashti et al. | Nature Communications | 2019
Genome-wide association study identifies genetic loci for self-reported habitual sleep duration supported by accelerometer-derived estimates. Identifies dozens of sleep-duration loci and compares questionnaire-based sleep measurements with objective accelerometer data.
| Heming Wang et al. | Nature Communications | 2019
Genome-wide association analysis of self-reported daytime sleepiness identifies 42 loci that suggest biological subtypes. Finds genetic pathways associated with daytime sleepiness and evidence for different biological mechanisms underlying excessive sleepiness.
| Bingxin Zhao et al. | Nature Genetics | 2019
Genome-wide association analysis of 19,629 individuals identifies variants influencing regional brain volumes and their overlap with neuropsychiatric traits. Maps common variants affecting structural brain measurements and investigates their relationships with neurological and psychiatric phenotypes.
| Jeanne E. Savage et al. | Nature Genetics | 2018
Genome-wide association meta-analysis in 269,867 individuals identifies new genetic and functional links to intelligence. Identifies hundreds of loci and genes related to cognitive ability and highlights neurodevelopmental and neuronal pathways.
| Samuel E. Jones et al. | Nature Communications | 2016
Genome-wide association analyses in 128,266 individuals identify new morningness and sleep duration loci. Early large-scale study maps genetic influences on morning preference and sleep duration and links circadian biology with human behavioral variation.
Reproductive Traits, Growth, and Development
| Loic Yengo et al. | Nature | 2022
A saturated map of common genetic variants associated with human height. Analysis of more than five million people identifies over 12,000 independent height-associated variants and approaches saturation of common-variant discovery.
| Nicole M. Warrington et al. | Nature Genetics | 2019
Maternal and fetal genetic effects on birth weight and their relevance to cardiometabolic risk factors. Separates maternal and fetal genetic influences to clarify how inherited variation and the intrauterine environment shape birth weight.
| Nicola Barban et al. | Nature Genetics | 2016
Genome-wide analysis identifies 12 loci influencing human reproductive behavior. Identifies genetic associations with age at first birth and number of children while emphasizing the substantial environmental component of reproductive behavior.
| Felix R. Day et al. | Nature Genetics | 2016
Physical and neurobehavioral determinants of reproductive onset and success. Identifies genetic loci influencing age at first sexual intercourse and examines connections among puberty, behavior, reproduction, and neurological traits.
| Momoko Horikoshi et al. | Nature | 2016
Genome-wide associations for birth weight and correlations with adult disease. Identifies numerous fetal-growth loci and examines genetic relationships between birth weight and later cardiometabolic disease.
| Momoko Horikoshi et al. | Nature Genetics | 2013
New loci associated with birth weight identify genetic links between intrauterine growth and adult height and metabolism. Expands birth-weight loci and reveals shared genetic influences connecting fetal growth with adult metabolic and anthropometric traits.
| Cathy E. Elks et al. | Nature Genetics | 2010
Thirty new loci for age at menarche identified by a meta-analysis of genome-wide association studies. Dramatically expands the known genetic architecture of female pubertal timing and implicates hormonal regulation and energy homeostasis.
| Rachel M. Freathy et al. | Nature Genetics | 2010
Variants in ADCY5 and near CCNL1 are associated with fetal growth and birth weight. Identifies common genetic variants influencing birth weight and establishes links between fetal development and later metabolic disease.
| Ken K. Ong et al. | Nature Genetics | 2009
Genetic variation in LIN28B is associated with the timing of puberty. Establishes LIN28B as an important genomic region influencing pubertal timing through association with age at menarche and related developmental traits.
| John R. B. Perry et al. | Nature Genetics | 2009
Meta-analysis of genome-wide association data identifies two loci influencing age at menarche. An early reproductive-trait GWAS identifies common variants involved in variation in female pubertal timing.
Cancer Genome-Wide Association Studies
| Jinyoung Byun et al. | Nature Genetics | 2022
Cross-ancestry genome-wide meta-analysis of lung cancer identifies novel susceptibility loci. Combines multiple ancestry groups to improve discovery and fine-mapping of inherited lung-cancer risk variants.
| David V. Conti et al. | Nature Genetics | 2021
Trans-ancestry genome-wide association meta-analysis of prostate cancer identifies new susceptibility loci and informs genetic risk prediction. Uses ancestry-diverse cohorts to identify additional prostate-cancer variants and investigate population differences in inherited risk.
| Haoyu Zhang et al. | Nature Genetics | 2020
Genome-wide association study identifies 32 novel breast cancer susceptibility loci from overall and subtype-specific analyses. Expands the inherited genetic architecture of breast cancer and identifies associations that differ among tumor subtypes.
| Takayuki Akamatsu et al. | Nature Communications | 2019
Genome-wide association study of prostate cancer in Japanese men identifies new susceptibility loci. Demonstrates the value of ancestry-specific cancer GWAS and identifies risk variants particularly informative in Japanese populations.
| Fredrick R. Schumacher et al. | Nature Genetics | 2018
Association analyses of more than 140,000 men identify 63 new prostate cancer susceptibility loci. Greatly expands the catalogue of common prostate-cancer variants and evaluates their cumulative contribution to inherited susceptibility.
| Gloria M. Petersen et al. | Nature Genetics | 2010
A genome-wide association study identifies pancreatic cancer susceptibility loci on chromosomes 13q22.1, 1q32.1 and 5p15.33. Expands inherited pancreatic-cancer risk loci and identifies regions containing genes with plausible roles in carcinogenesis.
| D. Timothy Bishop et al. | Nature Genetics | 2009
Genome-wide association study identifies three loci associated with melanoma risk. Identifies common variants affecting susceptibility to cutaneous melanoma and connects pigmentation biology with cancer risk.
| Rayjean J. Hung et al. | Nature | 2008
A susceptibility locus for lung cancer maps to nicotinic acetylcholine receptor genes on 15q25. Landmark GWAS links a chromosome 15 region containing nicotinic receptor genes with lung-cancer susceptibility and smoking-related biology.
| Brent W. Zanke et al. | Nature Genetics | 2007
Genome-wide association scan identifies a colorectal cancer susceptibility locus on chromosome 8q24. One of the earliest successful cancer GWAS identifies 8q24 as an important susceptibility region subsequently implicated in several cancers.
Autoimmune, Inflammatory, and Allergic Disease GWAS
| Kazuyoshi Ishigaki et al. | Nature Genetics | 2022
Multi-ancestry genome-wide association analyses identify novel genetic mechanisms in rheumatoid arthritis. Analysis of more than 276,000 participants identifies 124 loci, including 34 novel regions, and improves cross-population fine-mapping.
| Tomomitsu Hirota et al. | Nature Genetics | 2012
Genome-wide association study identifies eight new susceptibility loci for atopic dermatitis in the Japanese population. Discovers additional immune-related eczema loci and replicates several associations previously identified in European and Chinese populations.
| Liangdan Sun et al. | Nature Genetics | 2011
A genome-wide association study in the Chinese Han population identifies two new susceptibility loci for atopic dermatitis. Extends eczema genetics beyond European cohorts and demonstrates the importance of studying ancestry-diverse populations.
| Tomomitsu Hirota et al. | Nature Genetics | 2011
Genome-wide association study identifies three new susceptibility loci for adult asthma in the Japanese population. Identifies ancestry-specific and shared genetic influences on adult asthma and broadens understanding of allergic-disease genetics.
| Philip E. Stuart et al. | Nature Genetics | 2010
Genome-wide association analysis identifies three psoriasis susceptibility loci. Identifies psoriasis associations involving NOS2, FBXL19, and NFKBIA and strengthens evidence for immune and inflammatory signaling pathways.
| Andre Franke, Adrian Strange et al. | Nature Genetics | 2010
A genome-wide association study identifies new psoriasis susceptibility loci and an interaction between HLA-C and ERAP1. Demonstrates genetic interaction between antigen processing and presentation pathways in psoriasis susceptibility.
| Australo-Anglo-American Spondyloarthritis Consortium | Nature Genetics | 2010
Genome-wide association study of ankylosing spondylitis identifies non-MHC susceptibility loci. Identifies associations at 2p15, 21q22, ANTXR2, and IL1R2 and reinforces roles for IL-23 and IL-1 signaling.
| ANZgene Consortium | Nature Genetics | 2009
Genome-wide association study identifies new multiple sclerosis susceptibility loci on chromosomes 12 and 20. Expands the genetic architecture of multiple sclerosis and identifies additional immune-regulatory regions influencing disease risk.
| International Multiple Sclerosis Genetics Consortium | New England Journal of Medicine | 2007
Risk alleles for multiple sclerosis identified by a genomewide study. Landmark association study identifies immune-related susceptibility variants beyond the established HLA region and strengthens evidence for immune dysregulation in multiple sclerosis.
| David A. van Heel et al. | Nature Genetics | 2007
A genome-wide association study for celiac disease identifies risk variants in the region harboring IL2 and IL21. Establishes cytokine-regulatory loci outside the HLA region as important contributors to celiac-disease susceptibility.
Cardiovascular GWAS
| Derek Klarin et al. | Nature Genetics | 2023
Genome-wide association study of thoracic aortic aneurysm and dissection in the Million Veteran Program. Identifies numerous loci associated with thoracic aortic disease and connects inherited risk with vascular-development and extracellular-matrix biology.
| Krishna G. Aragam et al. | Nature Genetics | 2022
Discovery and systematic characterization of risk variants and genes for coronary artery disease in over a million participants. Identifies hundreds of coronary-disease associations and integrates genetic and functional evidence to prioritize candidate causal genes.
| Sonia Shah et al. | Nature Communications | 2020
Genome-wide association and Mendelian randomisation analysis provide insights into the pathogenesis of heart failure. Combines GWAS with causal-inference approaches to identify risk loci and investigate potentially modifiable contributors to heart failure.
| Derek Klarin et al. | Nature Genetics | 2019
Genome-wide association analysis of venous thromboembolism identifies new risk loci and genetic overlap with cardiovascular traits. Expands the genetic basis of venous thrombosis and highlights coagulation, blood-cell, and vascular mechanisms.
| Jonas B. Nielsen et al. | Nature Genetics | 2018
Biobank-driven genomic discovery yields new insight into atrial fibrillation biology. Analysis of more than one million participants identifies over 100 AF loci and implicates cardiac-developmental and electrophysiological pathways.
| Evangelos Evangelou et al. | Nature Genetics | 2018
Genetic analysis of over one million people identifies 535 new loci associated with blood pressure traits. Massively expands blood-pressure genetics and links associated variants with vascular, adrenal, renal, and metabolic biological pathways.
| Christopher P. Nelson et al. | Nature Genetics | 2017
Association analyses based on false discovery rate implicate new loci for coronary artery disease. Expands known CAD loci and demonstrates how large association datasets can uncover additional variants below conventional initial discovery thresholds.
| Panos Deloukas et al. | Nature Genetics | 2013
Large-scale association analysis identifies new risk loci for coronary artery disease. Large consortium analysis increases the number of established CAD loci and highlights lipid, inflammatory, vascular, and previously unknown pathways.
| Connie R. Bezzina et al. | Nature Genetics | 2010
Genome-wide association study identifies a susceptibility locus at 21q21 for ventricular fibrillation in acute myocardial infarction. Demonstrates that inherited common variation can influence the risk of life-threatening arrhythmia following myocardial infarction.
| Myocardial Infarction Genetics Consortium | Nature Genetics | 2009
Genome-wide association of early-onset myocardial infarction with common single nucleotide polymorphisms and copy number variants. Early large-scale cardiovascular GWAS identifies multiple common variants associated with premature myocardial infarction.
Metabolic, Liver, Kidney, and Biochemical Trait GWAS
| Vasiliki Lagou et al. | Nature Genetics | 2023
Sex-dimorphic genetic effects and novel loci for fasting and random glucose in large population studies. Expands the genetic architecture of blood-glucose regulation and investigates genetic effects that differ according to sex and metabolic context.
| Yanhua Chen et al. | Nature Genetics | 2023
Genome-wide association meta-analysis identifies new loci for nonalcoholic fatty liver disease. Expands known fatty-liver loci and integrates genomic and functional evidence to identify genes and biological pathways involved in hepatic fat accumulation.
| Anubha Mahajan et al. | Nature Genetics | 2022
Multi-ancestry genetic study of type 2 diabetes highlights the power of diverse populations for discovery and translation. Large ancestry-diverse analysis expands diabetes loci, improves fine-mapping, and identifies biological pathways relevant to disease development.
| Marijana Vujkovic et al. | Nature Genetics | 2022
A multiancestry genome-wide association study of unexplained chronic ALT elevation as a proxy for nonalcoholic fatty liver disease. Uses electronic health records and ancestry-diverse cohorts to discover dozens of genetic loci influencing chronic liver injury.
| Ji Chen et al. | Nature Genetics | 2021
The trans-ancestral genomic architecture of glycemic traits. Combines ancestry-diverse GWAS to discover and fine-map loci influencing glucose regulation and improves understanding of population-shared and population-specific signals.
| Vincent L. Chen et al. | Nature Communications | 2021
Genome-wide association study of serum liver enzymes implicates diverse metabolic and liver-specific pathways. Maps common genetic influences on liver-enzyme levels and provides clues to mechanisms affecting liver injury and metabolic disease.
| Adrienne Tin et al. | Nature Genetics | 2019
Target genes, variants, tissues and transcriptional pathways influencing human serum urate levels. Very large GWAS identifies numerous loci controlling urate concentrations and connects them with kidney transport and gout-related biology.
| Anna Köttgen et al. | Nature Genetics | 2013
Genome-wide association analyses identify 18 new loci associated with serum urate concentrations. Expands urate genetics and implicates kidney transporters, metabolic pathways, and regulatory mechanisms important to gout susceptibility.
| Robert A. Scott et al. | Nature Genetics | 2012
Large-scale association analyses identify new loci influencing glycemic traits and provide insight into the underlying biological pathways. Identifies additional variants controlling fasting glucose, insulin, and related metabolic phenotypes.
| Stefano Romeo et al. | Nature Genetics | 2008
Genetic variation in PNPLA3 confers susceptibility to nonalcoholic fatty liver disease. Landmark association study identifies the PNPLA3 locus as a major inherited determinant of liver-fat accumulation and liver-disease susceptibility.
Respiratory, Eye, Bone, and Structural Trait GWAS
| Nick Shrine et al. | Nature Genetics | 2023
Multi-ancestry genome-wide association analyses improve resolution of genes and pathways influencing lung function and COPD. Uses very large ancestry-diverse cohorts to discover and refine hundreds of respiratory-trait association signals.
| Louise V. Wain et al. | Nature Genetics | 2017
Genome-wide association analyses for lung function and chronic obstructive pulmonary disease identify new loci and potential drug targets. Integrates lung-function and COPD association data to expand risk loci and prioritize therapeutically relevant biological pathways.
| Daan W. Loth et al. | Nature Genetics | 2014
Genome-wide association analysis identifies six new loci associated with forced vital capacity. Identifies genetic determinants of lung volume and highlights pathways involved in development and pulmonary tissue structure.
| Karol Estrada et al. | Nature Genetics | 2012
Genome-wide meta-analysis identifies 56 bone mineral density loci and reveals 14 loci associated with fracture. Large GEFOS consortium analysis links common genetic variation with bone density, skeletal biology, and fracture susceptibility.
| María Soler Artigas et al. | Nature Genetics | 2011
Genome-wide association and large-scale follow-up identifies 16 new loci influencing lung function. Significantly expands pulmonary-function loci and implicates genes involved in lung growth, tissue remodeling, and airway biology.
| Emmanouela Repapi et al. | Nature Genetics | 2010
Genome-wide association study identifies five loci associated with lung function. Identifies common variants affecting spirometry measures and establishes several genes involved in pulmonary development and airway function.
| Dana B. Hancock et al. | Nature Genetics | 2010
Meta-analyses of genome-wide association studies identify multiple loci associated with pulmonary function. Demonstrates the polygenic basis of lung function and identifies genomic regions relevant to respiratory physiology and disease.
| Pirro G. Hysi et al. | Nature Genetics | 2010
A genome-wide association study for myopia and refractive error identifies a susceptibility locus at 15q25. Establishes common inherited variation contributing to refractive error and provides molecular clues to eye growth and myopia.
| Abbas M. Solouki et al. | Nature Genetics | 2010
A genome-wide association study identifies a susceptibility locus for refractive errors and myopia at 15q14. Independently identifies common genetic variation affecting refractive development and risk of nearsightedness.
| Fernando Rivadeneira et al. | Nature Genetics | 2009
Twenty bone-mineral-density loci identified by large-scale meta-analysis of genome-wide association studies. Expands genetic understanding of skeletal density and implicates biological pathways controlling bone formation and remodeling.
GWAS Statistical and Computational Methods
| Christiaan A. de Leeuw et al. | PLOS Computational Biology | 2015
MAGMA: generalized gene-set analysis of GWAS data. Develops a widely used framework for converting SNP-level association results into gene-level and pathway-level tests.
| Jian Yang et al. | American Journal of Human Genetics | 2011
GCTA: a tool for genome-wide complex trait analysis. Introduces software for estimating SNP-based heritability and examining the collective contribution of genome-wide common variants to complex traits.
| Teri A. Manolio | New England Journal of Medicine | 2010
Genomewide association studies and assessment of the risk of disease. Reviews how GWAS discoveries can and cannot be used for disease-risk estimation and explains challenges posed by small individual effect sizes and incomplete heritability.
| Jimmy Z. Liu et al. | American Journal of Human Genetics | 2010
A versatile gene-based test for genome-wide association studies. Introduces VEGAS, which combines evidence from multiple variants within genes while accounting for linkage disequilibrium.
| Cristen J. Willer, Yun Li and Gonçalo R. Abecasis | Bioinformatics | 2010
METAL: fast and efficient meta-analysis of genomewide association scans. Introduces widely used software for combining GWAS summary statistics across studies while efficiently handling millions of association results.
| Hyun Min Kang et al. | Nature Genetics | 2010
Variance component model to account for sample structure in genome-wide association studies. Introduces EMMAX, a mixed-model approach that efficiently controls population structure and relatedness in large association studies.
| John Hardy and Andrew Singleton | New England Journal of Medicine | 2009
Genomewide association studies and human disease. Explains the early transformation of human genetics brought about by GWAS and discusses biological interpretation of association signals.
| Bryan Howie, Peter Donnelly and Jonathan Marchini | PLOS Genetics | 2009
A flexible and accurate genotype imputation method for the next generation of genome-wide association studies. Introduces improved imputation methods that enabled researchers to infer millions of untyped genotypes from reference haplotypes.
| Brian L. Browning and Sharon R. Browning | American Journal of Human Genetics | 2009
A unified approach to genotype imputation and haplotype-phase inference for large data sets of trios and unrelated individuals. Presents computational methods that improve haplotype estimation and genotype imputation, both central components of modern GWAS pipelines.
| Shaun Purcell et al. | American Journal of Human Genetics | 2007
PLINK: a tool set for whole-genome association and population-based linkage analyses. Describes one of the most widely used software packages for genotype quality control, association testing, population structure, and genomic-data manipulation.
Proteomics, Metabolomics, Microbiome, and Molecular-Phenotype GWAS
| Diego A. Lopera-Maya et al. | Nature Genetics | 2022
Effect of host genetics on the gut microbiome in 7,738 participants of the Dutch Microbiome Project. Finds host-genetic associations with microbiome composition and highlights particularly strong effects involving LCT and ABO.
| Egil Ferkingstad et al. | Nature Genetics | 2021
Large-scale integration of the plasma proteome with genetics and disease. GWAS of thousands of plasma proteins identifies thousands of protein quantitative-trait loci and links protein variation with human diseases and therapeutic targets.
| Alexander Kurilshikov et al. | Nature Genetics | 2021
Large-scale association analyses identify host factors influencing human gut microbiome composition. International microbiome GWAS examines thousands of individuals and identifies human genetic variants affecting microbial taxa and biological pathways.
| Malte C. Rühlemann et al. | Nature Genetics | 2021
Genome-wide association study in 8,956 German individuals identifies influence of ABO histo-blood groups on gut microbiome. Demonstrates links between host blood-group genetics and intestinal microbial composition.
| Nobuo Fuse et al. | Communications Biology | 2020
Genome-wide association study of the human metabolome in a Japanese population. Maps inherited influences on a broad panel of metabolites and adds East Asian data to molecular-phenotype association research.
| Yunpeng Wang et al. | PLOS Genetics | 2020
Genome-wide association study identifies genetic loci influencing circulating cytokine concentrations at birth. Maps inherited influences on neonatal immune signaling and demonstrates how GWAS can investigate molecular traits early in human development.
| Benjamin B. Sun et al. | Nature | 2018
Genomic atlas of the human plasma proteome. Large proteomic association study maps genetic determinants of circulating proteins and provides a resource for connecting disease-associated loci with candidate biological mechanisms.
| Tingting Long et al. | Nature Genetics | 2017
Whole-genome sequencing identifies common-to-rare variants associated with human blood metabolites. Extends metabolic association mapping across the allele-frequency spectrum and identifies rare variants with substantial biochemical effects.
| So-Youn Shin et al. | Nature Genetics | 2014
An atlas of genetic influences on human blood metabolites. Maps hundreds of genetic associations with circulating metabolites and reveals links among genes, biochemical pathways, and complex diseases.
| David Melzer et al. | PLOS Genetics | 2008
A genome-wide association study identifies protein quantitative trait loci affecting circulating human proteins. Early proteomic GWAS demonstrates that common variants can strongly influence circulating protein concentrations.