Advances in Population Genetics: Difference between revisions

From WikiDemocracy
Jump to navigationJump to search
Lilly (talk | contribs)
Created page with "```wiki {{#seo: |title=Advances in Population Genetics |description=An overview of recent advances in population genetics, including ancient DNA, natural selection, demographic inference, ancestral recombination graphs, machine learning, human pangenomes, structural variation, and computational genomics. |keywords=population genetics, population genomics, ancient DNA, natural selection, demographic inference, human genetics, ancestral recombination graphs, machine learni..."
 
Lilly (talk | contribs)
Blanked the page
Tag: Blanking
 
Line 1: Line 1:
```wiki
{{#seo:
|title=Advances in Population Genetics
|description=An overview of recent advances in population genetics, including ancient DNA, natural selection, demographic inference, ancestral recombination graphs, machine learning, human pangenomes, structural variation, and computational genomics.
|keywords=population genetics, population genomics, ancient DNA, natural selection, demographic inference, human genetics, ancestral recombination graphs, machine learning, deep learning, pangenome, structural variation, genetic variation, admixture, gene flow, genomic diversity
|image=File:Placeholder.png
|image_width=300
|image_height=200
|type=article
}}


[[Category:Population genetics]]
[[Category:Population genomics]]
[[Category:Evolutionary genetics]]
[[Category:Human genetics]]
[[Category:Genomics]]
[[Category:Ancient DNA]]
[[Category:Bioinformatics]]
[[Category:Natural selection]]
__NOTOC__
== Advances in Population Genetics ==
Population genetics examines how genetic variation is distributed within and among populations and how evolutionary forces such as natural selection, genetic drift, mutation, recombination, migration, and demographic change shape that variation. The rapid expansion of genome sequencing, ancient DNA research, computational modeling, and statistical genetics has substantially increased the ability of researchers to reconstruct population histories and identify evolutionary processes.
Recent advances have transformed population genetics from a discipline that often relied on relatively small numbers of genetic markers into one capable of analyzing whole genomes from thousands of living and ancient individuals. These developments allow researchers to study evolutionary change over both long and relatively recent periods, reconstruct complex histories of population splitting and admixture, and estimate how natural selection has altered allele frequencies.
At the same time, increasingly sophisticated statistical and computational approaches have exposed important limitations in population-genetic inference. Mutation-rate variation, recombination-rate variation, linked selection, spatial population structure, sequencing quality, model assumptions, and sampling strategies can all affect estimates of demographic history and selection. Modern population genetics therefore increasingly combines large genomic datasets with simulations, model comparison, machine learning, and explicit tests of uncertainty.
=== Ancient DNA and Direct Observation of Evolutionary Change ===
Ancient DNA has become one of the most important developments in population genetics because it provides genetic information from individuals who lived at different points in the past. Traditional population-genetic reconstruction generally infers past events from patterns observed in living populations. Ancient genomes provide temporal observations that allow some evolutionary changes to be examined more directly.
Large collections of ancient genomes are now being used to study changes in allele frequencies through time. These datasets can reveal periods during which particular variants increased or decreased in frequency and can help researchers estimate the strength and timing of natural selection.
Studies of ancient populations have identified widespread evidence of directional selection and adaptation. Ancient genomes can also help distinguish changes caused by natural selection from changes caused by migration, admixture, demographic replacement, or genetic drift.
Combining ancient and modern genomes improves reconstruction of allele-frequency histories. Statistical methods increasingly use genetic time-series data to estimate selection coefficients while accounting for incomplete sampling and random genetic drift.
Ancient DNA therefore changes population genetics from a field based almost entirely on retrospective inference into one that can sometimes observe portions of evolutionary trajectories directly.
=== Natural Selection and Genomic Adaptation ===
Detecting natural selection remains a central problem in population genetics. Modern genomic datasets allow researchers to search for patterns created by positive selection, purifying selection, background selection, and other evolutionary processes.
Population-genomic scans examine patterns such as allele frequencies, haplotypes, linkage disequilibrium, and genetic differentiation to identify genomic regions that may have experienced selection. However, interpretation remains difficult because demographic events can produce patterns resembling selection.
Linked selection is particularly important. Purifying selection acting against harmful mutations can reduce genetic diversity in nearby genomic regions through background selection. Recombination rates and the organization of functional regions influence how widely these effects extend across the genome.
GC-biased gene conversion can also change patterns of genomic variation and potentially bias demographic inference. Research indicates that background selection and biased gene conversion affect substantial portions of the human genome.
Selection can also interact with geography. Beneficial variants spreading across spatially structured populations may create patterns that resemble local adaptation. Consequently, population structure and migration must be incorporated into interpretations of genomic selection scans.
Researchers are also developing methods for estimating simultaneous selection at multiple linked loci, particularly in admixed populations. These approaches recognize that adaptation may involve combinations of variants rather than isolated changes at single genetic positions.
=== Reconstructing Population History ===
Population genetics provides methods for reconstructing changes in population size, divergence, migration, admixture, and gene flow. Genomic datasets contain information about these processes because demographic events influence genetic diversity and genealogical relationships.
Coalescent theory remains a major foundation for demographic inference. Modern extensions allow researchers to estimate population histories using increasingly large genomic datasets and more complicated demographic models.
Admixture graphs represent population relationships as combinations of splits and gene-flow events. Bayesian methods can assign probabilities to different admixture histories and quantify uncertainty surrounding proposed population relationships.
Other methods analyze incomplete lineage sorting and genealogical discordance to reconstruct ancestral population sizes and divergence histories. These approaches are useful when relationships among populations or species cannot be represented by a simple branching tree.
Population-genetic methods are also being applied beyond human populations. Similar frameworks can reconstruct gene flow between species, estimate population trajectories in rapidly evolving organisms, and investigate the demographic histories of microbial populations.
Research on human gut microorganisms illustrates the expanding scope of the field. Population-genetic methods can be applied to microbial species associated with humans to examine demographic histories and selective pressures over human history.
=== Ancestral Recombination Graphs ===
Ancestral recombination graphs, often called ARGs, are becoming an important framework for population genetics. An ARG attempts to describe genealogical relationships across an entire recombining genome.
Because recombination causes different genomic regions to have different genealogical histories, a single evolutionary tree cannot accurately represent an entire genome. ARGs provide a framework for representing the network of genealogies produced by recombination.
Inferred ARGs can support analyses of demographic history, natural selection, mutation patterns, and ancestry. Improvements in computational efficiency are making ARG-based approaches increasingly practical for large genomic datasets.
However, ARG-based inference can also be affected by evolutionary processes that are not adequately represented in the underlying model. Linked natural selection, for example, can distort estimates of historical population size derived from genealogical reconstruction.
ARG methods therefore demonstrate both the increasing power and increasing complexity of population-genomic inference.
=== Machine Learning and Population-Genetic Inference ===
Machine learning and deep learning have become increasingly prominent in population genetics. Neural networks can be trained on simulated genomic datasets representing known evolutionary scenarios and then used to infer demographic or selective parameters from empirical data.
Machine-learning approaches have been applied to demographic reconstruction, natural-selection detection, ancestry analysis, and other population-genetic problems. Their ability to recognize complex patterns can make them useful when traditional summary statistics fail to capture all available genomic information.
Simulation-based supervised learning is particularly important. Researchers generate large numbers of datasets under different demographic parameters, train algorithms to recognize the resulting patterns, and then use those models to estimate parameters from observed genomic data.
Despite their potential, machine-learning methods introduce significant interpretability problems. A neural network may produce an accurate prediction without making it obvious which features of the genomic data contributed to that prediction.
Researchers have therefore developed methods for interpreting population-genetic machine learning, including experiments that rearrange haplotype matrices to determine which genomic patterns influence a model.
Preprocessing decisions can also affect what neural networks learn. Changes in the representation, ordering, or transformation of genomic data may alter model behavior and influence apparent accuracy.
These findings reinforce the importance of simulation, validation, model checking, and interpretability when machine learning is used for evolutionary inference.
=== Genome Simulation and Statistical Reliability ===
Genome simulation plays a major role in contemporary population genetics. Simulations allow researchers to generate genomic datasets under known evolutionary conditions and evaluate whether statistical methods can recover the correct history.
Standardized simulation frameworks are increasingly incorporating realistic genome structure and natural selection. These tools make it easier to compare inference methods using common demographic and evolutionary models.
Simulation studies have also revealed sources of bias. Variation in mutation and recombination rates can influence estimates of demographic history and the distribution of fitness effects.
Sequencing strategy can create additional complications. Low-pass genome sequencing provides a relatively inexpensive way to obtain population-scale genomic information, but low sequencing depth can introduce biases. Statistical models that explicitly account for those biases can improve inference.
Population-genetic researchers increasingly recommend testing multiple evolutionary models rather than relying on a single assumed history. Simulation-based model checking and explicit comparison of alternative demographic scenarios can help identify conclusions that depend excessively on model assumptions.
Competitions comparing demographic-inference strategies provide another method for evaluating reliability. By giving researchers simulated genomic datasets whose true evolutionary histories are initially concealed, these exercises test how successfully different methods reconstruct known histories.
=== Genetic Variation and Structural Variation ===
Advances in genome sequencing have broadened the concept of genetic variation. Population genetics once focused heavily on single-nucleotide variants, but modern studies increasingly examine structural variation, haplotypes, somatic variation, mosaicism, and other forms of genomic diversity.
Structural variants include deletions, duplications, inversions, insertions, and other large-scale genomic changes. These variants contribute to differences among individuals and populations and can influence both evolutionary processes and disease susceptibility.
Improved haplotype phasing attempts to determine which genetic variants occur together on individual chromosome copies. Genotype imputation uses observed variants and reference populations to predict genetic variants that were not directly measured.
More accurate phasing and imputation increase the amount of information that can be extracted from genomic datasets and improve many downstream population-genetic analyses.
Researchers are also increasingly examining variation within individuals. Germline mutations, somatic mutations, and genetic mosaicism demonstrate that genetic diversity exists not only among populations and individuals but also among different cells within the same person.
=== Human Pangenomes and Broader Genomic Diversity ===
Traditional human genomic research relied heavily on a single linear reference genome. Although enormously useful, a single reference cannot represent the full range of structural and sequence variation found among human populations.
The Human Pangenome Project seeks to build genomic references representing substantially greater human diversity. Pangenome approaches can incorporate alternative sequences and structural variants that are difficult to represent in a single linear genome.
Graph-based pangenome references allow multiple genomic paths to coexist within the same reference framework. This improves representation of genomic regions that vary substantially among individuals.
Pangenome resources may improve variant discovery, structural-variation analysis, ancestry research, and studies of population history by reducing biases associated with comparison against a single reference sequence.
The development of more representative genomic resources also highlights the importance of sampling diversity. Population-genetic conclusions depend on which populations are included in reference panels and genomic studies.
=== Cross-Ancestry Genomic Prediction ===
Population-genetic differences have important implications for genomic prediction. Predictive genetic models developed in one population may perform less accurately when applied to populations with different ancestry backgrounds.
Differences in allele frequencies and linkage disequilibrium contribute to this reduced transferability. Genetic associations identified in one population may therefore have weaker predictive value in another.
Research mapping cross-ancestry prediction accuracy demonstrates the importance of including diverse populations in genomic studies. More representative datasets can improve understanding of genetic variation while reducing the risk that genomic technologies work substantially better for some populations than others.
These findings connect population genetics with broader questions in medical genetics and genomic prediction, where understanding population structure is necessary for interpreting genetic associations responsibly.
=== Spatial Structure, Migration, and Admixture ===
Many traditional population-genetic models divide individuals into discrete populations. Real populations, however, are often distributed continuously across geography.
Continuous spatial structure can influence allele frequencies and genetic similarity. Individuals located near one another may be more genetically similar simply because migration occurs more frequently over short distances.
If continuous populations are incorrectly modeled as sharply separated groups, standard population-genetic statistics may give misleading results.
Migration and admixture further complicate population histories. Gene flow can introduce genetic variants into populations and alter patterns that might otherwise be interpreted as evidence of selection or population divergence.
Methods for identifying introgression and shared ancestral variation help distinguish genetic exchange from incomplete lineage sorting and other processes that produce similar genomic patterns.
Population genetics therefore increasingly treats geography, migration, population structure, and historical admixture as fundamental components of evolutionary inference rather than secondary complications.
=== Integrating Genomic and Epigenomic Information ===
Population-genetic inference is beginning to incorporate information beyond conventional DNA sequence variation. Genomic and epigenomic data can provide complementary information about evolutionary and demographic processes.
Methods extending sequentially Markovian coalescent approaches can incorporate highly mutable genomic or epigenomic markers. Because some of these markers change more rapidly than ordinary sequence variants, they may contain additional information about relatively recent demographic events.
Integrating multiple types of biological information represents a broader trend toward increasingly comprehensive evolutionary models. Future population genetics is likely to combine genome sequences, ancient DNA, epigenomic information, functional genomics, spatial data, and environmental information.
=== Challenges in Population-Genetic Inference ===
The increasing sophistication of population-genetic analysis does not eliminate uncertainty. Instead, larger datasets often reveal additional sources of complexity.
Demographic history can mimic natural-selection signals. Selection can distort demographic inference. Mutation and recombination rates vary across genomes. Population structure can resemble discrete ancestry groups even when genetic variation is geographically continuous.
Sequencing depth and data processing can introduce statistical biases. Machine-learning methods may identify predictive patterns without providing clear biological explanations. Different demographic models may also fit the same genomic observations.
These challenges make model testing and transparency increasingly important. Researchers must consider alternative demographic histories, evaluate assumptions with simulation, quantify uncertainty, and investigate whether apparent evolutionary signals remain robust under different analytical approaches.
Rather than producing a single unquestionable reconstruction of population history, modern population genetics increasingly provides probabilistic models whose strength depends on the quality of data, assumptions, and statistical validation.
=== Conclusion ===
Population genetics is being transformed by the convergence of large genomic datasets, ancient DNA, improved sequencing technology, sophisticated statistical models, genome simulation, ancestral recombination graphs, machine learning, and pangenome references.
Ancient DNA provides direct evidence of genetic change through time, while new approaches improve estimates of natural selection, demographic history, migration, admixture, and gene flow. ARGs and other genealogy-based methods allow increasingly detailed reconstruction of genomic ancestry, while machine learning provides tools for detecting complex patterns that may be difficult to capture with traditional statistics.
At the same time, recent research demonstrates that population-genetic inference remains sensitive to model assumptions, spatial structure, linked selection, mutation and recombination variation, sequencing strategies, and sampling. These findings have encouraged greater reliance on simulations, alternative-model testing, and explicit estimates of uncertainty.
The emerging picture is therefore not simply one of larger datasets and more powerful computers. Population genetics is moving toward increasingly realistic models of evolutionary history that recognize genomes as products of interacting processes operating across time, geography, populations, and genomic regions. As genomic resources become larger and more representative, these methods are likely to provide increasingly detailed insights into the origins, movements, adaptation, and genetic diversity of populations.
__TOC__
```
```wiki
=== Ancient DNA, Natural Selection, and Adaptation ===
'''1. Genomic Insights into Natural Selection in Recent Human History'''
[DOI:10.1038/s41576-026-01010-9 | Pontus Skoglund and Iain Mathieson | Nature Reviews Genetics | 2026]
Reviews how ancient and modern genomes are improving estimates of when, where, and how natural selection acted on human populations.
'''2. Insights into Human Adaptation from Ancient DNA'''
[DOI:10.1038/s41588-026-02562-6 | MemarMoshrefi et al. | Nature Genetics | 2026]
Examines how temporally sampled ancient genomes reveal adaptive changes that are difficult to reconstruct from living populations alone.
'''3. Ancient DNA Reveals Pervasive Directional Selection Across West Eurasia'''
[DOI:10.1038/s41586-026-10358-1 | Ali Akbari et al. | Nature | 2026]
Uses more than 15,000 ancient individuals to identify sustained allele-frequency changes consistent with directional selection across West Eurasia.
'''4. Selection Estimation from Genetic Time-Series Data: Effects of Limited Sampling and Genetic Drift'''
[DOI:10.1093/molbev/msaf301 | Cheng et al. | Molecular Biology and Evolution | 2025]
Quantifies how sampling noise and genetic drift influence selection-coefficient estimates obtained from allele-frequency trajectories.
'''5. Fast and Accurate Estimation of Selection Coefficients and Allele Histories from Ancient and Modern DNA'''
[DOI:10.1093/molbev/msae156 | Andrew H. Vaughn and Rasmus Nielsen | Molecular Biology and Evolution | 2024]
Introduces an efficient framework for reconstructing allele-frequency histories and estimating selection strength from ancient and contemporary genomes.
'''6. A Quantitative Genetic Model of Background Selection in Humans'''
[DOI:10.1371/journal.pgen.1011144 | Vince Buffalo and Andrew D. Kern | PLOS Genetics | 2024]
Develops a quantitative framework linking functional genomic architecture, recombination, and the broad effects of background selection on human diversity.
'''7. Inferring Multi-Locus Selection in Admixed Populations'''
[DOI:10.1371/journal.pgen.1011062 | Nicolas M. Ayala, Maximilian Genetti and Russell Corbett-Detig | PLOS Genetics | 2023]
Develops a framework for estimating simultaneous selection at multiple linked sites following population admixture.
'''8. Population Genetic Models for the Spatial Spread of Adaptive Variants: A Review in Light of SARS-CoV-2 Evolution'''
[DOI:10.1371/journal.pgen.1010391 | Margaret C. Steiner and John Novembre | PLOS Genetics | 2022]
Reviews mathematical models describing how beneficial alleles spread geographically and connects them with rapidly evolving pathogens.
'''9. Global Adaptation Complicates the Interpretation of Genome Scans for Local Adaptation'''
[PMCID:PMC7857299 | Lotterhos | Evolution Letters | 2021]
Demonstrates that globally beneficial alleles spreading through spatially structured populations can resemble signals normally attributed to local adaptation.
'''10. Background Selection and Biased Gene Conversion Affect More Than 95% of the Human Genome and Bias Demographic Inferences'''
[DOI:10.7554/eLife.36317 | Pouyet et al. | eLife | 2018]
Shows that linked purifying selection and GC-biased gene conversion influence much of the human genome and can distort demographic reconstruction.
=== Demographic History, Gene Flow, and Population Structure ===
'''11. Inference and Applications of Ancestral Recombination Graphs'''
[DOI:10.1038/s41576-024-00772-4 | Nielsen et al. | Nature Reviews Genetics | 2025]
Reviews ancestral recombination graphs as a framework for reconstructing genealogical relationships throughout recombining genomes.
'''12. GHIST 2024: The First Genomic History Inference Strategies Tournament'''
[DOI:10.1093/molbev/msaf257 | Travis J. Struck et al. | Molecular Biology and Evolution | 2025]
Benchmarks competing demographic-inference approaches using simulated genomic datasets whose evolutionary histories were initially concealed from participants.
'''13. Inference of Gene Flow between Species from Genomic Data When the Mode, Direction, and Lineages Are Misspecified'''
[DOI:10.1093/molbev/msaf121 | Thawornwattana et al. | Molecular Biology and Evolution | 2025]
Tests how assumptions about migration and introgression affect estimates of interspecific gene flow under multispecies coalescent models.
'''14. Bayesian Phylodynamic Inference of Multitype Population Trajectories Using Genomic Data'''
[DOI:10.1093/molbev/msaf130 | Müller et al. | Molecular Biology and Evolution | 2025]
Develops Bayesian approaches for reconstructing changes in multiple interacting population types directly from genealogical data.
'''15. Bayesian Model Averaging of Parametric Coalescent Models for Phylodynamic Inference'''
[DOI:10.1093/molbev/msaf297 | Magee et al. | Molecular Biology and Evolution | 2025]
Uses model averaging to reduce dependence on a single assumed demographic model when reconstructing population-size histories.
'''16. Inference of the Demographic Histories and Selective Effects of Human Gut Commensal Microbiota Over the Course of Human History'''
[DOI:10.1093/molbev/msaf010 | Mah, Lohmueller and Garud | Molecular Biology and Evolution | 2025]
Applies population-genetic demographic and fitness-effect inference to dozens of widespread human gut microbial species.
'''17. Biases in ARG-Based Inference of Historical Population Size in Populations Experiencing Selection'''
[DOI:10.1093/molbev/msae118 | Jacob I. Marsh and Parul Johri | Molecular Biology and Evolution | 2024]
Shows how linked natural selection can distort population-size histories inferred from ancestral recombination graphs.
'''18. The Effects of Mutation and Recombination Rate Heterogeneity on the Inference of Demography and the Distribution of Fitness Effects'''
[DOI:10.1093/gbe/evae004 | Vivak Soni, Susanne P. Pfeifer and Jeffrey D. Jensen | Genome Biology and Evolution | 2024]
Demonstrates that realistic variation in mutation and recombination rates can materially affect demographic and selection estimates.
'''19. TRAILS: Tree Reconstruction of Ancestry Using Incomplete Lineage Sorting'''
[DOI:10.1371/journal.pgen.1010836 | Rivas-González et al. | PLOS Genetics | 2024]
Introduces a hidden Markov model for recovering ancestral population sizes and divergence histories from genealogical discordance among species.
'''20. Leveraging Shared Ancestral Variation to Detect Local Introgression'''
[DOI:10.1371/journal.pgen.1010155 | Lopez Fang et al. | PLOS Genetics | 2024]
Introduces the D+ statistic, which uses ancestral and derived allele sharing to improve detection of localized introgressed genomic regions.
'''21. Developing an Evolutionary Baseline Model for Humans: Jointly Inferring Purifying Selection with Population History'''
[DOI:10.1093/molbev/msad100 | Parul Johri, Susanne P. Pfeifer and Jeffrey D. Jensen | Molecular Biology and Evolution | 2023]
Jointly models demography, purifying selection, recombination, and genome architecture to build more realistic human evolutionary baselines.
'''22. Bayesian Inference of Admixture Graphs on Native American and Arctic Populations'''
[DOI:10.1371/journal.pgen.1010410 | Nielsen et al. | PLOS Genetics | 2023]
Applies Bayesian inference to admixture graphs, providing probability-based estimates of complex population splits and gene-flow histories.
'''23. Space Is the Place: Effects of Continuous Spatial Structure on Analysis of Population Genetic Data'''
[Genetics 215:193–214 | Gideon S. Bradburd and Peter L. Ralph | Genetics | 2020]
Explains how continuous geography can distort standard population-genetic statistics when populations are incorrectly modeled as discrete units.
=== Machine Learning, Simulation, and Statistical Inference ===
'''24. Accessible, Realistic Genome Simulation with Selection Using stdpopsim'''
[DOI:10.1093/molbev/msaf236 | Graham Gower et al. | Molecular Biology and Evolution | 2025]
Extends standardized population-genetic simulation infrastructure to realistic models incorporating natural selection and complex genome structure.
'''25. Assessing Simulation-Based Supervised Machine Learning for Demographic Parameter Inference from Genomic Data'''
[DOI:10.1038/s41437-025-00773-x | Arnaud Quelin, Frédéric Austerlitz and Flora Jay | Heredity | 2025]
Tests how well supervised machine-learning approaches recover demographic parameters when trained on population-genetic simulations.
'''26. Interpreting Supervised Machine Learning Inferences in Population Genomics Using Haplotype Matrix Permutations'''
[DOI:10.1093/molbev/msaf250 | Linh N. Tran, David Castellano and Ryan N. Gutenkunst | Molecular Biology and Evolution | 2025]
Develops tools for determining which genomic patterns neural networks use when making population-genetic predictions.
'''27. Modeling Biases from Low-Pass Genome Sequencing to Enable Accurate Population Genetic Inferences'''
[DOI:10.1093/molbev/msaf002 | Crawford and Gutenkunst | Molecular Biology and Evolution | 2025]
Incorporates sequencing-depth biases directly into demographic models, allowing inexpensive low-pass genomic datasets to be used more reliably.
'''28. Harnessing Deep Learning for Population Genetic Inference'''
[DOI:10.1038/s41576-023-00636-3 | Xin Huang et al. | Nature Reviews Genetics | 2024]
Reviews neural-network applications for demographic reconstruction, selection detection, ancestry analysis, and other population-genetic problems.
'''29. Computationally Efficient Demographic History Inference from Allele Frequencies with Supervised Machine Learning'''
[DOI:10.1093/molbev/msae077 | Tran et al. | Molecular Biology and Evolution | 2024]
Develops machine-learning methods that rapidly infer demographic parameters from allele-frequency information while maintaining competitive accuracy.
'''30. Deep Learning in Population Genetics'''
[DOI:10.1093/gbe/evad008 | Kevin Korfmann, Oscar E. Gaggiotti and Matteo Fumagalli | Genome Biology and Evolution | 2023]
Surveys deep-learning architectures being applied to genomic datasets and discusses interpretability, simulations, and model uncertainty.
'''31. On Convolutional Neural Networks for Selection Inference: Revealing the Effect of Preprocessing on Model Learning and the Capacity to Discover Novel Patterns'''
[DOI:10.1371/journal.pcbi.1010979 | Ryan M. Cecil and Lauren A. Sugden | PLOS Computational Biology | 2023]
Shows that preprocessing decisions can strongly affect what neural networks learn when identifying genomic signatures of selection.
'''32. Recommendations for Improving Statistical Inference in Population Genomics'''
[DOI:10.1371/journal.pbio.3001669 | Johri et al. | PLOS Biology | 2022]
Examines common statistical pitfalls in population genomics and recommends simulation, model checking, and explicit consideration of alternative evolutionary histories.
=== Genomic Variation, Haplotypes, Pangenomes, and Cross-Ancestry Prediction ===
'''33. Advances in Haplotype Phasing and Genotype Imputation'''
[DOI:10.1038/s41576-025-00895-2 | Quan Sun and Yun Li | Nature Reviews Genetics | 2026]
Reviews improved statistical and computational approaches for reconstructing haplotypes and imputing unobserved genetic variants.
'''34. Diversity and Consequences of Structural Variation in the Human Genome'''
[DOI:10.1038/s41576-024-00808-9 | Collins and Talkowski | Nature Reviews Genetics | 2025]
Explains how deletions, duplications, inversions, and other structural variants contribute to human diversity, evolution, and disease.
'''35. Genetic Variation Across and Within Individuals'''
[DOI:10.1038/s41576-024-00709-x | Yu et al. | Nature Reviews Genetics | 2024]
Reviews germline, somatic, mosaic, and population-level genomic variation and the technologies increasingly capable of detecting it.
'''36. Mapping the Relative Accuracy of Cross-Ancestry Prediction'''
[DOI:10.1038/s41467-024-54727-8 | Alexa S. Lupi, Ana I. Vazquez and Gustavo de los Campos | Nature Communications | 2024]
Quantifies how genetic ancestry, allele frequencies, and linkage disequilibrium influence the transferability of genomic prediction across populations.
'''37. A Draft Human Pangenome Reference'''
[DOI:10.1038/s41586-023-05896-x | Wen-Wei Liao et al. | Nature | 2023]
Presents a graph-based human pangenome capturing substantially more structural and sequence diversity than conventional reference genomes.
'''38. The Human Pangenome Project: A Global Resource to Map Genomic Diversity'''
[DOI:10.1038/s41586-022-04601-8 | Wang et al. | Nature | 2022]
Describes an international effort to replace reliance on a single linear reference genome with references representing broader human diversity.
=== Integrative Population-Genomic Methods ===
'''39. Population Genomic Scans for Natural Selection and Demography'''
[DOI:10.1146/annurev-genet-111523-102651 | Xiaoheng Cheng and Matthias Steinrücken | Annual Review of Genetics | 2024]
Reviews modern approaches for detecting natural selection and reconstructing demographic history from population-scale genomic data.
'''40. Improved Inference of Population Histories by Integrating Genomic and Epigenomic Data'''
[DOI:10.7554/eLife.89470.4 | Thibaut Sellinger, Frank Johannes and Aurélien Tellier | eLife | 2024]
Extends sequentially Markovian coalescent inference so hypermutable genomic and epigenomic markers can improve reconstruction of recent demographic events.
```