Rare variant association testing

Rare variant association testing is a class of statistical methods used in genetic association studies to test whether rare genetic variants within a gene, genomic region, pathway, or other predefined variant set are associated with a phenotype. These methods are used because individual rare variants are observed in relatively few individuals, which can make single-variant association tests underpowered.[1][2]

Rare variant association tests are commonly applied to sequencing data from whole exome sequencing, whole genome sequencing, and targeted resequencing studies. Instead of testing each variant separately, they aggregate variants within a defined unit and test the combined evidence for association. Common approaches include burden tests, variance-component tests such as the sequence kernel association test (SKAT), combined tests such as SKAT-O, p-value combination methods such as the aggregated Cauchy association test (ACAT), and annotation-informed methods such as STAAR.[3][4][5][6]

Background

Genome-wide association studies based on common variants test genetic markers one at a time. This approach can be inefficient for rare variants because any individual rare variant may be carried by only a small number of people in a study. Rare variant association testing addresses this by grouping variants and testing their combined relationship with a trait or disease.[1] Variant sets may be defined by protein-coding genes, exons, regulatory regions, sliding genomic windows, pathways, or functional annotations. The variants included in a test are often filtered or weighted by properties such as minor allele frequency, predicted functional consequence, or external annotation scores.[7]

History

Early rare variant association methods focused on collapsing or aggregating variants within a region. Li and Leal proposed methods for detecting associations with rare variants using sequence data.[1] Madsen and Browning introduced a weighted-sum statistic for groupwise association testing of rare mutations.[8] Subsequent work evaluated rare variant association approaches and developed pooled or set-based tests for resequencing studies.[9][10] The sequence kernel association test introduced a variance-component framework for testing the association between a set of variants and a continuous or dichotomous trait while adjusting for covariates.[3] SKAT-O was later developed as an approach that combines burden and SKAT-type tests to improve performance across different genetic architectures.[4][11]

Later methods incorporated p-value combination and functional annotations. ACAT combines p-values using a Cauchy-based transformation and has been used to construct ACAT-V and ACAT-O tests.[5] STAAR extended rare variant association testing by incorporating multiple functional annotations and annotation-based weighting schemes for large whole-genome sequencing studies.[6]

Methodological classes

Burden tests

Burden tests collapse multiple variants within a set into a single genetic score for each individual. The score is then tested for association with the phenotype. These tests can be powerful when most causal variants in a set influence the phenotype in the same direction and have similar effects.[8][4]

A simple burden score may be written as:

where is the genotype of individual at variant , is a variant weight, and is the number of variants in the set.[citation needed]

Variance-component tests

Variance-component tests model variant effects as random effects rather than collapsing all variants into a single score. This allows variants within a set to have effects that differ in magnitude or direction. SKAT is a widely used variance-component test for rare variant association analysis.[3]

In SKAT, genetic similarity between individuals can be represented by a kernel matrix:

where is the genotype matrix and is a diagonal matrix of variant weights. Association is then tested using a variance-component score statistic.[3]

Combined tests

Combined tests are designed to retain power across different genetic architectures. SKAT-O combines burden and SKAT-type statistics, allowing the test to adapt between a burden-like model and a variance-component model.[4][11]

P-value combination methods

The aggregated Cauchy association test combines p-values using a Cauchy transformation. In rare variant analysis, ACAT can combine variant-level p-values to form ACAT-V, or combine p-values from different set-based tests to form omnibus tests such as ACAT-O.[5]

Annotation-informed tests

Annotation-informed methods use external information about variants, such as predicted functional impact, conservation, regulatory annotation, or allele frequency. STAAR incorporates both qualitative variant categories and multiple quantitative annotations using an omnibus weighting framework.[6] STAARpipeline extended this approach to large-scale whole-genome sequencing studies, including coding and noncoding analyses.[12]

Applications

Rare variant association testing is used in studies of complex traits, Mendelian and oligogenic disease risk, quantitative traits, and biobank-scale sequencing studies. It is commonly applied at the gene level in exome sequencing studies and at both gene-centric and non-gene-centric units in whole-genome sequencing studies.[2][7][12]

Considerations

The performance of rare variant association tests depends on the biological architecture of the trait, sample size, phenotype definition, allele-frequency thresholds, variant grouping, ancestry adjustment, relatedness, and functional annotation quality. Burden tests may lose power when variants have effects in opposite directions, while variance-component and omnibus tests may be more robust across heterogeneous effect patterns.[4][2]

See also

References

  1. ^ a b c Li, Bingshan; Leal, Suzanne M. (2008). "Methods for Detecting Associations with Rare Variants for Common Diseases: Application to Analysis of Sequence Data". The American Journal of Human Genetics. 83 (3): 311–321. doi:10.1016/j.ajhg.2008.06.024. PMC 2842185. PMID 18691683.
  2. ^ a b c Lee, Seunggeun; Abecasis, Gonçalo R.; Boehnke, Michael; Lin, Xihong (2014). "Rare-Variant Association Analysis: Study Designs and Statistical Tests". The American Journal of Human Genetics. 95 (1): 5–23. Bibcode:2014AmJHG..95....5L. doi:10.1016/j.ajhg.2014.06.009. PMC 4085641. PMID 24995866.
  3. ^ a b c d Wu, Michael C.; Lee, Seunggeun; Cai, Tianxi; Li, Yun; Boehnke, Michael; Lin, Xihong (2011). "Rare-Variant Association Testing for Sequencing Data with the Sequence Kernel Association Test". The American Journal of Human Genetics. 89 (1): 82–93. Bibcode:2011AmJHG..89...82W. doi:10.1016/j.ajhg.2011.05.029. PMC 3135811. PMID 21737059.
  4. ^ a b c d e Lee, Seunggeun; Wu, Michael C.; Lin, Xihong (2012). "Optimal tests for rare variant effects in sequencing association studies". Biostatistics. 13 (4): 762–775. doi:10.1093/biostatistics/kxs014. PMC 3440237. PMID 22699862.
  5. ^ a b c Liu, Yaowu; Chen, Sixing; Li, Zilin; Morrison, Alanna C.; Boerwinkle, Eric; Lin, Xihong (2019). "ACAT: A Fast and Powerful p Value Combination Method for Rare-Variant Analysis in Sequencing Studies". The American Journal of Human Genetics. 104 (3): 410–421. doi:10.1016/j.ajhg.2019.01.002. PMC 6407498. PMID 30849328.
  6. ^ a b c Li, Xihao; Li, Zilin; Zhou, Hufeng; Gaynor, Sheila M.; Liu, Yaowu; Chen, Han; Sun, Ryan; Dey, Rounak; Lin, Xihong (2020). "Dynamic incorporation of multiple in silico functional annotations empowers rare variant association analysis of large whole-genome sequencing studies at scale". Nature Genetics. 52 (9): 969–983. doi:10.1038/s41588-020-0676-4. PMC 7483769. PMID 32839606.
  7. ^ a b Povysil, Gundula; Petrovski, Slavé; Hostyk, Julia; Aggarwal, Varun; Allen, Andrew S.; Goldstein, David B. (2019). "Rare-variant collapsing analyses for complex traits: guidelines and applications". Nature Reviews Genetics. 20 (12): 747–759. doi:10.1038/s41576-019-0177-4. PMID 31605095.
  8. ^ a b Madsen, Bo Eskerod; Browning, Sharon R. (2009). "A Groupwise Association Test for Rare Mutations Using a Weighted Sum Statistic". PLOS Genetics. 5 (2) e1000384. doi:10.1371/journal.pgen.1000384. PMC 2633048. PMID 19214210.
  9. ^ Morris, Andrew P.; Zeggini, Eleftheria (2010). "An evaluation of statistical approaches to rare variant analysis in genetic association studies". Genetic Epidemiology. 34 (2): 188–193. doi:10.1002/gepi.20450. PMC 2962811. PMID 19810025.
  10. ^ Price, Alkes L.; Kryukov, Gregory V.; de Bakker, Paul I. W.; Purcell, Shaun M.; Staples, Jeff; Wei, Li-Jen; Sunyaev, Shamil R. (2010). "Pooled Association Tests for Rare Variants in Exon-Resequencing Studies". The American Journal of Human Genetics. 86 (6): 832–838. doi:10.1016/j.ajhg.2010.04.005. PMC 3032073. PMID 20471002.
  11. ^ a b Lee, Seunggeun; Emond, Mary J.; Bamshad, Michael J.; Barnes, Kathleen C.; Rieder, Mark J.; Nickerson, Deborah A.; Christian, iPSYCH; Wijsman, Ellen M.; Tabor, Holly K.; Leal, Suzanne M.; Browning, Brian L.; Lin, Xihong (2012). "Optimal Unified Approach for Rare-Variant Association Testing with Application to Small-Sample Case-Control Whole-Exome Sequencing Studies". The American Journal of Human Genetics. 91 (2): 224–237. doi:10.1016/j.ajhg.2012.06.007. PMC 3415556. PMID 22863193.
  12. ^ a b Li, Zilin; Li, Xihao; Liu, Yaowu; Shen, Jingjing; Chen, Han; Zhou, Hufeng; Morrison, Alanna C.; Boerwinkle, Eric; Lin, Xihong (2022). "A framework for detecting noncoding rare-variant associations of large-scale whole-genome sequencing studies". Nature Methods. 19 (12): 1599–1611. doi:10.1038/s41592-022-01640-x. PMC 10008172. PMID 36303018.

Content Disclaimer

Informasi ini disarikan dari Wikipedia dan disajikan kembali untuk tujuan edukasi. Konten tersedia di bawah lisensi CC BY-SA 3.0. Kami tidak bertanggung jawab atas ketidakakuratan data yang bersumber dari kontribusi publik tersebut.

  1. The information displayed on this website is sourced in part or in whole from Wikipedia and has been adapted for the purpose of restating it. We strive to provide accurate and relevant information, however:
  2. There is no guarantee of absolute accuracy. Wikipedia is an open, collaborative project that can be edited by anyone, so information is subject to change.
  3. It is not intended to constitute professional advice. The content displayed is for informational and educational purposes only. For important decisions (e.g., medical, legal, or financial), please consult a professional.
  4. Content copyright. Wikipedia is licensed under the Creative Commons Attribution-ShareAlike License (CC BY-SA). This means that content may be reused with appropriate attribution and shared under a similar license.
  5. Responsible use. Any risk arising from the use of information from this website is entirely the responsibility of the user.