The spectrum of copy number variation in the Pan-Canadian HostSeq databank

The integration of copy number variant (CNV) workflows into genome sequencing (GS) analysis pipelines allows for the identification of CNVs implicated in disease. Here, we generated a novel resource of CNVs identified in the HostSeq population cohort from Canada, identified the prevalence of recurrent CNVs associated with neurodevelopmental disorders, and determined CNVs of potential clinical relevance for reproductive planning and personal disease risk. GS data and CNV calls were generated for 10,488 participants from across Canada as part of the HostSeq initiative. The CNV calls were filtered to generate a rare dataset, which was further filtered into the OMIM morbid, ClinGen dosage, and DECIPHER datasets. CNV deletions were stratified into either Tier 1, 2, or 3 based on the inheritance pattern of the genes involved. A putatively pathogenic dataset was generated by identifying CNVs in the ClinGen dosage and DECIPHER datasets with at least 80% overlap with previously identified pathogenic CNVs. A total of 8,543,334 CNV calls were generated. Filtering for rare variants yielded 36,631 CNVs, of which 9,922 (27.08%) encompassed at least one OMIM gene, 728 (1.99%) had at least 10% overlap with a DECIPHER region, and 1,136 (3.10%) encompassed at least one ClinGen dosage-sensitive gene. CNV deletions were stratified into 1,833 Tier 1 deletions, 32 Tier 2 deletions, and 184 Tier 3 deletions. There were 82 CNVs identified in regions associated with neurodevelopmental disorders. Of the genes with either a ClinGen dosage sensitivity score or overlap with a DECIPHER region, 206 CNVs were deemed as putatively pathogenic. We were able to detect a wide range of CNVs, highlighting the use of integrating CNV workflows into the analysis pipeline, and generated a data resource for medical genomics for the Canadian population encompassing the full spectrum of CNVs.

Authors

Institutions

Publication Details

Journal
PLoS ONE
Published
2026-09-25
DOI
https://doi.org/10.1371/journal.pone.0359210
Primary Topic
Genomic variations and chromosomal abnormalities
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

The spectrum of copy number variation in the Pan-Canadian HostSeq databank

Steven Friedman, Tony Mazzulli, Allison McGeer, Erika Frangione et al.
PLoS ONE
Genomic variations and chromosomal abnormalities
article

The spectrum of copy number variation in the Pan-Canadian HostSeq databank

Steven Friedman, Tony Mazzulli, Allison McGeer, Erika Frangione, Bjug Borgundvaag, Lee William Goneau, Elisa Lapadula, Alexandra Binnie, Sunakshi Chowdhary, Zeeshan Khan, Yvonne Bombard, Bhooma Thiruvahindrapuram, Jordan P. Lerner‐Ellis, Marc Dagher, Hanna Faghfoury, David Di Iorio, Luke Devine, Shelley L. McLeod, Stephen W. Scherer, Saranya Arnoldo, Dawit Wolday, Seth Stern, Trevor J. Pugh, Selina Casalino, Abdul Noor, Chun Yiu Jordan Fung, Navneet Aujla, Chloe Mighton, Juliet Young, Gregory Morgan, Radhika Mahajan, Jared Simpson, Maahil Arshad, Anne-Claude Gingras, Ahmed Taher, David Richardson, Lochana Jayachandran, Elena Greenfeld, Marc Clausen, Jennifer Taher, CGEn HostSeq Initiative, Lisa Strug, Georgia MacDonald
article en

Abstract

The integration of copy number variant (CNV) workflows into genome sequencing (GS) analysis pipelines allows for the identification of CNVs implicated in disease. Here, we generated a novel resource of CNVs identified in the HostSeq population cohort from Canada, identified the prevalence of recurrent CNVs associated with neurodevelopmental disorders, and determined CNVs of potential clinical relevance for reproductive planning and personal disease risk. GS data and CNV calls were generated for 10,488 participants from across Canada as part of the HostSeq initiative. The CNV calls were filtered to generate a rare dataset, which was further filtered into the OMIM morbid, ClinGen dosage, and DECIPHER datasets. CNV deletions were stratified into either Tier 1, 2, or 3 based on the inheritance pattern of the genes involved. A putatively pathogenic dataset was generated by identifying CNVs in the ClinGen dosage and DECIPHER datasets with at least 80% overlap with previously identified pathogenic CNVs. A total of 8,543,334 CNV calls were generated. Filtering for rare variants yielded 36,631 CNVs, of which 9,922 (27.08%) encompassed at least one OMIM gene, 728 (1.99%) had at least 10% overlap with a DECIPHER region, and 1,136 (3.10%) encompassed at least one ClinGen dosage-sensitive gene. CNV deletions were stratified into 1,833 Tier 1 deletions, 32 Tier 2 deletions, and 184 Tier 3 deletions. There were 82 CNVs identified in regions associated with neurodevelopmental disorders. Of the genes with either a ClinGen dosage sensitivity score or overlap with a DECIPHER region, 206 CNVs were deemed as putatively pathogenic. We were able to detect a wide range of CNVs, highlighting the use of integrating CNV workflows into the analysis pipeline, and generated a data resource for medical genomics for the Canadian population encompassing the full spectrum of CNVs.

PLoS ONEVol. 21(9)
Ontario Institute for Cancer Research (CA), University Health Network (CA), Universidade Presbiteriana Mackenzie (BR), University of Toronto (CA), Lunenfeld-Tanenbaum Research Institute (CA), Hospital for Sick Children (CA), Women's College Hospital (CA), Health Net (US), Mount Sinai Hospital (US), William Osler Health System (CA), Schwartz/Reisman Emergency Medicine Institute (CA), Unity Health Toronto
Openalex Percentile: Top 12%
Genomic variations and chromosomal abnormalities
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.