	<rdf:RDF xmlns:admin="http://webns.net/mvcb/" xmlns="http://purl.org/rss/1.0/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:prism="http://purl.org/rss/1.0/modules/prism/" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:syn="http://purl.org/rss/1.0/modules/syndication/">
	<channel rdf:about="https://biorxiv.org">
	<admin:errorReportsTo rdf:resource="mailto:biorxiv@cshlpress.edu"/>
	<title>bioRxiv Channel: Somatic Mosaicism across the Human Tissues Network (SMaHT)</title>
	<link>https://biorxiv.org</link>
	<description>
	This feed contains articles for bioRxiv Channel "Somatic Mosaicism across the Human Tissues Network (SMaHT)"
	</description>

		<items>
	<rdf:Seq>
		</rdf:Seq>
	</items>
	<prism:eIssn/>
	<prism:publicationName>bioRxiv</prism:publicationName>
	<prism:issn/>

	<image rdf:resource=""/>
	</channel>
	<image rdf:about="">
	<title>bioRxiv</title>
	<url/>
	<link>https://biorxiv.org</link>
	</image>
	<item rdf:about="https://biorxiv.org/cgi/content/short/2024.12.18.629274v1?rss=1">
<title>
<![CDATA[
A personalized multi-platform assessment of somatic mosaicism in the human frontal cortex 
]]>
</title>
<link>
https://biorxiv.org/cgi/content/short/2024.12.18.629274v1?rss=1"
</link>
<description><![CDATA[
Somatic mutations in individual cells create genomic mosaicism, influencing genetic disorders and cancers. While clonal mutations in cancers are well-studied, rarer somatic variants in normal tissues remain poorly characterized. This study systematically evaluates detection methods using a personalized donor-specific assembly (DSA) from a neurotypical individuals dorsolateral prefrontal cortex assessed with Oxford Nanopore, NovaSeq, linked-read sequencing, Cas9-targeted long-read sequencing (TEnCATS), and single-neuron MALBAC amplification. The haplotype-resolved DSA improved cross-platform analysis, dramatically increasing phasing rates. Germline SNVs, structural variations (SVs), and transposable elements (TEs) were recalled with 99.4%-99.7% accuracy in bulk tissue, and phased haplotype analysis reduced false positives by 15.4%-75.1% for putative somatic candidates. Long-read single-neuron sequencing detected nine somatic SV candidates, demonstrating enhanced sensitivity for rare variants, while TEnCATS identified eight low-frequency somatic TE candidates. These findings highlight advanced methodologies for precise somatic variant detection, critical for understanding mosaicisms role in health and disease.
]]></description>
<dc:creator>Zhou, W.</dc:creator>
<dc:creator>Mumm, C.</dc:creator>
<dc:creator>Gan, Y.</dc:creator>
<dc:creator>Switzenberg, J. A.</dc:creator>
<dc:creator>Wang, J.</dc:creator>
<dc:creator>De Oliveira, P.</dc:creator>
<dc:creator>Kathuria, K.</dc:creator>
<dc:creator>Losh, S. J.</dc:creator>
<dc:creator>McDonald, T. L.</dc:creator>
<dc:creator>Bessell, B.</dc:creator>
<dc:creator>Van Deynze, K.</dc:creator>
<dc:creator>McConnell, M. J.</dc:creator>
<dc:creator>Boyle, A. P.</dc:creator>
<dc:creator>Mills, R. E.</dc:creator>
<dc:date>2024-12-21</dc:date>
<dc:identifier>doi:10.1101/2024.12.18.629274</dc:identifier>
<dc:title><![CDATA[A personalized multi-platform assessment of somatic mosaicism in the human frontal cortex]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory Press</dc:publisher>
<prism:publicationDate>2024-12-21</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://biorxiv.org/cgi/content/short/2024.11.07.619809v1?rss=1">
<title>
<![CDATA[
Image-based DNA Sequencing Encoding for Detecting Low-Mosaicism Somatic Mobile Element Insertions 
]]>
</title>
<link>
https://biorxiv.org/cgi/content/short/2024.11.07.619809v1?rss=1"
</link>
<description><![CDATA[
Active LINE-1 (L1), Alu, and SVA mobile elements in the human genome are capable of retrotransposition, resulting in novel mobile element insertions (MEIs) in both germline and somatic tissues. Detecting MEIs through DNA sequencing relies on supporting reads overlapping MEI junctions; however, artifacts from DNA amplification, sequencing, and alignment errors produce numerous false positives. Systematic detection of somatic MEIs, particularly those with low mosaicism, remains a significant challenge. Previous methods had required a high number of supporting reads which limits the detection sensitivity, or human inspections that are susceptible to biases. Here, we developed RetroNet, an algorithm that encodes MEI-supporting sequencing reads into images, and employs a deep neural network to identify somatic MEIs with as few as two reads. Trained on extensive and diverse datasets and benchmarked across various conditions, RetroNet surpasses previous methods and eliminates the need for extensive manual examinations. The RetroNet analysis on the Illumina sequencing of 161x or 195x of a cancer cell line achieved an average precision of 0.885 and recall of 0.579 for detecting somatic L1 insertions that are present in as few as 1.79% of the cells. Additionally, we demonstrated that RetroNet is effective for analyzing highly degraded DNA, such as circulating tumor DNA. RetroNet is applicable to the rapidly generated short-read sequencing data and has the potential to provide further insights into the functional and pathological implications of somatic retrotranspositions.
]]></description>
<dc:creator>Tan, M.</dc:creator>
<dc:creator>Lin, Z.</dc:creator>
<dc:creator>Chen, Z.</dc:creator>
<dc:creator>Park, J.</dc:creator>
<dc:creator>He, Z.</dc:creator>
<dc:creator>Zhou, H.</dc:creator>
<dc:creator>Lee, E. A.</dc:creator>
<dc:creator>Gao, Z.</dc:creator>
<dc:creator>Zhu, X.</dc:creator>
<dc:date>2024-11-08</dc:date>
<dc:identifier>doi:10.1101/2024.11.07.619809</dc:identifier>
<dc:title><![CDATA[Image-based DNA Sequencing Encoding for Detecting Low-Mosaicism Somatic Mobile Element Insertions]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory Press</dc:publisher>
<prism:publicationDate>2024-11-08</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://biorxiv.org/cgi/content/short/2024.11.06.622310v1?rss=1">
<title>
<![CDATA[
Deaminase-assisted single-molecule and single-cell chromatin fiber sequencing 
]]>
</title>
<link>
https://biorxiv.org/cgi/content/short/2024.11.06.622310v1?rss=1"
</link>
<description><![CDATA[
Gene regulation is mediated by the co-occupancy of numerous proteins along individual chromatin fibers. However, our tools for deeply profiling how proteins co-occupy individual fibers, especially at the single-cell level, remain limited. We present Deaminase-Assisted single-molecule chromatin Fiber sequencing (DAF-seq), which leverages a non-specific double-stranded DNA deaminase toxin A (SsDddA) to efficiently stencil protein occupancy along DNA molecules via selective deamination of accessible cytidines, which are preserved via C-to-T transitions upon DNA amplification. We demonstrate that DAF-seq enables [~]200,000-fold enrichment of target loci for single-molecule footprinting at near single-nucleotide resolution, enabling the precise delineation of the regulatory logic guiding neighboring proteins to cooperatively occupy chromatin fibers. Furthermore, DAF-seq enables the synchronous identification of single-molecule chromatin and genetic architectures - resolving the functional impact of rare somatic variants, as well as transitional chromatin states guiding haplotype-selective promoter actuation. Finally, we demonstrate that single-cell DAF-seq enables the accurate reconstruction of the diploid genome and epigenome from a single cell, revealing that a cells accessible regulatory landscape can diverge by as much as 63% while still retaining the cells identity. Overall, DAF-seq enables the comprehensive characterization of protein occupancy and chromatin accessibility across entire chromosomes with single-nucleotide, single-molecule, single-haplotype, and single-cell precision.
]]></description>
<dc:creator>Swanson, E. G.</dc:creator>
<dc:creator>Mao, Y.</dc:creator>
<dc:creator>Mallory, B. J.</dc:creator>
<dc:creator>Vollger, M. R.</dc:creator>
<dc:creator>Ranchalis, J.</dc:creator>
<dc:creator>Bohaczuk, S. C.</dc:creator>
<dc:creator>Parmalee, N. L.</dc:creator>
<dc:creator>Bennett, J. T.</dc:creator>
<dc:creator>Stergachis, A. B.</dc:creator>
<dc:date>2024-11-06</dc:date>
<dc:identifier>doi:10.1101/2024.11.06.622310</dc:identifier>
<dc:title><![CDATA[Deaminase-assisted single-molecule and single-cell chromatin fiber sequencing]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory Press</dc:publisher>
<prism:publicationDate>2024-11-06</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://biorxiv.org/cgi/content/short/2023.10.09.561356v1?rss=1">
<title>
<![CDATA[
Characterization of Cancer Evolution Landscape Based on Accurate Detection of Somatic Mutations in Single Tumor Cells 
]]>
</title>
<link>
https://biorxiv.org/cgi/content/short/2023.10.09.561356v1?rss=1"
</link>
<description><![CDATA[
Accurate detection of somatic mutations in single tumor cells is greatly desired as it allows us to quantify the single-cell mutation burden and construct the mutation-based phylogenetic tree. Here we developed scNanoSeq chemistry and profiled 842 single cells from 21 human breast cancer samples. The majority of the mutation-based phylogenetic trees comprise a characteristic stem evolution followed by the clonal sweep. We observed the subtype-dependent lengths in the stem evolution. To explain this phenomenon, we propose that the differences are related to different reprogramming required for different subtypes of breast cancer. Furthermore, we reason that the time that the tumor-initiating cell took to acquire the critical clonal-sweep-initiating mutation by random chance set the time limit for the reprogramming process. We refer to this model as a reprogramming and critical mutation co-timing (RCMC) subtype model. Next, in the sweeping clone, we observed that tumor cells undergo a branched evolution with rapidly decreasing selection. In the most recent clades, effectively neutral evolution has been reached, resulting in a substantially large number of mutational heterogeneities. Integrative analysis with 522-713X ultra-deep bulk whole genome sequencing (WGS) further validated this evolution mode. Mutation-based phylogenetic trees also allow us to identify the early branched cells in a few samples, whose phylogenetic trees support the gradual evolution of copy number variations (CNVs). Overall, the development of scNanoSeq allows us to unveil novel insights into breast cancer evolution.
]]></description>
<dc:creator>Niu, M.</dc:creator>
<dc:creator>Zhang, Y.</dc:creator>
<dc:creator>Luo, J.</dc:creator>
<dc:creator>Sinson, J. C.</dc:creator>
<dc:creator>Thompson, A. M.</dc:creator>
<dc:creator>Zong, C.</dc:creator>
<dc:date>2023-10-09</dc:date>
<dc:identifier>doi:10.1101/2023.10.09.561356</dc:identifier>
<dc:title><![CDATA[Characterization of Cancer Evolution Landscape Based on Accurate Detection of Somatic Mutations in Single Tumor Cells]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory Press</dc:publisher>
<prism:publicationDate>2023-10-09</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://biorxiv.org/cgi/content/short/2025.10.07.680917v1?rss=1">
<title>
<![CDATA[
Multi-platform framework for mapping somatic retrotransposition in human tissues 
]]>
</title>
<link>
https://biorxiv.org/cgi/content/short/2025.10.07.680917v1?rss=1"
</link>
<description><![CDATA[
Mobile element insertions (MEI) shape the human genome in both germline and somatic tissues. While inherited MEIs are well characterized, mapping somatic MEIs (sMEI) in non-cancer tissues remains challenging due to their low allelic fraction and repetitive nature. We established an integrative framework for sMEI analysis leveraging modern sequencing technologies and analytical innovations. We first benchmarked sMEI detection and demonstrated advantages of long-read and MEI-targeted sequencing for ultra-low-frequency events using a mixture of well-established cell lines. We then showed that haplotype phasing and donor-specific assemblies refine sMEI detection, effectively distinguishing from germline and false signals in in-silico tumor-normal mixtures. We further developed a source-tracing strategy based on internal sequence variation, expanding the catalogue of active source elements beyond traditional transduction-based methods. Applying this framework to donor tissues, we identified 18 rare somatic L1 insertions, revealing structural and source diversity. Our work provides a foundational framework and biological insight into sMEIs.
]]></description>
<dc:creator>Wang, S.</dc:creator>
<dc:creator>Bae, M.</dc:creator>
<dc:creator>Wang, J.</dc:creator>
<dc:creator>Zhao, B.</dc:creator>
<dc:creator>Nguyen, K.</dc:creator>
<dc:creator>Mallett, S.</dc:creator>
<dc:creator>Switzenberg, J. A.</dc:creator>
<dc:creator>Losh, S. J.</dc:creator>
<dc:creator>Sexton, C. E.</dc:creator>
<dc:creator>Miao, B.</dc:creator>
<dc:creator>Dong, S.</dc:creator>
<dc:creator>Zeng, X.</dc:creator>
<dc:creator>Wang, Z.</dc:creator>
<dc:creator>McDonald, T. L.</dc:creator>
<dc:creator>Mumm, C.</dc:creator>
<dc:creator>Gadde, R. K.</dc:creator>
<dc:creator>Tariq, A. M.</dc:creator>
<dc:creator>Chen, Z.</dc:creator>
<dc:creator>Feng, W. C.</dc:creator>
<dc:creator>Burn, A.</dc:creator>
<dc:creator>Park, J.</dc:creator>
<dc:creator>Chu, C.</dc:creator>
<dc:creator>Shen, H.</dc:creator>
<dc:creator>Wang, T.</dc:creator>
<dc:creator>Urban, A. E.</dc:creator>
<dc:creator>Zhu, X.</dc:creator>
<dc:creator>Li, H.</dc:creator>
<dc:creator>Burns, K. H.</dc:creator>
<dc:creator>Chun, H.-J. E.</dc:creator>
<dc:creator>Park, P. J.</dc:creator>
<dc:creator>SMaHT MEI Working Group,</dc:creator>
<dc:creator>Boyle, A. P.</dc:creator>
<dc:creator>Mills, R. E.</dc:creator>
<dc:creator>Zhou, W.</dc:creator>
<dc:creator>Lee, E. A.</dc:creator>
<dc:date>2025-10-07</dc:date>
<dc:identifier>doi:10.1101/2025.10.07.680917</dc:identifier>
<dc:title><![CDATA[Multi-platform framework for mapping somatic retrotransposition in human tissues]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory Press</dc:publisher>
<prism:publicationDate>2025-10-07</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://biorxiv.org/cgi/content/short/2025.10.09.678885v1?rss=1">
<title>
<![CDATA[
Comprehensive benchmarking of somatic mutation detection by the SMaHT Network 
]]>
</title>
<link>
https://biorxiv.org/cgi/content/short/2025.10.09.678885v1?rss=1"
</link>
<description><![CDATA[
Somatic mosaicism is increasingly recognized as a fundamental feature of human biology, yet the detection of somatic mutations remains challenging. The SMaHT Network conducted four large-scale benchmarking experiments involving cell-lines and donor tissues, to evaluate sequencing technologies, experimental approaches, and computational methods for detecting different types of somatic mutations, generating community resource with >1,000x short-read and 100-400x long-read data for each of the nine analyzed samples. We determined effective strategies for utilizing short- and long-reads sequencing for mutation detection and demonstrated that using donor-specific assemblies and human pangenome improved calling, extending mutation catalogs to challenging genomic regions. We benchmarked six duplex technologies and showed that single-cell sequencing resolves cell type-specific mutational patterns and heterogeneity. Our results indicate that bulk, single-cell, and duplex analyses are complementary - and leveraging all three provides comprehensive characterization of mosaicism within tissues. Together, these findings provide a roadmap for accurate, genome-wide somatic mutation discovery and analysis.
]]></description>
<dc:creator>The Somatic Mosaicism across Human Tissues Network (SMaHT),</dc:creator>
<dc:creator>Abyzov, A.</dc:creator>
<dc:date>2025-10-10</dc:date>
<dc:identifier>doi:10.1101/2025.10.09.678885</dc:identifier>
<dc:title><![CDATA[Comprehensive benchmarking of somatic mutation detection by the SMaHT Network]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory Press</dc:publisher>
<prism:publicationDate>2025-10-10</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://biorxiv.org/cgi/content/short/2025.10.10.681725v1?rss=1">
<title>
<![CDATA[
A telomere-to-telomere map of somatic mutation burden and functional impact in cancer 
]]>
</title>
<link>
https://biorxiv.org/cgi/content/short/2025.10.10.681725v1?rss=1"
</link>
<description><![CDATA[
Oncogenesis involves widespread genetic and epigenetic alterations, yet the full spectrum of somatic variation genome-wide remains unresolved. We generated a near-telomere-to-telomere (T2T) diploid assembly of a donor paired with deep short- and long-read sequencing of their melanoma. This revealed that 16% of somatic variants occur in sequences absent from GRCh38, with satellite repeats acting as hotspots for UV-induced damage due to sequence-intrinsic mutability and inefficient repair. Centromere kinetochore domains emerged as focal sites of structural, genetic, and epigenetic variation, leading to remodeling of centromere kinetochore binding domains during tumor evolution. Single-molecule telomere reconstructions uncovered cycles of attrition, deletion, and telomerase-mediated extension that shape cancer telomeres. Finally, diploid chromatin maps exposed that copy number alterations and epimutations, rather than point mutations, predominate in rewiring cancer regulatory programs. These findings define the full landscape of a cancers somatic variation and their functional impact, establishing a blueprint for T2T studies of mosaicism.
]]></description>
<dc:creator>Sohn, M.-H.</dc:creator>
<dc:creator>Dubocanin, D.</dc:creator>
<dc:creator>Vollger, M. R.</dc:creator>
<dc:creator>Kwon, Y.</dc:creator>
<dc:creator>Minkina, A.</dc:creator>
<dc:creator>Munson, K. M.</dc:creator>
<dc:creator>Hart, S. F.</dc:creator>
<dc:creator>Ranchalis, J. E.</dc:creator>
<dc:creator>Parmalee, N. L.</dc:creator>
<dc:creator>Sedeno-Cortes, A. E.</dc:creator>
<dc:creator>Ou, J.</dc:creator>
<dc:creator>Au, N. Y.</dc:creator>
<dc:creator>Bohaczuk, S.</dc:creator>
<dc:creator>Carroll, B.</dc:creator>
<dc:creator>Frazar, C. D.</dc:creator>
<dc:creator>Harvey, W. T.</dc:creator>
<dc:creator>Hoekzema, K.</dc:creator>
<dc:creator>Huang, M.-F.</dc:creator>
<dc:creator>Jacques, C. N.</dc:creator>
<dc:creator>Jensen, D. M.</dc:creator>
<dc:creator>Kolar, J. T.</dc:creator>
<dc:creator>Lee, R.</dc:creator>
<dc:creator>Lin, J.</dc:creator>
<dc:creator>Loy, K.</dc:creator>
<dc:creator>Mack, T.</dc:creator>
<dc:creator>Mao, Y.</dc:creator>
<dc:creator>Pham, M. M.</dc:creator>
<dc:creator>Ryke, E.</dc:creator>
<dc:creator>Smith, J. D.</dc:creator>
<dc:creator>Sutherlin, L.</dc:creator>
<dc:creator>Swanson, E. G.</dc:creator>
<dc:creator>Weiss, J. M.</dc:creator>
<dc:creator>SMaHT Assembly Working Group,</dc:creator>
<dc:creator>Carvalho, C.</dc:creator>
<dc:creator>Coorens, T. H.</dc:creator>
<dc:creator>Harris, K.</dc:creator>
<dc:creator>Wei, C.-L.</dc:creator>
<dc:creator>Eichler, E. E.</dc:creator>
<dc:creator>Altemose, N.</dc:creator>
<dc:creator>Bennett, J. T.</dc:creator>
<dc:creator>Stergachis, A. B.</dc:creator>
<dc:date>2025-10-13</dc:date>
<dc:identifier>doi:10.1101/2025.10.10.681725</dc:identifier>
<dc:title><![CDATA[A telomere-to-telomere map of somatic mutation burden and functional impact in cancer]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory Press</dc:publisher>
<prism:publicationDate>2025-10-13</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://biorxiv.org/cgi/content/short/2025.10.31.685648v1?rss=1">
<title>
<![CDATA[
A comprehensive view of somatic mosaicism by single-cell DNA analysis 
]]>
</title>
<link>
https://biorxiv.org/cgi/content/short/2025.10.31.685648v1?rss=1"
</link>
<description><![CDATA[
Single-cell DNA sequencing offers a powerful means of studying somatic mosaicism but requires careful analysis to mitigate DNA amplification-related artifacts. We performed primary template-directed amplification (PTA) and sequencing of 102 nuclei from postmortem lung and colon tissues of a 74-year-old male. Single-cell mutation burdens and spectra were validated by duplex sequencing and revealed heterogeneity across organs and cells, including signatures of APOBEC activity and tobacco exposure. Cells from both tissues exhibited chromosomal aneuploidies, loss of chromosome Y, and chromosomal rearrangements including rearrangements of the T-cell receptor loci indicative of T-cells. Shared embryonic mutations between cells enabled reconstruction of cellular ancestries from the zygote, which were validated by bulk sequencing. Collectively, we demonstrate a comprehensive approach for single-cell genomics that yields an expansive view of diverse somatic mutation types from development through aging across diverse tissues--insights that are obscured in bulk sequencing and only partially captured by other single-cell methods.
]]></description>
<dc:creator>Luquette, L. J.</dc:creator>
<dc:creator>Coorens, T. H. H.</dc:creator>
<dc:creator>Natu, A.</dc:creator>
<dc:creator>Suvakov, M.</dc:creator>
<dc:creator>Caplin, A.</dc:creator>
<dc:creator>Jun, M. S.</dc:creator>
<dc:creator>Mo, A.</dc:creator>
<dc:creator>Pelt, J.</dc:creator>
<dc:creator>Anderson, L.</dc:creator>
<dc:creator>Berselli, M.</dc:creator>
<dc:creator>Bhamidipati, S.</dc:creator>
<dc:creator>Blanchard, T.</dc:creator>
<dc:creator>Brew, J.</dc:creator>
<dc:creator>Chun, H.-J. E.</dc:creator>
<dc:creator>Chun, H.</dc:creator>
<dc:creator>Dehankar, M. K.</dc:creator>
<dc:creator>Feng, W. C.</dc:creator>
<dc:creator>Furatero, R.</dc:creator>
<dc:creator>Grochowski, C. M.</dc:creator>
<dc:creator>Ho, E.</dc:creator>
<dc:creator>Jang, Y.</dc:creator>
<dc:creator>Kottapalli, K.</dc:creator>
<dc:creator>Leonard, M. K.</dc:creator>
<dc:creator>Lim, N. S.</dc:creator>
<dc:creator>Lindsay, T.</dc:creator>
<dc:creator>Nicholson, S.</dc:creator>
<dc:creator>Raimondi, I.</dc:creator>
<dc:creator>Runnels, A.</dc:creator>
<dc:creator>Scharlee, C.</dc:creator>
<dc:creator>Shin, J.</dc:creator>
<dc:creator>Veit, A. D.</dc:creator>
<dc:creator>VonDran, M.</dc:creator>
<dc:creator>Wang, Y.</dc:creator>
<dc:creator>Yuan, D. J.</dc:creator>
<dc:creator>Zhao, Y.</dc:creator>
<dc:creator>Bell, T. J.</dc:creator>
<dc:creator>Ardlie, K.</dc:creator>
<dc:creator>Doddapaneni, H.</dc:creator>
<dc:creator>Fulton, R.</dc:creator>
<dc:creator>Germer, S.</dc:creator>
<dc:creator>Landau, D.</dc:creator>
<dc:creator>Oh, J. W.</dc:creator>
<dc:creator>Park, P. J.</dc:creator>
<dc:creator>Vaccarino, F. M.</dc:creator>
<dc:creator>Walsh, C. A.</dc:creator>
<dc:creator>Abyzov</dc:creator>
<dc:date>2025-11-03</dc:date>
<dc:identifier>doi:10.1101/2025.10.31.685648</dc:identifier>
<dc:title><![CDATA[A comprehensive view of somatic mosaicism by single-cell DNA analysis]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory Press</dc:publisher>
<prism:publicationDate>2025-11-03</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://biorxiv.org/cgi/content/short/2025.10.13.681545v1?rss=1">
<title>
<![CDATA[
Comprehensive benchmarking of somatic single-nucleotide variant and indel detection at ultra-low allele fractions using short- and long-read data 
]]>
</title>
<link>
https://biorxiv.org/cgi/content/short/2025.10.13.681545v1?rss=1"
</link>
<description><![CDATA[
Mosaic mutations in normal tissues occur at low variant allele fractions (VAFs), complicating detection. To benchmark strategies, the SMaHT Network created a cell-line mixture (1:49) and produced ultra-deep whole-genome sequencing using short and long reads (five centers, 180-500x each). We assembled a reference of 44,008 mosaic SNVs and 2,059 Indels, cross-validation between platforms to expose limits of short-read analysis. We also partitioned the genome by mappability to examine the impact of genomic context, added a negative reference set, and accounted for culture-derived mutations. When seven institutions applied eleven algorithms to mixture data, call sets were largely discordant across tools and replicates, partly reflecting stochastic presence of low-VAF mutations in biological replicants. For >2% VAF SNVs, sensitivity and precision approached [~]80% at [&ge;]300x, with little gain from additional sequencing. This work provides a comprehensive framework for reliable detection of low-VAF mutations in non-cancer tissues and a valuable resource for the community.
]]></description>
<dc:creator>Ha, Y.-J. J.</dc:creator>
<dc:creator>Maziec, D.</dc:creator>
<dc:creator>Markowski, J.</dc:creator>
<dc:creator>Georges, S. J.</dc:creator>
<dc:creator>Parmalee, N. L.</dc:creator>
<dc:creator>Berselli, M.</dc:creator>
<dc:creator>Coorens, T. H.</dc:creator>
<dc:creator>Dong, S.</dc:creator>
<dc:creator>Gardiner, S.</dc:creator>
<dc:creator>Kalra, D.</dc:creator>
<dc:creator>Li, D.</dc:creator>
<dc:creator>Miao, B.</dc:creator>
<dc:creator>Musunuri, R.</dc:creator>
<dc:creator>Xue, L.</dc:creator>
<dc:creator>Yu, Z.</dc:creator>
<dc:creator>Walker, K.</dc:creator>
<dc:creator>Anderson, L.</dc:creator>
<dc:creator>Au, N. Y.</dc:creator>
<dc:creator>Cibulskis, C.</dc:creator>
<dc:creator>Doddapaneni, H.</dc:creator>
<dc:creator>Grochowski, C. M.</dc:creator>
<dc:creator>Jensen, D. M.</dc:creator>
<dc:creator>Lindsay, T.</dc:creator>
<dc:creator>Loy, K.</dc:creator>
<dc:creator>Narayan, A.</dc:creator>
<dc:creator>Narzisi, G.</dc:creator>
<dc:creator>Ou, J.</dc:creator>
<dc:creator>Pham, M. M.</dc:creator>
<dc:creator>Runnels, A. M.</dc:creator>
<dc:creator>Stergachis, A. B.</dc:creator>
<dc:creator>Sutherlin, L. M.</dc:creator>
<dc:creator>Wang, T.</dc:creator>
<dc:creator>Jin, H.</dc:creator>
<dc:creator>Feng, W. C.</dc:creator>
<dc:creator>Zhang, Y.</dc:creator>
<dc:creator>Veit, A. D.</dc:creator>
<dc:creator>Kim, C. T.</dc:creator>
<dc:creator>Chun, H.-J. E.</dc:creator>
<dc:creator>Ardlie, K.</dc:creator>
<dc:creator>Fulton, R. S.</dc:creator>
<dc:creator>Germer, S.</dc:creator>
<dc:creator>Gibbs, R. A.</dc:creator>
<dc:creator>Marth, G. T.</dc:creator>
<dc:creator>Bennett, J. T.</dc:creator>
<dc:creator>Park, P. J.</dc:creator>
<dc:date>2025-10-14</dc:date>
<dc:identifier>doi:10.1101/2025.10.13.681545</dc:identifier>
<dc:title><![CDATA[Comprehensive benchmarking of somatic single-nucleotide variant and indel detection at ultra-low allele fractions using short- and long-read data]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory Press</dc:publisher>
<prism:publicationDate>2025-10-14</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://biorxiv.org/cgi/content/short/2025.09.18.677206v1?rss=1">
<title>
<![CDATA[
Comprehensive benchmarking of somatic structural variant detection at ultra-low allele fractions 
]]>
</title>
<link>
https://biorxiv.org/cgi/content/short/2025.09.18.677206v1?rss=1"
</link>
<description><![CDATA[
Postzygotic mosaicism gives rise to somatic structural variants (SVs) at ultra-low variant allele fractions (VAFs), which pose challenges for detection due to the high-coverage sequencing required and noise introduced by sequencing artifacts. Although somatic SV detection has been extensively studied in cancer, these studies are not directly applicable to the study of tissue mosaicism, as they rely on matched normals, target higher VAF ranges, and are enriched for different types of SVs.

We present comprehensive benchmark data and best practices for non-cancer somatic SV detection. We created a synthetic mosaic sample by combining six HapMap individuals at varying proportions, generating allele fractions as low as 0.25%. This sample was sequenced to [~]2,300x total coverage using Illumina, PacBio, and Nanopore technologies across multiple sequencing centers. A high-confidence benchmark SV set containing over 21,000 pseudo-somatic insertions and deletions [&ge;]50bp was derived from haplotype-resolved assemblies.

We evaluated 12 SV discovery pipelines and identified caller-specific strengths and sequencing platform-specific shortcomings. We find that short read-based approaches show reduced recall for insertions and repeat-associated SVs, whereas long-read sequencing achieves high accuracy throughout the genome, increasing linearly with coverage. The best algorithms sensitivity exceeded 80% for VAFs [&ge;]4% and 15% for VAFs of 0.5-1% with 60x coverage.

The publicly available benchmarking data and comparative analysis of current methods provide a foundation for robust discovery of SV mosaicism in non-cancer tissues..
]]></description>
<dc:creator>Zhang, Y.</dc:creator>
<dc:creator>English, A. C.</dc:creator>
<dc:creator>Paulin, L. F.</dc:creator>
<dc:creator>Grochowski, C. M.</dc:creator>
<dc:creator>Maheshwari, S.</dc:creator>
<dc:creator>Mack, T.</dc:creator>
<dc:creator>Berselli, M.</dc:creator>
<dc:creator>Veit, A. D.</dc:creator>
<dc:creator>Fu, Y.</dc:creator>
<dc:creator>SMAHT SV working group,</dc:creator>
<dc:creator>Park, P. J.</dc:creator>
<dc:creator>Sedlazeck, F. J.</dc:creator>
<dc:date>2025-09-20</dc:date>
<dc:identifier>doi:10.1101/2025.09.18.677206</dc:identifier>
<dc:title><![CDATA[Comprehensive benchmarking of somatic structural variant detection at ultra-low allele fractions]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory Press</dc:publisher>
<prism:publicationDate>2025-09-20</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://biorxiv.org/cgi/content/short/2025.09.29.679336v1?rss=1">
<title>
<![CDATA[
A Pangenomic Method for Establishing a Somatic Variant Detection Resource in HapMap Mixtures 
]]>
</title>
<link>
https://biorxiv.org/cgi/content/short/2025.09.29.679336v1?rss=1"
</link>
<description><![CDATA[
Somatic mosaicism is essential in human biology and disease, yet robust benchmarks are scarce. The SMaHT Consortium mixed six HapMap cell lines to create artificial somatic variants spanning 0.25% to 16.5% variant allele fractions. We developed a technology-agnostic method that builds pangenome graphs from individual assemblies to create unified benchmarking sets: > 6M single-nucleotide variants, 1.8M small insertions/deletions, 49K structural variations, and 10K mobile element insertions across autosomes, X, and mitochondrial chromosomes. We validated the variants using ultra-deep simulated reads and developed a binomial-based model to estimate coverage requirements for variant detection. Evaluating multiple callers showed CHM13 alignment improves structural variant detection and offers advantages in difficult-to-map regions compared to GRCh38. Systematic characterization showed regions with low detection rate are enriched in centromeres, satellite sequences, tandem repeats, and falsely duplicated genes. This accurate, versatile resource enables systematic evaluation of somatic variant detection technologies.
]]></description>
<dc:creator>Kong, N.</dc:creator>
<dc:creator>Tang, Z.</dc:creator>
<dc:creator>Ruttenberg, A.</dc:creator>
<dc:creator>Macias-Velasco, J. F.</dc:creator>
<dc:creator>Li, Z.</dc:creator>
<dc:creator>Zhang, W.</dc:creator>
<dc:creator>Miao, B.</dc:creator>
<dc:creator>Xin, Z.</dc:creator>
<dc:creator>Fu, Q.</dc:creator>
<dc:creator>Park, H.</dc:creator>
<dc:creator>Zhuo, X.</dc:creator>
<dc:creator>Mehinovic, E.</dc:creator>
<dc:creator>Belter, E.</dc:creator>
<dc:creator>Garza, J. E.</dc:creator>
<dc:creator>Dong, S.</dc:creator>
<dc:creator>Casey, E.</dc:creator>
<dc:creator>Johnson, B. K.</dc:creator>
<dc:creator>Majewski, M. F.</dc:creator>
<dc:creator>Palmer, T.</dc:creator>
<dc:creator>Cheng, Y.</dc:creator>
<dc:creator>Lindsay, T.</dc:creator>
<dc:creator>Schedl, T.</dc:creator>
<dc:creator>Li, D.</dc:creator>
<dc:creator>Shen, H.</dc:creator>
<dc:creator>SMaHT Network Assembly/Pangenome Working Group,</dc:creator>
<dc:creator>Fulton, R.</dc:creator>
<dc:creator>Wang, T.</dc:creator>
<dc:creator>Jin, S. C.</dc:creator>
<dc:date>2025-10-01</dc:date>
<dc:identifier>doi:10.1101/2025.09.29.679336</dc:identifier>
<dc:title><![CDATA[A Pangenomic Method for Establishing a Somatic Variant Detection Resource in HapMap Mixtures]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory Press</dc:publisher>
<prism:publicationDate>2025-10-01</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://biorxiv.org/cgi/content/short/2025.12.05.692678v1?rss=1">
<title>
<![CDATA[
Expanding the Genome in a Bottle Truth Set: Detection and Validation of Novel Low-frequency Variants Using High-accuracy NanoSeq 
]]>
</title>
<link>
https://biorxiv.org/cgi/content/short/2025.12.05.692678v1?rss=1"
</link>
<description><![CDATA[
HighlightsO_LINanoSeq-MBN achieves near-genome, Poisson-like coverage with minimal trinucleotide bias.
C_LIO_LIExpands the GIAB truth set by up to 160k de novo variants.
C_LIO_LIAdds a somatic layer to GIAB, enabling benchmarking and calibration of rare variants.
C_LIO_LIHigh-CADD exonic and splice variants highlight value for surveillance and clinical triage.
C_LI

Somatic mutations record tissue molecular history and inform risk, prognosis, and therapy, yet their variant allele fractions often fall below the reliable detection limit of conventional short-read sequencing. In contrast, duplex sequencing technology featured by NanoSeq applies the principle of single molecule detection and thereby overcomes the limitation. However, the original NanoSeq protocol relies on the restriction enzyme-based genome fragmentation, which constrained its genome coverage to 30-40%. To enable whole-genome discovery with duplex-level fidelity, we pursued two complementary approaches to optimize the NanoSeq protocol: (i) a restriction-enzyme strategy densifies accessible sites using orthogonal 4-bp cutters; and (ii) a workflow using sonication followed by mung bean nuclease with T4 polynucleotide kinase, Klenow fragment and dATP/ddBTP mixture (NanoSeq-MBN) to blunt and repair/A-tailing DNA, while minimizing repair artifacts. We systematically benchmarked their performance using Genome in a Bottle (GIAB) gold-standard sample mixtures. As a result, NanoSeq-MBN achieved near genome-wide, Poissonlike coverage with minimal trinucleotide-context bias and ultra-high accuracy. Beyond variants already present in the GIAB truth set, NanoSeq-MBN identified approximately 120,000-160,000 de novo mutations per sample missing in the truth set, Notably, over 98% had orthogonal support in reanalyzed GIAB bulk Illumina HiSeq libraries. These novel variants extended GIAB from germline benchmarking to rare-variant discovery and calibration of subclonal detection. Functional annotation revealed enrichment of high Combined Annotation Dependent Depletion (CADD) scores mutations in exonic and splice-related regions. Variants intersecting ClinVar entries and OMIM genes highlighted potential for surveillance and clinical triage. Collectively, these results add a somatic layer to GIAB, enabling calibration of burdens and mutational signatures in lymphoblastoid lines and provide reference material for rare-variant assays. The NanoSeq-MBN workflow offers a path to whole-genome, high-fidelity discovery of ultra-rare somatic variation with relevance to clinical assay validation.
]]></description>
<dc:creator>Zhang, Y.</dc:creator>
<dc:creator>Chao, H.</dc:creator>
<dc:creator>Niu, M.</dc:creator>
<dc:creator>Grochowski, C. M.</dc:creator>
<dc:creator>Kottapalli, K.</dc:creator>
<dc:creator>Bhamidipati, S. V.</dc:creator>
<dc:creator>muzny, d. m.</dc:creator>
<dc:creator>Gibbs, R. A.</dc:creator>
<dc:creator>Zong, C.</dc:creator>
<dc:creator>Doddapaneni, H. V.</dc:creator>
<dc:date>2025-12-06</dc:date>
<dc:identifier>doi:10.64898/2025.12.05.692678</dc:identifier>
<dc:title><![CDATA[Expanding the Genome in a Bottle Truth Set: Detection and Validation of Novel Low-frequency Variants Using High-accuracy NanoSeq]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory Press</dc:publisher>
<prism:publicationDate>2025-12-06</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://biorxiv.org/cgi/content/short/2025.12.03.692191v1?rss=1">
<title>
<![CDATA[
MosaicSim: A Novel Mosaic Variant Simulator Reveals Diminishing Returns of Ultra-High Coverage for Mosaic Variant Detection 
]]>
</title>
<link>
https://biorxiv.org/cgi/content/short/2025.12.03.692191v1?rss=1"
</link>
<description><![CDATA[
Genetic mutations within select cells of a tissue, termed mosaic variants (MV), are being increasingly recognized for their role in human disease. This growing interest underscores the need for specialized tools to detect and analyze MVs. However, such detection methods still lack thorough evaluation, largely due to missing benchmarking datasets that are large, reliable, and reflective of the complexity of biological samples. To address this gap, we developed MosaicSim, a tool for simulating variants in realistic sequencing data. The TweakVar workflow is at the tools core and represents a unique simulation pipeline that layers simulated MVs onto empirical whole genome sequencing data, generating a large, realistic ground truth dataset that combines the strengths of both simulation and biological data. To demonstrate the functionality of the workflow, we simulated 1,000 mosaic single nucleotide polymorphisms using TweakVar within whole genome sequencing files of different coverages. MVs were called with Illuminas DRAGEN and compared to the ground truth. Our results show 150x-445x coverage performed comparably, with a true-positive rate between 50.4% (300x) and 54.9% (150x) and no false-positives detected. Across all samples, increasing variant allele frequency had a significant positive effect on call success. Additionally, we observed that call rates for variants in lower complexity regions improved with increasing read depth. We did not find significant effects attributable to specific mutation patterns or mean read map quality. MosaicSim fills a critical unmet need by providing representative, customizable ground truth datasets for MV benchmarking, enabling systematic evaluation and optimization of variant calling methods.
]]></description>
<dc:creator>Stricker, E.</dc:creator>
<dc:creator>Jaryani, F.</dc:creator>
<dc:creator>Izydorczyk, M.</dc:creator>
<dc:creator>Poon, C.-L.</dc:creator>
<dc:creator>Sanio, P.</dc:creator>
<dc:creator>Alexander, A.</dc:creator>
<dc:creator>Deb, S.</dc:creator>
<dc:creator>Sedlazeck, F.</dc:creator>
<dc:creator>Rogers, J.</dc:creator>
<dc:creator>Atkinson, E. G.</dc:creator>
<dc:date>2025-12-07</dc:date>
<dc:identifier>doi:10.64898/2025.12.03.692191</dc:identifier>
<dc:title><![CDATA[MosaicSim: A Novel Mosaic Variant Simulator Reveals Diminishing Returns of Ultra-High Coverage for Mosaic Variant Detection]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory Press</dc:publisher>
<prism:publicationDate>2025-12-07</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://biorxiv.org/cgi/content/short/2025.10.28.685157v1?rss=1">
<title>
<![CDATA[
Single cell whole genome and transcriptome sequencing links somatic mutations to cell identity and ancestry 
]]>
</title>
<link>
https://biorxiv.org/cgi/content/short/2025.10.28.685157v1?rss=1"
</link>
<description><![CDATA[
The role of somatic mutations in human development and disease is obscured by difficulties in characterizing mutations at the single cell level and identifying cell types carrying them. Here we analysed somatic genomes of clonal iPSC lines and of single-cells after whole-genome amplification (scWGA) by PTA and ResolveOme from skin fibroblasts, blood and urine of a live donor. Mutation burden and spectra converged across approaches, revealing heterogeneous mutational footprints across cells driven by environmental exposures (UV damage and chemotherapy) and lymphocyte differentiation. Aneuploidies in single cells were detected by all the approaches and were orthogonally validated by Strand-seq. Uniquely, ResolveOme enabled cell-type identification using single-cell transcriptomes. Using a newly developed method accounting for noise and allele drop-out in scWGA, we de novo reconstructed the cell phylogenetic tree for this donor. Together, scWGA establishes a powerful foundation for comprehensive, cell type-aware, lineage-aware profiling of somatic mutations at single cell level.
]]></description>
<dc:creator>Natu, A.</dc:creator>
<dc:creator>Dehankar, M. K.</dc:creator>
<dc:creator>Pattni, R.</dc:creator>
<dc:creator>Suvakov, M.</dc:creator>
<dc:creator>Olisov, D.</dc:creator>
<dc:creator>Tomasini, L.</dc:creator>
<dc:creator>Jang, Y.</dc:creator>
<dc:creator>Huang, Y.</dc:creator>
<dc:creator>Benito-Garragori, E.</dc:creator>
<dc:creator>Hasenfeld, P.</dc:creator>
<dc:creator>Korbel, J.</dc:creator>
<dc:creator>Urban, A. E.</dc:creator>
<dc:creator>Abyzov, A.</dc:creator>
<dc:creator>Vaccarino, F. M.</dc:creator>
<dc:date>2025-10-29</dc:date>
<dc:identifier>doi:10.1101/2025.10.28.685157</dc:identifier>
<dc:title><![CDATA[Single cell whole genome and transcriptome sequencing links somatic mutations to cell identity and ancestry]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory Press</dc:publisher>
<prism:publicationDate>2025-10-29</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://biorxiv.org/cgi/content/short/2025.07.28.667228v1?rss=1">
<title>
<![CDATA[
ToxiTaRGET: a multi-omics resource for toxicant-responsive molecular targets 
]]>
</title>
<link>
https://biorxiv.org/cgi/content/short/2025.07.28.667228v1?rss=1"
</link>
<description><![CDATA[
Environmental toxicant exposures can induce widespread alterations in both the transcriptome and epigenome of mammals, and directly contribute to the increased risk of various diseases, including cardiovascular disorders, cancer, and neurological disorders. To evaluate how early-life toxicants produce long-term impacts on the transcriptome and epigenome in mice, the Toxicant Exposures and Responses by Genomic and Epigenomic Regulators of Transcription II (TaRGET II) Consortium generated a landmark resource comprising 3,607 multi-omics from longitudinal studies in mice. The molecular changes in responding to distinct environmental toxicants, including arsenic (As), lead (Pb), bisphenol A (BPA), tributyltin (TBT), di-2-ethylhexyl phthalate (DEHP), dioxin (TCDD), and fine particulate matter (PM2.5), were systematically identified and visualized on an integrative platform, ToxiTaRGET, to allow quickly search and browse by researchers. ToxiTaRGET houses a rich repository of molecular signatures, including gene expression, chromatin accessibility, and DNA methylation profiles, in response to early-life toxicant exposures. These molecular signatures span multiple biologically important tissues in both male and female mice at three distinct life stages, offering a valuable resource for the environmental health and toxicogenomic research communities.
]]></description>
<dc:creator>Kumar, R.</dc:creator>
<dc:creator>Fu, T.</dc:creator>
<dc:creator>Kuntala, P. K.</dc:creator>
<dc:creator>Fu, S.</dc:creator>
<dc:creator>Li, D.</dc:creator>
<dc:creator>Bartolomei, M. S.</dc:creator>
<dc:creator>Walker, C. L.</dc:creator>
<dc:creator>Wang, T.</dc:creator>
<dc:creator>Zhang, B. A.</dc:creator>
<dc:date>2025-08-02</dc:date>
<dc:identifier>doi:10.1101/2025.07.28.667228</dc:identifier>
<dc:title><![CDATA[ToxiTaRGET: a multi-omics resource for toxicant-responsive molecular targets]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory Press</dc:publisher>
<prism:publicationDate>2025-08-02</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://biorxiv.org/cgi/content/short/2026.04.13.718259v1?rss=1">
<title>
<![CDATA[
Long-read MitoScope reveals tissue-resolved somatic mitochondrial variation and landscape of nuclear-embedded mitochondrial sequences 
]]>
</title>
<link>
https://biorxiv.org/cgi/content/short/2026.04.13.718259v1?rss=1"
</link>
<description><![CDATA[
The mitochondrial genome (mtDNA), rich in repeats and prone to nuclear mitochondrial DNA segments (NUMTs), drives somatic mosaicism implicated in cancer, metabolic syndromes, and neurodegeneration, yet short-read sequencing yields incomplete catalogs, mapping artifacts, and false heteroplasmies. Here, we introduce MitoScope, a scalable long-read workflow to assemble mtDNA, perform high-fidelity variant calling, resolve heteroplasmy, and characterize NUMTs in benchmarking tissues from the Somatic Mosaicism Across Human Tissues (SMaHT) Network. MitoScope shows high sensitivity and precision, determines copy number, and uncovers low-frequency variants. We define an age- and tissue-dependent landscape of mtDNA mosaicism, including low-frequency pathogenic heteroplasmies, a bimodal heteroplasmy spectrum shaped by purifying selection, and age-accumulating deletions enriched for microhomology. Parallel profiling of NUMTs identifies high-confidence events with >2-fold more NUMTs than short-read surveys--with evidence of nonrandom trinucleotide contexts at breakpoints. These findings expose pervasive, tissue-resolved somatic mtDNA and NUMT instability with direct relevance for variant interpretation, aging, and human disease.
]]></description>
<dc:creator>Zakarian, C.</dc:creator>
<dc:creator>Smith, J. D.</dc:creator>
<dc:creator>Wong, C. H.</dc:creator>
<dc:creator>Frazar, C. D.</dc:creator>
<dc:creator>Ryke, E.</dc:creator>
<dc:creator>McGee, S. R.</dc:creator>
<dc:creator>Richardson, M.</dc:creator>
<dc:creator>Weiss, J. M.</dc:creator>
<dc:creator>Munson, K. M.</dc:creator>
<dc:creator>Hoekzema, K.</dc:creator>
<dc:creator>Mack, T.</dc:creator>
<dc:creator>Kwon, Y.</dc:creator>
<dc:creator>Ou, J.</dc:creator>
<dc:creator>Neph, S. J.</dc:creator>
<dc:creator>Sohn, M.-H.</dc:creator>
<dc:creator>Minkina, A.</dc:creator>
<dc:creator>Bennett, J. T.</dc:creator>
<dc:creator>Stergachis, A. B.</dc:creator>
<dc:creator>Eichler, E. E.</dc:creator>
<dc:creator>Wei, C.-L.</dc:creator>
<dc:date>2026-04-15</dc:date>
<dc:identifier>doi:10.64898/2026.04.13.718259</dc:identifier>
<dc:title><![CDATA[Long-read MitoScope reveals tissue-resolved somatic mitochondrial variation and landscape of nuclear-embedded mitochondrial sequences]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory Press</dc:publisher>
<prism:publicationDate>2026-04-15</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://biorxiv.org/cgi/content/short/2026.01.04.697580v1?rss=1">
<title>
<![CDATA[
Leveraging Human Pangenome for Improved Somatic Variant Detection 
]]>
</title>
<link>
https://biorxiv.org/cgi/content/short/2026.01.04.697580v1?rss=1"
</link>
<description><![CDATA[
Somatic variant detection is technically challenging due to low variant allele fractions, the confounding presence of germline variation, and reference bias. Linear references such as GRCh38 miss sample-specific variation, causing misalignments and incorrect variant calls. Although telomere-to-telomere donor-specific assemblies (DSAs) accurately represent individual genomes, their application is limited by cost and technical barriers. Alternatively, the graph-based human pangenome provides a scalable framework to improve read alignment and perform genome inference. Here, we benchmarked somatic variant detection using GRCh38, graph-based pangenomes, and pangenome-inferred DSAs with a HapMap mixture dataset and the COLO829 melanoma cell line. Pangenome-guided alignment improves read mapping and somatic variant calling accuracy. Furthermore, personalized pangenomes partially reconstruct donor-specific genomic content, improving accuracy, reducing germline contamination, and enabling detection of events in loci absent or poorly represented in GRCh38. These findings demonstrate that graph-based and personalized pangenomes are effective strategies for enhancing somatic variant detection compared with GRCh38.
]]></description>
<dc:creator>Fu, Q.</dc:creator>
<dc:creator>Xin, Z.</dc:creator>
<dc:creator>Miao, B.</dc:creator>
<dc:creator>Zhang, W.</dc:creator>
<dc:creator>Kong, N.</dc:creator>
<dc:creator>Tang, Z.</dc:creator>
<dc:creator>Ruttenberg, A.</dc:creator>
<dc:creator>Albracht, D.</dc:creator>
<dc:creator>Belter, E. A.</dc:creator>
<dc:creator>Garza, J. E.</dc:creator>
<dc:creator>Tomlinson, C.</dc:creator>
<dc:creator>Mehinovic, E.</dc:creator>
<dc:creator>Shen, J.</dc:creator>
<dc:creator>Zhuo, X.</dc:creator>
<dc:creator>Dong, S.</dc:creator>
<dc:creator>Johnson, B. K.</dc:creator>
<dc:creator>Majewski, M. F.</dc:creator>
<dc:creator>Palmer, T.</dc:creator>
<dc:creator>Jang, H. J.</dc:creator>
<dc:creator>Cheng, Y.</dc:creator>
<dc:creator>Li, Z.</dc:creator>
<dc:creator>Lawson, H. A.</dc:creator>
<dc:creator>Lindsay, T.</dc:creator>
<dc:creator>Li, D.</dc:creator>
<dc:creator>Fulton, R.</dc:creator>
<dc:creator>Shen, H.</dc:creator>
<dc:creator>Jin, S. C.</dc:creator>
<dc:creator>SMaHT Network Assembly/Pangenome Working Group,</dc:creator>
<dc:creator>Macias-Velasco, J. F.</dc:creator>
<dc:creator>Wang, T.</dc:creator>
<dc:date>2026-01-04</dc:date>
<dc:identifier>doi:10.64898/2026.01.04.697580</dc:identifier>
<dc:title><![CDATA[Leveraging Human Pangenome for Improved Somatic Variant Detection]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory Press</dc:publisher>
<prism:publicationDate>2026-01-04</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://biorxiv.org/cgi/content/short/2026.02.20.707061v1?rss=1">
<title>
<![CDATA[
Donor-specific assemblies enhance somatic structural variant detection in complex genomic regions 
]]>
</title>
<link>
https://biorxiv.org/cgi/content/short/2026.02.20.707061v1?rss=1"
</link>
<description><![CDATA[
Structural variants (SVs) contribute substantially to genomic variation and disease, but detecting somatic SVs (sSVs) remains difficult due to reference bias, mosaicism, and enrichment in repetitive regions. Linear reference genomes, like GRCh38 and CHM13, do not fully capture individual genomic structure, which can obscure true somatic variation. Donor-specific assemblies (DSAs) generated from the same genome where sSVs are being assayed provide a personalized alternative, yet their performance for sSV detection has not been systematically assessed. As part of the Somatic Mosaicism across Human Tissues (SMaHT) Network, we benchmark a DSA for sSV discovery in the COLO829 melanoma cell line with a matched normal sample from the same individual. We compare sSV detection across GRCh38, CHM13, and the COLO829BL_DSA using three different sSV callers (Delly, Severus, and Sniffles2) and sequence data from multiple long-read platforms. The COLO829BL_DSA identifies 1.8-fold more manually validated sSVs than linear references, in regions both shared with GRCh38 and CHM13 and unique to the COLO829BL_DSA. Variants detected only with the COLO829BL_DSA are often found in satellite and other repeat-rich regions that are difficult to resolve using standard references. In addition, several COLO829BL_DSA-specific sSVs are located in genes, some of which are associated with cancer. Overall, these results underscore the utility of DSAs in improving sSV detection.
]]></description>
<dc:creator>Mack, T. M.</dc:creator>
<dc:creator>Lin, J.</dc:creator>
<dc:creator>Ren, L.</dc:creator>
<dc:creator>Sohn, M.-H.</dc:creator>
<dc:creator>Minkina, A.</dc:creator>
<dc:creator>Kwon, Y.</dc:creator>
<dc:creator>Yoo, D.</dc:creator>
<dc:creator>Sui, Y.</dc:creator>
<dc:creator>Munson, K. M.</dc:creator>
<dc:creator>Hoekzema, K.</dc:creator>
<dc:creator>Mastrorosa, F. K.</dc:creator>
<dc:creator>Sorensen, M.</dc:creator>
<dc:creator>Ayllon, M.</dc:creator>
<dc:creator>Sun, K. A.</dc:creator>
<dc:creator>Koundiya, N.</dc:creator>
<dc:creator>Ou, J.</dc:creator>
<dc:creator>Noyes, M. D.</dc:creator>
<dc:creator>Sedeno-Cortes, A.</dc:creator>
<dc:creator>Leonardson, A.</dc:creator>
<dc:creator>Jacques, C. N.</dc:creator>
<dc:creator>Oliviera, C.</dc:creator>
<dc:creator>Frazar, C. D.</dc:creator>
<dc:creator>Zakarian, C.</dc:creator>
<dc:creator>Jensen, D. M.</dc:creator>
<dc:creator>Swanson, E. G.</dc:creator>
<dc:creator>Ryke, E.</dc:creator>
<dc:creator>Kolar, J. T.</dc:creator>
<dc:creator>Ranchalis, J.</dc:creator>
<dc:creator>Sutherlin, L.</dc:creator>
<dc:creator>Vollger, M. R.</dc:creator>
<dc:creator>Loy, K.</dc:creator>
<dc:creator>Pham, M. M.</dc:creator>
<dc:creator>Huang, M.-F.</dc:creator>
<dc:creator>Au, N. Y.</dc:creator>
<dc:creator>Nielsen, P. M.</dc:creator>
<dc:creator>McGee, S. R.</dc:creator>
<dc:creator>Neph, S.</dc:creator>
<dc:creator>Bohaczuk, S.</dc:creator>
<dc:creator>Shaffer, T.</dc:creator>
<dc:creator>Freeman, V.</dc:creator>
<dc:creator>Mao, Y.</dc:creator>
<dc:creator>Cohen Stillman, B.</dc:creator>
<dc:creator>Richardson, M.</dc:creator>
<dc:creator>Smith, J. D.</dc:creator>
<dc:creator>Weiss, J. M.</dc:creator>
<dc:date>2026-02-20</dc:date>
<dc:identifier>doi:10.64898/2026.02.20.707061</dc:identifier>
<dc:title><![CDATA[Donor-specific assemblies enhance somatic structural variant detection in complex genomic regions]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory Press</dc:publisher>
<prism:publicationDate>2026-02-20</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://biorxiv.org/cgi/content/short/2025.10.03.680364v1?rss=1">
<title>
<![CDATA[
Protamine lacunae preserve the paternal chromatin landscape in sperm 
]]>
</title>
<link>
https://biorxiv.org/cgi/content/short/2025.10.03.680364v1?rss=1"
</link>
<description><![CDATA[
The transmission of the paternal genome requires extensive chromatin reorganization, in which nucleosomes are largely replaced by protamines that drive extreme condensation of the genome in the sperm head. Using Fiber-seq, we resolve patterns of paternal chromatin repackaging in sperm at single-molecule resolution, revealing the interplay between protamination and nucleosome retention. We find that nucleosome retention is probabilistic, with no locus universally occupied. Although promoters of spermatogenic genes preferentially harbor retained nucleosomes, the predominant carrier of paternal epigenetic information is protamine lacunae, accessible discontinuities in the protamine coat that preferentially mark critical regulatory elements. By contrast, centromere kinetochore binding regions robustly retain CENP-A mono-nucleosomes, providing a mechanism for the focal transmission of paternal centromeres. Finally, we find that paternal chromatin repackaging is altered in low-motility sperm. Together, these findings reveal distinct modes of paternal chromatin epigenetic inheritance with broad implications for development and infertility.
]]></description>
<dc:creator>Tullius, T. W.</dc:creator>
<dc:creator>Heuer, R. A.</dc:creator>
<dc:creator>Bohaczuk, S. C.</dc:creator>
<dc:creator>Mallory, B.</dc:creator>
<dc:creator>Dubocanin, D.</dc:creator>
<dc:creator>Ranchalis, J.</dc:creator>
<dc:creator>Ayaz, A.</dc:creator>
<dc:creator>Mason, C.</dc:creator>
<dc:creator>Seli, E.</dc:creator>
<dc:creator>Phillippy, A. M.</dc:creator>
<dc:creator>Stergachis, A.</dc:creator>
<dc:creator>Lesch, B. J.</dc:creator>
<dc:date>2025-10-05</dc:date>
<dc:identifier>doi:10.1101/2025.10.03.680364</dc:identifier>
<dc:title><![CDATA[Protamine lacunae preserve the paternal chromatin landscape in sperm]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory Press</dc:publisher>
<prism:publicationDate>2025-10-05</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://biorxiv.org/cgi/content/short/2025.12.12.692823v1?rss=1">
<title>
<![CDATA[
Benchmarking of duplex sequencing approaches to reveal somatic mutation landscapes 
]]>
</title>
<link>
https://biorxiv.org/cgi/content/short/2025.12.12.692823v1?rss=1"
</link>
<description><![CDATA[
Detecting somatic mutations in normal tissues is challenging due to sequencing errors and the low allele fractions of post-zygotic variants. Duplex sequencing greatly reduces errors and can detect mutations at any allele fraction, but systematic, cross-platform comparisons are lacking. We present a comprehensive benchmarking of six duplex sequencing technologies used by the SMaHT Network: CODEC, CompDuplex-seq, HiDEF-seq, NanoSeq, ppmSeq, and VISTA-seq. We evaluated their performance using cord blood DNA, a tumor-normal cell line mixture, and homogenates from six human tissues. Each method shows distinct profiles in genomic footprint, sensitivity, and cost. Despite differences in library construction and sequencing platforms, estimates of mutation rates and mutational signatures are highly concordant. Integration with ultra-deep whole-genome sequencing shows that duplex approaches sensitively capture mutations and signatures beyond embryonic or clonally expanded variants. These results provide a foundation for selecting duplex methods and interpreting their data, enabling scalable single-molecule analyses of somatic mutation landscapes.

Graphical Abstract

O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=101 SRC="FIGDIR/small/692823v1_ufig1.gif" ALT="Figure 1">
View larger version (46K):
org.highwire.dtl.DTLVardef@123b221org.highwire.dtl.DTLVardef@83ba0aorg.highwire.dtl.DTLVardef@2b0ab0org.highwire.dtl.DTLVardef@1cae35e_HPS_FORMAT_FIGEXP  M_FIG C_FIG
]]></description>
<dc:creator>Zhang, Y.</dc:creator>
<dc:creator>Viswanadham, V.</dc:creator>
<dc:creator>Andreopoulos, M.</dc:creator>
<dc:creator>Glodzik, D.</dc:creator>
<dc:creator>Liu, R.</dc:creator>
<dc:creator>Luquette, L.</dc:creator>
<dc:creator>Jo, S.-Y.</dc:creator>
<dc:creator>Narayan, A.</dc:creator>
<dc:creator>Niu, M.</dc:creator>
<dc:creator>Anderson, L.</dc:creator>
<dc:creator>Brew, J.</dc:creator>
<dc:creator>Chao, H.</dc:creator>
<dc:creator>Cibulskis, C.</dc:creator>
<dc:creator>Dong, G.</dc:creator>
<dc:creator>Evani, U.</dc:creator>
<dc:creator>Feng, W.</dc:creator>
<dc:creator>Gronska-Peski, M.</dc:creator>
<dc:creator>Helland, A.</dc:creator>
<dc:creator>Hilal, N.</dc:creator>
<dc:creator>Jabara, N.</dc:creator>
<dc:creator>Jin, H.</dc:creator>
<dc:creator>Li, N.</dc:creator>
<dc:creator>Manam, M.</dc:creator>
<dc:creator>Mallett, S.</dc:creator>
<dc:creator>Runnels, A.</dc:creator>
<dc:creator>Scharlee, C.</dc:creator>
<dc:creator>Smith, C.</dc:creator>
<dc:creator>The SMaHT Duplex Sequencing Working Group,</dc:creator>
<dc:creator>Shao, D.</dc:creator>
<dc:creator>Walsh, C.</dc:creator>
<dc:creator>Adalsteinsson, V.</dc:creator>
<dc:creator>Germer, S.</dc:creator>
<dc:creator>Gibbs, R.</dc:creator>
<dc:creator>Choudhury, S.</dc:creator>
<dc:creator>Doddapaneni, H.</dc:creator>
<dc:creator>Evrony, G.</dc:creator>
<dc:creator>Zong, C.</dc:creator>
<dc:creator>Coorens, T.</dc:creator>
<dc:date>2025-12-15</dc:date>
<dc:identifier>doi:10.64898/2025.12.12.692823</dc:identifier>
<dc:title><![CDATA[Benchmarking of duplex sequencing approaches to reveal somatic mutation landscapes]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory Press</dc:publisher>
<prism:publicationDate>2025-12-15</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://biorxiv.org/cgi/content/short/2025.08.23.671919v1?rss=1">
<title>
<![CDATA[
As48, a First-in-Class Dual-Function TREM2 Modulator: Receptor Activation and Shedding Inhibition 
]]>
</title>
<link>
https://biorxiv.org/cgi/content/short/2025.08.23.671919v1?rss=1"
</link>
<description><![CDATA[
Triggering receptor expressed on myeloid cells 2 (TREM2) dysfunction contributes to Alzheimers disease pathogenesis, yet current therapeutics cannot prevent ADAM-mediated receptor shedding that diminishes signaling efficacy. Using Affinity Selection-Mass Spectrometry (AS-MS) screening, we identified As48, a novel small molecule that binds TREM2 with high affinity. Biophysical validation confirmed s 7-fold selectivity over TREM1. Cellular assays demonstrated that As48 functions as a TREM2 agonist, activating SYK phosphorylation and enhancing microglial phagocytosis. Molecular docking and molecular dynamics simulations revealed that As48 binds near the cleavage region, establishing hydrogen bonds with Gly68 and reducing conformational flexibility in regions 58-102. Based on this structural insight, we investigated the effect of As48 on TREM2 ectodomain shedding and discovered inhibition of receptor shedding without affecting ADAM10/17 protease activities, representing the first small molecule with anti-shedding properties through conformational restriction of protease accessibility. Importantly, As48 displayed favorable pharmacokinetics with potential for brain permeability, supporting its translational relevance. Through its dual and simultaneous promotion of receptor activation and prevention of shedding, As48 represents a paradigm shift in TREM2 modulation and neuroinflammatory drug discovery.

Abstract figure

O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=200 SRC="FIGDIR/small/671919v1_ufig1.gif" ALT="Figure 1">
View larger version (38K):
org.highwire.dtl.DTLVardef@1b67c28org.highwire.dtl.DTLVardef@1913034org.highwire.dtl.DTLVardef@f3bf58org.highwire.dtl.DTLVardef@97448a_HPS_FORMAT_FIGEXP  M_FIG C_FIG
]]></description>
<dc:creator>Cho, S.</dc:creator>
<dc:creator>El gaamouch, F.</dc:creator>
<dc:creator>Upadhyay, S.</dc:creator>
<dc:creator>Nada, H.</dc:creator>
<dc:creator>Kuncewicz, K.</dc:creator>
<dc:creator>Gabr, M.</dc:creator>
<dc:date>2025-08-28</dc:date>
<dc:identifier>doi:10.1101/2025.08.23.671919</dc:identifier>
<dc:title><![CDATA[As48, a First-in-Class Dual-Function TREM2 Modulator: Receptor Activation and Shedding Inhibition]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory Press</dc:publisher>
<prism:publicationDate>2025-08-28</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://biorxiv.org/cgi/content/short/2025.12.10.693405v1?rss=1">
<title>
<![CDATA[
Nanopore whole-genome sequencing reveals conserved chromosome-specific telomere architecture across tissues and populations 
]]>
</title>
<link>
https://biorxiv.org/cgi/content/short/2025.12.10.693405v1?rss=1"
</link>
<description><![CDATA[
Telomeres are repetitive nucleoprotein structures that cap the ends of linear chromosomes and are essential for maintaining genomic stability. While individual chromosome ends maintain distinct telomere lengths, the extent of this conservation across tissues and populations remains unclear due to the difficulty of analyzing repetitive telomeric sequences. Here, we show that high-coverage whole-genome Nanopore sequencing enables robust measurement of telomere length at the level of individual chromosomes. Nanopore reads yield reproducible telomere length estimates across replicates, in contrast to PacBio HiFi reads. Across > 250 individuals from 1000 Genomes and the SMaHT projects, chromosome-specific telomere length patterns are conserved across individuals and tissues, with tissues from the same individual showing highly similar patterns. This conserved landscape suggests coordinated regulation, whose disruption may contribute to genomic instability. Nanopore sequencing also allows simultaneous detection of structural variants, including disruption of TERT and NHP2 that drive global telomere shortening. Furthermore, our quantification of telomere variant repeats in positional context indicates active telomerase-mediated elongation. Our integrated profiling of telomere length and structural variation enables inference of variant effects on chromosome-specific telomere dynamics and may uncover risk factors for short telomere syndromes and cancer. Importantly, positionally fully resolved telomeric variant repeat patterns may predict activated telomere maintenance mechanisms with high accuracy.

SignificanceResolving chromosome-specific telomere length and variant-repeat architecture across tissues and individuals provides a framework to dissect coordinated telomere maintenance, its disruption by genetic variants, and how this shapes telomere mosaicism and disease risk.
]]></description>
<dc:creator>Engel, N. L.</dc:creator>
<dc:creator>Brors, B.</dc:creator>
<dc:creator>Feuerbach, L.</dc:creator>
<dc:creator>Park, P. J.</dc:creator>
<dc:date>2025-12-12</dc:date>
<dc:identifier>doi:10.64898/2025.12.10.693405</dc:identifier>
<dc:title><![CDATA[Nanopore whole-genome sequencing reveals conserved chromosome-specific telomere architecture across tissues and populations]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory Press</dc:publisher>
<prism:publicationDate>2025-12-12</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://biorxiv.org/cgi/content/short/2026.09.01.748636v1?rss=1">
<title>
<![CDATA[
Integrated map of somatic mosaicism across human tissues in 25 individuals 
]]>
</title>
<link>
https://biorxiv.org/cgi/content/short/2026.09.01.748636v1?rss=1"
</link>
<description><![CDATA[
Although all cells in the body descend from one genome, they accumulate distinct genetic and epigenetic changes over a lifetime, producing a mosaic of somatic variation that can shape development, aging, and disease. This mosaicism is often studied in isolation, leaving unclear how these forms relate within and between individuals. Here, we present the first integrated analysis of the Somatic Mosaicism across Human Tissues (SMaHT) Network's production resource, profiling up to 20 tissues from 25 donors using short- and long-read, duplex, single-cell, transcriptomic, and epigenomic sequencing, alongside donor-specific near-telomere-to-telomere assemblies. Somatic mutation burden cannot be captured by a single data type or metric, as tissues accumulate distinct variant classes largely independently of one another. Long-read and single-cell data resolved cell-type-specific mutational processes, traced mobile element insertions to source loci, and revealed the developmental timing and functional consequences of individual mutations. Donor-specific assemblies uncovered elevated mutation rates within centromeres and segmental duplications inaccessible to standard reference genomes, while haplotype-resolved chromatin and methylation data showed that nongenetically-deterministic epigenetic states are pervasive across tissues. Together, these findings provide an integrated, multi-scale portrait of somatic mosaicism across the human body, establishing a baseline against which its contributions to aging and disease can be measured.
]]></description>
<dc:creator>The Somatic Mosaicism across Human Tissues Network,</dc:creator>
<dc:creator>Sedlazeck, F. J.</dc:creator>
<dc:creator>Coorens, T. H. H.</dc:creator>
<dc:creator>Park, P. J.</dc:creator>
<dc:creator>Stergachis, A. B.</dc:creator>
<dc:date>2026-09-03</dc:date>
<dc:identifier>doi:10.64898/2026.09.01.748636</dc:identifier>
<dc:title><![CDATA[Integrated map of somatic mosaicism across human tissues in 25 individuals]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory Press</dc:publisher>
<prism:publicationDate>2026-09-03</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://biorxiv.org/cgi/content/short/2026.05.27.726245v1?rss=1">
<title>
<![CDATA[
Pansoma, a machine learning tool for identifying somatic variants using pangenome graphs 
]]>
</title>
<link>
https://biorxiv.org/cgi/content/short/2026.05.27.726245v1?rss=1"
</link>
<description><![CDATA[
Somatic variant calling, the identification of mutations in non-germline cells acquired over an individuals lifetime, is critical for studying diseases, including cancer, and for developing precision oncology strategies. Traditional somatic variant calling methods rely on linear reference genomes, which do not adequately capture human genetic diversity and result in reference bias, compromising the accuracy of somatic variant detection. Recently developed graph-based human pangenome reference represents diverse genetic variants across human populations and has promised to drive advances in many genetics and genomics studies. In this study, we introduced Pansoma, a novel pangenome-native and machine learning-based tool specifically designed for somatic variant calling using a pangenome graph reference. Pansoma performs somatic variant detection from both short- and long-read sequencing data by learning tensor representations of alignment on graph nodes rather than on a linear reference. Pansoma outputs variant representations anchored to the pangenome graph paths and conventional somatic variant calls remapped to the linear reference. Additionally, we provide accompanying bioinformatics tools tailored for graph-based genomic data management and variant calling results analysis. Benchmarking shows that Pansoma not only improves tumor-only somatic variant detection but also preserves graph-specific variant representations that are not directly recoverable from linear- reference outputs.
]]></description>
<dc:creator>Shen, J.</dc:creator>
<dc:creator>Fu, Q.</dc:creator>
<dc:creator>Macias, J. F.</dc:creator>
<dc:creator>Human Pangenome Reference Consortium,</dc:creator>
<dc:creator>Li, D.</dc:creator>
<dc:creator>Wang, T.</dc:creator>
<dc:date>2026-05-29</dc:date>
<dc:identifier>doi:10.64898/2026.05.27.726245</dc:identifier>
<dc:title><![CDATA[Pansoma, a machine learning tool for identifying somatic variants using pangenome graphs]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory Press</dc:publisher>
<prism:publicationDate>2026-05-29</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://biorxiv.org/cgi/content/short/2026.06.24.734333v1?rss=1">
<title>
<![CDATA[
Short-Read Sequencing Benchmarking with Donor-Specific Assemblies 
]]>
</title>
<link>
https://biorxiv.org/cgi/content/short/2026.06.24.734333v1?rss=1"
</link>
<description><![CDATA[
BackgroundHigh-throughput short-read sequencing has become a core technology for genomics, but the rapid expansion of available platforms has made it increasingly important to benchmark them under standardized conditions. A major challenge is that conventional reference-based comparisons confound true sequencing errors with inherited variation and reference bias, making it difficult to isolate platform-intrinsic performance.

ResultsWe benchmarked nine short-read chemistries across seven DNA sequencers using two highly characterized benchmark samples, HG002 and COLO829BL, together with donor-specific assemblies to measure sequencing errors against sample-matched genomic references. This strategy separated authentic platform errors from biological divergence and revealed substantial differences in substitution, indel, read-position, and sequence-context error profiles. Element AVITI UltraQ and Roche SBX-D showed the lowest substitution error rates, whereas Ultima and Roche chemistries exhibited the strongest indel-associated biases. We also found pronounced platform-specific effects in low-complexity regions and trinucleotide contexts, including homopolymer-associated errors and context-dependent substitution skews that are directly relevant to rare-variant detection. In addition, we show that donor-specific references are essential for unbiased base-quality recalibration because they minimize reference bias and more faithfully support cross-platform comparison and low-frequency variant-calling thresholds.

ConclusionsDonor-specific assembly-based benchmarking provides a robust framework for measuring true short-read sequencing errors and comparing platforms on a common, sample-matched basis. Our results establish a comprehensive reference for the community and show that authentic error profiles can guide platform selection, quality filtering, and improved detection of rare somatic variation.
]]></description>
<dc:creator>McGee, S. R.</dc:creator>
<dc:creator>Smith, J. D.</dc:creator>
<dc:creator>Frazar, C. D.</dc:creator>
<dc:creator>Ryke, E.</dc:creator>
<dc:creator>Vollger, M. R.</dc:creator>
<dc:creator>Kwon, Y.</dc:creator>
<dc:creator>Bennett, J. T.</dc:creator>
<dc:creator>Eichler, E. E.</dc:creator>
<dc:creator>Stergachis, A.</dc:creator>
<dc:creator>Wei, C.-L.</dc:creator>
<dc:date>2026-06-28</dc:date>
<dc:identifier>doi:10.64898/2026.06.24.734333</dc:identifier>
<dc:title><![CDATA[Short-Read Sequencing Benchmarking with Donor-Specific Assemblies]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory Press</dc:publisher>
<prism:publicationDate>2026-06-28</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://biorxiv.org/cgi/content/short/2026.06.10.731451v1?rss=1">
<title>
<![CDATA[
Somatic variant detection in normal tissues from single-cell sequencing data 
]]>
</title>
<link>
https://biorxiv.org/cgi/content/short/2026.06.10.731451v1?rss=1"
</link>
<description><![CDATA[
A crucial advantage of single-cell sequencing (SCS) is its ability to identify somatic variants in individual cells, enabling phylogenetic analysis of cellular populations within bulk tissues. While identifying somatic variants in tumor tissues via SCS has become a common practice, doing so in normal tissues remains challenging due to the rarity of somatic variants in normal cells. To evaluate the feasibility of somatic variant calling from widely available single-nucleus RNA-seq (snRNA-seq) and single-nucleus ATAC-seq (snATAC-seq) data, we profiled a Cell-line mix of six HapMap samples prepared by the SMaHT consortium using 10x Genomics 5 snRNA-seq (12k cells with 36k mean reads per cell) and snATAC-seq (11k cells with 14k median high-quality fragments per cell) for variant calling. PacBio long-read whole genome sequencing (WGS) data (109x) generated from individual cell lines were used as ground truth. Two computational tools, Monopogen and SComatic, were used for somatic variant calling from the SCS data. Monopogen achieved single nucleotide variant (SNV) detection accuracies of 93.30% in the snRNA-seq and 99.64% in the snATAC-seq data, both of which outperformed SComatic (74.35% and 94.29%, respectively). Monopogen also consistently detected somatic SNVs at cellular fractions as low as 0.5% (2.54% in snRNA and 0.81% in snATAC) in individual samples. Notably, snATAC-seq exhibited higher genomic coverage breadth and larger number of variants detected than snRNA-seq. While the SCS data have lower overall genome coverage than that of the bulk WGS, the single-cell level variant resolution allows Monopogen to assign variants to their cells of origin with over 80% accuracy in both RNA and ATAC modalities, thereby facilitating studies of clonal evolution and cell-type-specific mutagenesis. Other benchmarking methods were also evaluated (DeepVariant, Cellsnp-lite and Mutect2) for comparison. In conclusion, our study demonstrated the feasibility of performing reliable single-cell somatic mutation calling in a cell-line mixture and discussed the strengths and limitations of current computational methods when applied to normal tissues.
]]></description>
<dc:creator>Luo, R.</dc:creator>
<dc:creator>Wang, Z.</dc:creator>
<dc:creator>Dou, J.</dc:creator>
<dc:creator>Bhamidipati, S. V.</dc:creator>
<dc:creator>Kalra, D.</dc:creator>
<dc:creator>Grochowski, C. M.</dc:creator>
<dc:creator>Doddapaneni, H. V.</dc:creator>
<dc:creator>Gibbs, R. A.</dc:creator>
<dc:creator>Chen, K.</dc:creator>
<dc:creator>Chen, R.</dc:creator>
<dc:date>2026-06-14</dc:date>
<dc:identifier>doi:10.64898/2026.06.10.731451</dc:identifier>
<dc:title><![CDATA[Somatic variant detection in normal tissues from single-cell sequencing data]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory Press</dc:publisher>
<prism:publicationDate>2026-06-14</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://biorxiv.org/cgi/content/short/2026.03.16.712141v1?rss=1">
<title>
<![CDATA[
MosaicTR: tandem repeat somatic instability quantification from long-read sequencing 
]]>
</title>
<link>
https://biorxiv.org/cgi/content/short/2026.03.16.712141v1?rss=1"
</link>
<description><![CDATA[
SummarySomatic instability of tandem repeats modifies disease onset and progression in repeat expansion disorders and serves as a biomarker for mismatch repair deficiency in cancer. MosaicTR quantifies per-locus somatic instability from haplotype-tagged long-read sequencing data without the read-length and PCR stutter limitations of short-read approaches. A motif-unit-weighted metric reduces platform-specific sequencing noise on both PacBio HiFi and Oxford Nanopore data, and pairwise comparison modes support detection of tissue-specific or longitudinal instability changes.

AvailabilityMosaicTR is freely available at https://github.com/junsoopablo/mosaictr under the MIT license.

Contactjunsoopablo@snu.ac.kr
]]></description>
<dc:creator>Kim, J.</dc:creator>
<dc:date>2026-03-18</dc:date>
<dc:identifier>doi:10.64898/2026.03.16.712141</dc:identifier>
<dc:title><![CDATA[MosaicTR: tandem repeat somatic instability quantification from long-read sequencing]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory Press</dc:publisher>
<prism:publicationDate>2026-03-18</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://biorxiv.org/cgi/content/short/2025.11.03.686348v1?rss=1">
<title>
<![CDATA[
Himito: a Graph-based Toolkit for Mitochondrial Genome Analysis using Long Reads 
]]>
</title>
<link>
https://biorxiv.org/cgi/content/short/2025.11.03.686348v1?rss=1"
</link>
<description><![CDATA[
Understanding the genetic and epigenetic regulation of mitochondrial DNA (mtDNA) is essential for elucidating mechanisms of aging and disease. Long-read sequencing can span the entire mitochondrial genome and directly capture base-modification signals, yet analytical tools for such data remain limited. We developed Himito, a graph-based toolkit for analyzing mitochondrial genome using long reads. Himito filters reads originating from nuclear mitochondrial insertions (NUMTs), constructs a sequence graph to represent mtDNA diversity, assembles primary haplotypes, calls variants, and analyzes 5-methylcytosine (5mC) modifications within a unified framework. Benchmarking on high-quality reference datasets shows Himito achieves superior performance in assembly and variant calling compared with existing tools. Applied to the All of Us (AoU) v8 dataset, Himito identified pathogenic mtDNA variants, revealed population-scale haplogroup diversity, and uncovered age-related genetic and epigenetic patterns. These results demonstrate that long-read sequencing, combined with graph-based analysis, enables integrated characterization of mitochondrial genomic and epigenomic variation. Himito is available at https://github.com/broadinstitute/Himito.
]]></description>
<dc:creator>Su, H.</dc:creator>
<dc:creator>Huang, Y.</dc:creator>
<dc:creator>Durham, T.</dc:creator>
<dc:creator>Kong, N.</dc:creator>
<dc:creator>Casey,, E.</dc:creator>
<dc:creator>Benjamin, D.</dc:creator>
<dc:creator>Jin, S. C.</dc:creator>
<dc:creator>Garimella, K. V.</dc:creator>
<dc:date>2025-11-05</dc:date>
<dc:identifier>doi:10.1101/2025.11.03.686348</dc:identifier>
<dc:title><![CDATA[Himito: a Graph-based Toolkit for Mitochondrial Genome Analysis using Long Reads]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory Press</dc:publisher>
<prism:publicationDate>2025-11-05</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://biorxiv.org/cgi/content/short/2026.03.20.713111v1?rss=1">
<title>
<![CDATA[
LongcallD: joint calling and phasing of small, structural and mosaic variants from long reads 
]]>
</title>
<link>
https://biorxiv.org/cgi/content/short/2026.03.20.713111v1?rss=1"
</link>
<description><![CDATA[
Long-read sequencing is a powerful technique capturing multiple variants within single continuous reads. This length allows individual reads to bridge small and structural variants while carrying crucial phasing information. However, current computational tools treat small variant calling, structural variant (SV) detection and phasing as largely disconnected problems, failing to unleash the full potential of long reads. Here, we present longcallD, a unified framework utilizing local multiple-sequence alignment to simultaneously call and phase small and structural variants. By integrating germline phasing and retrotransposition hallmarks, longcallD also identifies low-fraction mosaic variants and detects mobile element insertions supported by a single read. Compared to existing methods, our unified approach substantially improves SV discovery and mosaic variants accuracy while maintaining competitive small variant calling. We anticipate that longcallD will provide a robust foundation for resolving complex genetic architectures in clinical and evolutionary applications.
]]></description>
<dc:creator>Gao, Y.</dc:creator>
<dc:creator>Liao, W.-W.</dc:creator>
<dc:creator>Qin, Q.</dc:creator>
<dc:creator>Hall, I. M.</dc:creator>
<dc:creator>Li, H.</dc:creator>
<dc:date>2026-03-22</dc:date>
<dc:identifier>doi:10.64898/2026.03.20.713111</dc:identifier>
<dc:title><![CDATA[LongcallD: joint calling and phasing of small, structural and mosaic variants from long reads]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory Press</dc:publisher>
<prism:publicationDate>2026-03-22</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://biorxiv.org/cgi/content/short/2025.10.17.683157v1?rss=1">
<title>
<![CDATA[
Timing the onset of homologous recombination deficiency before cancer diagnosis 
]]>
</title>
<link>
https://biorxiv.org/cgi/content/short/2025.10.17.683157v1?rss=1"
</link>
<description><![CDATA[
Mutations in BRCA1 and BRCA2 genes, whether inherited or somatically acquired, cause homologous recombination deficiency (HRD) in tumor cells. The timing of HRD onset in the emerging tumor lineage is unknown. Here, we present HRDTimer, an algorithm to infer the onset of HRD-driven mutagenesis prior to cancer diagnosis. We estimate that HRD arises at 34% of SBS1-based molecular time---corresponding to a median of 8.3 years (IQR 7.1--10.4) prior to diagnosis in triple-negative breast cancers, and 15.0 years (IQR 12.0--20.6) in ER-positive breast cancers. Bulk sequencing reveals accelerated SBS1 accumulation following neoplastic transformation compared to normal tissue, influencing the estimated age of HRD onset. Single-cell duplex sequencing confirms SBS1 acceleration in tumors and further shows that non-tumor cells largely lack the HRD signature, indicating that HRD is rare in pre-malignant cells, even in BRCA1/2 mutation carriers. Together, our analysis pinpoints the onset of HRD before diagnosis, defining a window for detection and potential interception.
]]></description>
<dc:creator>Andreopoulos, M.</dc:creator>
<dc:creator>Niu, M.</dc:creator>
<dc:creator>Zhang, Y.</dc:creator>
<dc:creator>Viswanadham, V. V.</dc:creator>
<dc:creator>Gulhan, D. C.</dc:creator>
<dc:creator>Jin, H.</dc:creator>
<dc:creator>Batalini, F.</dc:creator>
<dc:creator>Wulf, G.</dc:creator>
<dc:creator>Zong, C.</dc:creator>
<dc:creator>Park, P. J.</dc:creator>
<dc:creator>Glodzik, D.</dc:creator>
<dc:date>2025-10-19</dc:date>
<dc:identifier>doi:10.1101/2025.10.17.683157</dc:identifier>
<dc:title><![CDATA[Timing the onset of homologous recombination deficiency before cancer diagnosis]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory Press</dc:publisher>
<prism:publicationDate>2025-10-19</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://biorxiv.org/cgi/content/short/2026.01.20.700228v1?rss=1">
<title>
<![CDATA[
SigFormer: an Attention-Based Framework for Robust Single-Sample Mutational Signature Decomposition 
]]>
</title>
<link>
https://biorxiv.org/cgi/content/short/2026.01.20.700228v1?rss=1"
</link>
<description><![CDATA[
Somatic mutational signatures imprint the history of exogenous exposures and endogenous processes on the genome, offering critical insights into pathologic etiology and disease risk. However, accurate signature decomposition at the single-sample level is still challenging when mutation burden is low, sampling noise is high, and candidate catalogs are large and redundant. Here, we present SigFormer, a set-conditioned transformer framework designed to facilitate robust somatic mutation analysis without reliance on large cohorts. By leveraging a cross-attention mechanism between customized reference input and sample mutation profile, SigFormer improves exposure recovery and detection accuracy compared with likelihood-driven refitting (MuSiCal) with the largest performance gains in high-noise and overcomplete settings. On PCAWG genomes, SigFormer preserves major tissue-level structure while sensitively and accurately capturing cooccurrence of low-abundance signatures but without the need for tumor-type-specific gating. In low-burden normal-tissue datasets spanning clonal expansion and microdissection studies, SigFormer maintains the high accuracy and recovers stable tissue-dependent patterns of SBS1/SBS5/SBS40a, pointing to underlying tissue-specific mutagenic heterogeneity in normal tissues. Finally, SigFormer quantifies an explicit unattributable residual component when the catalogue is incomplete, preventing forced allocation into flexible flat signatures and providing a useful signal for downstream analyses.
]]></description>
<dc:creator>Zhang, Y.</dc:creator>
<dc:creator>Niu, M.</dc:creator>
<dc:creator>Zong, C.</dc:creator>
<dc:date>2026-01-21</dc:date>
<dc:identifier>doi:10.64898/2026.01.20.700228</dc:identifier>
<dc:title><![CDATA[SigFormer: an Attention-Based Framework for Robust Single-Sample Mutational Signature Decomposition]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory Press</dc:publisher>
<prism:publicationDate>2026-01-21</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://biorxiv.org/cgi/content/short/2026.08.05.743125v1?rss=1">
<title>
<![CDATA[
Pangenome discovery and characterization of human protein-coding duplicated genes 
]]>
</title>
<link>
https://biorxiv.org/cgi/content/short/2026.08.05.743125v1?rss=1"
</link>
<description><![CDATA[
Protein-coding genes mapping to high-identity segmental duplications (SDs) have been difficult to annotate and characterize and are the source of most previously unknown protein-coding genes being discovered as part of the human pangenome. Here, we combine long-read assembled human genomes (298) and long-read transcriptome data (5.6 billion full-length cDNA from 83 tissues) to phylogenetically interrogate 493 gene families discovering 2713 potentially copy number polymorphic genes not present in the human reference genome. For reference SD gene families where paralog specificity can be assigned, we find that 60.0% are expressed and maintain open reading frames, with 45.7% showing high expression in brain, embryo, or testis. We revise 386 gene models, including 150 that absent or different from current T2T-CHM13 gene annotation and 236 (35.1%) pseudogenes as protein-coding where we find evidence of transcription, an open reading frame, and chromatin-accessible promoters. We find that 24.2% of SD genes show evidence of constraint for both copy number and amino acid mutation. The majority of these constraint genes are ancestral, whereas only 16.2% of derived duplicated genes that emerged recently in the human lineage show evidence of constraint. The pangenome provides unparalleled specificity to understand genetic variation in SD genes allowing us to distinguish functional genes from pseudogenes and highlighting potential gene innovations that arose most recently in human evolution.
]]></description>
<dc:creator>Ren, L.</dc:creator>
<dc:creator>Yoo, D.</dc:creator>
<dc:creator>Vlajic, K.</dc:creator>
<dc:creator>Dishuck, P. C.</dc:creator>
<dc:creator>Guitart, X.</dc:creator>
<dc:creator>Kwon, Y.</dc:creator>
<dc:creator>Lin, J.</dc:creator>
<dc:creator>Munson, K. M.</dc:creator>
<dc:creator>Hoekzema, K.</dc:creator>
<dc:creator>Stergachis, A.</dc:creator>
<dc:creator>Vollger, M. R.</dc:creator>
<dc:creator>Schweppe, D. K.</dc:creator>
<dc:creator>Eichler, E. E.</dc:creator>
<dc:date>2026-08-06</dc:date>
<dc:identifier>doi:10.64898/2026.08.05.743125</dc:identifier>
<dc:title><![CDATA[Pangenome discovery and characterization of human protein-coding duplicated genes]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory Press</dc:publisher>
<prism:publicationDate>2026-08-06</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://biorxiv.org/cgi/content/short/2025.09.14.676124v1?rss=1">
<title>
<![CDATA[
SomaMutDB 2.0: A comprehensive database for exploring somatic mutations and their functional impact in normal human tissues 
]]>
</title>
<link>
https://biorxiv.org/cgi/content/short/2025.09.14.676124v1?rss=1"
</link>
<description><![CDATA[
Recent advances in ultra-accurate sequencing technologies have revealed that somatic mutations accumulate across the human lifespan and may contribute to both normal aging and disease. These mutations are highly diverse, often non-recurrent, and functionally heterogeneous, which makes their biological impact difficult to evaluate systematically. Although many studies have profiled somatic mutations in individual tissues or limited cohorts, a centralized and scalable platform that integrates discoveries and supports functional interpretation has been lacking. To address this gap, we present SomaMutDB 2.0 (https://somamutdb.org/SomaMutDB/), a substantially expanded database that catalogs 8.9 million mutations (8.57 million SNVs and 0.29 million INDELs) from 10,852 samples of 607 human subjects across 47 studies. Beyond expanded data coverage, SomaMutDB 2.0 introduces a comprehensive functional annotation framework that applies 22 predictive models, encompassing coding, regulatory, expression-based, and ensemble predictors, to systematically assess mutational impact. Users can browse pre-annotated variants through an interactive interface or upload their own variants for real-time analysis, with results contextualized against all mutations from normal, non-diseased tissues in the database. Together, these advances establish SomaMutDB 2.0 as the most comprehensive resource currently available for characterizing somatic mosaicism and functional interpretation in human health and aging.

Graphical AbstractSomaMutDB 2.0 provides an expanded catalog of 8.9 million somatic mutations across 30 tissues, along with pipelines for mutational signature analysis and 22-tool functional annotation that enable user-submitted variant interpretation.



O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=113 SRC="FIGDIR/small/676124v1_ufig1.gif" ALT="Figure 1">
View larger version (39K):
org.highwire.dtl.DTLVardef@1c76484org.highwire.dtl.DTLVardef@1984f71org.highwire.dtl.DTLVardef@8789a9org.highwire.dtl.DTLVardef@5ed54a_HPS_FORMAT_FIGEXP  M_FIG C_FIG
]]></description>
<dc:creator>Shea, A.</dc:creator>
<dc:creator>Sun, S.</dc:creator>
<dc:creator>Kennedy, J.</dc:creator>
<dc:creator>Zhang, L.</dc:creator>
<dc:creator>Vijg, J.</dc:creator>
<dc:creator>Dong, X.</dc:creator>
<dc:date>2025-09-17</dc:date>
<dc:identifier>doi:10.1101/2025.09.14.676124</dc:identifier>
<dc:title><![CDATA[SomaMutDB 2.0: A comprehensive database for exploring somatic mutations and their functional impact in normal human tissues]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory Press</dc:publisher>
<prism:publicationDate>2025-09-17</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://biorxiv.org/cgi/content/short/2025.05.30.656844v1?rss=1">
<title>
<![CDATA[
Cell-type-specific patterns and consequences of somatic mutation in development and aging brain 
]]>
</title>
<link>
https://biorxiv.org/cgi/content/short/2025.05.30.656844v1?rss=1"
</link>
<description><![CDATA[
Elucidating the role of somatic mutations in cancer, healthy tissues, and aging depends on methods that can accurately characterize somatic mosaicism across different cell types, as well as assay their impact on cellular function. Current technologies to study cell-type-specific somatic mutations within tissues are low-throughput. We developed Duplex-Multiome, incorporating duplex consensus sequencing to accurately identify somatic single-nucleotide variants (sSNV) from the same nucleus simultaneously analyzed for single-nucleus ATAC-seq (snATAC-seq) and RNA-seq (snRNA-seq). By introducing strand-tagging into the construction of snATAC-seq libraries, duplex sequencing reduces sequencing error by >10,000-fold while eliminating artifactual mutational signatures. When applied to 98%/2% mixed cell lines, Duplex-Multiome identified sSNVs present in 2% of cells with 92% precision and accurately captured known sSNV mutational spectra, while revealing unexpected subclonal lineages. Duplex-Multiome of > 51,400 nuclei from postmortem brain tissue captured sSNV burdens and spectra across all major brain cell types and subtypes, including those difficult to assay by single-cell whole-genome sequencing (scWGS). This revealed for the first time that diverse neuronal and glial cell types show distinct rates and patterns of age-related mutation, while also directly discovering developmental cell lineage relationships. Duplex-Multiome identified clonal sSNVs occurring at increased rates in glia of certain aged brains, as well as clonal sSNVs that correlated with changes in expression of nearby genes, in both neurotypical and autism spectrum disorder (ASD) individuals, directly demonstrating that somatic mutagenesis can contribute to gene expression phenotypes. Duplex-Multiome can be easily adopted into the 10X Multiome protocol and will bridge somatic mosaicism to a wide range of phenotypic readouts across cell types and tissues.
]]></description>
<dc:creator>Kriz, A. J.</dc:creator>
<dc:creator>Mao, S.</dc:creator>
<dc:creator>Shao, D. D.</dc:creator>
<dc:creator>Snellings, D. A.</dc:creator>
<dc:creator>Andersen, R.</dc:creator>
<dc:creator>Dong, G.</dc:creator>
<dc:creator>Ma, C. C.</dc:creator>
<dc:creator>Cline, H. E.</dc:creator>
<dc:creator>Huang, A. Y.</dc:creator>
<dc:creator>Lee, E. A.</dc:creator>
<dc:creator>Walsh, C. A.</dc:creator>
<dc:date>2025-06-01</dc:date>
<dc:identifier>doi:10.1101/2025.05.30.656844</dc:identifier>
<dc:title><![CDATA[Cell-type-specific patterns and consequences of somatic mutation in development and aging brain]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory Press</dc:publisher>
<prism:publicationDate>2025-06-01</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://biorxiv.org/cgi/content/short/2024.06.14.599122v1?rss=1">
<title>
<![CDATA[
A haplotype-resolved view of human gene regulation 
]]>
</title>
<link>
https://biorxiv.org/cgi/content/short/2024.06.14.599122v1?rss=1"
</link>
<description><![CDATA[
Diploid human cells contain two non-identical genomes, and differences in their regulation underlie human development and disease. We present Fiber-seq Inferred Regulatory Elements (FIRE) and show that FIRE provides a more comprehensive and quantitative snapshot of the accessible chromatin landscape across the 6 Gbp diploid human genome, overcoming previously unrecognized biases in existing regulatory element catalogs. FIRE enables comprehensive detection of haplotype-selective chromatin accessibility (HSCA), exposing novel imprinted elements lacking underlying parent-of-origin CpG methylation differences, and gene regulatory modules that permit genes to escape X chromosome inactivation. We uncover that the human leukocyte antigen (HLA) locus harbors the most HSCA in immune cells, where we resolve specific transcription factor (TF) binding events disrupted by disease-associated variants. Finally, we demonstrate that the regulatory landscape of a cell is littered with autosomal somatic chromatin epimutations that are propagated by clonal expansions to create mitotically stable and non-genetically deterministic chromatin alterations.
]]></description>
<dc:creator>Vollger, M. R.</dc:creator>
<dc:creator>Swanson, E. G.</dc:creator>
<dc:creator>Neph, S. J.</dc:creator>
<dc:creator>Ranchalis, J.</dc:creator>
<dc:creator>Munson, K. M.</dc:creator>
<dc:creator>Ho, C.-H.</dc:creator>
<dc:creator>Sedeno-Cortes, A. E.</dc:creator>
<dc:creator>Fondrie, W. E.</dc:creator>
<dc:creator>Bohaczuk, S. C.</dc:creator>
<dc:creator>Mao, Y.</dc:creator>
<dc:creator>Parmalee, N. L.</dc:creator>
<dc:creator>Mallory, B. J.</dc:creator>
<dc:creator>Harvey, W. T.</dc:creator>
<dc:creator>Kwon, Y.</dc:creator>
<dc:creator>Garcia, G. H.</dc:creator>
<dc:creator>Hoekzema, K.</dc:creator>
<dc:creator>Meyer, J. G.</dc:creator>
<dc:creator>Cicek, M.</dc:creator>
<dc:creator>Eichler, E. E.</dc:creator>
<dc:creator>Noble, W. S.</dc:creator>
<dc:creator>Witten, D. M.</dc:creator>
<dc:creator>Bennett, J. T.</dc:creator>
<dc:creator>Ray, J. P.</dc:creator>
<dc:creator>Stergachis, A. B.</dc:creator>
<dc:date>2024-06-16</dc:date>
<dc:identifier>doi:10.1101/2024.06.14.599122</dc:identifier>
<dc:title><![CDATA[A haplotype-resolved view of human gene regulation]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory Press</dc:publisher>
<prism:publicationDate>2024-06-16</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://biorxiv.org/cgi/content/short/2024.06.18.599438v1?rss=1">
<title>
<![CDATA[
Transgenerational transmission of post-zygotic mutations suggests symmetric contribution of first two blastomeres to human germline 
]]>
</title>
<link>
https://biorxiv.org/cgi/content/short/2024.06.18.599438v1?rss=1"
</link>
<description><![CDATA[
Little is known about the origin of germ cells in humans. We previously leveraged post-zygotic mutations to reconstruct zygote-rooted cell lineage ancestry trees in a phenotypically normal woman, termed NC0. Here, by sequencing the genome of her children and their father, we analyzed the transmission of early pre-gastrulation lineages and corresponding mutations across human generations. We found that the germline in NC0 is polyclonal and is founded by at least two cells likely descending from the two blastomeres arising from the first zygotic cleavage. Analyses of public data from several multi-children families and from 1,934 familial quads confirmed this finding in larger cohorts, revealing that known imbalances of up to 90:10 in early lineages allocation in somatic tissues are not reflected in transmission to offspring, establishing a fundamental difference in lineage allocation between the soma and the germline. Analyses of all the data consistently suggest that germline has a balanced 50:50 lineage allocation from the first two blastomeres.
]]></description>
<dc:creator>Jang, Y.</dc:creator>
<dc:creator>Tomasini, L.</dc:creator>
<dc:creator>Bae, T.</dc:creator>
<dc:creator>Szekely, A.</dc:creator>
<dc:creator>Vaccarino, F. M.</dc:creator>
<dc:creator>Abyzov, A.</dc:creator>
<dc:date>2024-06-22</dc:date>
<dc:identifier>doi:10.1101/2024.06.18.599438</dc:identifier>
<dc:title><![CDATA[Transgenerational transmission of post-zygotic mutations suggests symmetric contribution of first two blastomeres to human germline]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory Press</dc:publisher>
<prism:publicationDate>2024-06-22</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://biorxiv.org/cgi/content/short/2023.12.30.573716v1?rss=1">
<title>
<![CDATA[
Threshold of somatic mosaicism disrupting the brain function 
]]>
</title>
<link>
https://biorxiv.org/cgi/content/short/2023.12.30.573716v1?rss=1"
</link>
<description><![CDATA[
Somatic mosaicism in a fraction of brain cells causes neurodevelopmental disorders, including childhood intractable epilepsy. However, the threshold for somatic mosaicism leading to brain dysfunction is unknown. In this study, we induced various mosaic burdens in mice of focal cortical dysplasia type II (FCD II), featuring mTOR somatic mosaicism and spontaneous behavioral seizures. Mosaic burdens ranged from approximately 1,000 to 40,000 neurons expressing the mTOR mutant in the somatosensory (SSC) or medial prefrontal (PFC) cortex. Surprisingly, just [~]8,000-9,000 neurons expressing the MTOR mutant were sufficient to trigger epileptic seizures. Mutational burden correlated with seizure frequency and onset, with a higher tendency for electrographic inter-ictal spikes and beta- and gamma-frequency oscillations in FCD II mice exceeding the threshold. Moreover, mutation-negative FCD II patients in deep sequencing of their bulky brain tissues revealed somatic mosaicism of mTOR pathway genes as low as 0.07% in resected brain tissues through ultra-deep targeted sequencing (up to 20 million reads). Thus, our study suggests that extremely low levels of somatic mosaicism can contribute to brain dysfunction.
]]></description>
<dc:creator>Kim, J.</dc:creator>
<dc:creator>Park, S. M.</dc:creator>
<dc:creator>Koh, H. Y.</dc:creator>
<dc:creator>Ko, A.</dc:creator>
<dc:creator>Kang, H.-C.</dc:creator>
<dc:creator>Chang, W. S.</dc:creator>
<dc:creator>Kim, D. S.</dc:creator>
<dc:creator>Lee, J. H.</dc:creator>
<dc:date>2023-12-30</dc:date>
<dc:identifier>doi:10.1101/2023.12.30.573716</dc:identifier>
<dc:title><![CDATA[Threshold of somatic mosaicism disrupting the brain function]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory Press</dc:publisher>
<prism:publicationDate>2023-12-30</prism:publicationDate>
<prism:section></prism:section>
</item>
</rdf:RDF>
