<?xml version="1.0" encoding="UTF-8" ?>
<rdf:RDF xmlns:admin="http://webns.net/mvcb/" xmlns="http://purl.org/rss/1.0/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:prism="http://purl.org/rss/1.0/modules/prism/" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:syn="http://purl.org/rss/1.0/modules/syndication/">
<channel rdf:about="https://biorxiv.org">
<admin:errorReportsTo rdf:resource="mailto:biorxiv@cshlpress.edu"/>
<title>bioRxiv Subject Collection: Genomics</title>
<link>https://biorxiv.org</link>
<description>
This feed contains articles for bioRxiv Subject Collection "Genomics"
</description>

<items>
<rdf:Seq>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.26.754726v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.27.754747v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.30.755763v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.30.755832v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.26.754650v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.25.753922v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.30.755794v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.23.752440v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.23.753890v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.30.755774v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.30.755709v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.25.754321v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.28.754967v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.27.754784v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.27.754713v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.30.755330v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.25.754218v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.28.755014v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.24.754104v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.24.753534v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.24.754018v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.24.753990v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.24.753865v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.23.753885v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.23.753695v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.26.754026v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.28.754846v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.22.753638v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.22.753313v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.26.754616v1?rss=1"/>
</rdf:Seq>
</items>
<prism:eIssn/>
<prism:publicationName>bioRxiv</prism:publicationName>
<prism:issn/>

<image rdf:resource=""/>
</channel>
<image rdf:about="">
<title>bioRxiv</title>
<url>https://www.biorxiv.org/sites/default/files/bioRxiv_article.jpg</url>
<link>https://www.biorxiv.org</link>
</image>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.26.754726v1?rss=1">
<title>
<![CDATA[
X-Y divergence of house fly (Musca domestica) proto-sex chromosomes follows distinct evolutionary trajectories despite residing in the same genome 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.26.754726v1?rss=1
</link>
<description><![CDATA[
Sex determination systems and sex chromosomes frequently differ between species. One cause of these differences is new sex determining genes that drive evolutionary turnover of sex chromosomes. At the earliest stages of this turnover, X and Y (or Z and W) chromosomes start out as nearly identical homologs, and they can diverge via chromosomal rearrangements (e.g., inversions) that suppress X-Y recombination. However, multiple examples from across animals and plants provide exceptions to this canonical model of sex chromosome evolution. For example, some X-Y pairs remain undifferentiated for long evolutionary time periods, while other sex chromosomes become differentiated without chromosomal rearrangements. How or why these non-canonical trajectories occur remains elusive, despite increasing evidence of their pervasiveness. The house fly, Musca domestica, is a well-suited system to address this gap because all six chromosomes can be a Y, providing multiple replicates of a natural experiment within a single genomic environment. To test for canonical and non-canonical evolutionary trajectories, we generated haplotype-resolved chromosome-level assemblies from five strains of the house fly, each of which carries a different Y chromosome (IM, IIM, IIIM, VM, and YM). We identified an inversion on only one of the sex chromosomes (IIM), which was associated with elevated X-Y divergence but did not capture the male-determining locus. In contrast, there was X-Y divergence across almost the entire length of the IM, IIIM, and VM sex chromosomes, despite no detectable inversions. YM was the only sex chromosome to contain substantial Y-specific sequences, which were limited to a segment on one end of the chromosome containing the male-determining gene. This YM chromosome, and its corresponding X, were highly diverged from the X chromosome found in many other flies (Muller element F), despite a strong cytological resemblance. This study highlights how multiple different canonical and non-canonical modes of sex chromosome evolution can co-exist within a single genome.
]]></description>
<dc:creator><![CDATA[ Son, J. H., Luecke, D., Bista, B., Manat, Y., Cao, W., Beukeboom, L., Bopp, D., Ellison, C. E., Saelao, P., Meisel, R. P. ]]></dc:creator>
<dc:date>2026-10-02</dc:date>
<dc:identifier>doi:10.64898/2026.09.26.754726</dc:identifier>
<dc:title><![CDATA[X-Y divergence of house fly (Musca domestica) proto-sex chromosomes follows distinct evolutionary trajectories despite residing in the same genome]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-02</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.27.754747v1?rss=1">
<title>
<![CDATA[
Method-dependent biases in cell type detection between single-cell and single-nucleus RNA sequencing in the photosymbiotic acoel Praesagittifera naikaiensis 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.27.754747v1?rss=1
</link>
<description><![CDATA[
Background Comparisons of single-cell and single-nucleus RNA sequencing (scRNA-seq and snRNA-seq) data have been described in some mammalian tissues and, subsequently, in Drosophila, but remain unexplored in most invertebrate lineages. The xenacoelomorphs occupy key phylogenetic positions, yet they differ anatomically from mammals. They have a reduced extracellular matrix, high-salt body fluid, and no circulatory system. Despite these differences, they possess a well-developed nervous system. One of the xenacoelomorphs, the photosymbiotic acoel (Praesagittifera naikaiensis) also harbours symbiotic Tetraselmis algae, whose RNA can be co-captured with host RNA. Results We compared scRNA-seq and snRNA-seq data from whole P. naikaiensis specimens. Both methods yielded high-quality data with comparable gene detection but a larger share of scRNA-seq reads derived from symbiotic algae. Gene-level analyses revealed that neural genes were enriched in snRNA-seq relative to non-neural genes. Cross-method label transfer and integration-based validation identified six snRNA-seq clusters, including some neural populations, that lacked a clear scRNA-seq counterpart. In contrast, three cell populations, including muscle and metabolically active clusters, were reciprocally validated as captured by both methods. Conclusions Our results show that key snRNA-seq advantages, particularly the enhanced recovery of neural transcripts, are recapitulated in our dataset, which is consistent with previous reports in mammals. Furthermore, snRNA-seq reduces symbiont-derived reads and recovers several cell populations underrepresented in scRNA-seq. These findings provide practical guidance for cell atlas construction in non-model, symbiotic invertebrates.
]]></description>
<dc:creator><![CDATA[ Nakamura, R., Hamada, M., Kobayashi, K., Ansai, S., Dindo, M., Sakagami, T., Luo, Y.-J., Satou, Y., Sekiguchi, T., Sakamoto, T. ]]></dc:creator>
<dc:date>2026-10-02</dc:date>
<dc:identifier>doi:10.64898/2026.09.27.754747</dc:identifier>
<dc:title><![CDATA[Method-dependent biases in cell type detection between single-cell and single-nucleus RNA sequencing in the photosymbiotic acoel Praesagittifera naikaiensis]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-02</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.30.755763v1?rss=1">
<title>
<![CDATA[
Genome-wide mapping of endogenous topoisomerase cleavage complexes reveals TOP1 at the interface of DNA topological regulation and genome instability 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.30.755763v1?rss=1
</link>
<description><![CDATA[
Topoisomerase cleavage complexes (TOPcc) are covalent protein-DNA intermediates formed transiently during topoisomerase catalysis and constitute a major class of DNA-protein crosslinks in human cells, yet their endogenous genomic distributions remain difficult to resolve. Here, we developed protein-DNA adduct immunoprecipitation followed by sequencing (pDAIP-seq), a highly sensitive method for detecting endogenous TOPcc that enables their genome-wide profiling under normal physiological conditions without stabilization by topoisomerase poisons. Using pDAIP-seq, we found that camptothecin treatment, required by existing TOP1cc profiling methods, not only increased TOP1cc abundance but also substantially altered its genomic distribution. Endogenous TOP1cc were globally reduced upon transcriptional suppression and increased upon TOP2A depletion, supporting TOP1cc regulation by the maintenance of genome-wide DNA topological homeostasis. Interestingly, both the regional distribution and nucleotide-scale positioning of TOP1cc, including the positions of TOP1 catalytic cleavage sites, were largely specified by local DNA sequence context. By extending pDAIP-seq to capture transient covalent protein-DNA intermediates involved in DNA repair, including TDP1-, PARP1-, and KU70-DNA adducts, we further identified endogenous TOP1cc as a potential major source of spontaneous DNA breaks in human cells. Finally, pDAIP-seq also enabled genome-wide mapping of TOP2cc, TOP3cc, and SPO11-DNA covalent intermediates in mouse testis undergoing programmed meiotic DNA cleavage. Together, these findings establish pDAIP-seq as a broadly applicable method for the mapping of covalent protein-DNA intermediates and reveal endogenous TOP1cc as a key intermediate between DNA topological regulation and human genome instability.
]]></description>
<dc:creator><![CDATA[ Li, S., Chen, C. ]]></dc:creator>
<dc:date>2026-10-02</dc:date>
<dc:identifier>doi:10.64898/2026.09.30.755763</dc:identifier>
<dc:title><![CDATA[Genome-wide mapping of endogenous topoisomerase cleavage complexes reveals TOP1 at the interface of DNA topological regulation and genome instability]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-02</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.30.755832v1?rss=1">
<title>
<![CDATA[
Topoisomerase-mediated resolution of negative DNA supercoiling sustains human transcription and gene activation 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.30.755832v1?rss=1
</link>
<description><![CDATA[
Transcription generates DNA supercoiling that can influence gene expression, but how transcription-associated supercoiling shapes RNA polymerase II (RNAPII) dynamics in human cells remain poorly understood. Here, we show that topoisomerase-mediated resolution of negative DNA supercoiling is required for productive transcription. Acute topoisomerase perturbation causes negative supercoiling to accumulate around transcription start sites (TSS), with greater negative supercoiling accumulation associated with stronger transcription reduction. Elevated negative supercoiling behind the elongating RNAPII, but not positive supercoiling ahead, is associated with reduced RNAPII elongation rates. Using NOTE-seq2 to simultaneously profile newly synthesized and total RNA in single cells, we further demonstrate that acute TOP1 depletion increases gene expression noise by reducing both transcriptional burst size and frequency. Mechanistically, negative supercoiling accumulation at TSS increases RNAPII promoter occupancy and promoter-proximal pausing, providing a potential basis for reduced burst size. Micro-C analysis shows that TOP1 depletion largely preserves chromatin compartments and TADs but weakens short-range chromatin loops, including enhancer-promoter interactions, likely contributing to reduced burst frequency. Finally, during T-cell activation, impaired resolution of negative supercoiling compromises rapid gene upregulation and attenuates the transcriptional activation program. Together, these findings establish topoisomerase-mediated resolution of negative DNA supercoiling as a key mechanism linking DNA topological homeostasis to RNAPII dynamics, chromatin looping, and rapid gene activation.
]]></description>
<dc:creator><![CDATA[ Lyu, J., Ahmed, A., Biswas, I., Li, S., Chen, C. ]]></dc:creator>
<dc:date>2026-10-02</dc:date>
<dc:identifier>doi:10.64898/2026.09.30.755832</dc:identifier>
<dc:title><![CDATA[Topoisomerase-mediated resolution of negative DNA supercoiling sustains human transcription and gene activation]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-02</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.26.754650v1?rss=1">
<title>
<![CDATA[
Applying AlphaGenome Variant Impact for SNV Prioritization 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.26.754650v1?rss=1
</link>
<description><![CDATA[
AlphaGenome Atlas, released in September 2026, provides functional predictions for nearly all possible human single-nucleotide variants using the AlphaGenome Variant Impact (AVI) score. However, empirical thresholds relating AVI quantile rank (AVI_QUANTILE) to clinical variant classifications have not been established. We derived AVI_QUANTILE thresholds distinguishing ClinVar pathogenic, likely pathogenic, uncertain, conflicting, likely benign, and benign SNVs using receiver operating characteristic analyses. Pathogenic/likely pathogenic variants were strongly enriched at the upper end of the AVI_QUANTILE distribution, yielding a Youden-optimal ClinVar-derived cutoff of 0.9938. Optimal cutoffs were generally consistent across several major molecular consequence classes, although lower thresholds were observed for selected noncoding categories. Similarly, most well-powered gene-disease groups showed cutoffs close to the global threshold, with modest gene-specific variation. Among 24,989 variants of uncertain significance, 11,357 (45.45%) exceeded the global cutoff, identifying a substantial subset with AlphaGenome-predicted functional effects comparable to those observed among pathogenic/likely pathogenic variants and potentially warranting further evaluation. In a steroid-dependent asthma fine-mapping analysis, AVI_QUANTILE showed little correspondence with SuSiE posterior inclusion probabilities, indicating that AlphaGenome functional predictions capture information distinct from statistical fine-mapping.
]]></description>
<dc:creator><![CDATA[ Qu, H.-Q., Hakonarson, H. ]]></dc:creator>
<dc:date>2026-10-01</dc:date>
<dc:identifier>doi:10.64898/2026.09.26.754650</dc:identifier>
<dc:title><![CDATA[Applying AlphaGenome Variant Impact for SNV Prioritization]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-01</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.25.753922v1?rss=1">
<title>
<![CDATA[
Genetic risk and inflammatory signaling converge on cell type-specific regulatory programs in type 1 diabetes 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.25.753922v1?rss=1
</link>
<description><![CDATA[
Type 1 diabetes (T1D) is a complex autoimmune disease characterized by the destruction of insulin-producing pancreatic {beta} cells. Both genetic susceptibility and epigenetic dysregulation contribute to T1D risk. Recent studies have reported {beta}-cell dysfunction before disease onset, suggesting that {beta}-cell-intrinsic mechanisms contribute to pathogenesis. However, the mechanistic link between disease-associated variants and {beta}-cell dysfunction remains poorly understood. We hypothesized that a subset of T1D-associated variants directly influences cell type-specific candidate cis-regulatory elements (cCREs) by disrupting transcription factor (TF) binding and altering gene expression. To define cell type-specific cCREs in human islets, we integrated publicly available epigenomic datasets, including bulk histone modification profiles and single-cell multiomic data from human pancreatic islets. We combined these data with T1D genome-wide association study (GWAS) variants to identify disease-associated SNPs located within enhancer regions. Using multiomic datasets from PANC-DB, we characterized cell type-specific chromatin accessibility in nondiabetic and T1D islets and identified regulatory regions that may control gene expression in {beta} cells and other islet cell types. We then applied ChromBPNet, a deep learning model that predicts base-pair-resolution chromatin accessibility from scATAC-seq data, to evaluate how specific variants may alter local regulatory activity. In parallel, we used TF footprinting to nominate TFs likely to bind these variant-containing regions. These analyses identified several T1D-associated SNPs predicted to alter chromatin accessibility at candidate TF binding sites, suggesting mechanisms by which noncoding variants may contribute to {beta}-cell dysfunction and T1D susceptibility. Our findings link T1D-associated variants to cell type-specific enhancer activity and regulatory pathways and provide candidates for future mechanistic studies.
]]></description>
<dc:creator><![CDATA[ Wang, L., Wei, Z. ]]></dc:creator>
<dc:date>2026-10-01</dc:date>
<dc:identifier>doi:10.64898/2026.09.25.753922</dc:identifier>
<dc:title><![CDATA[Genetic risk and inflammatory signaling converge on cell type-specific regulatory programs in type 1 diabetes]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-01</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.30.755794v1?rss=1">
<title>
<![CDATA[
Scalable saturation mutagenesis reveals gene regulatory architecture and rare variant effects 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.30.755794v1?rss=1
</link>
<description><![CDATA[
Long-context sequence-to-function models enable nucleotide-resolution prediction of regulatory variant effects, motivating comprehensive mutational interrogation across broad genomic contexts. Yet conventional in silico saturation mutagenesis (ISM) scores mutations one at a time, requiring millions of model evaluations for a single gene and billions to trillions at genome scales. Here, we introduce Multi-ISM, a scalable framework that reformulates ISM as a sparse recovery problem. Multi-ISM generates mutational maps using 45-fold fewer model evaluations than exhaustive single-variant ISM and matches or exceeds its accuracy on variant-effect benchmarks. Multi-ISM is architecture-agnostic and transfers across large-scale sequence-to-function models. We applied Multi-ISM to 5,000 protein-coding genes, including 3,317 OMIM disease genes, generating base-pair-resolution, tissue-resolved attribution maps across 500-kb windows. These maps supported enhancer--gene prioritization and identification of cell-type-specific regulatory elements. Gene-level summaries of the maps captured regulatory complexity and showed that more constrained genes had smaller predicted mutational effects. Aggregating Multi-ISM predictions into gene-level rare-variant burdens improved personalized expression prediction over a common-variant elastic net, with the largest gains at expression outliers. Multi-ISM makes nucleotide-resolution interpretation of long-context sequence models a routine computation rather than a dedicated effort, so that new architectures, functional readouts, and cellular contexts can be mapped as they appear.
]]></description>
<dc:creator><![CDATA[ Yuan, H., Huang, X., Auerbach, B., Linder, J., Srivastava, D., Kelley, D. R. ]]></dc:creator>
<dc:date>2026-10-01</dc:date>
<dc:identifier>doi:10.64898/2026.09.30.755794</dc:identifier>
<dc:title><![CDATA[Scalable saturation mutagenesis reveals gene regulatory architecture and rare variant effects]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-01</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.23.752440v1?rss=1">
<title>
<![CDATA[
The Genome In A Bottle HG002 assembly-based variant benchmark set enables comprehensive benchmarking of small and structural variants. 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.23.752440v1?rss=1
</link>
<description><![CDATA[
High-accuracy, high-resolution benchmark data are fundamental to improving the accuracy and generalizability of sequencing technologies and associated algorithms. Most existing benchmarks for DNA sequencing rely on mapping to a single linear reference, omitting inaccessible and highly repetitive complex genomic regions. Here, we present the Genome in a Bottle (GIAB) Consortium's HG002 v5.0q benchmark set, a more comprehensive variant benchmark derived from a fully phased, highly curated, and accurate diploid assembly of HG002. Our approach leverages recent advances in long-read sequencing and assembly to provide confident small- and structural-variant calls across previously inaccessible regions, including segmental duplications, homopolymers, tandem repeats, and complex immune loci. The v5.0 benchmark significantly expands the callable genome relative to v4.2.1, adding x200 Mbp of sequence and increasing variant detection by 18 % for small variants and 3-fold for structural variants compared to previous benchmarks. The benchmark regions include more variants in repetitive regions, including 1.2 x segmental duplications, 1.8 x tandem repeats, and 2.0 x homopolymers, compared to v4.2.1. Separate and combined small- and structural-variant benchmarks are available for three versions of the human reference genome (GRCh37, GRCh38, and CHM13). We validate the reliability of this benchmark set by comparing multiple sequencing technologies and variant calling pipelines and curating differences. The v5.0q benchmark provides a critical resource for the development and validation of next-generation sequencing technologies and analysis methods.
]]></description>
<dc:creator><![CDATA[ Olson, N. D., Dwarshuis, N., Hansen, N. F., Koren, S., Rhie, A., McDaniel, J. H., Wagner, J., Berry, G., Shafin, K., English, A., Saunders, C. T., Kundu, R., Holt, J. M., Freed, D., Abdulkadir, A. A., Daniels, C., Majidian, S., Paulin, L. F., Sedlazeck, F. J., Fleharty, M., Gao, Y., Denti, L., Brambrink, L., Carroll, A., Chang, P.-C., Cook, D. E., Kolesnikov, A., Catreux, S., Edlund, C., Han, J., Murray, L., Qiu, Y., Rossi, M., Scheffler, K., Truong, S., Wang, Y., Chikhi, R., Flores, C., Jaspez, D., Lorenzo-Salazar, J. M., Munoz-Barrera, A., Rubio-Rodriguez, L. A., Lin, M.-J., Vaddadi, K., Lan ]]></dc:creator>
<dc:date>2026-10-01</dc:date>
<dc:identifier>doi:10.64898/2026.09.23.752440</dc:identifier>
<dc:title><![CDATA[The Genome In A Bottle HG002 assembly-based variant benchmark set enables comprehensive benchmarking of small and structural variants.]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-01</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.23.753890v1?rss=1">
<title>
<![CDATA[
Chromosome-level genome assembly of the Olympia oyster (Ostrea lurida), the only oyster species native to the West Coast of North America 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.23.753890v1?rss=1
</link>
<description><![CDATA[
The Olympia oyster (Ostrea lurida) was once a major food source and a key component of coastal marine ecosystems and culture on the US West Coast. Overfishing and habitat contamination in the early to mid-20th century severely depleted populations of Olympia oysters, and the species is currently considered a species of concern. Although multiple restoration efforts are underway, progress understanding the species biology and recovery potential has been limited by the historical lack of genomic resources. Here, we report a chromosome-level assembly of the Olympia oyster. The assembly was built using Oxford Nanopore reads and Hi-C to obtain a highly contiguous 1.03 Gb-assembly consisting of 10 chromosomes. RNA from gills, mantle, adductor muscle, and a mix of brooding eggs and larvae were also sequenced as 150 nt long paired-end reads which were used to guide the genome annotation. Over 52 K genes were predicted in the genome, of which 34K were protein-coding. A BUSCO analysis on the assembly showed 98.9% of the mollusca_odb12 gene set were complete (with 2.1% duplicated, and 0.8% missing). The short RNA reads were also used to identify tissue-level gene expression signatures in adductor muscles, egg and brooding larvae, and mantle. Development-related genes were highly expressed in eggs and brooding larvae whereas many xenobiotic metabolism genes were present in mantle tissues, highlighting mantles role in the interaction of the oyster with external environments. This new genomic resource will greatly advance our understanding of this keystone mollusk and provide a critical tool for guiding restoration efforts and the development of sustainable commercial production.
]]></description>
<dc:creator><![CDATA[ Calla, B., Sim, S. B., Yeats, M. S., Johnson, K. M., Chavez-Congrains, K., Kauwe, A. N., Plough, L. V., Geib, S. M. ]]></dc:creator>
<dc:date>2026-10-01</dc:date>
<dc:identifier>doi:10.64898/2026.09.23.753890</dc:identifier>
<dc:title><![CDATA[Chromosome-level genome assembly of the Olympia oyster (Ostrea lurida), the only oyster species native to the West Coast of North America]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-01</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.30.755774v1?rss=1">
<title>
<![CDATA[
Factors underlying creativity show distinct genetic architecture and genetic correlations with psychiatric disorders 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.30.755774v1?rss=1
</link>
<description><![CDATA[
While creativity is widely recognized as multifaceted, genome-wide association studies have largely focused on individual measures. Here, we used genomic structural equation modeling of six creativity-related GWASs and identified two latent genetic factors: a general factor and an artistic factor, with 22 and 13 genome-wide significant loci, respectively, nine of which are novel. The general factor was almost perfectly genetically correlated with educational attainment (rg = 0.96), suggesting that it primarily captured genetic influences shared with cognitive and educational processes rather than creativity-specific variation. In contrast, the artistic factor showed only a weak genetic correlation with educational attainment (rg = 0.11) and consistent positive genetic correlations with psychiatric disorders, whereas associations for the general factor were largely null or negative. Biological annotation analyses further identified distinct enrichment patterns unique to the artistic factor, providing initial evidence that different dimensions of creativity may be characterized by partially distinct neurobiological pathways.
]]></description>
<dc:creator><![CDATA[ Xia, P., Wesseldijk, L. W., Bignardi, G., Lu, Y., Ullen, F., Mosing, M. A. ]]></dc:creator>
<dc:date>2026-09-30</dc:date>
<dc:identifier>doi:10.64898/2026.09.30.755774</dc:identifier>
<dc:title><![CDATA[Factors underlying creativity show distinct genetic architecture and genetic correlations with psychiatric disorders]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-30</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.30.755709v1?rss=1">
<title>
<![CDATA[
Intelligence and the Big Five personality traits share a genetic and biological basis 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.30.755709v1?rss=1
</link>
<description><![CDATA[
Intelligence is often considered to be fundamentally different from other personality traits. Twin studies show that intelligence and the Big Five personality traits are partly influenced by the same genetic factors, but it remains unclear to what degree shared genes, SNPs, and biological pathways influence them. To address these questions, we analyzed the genetic associations of all six traits. First, we calculated their genetic correlations using LD score regression. We then quantified polygenic overlap using a bivariate causal mixture model, finding that approximately 90% of the genetic variants influencing the Big Five personality traits also impact intelligence. Finally, we performed gene prioritization and gene set enrichment analysis on each trait, once again finding substantial overlap. At each level, intelligence was as similar to the Big Five traits as they were to one another. Genetically and biologically, intelligence fits in with the rest of personality.
]]></description>
<dc:creator><![CDATA[ Buonaccorsi, A. C., DeYoung, C. G., Maher, D., Edwards, T. ]]></dc:creator>
<dc:date>2026-09-30</dc:date>
<dc:identifier>doi:10.64898/2026.09.30.755709</dc:identifier>
<dc:title><![CDATA[Intelligence and the Big Five personality traits share a genetic and biological basis]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-30</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.25.754321v1?rss=1">
<title>
<![CDATA[
Genomic structural variation in clownfish adaptive radiation 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.25.754321v1?rss=1
</link>
<description><![CDATA[
Clownfishes are a complex of 28 species that rapidly diversified after acquiring mutualism with host sea anemones around 15 million years ago. The genomic mechanisms behind this radiation are only beginning to emerge, and structural variants (SVs) have received little attention. Previous phylogenomic studies suggested two large inversions on chromosome 18, but their distribution across the clownfish phylogeny and their potential effects have not been formally explored. SVs can disrupt gene function and regulation or alter gene dosage, and they have been shown to play a central role in adaptive evolution and diversification across many taxa. Here, we characterized SVs across the clownfish radiation using long-read PacBio sequencing, generating high-quality genome assemblies for sixteen species spanning the main lineages of the genus. We identified and catalogued nearly 130,000 SVs spanning close to 570 Mb. Short insertions and deletions were the most numerous, but large inversions accounted for most of the affected sequence. No SV was consistently associated with host specialization. We nevertheless confirmed a single inversion of approximately 17 Mb on chromosome 18, clarifying earlier reports of two separate rearrangements at this locus, and identified an additional inversion on chromosome 9 whose distribution suggests a history of hybridization. This study provides the first genome-wide characterization of structural variation in clownfishes, revealing extensive SVs despite the group's rapid and recent radiation. It also delivers new, high-quality genomic resources that will support future research into this iconic group of coral reef fishes and their adaptation to sea anemones.
]]></description>
<dc:creator><![CDATA[ Marcionetti, A., Inacio Martins, E., Hartasanchez, D. A., Cortesi, F., Fitzgerald, L. M., Frederich, B., Gaboriau, T., Garcia Jimenez, A., Laudet, V., Schmid, S., Salamin, N. ]]></dc:creator>
<dc:date>2026-09-30</dc:date>
<dc:identifier>doi:10.64898/2026.09.25.754321</dc:identifier>
<dc:title><![CDATA[Genomic structural variation in clownfish adaptive radiation]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-30</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.28.754967v1?rss=1">
<title>
<![CDATA[
Phylodynamic analysis reveals asymmetric and accelerating global spread of dengue virus 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.28.754967v1?rss=1
</link>
<description><![CDATA[
Dengue virus (DENV) continues to expand globally, with increasing incidence in both endemic and non-endemic countries. Leveraging a large global dataset of DENV genomes starting from the 1940s, we reconstructed global and regional patterns of DENV spread using phylogeographic and phylodynamic methods. Our analyses reveal that the current global diversity of urban DENV originated in Asia in the first half of the 20th century. During the more than 100 years since, we find that the highest density of between-country spread occurs within Asia and the region continues to be a global source, seeding transmission in Africa and the Americas. Global spread is, however, not balanced. Even though there is more travel to South America, we show that DENV lineages primarily enter North America (mostly from South Asia) through a link defined by seasonal synchrony across the Northern Hemisphere. Along with a substantial increase in global travel starting ~2010, the rate of spread to the Americas is increasing, establishing new lineages associated with large recent outbreaks. DENV lineages do not frequently spread back to Asia from the Americas or Africa, and we posit this is due to higher genetic diversity in Asia, leading to stiffer competition for invading lineages. Within the Americas, the network of DENV dispersal initiates in North America, primarily the Caribbean, and then is maintained by major hubs in South America. Overall, we find that travel volume, rather than climate factors, is mostly responsible for long-term patterns of DENV spread globally and within the Americas. Our results highlight key nodes in the global dynamics of DENV lineages, and suggest spread will continue to accelerate in a more connected world.
]]></description>
<dc:creator><![CDATA[ Hill, V., Dudas, G., Ji, X., Huits, R., Hamer, D. H., Joshi, H.-K., Porzucek, A. J., Lopes, R., Robertson, H., Whittaker, C. F., Kraemer, M. U. G., Johansson, M. A., Faria, N. R., Brady, O., Dellicour, S., Suchard, M. A., Lemey, P., Carlson, C. J., Baele, G., Grubaugh, N. D. ]]></dc:creator>
<dc:date>2026-09-30</dc:date>
<dc:identifier>doi:10.64898/2026.09.28.754967</dc:identifier>
<dc:title><![CDATA[Phylodynamic analysis reveals asymmetric and accelerating global spread of dengue virus]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-30</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.27.754784v1?rss=1">
<title>
<![CDATA[
Whole-genome characterization of Plasmodium falciparum population structure and drug resistance across Uganda and neighboring countries 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.27.754784v1?rss=1
</link>
<description><![CDATA[
Uganda has one of the world's highest malaria burdens, but parasite population structure and connectivity have not been characterized nationally with whole-genome data. We analyzed 249 monoclonal Plasmodium falciparum genomes from 26 sentinel sites across Uganda in 2023 and compared them with publicly available genomes from five neighboring countries. Genetic diversity was high and differentiation low in Uganda, but identity-by-descent detected limited structure. This was most pronounced in the low-transmission southwest, where parasites shared ancestry and recent relatedness with isolates from western Tanzania. Selection scans identified signals at loci associated with resistance to artemisinins and their partner drugs, and to chloroquine and antifolates. The strongest signal was at PfPX1 where the PIN haplotype (L1222P, M1701I, D1705N) reached 70% prevalence nationwide, but selection was carried by PIN haplotypes that had additionally acquired D384A, largely confined to northern and eastern Uganda. A second substitution at the same codon, D384Y, formed a haplotype shared across the Tanzanian border. PfKelch13 variants were partitioned across multiple backgrounds, consistent with repeated independent emergence, and C469Y was enriched on the PfPX1 D384A-PIN background. These genomes provide a 2023 baseline for tracking potential mediators of drug resistance and support coordinated genomic surveillance across Uganda and its neighbors.
]]></description>
<dc:creator><![CDATA[ Kiyaga, S., Mbabazi, M., Katairo, T., Asua, V., Semakuba, F. D., Nsengimaana, B., Tukwasibwe, S., Wiringilimaana, I., Nakasaanyaa, J., Mulondoa, J., Watyekele, E., Maiteki, C. S., Katete, D. P., Nabende, J. N., Jjingo, D., Mboowa, G., Batte, C., Sam, N., Kamya, M. R., Agaba, B. B., Ssewanyana, I., Aranda-Diazi, A., Ajogbasile, F. V., Davlieva, M., Dorsey, G., Rosenthal, P. J., Conrad, M. D., Greenhouse, B., Hathaway, N. J., Briggs, J. ]]></dc:creator>
<dc:date>2026-09-30</dc:date>
<dc:identifier>doi:10.64898/2026.09.27.754784</dc:identifier>
<dc:title><![CDATA[Whole-genome characterization of Plasmodium falciparum population structure and drug resistance across Uganda and neighboring countries]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-30</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.27.754713v1?rss=1">
<title>
<![CDATA[
Rapid whole genome- and annotation-based multilocus sequence typing of the fungal pathogen Histoplasma capsulatum 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.27.754713v1?rss=1
</link>
<description><![CDATA[
Introduction. Histoplasma capsulatum is an important fungal pathogen of humans, especially in immunocompromised patients, and the World Health Organisation has listed Histoplasma as a High Priority fungal pathogen. The taxonomy of the genus Histoplasma is still under development, but its only species H. capsulatum has at least of seven lineages, and a range of sublineages. Gap Statement. Allelic multilocus sequence typing (MLST) has been very successfully applied for typing and epidemiology of pathogenic bacteria, but its adoption for fungal pathogens is more limited. Genome sequence-based MLST approaches promise much a higher resolution for comparative phylogenetics, but there are only a small number of genome assemblies available for H. capsulatum. Aim. To develop and validate genome-based MLST typing for H. capsulatum. Methodology. A total of 406 genome assemblies were generated from publicly available Illumina short-read sequencing samples. Selected genome assemblies were subsequently used to generate MLST schemes with chewBBACA, based on whole genome sequences (wgMLST), NCBI-obtained annotations (annMLST) and annotations obtained with Funannotate (fanMLST), and used for lineage identification, and comparative analyses of population structure. Results. Genome assemblies were successfully generated for 406 H. capsulatum samples, with BUSCO completeness scores of over 98%. A subset of 60 genome assemblies was used to generate and validate MLST schemes based on the 60 genome assemblies, 5 genome annotations obtained from Genbank, and 13 genome annotations generated using Funannotate, resulting in schemes with > 8,000 markers. Each of the MLST schemes allowed reproducible separation of all 7 lineages and 6 H. capsulatum Suramericanum sublineages, without the need for high-performance clusters or powerful computing equipment. Comparison of the different schemes using tanglegrams and scatter plots showed good reproducibility (r2 values of 0.8-1.0). Comparison of total sizes of genome assemblies showed significant lineage-dependent size differences, with the Suramericanum, LAmB, India, Capsulatum and Mississippiense lineages averaging a total genome size of 30-33 million bp, and the Africa and Ohiense lineages averaging at 39 million bp. Conclusion. Genome sequence- and especially genome annotation-based MLST analyses are a promising approach that can be adapted as a rapid, high-throughput method for genomic clustering and epidemiological tracing of pathogenic fungi.
]]></description>
<dc:creator><![CDATA[ van Vliet, A. H. M. ]]></dc:creator>
<dc:date>2026-09-30</dc:date>
<dc:identifier>doi:10.64898/2026.09.27.754713</dc:identifier>
<dc:title><![CDATA[Rapid whole genome- and annotation-based multilocus sequence typing of the fungal pathogen Histoplasma capsulatum]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-30</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.30.755330v1?rss=1">
<title>
<![CDATA[
Evo 2 as a classification machine: evidence from in-context learning and mechanistic interpretability 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.30.755330v1?rss=1
</link>
<description><![CDATA[
In-context learning (iCL) is an emergent capability of Large Language Models (LLMs), allowing them to perform new tasks at inference time using prompt-injected examples. While extensively studied in LLMs, the boundaries of its capabilities and underlying mechanisms remain poorly characterised in genomic Language Models (gLMs). Here, we map the operating regime of Evo 2, a nucleotide-level foundation gLM, across five binary classification tasks spanning biological and artificial sequences. We observe robust iCL on shorter natural sequences (F1=0.902 for miRNA, 0.785 for Toxins), degrading with sequence length, and collapsing at kilobase scale. Strikingly, we find no benefit in model scaling, as the 7B model systematically outperforms the 40B variant. We further find that perplexity, a widely used gLM performance proxy, poorly predicts accuracy. Mechanistic interpretability reconciles these observations: the logit-lens profiling suggests a prediction-generalisation trade-off, while Jacobian Scope indicates models might track prompts' structure rather than the signal-carrying content.
]]></description>
<dc:creator><![CDATA[ Bertolini Agnoletto, L., Curion, F., Petrillo, M., Leoni, G., Ronco, M., Ruiz Serra, V., Consoli, S., Ceresa, M. ]]></dc:creator>
<dc:date>2026-09-30</dc:date>
<dc:identifier>doi:10.64898/2026.09.30.755330</dc:identifier>
<dc:title><![CDATA[Evo 2 as a classification machine: evidence from in-context learning and mechanistic interpretability]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-30</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.25.754218v1?rss=1">
<title>
<![CDATA[
Reconstruction of the Prox gene family evolution in vertebrates reveals multiple lineage-specific gene losses 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.25.754218v1?rss=1
</link>
<description><![CDATA[
Prospero-related homeobox (Prox) genes encode a family of transcription factors that play essential roles in the development of several organs and systems, including the central nervous system, lymphatic endothelium, musculature, and liver. Despite their developmental importance, the evolutionary history of the vertebrate Prox gene family remains poorly understood. In this study we combined phylogenetic and synteny analysis to characterise the evolution of the Prox family in vertebrates. Our results reveal that two to three Prox subfamilies were already present in the last common ancestor of jawed vertebrates. We clarify the identity and evolutionary relationships of well-studied members of this family and identify multiple independent losses of Prox2 and Prox3 genes in specific vertebrate lineages. Furthermore, we uncover evidence for the existence of a fourth Prox gene in the ancestral vertebrate genome, which was subsequently lost. Overall, this study provides the first comprehensive analysis of the evolutionary history of the vertebrate Prox gene family and establishes a foundations for future studies on the functional roles of these genes.
]]></description>
<dc:creator><![CDATA[ Panara, V., Leyhr, J., Koltowska, K., Haitina, T. ]]></dc:creator>
<dc:date>2026-09-30</dc:date>
<dc:identifier>doi:10.64898/2026.09.25.754218</dc:identifier>
<dc:title><![CDATA[Reconstruction of the Prox gene family evolution in vertebrates reveals multiple lineage-specific gene losses]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-30</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.28.755014v1?rss=1">
<title>
<![CDATA[
Gene flux shapes diversity and evolution of the ancient 17q21.31 inversion polymorphism 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.28.755014v1?rss=1
</link>
<description><![CDATA[
A hallmark of chromosomal inversions is that they suppress recombination between haplotypes, allowing inversion haplotypes to persist as single co-inherited units. To determine the extent to which inversions nevertheless permit genetic exchange, we investigated a common 979-kb inversion polymorphism at the human 17q21.31 locus. This locus exhibits deep divergence between the reference (H1) and inverted (H2) haplotypes, extensive segmental duplications (SDs) flanking the inversion, and association with neurodegenerative diseases, developmental disorders, and fertility-related phenotypes. Using single-cell sperm genome sequencing data, we directly measured recombination rates between H1 and H2 haplotypes and found near-complete suppression of single-crossover events between the haplotypes. The rare single crossovers that did occur were mediated by non-allelic homologous recombination between shared H1 and H2 SDs, generating novel duplication architectures. In contrast, two-switch events consistent with gene conversion or double crossovers, spanning 17-150 kb, occurred throughout the inversion at rates exceeding genome-wide estimates for events of comparable size. Consistent with recurring genetic exchange, we identified 99 distinct H1-H2 recombinant haplotypes segregating in All of Us genomes, including 26 with combinations of H1 and H2 SDs. These recombinant haplotypes facilitated dissection of the inversion's effects on fertility-related phenotypes; using a large parent-embryo dataset, we found that H2 additively increases female crossover rates across chromosomes and that KANSL1 duplications do not explain this effect. Finally, ancestral recombination graphs dated H1-H2 gene flux (the exchange of genetic material between alternative arrangements) to approximately 100-500 thousand years ago, revealing that H1 and H2 haplotypes have co-segregated for at least half a million years. Together, these results demonstrate that inversions can be permeable barriers to recombination, with ongoing gene flux influencing the diversity and evolution of inversion polymorphisms.
]]></description>
<dc:creator><![CDATA[ Harringmeyer, O. S., Biddanda, A., McCoy, R. C., Akey, J. M. ]]></dc:creator>
<dc:date>2026-09-30</dc:date>
<dc:identifier>doi:10.64898/2026.09.28.755014</dc:identifier>
<dc:title><![CDATA[Gene flux shapes diversity and evolution of the ancient 17q21.31 inversion polymorphism]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-30</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.24.754104v1?rss=1">
<title>
<![CDATA[
Integrated transcriptomic analysis identifies Paternally Expressed Gene 10 as a novel candidate biomarker associated with aggressive features in liposarcoma 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.24.754104v1?rss=1
</link>
<description><![CDATA[
Abstract Background Liposarcoma is a rare cancer, comprising four major histological subtypes: well-differentiated, dedifferentiated, myxoid and pleomorphic liposarcoma. Therapeutic options remain limited for advanced disease, and additional biomarkers that may support biological characterization, diagnosis, or risk stratification are needed. Methods Bulk RNA sequencing (RNA-seq) was performed on paired tumour and adjacent non-tumour tissues from an analysable subset of 16 patients (22 tumour specimens) within a clinicopathological cohort of 21 patients. Differential gene expression analysis, principal component analysis, gene set enrichment (GSEA) analysis and single-sample GSEA (ssGSEA) were used to define subtype-specific transcriptional programs and imprinted-gene dysregulation. PEG10 expression was validated by RT-qPCR and RNA fluorescence in situ hybridization (FISH), and its cellular localization was explored by combined FISH and immunofluorescence for P53, vimentin, Ki67, and CD34. PEG10 expression in liposarcoma was further investigated in an external single-cell RNA-sequencing dataset of dedifferentiated liposarcoma. An exploratory survival analysis was performed in the pan-sarcoma TCGA-SARC cohort. Results Transcriptomic profiling identified distinct subtype-specific gene expression patterns and broad dysregulation of imprinted genes in liposarcoma compared with adjacent non-tumour tissue. Among all the imprinted genes, PEG10 showed a marked tumour-to-normal expression difference and good discriminatory performance between tumour and normal tissues in the discovery cohort (AUC=0.867). In the TCGA sarcoma cohort high PEG10 expression was associated with shorter overall survival. Immuno-FISH and single-cell analyses revealed that PEG10 was predominantly detected in tumour and proliferating cell populations. Importantly, PEG10 expression was higher in high-grade tumours. In bulk RNA-seq data, PEG10 expression correlated with hypoxia, extracellular matrix remodelling and invasion ssGSEA signatures. In the external single-cell RNA-seq dataset PEG10high tumours showed a reduced immune cell infiltration. Conclusions This study provides an integrated transcriptomic characterisation of liposarcoma and identifies PEG10 as a candidate tumour-associated biomarker linked to aggressive disease features. The diagnostic and prognostic value of PEG10 requires validation in larger, independent, subtype-balanced liposarcoma cohorts, and its potential therapeutic relevance requires direct functional investigation.
]]></description>
<dc:creator><![CDATA[ Turiello, R., Scognamiglio, G., Bonicelli, A., Gitano, G., Manco, E., Olivieri, S., Barbato, A., Fruggiero, C., Salati, M., Iuliano, A., Cerrone, M., Arcucci, A., Martorana, A., De Pietro, G., Santella, E., Picozzi, F., Cantile, M., Tafuto, S., Budillon, A., Franco, B., Carotenuto, P. ]]></dc:creator>
<dc:date>2026-09-29</dc:date>
<dc:identifier>doi:10.64898/2026.09.24.754104</dc:identifier>
<dc:title><![CDATA[Integrated transcriptomic analysis identifies Paternally Expressed Gene 10 as a novel candidate biomarker associated with aggressive features in liposarcoma]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-29</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.24.753534v1?rss=1">
<title>
<![CDATA[
Disease relevance and replicability of deep learning gene expression prediction 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.24.753534v1?rss=1
</link>
<description><![CDATA[
Recent deep learning (DL) models predict average gene expression levels from DNA sequences with high overall correlation to measured values. We examine these DL models through the lens of disease research. The seemingly high overall performance of DL models is largely due to capturing whether genes are "On" or "Off", and to a lesser extent, disease-relevant expression level changes for "On" genes. Indeed, the more bimodal the gene expression distribution, the better the reported performance. We track the extent of this issue across tissues, molecular systems, and cancers. Compounding the model evaluation issue, we find inconsistencies between the published code and reported performance, highlighting the importance of versioning and publishing performance evaluation code. These findings indicate that the high reported performance of popular DL models falls unexpectedly short in disease applications, and that the problem of personalized genomic prediction remains far from solved in a disease context.
]]></description>
<dc:creator><![CDATA[ Zhang, A., Tasaki, S., Connell, D., Ng, B., Gaiteri, C. ]]></dc:creator>
<dc:date>2026-09-29</dc:date>
<dc:identifier>doi:10.64898/2026.09.24.753534</dc:identifier>
<dc:title><![CDATA[Disease relevance and replicability of deep learning gene expression prediction]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-29</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.24.754018v1?rss=1">
<title>
<![CDATA[
SCAR: Controlled mutations, ancient DNA damage, and fragmentation of fasta and fastq sequences 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.24.754018v1?rss=1
</link>
<description><![CDATA[
Summary: Controlled modification of sequencing data is important for reproducible benchmarking, particularly when evaluating analyses affected by read fragmentation, divergent reference genomes and ancient DNA damage. SCAR introduces controlled mutations, fragmentation and position-specific damage into FASTA and FASTQ data. All features can be used separately or in combination, and empirical fragment-length and mismatch profiles can be supplied to closely reproduce the properties of targeted datasets. Validation analyses showed the successful implementation of the requested sequence changes, and their impact on read mapping and heterozygosity analyses. Availability and Implementation: SCAR is implemented in C++ and the latest code is available at https://github.com/Madshartmann1/SCAR. The version described in this article is archived at Zenodo at https://doi.org/10.5281/zenodo.22789736 , with validation material available at https://doi.org/10.5281/zenodo.21915766.
]]></description>
<dc:creator><![CDATA[ Hartmann, M., Westbury, M. V. ]]></dc:creator>
<dc:date>2026-09-29</dc:date>
<dc:identifier>doi:10.64898/2026.09.24.754018</dc:identifier>
<dc:title><![CDATA[SCAR: Controlled mutations, ancient DNA damage, and fragmentation of fasta and fastq sequences]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-29</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.24.753990v1?rss=1">
<title>
<![CDATA[
Alcohol Use Disorder and Smoking-Associated Molecular Alterations in Human Prefrontal Cortex in Single Nucleus (sn) RNA-Seq 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.24.753990v1?rss=1
</link>
<description><![CDATA[
Chronic smoking worsens alcohol related brain injury and impairs neurocognitive recovery, yet the cell type specific molecular interactions between alcohol use disorder (AUD) and smoking in the human brain remain largely unexplored. We analyzed single-nucleus RNA-seq from the prefrontal cortex of 73 individuals (614,932 nuclei, 24 cell types), comparing AUD versus controls, smoking versus controls, and combined AUD+smoking versus controls. In excitatory neurons, neuroinflammatory, complement, GPCR, kinase, proteostatic, and extracellular-matrix pathways confined to one or two subtypes in either condition alone expanded across nearly all subtypes when AUD and smoking co-occurred. Inhibitory neurons showed a parallel but distinct expansion, additionally recruiting translational regulation, RNA splicing, and deubiquitination pathways. Smoking alone produced widespread repression of bioenergetic, proteostatic, and synaptic pathways in excitatory neurons, whereas AUD produced mixed pathway activation and repression in excitatory neurons and broad inflammatory activation in inhibitory neurons. In the combined condition, microglia exhibited paradoxical repression of phagocytosis, TYROBP signaling, and translational machinery, consistent with an exhaustion like state, accompanied by an apparently compensatory shift in inflammatory signaling toward astrocytes, oligodendrocyte precursor cells, and oligodendrocytes, which showed coordinated inflammatory activation. Vascular cells showed an endothelial-VLMC dichotomy, with endothelial repression of RNA processing and proteostasis contrasting with perivascular activation. Our analysis identifies bioenergetic and proteostatic stress as the predominant signature of smoking, whereas immune and metabolic activation characterize AUD. In AUD+Smoking, these alterations extend across more neuronal subtypes and include neuroinflammation, complement activation, altered signaling, proteostatic stress, and post-transcriptional remodeling associated with greater cortical dysfunction in co-occurring AUD and smoking.
]]></description>
<dc:creator><![CDATA[ Joshi, A., Sanna, P. P. ]]></dc:creator>
<dc:date>2026-09-29</dc:date>
<dc:identifier>doi:10.64898/2026.09.24.753990</dc:identifier>
<dc:title><![CDATA[Alcohol Use Disorder and Smoking-Associated Molecular Alterations in Human Prefrontal Cortex in Single Nucleus (sn) RNA-Seq]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-29</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.24.753865v1?rss=1">
<title>
<![CDATA[
Evolutionary dynamics of the insertion sequence IS6110 in the Mycobacterium tuberculosis complex 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.24.753865v1?rss=1
</link>
<description><![CDATA[
Insertion sequences (IS) are the most common type of transposable element in prokaryotes and shape the structure of genomes through transposition and by providing a substrate for recombination. Despite the mutational impact of IS, the evolutionary dynamics of most elements in host species remain unknown. Here we study the dynamics of IS6110 in 10,000 strains of the Mycobacterium tuberculosis complex (MTBC). We developed a tool that allows the detection and comparison of IS insertions from short reads without using a reference genome. Using ancestral state reconstruction (ASR) on presence-absence patterns of IS6110, we describe the distribution of copy numbers (CNs) in the MTBC, infer birth rates of the element, and identify genomic regions with large numbers of parallel IS6110 insertions. Copy numbers in the MTBC range from 1 in some clades to more than 30 in strains of La3 (M. orygis). IS6110 birth rates scale approximately linearly with copy number and are elevated on terminal branches, consistent with the delayed action of purifying selection. A key characteristic of IS6110 is its occurrence in hotspots: the 5% most frequently targeted regions account for half of all independent insertion events. The motif 5'-TCTCAAAW-3' is enriched around target sites and in hotspots, suggesting that the accumulation of insertions in these regions results through a combination of non-random insertion and purifying selection in other regions. To conclude the study, we propose a niche constraints model according to which the distribution of IS6110 in the MTBC is governed by the rarity of regions that have both suitable DNA properties and little functional value for the host.
]]></description>
<dc:creator><![CDATA[ Stritt, C., Loiseau, C., Kalkan, S., Borrell, S., Brites, D., Gagneux, S. ]]></dc:creator>
<dc:date>2026-09-29</dc:date>
<dc:identifier>doi:10.64898/2026.09.24.753865</dc:identifier>
<dc:title><![CDATA[Evolutionary dynamics of the insertion sequence IS6110 in the Mycobacterium tuberculosis complex]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-29</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.23.753885v1?rss=1">
<title>
<![CDATA[
Identifying, phasing, and structurally annotating sex chromosomes for genome assemblies using CBS-tools 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.23.753885v1?rss=1
</link>
<description><![CDATA[
A complete reference genome for species with chromosomally-determined separate sexes should contain scaffolds for all sex chromosome homologs. However, sex chromosomes present distinct computational challenges compared to autosomes. Here we present a k-mer based analysis that utilizes whole-genome sequencing of a few sex-identified isolates: Cytogenetics-By-Sequencing (CBS) tools. Unlike other approaches that typically address one aspect of the sex chromosomes, CBS-tools strives to guide users from the discovery of the heterogametic sex through identifying the sex-determination region (SDR). The core of CBS-tools is automated quantification of sex-specific k-mers in order to predict the heterogametic sex. Using publicly-available datasets, CBS-tools correctly identified the known sex-system of the 31 species tested. Additionally, we used these k-mers to verify and correct phasing of sex chromosomes between haplotypes in species representing different sex-systems. Finally, we used these k-mers to delimit the SDR boundary using an interactive web platform. CBS-tools was developed with previously unexplored sex chromosome systems in mind, but is also suitable for well-examined sex chromosome pairs.
]]></description>
<dc:creator><![CDATA[ Whitt, L., Akozbek, L., Bentz, P. C., Armstrong, E., Harkess, A., Carey, S. B. ]]></dc:creator>
<dc:date>2026-09-29</dc:date>
<dc:identifier>doi:10.64898/2026.09.23.753885</dc:identifier>
<dc:title><![CDATA[Identifying, phasing, and structurally annotating sex chromosomes for genome assemblies using CBS-tools]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-29</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.23.753695v1?rss=1">
<title>
<![CDATA[
Genomic correlates of metastatic competence and progression in human melanoma 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.23.753695v1?rss=1
</link>
<description><![CDATA[
Genomic events and their timing that grant a primary tumour the competence to disseminate remain poorly defined. We performed sequencing of 247 stage I/II primary cutaneous melanomas (CMs) and 60 matched metastases without intervening therapy from a prospectively followed registry cohort with a median followup of 92 months, integrating copy-number, mutational, protein and spatial-transcriptomic analyses. Relapse was not distinguished by oncogenic point mutations, which were largely shared between primaries and metastases, but by somatic copy-number alterations (SCNAs) and global chromosomal instability. We defined OncoCycle, a six-gene copy-number signature (amplification of CDK4, MCL1 and CD276; biallelic loss of CDKN2A, CDKN2B and TP53BP1) that predicted relapse independently of established clinicopathological features in melanoma, and a pan-cancer analysis. In matched pairs, metastatic progression was driven by continued copy-number evolution and reduction in intra-tumoural heterogeneity, rather than by acquired point mutations, and OncoCycle alterations from primary tumours were preserved in metastasis seeding clones. Clonal reconstruction revealed both monoclonal and polyclonal metastasis seeding, and spatial transcriptomics resolved copy-number-defined metastatic subclones occupying and programming distinct immune and stromal niches. Thus, metastatic competence was primed early by focal SCNAs on a background of chromosomal instability, elaborated by continued copy-number evolution during dissemination and spatio-temporal interactions with the tumour-microenvironment.
]]></description>
<dc:creator><![CDATA[ Chatziioannou, E., Luthria, K., Shah, P., Armeanu-Ebinger, S., Admard, J., Schroeder, C., Forchhammer, S., Stihler, F., Zhang, X., Maiche, S., Mankel, A. B., Stadelmaier, J., Amaral, T., Leiter, U., Maczey, E., Mattern, S., Bonzheim, I., Fend, F., Garbe, C., Heneka, Y., Krauss, M., Pop, O. T., Schreieder, L., Haferkamp, S., Stratigos, A. J., Schuerch, C. M., Ossowski, S., Martus, P., Nahnsen, S., Sinnberg, T., Flatz, L., Riess, O., Izar, B., Roecken, M. ]]></dc:creator>
<dc:date>2026-09-29</dc:date>
<dc:identifier>doi:10.64898/2026.09.23.753695</dc:identifier>
<dc:title><![CDATA[Genomic correlates of metastatic competence and progression in human melanoma]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-29</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.26.754026v1?rss=1">
<title>
<![CDATA[
PfPHAST: Plasmodium falciparum Public Health Amplicon Sequencing Tool, a Streamlined Panel for Malaria Genomic Surveillance 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.26.754026v1?rss=1
</link>
<description><![CDATA[
Genomic tools can support malaria control policy through surveillance of Plasmodium falciparum populations, tracking antimalarial drug resistance, pfhrp2/3 deletions that compromise rapid diagnostic tests, and selection at the circumsporozoite protein (PfCSP) vaccine target, as well as through molecular correction of therapeutic efficacy studies (TES). Multiplex Amplicons for Drug, Diagnostic, Diversity, and Differentiation Haplotypes using Targeted Resequencing (MAD4HatTeR), a comprehensive amplicon sequencing panel covering up to 276 targets, supports these applications but is tailored to research rather than routine programmatic use. We developed P. falciparum Public Health Amplicon Sequencing Tool (PfPHAST), a 56-target derivative of MAD4HatTeR spanning drug resistance loci, pfhrp2/3 deletion, PfCSP genotyping, non-falciparum species identification, and 20 high-heterozygosity microhaplotype loci for TES classification. We compared PfPHAST and MAD4HatTeR using laboratory strain controls, including two-strain dilution series and a five-strain mixture, across parasite densities of 100 to 10,000 parasites/L. At matched per-target depth, PfPHAST achieved a higher quality-control pass rate than MAD4HatTeR (94.4% versus 90.0%) and distributed reads more evenly across targets. The panels showed comparable recall and precision for drug resistance codons and microhaplotypes, reaching near-complete recall above 40% within-sample allele frequency (WSAF) at all densities, with reduced sensitivity for minor alleles below 10% WSAF at low parasite density in both panels. Observed and expected WSAF correlated strongly for both panels, and both resolved a five-strain polyclonal mixture, including a 5% minor strain. By concentrating sequencing capacity on targets of greatest programmatic relevance, PfPHAST offers a scalable, lower-cost alternative to comprehensive research panels without sacrificing performance on shared targets, complementing MAD4HatTeR for routine molecular malaria surveillance.
]]></description>
<dc:creator><![CDATA[ Semakuba, F. D., Nsengimaana, B., Katairo, T., Mbabazi, M., Murie, K., Hathaway, N., Wiringilimaana, I., Muwanika, J. V., Lum, K., Davlieva, M., Asua, V., Kiyaga, S., Tukwasibwe, S., Nakasaanya, J., Achom, K. B., Ayitewala, A., Mulondo, J., Watyekele, E., Nsobya, S. L., Kamya, M. R., Agaba, B. B., Rosenthal, P. J., Conrad, M. D., Ssewanyana, I., Greenhouse, B., Aranda-Diaz, A., Briggs, J. ]]></dc:creator>
<dc:date>2026-09-29</dc:date>
<dc:identifier>doi:10.64898/2026.09.26.754026</dc:identifier>
<dc:title><![CDATA[PfPHAST: Plasmodium falciparum Public Health Amplicon Sequencing Tool, a Streamlined Panel for Malaria Genomic Surveillance]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-29</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.28.754846v1?rss=1">
<title>
<![CDATA[
The Flame Retardant Triphenyl Phosphate Induces Varying Responses in Genetically Diverse Mouse Induced Pluripotent Stem Cells 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.28.754846v1?rss=1
</link>
<description><![CDATA[
Genetic variation impacts biological response to chemical and environmental exposures, contributing to differences in resistance or susceptibility. Forward genetic screens of genetically diverse cell populations can be used to identify the precise genetic variants that drive response variation. Cell lines from laboratory mouse genetic reference populations like the Diversity Outbred (DO) population are well powered for genetic mapping. Diverse stem cell panels are especially beneficial for studying early developmental effects of chemical exposures like triphenyl phosphate (TPHP), an organophosphate flame retardant linked to adverse developmental effects, including altered cell cycle in stem cells. Self-renewal and pluripotency are defining properties of stem cells during development; therefore, it is important to understand how chemicals like TPHP may alter these critical characteristics. To better understand the influence of genetic variation on TPHP exposure response, we utilized a genetically diverse panel of DO induced pluripotent stem cells (DO iPSCs). Our analyses characterize interline variation in gene expression through differential gene expression analysis and expression quantitative trait locus (eQTL) mapping. We identified genomic loci that contribute to variation in gene expression following TPHP exposure, including a regulatory hotspot on chromosome 15. Gene set enrichment highlighted several pathways affected by TPHP including lipid metabolism, steroid hormone signaling, and cell cycle. Our study shows that genetic variation modulates the effects of TPHP on gene expression and cell cycle in stem cells. Additionally, our work demonstrates that genetically diverse cell panels like the DO iPSC resource offer a tractable and scalable approach for identifying the genetic determinants of chemical toxicity.
]]></description>
<dc:creator><![CDATA[ Armstrong, M., Janeczko, K., Dewey, H. B., Czechanski, A., Chen, Q., Swanzey, E., O'Connor, C., Martin, W., Aydin, S. B., Munger, S. C., Reinholdt, L. G. ]]></dc:creator>
<dc:date>2026-09-28</dc:date>
<dc:identifier>doi:10.64898/2026.09.28.754846</dc:identifier>
<dc:title><![CDATA[The Flame Retardant Triphenyl Phosphate Induces Varying Responses in Genetically Diverse Mouse Induced Pluripotent Stem Cells]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-28</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.22.753638v1?rss=1">
<title>
<![CDATA[
Structural polymorphism and population-variable coding capacity of HERV-K(HML-2) in human pangenomes 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.22.753638v1?rss=1
</link>
<description><![CDATA[
Approximately 8% of the human genome is derived from ancient retroviral infections. The most recently integrated of these endogenous retroviruses is the HERV-K(HML-2) clade, whose expression has been associated with cancer, amyotrophic lateral sclerosis, and embryogenesis. Studies of HERV expression, particularly HML-2, have relied predominantly on short-read sequencing. However, the high similarity among HML-2 proviruses prevents many short reads from being assigned uniquely to individual loci. We therefore compared haplotype-resolved long-read genome assemblies from 292 donors to resolve variation in proviral structure and coding capacity. Several loci previously thought to be fixed were structurally polymorphic. Tandem arrays occurred at 13 loci and contained up to six proviral copies in a single array. At 8q11.23, we identified a previously undescribed full-length provirus in one haplotype. All 583 other haplotypes carried a solo-LTR. We found that standard reference genomes failed to represent the coding capacity retained in many individuals, whose proviruses contained intact open reading frames despite disruptive mutations in the reference sequences. Short-read genotypes left 32.5% of the tested donor-variant pairs unresolved at sites associated with viral reading frames. These findings show why HML-2 expression must be interpreted in the context of the structural and coding alleles each individual carries.
]]></description>
<dc:creator><![CDATA[ Murimi-Worstell, D. A., Freeman, M., Minkina, A., Vollger, M. R., Stergachis, A. B., Coffin, J. M. ]]></dc:creator>
<dc:date>2026-09-28</dc:date>
<dc:identifier>doi:10.64898/2026.09.22.753638</dc:identifier>
<dc:title><![CDATA[Structural polymorphism and population-variable coding capacity of HERV-K(HML-2) in human pangenomes]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-28</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.22.753313v1?rss=1">
<title>
<![CDATA[
Structural variation in repeat elements is widespread in normal human tissues and in tumorigenesis 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.22.753313v1?rss=1
</link>
<description><![CDATA[
Somatic mosaicism contributes to genomic variation, yet postzygotic structural variants remain under-characterized. We performed long- and short-read WGS from multiple individuals (n=47 normal tissues; n=168 samples) and identified mosaic structural variants in all individuals and germ layers, impacting a median 285.2 kb/genome. Nearly half of breakpoints were independently validated, with tissue distributions reflecting both early and late developmental origins. Most mosaic variants were repeat-mediated and 8.3% overlapped functional elements, an enrichment compared to germline variants. To extend these analyses in samples where long-read sequencing is infeasible, we measured repeat alterations from short-read sequencing, recapitulating mosaic tissue-specific differences. We characterized tumor- and tissue- specific variation in repeats across 15 cancer types and found tumor-related repeat variation to be similar in scale to that of normal mosaic variation. Tracking repeat changes in cell-free DNA provided a noninvasive approach for tumor monitoring. Our analyses revealed widespread repeat-driven structural variation in health and disease.
]]></description>
<dc:creator><![CDATA[ Annapragada, A. V., White, J., Orjuela, H., Bartolomucci, A., Eastman, A., Koul, S., Lebarbenchon, K., Bruhm, D., Short, S., Boyapati, K., Niknafs, N., Norton, C., Girish, V., Vulpescu, N., Velculescu, S., Velculescu, J., Adleff, V., Nelson, A., Foda, Z., Winterhoff, B., Drapkin, R., Schatz, M., Phallen, J., Scharpf, R., Velculescu, V. ]]></dc:creator>
<dc:date>2026-09-28</dc:date>
<dc:identifier>doi:10.64898/2026.09.22.753313</dc:identifier>
<dc:title><![CDATA[Structural variation in repeat elements is widespread in normal human tissues and in tumorigenesis]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-28</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.26.754616v1?rss=1">
<title>
<![CDATA[
Whole-Genome Metagenomics Insights Revealed a Uranium Bio-remediating Cross-Domain Microbiome in the Dhala Impact Structure, India 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.26.754616v1?rss=1
</link>
<description><![CDATA[
Meteorite impact structures on earth represent unique terrestrial analogues of extreme planetary environments where fractured lithologies, hydrothermal alteration, elevated concentrations of radionuclides, and prolonged water-rock interactions create ecological niches for specialized microbial communities. The Palaeoproterozoic Dhala impact structure in north-central India is characterized by uranium-bearing impactites and possesses a very high concentration of Uranium (up to 99.5 ppm) that makes it a unique natural laboratory to investigate microbial adaptation to radioactive and metal-rich geological environments. In addition, the Dhala structure harbors specialized radiotoxic hydrothermal mineral phases such as coffinite and pitchblende, which show a globally distinct and biologically unexplored niche. The present investigation is the first in-depth, whole-genome metagenomic study of soil and water samples of uranium-rich sites from the Dhala structure, exploring the whole microbial community (Bacteria, Archaea, Eukaryotes, and Viruses) up to the species level present in the samples. The cross-domain metagenomics analysis identified a highly unique and resilient microbiome that consists of bacteria, archaea, eukaryotes, and viruses, coevolved under prolonged radiological and multiple metal stress. Functional metagenomic analysis revealed numerous metabolic adaptations responsible for their survival, homeostasis, and biomineralization in situ. The deep sequencing method further enabled the identification of 58 functional genes directly involved in uranium bioremediation. Furthermore, this study revealed important pathways for radionuclide reduction, cellular efflux, and biotransformation. Novel findings of the present study point out the in-depth evolutionary relations of biosphere-geosphere interaction within an ancient hypervelocity impact-generated ecosystem. In a nutshell, the Dhala extremophiles with novel genetic makeup present an untapped and promising reservoir for the development of eco-friendly, next-generation microbial bioremediation strategies to efficiently tackle the global anthropogenic contamination of uranium and the challenges of nuclear waste management. Thus, the present study establishes Dhala structure as an unparalleled site for environmental geomicrobiology, microbial evolution, and astrobiology.
]]></description>
<dc:creator><![CDATA[ Kapinder,, Dwivedi, S., Singh, A. K., Pati, J. K., Kumar, P. ]]></dc:creator>
<dc:date>2026-09-28</dc:date>
<dc:identifier>doi:10.64898/2026.09.26.754616</dc:identifier>
<dc:title><![CDATA[Whole-Genome Metagenomics Insights Revealed a Uranium Bio-remediating Cross-Domain Microbiome in the Dhala Impact Structure, India]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-28</prism:publicationDate>
<prism:section></prism:section>
</item>
</rdf:RDF>
