<?xml version="1.0" encoding="UTF-8" ?>
<rdf:RDF xmlns:admin="http://webns.net/mvcb/" xmlns="http://purl.org/rss/1.0/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:prism="http://purl.org/rss/1.0/modules/prism/" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:syn="http://purl.org/rss/1.0/modules/syndication/">
<channel rdf:about="https://biorxiv.org">
<admin:errorReportsTo rdf:resource="mailto:biorxiv@cshlpress.edu"/>
<title>bioRxiv Subject Collection: Genomics Bioinformatics</title>
<link>https://biorxiv.org</link>
<description>
This feed contains articles for bioRxiv Subject Collection "Genomics Bioinformatics"
</description>

<items>
<rdf:Seq>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.31.747971v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.31.748368v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.31.748303v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.31.748319v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.31.748316v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.31.748304v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.31.748226v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.31.747916v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.31.748244v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.31.748380v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.31.748410v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.31.748429v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.31.748296v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.31.748400v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.31.748267v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.02.748787v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.28.747928v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.28.747864v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.30.748130v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.30.748136v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.30.743167v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.28.747909v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.31.748056v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.30.748103v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.31.748297v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.31.747784v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.30.748028v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.30.748098v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.30.748061v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.30.748081v1?rss=1"/>
</rdf:Seq>
</items>
<prism:eIssn/>
<prism:publicationName>bioRxiv</prism:publicationName>
<prism:issn/>

<image rdf:resource=""/>
</channel>
<image rdf:about="">
<title>bioRxiv</title>
<url>https://www.biorxiv.org/sites/default/files/bioRxiv_article.jpg</url>
<link>https://www.biorxiv.org</link>
</image>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.31.747971v1?rss=1">
<title>
<![CDATA[
Global tree encoding of atlas-scale single-cell genomics 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.31.747971v1?rss=1
</link>
<description><![CDATA[
The rapid expansion of single-cell genomic datasets has led to the compilation of biological resources comprising hundreds of millions of cells across tissues, developmental stages, and disease states. This has underscored the need for scalable and interpretable data representations that preserve the complex relationships and multi-scale organization of cellular states, while remaining computationally tractable at atlas scale. Existing approaches based on discrete abstractions have enabled cell annotation, clustering, and trajectory inference, but are often optimized for local inference tasks and may obscure continuous cellular relationships and multi-resolution structure within complex transcriptional and other genomic landscapes. Moreover, increasing dataset sizes often require information-reduction strategies such as random downsampling, limiting the resolution of rare cell populations and heterogeneous cellular states. Here, we present MILK, a scalable computational framework that organizes high-dimensional single-cell populations into unified tree representations. Across large-scale transcriptomic atlases, MILK enables representative subsampling with preserved information, supporting the tractable application of existing algorithms for tasks including deep generative model training and foundation model benchmarking. Additionally, MILK enables holistic, multi-resolution analyses that capture global developmental trajectories, characterize disease-associated cellular perturbations across tissues, and facilitate comparison of transcriptional programs across species within a coherent hierarchical framework. Together, these results establish the hierarchical organization of biological data as a scalable and unifying representation of cellular identity, enabling integrative analysis of single-cell genomic data across diverse contexts.
]]></description>
<dc:creator><![CDATA[ Kiyota, B., Lee, C., Yao, H., Yachie, N. ]]></dc:creator>
<dc:date>2026-09-04</dc:date>
<dc:identifier>doi:10.64898/2026.08.31.747971</dc:identifier>
<dc:title><![CDATA[Global tree encoding of atlas-scale single-cell genomics]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-04</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.31.748368v1?rss=1">
<title>
<![CDATA[
Visualizing and integrating linear and graph pangenomes at the Maize Genetics and Genomics Database 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.31.748368v1?rss=1
</link>
<description><![CDATA[
Genome visualization has existed for as long as assembled genomes, with an array of tools and approaches to meet researchers' needs. The genome browser is one of these tools, and it has become a critical analytical resource for geneticists, genome biologists, and breeders. The evolution of genome browser capabilities is ongoing, and the USDA-ARS Maize Genetics and Genomics Database (MaizeGDB) has implemented JBrowse2, the most recent iteration of JBrowse. It includes multiple genome browser and alignment views, demonstrating pangenome visualization capability and permitting sophisticated functional characterization of maize loci. Described here also is MaizeGDB's new Pangenome Viewer, which integrates pangenome graph views with linear browser visualization to obtain an on-the-fly, interactive visual snapshot of structural variation across a pangenome at user-selected loci, with zoom capabilities and statistical and variant information for each subgraph. Finally, we demonstrate how different MaizeGDB pangenome pipelines complement one another to help guide accurate analyses. This functionality can lead to more precise characterization of loci that confer important agronomic traits, resulting in better outcomes for farmers and the public.
]]></description>
<dc:creator><![CDATA[ Portwood, J. L., Cannon, E. K., Haley, O. C., Tibbs-Cortes, L. E., Andorf, C. M., Woodhouse, M. R. ]]></dc:creator>
<dc:date>2026-09-04</dc:date>
<dc:identifier>doi:10.64898/2026.08.31.748368</dc:identifier>
<dc:title><![CDATA[Visualizing and integrating linear and graph pangenomes at the Maize Genetics and Genomics Database]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-04</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.31.748303v1?rss=1">
<title>
<![CDATA[
siProGenA: Generative siRNA Candidate Construction via Position Proposal and Guide Generation 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.31.748303v1?rss=1
</link>
<description><![CDATA[
Small interfering RNAs (siRNAs) are short guide RNAs that recruit the RNA-induced silencing complex (RISC) to complementary target sites on messenger RNAs (mRNAs), triggering Ago2-mediated cleavage and gene silencing. siRNA design requires compact candidate sets that cover a target while preserving efficacy, specificity, and practical sequence constraints. Existing pipelines usually enumerate candidate windows, assign a canonical guide to each window, and then rank preconstructed siRNA--mRNA pairs. This has produced strong pairwise efficacy predictors, but leaves a candidate-construction gap: candidate positions and guide sequences are fixed before the model begins to rank them. We address this gap by decomposing siRNA candidate construction into two generative decisions: where to place candidates within an mRNA segment, and what constrained guide variants to consider at a candidate position. We instantiate this framework as siProGenA, using a Discrete Denoising Diffusion Probabilistic Model (D3PM) for mRNA-conditioned position proposal and a Bayesian Flow Network (BFN) for temperature-controlled guide generation. On 62 positive test segments, the diversity-aware final library reaches Hit@1 = 0.790 and Hit@5 = 0.903. In a measured-site controlled Stage~2 evaluation, seed- and cleavage-preserving variants outscore the canonical complement for 89.8% of measured sites, with supporting gains across additional computational scorers, random-mismatch controls, and biophysical diagnostics. Together, the results support a modular proposal--generation view of siRNA candidate construction for prioritizing compact candidate sets.
]]></description>
<dc:creator><![CDATA[ Ma, Z., Zhou, J., Wang, R., Deng, Z., Wu, Z., Zheng, Y. ]]></dc:creator>
<dc:date>2026-09-04</dc:date>
<dc:identifier>doi:10.64898/2026.08.31.748303</dc:identifier>
<dc:title><![CDATA[siProGenA: Generative siRNA Candidate Construction via Position Proposal and Guide Generation]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-04</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.31.748319v1?rss=1">
<title>
<![CDATA[
BROOQS: Spectral Methods Resolve Level-1 Hybridization Cycles without Tests of Symmetry 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.31.748319v1?rss=1
</link>
<description><![CDATA[
Modern phylogenomic analyses often seek to reconstruct both vertical and reticulate evolutionary histories. While the prevalence of non-vertical evolution is increasingly appreciated, inferring networks remains conceptually challenging and computationally demanding. Following the success of quartet-based methods for handling gene tree discordance, several quartet-based network inference methods have been developed. A key insight of these methods is that level-1 networks can be constructed by first building a multifurcating tree called tree-of-blobs and then resolving each polytomy into a cycle. This two-step approach makes the problem easier both conceptually and computationally. However, these quartet-based methods often rely on noisy statistical tests of asymmetry in quartet frequencies. Moreover, they either enumerate all quartets, losing some scalability, or subsample them, losing information. We introduce BROOQS, a quartet-based method for resolving trees of blobs into a level-1 phylogenetic network. BROOQS efficiently aggregates information from all quartets around a blob without enumerating them, builds a pairwise similarity matrix, and uses robust spectral ordering algorithms to recover the cyclic ordering without relying on individual quartet symmetry tests. We prove theoretically that our spectral method is consistent under the network multi-species coalescent (NMSC) model. Across simulated and empirical datasets, BROOQS consistently improves accuracy and scalability compared to existing methods and extends to thousands of taxa.
]]></description>
<dc:creator><![CDATA[ Arasti, S., Mirarab, S. ]]></dc:creator>
<dc:date>2026-09-04</dc:date>
<dc:identifier>doi:10.64898/2026.08.31.748319</dc:identifier>
<dc:title><![CDATA[BROOQS: Spectral Methods Resolve Level-1 Hybridization Cycles without Tests of Symmetry]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-04</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.31.748316v1?rss=1">
<title>
<![CDATA[
LEAPING A CONSERVATION GENETICS GAP: COMBINING POPULATION GENETICS AND SOCIAL NETWORK ANALYSIS FOR LUPINUS PERENNIS RESTORATION 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.31.748316v1?rss=1
</link>
<description><![CDATA[
Genomic data can assist conservation planning, but implementation of genetic-based recommendations is rare. Closing this "conservation genetics gap" is crucial for restoration and translocation decisions, particularly for highly fragmented ecosystems like pine barrens. We first analyze the conservation genetics gap in pine barren management organizations using a survey. We then analyze whole-genome data to characterize population structure and diversity in the pine barren keystone plant Lupinus perennis. We find little genetic difference among L. perennis populations in the Northeast United States. While populations show signs of isolation by distance (rho = 0.785, p = 0.001), we find no correlation between population size and heterozygosity (Estimate = 0.003, p = 0.357) or inbreeding (Estimate = 0.028, p = 0.416). We also find no effect of seed source population on forecasted effective population size in restoration simulations. These findings and our recommendation that managers consider non-local seed provenancing for L. perennis restorations in the Northeast have been integrated into SWAP policy in two states. We generally recommend practitioners record seed provenance, save genetic material for future testing, and incorporate fitness metrics into restoration monitoring. To effectively close the conservation genetics gap, researchers providing conservation recommendations must consult with practitioners before designing experiments.
]]></description>
<dc:creator><![CDATA[ Kimball-Rhines, C., Taveras-Guzman, S., Mavrommati, G., Moyers, B. T. ]]></dc:creator>
<dc:date>2026-09-04</dc:date>
<dc:identifier>doi:10.64898/2026.08.31.748316</dc:identifier>
<dc:title><![CDATA[LEAPING A CONSERVATION GENETICS GAP: COMBINING POPULATION GENETICS AND SOCIAL NETWORK ANALYSIS FOR LUPINUS PERENNIS RESTORATION]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-04</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.31.748304v1?rss=1">
<title>
<![CDATA[
Poly Pipeline: A Polyvalent Spatial Transcriptomics Workflow Validated Across Polyploid and Diploid Organisms 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.31.748304v1?rss=1
</link>
<description><![CDATA[
Spatial transcriptomics (ST) has emerged as a transformative approach for visualizing tissue landscapes, yet it faces significant challenges regarding data standardization, sparsity, and the analysis of complex genomes, particularly polyploid plants. To address these limitations, we introduce Poly Pipeline, a robust and universal bioinformatic workflow designed to streamline analysis across diverse plant and animal genomes. The pipeline integrates a comprehensive converter for proprietary formats, clustering algorithms, and hdWGCNA co-expression networks, which indirectly preserves the expression signatures of low-expressed duplicated genes. Benchmarking across datasets from wheat, rice, Arabidopsis, and mouse demonstrated the broad applicability of the pipeline in identifying relevant clusters, showing effectiveness across diverse organisms and data types. By providing a unified and reproducible framework, Poly Pipeline addresses a critical gap in analyzing genomic redundancy, especially that related to polyploidy, and promotes FAIR data principles for the broader scientific community.
]]></description>
<dc:creator><![CDATA[ Carvalho, P. C., Millsteed, T., Henry, R. J. ]]></dc:creator>
<dc:date>2026-09-04</dc:date>
<dc:identifier>doi:10.64898/2026.08.31.748304</dc:identifier>
<dc:title><![CDATA[Poly Pipeline: A Polyvalent Spatial Transcriptomics Workflow Validated Across Polyploid and Diploid Organisms]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-04</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.31.748226v1?rss=1">
<title>
<![CDATA[
AltraFlowSOM: A Semi-Supervised Framework for Imaging Mass Cytometry Phenotyping 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.31.748226v1?rss=1
</link>
<description><![CDATA[
Imaging Mass Cytometry (IMC) enables the simultaneous quantification of 40+ protein markers at single cell resolution in tissue, however biologically faithful phenotyping at scale remains a critical bottleneck. Unsupervised clustering fragments coherent populations or conversely merges biologically incoherent ones into a single cluster, supervised classifiers impose a closed vocabulary, and the presence of rare subsets (encoding clinically relevant biology) in conjunction with abundant subsets may be detrimental to detection performances. We present AltraFlowSOM, a semi-supervised extension of FlowSOM that embeds partial expert annotations directly into self-organizing map training via a two-layer SuperSOM architecture, balancing label-guided topology anchoring with unsupervised discovery. By anchoring the map to biologically labelled reference points, AltraFlowSOM circumvents the canonical dependency between batch correction and clustering. Evaluated under Leave-one-out cross validation on two independent IMC cohorts, Lupus Nephritis (n=22 ROIs) and Sjogren syndrome (n=10 ROIs), AltraFlowSOM outperformed all unsupervised and supervised baseline on Adjusted Rand Index, F1 scores (macro and weighted), weighted purity and in the identification of rare populations. The median Treg cell recovery exceeded that of all comparator methods. AltraFlowSOM resolves the scalability-alignment-discovery trilemma, by establishing a semi-supervised SOM as a generalizable method for high dimensional IMC phenotyping.
]]></description>
<dc:creator><![CDATA[ ANILKUMAR REKHA, A., Bettacchioli, E., Le Dantec, C., Hemon, P., Jouve, P. E., Hillion, S. ]]></dc:creator>
<dc:date>2026-09-04</dc:date>
<dc:identifier>doi:10.64898/2026.08.31.748226</dc:identifier>
<dc:title><![CDATA[AltraFlowSOM: A Semi-Supervised Framework for Imaging Mass Cytometry Phenotyping]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-04</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.31.747916v1?rss=1">
<title>
<![CDATA[
QuickSeg: A fast, versatile and accurate algorithm for genomic copy number segmentation using dynamic programming 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.31.747916v1?rss=1
</link>
<description><![CDATA[
Copy number alterations are among the most common genomic aberrations in cancer and their accurate identification relies on robust segmentation of sequencing read-depth signals. Existing segmentation methods typically balance computational efficiency against segmentation accuracy and remain sensitive to technical artifacts present in sequencing data. Here, we present QuickSeg, a fast and versatile methodology that uses an exact dynamic programming algorithm to detect copy number segments using median-based error function. Motivated by the observation that sequencing depth distributions contain a small but pervasive population of outlying observations, this approach provides increased robustness to technical noise while simultaneously reducing the computational complexity of the segmentation problem. Across whole-genome sequencing of cancer cohorts, using breakpoint-supported somatic copy number alterations, we demonstrate improved segmentation precision over two widely used baseline methods, Circular Binary Segmentation (CBS) and Piecewise Constant Fitting (PCF), across a broad range of sensitivity thresholds. QuickSeg also consistently outperformed both methods with respect to runtime and memory usage. Collectively, our results show that robust median-based optimization provides both biological and computational advantages for copy number segmentation, enabling accurate analysis of large sequencing cohorts with minimal computational requirements.
]]></description>
<dc:creator><![CDATA[ Schlotmann, B., Favero, F., Locallo, A., Weischenfeldt, J. L. ]]></dc:creator>
<dc:date>2026-09-04</dc:date>
<dc:identifier>doi:10.64898/2026.08.31.747916</dc:identifier>
<dc:title><![CDATA[QuickSeg: A fast, versatile and accurate algorithm for genomic copy number segmentation using dynamic programming]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-04</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.31.748244v1?rss=1">
<title>
<![CDATA[
Modelling interpretable patient-level representationsfrom structured and simple multimodal data 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.31.748244v1?rss=1
</link>
<description><![CDATA[
Patient cohort profiling increasingly includes structured views for multiple modalities, such as single-cell RNA sequencing, spatial transcriptomics or proteomics, and histology, each providing multiple subobservations per patient, including single cells, spatial spots or patches. To model such data along with simple patient-level views, current multimodal integration methods typically rely on separately precomputed summaries and fail to fully leverage information in structured views. Here we present FACTMx, a variational framework that jointly models structured and simple views to learn interpretable patient-level representations. FACTMx couples latent patient factors with subobservation clustering and per-patient component proportions, enabling direct interpretation and downstream association analyses. The framework supports different structured-view mixture assumptions, including topic- and Gaussian-structured data, while retaining modular encoder-decoder parameterisations. In simulations spanning sparse and dense dependencies and multiple noise regimes, FACTMx improved reconstruction, integration and recovery of structured components relative to previous methods. Applied to non-small cell lung cancer cohorts, FACTMx captured survival-associated latent signals linked to immune microenvironments, gene expression pathways and spatially coherent histological patterns. In a longitudinal coronary syndrome cohort, FACTMx highlighted an outcome-associated axis connected to ejection-fraction change, immune cell states, soluble mediators and cardiac injury markers. These results support joint structured-simple modelling for interpretable multimodal patient stratification.
]]></description>
<dc:creator><![CDATA[ Oksza-Orzechowski, K., Lazecka, M., Koperski, L., Wojtowicz, D., Mozejko, M., Schulz, D., Liechti, R., Marzetta, F., Morfouace, M., Hong, H. S., Tissot, S., Bodenmiller, B., Staub, E., Szczurek, E. ]]></dc:creator>
<dc:date>2026-09-04</dc:date>
<dc:identifier>doi:10.64898/2026.08.31.748244</dc:identifier>
<dc:title><![CDATA[Modelling interpretable patient-level representationsfrom structured and simple multimodal data]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-04</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.31.748380v1?rss=1">
<title>
<![CDATA[
Gene function prediction from bulk coexpression is bounded by cell-type-level signal 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.31.748380v1?rss=1
</link>
<description><![CDATA[
It is widely accepted in genomics that coexpression of RNA transcripts suggests a commonality of function. This intuition is explicitly leveraged in machine learning methods that predict gene function, where it is often combined with other features such as protein interactions and sequence similarity. For example, including coexpression data from human tissue expression boosts performance for predicting Gene Ontology annotations. However, the biological underpinnings of this observation have not been well-investigated. Building on earlier results from our group, in this work we show that gene function is predictable from coexpression substantially because it reflects differences in expression between cell types, and these differences are also intrinsic to the ground truth labels. Using simulations and analyses of real data, we show that variance in the cellular composition of bulk samples impacts function learnability and attribute this to cell type marker gene content in the GO terms. We further show that cell type profiles, where the relationship between gene expression and cell type is made transparent, are effective for predicting gene function while increasing interpretability. These results indicate that function prediction models trained on bulk coexpression are largely limited to cell-type-level resolution rather than fine-grained biochemical function, with direct consequences for how such predictions should be interpreted.
]]></description>
<dc:creator><![CDATA[ Adrian-Hamazaki, A., Pavlidis, P. ]]></dc:creator>
<dc:date>2026-09-04</dc:date>
<dc:identifier>doi:10.64898/2026.08.31.748380</dc:identifier>
<dc:title><![CDATA[Gene function prediction from bulk coexpression is bounded by cell-type-level signal]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-04</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.31.748410v1?rss=1">
<title>
<![CDATA[
On doubting image quality assessment metrics for microscopy virtual staining 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.31.748410v1?rss=1
</link>
<description><![CDATA[
Pairing label-free microscopy with virtual staining could reduce the cost and experimental burden of fluorescence microscopy, but its impact is conditional on generalizable inference. Most virtual staining studies assess performance using image quality assessment (IQA) metrics developed for natural images, yet how well these metrics translate to microscopy remains unknown. Here, we examined the behavior of seven commonly-used full-reference training objectives and metrics, MAE, PSNR, SSIM, foreground PSNR and SSIM, LPIPS, and DISTS, under controlled image degradation and realistic out-of-distribution virtual staining. We applied graded intensity, textural, and morphological transformations to Cell Painting images spanning 18 cell lines, seeding densities, and fluorescence channels. Channel, cell line identity and seeding density explained substantial metric variation after controlling for degradation magnitude. DISTS and foreground metrics showed more favorable balance between degradation sensitivity and biological invariance, although no metric reported performance independent of biological context. Incrementally degrading images and evaluating concomitant metric degradation further revealed that most metrics used only a small fraction of their nominal numerical ranges and frequently plateaued while image degradation visibly continued. We next trained three popular virtual staining model architectures (UNet, WGAN-GP, UNeXt) on five U2-OS seeding densities separately, and computed metrics on model predictions across 17 unseen cell lines. We observed that architecture and training U2-OS seeding density together explain less than 2% of metric variation. Visual inspection suggested comparable scores across cell lines correspond to qualitatively distinct errors, such as differences in cell morphology and marker intensity. These findings show that conventional IQA metrics do not effectively translate to virtual staining applications. Selection or optimization of virtual staining models against real application such as in label-free high content drug screening should instead be approached in an application-oriented fashion.
]]></description>
<dc:creator><![CDATA[ Li, W.-s., Way, G. P. ]]></dc:creator>
<dc:date>2026-09-04</dc:date>
<dc:identifier>doi:10.64898/2026.08.31.748410</dc:identifier>
<dc:title><![CDATA[On doubting image quality assessment metrics for microscopy virtual staining]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-04</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.31.748429v1?rss=1">
<title>
<![CDATA[
Creating DNAm Algorithms Using the Illumina Methylation Screening Array (MSA) 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.31.748429v1?rss=1
</link>
<description><![CDATA[
Most established DNA methylation (DNAm) biomarkers were developed on legacy Illumina EPIC arrays. The Infinium Methylation Screening Array (MSA) offers a lower-cost, higher-throughput alternative with reduced probe content, but EPIC-trained algorithms cannot be assumed to transfer directly. Here we present a reproducibility-based framework for developing and transferring DNAm algorithms on the MSA. Using paired biological replicates profiled on EPICv1 and MSA (1,764 EPICv1-MSA sample pairs, plus within-array MSA replicates on the same and different beadchips), we quantified probe-level agreement using mean absolute error (MAE) and intraclass correlation coefficients (ICC). Of 140,150 CpG sites shared between EPICv1 and MSA, 40,786 (29.1%) met both stability criteria (MAE < 0.05 and ICC(2,k) > 0.6). This stable feature space supported two modelling streams. First, we trained 134 epigenetic biomarker proxies (EBPs) natively on MSA, with and without kernel principal component analysis (kPCA) for sample-level harmonisation. All 134 reached same-beadchip ICC(2,1) >= 0.80 (median 0.97) and 96.3% reached different-beadchip ICC(2,1) >= 0.60 (median 0.81), with a median Spearman correlation of 0.48 against observed values. Among the 72 kPCA-selected models with a comparable stable-probe baseline, 70 (97%) showed higher cross-beadchip ICC (median improvement +0.18). Second, we transferred three established clocks using model-specific strategies: OMICmAge and SystemsAge were retrained to estimate their EPICv1-derived values (held-out test-set rho = 0.944 and 0.912-0.949), whereas DunedinPACE required stable-probe normalisation and robust linear calibration, which raised cross-array ICC(2,1) from 0.784-0.810 to 0.891-0.925 and reduced MAE from 0.085-0.089 to 0.041-0.050 across three sample sets. Reduced probe content does not preclude reproducible DNAm biomarker measurement, and transfer strategy must be matched to model architecture.
]]></description>
<dc:creator><![CDATA[ Seale, K., Hassouneh, S., Giosan, I., Sugden, K., Balague-Dobon, L., Dwaraka, V., Lasky-Su, J. A. B., Mallin, M., Caspi, A., Moffitt, T., Smith, R., Carreras-Gallo, N. ]]></dc:creator>
<dc:date>2026-09-04</dc:date>
<dc:identifier>doi:10.64898/2026.08.31.748429</dc:identifier>
<dc:title><![CDATA[Creating DNAm Algorithms Using the Illumina Methylation Screening Array (MSA)]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-04</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.31.748296v1?rss=1">
<title>
<![CDATA[
A pan-cohort transcriptional landscape of breast cancer maps subtype and microenvironmental programs 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.31.748296v1?rss=1
</link>
<description><![CDATA[
Breast cancer comprises heterogeneous transcriptional states that are incompletely captured by discrete clinical or molecular subtype labels. To visualize this heterogeneity in a unified framework, we integrated bulk RNA-seq data from 2,284 patient samples across 13 studies using 18,089 protein coding genes, a harmonized processing pipeline, batch correction, consensus clustering and PaCMAP dimensionality reduction to construct an interactive breast cancer transcriptional landscape. Consensus clustering identified five major regions, which were annotated using PAM50 scores calculated for each sample: Luminal A, Luminal B, HER2 enriched, and two basal associated clusters. The basal clusters separated into an immune rich region marked by T cell-inflamed, tumor-associated macrophages (TAM), and low-purity signatures, and a cell-cycle driven region enriched for proliferation and DNA replication programs. Overlay of marker genes, pathways, kinases, neuronal like signaling programs, cancer associated fibroblasts (CAF) states, and TAM programs revealed spatially organized subtype biology and microenvironmental heterogeneity. Finally, projection of therapy associated resistance signatures identified landscape regions linked to predicted resistance to HER2-targeted therapy and hormone receptor directed endocrine therapies. By enabling interactive exploration of transcriptional states, marker genes, pathways, and therapeutic response programs, this resource provides a community framework for biomarker discovery in breast cancer.
]]></description>
<dc:creator><![CDATA[ Arora, S., Suresh, R., Holland, N., Glatzer, G., Jensen, M., Konnick, E. Q., Pritchard, C., Li, Y., Parsons, H. A., Hurvitz, S. A., Holland, E. C. ]]></dc:creator>
<dc:date>2026-09-04</dc:date>
<dc:identifier>doi:10.64898/2026.08.31.748296</dc:identifier>
<dc:title><![CDATA[A pan-cohort transcriptional landscape of breast cancer maps subtype and microenvironmental programs]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-04</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.31.748400v1?rss=1">
<title>
<![CDATA[
A Compendium of 49 Experimental SBS Signatures for Decoding Human Cancer Mutational Processes 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.31.748400v1?rss=1
</link>
<description><![CDATA[
Human cancer genomes harbor distinct mutational patterns that reflect past processes of DNA damage and repair. However, the precise attribution of these signatures to specific chemical carcinogens lacks a standardized experimental reference framework. To address this gap, we curated 4,282 genome-wide sequencing datasets from 42 model systems across five species exposed to 146 cancer-risk agents. This platform yielded 49 robust experimental single-base substitution signatures (eSS), with 28 matching 19 established COSMIC signatures and 21 defining novel mutational processes. We reconstructed 24 COSMIC signatures, assigning candidate etiologies to five signatures of unknown origin and revising two contested assignments. Pan-cancer decomposition detected four eSS-like mutational processes enriched in smokers across 4,951 tumors. Lastly, independent single-molecule sequencing of primary human organoids reproduced these profiles with high fidelity, confirming true platform-independent biological reproducibility across complex human models. This eSS repertoire provides a reference that links human mutational processes to mechanistic classes of DNA damage.
]]></description>
<dc:creator><![CDATA[ Zhivagui, M., Au, J. N., Sharma, S., Nguyen, P. T., Al-Azzam, S., Zhang, J., Barnes, M., Alexandrov, L. B. ]]></dc:creator>
<dc:date>2026-09-04</dc:date>
<dc:identifier>doi:10.64898/2026.08.31.748400</dc:identifier>
<dc:title><![CDATA[A Compendium of 49 Experimental SBS Signatures for Decoding Human Cancer Mutational Processes]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-04</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.31.748267v1?rss=1">
<title>
<![CDATA[
Mitotic chromosomes are mechanically connected by chromatin-based interchromosome linkers 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.31.748267v1?rss=1
</link>
<description><![CDATA[
The spatial organization of chromosomes is crucial for gene regulation and genome stability. Metaphase chromosomes appear physically discrete from one another in conventional chromosome preparations. However, interchromosomal connections, or "linkers", have been observed among mitotic chromosomes, and may play important roles in coordinated chromosome movements and genome stability. We use micropipette-based isolation and manipulation to analyze interchromosome linkers between mammalian mitotic chromosomes. Pulling an isolated chromosome reveals that thin linkers connect them; CREST staining indicates their locations to be at centromeres. The linkers display linear elasticity and have a length-doubling force of approximately 300 pN, comparable to the length-doubling force of an entire metaphase chromosome. Enzymatic treatments demonstrate that the linkers are disrupted by DNase, but not by RNase, protease, or microtubule polymerization (spindle fiber) inhibitors, indicating that linker connectivity is based on DNA. Immunofluorescence experiments indicate that histones and topo II are on the linkers, with topo II organized into discrete foci spaced by approximately 0.4 Mbp. After prolonged metaphase arrest using spindle inhibitors isolated genomes do not display interchromosome linkers. Finally, experiments on intact cells show that interchromosome linkers containing CENP-B can be observed between mitotic chromosomes, indicating they are not an artifact of genome isolation.
]]></description>
<dc:creator><![CDATA[ Marko, J. F., Hua, L. L., Akhtar, O., Biggs, R. J., Plancarte, G. Q., Sun, M. ]]></dc:creator>
<dc:date>2026-09-04</dc:date>
<dc:identifier>doi:10.64898/2026.08.31.748267</dc:identifier>
<dc:title><![CDATA[Mitotic chromosomes are mechanically connected by chromatin-based interchromosome linkers]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-04</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.02.748787v1?rss=1">
<title>
<![CDATA[
Experimental and In Silico Analysis of the Structural Dynamics of Dengue NS2B-NS3 Protease 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.02.748787v1?rss=1
</link>
<description><![CDATA[
The Dengue virus (DENV), a major global health concern, causes dengue fever, predominantly affecting tropical and subtropical regions. The NS2B-NS3 protease complex of DENV is critical for viral replication, making it a promising target for antiviral drug development. This study investigates the structural dynamics of the NS2B-NS3 protease under varying pH conditions, integrating experimental and computational approaches. The recombinant NS2B-NS3 protease was expressed in E. coli, purified using affinity chromatography, and analysed for purity via SDS-PAGE. Circular dichroism (CD) spectroscopy revealed pH-dependent secondary structural changes, indicating stability at neutral to slightly basic pH and destabilization under acidic and highly basic conditions. Dynamic light scattering (DLS) analysis demonstrated protein aggregation and structural heterogeneity under extreme pH levels. Complementary in silico techniques, including homology modelling and molecular dynamics (MD) simulations, provided detailed insights into the conformational changes of the protease. The modelled structure, validated and refined through computational tools, revealed structural stability at physiological pH, with notable disruptions at pH extremes. DSSP (Dictionary of Secondary Structure of Proteins) and principal component analyses highlighted significant secondary structural transitions, especially at acidic pH where -helices and {beta}-sheets transformed into random coils Docking and MD simulation was carried out with the small molecule for therapeutic analysis between Dengue NS2bNS3 and small molecule. This study emphasizes the pH-dependent conformational plasticity of the NS2B-NS3 protease, contributing to the understanding of its functional mechanisms and providing a foundation for the rational design of pH-specific inhibitors. These findings underscore the importance of structural biology in advancing therapeutic strategies against dengue fever.
]]></description>
<dc:creator><![CDATA[ Muthuvel, s. k., BALAKRISHNAN, S. S., MANJINI, S., DHAL, K., DAS, S., S, D. ]]></dc:creator>
<dc:date>2026-09-04</dc:date>
<dc:identifier>doi:10.64898/2026.09.02.748787</dc:identifier>
<dc:title><![CDATA[Experimental and In Silico Analysis of the Structural Dynamics of Dengue NS2B-NS3 Protease]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-04</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.28.747928v1?rss=1">
<title>
<![CDATA[
GLORB: Robust Bayesian inference for differential expression underglobal expression shifts 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.28.747928v1?rss=1
</link>
<description><![CDATA[
Estimating differential gene expression is a common task in RNA-seq. Current methods mostly rely on normalization to reduce variance and increase accuracy. These methods are widely used and provide invaluable information about transcriptomic changes between biological conditions. However, widely used normalization methods are known to distort estimates of transcript differential expression when a majority of genes are upregulated or downregulated, or when the total RNA content per cell changes. Despite the presence of global expression shifts in a number of contexts, few methods exist that can provide accurate normalization and estimate linear models under this context without spike-in controls. Here, we present textbf{GLORB} (textbf{G}textbf{L}textbf{O}bal-shift textbf{R}obust textbf{B}ayesian model), a method for estimating generalized linear models under global upregulation. We develop two models that are able to recover differentially expressed genes and linear model coefficients with lower distortion of results. We show that our model's method of accounting for library size variance is consistent with DESeq2's median of ratios, and edgeR's trimmed mean of M-values under conditions when a minority of genes are upregulated and outperforms them under circumstances when most genes are either increased or decreased between groups. Finally, because our method does not rely on calculating geometric means for each gene it is able to work in datasets with much higher sparsity.
]]></description>
<dc:creator><![CDATA[ Callahan, R. L., Coleman, S. D., Ngo, T. T. M. ]]></dc:creator>
<dc:date>2026-09-03</dc:date>
<dc:identifier>doi:10.64898/2026.08.28.747928</dc:identifier>
<dc:title><![CDATA[GLORB: Robust Bayesian inference for differential expression underglobal expression shifts]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-03</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.28.747864v1?rss=1">
<title>
<![CDATA[
ContrasTED: contrastive domain embeddings for scalable remote homology classification 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.28.747864v1?rss=1
</link>
<description><![CDATA[
Protein structure prediction has expanded structural databases to hundreds of millions of domains. Classifying these domains into homologous superfamilies reveals evolutionary and functional relationships that can persist despite low sequence similarity. As the size of structural databases continues to grow, homology classification requires methods that combine scalability with accuracy. Here we present ContrasTED, which uses CATH-supervised center-contrastive learning to project structure-aware embeddings into a domain-level metric space for nearest-centroid superfamily assignment. On a sequence-filtered S20 benchmark (n = 1,028), superfamily assignment accuracy reached 92.9% (1-NN) and 91.4% (nearest centroid), exceeding sequence search, profile HMMs, Foldseek, and a classifier trained on embeddings. The learned latent space separates superfamilies while retaining structural information below 20% sequence identity, with the largest gains among sparsely represented superfamilies. ContrasTED produces 4.67 million new candidate assignments across 3,796 superfamilies in The Encyclopedia of Domains (TED), extending annotation coverage beyond previous structure-based methods.
]]></description>
<dc:creator><![CDATA[ Miller, D. M., Bordin, N., Jeyananthan, J., Waman, V., Heinzinger, M., Orengo, C. ]]></dc:creator>
<dc:date>2026-09-03</dc:date>
<dc:identifier>doi:10.64898/2026.08.28.747864</dc:identifier>
<dc:title><![CDATA[ContrasTED: contrastive domain embeddings for scalable remote homology classification]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-03</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.30.748130v1?rss=1">
<title>
<![CDATA[
HiC-LEGO: Biologically Guided High-Resolution 3D Genome Reconstruction Preserves Chromatin Organization at Kilobase Resolution 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.30.748130v1?rss=1
</link>
<description><![CDATA[
Three-dimensional (3D) chromosome reconstruction from Hi-C contact maps remains challenging because genome organization is hierarchical and fine-resolution models must reconcile local structure with chromosome-scale constraints. Here we present HiC-LEGO, a domain-aware hierarchical framework integrating ensemble chromatin domains with graph-based structural learning and progressive chromosome assembly. By combining ensemble domain selection with hierarchical reconstruction, HiC-LEGO reduces dependence on individual domain definitions while maintaining local organization during chromosome-scale assembly. Across five human cell lines at 5-kb resolution, HiC-LEGO achieves higher reconstruction concordance than evaluated state-of-the-art methods while better preserving domain organization. At 1-kb resolution, HiC-LEGO reconstructs complete GM12878 chromosome 8 and recovers close spatial proximity between an epigenomically supported distal MYC enhancer and its promoter. Reconstructions from 5-kb Micro-C data show that 249 experimentally defined RCMC microcompartment interactions at the Ppm1g locus occupy compact 3D configurations. In the breast cancer dataset, HiC-LEGO reconstructs structures that maintain stable TAD organization across healthy breast, primary tumors, and liver metastases, while revealing greater inter-patient structural heterogeneity in malignant pleural effusion samples. Pore-C validation shows that experimentally observed multi-way contacts spanning 1-5 Mb are enriched in compact reconstructed configurations across all 23 chromosomes. Thus, HiC-LEGO preserves regulatory interactions, disease-associated chromatin organization and higher-order spatial relationships beyond pairwise contact-map concordance.
]]></description>
<dc:creator><![CDATA[ Pandeya, A., Chowdhury, M. F. K., Oluwadare, O. ]]></dc:creator>
<dc:date>2026-09-03</dc:date>
<dc:identifier>doi:10.64898/2026.08.30.748130</dc:identifier>
<dc:title><![CDATA[HiC-LEGO: Biologically Guided High-Resolution 3D Genome Reconstruction Preserves Chromatin Organization at Kilobase Resolution]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-03</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.30.748136v1?rss=1">
<title>
<![CDATA[
Retrieval of binding sites across the AlphaFold human proteome using protein language model representations 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.30.748136v1?rss=1
</link>
<description><![CDATA[
Protein language models (PLMs) provide powerful representations of protein sequence, but their utility for proteome-scale binding-site retrieval remains unclear. Here, we present PocketScope, a training-free framework that represents cavity-lining residues using frozen ESM-C 600M embeddings and retrieves related binding sites through exhaustive late-interaction MaxSim, without pooling or approximate nearest-neighbor search. PocketScope identified 153,805 cavities across 37,682 proteins in the AlphaFold human proteome and recovered documented drug off-targets across a curated set of pharmacological pairs. On the ProSPECCTs benchmark, PocketScope ranks 1st of 23 methods by mean rank across the ten collections. PocketScope provides a practical framework for proteome-scale off-target prediction. PocketScope is open source and also freely available as a web server at https://www.bhargavaresearch.org/pocketscope.
]]></description>
<dc:creator><![CDATA[ Mohan, K., Bhargava, Y. ]]></dc:creator>
<dc:date>2026-09-03</dc:date>
<dc:identifier>doi:10.64898/2026.08.30.748136</dc:identifier>
<dc:title><![CDATA[Retrieval of binding sites across the AlphaFold human proteome using protein language model representations]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-03</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.30.743167v1?rss=1">
<title>
<![CDATA[
Bio-Babel: autonomous cross-language reconstruction of agent-ready computational biology ecosystems 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.30.743167v1?rss=1
</link>
<description><![CDATA[
Scientific software is often locked in its source language, callable from others but not natively extensible or built upon. We present Bio-Babel (biobabel.stanford.edu), an autonomous framework that rebuilds software natively in target ecosystems and makes it agent-callable, shipping each package with an agent-readable contract of its usage. Reconstructing a hierarchical, interdependent R-to-Python single-cell stack, Bio-Babel reproduced the originals faithfully and revealed pancreatic differentiation dynamic defects under graded SWI/SNF loss.
]]></description>
<dc:creator><![CDATA[ Liu, N., Chen, X., Cui, M., Liao, X., Song, X., Pantham, S., Xu, W., Qiu, X. ]]></dc:creator>
<dc:date>2026-09-03</dc:date>
<dc:identifier>doi:10.64898/2026.08.30.743167</dc:identifier>
<dc:title><![CDATA[Bio-Babel: autonomous cross-language reconstruction of agent-ready computational biology ecosystems]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-03</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.28.747909v1?rss=1">
<title>
<![CDATA[
The human metabolite - protein interactome reveals a global layer of cellular coordination 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.28.747909v1?rss=1
</link>
<description><![CDATA[
Metabolites are substrates, products, cofactors, and regulators, but protein-protein interaction networks do not represent their potential to organize proteins across conventional pathway boundaries. Using the LIGMAP virtual-screening algorithm, we mapped 308 human metabolite codes to pockets in monomers, dimer interfaces, and non-interface sites in dimers and represented attractive or repulsive COLIG states involving pairs of metabolites in the same pocket. On a fixed cohort of 3,938 proteins, the mean coverage of 104 strict non-enzyme pathways was 44.6% for LIGMAP, 70.1% for STRING, and 82.5% for STRING+LIGMAP; the union placed 86.3% of eligible proteins in the largest connected component and 95.1% in the two largest components. STRING+LIGMAP protein coverage was 88.0% for 68 enzyme-only pathways and 88.9% for 1,283 mixed pathways. In pathway-held-out, degree-matched prediction, adding LIGMAP to degree plus STRING increased the mean area under the precision-recall curve from 0.651 to 0.660 (paired P = 0.024); adding BioLiP2 increased it to 0.663 (paired P = 0.005). Ancient-only and non-ancient-only subnetworks were each globally connected; ancient features were denser, whereas non-ancient features covered more proteins and pathways. At the full 5,426-protein scale, retaining only features assigned to 2-100 proteins recovered 698 of 1,691 strict non-enzyme reference edges (41.3%) and exceeded both protein-label and exact bipartite degree-preserving nulls. Uncapped recovery approached saturation and lost identity-selective enrichment. Experimentally established metabolite-dependent complexes validate the local mechanism independently of LIGMAP; LIGMAP fully recovered two of seven stringent direct mechanisms and all three broader serial axes examined. Our findings reveal a global metabolite-mediated architecture with the capacity to coordinate proteins across otherwise distinct cellular systems.
]]></description>
<dc:creator><![CDATA[ Skolnick, J., Srinivasan, B. ]]></dc:creator>
<dc:date>2026-09-03</dc:date>
<dc:identifier>doi:10.64898/2026.08.28.747909</dc:identifier>
<dc:title><![CDATA[The human metabolite - protein interactome reveals a global layer of cellular coordination]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-03</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.31.748056v1?rss=1">
<title>
<![CDATA[
Comprehensive Evaluation of Protein Language Model Embeddings for Drug-Target Affinity Prediction 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.31.748056v1?rss=1
</link>
<description><![CDATA[
Accurate identification of drug-target interactions is consequential for novel drug discovery and development. Deep learning methods for drug-target affinity (DTA) prediction have shown great promise in accelerating drug discovery and reducing development costs. Although graph neural networks have improved drug representation learning for DTA prediction tasks, many models still struggle to effectively and efficiently capture protein information, limiting overall prediction accuracy. In this work, we systematically evaluate the impact of pre-trained protein language models (PLMs) on the downstream task of predicting binding affinity between drugs and target proteins. We design multiple experiments across four different molecular representation backbones and assess the effect of incorporating PLM embeddings, comparing their performance to classical 1D convolution methods. We evaluate four families of PLMs which we integrate into PLM-GraphDTA, each built on distinct architectures and optimized for different tasks, including structure prediction, function prediction, and sequence unmasking. Additionally, we evaluate DeepGraphDTA, an architectural modification of the baseline convolution method designed to improve protein representation learning. The models are evaluated on two benchmark datasets, Davis and KIBA, using concordance index (CI) and mean squared error (MSE) as performance metrics. We further evaluate the generalization power of each model using cold-start train and test splits, and analyze the per-protein contribution to total CI. The results indicate simple architectural modifications to traditional convolution methods may be sufficient to bridge the gap to large pre-trained PLMs.
]]></description>
<dc:creator><![CDATA[ Marijan, M., Tanasijevic, I. ]]></dc:creator>
<dc:date>2026-09-03</dc:date>
<dc:identifier>doi:10.64898/2026.08.31.748056</dc:identifier>
<dc:title><![CDATA[Comprehensive Evaluation of Protein Language Model Embeddings for Drug-Target Affinity Prediction]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-03</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.30.748103v1?rss=1">
<title>
<![CDATA[
Hi-cGAN: Prediction of Hi-C interaction matrices with conditional generative adversarial networks 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.30.748103v1?rss=1
</link>
<description><![CDATA[
Background: The three-dimensional organization of the genome is a fundamental aspect of its function and regulation. High-throughput chromosome conformation capture techniques, such as Hi-C, have revolutionized our understanding of spatial genome organization. However, 3C-based methods are resource-intensive and technically demanding. This has driven the development of computational approaches for predicting Hi-C interaction matrices. Hi-cGAN, a novel approach based on conditional generative adversarial networks, offers a computational alternative to extensive wet-lab work by predicting Hi-C interaction matrices. This computational approach contributes to a broader exploration and understanding of genome architecture. Findings: The network pairs a convolutional generator with a convolutional discriminator, evaluated across bin sizes, inputs and cell types. It predicts a whole genome as a cool file at bin sizes from 2 to 25 kb, where Akita, C.Origami and Epiphany emit fixed windows of 1 Mb, 2 Mb and 990 kb. With the input chosen on a validation chromosome, agreement approaches Epiphany's and stays below the sequence-based C.Origami and Akita: over Akita's 411 held-out windows the mean correlation is 0.238 against 0.506. Boundary and loop calls agree less closely, placing the maps at the domain scale. Conclusions: Chromatin factor occupancy determines a substantial part of contact structure, and two tracks capture most of it. The most informative track depends on the resolution: CTCF and the cohesin subunits at 5 to 10 kb, active histone marks at 25 kb. Transfer to an unseen cell type costs about 0.12 SCC, and which method leads depends on the measure.
]]></description>
<dc:creator><![CDATA[ Krauth, R., Kumar, A., Wolff, J. ]]></dc:creator>
<dc:date>2026-09-03</dc:date>
<dc:identifier>doi:10.64898/2026.08.30.748103</dc:identifier>
<dc:title><![CDATA[Hi-cGAN: Prediction of Hi-C interaction matrices with conditional generative adversarial networks]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-03</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.31.748297v1?rss=1">
<title>
<![CDATA[
ClassifyITS: An R Package for assigning taxonomy to fungal ITS sequences using taxon-specific cutoff values 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.31.748297v1?rss=1
</link>
<description><![CDATA[
1. Fungi are key drivers of decomposition and nutrient cycling across the globe, yet accurate classification of environmental fungal internal transcribed spacer (ITS) sequences remains challenging. These persistent challenges reflect the variable evolutionary properties of ITS, limited representation of fungal diversity in reference databases, and the application of classifiers originally developed for more conserved prokaryotic markers. 2. Here, we present ClassifyITS, an R package that performs alignment-based taxonomic classification of full length fungal ITS sequences or individual ITS subregions (ITS1 or ITS2) using taxon-specific sequence identity thresholds. In addition to taxonomic assignments, ClassifyITS generates summary statistics and diagnostic visualizations to support interpretation and quality control. 3. Using a deep subsurface fungal ITS dataset containing many poorly characterized taxa, ClassifyITS outperformed the common classifiers SINTAX and DADA2, with higher agreement to expert curated assignments and lower rates of over classifying and under classifying sequences to taxonomic ranks. Across all classifiers and approaches, taxonomic accuracy increased strongly with sequence similarity to the reference database, emphasizing the importance of continued expansion and curation of fungal sequence databases. 4. By providing an accessible and reproducible R based workflow that improves taxonomic classification, ClassifyITS supports more accurate biodiversity monitoring and enhances downstream functional interpretation of fungal communities.
]]></description>
<dc:creator><![CDATA[ Moon, Q., Tedersoo, L., Griebler, C., Karwautz, C., Cukusic, A., James, T. ]]></dc:creator>
<dc:date>2026-09-03</dc:date>
<dc:identifier>doi:10.64898/2026.08.31.748297</dc:identifier>
<dc:title><![CDATA[ClassifyITS: An R Package for assigning taxonomy to fungal ITS sequences using taxon-specific cutoff values]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-03</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.31.747784v1?rss=1">
<title>
<![CDATA[
scRep: A Latent-Space Self-Distilled Foundation Model for Single-Cell Representation Learning 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.31.747784v1?rss=1
</link>
<description><![CDATA[
Single-cell foundation models have shown strong potential for learning transferable representations from large-scale transcriptomic data. However, many existing approaches rely on reconstructing masked gene expression values, creating a potential mismatch between observation-space reconstruction and the goal of learning stable biological representations. This challenge is particularly relevant to single-cell RNA sequencing, where sparsity, incomplete gene detection, and technical variation can obscure the underlying biological state. Here, we introduce scRep, a compact latent-space self-distillation framework for single-cell representation learning. Rather than reconstructing raw expression values, scRep aligns differently perturbed views of the same cell through a momentum-updated teacher--student architecture, with self-distillation objectives at both the cell and gene levels. This representation-centered formulation encourages the model to capture biological information that remains stable across incomplete and perturbed transcriptomic observations. Using frozen representations without task-specific fine-tuning, scRep pretrained on approximately 2.8 million cells achieves the strongest overall performance across the evaluated frozen-representation benchmarks, demonstrating strong sample efficiency. A larger-scale scRep model pretrained on 30.72 million cells further demonstrates that the framework remains effective when scaled to a substantially larger and more diverse corpus. Beyond cell identity, scRep prioritizes established marker genes, recovers transcription factor--associated gene programs with cell-type-specific activity, and preserves continuous developmental structure that supports graph-based pseudotime inference. We further show that pretraining performance is closely associated with biological diversity: reducing redundant cells while improving cell-type coverage can match or exceed the performance of larger, less balanced training corpora. Together, these results establish latent-space self-distillation as an effective alternative to expression reconstruction for single-cell foundation modeling and suggest that efficient scaling depends not only on the number of cells, but also on the learning objective and the biological diversity of the pretraining corpus.
]]></description>
<dc:creator><![CDATA[ Wang, S., Hu, Z., Bie, Y., Yin, Q., Chen, H., Li, Q. ]]></dc:creator>
<dc:date>2026-09-03</dc:date>
<dc:identifier>doi:10.64898/2026.08.31.747784</dc:identifier>
<dc:title><![CDATA[scRep: A Latent-Space Self-Distilled Foundation Model for Single-Cell Representation Learning]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-03</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.30.748028v1?rss=1">
<title>
<![CDATA[
An annotation-overlap-flagged rare-disease gene-prioritisation benchmark and PMC index recipe 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.30.748028v1?rss=1
</link>
<description><![CDATA[
Benchmarks for rare-disease gene prioritisation are assembled from published clinical cases. Those cases often come from the same publications used to build knowledge-base ("curated") tools, so a curated tool can be scored on its own source literature. This resource makes that circularity measurable. We release a stratified benchmark of 1,047 rare-disease cases from the GA4GH Phenopacket Store v0.1.26. Each case pairs a Human Phenotype Ontology profile with a 50-gene candidate list (one causal gene, 49 distractors) and the causal-gene label, sampled across four operational MONDO-derived disease strata and issued in two case-paired variants: random distractors, and phenotype-similar distractors selected by HPO Resnik similarity. Two case-level metadata layers support fairer evaluation: a per-case flag recording whether a case's source publication is cited in the HPO disease-annotation file, defining an overlap-absent subset (n = 282), and publication-recency strata. We also specify a deterministic, version-pinned recipe for a hybrid dense-plus-sparse retrieval index over approximately 2.25 million PMC Open Access articles (52,777,395 chunks). The resource reports no tool comparisons.
]]></description>
<dc:creator><![CDATA[ Angulo, J., Yeste, V., Espinos-Morato, H. ]]></dc:creator>
<dc:date>2026-09-03</dc:date>
<dc:identifier>doi:10.64898/2026.08.30.748028</dc:identifier>
<dc:title><![CDATA[An annotation-overlap-flagged rare-disease gene-prioritisation benchmark and PMC index recipe]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-03</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.30.748098v1?rss=1">
<title>
<![CDATA[
Ageas enables time-agnostic cell fate inference from single-cell and spatial multi-omics data 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.30.748098v1?rss=1
</link>
<description><![CDATA[
Understanding cell fate decisions is fundamental to developmental biology and disease research. However, experimental lineage tracing requires genetic manipulation, which is impractical in many systems, particularly in humans. Computational approaches often rely on time-resolved measurements, which single-cell and spatial omics studies rarely provide. Here, we present Ageas, a time-agnostic transfer learning framework for cell fate inference from single-cell and spatial multi-omics data. Ageas overcomes these limitations by learning fate memory from terminal cell populations and transferring this information to progenitor or intermediate cells, enabling fate bias inference from static molecular snapshots. To enable robust generalization across molecular modalities, Ageas employs a data-adaptive ensemble strategy with automated model selection. In benchmark datasets with lineage-traced single-cell transcriptomic and epigenomic profiles, as well as spatial transcriptomics, Ageas achieves strong performance compared to existing methods. Applying Ageas to a 3D human embryo reveals a spatially organized anterior-posterior gradient of epiblast fate priming, with anterior epiblast cells biased toward ectodermal fates and posterior cells toward primitive streak-derived lineages, accompanied by regionally graded fate-associated regulatory programs. Together, these results establish Ageas as a general framework for decoding cell fate decisions from static molecular snapshots.
]]></description>
<dc:creator><![CDATA[ Jiang, J., Kong, A., Yu, G. ]]></dc:creator>
<dc:date>2026-09-03</dc:date>
<dc:identifier>doi:10.64898/2026.08.30.748098</dc:identifier>
<dc:title><![CDATA[Ageas enables time-agnostic cell fate inference from single-cell and spatial multi-omics data]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-03</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.30.748061v1?rss=1">
<title>
<![CDATA[
Kintsugi maps nucleus-poor RNA compartments in subcellular spatial transcriptomics 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.30.748061v1?rss=1
</link>
<description><![CDATA[
Subcellular spatial transcriptomics resolves RNA where tissue structures extend beyond nuclei, but current aggregation strategies either impose fixed areal units or use segmented cells as anchors. This leaves process-rich and extracellular compartments difficult to measure directly. Here we report Kintsugi, a deterministic tessellation method that partitions sparse count matrices by captured-UMI density and assigns every in-tissue bin. It separates Pearson-residual gene composition from captured transcript density, preserving complete tissue representation without histology or nuclear segmentation. In mouse brain Visium HD, Kintsugi recovered nucleus-poor neuropil compartments enriched for glial and dendritic transcripts, and Xenium molecule coordinates supported dendritic mRNA localisation away from nuclei. The same representation identified nucleus-poor fibrotic scar in idiopathic pulmonary fibrosis and matrix structure in fetal cartilage. CODEX proteomics further showed that captured density was not reducible to nuclear packing alone. These results show that nucleus-poor tissue spaces contain structured, interpretable RNA biology that becomes accessible through complete tissue tessellation.
]]></description>
<dc:creator><![CDATA[ Yang, C., Zhang, X., Chen, J. ]]></dc:creator>
<dc:date>2026-09-03</dc:date>
<dc:identifier>doi:10.64898/2026.08.30.748061</dc:identifier>
<dc:title><![CDATA[Kintsugi maps nucleus-poor RNA compartments in subcellular spatial transcriptomics]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-03</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.30.748081v1?rss=1">
<title>
<![CDATA[
Multiparental RNA-seq driven eQTL screening identifies loci underlying host plant fitness in a generalist herbivore 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.30.748081v1?rss=1
</link>
<description><![CDATA[
The two-spotted spider mite (Tetranychus urticae) is an extremely polyphagous pest, yet the genetic basis of this adaptive potential remains to be fully elucidated. Since expression quantitative trait loci (eQTLs) provide the genetic basis of numerous phenotypes, we aimed to identify trans-eQTL hotspots underlying T. urticae fitness upon transfer from a common (bean) to a challenging (tomato) host plant. Nonetheless, the identification of trans-eQTLs is complex and often constrained by methodological challenges and high costs. Therefore, we employed a multiparental mapping strategy driven by RNA-seq, enabling us to leverage extensive genetic variation in a cost-efficient manner. A randomly mating population was generated on bean from a small number of genetically diverse, often heterozygous parents, and subsequently transferred to tomato prior to RNA-seq. Upon whole genome sequencing of the parents, RNA-seq of the mapping population individuals was sufficient to reconstruct their genomes as a combination of parental haploblocks. Subsequent eQTL mapping identified 23 distinct trans-eQTL hotspot regions associated with the expression of numerous target genes. The most prominent hotspot on chromosome 3 was associated with approximately 900 genes and showed enrichment for functions related to detoxification and digestion. Furthermore, 4 of these trans-eQTL hotspot genotypes explained significant variation in mite fitness on the challenging host plant. This contrasted a traditional QTL mapping approach, where these genotypes could not be detected due to multiple testing correction. Our study hence offers a powerful strategy for trans-eQTL hotspot screening and the discovery of trait-associated loci in complex, multiparental genetic backgrounds.
]]></description>
<dc:creator><![CDATA[ De Graeve, F., De Beer, B., De Rouck, S., Naessens, S., Herpoel, M., Fostier, J., Clark, R., Van Leeuwen, T., De Meyer, T. ]]></dc:creator>
<dc:date>2026-09-03</dc:date>
<dc:identifier>doi:10.64898/2026.08.30.748081</dc:identifier>
<dc:title><![CDATA[Multiparental RNA-seq driven eQTL screening identifies loci underlying host plant fitness in a generalist herbivore]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-09-03</prism:publicationDate>
<prism:section></prism:section>
</item>
</rdf:RDF>
