<?xml version="1.0" encoding="UTF-8" ?>
<rdf:RDF xmlns:admin="http://webns.net/mvcb/" xmlns="http://purl.org/rss/1.0/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:prism="http://purl.org/rss/1.0/modules/prism/" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:syn="http://purl.org/rss/1.0/modules/syndication/">
<channel rdf:about="https://biorxiv.org">
<admin:errorReportsTo rdf:resource="mailto:biorxiv@cshlpress.edu"/>
<title>bioRxiv Subject Collection: Bioinformatics</title>
<link>https://biorxiv.org</link>
<description>
This feed contains articles for bioRxiv Subject Collection "Bioinformatics"
</description>

<items>
<rdf:Seq>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.27.754762v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.27.753847v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.27.754758v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.27.754054v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.27.754741v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.26.754739v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.27.754743v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.30.755594v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.30.755779v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.28.754049v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.29.755321v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.26.754620v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.26.754622v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.26.754587v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.26.754480v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.26.754596v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.26.754618v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.25.754557v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.25.754454v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.25.754484v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.25.754510v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.25.754276v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.27.754822v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.25.754486v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.25.754455v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.30.755756v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.25.754409v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.25.754345v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.25.754162v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.09.25.754380v1?rss=1"/>
</rdf:Seq>
</items>
<prism:eIssn/>
<prism:publicationName>bioRxiv</prism:publicationName>
<prism:issn/>

<image rdf:resource=""/>
</channel>
<image rdf:about="">
<title>bioRxiv</title>
<url>https://www.biorxiv.org/sites/default/files/bioRxiv_article.jpg</url>
<link>https://www.biorxiv.org</link>
</image>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.27.754762v1?rss=1">
<title>
<![CDATA[
Why rescue-based transcript ranking can mislead: normalization displacement and shared-control coupling in perturbation transcriptomics 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.27.754762v1?rss=1
</link>
<description><![CDATA[
When a disease-associated transcript moves toward control after a phenotypically effective rescue, the pattern can be read as mediator evidence. We show why this inference can fail even when the computation is reproducible. Across normalization schemes in PLP1-mutant Jimpy brain (GSE277705), the original negative-interaction criterion classified 47.6-76.0% of all expressed genes and 77.6-86.6% of disease-eligible genes. Under total CPM, the excess was driven largely by a +1.016-SD median WT treatment response; median-of-ratios (MOR) reduced it to +0.154 SD while leaving the Jimpy median near zero. Under MOR, mirror enrichment among disease-increased (77.6%) and disease-decreased (74.0%) genes showed selection dependence but could not distinguish mechanical coupling from genuine state restoration. Slope decomposition gave a 48.7% mechanical-to-observed ratio (-0.125 of -0.256); deleting one of four untreated-Jimpy samples moved it from 37.8% to 90.1%. Candidate-local point estimates were 60.8% near TRIB3 and 38.6% near DDIT4; the TRIB3-minus-DDIT4 difference stayed positive in all four deletions (8.6--27.7 points), but its magnitude remained unstable. Oligodendrocyte-specific Perk deletion robustly suppressed TRIB3 (R^{JP}=-0.817, 95% CI [-1.048,-0.587]) and DDIT4 (-0.292, [-0.419,-0.169]), whereas 2BAct resolved neither transcript-specific rescue. Rescue transcriptomics can therefore establish broad state restoration without identifying which downstream transcript mediates it.
]]></description>
<dc:creator><![CDATA[ Jamil, H. M., Gow, A. ]]></dc:creator>
<dc:date>2026-10-02</dc:date>
<dc:identifier>doi:10.64898/2026.09.27.754762</dc:identifier>
<dc:title><![CDATA[Why rescue-based transcript ranking can mislead: normalization displacement and shared-control coupling in perturbation transcriptomics]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-02</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.27.753847v1?rss=1">
<title>
<![CDATA[
Downstream mRNA Secondary Structure, Not Codon Elongation Supply, Coordinates Co-Translational Protein Folding Across the Human Ribosomal Exit Tunnel 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.27.753847v1?rss=1
</link>
<description><![CDATA[
How ribosomes pace translation to assist nascent protein folding remains an open question in molecular biology. While synonymous codon selection is widely hypothesized to regulate elongation rates to facilitate domain organization, distinguishing genuine translational kinetics from baseline amino acid preferences has proven technically difficult. Here, we analyze a non-redundant cohort of 1,270 high-resolution human crystal structures (410,151 residues) mapped to their native mRNA transcripts. When we mathematically isolate synonymous codon choices from amino acid identity using orthogonal linear projection, standard codon-supply metrics i.e. the tRNA Adaptation Index (tAI) and the Codon Adaptation Index (CAI), show negligible independent spatial coupling with downstream protein structure. Their uncorrected correlations predominantly reflect local amino acid chemistry rather than physical translation pacing. In contrast, downstream mRNA secondary structure stability (minimum free energy, MFE) displays a subtle but consistent correlation that survives amino acid control. Across 100,000 whole-proteome permutations per pair, this MFE signal centers at an offset of +15 to +16 codons across multiple independent physical properties, including residue packing density (r = -0.0699, Z = -23.64, p < 10-5), crystallographic rigidity (B-factor, r = +0.0614, Z = +13.74, p < 10-5), and solvent burial (SASA, r = +0.0492, Z = +18.98, p < 10-5). This +15 codon offset corresponds directly to the physical dimensions of the eukaryotic 80S ribosome: the path from the peptidyl transferase center to the internal uL4/uL22 constriction neck (~10 amino acids) plus the downstream mRNA helicase entry channel (~5 codons). While the overall effect size is modest, accounting for approximately 0.49% of local packing variance, its spatial specificity and consistency across independent structural metrics suggest that downstream mRNA stability acts as a localized mechanical brake during early chain compaction.
]]></description>
<dc:creator><![CDATA[ Saad Ul Hassan, M., Jabeen, I., Kiani, Y. S. ]]></dc:creator>
<dc:date>2026-10-02</dc:date>
<dc:identifier>doi:10.64898/2026.09.27.753847</dc:identifier>
<dc:title><![CDATA[Downstream mRNA Secondary Structure, Not Codon Elongation Supply, Coordinates Co-Translational Protein Folding Across the Human Ribosomal Exit Tunnel]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-02</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.27.754758v1?rss=1">
<title>
<![CDATA[
Auditing Protein-Protein Interaction Signals with Sparse Autoencoder Fingerprints 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.27.754758v1?rss=1
</link>
<description><![CDATA[
Protein language models have become a dominant foundation for sequence-based protein-protein interaction (PPI) prediction, but their generalization remains limited under stringent evaluation, and benchmark accuracy alone cannot reveal whether a PPI predictor learns partner-specific biological signals or exploits contextual shortcuts. Here we introduce AuditPPI, an interpretable framework that transforms sparse-autoencoder (SAE) features from a frozen protein language model into order-invariant pair fingerprints, recasting PPI prediction as an auditable tabular-learning problem. AuditPPI achieves competitive predictive performance across diverse PPI benchmarks while enabling multiscale auditing of the information supporting its predictions. Protein-level audits show that partner-independent participation signals remain substantial in conventional benchmarks, including protein-disjoint benchmark, but are largely insufficient for pair-level discrimination in topology- and degree-controlled benchmark. Pair-level analyses further show that removing protein identity overlap does not eliminate contextual structure: protein-disjoint benchmarking retains strong subcellular co-localization signals and benchmark-dependent feature matching captured by SAE co-activation and absolute-difference. Structural analyses, however, provide little evidence that these predictive signals are specifically grounded in PPI interfaces: feature enrichment at interfaces for protein-disjoint benchmark was no longer detectable after controlling for surface exposure, and contact-specific enrichment was observed for only a small fraction of testable SAE feature pairs. AuditPPI therefore separates predictive success from mechanistic fidelity and provides a practical framework for evaluating whether sequence-based PPI predictors capture partner-specific biological signals instead of contextual shortcuts.
]]></description>
<dc:creator><![CDATA[ Zhu, W., Wang, S., Liu, X., Xue, Y., Shen, H.-B., Chen, B., Pan, X. ]]></dc:creator>
<dc:date>2026-10-02</dc:date>
<dc:identifier>doi:10.64898/2026.09.27.754758</dc:identifier>
<dc:title><![CDATA[Auditing Protein-Protein Interaction Signals with Sparse Autoencoder Fingerprints]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-02</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.27.754054v1?rss=1">
<title>
<![CDATA[
GGE: General-purpose deep meta-learning for classification of human transcriptomes with limited data 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.27.754054v1?rss=1
</link>
<description><![CDATA[
Transcriptomic classification is often hindered by the small number of samples relative to the high dimensionality of gene expression data. We introduce General Gene Expression (GGE), a deep meta-learning framework designed to support robust classification in this limited-sample setting. By training across 5,220 distinct biomedical prediction objectives drawn from 1,779 different human datasets, GGE learns a model initialization that captures biological patterns shared across heterogeneous classification tasks. This learned initialization has two key advantages. First, it enables improved performance to new datasets using only a small number of labeled samples. Second, because it is learned across diverse prediction objectives, it can be applied to a broad range of biomedical problems. We show that GGE outperforms established classifiers in data-limited settings across a wide range of applications, including datasets generated using different RNA-seq platforms and preprocessing pipelines. In addition, attention-based analysis identifies recurrent genes that contribute to performance across multiple biological objectives, providing insight into shared determinants of human biological states. Together, these results establish GGE is a general-purpose framework for human transcriptome-based classification in biomedical settings where labeled data are scarce.
]]></description>
<dc:creator><![CDATA[ Yankovitz, G., Gat-Viks, I. ]]></dc:creator>
<dc:date>2026-10-02</dc:date>
<dc:identifier>doi:10.64898/2026.09.27.754054</dc:identifier>
<dc:title><![CDATA[GGE: General-purpose deep meta-learning for classification of human transcriptomes with limited data]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-02</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.27.754741v1?rss=1">
<title>
<![CDATA[
GroundAnnot: a closed-vocabulary contract for grounding LLM gene-set annotation in live enrichment backends 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.27.754741v1?rss=1
</link>
<description><![CDATA[
Motivation: LLM agents increasingly draft functional interpretations of gene lists, but can cite Gene Ontology (GO) terms that no current enrichment backend returned for that list, and can pair real GO accessions with fabricated labels. Results: We present GroundAnnot, a client for PANTHER, Enrichr, and g:Profiler that returns a closed vocabulary of GO term IDs and their backend labels, and enforces two contracts: Contract A (no ID or label outside the backend payload) and Contract B (no enrichment claim outside the backend's significant set). Across three open-weight models (Qwen2.5-7B local; Qwen3.8-27B and GPT-OSS-120B via Groq) on six curated disease gene lists, mean valid-enriched rates were 15.3%, 20.1%, and 35.1%. All unsupported IDs were real GO terms classified by QuickGO as wrong-biology, obsolete, or wrong-branch (zero fabricated accessions). A distinct failure mode emerged: for the Alzheimer's gene list, the local 7B model paired every one of its ten stable picks with a label that does not match the current GO term for that accession (10/10 mismatches), producing a coherent synaptic-signalling narrative that the IDs do not support. Mean overlap with the backends' FDR top-10 union was 0.33-0.67/10. On 50 MSigDB Hallmark gene sets, the strongest model named the defining pathway in 21/44 (47.7%) directly named cases. When an enriched shortlist was placed in the prompt, all three models complied fully (60/60 for each model; 180/180 overall), showing that a bounded vocabulary removes these output classes when the downstream agent is required to use it.
]]></description>
<dc:creator><![CDATA[ Malima, M. B. ]]></dc:creator>
<dc:date>2026-10-02</dc:date>
<dc:identifier>doi:10.64898/2026.09.27.754741</dc:identifier>
<dc:title><![CDATA[GroundAnnot: a closed-vocabulary contract for grounding LLM gene-set annotation in live enrichment backends]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-02</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.26.754739v1?rss=1">
<title>
<![CDATA[
EnsPlex: Integrative Protein Complex Prediction through Multi-source Complementary Structural Sampling and Topology-Aware Candidate Selection 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.26.754739v1?rss=1
</link>
<description><![CDATA[
Protein complex structure prediction depends on both the breadth of candidate coverage and the ability to select accurate models within a limited output budget. Most existing approaches emphasize only one stage. End-to-end models and molecular docking workflows generate candidate structures, whereas quality-assessment models primarily rerank a predefined candidate pool. A complete strategy must account for complementary sampling across generators, differences among scoring scales and the allocation of final candidate quotas. Here, EnsPlex is presented as a multi-source framework that couples structural sampling with candidate selection for protein-protein interaction complex prediction. EnsPlex combines conformation-expanded docking, AlphaFold-Multimer, AlphaFold3 and Boltz-1 to expand the sampled conformational space. Conformation-expanded docking comprises monomer conformational expansion followed by flexible HADDOCK docking. FACET is trained to predict candidate quality using DockQ and its component metrics as supervision. Across 102 antigen-antibody systems, EnsPlex achieved up to a 27.8% relative improvement in target success rate over AlphaFold3 with its built-in ranking under matched output budgets. FACET also improved within-source ranking in the internal candidate pools and in homology-filtered external data. These findings support the utility of EnsPlex for structure prediction in the evaluated antigen-antibody systems.
]]></description>
<dc:creator><![CDATA[ Hu, B., Lu, Z., Zaman, K., Sun, Z., He, X. ]]></dc:creator>
<dc:date>2026-10-02</dc:date>
<dc:identifier>doi:10.64898/2026.09.26.754739</dc:identifier>
<dc:title><![CDATA[EnsPlex: Integrative Protein Complex Prediction through Multi-source Complementary Structural Sampling and Topology-Aware Candidate Selection]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-02</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.27.754743v1?rss=1">
<title>
<![CDATA[
AbRefine: Framework-Conditioned Refinement for Structure-Guided Antibody CDR Design 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.27.754743v1?rss=1
</link>
<description><![CDATA[
Structure-guided antibody CDR design constrains residue selection using backbone geometry, but structural evidence alone may leave multiple amino acids plausible. We investigate whether the fixed antibody framework provides complementary sequence information for resolving this ambiguity. We introduce AbRefine, an information-aware fusion framework that combines backbone-conditioned structural predictions with framework-conditioned sequence distributions derived from a pretrained antibody language model with all designed CDRs masked. Its Information-Aware Fusion (IAF) module learns residue-specific fusion weights from the uncertainty, disagreement, and information content of the two predictive distributions, while keeping both pretrained experts frozen. Controlled interventions show that the signal is target specific: matched frameworks outperform shuffled frameworks by 7.37 percentage points, and the largest gain occurs at positions with high structural uncertainty and informative framework evidence (+5.77$ points). Compared with static fusion, IAF further improves AbMPNN recovery by +1.53 points (95% CI +0.93$ to +2.14$) and significantly reduces native negative log-likelihood for both AbMPNN and AntiFold, without detectable degradation in fold-back structural compatibility. These results show that framework context provides information beyond backbone geometry and that learning its residue-specific utility enables effective adaptive fusion for antibody CDR design.
]]></description>
<dc:creator><![CDATA[ Liu, Q., Xu, X., Cheng, Q., Meng, L., Wu, C., Yang, D., Gao, X., Chen, X. ]]></dc:creator>
<dc:date>2026-10-02</dc:date>
<dc:identifier>doi:10.64898/2026.09.27.754743</dc:identifier>
<dc:title><![CDATA[AbRefine: Framework-Conditioned Refinement for Structure-Guided Antibody CDR Design]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-02</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.30.755594v1?rss=1">
<title>
<![CDATA[
Inference on paths of yeast transcription pre-initiation complex assembly by model-based analysis of altered occupancy profiles 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.30.755594v1?rss=1
</link>
<description><![CDATA[
Dysfunctions in transcription cause severe human diseases. In eukaryotes, transcription initiation by RNA polymerase (Pol) II requires the formation of the pre-initiation complex (PIC) composed of general transcription factors (GTFs), Pol II, and an essential multiprotein coactivator, Mediator. In vitro experiments with GTFs and Pol II have depicted a linear sequence of PIC assembly. However, PIC assembly paths in vivo at the genomic scale and Mediator's role in these paths remain poorly understood. One strategy characterizes ChIP-seq occupancy alterations in Mediator mutants, but analysis of these data still lacks a formalized interpretation framework. In this study, we address this need by developing a quantitative framework that scores candidate assembly paths according to their ability to explain these alterations. This framework combines a model of PIC assembly based on a system of differential equations with parameter optimization based on a penalized criterion. We applied this new framework to datasets from two yeast Mediator mutants, in the Med17 and Med10 subunits. Our results show that the best-supported PIC assembly paths differ from those inferred in vitro. In particular, we show that it is important to distinguish the core TFIIH module from the TFIIK kinase module, and that similar paths starting with Mediator or TBP receive different levels of support. Furthermore, although multiple paths fit the data equally well, they differ markedly for specific gene groups. Taken together, the results show that the path can depend on the gene, and are also compatible with more than one path being operational in vivo for a single gene. This work improves our understanding of PIC assembly paths and transcription initiation, and opens perspectives for future modeling and experimental studies.
]]></description>
<dc:creator><![CDATA[ Novikova, E. A., Goldar, A., Nicolas, P., Soutourina, J. ]]></dc:creator>
<dc:date>2026-10-02</dc:date>
<dc:identifier>doi:10.64898/2026.09.30.755594</dc:identifier>
<dc:title><![CDATA[Inference on paths of yeast transcription pre-initiation complex assembly by model-based analysis of altered occupancy profiles]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-02</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.30.755779v1?rss=1">
<title>
<![CDATA[
SCITRAM: a Single-cell Integrated transcription regulation modeler for predicting master transcription factors 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.30.755779v1?rss=1
</link>
<description><![CDATA[
Single-cell omics technologies have provided unprecedented access to molecular cell heterogeneity, revolutionizing our understanding of complex living systems. Although generating such type of data has become a streamlined process, data processing still relies on command-line driven tools, or commercial solutions. Furthermore, single-cell transcriptomics analysis is usually reduced to stratifying cells on the grounds of gene expression signatures, while deconvolving gene regulatory programs (and their associated master transcription factors) responsible for cell heterogeneity are not classically addressed with the available tools, or performed under different data processing platforms. Herein, we present SCITRAM (Single-Cell Integrated TRAnscription regulation Modeler), a user-friendly stand-alone computational solution for predicting master transcription factors from single-cell transcriptomics data. SCITRAM first reconstructs a primary gene regulatory network and subsequently uses this network as template to model transcriptional cascades driven by the in-silico activation of transcription factors identified within the system. We validated SCITRAM performance across a variety of cell fate transition events including T cell exhaustion during solid tumor treatment, lineage stratification during pancreatic endocrinogenesis, as well as nervous tissue formation in cerebral organoid models. Overall, SCITRAM provides an accessible and integrated platform for processing (large amounts of) single-cell transcriptomics data, enabling the investigation of transcriptional regulatory programs and their master transcription factors without requiring a command-line driven environment.
]]></description>
<dc:creator><![CDATA[ Mendoza-Parra, M. A., Galindo-Albarra, A., Duvina, M., Barbao, P., Rodriguez-Garcia, A., Guedan, S., Heuser-Loy, C., Gattinoni, L. ]]></dc:creator>
<dc:date>2026-10-01</dc:date>
<dc:identifier>doi:10.64898/2026.09.30.755779</dc:identifier>
<dc:title><![CDATA[SCITRAM: a Single-cell Integrated transcription regulation modeler for predicting master transcription factors]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-01</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.28.754049v1?rss=1">
<title>
<![CDATA[
A cross-species meta-transcriptomic analysis of viral infection-induced shifts in microbial transcript profiles across insects. 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.28.754049v1?rss=1
</link>
<description><![CDATA[
Microbiome has co-evolved with insects over millions of years, establishing symbiotic relationships that influence their fitness and shape evolutionary trajectories. These long-term associations are shaped by the hosts ecology, diet, and physiology which determines microbial diversity. In turn, external stresses such as pathogenic infections can adversely affect the host microbiome, further compounding the effects of pathogenic infections on the host. Using a computational pipeline and publicly available transcriptomics datasets, we analyzed the effect of viral infections on transcript derived microbiome diversity and composition in various insects including Drosophila melanogaster, Acyrthosiphon pisum, Antheraea pernyi, and Culex pipiens. The meta-transcriptomic reanalysis provides an indirect but comprehensive comparison of infection-associated transcriptional microbiome changes across hosts with distinct life histories. Our study reveals that viral infections influence the host-associated microbial transcripts, leading to shifts in transcript-derived microbial profiles and offering insights into host-microbe dynamics. In Cx. pipiens, microbial transcript profiles remained relatively stable following RVFV exposure and were dominated by the bacterial endosymbiont Wolbachia. Our study highlights viral infections as key drivers of infection-associated shifts in microbial profiles and suggests a potential association between Wolbachia dominance and transcriptional microbiome stability in insects, with age as a contributing factor in shaping microbial transcript profiles in Culex. This cross-species framework uncovers conserved and host-specific patterns of virus-microbiome interactions, expanding our understanding of microbial resilience mechanisms in insects.
]]></description>
<dc:creator><![CDATA[ Mishra, A., Shelke, T., Gupta, V. ]]></dc:creator>
<dc:date>2026-10-01</dc:date>
<dc:identifier>doi:10.64898/2026.09.28.754049</dc:identifier>
<dc:title><![CDATA[A cross-species meta-transcriptomic analysis of viral infection-induced shifts in microbial transcript profiles across insects.]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-01</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.29.755321v1?rss=1">
<title>
<![CDATA[
Systematic exploration of predicted quaternary structures within pandemic-relevant viral proteomes 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.29.755321v1?rss=1
</link>
<description><![CDATA[
Mechanistic understanding of viral protein-protein interactions enables global health security and pandemic preparedness, yet experimental characterisation remains difficult, costly, and often restricted to specialised laboratories. Here, we describe an in silico campaign using AlphaFold2 and AlphaFold-Multimer to predict monomers and dimer structures within 2,812 viral proteomes from 23 viral families relevant to human health. We report high-confidence predicted structures of 5,279 hetero-, and 2,749 homo-dimers. Structural clustering of interfaces reduces the set fivefold to 1,598 consolidated by form and function, of which 471 (29.5%) lacked detectable similarity to experimentally determined interfaces in the PDB. These clusters yielded two notable findings: new data on vaccinology-relevant glycoproteins, and accurate prediction of the specificity and cleavage-site recognition of viral proteases, which are targets for anti-viral development. Viral monomers and high-confidence dimers are available for interactive browsing in the Pandemic Preparedness Portal of the AlphaFold Database (https://alphafold.ebi.ac.uk/), and the full dataset is available for bulk download.
]]></description>
<dc:creator><![CDATA[ Han, Y., Narain, R., Abbara, R., Afonso, M. Q. L., Austin, J., Bertoni, D., Cha, S., Cowen-Rivers, A. I., Ellaway, J. I. J., Fransos, D., Gion, K., Hsu, D., Hulo, C., Kim, R. S., Kovalevskiy, O., Laydon, A., Mason, J., Nair, S., Paramval, U., Patel, N., Patel, R., Pidruchna, I., Ratcliff, J., Tsenkov, M. I., Vanenzi, N. A. E., Vollmar, M., Wahome, N., Wong, E. L. H., Fleming, J. R., Velankar, S., Mirdita, M., Le Mercier, P., Steinegger, M., Dallago, C., Grove, J. ]]></dc:creator>
<dc:date>2026-10-01</dc:date>
<dc:identifier>doi:10.64898/2026.09.29.755321</dc:identifier>
<dc:title><![CDATA[Systematic exploration of predicted quaternary structures within pandemic-relevant viral proteomes]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-01</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.26.754620v1?rss=1">
<title>
<![CDATA[
Widespread 3-Base Periodicity in Complex DNA Mixtures and an Alignment-Free Algorithm to Uncover its Source 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.26.754620v1?rss=1
</link>
<description><![CDATA[
Background: Protein-coding DNA exhibits characteristic three-base periodicity. In unaligned DNA libraries, however, this signal should cancel if fragment boundaries are uniformly distributed across codon phases. Yet, we find that in multiple metagenomic libraries, base frequencies depended systematically on distance from fragment boundaries, producing a period-3 pattern. We investigated the origin of this unexpected signal. Results: The signal was observed in modern metagenomic datasets and widely across ancient-DNA datasets. We developed a new alignment-free algorithm to infer latent triplet-phase structure from nucleotide composition. The inferred shifts were strongly nonuniform at fragment boundaries, explaining the read-coordinate periodicity. We initially suspected a protocol artifact. To test an alternative, we derived a probabilistic framework linking nucleotide-dependent boundary selection to codon-phase frequencies. The framework showed that, because nucleotide composition differs among codon positions, preferential breakage at particular nucleotides biases the phases represented at fragment boundaries. Simulations confirmed this mechanism: purine-associated fragmentation produced strong period-3 coherence in a coding-dense bacterial genome, whereas no comparable genome-wide effect appeared in the human reference genome. Conclusions: The observed read-coordinate three-base periodicity does not require an intrinsic phase bias in the source DNA. Instead, sequence-dependent fragment formation or recovery can couple boundaries to codon phase. Purine-associated post-mortem fragmentation plausibly explains the widespread phenomenon in ancient-DNA libraries, although PCR- or library-specific mechanisms may contribute to it in modern datasets. Our alignment-free algorithm provides a practical method for estimating and normalizing phase shifts before searching for positional nucleotide patterns, preventing boundary-associated bias from being mistaken for an intrinsic biological pattern.
]]></description>
<dc:creator><![CDATA[ hecht, n., Duek, N., Dotan, Y., Rosset, S., Slon, V., Safra, M. ]]></dc:creator>
<dc:date>2026-10-01</dc:date>
<dc:identifier>doi:10.64898/2026.09.26.754620</dc:identifier>
<dc:title><![CDATA[Widespread 3-Base Periodicity in Complex DNA Mixtures and an Alignment-Free Algorithm to Uncover its Source]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-01</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.26.754622v1?rss=1">
<title>
<![CDATA[
In-silico Exploration of Drug Leads Targeting Mycobacterial Ag85C 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.26.754622v1?rss=1
</link>
<description><![CDATA[
This study aims to identify potential inhibitors of the Ag85C enzyme from Mycobacterium tuberculosis through primary docking, pharmacophore modelling, and ADME/pharmacokinetics analysis. Delamanid and Pretomanid emerged as the top hits in primary docking, demonstrating strong binding affinities and interactions with key active site residues. A robust pharmacophore model was developed, identifying essential features for binding and used to screen 4254 antituberculotic compounds, yielding three promising candidates: F0314-0040, F0398-0505, and F0349-2716. These compounds exhibited high binding affinities in major docking studies. ADME and pharmacokinetics analysis revealed favorable Drug likeness and absorption profiles, supporting their potential as effective antituberculotic agents. The integration of computational methods provided a comprehensive approach for rational drug design.
]]></description>
<dc:creator><![CDATA[ Singh, K. R., Anuradha, D. C. M., Suresh kumar, D. C. ]]></dc:creator>
<dc:date>2026-10-01</dc:date>
<dc:identifier>doi:10.64898/2026.09.26.754622</dc:identifier>
<dc:title><![CDATA[In-silico Exploration of Drug Leads Targeting Mycobacterial Ag85C]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-01</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.26.754587v1?rss=1">
<title>
<![CDATA[
QuImputer: an LD-informed QUBO/Ising formulation of genotype imputation with quantum-hardware execution 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.26.754587v1?rss=1
</link>
<description><![CDATA[
Genotype imputation offers a biological application for optimization methods that may eventually run on larger quantum computers. We developed QuImputer to encode allele-frequency priors and signed linkage disequilibrium (LD) in a reduced Ising/QUBO objective, with only missing genotypes represented as optimization variables. The same objective can be passed to exhaustive classical search, quantum approximate optimization algorithm simulation, or quantum hardware. We evaluated the formulation on 100 rice accessions across 12 chromosomes, estimating LD and allele frequency from the same post-masking observed matrix supplied to Beagle 5.4. Evaluation was restricted to known truth at sites with exact dosage coding and an alternate allele that was not the major allele, with MAF below 0.50, yielding 1.88 billion scored genotypes. Beagle achieved higher rare-ALT sensitivity (2.52% versus 0.54%), F1 (0.047 versus 0.011), Matthews correlation coefficient (0.088 versus 0.044), and pooled hard-call dosage R-squared (0.231 versus 0.014). QuImputer produced fewer false ALT calls, with Type I error of 0.038% versus 0.223%. Analyses of a chromosome-1 subset showed limited local LD support for many masked rare-ALT carriers. In five small IBM Fez validation problems, the selected minimum-energy outputs matched unique all-reference exact optima; the pooled frequency of those optima was 1.21-2.63%. Separate component-wise analysis of 12 sparse instances, 11 of which have a recorded hardware result, found all-reference exact optima, delimiting the interpretation of their hardware concordance. QuImputer provides an LD-informed formulation and a route to quantum-hardware execution, together with an empirical assessment of the limitations of the evaluated configuration. These results establish a basis for studying richer information models and alternative classical or quantum solvers.
]]></description>
<dc:creator><![CDATA[ Kim, M., Lee, T.-H. ]]></dc:creator>
<dc:date>2026-10-01</dc:date>
<dc:identifier>doi:10.64898/2026.09.26.754587</dc:identifier>
<dc:title><![CDATA[QuImputer: an LD-informed QUBO/Ising formulation of genotype imputation with quantum-hardware execution]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-01</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.26.754480v1?rss=1">
<title>
<![CDATA[
IQC: A Novel Criterion for Assessing Feature Selection Stability in High-Dimensional Analyses 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.26.754480v1?rss=1
</link>
<description><![CDATA[
In recent years, the issue of reproducibility in scientific studies has been at the forefront of scientific discussion. Many of these issues arise due to poor employment of statistical methods, making it crucial to address these knowledge gaps for frequently used methods. Elastic net regression (ENET) is one such algorithm used frequently in biology and public health due to its ability to select important predictors (also known as "features") from large data sets. However, ENET struggles with consistent feature selection and hyperparameter choice in these applications due to low observation counts and noisy data. This limitation hinders ENET's effectiveness in biological research due to low model reproducibility and predictor interpretability. We propose the Instability Quotient Criterion (IQC) as a metric to quantify the consistency of selected features for feature selection models such as ENET, with a specific emphasis on biological applications. IQC leverages ensemble modeling and Shannon's entropy in order to determine the stability of selected features derived from ENET models. We demonstrate IQC's robustness via simulation studies, showing that IQC is capable of detecting significant decreases in model stability with lower observation counts, higher dimensionality, and higher proportions of noisy signal. We additionally utilize simulation to provide a naive threshold for low versus high stability. Finally, we applied IQC to a case study to illustrate how our method can be integrated into a larger analysis pipeline when looking at microarray gene expression data. IQC indicated stability issues for differentially-expressed genes when classifying acute leukemia subtypes in spite of high classification success. Lastly, we discuss how model instability impedes interpretation of biologically significant genes for discerning cancer subtypes. Overall, we present IQC as a robust evaluation metric that can easily be incorporated into high-dimensional biological analyses.
]]></description>
<dc:creator><![CDATA[ Moxley, T. A., Ridenhour, B. J. ]]></dc:creator>
<dc:date>2026-10-01</dc:date>
<dc:identifier>doi:10.64898/2026.09.26.754480</dc:identifier>
<dc:title><![CDATA[IQC: A Novel Criterion for Assessing Feature Selection Stability in High-Dimensional Analyses]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-01</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.26.754596v1?rss=1">
<title>
<![CDATA[
LOL: a Python package for leakage-aware phenotype prediction and confounding diagnostics in GEO transcriptomics datasets 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.26.754596v1?rss=1
</link>
<description><![CDATA[
Phenotype prediction from public gene-expression data is straightforward to implement but reliable prediction is harder. Leakage during feature selection can inflate performance estimates. Class imbalance can hide poor performance in minority classes. Confounds like tissue type, repeated sampling, and demographic differences can also create patterns that resemble phenotype-associated biology. Here, I present LOL (Latent Omics Learning), a Python package for phenotype prediction from bulk GEO RNA-seq and microarray data. LOL uses leakage-safe preprocessing, imbalance-aware evaluation, and group-aware validation where repeated samples are present. Each analysis produces two audit diagnostics. The leakage inflation index ({Delta}LII) measures the change in performance between naive and leakage-safe evaluation. PVCA-lite examines associations between the latent representation and available metadata covariates, including phenotype. I evaluated the workflow across six independent GEO cohorts. The datasets covered microarray and RNA-seq platforms, binary and multi-class phenotypes, and sample sizes ranging from 18 to 566. The analyses produced a varied set of outcomes. Across diverse GEO datasets, LOL can reveal phenotype-associated expression patterns while identifying factors distort their interpretation. These include information leakage, class imbalance, repeated sampling, tissue differences, and demographic or other metadata-associated confounding. The resulting diagnostics help distinguish robust predictions from results that require further scrutiny. LOL does not introduce a new representation-learning algorithm. Its contribution is a reproducible workflow that places validation and diagnostic checks alongside prediction. The resulting analysis can expose weaknesses in an apparently convincing result and make those limitations visible. LOL is freely available under the MIT license from https://github.com/ngangao/lol-omics.
]]></description>
<dc:creator><![CDATA[ Muigano, M. N. ]]></dc:creator>
<dc:date>2026-10-01</dc:date>
<dc:identifier>doi:10.64898/2026.09.26.754596</dc:identifier>
<dc:title><![CDATA[LOL: a Python package for leakage-aware phenotype prediction and confounding diagnostics in GEO transcriptomics datasets]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-01</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.26.754618v1?rss=1">
<title>
<![CDATA[
Ultrafast and Accurate Selection of High-Quality Protein Complex Models from Large-Scale Prediction 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.26.754618v1?rss=1
</link>
<description><![CDATA[
In the post-AlphaFold era, advances in protein complex structure prediction have enabled the generation of thousands of candidate models per target, making the accurate and efficient selection of high-quality models from large candidate pools a critical challenge. Existing model-selection approaches commonly employ estimated model accuracy (EMA) methods, which include single-model and consensus-based methods. The former lack cross-model comparative information, whereas the latter often incur high computational costs from pairwise structural alignments and are influenced by candidate-pool quality. Here, we present DeepUMQA-Selection, a two-stage framework combining single-model accuracy estimators for preliminary selection with Ultrafast Shape Recognition (USR)-based structural alignment for consensus selection. Preliminary selection enriches the candidate pool with high-quality models and reduces subsequent structural comparisons, while USR-based structural alignment enables efficient consensus selection without explicit structural superposition. On the retrospective CASP16 QMODE3 benchmark dataset, DeepUMQA-Selection outperformed all participating methods and predictor-derived confidence metrics, as measured by mean Top-5 weighted penalty. The advantage was particularly pronounced for heteromeric targets, with a 13.9% improvement in model-selection performance relative to the best-performing CASP16 participating method. Furthermore, a confidence-assisted strategy enabled the framework to accommodate larger protein complexes while maintaining superior model-selection performance over competing methods, supporting its scalability. Runtime benchmarking showed that the proposed USR-based structural alignment achieved a substantial 317-fold speedup over US-align for candidate pools containing 1,000 models, demonstrating its computational efficiency for large-scale model selection. Our results demonstrate that DeepUMQA-Selection enables accurate and computationally efficient selection of high-quality protein complex models from large-scale candidate pools.
]]></description>
<dc:creator><![CDATA[ Liang, F., Xie, L., Ye, E., Pan, Z., Xu, T., Wang, H., Wei, P., Liu, J., Zhang, G. ]]></dc:creator>
<dc:date>2026-10-01</dc:date>
<dc:identifier>doi:10.64898/2026.09.26.754618</dc:identifier>
<dc:title><![CDATA[Ultrafast and Accurate Selection of High-Quality Protein Complex Models from Large-Scale Prediction]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-01</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.25.754557v1?rss=1">
<title>
<![CDATA[
Quantifying the Provenance-to-Function Gap in Antidiabetic Peptide Prediction: Homology-Aware Evaluation and the ADP-Hybrid Baseline 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.25.754557v1?rss=1
</link>
<description><![CDATA[
Antidiabetic peptides (ADPs) are short bioactive sequences of therapeutic interest, and sequence-based classifiers prioritise experimental candidates. A classifier is useful only if its accuracy transfers to unseen sequences, which depends on how the benchmark was as sembled and partitioned. We term the distance between what such a classifier is scored on and the function it is meant to predict the provenance-to-function gap, and we quantify it. We re-evaluate the two-layer ADP benchmark of Basith et al. under a protocol that groups homologous peptides rather than splitting them at random. Of the 877 ADPs, 218 attribute to a precursor protein by exact substring containment; within that subset the second-layer label coincides exactly with precursor identity, all 140 human-insulin fragments carrying the type-1 label and all 76 bovine milk-protein fragments the type-2 label, without exception. Accuracy tracks identity to the training set, rising from a Matthews correlation coefficient (MCC) of 0.39-0.43 below 50% identity to 0.92-0.96 between 70% and 90%, and peptide length alone reaches MCC 0.619 on held-out data. Auditing sixteen further peptide bench marks shows the coupling is not confined to this resource: fragment families share a label more often than chance in eleven of fourteen testable datasets, and not in three, so the prop erty is common rather than universal. Rebuilding the published architecture on identical folds shows its advantage over a single classifier is a function of the split: present under random partitioning, absent once homologues are separated. ADP-Hybrid, one tree-ensemble classifier per layer, reaches MCC 0.843 on layer 1 against a published 0.841 at 21-38 times the inference throughput, and 0.801 against 0.858 on layer 2. We release the protocol and the controls that expose these properties.
]]></description>
<dc:creator><![CDATA[ Islam, A., Hosen, M. F., Basar, M. A., Mollah, M. S. H., Uddin, M. S. ]]></dc:creator>
<dc:date>2026-10-01</dc:date>
<dc:identifier>doi:10.64898/2026.09.25.754557</dc:identifier>
<dc:title><![CDATA[Quantifying the Provenance-to-Function Gap in Antidiabetic Peptide Prediction: Homology-Aware Evaluation and the ADP-Hybrid Baseline]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-01</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.25.754454v1?rss=1">
<title>
<![CDATA[
Three-Year Longitudinal Analysis of Splicing Changes in Myotonic Dystrophy Type 1 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.25.754454v1?rss=1
</link>
<description><![CDATA[
Myotonic dystrophy type 1 (DM1) is caused by a CTG expansion in the 3' untranslated region of the DMPK gene that results in the expression of expanded CUG RNA which trap splicing factors resulting in widespread alternative splicing dysregulation. A panel of these splicing events, or splicing index (SI), is currently used to measure splicing dysregulation in DM1 patients. Despite significant interest in splicing, little is known about its progression over time in DM1. The objective of this study was to evaluate the changes in SI over 3 years in a well-phenotyped cohort of DM1 participants and identify factors contributing to these changes. RNA was extracted from vastus lateralis (VL) biopsies of 20 DM1 participants taken 3 years apart and the SI score was measured using targeted RNA sequencing. Baseline SI scores correlated with percent predicted muscle strength of the ankle dorsiflexors, knee extensors and hand grip. On average, SI scores worsened over 3 years, but the progression varied markedly among participants. Adult/juvenile phenotypes and lower baseline SI scores were among the factors associated with higher progression of SI scores over time. These findings emphasize the importance of integrating clinical phenotype and baseline molecular severity when interpreting longitudinal SI change, particularly in the context of clinical trials.
]]></description>
<dc:creator><![CDATA[ Legare, C., Planco, L., Merritt, R., Ripollone, J., Conner, S., Desrochers, L., Nigim, F., Roussel, M.-P., Gagnon, C., Cleary, J. D., Berglund, A., Duchesne, E. ]]></dc:creator>
<dc:date>2026-10-01</dc:date>
<dc:identifier>doi:10.64898/2026.09.25.754454</dc:identifier>
<dc:title><![CDATA[Three-Year Longitudinal Analysis of Splicing Changes in Myotonic Dystrophy Type 1]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-01</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.25.754484v1?rss=1">
<title>
<![CDATA[
Octave: Scale-resolved Evaluation of Spatial Gene Expression Prediction from Histology 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.25.754484v1?rss=1
</link>
<description><![CDATA[
Spatial gene expression prediction from histology is typically evaluated by the mean per-gene Pearson correlation (PCC). PCC does not distinguish spatial scales: coarse tissue organization alone can earn most of the score. A domain oracle makes this precise: given an image-derived partition of the tissue, the predictor that assigns each domain its measured mean expression is the best that any predictor constant on that partition can do. On a widely used benchmark with a ~100 um pitch, a learnable version that estimates each domain mean from training data reaches 82% of the PCC of a trained model, and a partition of the spot coordinates alone, given measured means, reaches 123%. At that pitch the score rewards coarse spatial structure whatever its source. We introduce OCTAVE, a scale-resolved evaluation that filters prediction and measurement into spatial bands with widths in micrometers and scores their agreement in each band. On 16 um Xenium data, differences that look modest under PCC grow at finer scales: at ~20 um the oracle shortfall relative to the trained model is 2.4 times its PCC shortfall at the median and exceeds it in all 13 specimens. Across 57 image encoders the finest band reorders 261 of the 1596 encoder pairs relative to PCC. The benchmark own finest band is 50 um wide, coarser than the scale at which these differences appear, so no score on that pitch can directly validate finer structure. OCTAVE reports how well a model predicts spatial expression at each spatial scale.
]]></description>
<dc:creator><![CDATA[ Wang, Q., Gu, S., Lan, Q., Jiang, X., Chen, Y., Song, Q. ]]></dc:creator>
<dc:date>2026-10-01</dc:date>
<dc:identifier>doi:10.64898/2026.09.25.754484</dc:identifier>
<dc:title><![CDATA[Octave: Scale-resolved Evaluation of Spatial Gene Expression Prediction from Histology]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-01</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.25.754510v1?rss=1">
<title>
<![CDATA[
Rank-Preserving Alignment Enables Cross-platform Learning and Phenotyping for Single-Cell and Spatial Proteomics 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.25.754510v1?rss=1
</link>
<description><![CDATA[
Decades of antibody-based protein profiling have generated valuable cohorts across evolving cytometry, sequencing and spatial imaging platforms. Differences in marker coverage, signal scale and antibody performance hinder joint analysis and reuse of cohorts with long-term clinical follow-up. Widely used RNA-based integration methods prioritize shared embeddings rather than directly comparable protein measurements. To fill in the gap, we developed RAMP (Rank-preserving Alignment of Multi-platform Proteomics), an unsupervised framework that aligns shared proteins without cell-type annotations. It combines sample-specific anchoring and cross-platform cell matching with bounded monotone transformations. The corrected values retain each protein marker's expression order across cells within a sample. Across seven integration tasks, RAMP achieved the strongest overall balance of batch correction and cell-type preservation. RAMP also improved marker-threshold transfer while preventing expression-order reversals observed with some comparison methods. These reversals made cell types with originally high marker expression appear lower-expressing than other populations. Such overcorrection can mislead cell identification and protein comparisons, even when the integrated embedding shows improved dataset mixing. By preserving each marker's expression order within a sample, RAMP protects the relationships needed to interpret normalized protein measurements. To investigate whether better harmonized inputs improve downstream learning, we used linear, nearest-neighbor and small neural-network models as a controlled, small-scale testbed. This approach tests an input quality question relevant to proteomic foundation models without undertaking large-scale pretraining. Under matched supervision, RAMP improved cross-platform annotation across all three model families and improved missing-marker prediction over unaligned inputs. These gains demonstrate the value of harmonized inputs for predictive learning and motivate evaluating RAMP as preprocessing for proteomic foundation models. In bone marrow, we used high-quality single-cell protein measurements to impute the failed CODEX CD19 channel, recovering B-cell spatial distributions. The completed panel improved B-cell separation from natural killer cells in the joint embedding. In melanoma, a proliferating CD8 T-cell state near tumor cells was associated with longer survival. Transfer such a spatial information to CITE-seq connected tumor proliferation and infiltration state to an adhesion and activation program involving PD-1, LFA-1 and CD2. In lung cancer, we used single-cell CyTOF measurements to predict unmeasured HLA-ABC states in vascular cancer-associated fibroblasts profiled by spatial imaging mass cytometry (IMC). Higher predicted scores were associated with lower immune-cell fractions and worse disease-free survival in the IMC cohort. In head and neck cancer, transferring a fibroblast/macrophage spatial niche score to CITE-seq identified a CD39/TIGIT-associated CD8 T-cell phenotype. Protein and RNA profiles supported checkpoint-associated features, and transfer back to CODEX showed correspondence with measured LAG3. RAMP enables existing single-cell and spatial cohorts to support new discoveries by connecting molecular programs, tissue organization and clinical outcomes across platforms.
]]></description>
<dc:creator><![CDATA[ Xu, Q., Le, T., Luo, Z., Ly, C., Yan, Y., Zheng, Y. ]]></dc:creator>
<dc:date>2026-10-01</dc:date>
<dc:identifier>doi:10.64898/2026.09.25.754510</dc:identifier>
<dc:title><![CDATA[Rank-Preserving Alignment Enables Cross-platform Learning and Phenotyping for Single-Cell and Spatial Proteomics]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-01</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.25.754276v1?rss=1">
<title>
<![CDATA[
Sparse dynamic graphical models for longitudinal non-Gaussian data 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.25.754276v1?rss=1
</link>
<description><![CDATA[
Longitudinal non-Gaussian data are common in biomedical and microbiome studies, where responses may not be normally distributed. However, many existing dynamic graphical models are developed for Gaussian responses, while many non-Gaussian graphical models focus on cross-sectional dependence and do not explicitly estimate temporal transition effects. We propose a latent Gaussian dynamic graphical model for longitudinal non-Gaussian responses, including binary, count, and compositional outcomes. The proposed framework jointly estimates the contemporaneous conditional dependence network and the lag-one temporal dependence structure. We further develop a penalized Monte Carlo expectation maximization procedure for estimation. Simulation studies across multiple graph structures and outcome types show that the proposed method can recover both dependence structures, with more stable recovery for the lag-one temporal dependence structure and more structure-dependent recovery for the contemporaneous conditional dependence network. We apply the proposed method to longitudinal data from the Integrative Human Microbiome Project (HMP2) using a multinomial observation model. The results reveal persistent lag-one temporal dependence and contemporaneous conditional dependence among dominant microbial taxa, demonstrating the utility of the proposed framework for dynamic network analysis of longitudinal non-Gaussian biomedical data.
]]></description>
<dc:creator><![CDATA[ Zheng, Y., Solis-Lemus, C. ]]></dc:creator>
<dc:date>2026-10-01</dc:date>
<dc:identifier>doi:10.64898/2026.09.25.754276</dc:identifier>
<dc:title><![CDATA[Sparse dynamic graphical models for longitudinal non-Gaussian data]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-01</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.27.754822v1?rss=1">
<title>
<![CDATA[
Evidence Scaling for Zero-Shot Protein Reasoning with Large Language Models 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.27.754822v1?rss=1
</link>
<description><![CDATA[
Large language models (LLMs) show emerging zero-shot capability for protein variant prediction, yet still lag behind specialized protein models. We ask whether this gap can be reduced by scaling access to biological evidence rather than adapting model parameters. We introduce BioEvidence, a training-free and model-agnostic interface that converts structural and evolutionary information from standard biological tools into compact evidence for frozen LLMs. On the ProteinGym benchmark, we observe evidence scaling: performance improves as evidence becomes richer. Structural and evolutionary evidence each improve performance, and combining them yields further gains, while mismatching the same evidence to the wrong variants degrades performance below the no-evidence baseline. Notably, BioEvidence enables zero-shot ranking to reach strong specialized protein predictors on matched evaluations, and the improvement persists on post-cutoff data released after the model's knowledge cutoff. Evidence also interacts with conventional scaling: for GPT-5.6 Sol, evidence at low reasoning effort outperforms the no-evidence condition at medium effort, while a six-model analysis associates stronger no-evidence performance with larger margins over evolutionary rank fusion. These results identify external evidence as a complementary scaling axis for scientific prediction alongside model capability and inference effort.
]]></description>
<dc:creator><![CDATA[ Hao, Z., Wang, C., Li, D., Wang, Y. ]]></dc:creator>
<dc:date>2026-10-01</dc:date>
<dc:identifier>doi:10.64898/2026.09.27.754822</dc:identifier>
<dc:title><![CDATA[Evidence Scaling for Zero-Shot Protein Reasoning with Large Language Models]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-01</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.25.754486v1?rss=1">
<title>
<![CDATA[
Limitations of differential expression for cell-type marker discovery in single-cell RNA sequencing atlases 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.25.754486v1?rss=1
</link>
<description><![CDATA[
Cell-type marker genes are definitional characteristics of a cell type, and in single-cell RNA sequencing atlases they are commonly nominated through differential expression (DE) testing. The Wilcoxon rank-sum test is the default method for marker gene discovery in Seurat and a recommended method in Scanpy, two widely used computational software packages for scRNA-seq analysis. Structurally, Wilcoxon's standardized statistic grows with sample size at a fixed underlying effect, meaning computed rankings across clusters of unequal size are incomparable. We quantified this effect across three atlas datasets - lung, kidney, and brain. Mean Wilcoxon score scaled with log cluster size at Pearson correlation r = 0.90-0.95 in each organ. In contrast, a classification-based comparator, NS-Forest, showed low correlation with cluster size (r = 0.10-0.28). This DE-cluster size relationship remained essentially unchanged after controlling for classification performance through partial correlation, decayed monotonically under cluster size downsampling, and was reproduced in simulation with marker quality fixed. The Wilcoxon rank-sum test can be transformed into a rank-based effect size by normalizing by the target-background pair count, removing the positive size dependence. However, this effect size nominates an identical set of candidate genes as the original statistic, indicating the cluster size dependence and poor specificity of DE nominated genes are separate problems. Genes nominated under default settings were substantially less specific to their target clusters, with median on-target fractions of 0.204 (lung), 0.166 (kidney), and 0.083 (brain), compared to NS-Forest's marker genes (0.580, 0.687, 0.364), revealing the relatively poor specificity that highly scoring DE genes exhibit.
]]></description>
<dc:creator><![CDATA[ Doggett, K., Pintard, D., Scheuermann, R. H., Zhang, Y. ]]></dc:creator>
<dc:date>2026-10-01</dc:date>
<dc:identifier>doi:10.64898/2026.09.25.754486</dc:identifier>
<dc:title><![CDATA[Limitations of differential expression for cell-type marker discovery in single-cell RNA sequencing atlases]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-01</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.25.754455v1?rss=1">
<title>
<![CDATA[
AmyloCore-ML: AI/Machine Learning Enabled Identification of Amyloid Fibril Core Regions Using Protein Language Models 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.25.754455v1?rss=1
</link>
<description><![CDATA[
In many neurodegenerative and systemic disorders, proteins can form insoluble protein aggregates called amyloids. Identification of amyloid forming regions in a protein remains central to understanding of its aggregation behaviour. Many predictors have been formed in the past and until recently which to estimate the aggregation-prone regions, which may not correspond to total residues found in actual disease associated fibril structures. To address this, we formed a sequence-based based predictor to ascertain structural amyloid-core propensity. The predictor was developed using experimentally determined fibril structures obtained from Amyloid Atlas. Ordered residues in the experimental structures were designated as core, and unresolved sampled from the same proteins as matched controls as non-core. Core and non-core regions were encoded using various physicochemical descriptors and protein language model (PLM) embeddings (ESM-2, ANKH, ProtT5) and evaluated using protein-grouped cross-validation. PLMs consistently outperformed physicochemical descriptors, with the best models reaching AUROC values of ~0.88 and AUPRC values of ~0.85. Locked full-length protein scans further localized experimental cores, with ESM-2/ExtraTrees W21 achieving AUROC 0.833, AUPRC 0.751, and a mean peak distance of 12.4 residues. In comparative benchmarking, our predictor showed better performance metrics than CrossBeta and AggrescanAI. The framework therefore provides residue-resolved prediction of structurally incorporated amyloid-core regions directly from sequence. AmyloCore-ML is easily accessible through an interactive Google Colab notebook, enabling sequence-based amyloid fibril-core prediction without local installation or dedicated computing infrastructure.
]]></description>
<dc:creator><![CDATA[ Singh, J., Jangid, R., Srivastava, A. ]]></dc:creator>
<dc:date>2026-10-01</dc:date>
<dc:identifier>doi:10.64898/2026.09.25.754455</dc:identifier>
<dc:title><![CDATA[AmyloCore-ML: AI/Machine Learning Enabled Identification of Amyloid Fibril Core Regions Using Protein Language Models]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-01</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.30.755756v1?rss=1">
<title>
<![CDATA[
The recoverable resolution of cellular perturbation-response prediction 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.30.755756v1?rss=1
</link>
<description><![CDATA[
Virtual cell models aim to predict how cells respond to perturbations, yet accurate prediction of post-perturbation gene expression can conceal failures to recover context-specific responses. Here we identify context geometry compression (CGC), in which models preserve cellular identity and shared perturbation effects while collapsing differences in how the same perturbation acts across contexts. Across biological systems, reproducible context-specific responses remained difficult to recover; in a large drug-response matrix, this persisted even with nearly complete perturbation coverage. The same drug-response predictions achieved higher normalized recovery in predefined pathway summaries than in gene-level response profiles, while coarse response programs were more predictable in human lymphoblastoid cell lines. Compact baseline transcriptomic representations retained most full-RNA predictive performance for patient-specific ex vivo drug responses. Together, these findings establish recoverable biological resolution, rather than measurement dimensionality, as a criterion for evaluating virtual cell predictions and matching model outputs to the biological detail supported by evidence.
]]></description>
<dc:creator><![CDATA[ Huang, Y., Wang, H., Li, C., Wilson, P. C. ]]></dc:creator>
<dc:date>2026-10-01</dc:date>
<dc:identifier>doi:10.64898/2026.09.30.755756</dc:identifier>
<dc:title><![CDATA[The recoverable resolution of cellular perturbation-response prediction]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-01</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.25.754409v1?rss=1">
<title>
<![CDATA[
SFUMATO: Bayesian probabilistic clustering for uncertainty-aware spatialtranscriptomics analysis and mapping 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.25.754409v1?rss=1
</link>
<description><![CDATA[
In situ sequencing methods provide subcellular-resolution gene expression data while preserving spatial context, enabling the integrated analysis of genomics and tissue morphology. However, many clustering workflows represent spatial transcriptomic structure through hard labels, imposing sharp boundaries between transcriptional domains and limiting the visualisation of gradual transitions, mixed signals, and uncertainty in cluster assignment, resulting in potentially misleading interpretations of tissue organisation. Here we present SFUMATO (Segmentation-Free Uncertainty Mapping and Analysis of Transcriptomic Organisation), a GPU-accelerated Bayesian probabilistic clustering and visualisation framework for spatial transcriptomics. SFUMATO builds on segmentation-free, multi-scale transcript binning and applies a Bayesian Gaussian mixture model to infer posterior probabilities over transcriptional components at each spatial location. These posterior probabilities provide a unified representation for hierarchical clustering, uncertainty-aware visualisation, and quantitative estimation of transcript-defined area fractions. SFUMATO assigns related colours to transcriptionally related components and blends colours according to posterior probabilities, producing maps in which sharp boundaries, gradual transitions, and ambiguous regions are represented directly in the visualisation. Classical hard labels remain recoverable from the same posterior distributions, preserving compatibility with downstream analyses that require discrete clusters. We evaluate SFUMATO on public 10x Genomics Xenium mouse brain and human breast cancer datasets. In mouse brain, SFUMATO preserves local transcriptional similarity in its colour encoding, maintains hierarchically consistent visualisations across different numbers of clusters, and recovers cell-associated spatial organisation without requiring prior cell segmentation. In breast cancer, SFUMATO-derived posterior maps can be converted into semantic masks concordant with pathology-associated regions, particularly invasive carcinoma, while highlighting the greater heterogeneity of DCIS-associated tissue. Together, these results show that SFUMATO provides an interpretable probabilistic representation of spatial transcriptomics that connects clustering, visualisation, and semantic segmentation across cell-associated and niche-level scales. More broadly, SFUMATO also provides a compact and continuous representation of highly multiplexed spatial transcriptomic information, offering a flexible input for future downstream analyses and multimodal integration.
]]></description>
<dc:creator><![CDATA[ Giustolisi, A., Avenel, C., Wählby, C. ]]></dc:creator>
<dc:date>2026-10-01</dc:date>
<dc:identifier>doi:10.64898/2026.09.25.754409</dc:identifier>
<dc:title><![CDATA[SFUMATO: Bayesian probabilistic clustering for uncertainty-aware spatialtranscriptomics analysis and mapping]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-01</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.25.754345v1?rss=1">
<title>
<![CDATA[
biomes: An R package for reproducible occurrence-to-biome classification using 31 global biome schemes 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.25.754345v1?rss=1
</link>
<description><![CDATA[
Abstract 1. Biomes are continental-scale vegetation units widely used in ecological and evolutionary research, often together with large-scale species-distribution data. Many conceptually distinct biome schemes exist, differing in the underlying definition, delineation methodology and spatial resolution, yet most studies using biomes default to a few readily available schemes rather than choosing the one best suited to the research question. 2. We present biomes, an R package distributing spatially explicit raster layers of 31 published global biome schemes in a harmonised format at a 10 x 10 km resolution, together with metadata and a concise description. To facilitate their use, biomes provides a user-friendly classification workflow to classify occurrence records into specific biomes and a data-driven ranking algorithm to help researchers to choose the most suitable biome scheme for their data and research question in a transparent, reproducible way. 3. The biomes workflow can be applied to any dataset of georeferenced terrestrial occurrences. It (1) accepts either user-provided occurrence records or records downloaded by the package, (2) ranks the schemes according to their fit to the occurrence data, (3) assigns the records to specific biomes, and (4) summarises the resulting taxon-biome patterns. We demonstrate this workflow using 17,030 occurrence records of 185 species of Bombacoideae (Malvaceae). The analysis identified the vegetation-based biome scheme of Ramankutty and Foley (1999) as the best fit and recovered the forest and savanna biome signal known in the group from the literature, with tropical evergreen woodland being the biome with the largest share of records (46.1%) and highest number of species (162 of 185 species; 87.6%). 4. biomes provides user-friendly access to technically standardised and spatially explicit layers of 31 biome schemes, and thereby enables the transparent, data-driven choice of biome schemes tailored to a specific research data and questions, improving the reproducibility and the specificity of biome use across ecological and evolutionary research.
]]></description>
<dc:creator><![CDATA[ Gross, H. C., Zizka, A., Fischer, J.-C., Walentowitz, A. ]]></dc:creator>
<dc:date>2026-10-01</dc:date>
<dc:identifier>doi:10.64898/2026.09.25.754345</dc:identifier>
<dc:title><![CDATA[biomes: An R package for reproducible occurrence-to-biome classification using 31 global biome schemes]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-01</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.25.754162v1?rss=1">
<title>
<![CDATA[
Integrative Multi-Omics Analysis Reveals Molecular Signatures of Age-Related Decline in Vervet Leg Skeletal Muscle 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.25.754162v1?rss=1
</link>
<description><![CDATA[
Age-related decline in skeletal muscle function, often manifesting as slower gait speed, contributes significantly to frailty and loss of independence. To uncover the molecular basis of this decline, we investigated coordinated changes across the transcriptome, miRNome, and methylome in vastus lateralis muscle from female vervet monkeys (Chlorocebus aethiops sabaeus), contrasting middle-aged animals with normal gait speed versus older animals with declining gait speed. Through multi-omics integration, we identified a molecular signature that sharply distinguishes these two groups. This signature involves key regulatory molecules, including miR-181a-5p, miR-320b and miR-425 control of SYNCRIP, miR-193b-3p regulation of MCL1 and PLAU, as well as a multi-omic regulatory network governing ZNF274 and ANP32E gene expression, and altered levels of mRNAs such as VEZT. Functional annotation implicates miRNA signatures with inverse expression of their mRNA targets operating within systems of skeletal muscle biology, metabolism, and potential immune-inflammation mechanisms. Collectively, these findings reveal an interconnected network of epigenetic, post-transcriptional, and transcriptional alterations associated with impaired muscle function during aging. This molecular signature provides novel insights into potential mechanisms driving sarcopenia and highlights specific molecules as potential targets for interventions aimed at preserving muscle health.
]]></description>
<dc:creator><![CDATA[ Patel, M. S., Frye, B., Negrey, J. D., Cox, L. A., Li, G., Jorgensen, M. J., Register, T., Kavanagh, K., Shively, C. A., Quillen, E. E. ]]></dc:creator>
<dc:date>2026-10-01</dc:date>
<dc:identifier>doi:10.64898/2026.09.25.754162</dc:identifier>
<dc:title><![CDATA[Integrative Multi-Omics Analysis Reveals Molecular Signatures of Age-Related Decline in Vervet Leg Skeletal Muscle]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-01</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.09.25.754380v1?rss=1">
<title>
<![CDATA[
Rank-based integration identifies convergent disease mechanisms across omics 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.09.25.754380v1?rss=1
</link>
<description><![CDATA[
The rapid growth of omics studies offers new opportunities to uncover disease mechanisms, yet findings from individual studies or single omics may be influenced by cohort-, platform- and layer-specific variation. Integrating evidence across studies and omics layers can reveal robust biological signals that are reproducible across molecular layers, but differences in assay platforms, sample sizes and feature coverage complicate direct combination of summary statistics. Here we present the Omics Rank-Based Integration Tool (ORBIT), a direction-aware framework for prioritizing concordant signals across omics layers using only within-dataset feature ranks and effect directions. ORBIT tests directional rank consistency across omics layers against a closed-form variance-gamma null distribution while accounting for inter-dataset correlation and incomplete feature overlap. Simulations show that ORBIT maintains well-calibrated false-positive rates under inter-omics correlation, increasing numbers of omics layers and missingness, while detecting concordant signals with increasing power as the number of layers grows. We applied ORBIT to tubulointerstitial transcriptomic data from patients with chronic kidney disease in the C-PROBE cohort and proteomic data from the Kidney Precision Medicine Project, identifying a concordant injury programme marked by inflammatory and fibrotic remodeling together with impaired energy metabolism. We also used ORBIT to integrate ten transcriptomic, four proteomic and one translatomic datasets in dilated cardiomyopathy heart tissue, identifying extracellular matrix remodeling and mitochondrial energy failure as the dominant concordant biological programmes. In both diseases, ORBIT prioritized concordant genes and identified pathways that were not significant in any single omics layer alone. These results show that directional rank integration can be used for integrating heterogeneous omics summary data and prioritizing robust molecular mechanisms.
]]></description>
<dc:creator><![CDATA[ Qiu, Z., Palmer, D., Jostins-Dean, L., Lewis, A. J., Bull, K., Nanchahal, J., Luo, Y. ]]></dc:creator>
<dc:date>2026-10-01</dc:date>
<dc:identifier>doi:10.64898/2026.09.25.754380</dc:identifier>
<dc:title><![CDATA[Rank-based integration identifies convergent disease mechanisms across omics]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-01</prism:publicationDate>
<prism:section></prism:section>
</item>
</rdf:RDF>
