<?xml version="1.0" encoding="UTF-8" ?>
<rdf:RDF xmlns:admin="http://webns.net/mvcb/" xmlns="http://purl.org/rss/1.0/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:prism="http://purl.org/rss/1.0/modules/prism/" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:syn="http://purl.org/rss/1.0/modules/syndication/">
<channel rdf:about="https://biorxiv.org">
<admin:errorReportsTo rdf:resource="mailto:biorxiv@cshlpress.edu"/>
<title>bioRxiv Subject Collection: Bioinformatics</title>
<link>https://biorxiv.org</link>
<description>
This feed contains articles for bioRxiv Subject Collection "Bioinformatics"
</description>

<items>
<rdf:Seq>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.06.743421v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.06.743410v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.04.742705v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.04.742550v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.04.742759v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.05.742796v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.07.743440v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.04.742799v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.04.742788v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.04.742789v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.03.742642v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.03.742627v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.03.742604v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.03.742561v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.05.743110v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.04.742537v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.06.743405v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.06.743374v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.03.742620v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.03.742388v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.07.743468v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.04.742344v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.07.743473v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.03.741541v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.03.742514v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.03.741871v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.03.742524v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.03.742471v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.03.742474v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.08.03.742406v1?rss=1"/>
</rdf:Seq>
</items>
<prism:eIssn/>
<prism:publicationName>bioRxiv</prism:publicationName>
<prism:issn/>

<image rdf:resource=""/>
</channel>
<image rdf:about="">
<title>bioRxiv</title>
<url>https://www.biorxiv.org/sites/default/files/bioRxiv_article.jpg</url>
<link>https://www.biorxiv.org</link>
</image>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.06.743421v1?rss=1">
<title>
<![CDATA[
PandaMap: A Python Package for Comprehensive Visualization of Protein-Ligand Interaction Networks 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.06.743421v1?rss=1
</link>
<description><![CDATA[
Protein-ligand interaction diagrams are a routine part of structural and medicinal chemistry, but the tools that produce them tend to force a choice: comprehensive detection with tabular output, publication-quality figures behind a licence, or a scripting environment that assumes expertise. PandaMap Protein AND ligAnd interaction MAPper is an open-source Python package that produces a 2D interaction diagram, an interactive 3D viewer, a text report, a machine-readable CSV, and a four-panel graphical summary from a single command. It reads PDB, mmCIF and PDBQT files, detects 15 interaction classes using crystallographically validated distance thresholds, and depends only on NumPy, Matplotlib, BioPython and Requests; RDKit improves the 2D ligand layout when present but is not required. Hydrogen bonds are filtered on the true D--H-A angle when the structure contains explicit hydrogens, matching PLIP's 100 criterion on the same evidence, and on distance alone otherwise, with the provenance of each measurement recorded. We benchmarked the package on three complexes chosen for different chemistry: enolase with a phosphonate transition-state analogue (PDB 1ELS), the EGFR kinase with erlotinib (1M17), and aldose reductase with IDD594 (1US0). PandaMap recovers the contacts these structures are known for, including the EGFR hinge hydrogen bond to MET769 and the IDD594 bromine THR113 halogen bond, both at distances identical to PLIP's. All detection thresholds, scoring weights and the exact commands used are given in the Supplementary Information, and the release carries a regression suite covering each interaction class. PandaMap 4.3.0 is available on PyPI under the MIT licence.
]]></description>
<dc:creator><![CDATA[ Panda, P. K. ]]></dc:creator>
<dc:date>2026-08-09</dc:date>
<dc:identifier>doi:10.64898/2026.08.06.743421</dc:identifier>
<dc:title><![CDATA[PandaMap: A Python Package for Comprehensive Visualization of Protein-Ligand Interaction Networks]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-08-09</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.06.743410v1?rss=1">
<title>
<![CDATA[
nf-cavalier: A Nextflow Pipeline for Rare Disease Variant Prioritization and Reporting 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.06.743410v1?rss=1
</link>
<description><![CDATA[
nf-cavalier is a Nextflow pipeline that automates genomic variant annotation, filtering, and reporting for individuals with rare Mendelian diseases. The pipeline takes as input variant callsets for an individual, family, or rare disease cohort, together with a target gene panel or a phenotype of interest. Variants are then filtered using various customisable criteria, including predicted gene consequence, computational pathogenicity predictions, population frequency, and familial segregation. The sequencing data for candidate variants is then visualised for human review. Candidate variant results are returned in user-friendly output formats, including interactive HTML reports and PowerPoint slide decks, with embedded links to external resources that enable rapid review by clinical research teams. nf-cavalier is maintained on GitHub (bahlolab/nf-cavalier) and licensed under the permissive MIT open-source licence.
]]></description>
<dc:creator><![CDATA[ Munro, J. E., Reid, J., Bahlo, M. E., Bennett, M. F. ]]></dc:creator>
<dc:date>2026-08-10</dc:date>
<dc:identifier>doi:10.64898/2026.08.06.743410</dc:identifier>
<dc:title><![CDATA[nf-cavalier: A Nextflow Pipeline for Rare Disease Variant Prioritization and Reporting]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-08-10</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.04.742705v1?rss=1">
<title>
<![CDATA[
FOCUS: end-to-end preprocessing, alignment and resolution-matched integration of spatial multi-omics data 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.04.742705v1?rss=1
</link>
<description><![CDATA[
Integrating spatial multi-omics data requires coordinated preprocessing, cross-modality alignment and feature registration across modalities that differ in file format, coordinate system and spatial resolution. No existing tool addresses this pipeline end-to-end from raw experimental files till aligned data object. We present FOCUS, an open-source Python package that takes raw data from spatial transcriptomics, mass spectrometry imaging, Raman spectroscopy imaging and brightfield or fluorescence microscopy through modality-specific preprocessing, interactive spatial alignment and resolution-matching registration to a unified MuData object, driven by a single configuration file. Its modular, registry-based architecture allows straightforward extension to additional modalities. FOCUS is accessible via a command-line interface, a browser-based GUI and a Python API.
]]></description>
<dc:creator><![CDATA[ Venturelli, L., Jacobs, J., Sifrim, A. ]]></dc:creator>
<dc:date>2026-08-09</dc:date>
<dc:identifier>doi:10.64898/2026.08.04.742705</dc:identifier>
<dc:title><![CDATA[FOCUS: end-to-end preprocessing, alignment and resolution-matched integration of spatial multi-omics data]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-08-09</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.04.742550v1?rss=1">
<title>
<![CDATA[
A Framework for Benchmarking Pathway Reconstruction Algorithms 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.04.742550v1?rss=1
</link>
<description><![CDATA[
Cells coordinate diverse biological processes through interactions among thousands of molecules, but mapping these interactions comprehensively and systematically remains an open problem. Pathway reconstruction algorithms address this problem by linking molecules of interest, identified from high-throughput omics experiments, using prior knowledge encoded as background interaction networks. This process recovers intermediate molecules and interactions that were not directly measured in the experimental data but plausibly connect the observed molecules. It generates testable hypotheses about interactions that drive cell behavior and informs the choice of follow-up experiments. Many algorithms have been created over decades, each optimizing different computational objectives and relying on different assumptions. The resulting heterogeneity has made benchmarking challenging, limiting systematic comparisons. Therefore, selecting an algorithm for a given biological context remains a non-trivial and poorly informed task. This registered report presents a large-scale benchmark of pathway reconstruction algorithms, evaluating 14 algorithms across 822 datasets from four biological settings. To enable this benchmark, we introduce Signaling Pathway Reconstruction Analysis Streamliner (SPRAS), which standardizes algorithm inputs, outputs, and execution into a formal framework, enabling systematic comparison that was previously infeasible. We will assess each algorithm on reconstruction performance against gold standard pathways, algorithm similarity, and computational performance across different biological contexts. Together, these evaluations will provide quantitative evidence for understanding pathway reconstruction algorithm behavior and guiding algorithm selection.
]]></description>
<dc:creator><![CDATA[ Talluri, N., Figueroa-Reid, T., Hiemstra, J., Magnano, C. S., Shedivy, A., Panda, N., Liu, Y., Sanjeev, S., Anderson, O. F., Barelvi, A., O'Brien, A., Johnson, O. T., Haddad, J. A., Halberg-Spencer, S. A., Nurbol, A., Jan, I., Degbelo, M., Nachreiner, D., Llera-Magord, C., Howland, G., Li, G. H., Ritz, A., Gitter, A. ]]></dc:creator>
<dc:date>2026-08-09</dc:date>
<dc:identifier>doi:10.64898/2026.08.04.742550</dc:identifier>
<dc:title><![CDATA[A Framework for Benchmarking Pathway Reconstruction Algorithms]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-08-09</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.04.742759v1?rss=1">
<title>
<![CDATA[
Mandatory use of mock communities highlighted by the descriptive comparison of Epi2Me 16S and EMU bioinformatic workflows for full-length 16S rRNA Nanopore sequencing. 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.04.742759v1?rss=1
</link>
<description><![CDATA[
Abstract Full-length 16S rRNA gene sequencing using Oxford Nanopore Technologies has emerged as a promising approach to improve species-level resolution in microbiota studies. However, the accuracy of taxonomic assignment remains highly dependent on the bioinformatics s used to process Nanopore long-read data. Therefore, the only way to ensure a good level of certainty in obtained results is to use positive controls in the form of mock communities in the experimental designs. In this study, we compared the performance of Epi2Me 16S (using Minimap2 or Kraken2) workflows provided by Oxford Nanopore Technologies and an EMU workflow for full-length 16S rRNA gene analysis. Using a commercial mock community sequenced across multiple Nanopore runs, taxonomic assignment accuracy and reproducibility was evaluated. Epi2Me-Kraken2 exhibited 18 % of incorrect genus-level assignments and failed to identify 3 species present in the mock community. While Epi2Me-Minimap2 achieved an excellent genus-level classification, reporting 9 % of sequences assigned to a genus not in the mock community, species-level assignments were inconsistent for several community members such as Listeria. In contrast, EMU provided accurate and consistent species-level taxonomic profiles, with all species correctly identified while keeping the number of genus absent from the mock community at 1.2%. Importance These results highlight that Epi2Me integrated workflows are not the best option for specie-level taxonomic assignation. More importantly, this paper underscores the importance of routine inclusion of positive controls for microbiota studies, in the form of mock communities, as a critical safeguard for accurate data interpretation. Without the use of a mock community, a paper published would be at risk of reporting wrong observations and inaccurate conclusions.
]]></description>
<dc:creator><![CDATA[ Shedleur-Bourguignon, F., Theriault, W. P., Thibodeau, A. ]]></dc:creator>
<dc:date>2026-08-09</dc:date>
<dc:identifier>doi:10.64898/2026.08.04.742759</dc:identifier>
<dc:title><![CDATA[Mandatory use of mock communities highlighted by the descriptive comparison of Epi2Me 16S and EMU bioinformatic workflows for full-length 16S rRNA Nanopore sequencing.]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-08-09</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.05.742796v1?rss=1">
<title>
<![CDATA[
PanGBank: a large-scale resource of precomputed microbial pangenomes built with PPanGGOLiN 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.05.742796v1?rss=1
</link>
<description><![CDATA[
PanGBank (https://pangbank.genoscope.cns.fr) is a comprehensive open-access database providing precomputed prokaryotic pangenomes at a broad taxonomic scale. Built upon PPanGGOLiN partitioned pangenome graphs, PanGBank addresses the growing need for large-scale comparative genomics through a standardized, regularly updated, and fully accessible resource. The initial release comprises two complementary collections covering more than 4,600 prokaryotic species from the Genome Taxonomy Database (GTDB), encompassing over 393,000 genomes: GTDB_all, maximizing taxonomic and environmental diversity through the inclusion of MAGs and SAGs, and GTDB_refseq, focusing on high-quality, annotation-rich genomes. Each species-level pangenome integrates graph-based statistical partitions into persistent, shell, and cloud gene families, together with regions of genomic plasticity (panRGP) and co-localized functional modules (panModule). PanGBank offers multiple access modes, including a REST API, a command-line interface (PanGBank-cli), and an interactive web interface. By combining large-scale pangenome resources with advanced graph-based analyses, PanGBank provides a scalable framework for exploring microbial diversity, genome evolution, functional variation, and the dissemination of adaptive traits across prokaryotic populations, as illustrated by a use case on Acinetobacter baumannii pangenome investigating the distribution and evolution of antimicrobial resistance determinants.
]]></description>
<dc:creator><![CDATA[ Mainguy, J., Lemane, T., Bazin, A., Arnoux, J., Gautreau, G., Medigue, C., Calteau, A., Vallenet, D. ]]></dc:creator>
<dc:date>2026-08-09</dc:date>
<dc:identifier>doi:10.64898/2026.08.05.742796</dc:identifier>
<dc:title><![CDATA[PanGBank: a large-scale resource of precomputed microbial pangenomes built with PPanGGOLiN]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-08-09</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.07.743440v1?rss=1">
<title>
<![CDATA[
A map of human protein-protein interaction embeddings for functional discovery 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.07.743440v1?rss=1
</link>
<description><![CDATA[
A protein's function depends not just on its own structure and localization, but also on the interactions with its partners. Many proteins are therefore better described by a set of partner-dependent roles than by a single annotation. Yet most approaches to the functional interpretation of protein-protein interactions (PPIs) remain protein or set-centric. They rely on pre-existing annotations, and perform worst where knowledge is sparse. Here, we present MAPPIE (Map of Protein-Protein Interaction Embeddings), a method that treats each PPI, rather than each protein, as a unit of representation. From 199,137 human interactions spanning 15,503 proteins, we build a two-dimensional map of the human PPI landscape for functional discovery. Protein language model embeddings for two protein interaction partners are combined and compressed into a latent space, with model selection guided by domain-domain interactions used as a structural proxy for interaction similarity. The resulting geometry separates domain defined interaction classes, organizes disorder associated interactions spatially, and splits interactions involving the same protein by partner. A query PPI's latent neighbourhood recovers its own annotated functions across molecular, complex, pathway, and biological processes. MAPPIE contributes most where existing functional evidence is weakest, outperforming interactome and sequence identity baselines for sparsely connected interactions. MAPPIE neighbours of query PPIs are enriched for partners in independent protein networks, recovering curated complex-level function even when subunits are spread across the map. Applied to a human dark interactome, MAPPIE assigns specific, experimentally supported functions to dark hub proteins.
]]></description>
<dc:creator><![CDATA[ Cihan, M., Distler, U., Andrade-Navarro, M. A. ]]></dc:creator>
<dc:date>2026-08-09</dc:date>
<dc:identifier>doi:10.64898/2026.08.07.743440</dc:identifier>
<dc:title><![CDATA[A map of human protein-protein interaction embeddings for functional discovery]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-08-09</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.04.742799v1?rss=1">
<title>
<![CDATA[
A distribution-aware and functionally relevant novel framework for generation and discovery of bioactive peptides 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.04.742799v1?rss=1
</link>
<description><![CDATA[
Recent advances in artificial intelligence have accelerated the discovery of bioactive peptides by enabling computational exploration of the vast peptide sequence space. However, existing peptide generation approaches generally rely on either distribution-learning models, which generate biologically realistic sequences but do not consistently optimize functional activity, or optimization based methods, which maximize prediction confidence while often deviating from the underlying distribution of experimentally validated peptides. To address this limitation, a two-phase generative evolutionary framework is proposed that integrates distribution learning with evolutionary optimization. In the first phase, Variational Autoencoders (VAE), Autoregressive Transformers (ART), and Token Diffusion Transformers (TDT) are used to generate biologically plausible seed peptides. In the second phase, these peptides were used as initial seed for Hill Climbing optimization procedure that iteratively improves fitness function score. The proposed two-phase framework was evaluated using a dataset of experimentally validated IL2 inducing peptides. Evaluation using independent IL2 prediction models showed that Autoregressive Transformer combined with Hill Climbing achieved the best overall performance, achieving the mean IL2 induction confidence score of 0.96 while reducing KL divergence from 2.26 for standalone Hill Climbing to 0.75. A case study on an independent IL13 inducing peptide dataset showed similar trends, with ART initialized Hill Climbing achieving the mean IL13 induction score of 0.99 while reducing KL divergence from 1.76 to 0.59. Overall, the framework provides a generalizable approach for balancing functional optimization and distributional realism and can be applied to peptide discovery and data augmentation in imbalanced biological datasets thereby generating high confidence peptides for wet lab validation.
]]></description>
<dc:creator><![CDATA[ Abhigyan, R., Sood, V., Arora, P., Kaur, B. ]]></dc:creator>
<dc:date>2026-08-09</dc:date>
<dc:identifier>doi:10.64898/2026.08.04.742799</dc:identifier>
<dc:title><![CDATA[A distribution-aware and functionally relevant novel framework for generation and discovery of bioactive peptides]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-08-09</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.04.742788v1?rss=1">
<title>
<![CDATA[
A generative model for dimensionality reduction with millions of features and few samples 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.04.742788v1?rss=1
</link>
<description><![CDATA[
In this paper, we demonstrate that it is feasible to train a deep generative model for dimensionality reduction with millions of features using relatively few samples, making this type of generative model a more versatile alternative to standard dimensionality reduction methods. Specifically, we hypothesize that, for a decoder-only model, the number of training samples required is almost independent of feature dimensionality across most network architectures. Through an extensive set of experiments on synthetic nonlinear data, we validate this hypothesis. We further train the model on a downsampled version of the 1000 Genomes Project (1KGP) dataset to assess its behavior under controlled reductions in sample size. Finally, we train a Deep Generative Decoder (DGD) on a curated dataset from the International Cancer Genome Consortium (ICGC) containing 4.4 million features, using approximately 4,000 training samples and 1,000 test samples. The resulting latent representation exhibits clear clustering and, when compared at the same latent dimensionality, outperforms PCA and variational autoencoders (VAEs) for tumor type classification. In addition, the DGD is computationally efficient and can be trained on a single 16 GB GPU. The implementation is available at https://github.com/cpancott/ReceptiveDGD.
]]></description>
<dc:creator><![CDATA[ Pancotti, C., Fariselli, P., Meisner, J., Krogh, A. ]]></dc:creator>
<dc:date>2026-08-09</dc:date>
<dc:identifier>doi:10.64898/2026.08.04.742788</dc:identifier>
<dc:title><![CDATA[A generative model for dimensionality reduction with millions of features and few samples]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-08-09</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.04.742789v1?rss=1">
<title>
<![CDATA[
Scop3P-Toolkit: executable structure-aware workflows linking PTMs, peptides, and mutations to protein function 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.04.742789v1?rss=1
</link>
<description><![CDATA[
Post-translational modifications (PTMs) and genetic variants regulate protein function, signalling, and disease, but their interpretation requires integration of sequence annotations with structural, interaction, and biophysical context. Although resources such as Scop3P, UniProt, the Protein Data Bank, and AlphaFold provide extensive annotations and structural information, integrating these data into reproducible structure-aware analyses still requires custom scripting and manual coordination between multiple independent tools. To address this challenge, we developed Scop3P-Toolkit, an open-source executable analytical environment for interactive analysis of PTMs, mutations, and proteomics-derived peptides in their structural context. The toolkit integrates protein annotation retrieval with structural mapping, residue interaction network analysis, comparative structural analysis, and residue-level biophysical profiling within a unified framework. Experimentally supported phosphosites, phosphopeptides, and phosphoproteomics evidence are provided for human proteins through Scop3P, with optional integration of curated UniProt PTM annotations. UniProt-derived PTMs, sequence features, and genetic variants are available for proteins from any species, extending the framework beyond the human phosphoproteome. Scop3P-Toolkit supports structure-centric analyses including interpretation of PTMs and disease-associated variants, analysis of residue interaction networks and their rewiring across alternative conformations, structural localisation of peptides, and exploration of protein-protein, protein-ligand, and host-pathogen interfaces. Interactive visualisation links sequence annotations, three-dimensional structures, residue interaction networks, and biophysical profiles, enabling coordinated exploration across multiple molecular representations. The toolkit is distributed as Jupyter notebooks, browser-based Voila applications, and a Galaxy interactive tool, providing transparent, accessible, and reproducible workflows for both computational and experimental researchers. By integrating biological annotation resources into executable, structure-aware workflows, Scop3P-Toolkit enables reproducible interpretation of PTMs, mutations, and proteomics data.
]]></description>
<dc:creator><![CDATA[ Diaz, A., Tichshenko, N., Depoortere, B. G. J., Andrade Buono, R., De Geest, P., Vranken, W. F., Martens, L., Ramasamy, P. ]]></dc:creator>
<dc:date>2026-08-09</dc:date>
<dc:identifier>doi:10.64898/2026.08.04.742789</dc:identifier>
<dc:title><![CDATA[Scop3P-Toolkit: executable structure-aware workflows linking PTMs, peptides, and mutations to protein function]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-08-09</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.03.742642v1?rss=1">
<title>
<![CDATA[
Elucidating biosynthetic pathways related to the synthesis of small halogenated peptidic natural products in marine sponge microbiomes 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.03.742642v1?rss=1
</link>
<description><![CDATA[
Marine sponges are known sources of bioactive natural products (NPs), many of which are produced by associated bacterial symbionts via encoded biosynthetic gene clusters (BGCs). A particularly interesting subclass of sponge-derived NPs is comprised of small, brominated alkaloids, which are recovered from diverse habitats and host sponge taxonomies. Despite having been described decades ago, most of these NPs do not have an elucidated biosynthetic origin. We queried metagenomes of several sponge species by making use of a minimal set of core enzymes that we postulate to be necessary to produce these small peptidic NPs: an FADH2-dependent halogenase and an AMP-binding adenylation enzyme. This revealed a variety of novel BGC architectures, many of which showed conservation among sponge host phylogenies and were encoded in the genomes of diverse sponge-associated bacteria. Furthermore, we identified a BGC in the sponge G. barretti that is potentially linked to the production of the iconic barettins, given its enzymatic machinery and specific acidobacterial origin. The present work contributes to the challenging quest to link orphan brominated NPs to their parent BGCs in the sponge holobiont and beyond.
]]></description>
<dc:creator><![CDATA[ Loureiro, C., Schorn, M. A., Alanjary, M., Kuipers, B., Louwen, J. J. R., van der Oost, J., Medema, M. H., Sipkema, D. ]]></dc:creator>
<dc:date>2026-08-09</dc:date>
<dc:identifier>doi:10.64898/2026.08.03.742642</dc:identifier>
<dc:title><![CDATA[Elucidating biosynthetic pathways related to the synthesis of small halogenated peptidic natural products in marine sponge microbiomes]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-08-09</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.03.742627v1?rss=1">
<title>
<![CDATA[
A Pseudo-Longitudinal Methylome Projection Framework Defines a Buccal PACE-like Aging-Rate Score from Cross-Sectional DNA Methylation Data 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.03.742627v1?rss=1
</link>
<description><![CDATA[
Background: DNA methylation-based biomarkers have enabled robust estimation of biological age across tissues, and longitudinally trained measures such as DunedinPACE provide estimates of the pace of aging from blood methylomes. However, longitudinal methylation data are often unavailable, particularly for minimally invasive tissues such as buccal mucosa. Here, we developed a pseudo-longitudinal framework to estimate a buccal mucosa-derived PACE-like aging-rate score from cross-sectional methylome data. Methods: We used a buccal biological age estimator as an internal pseudo-time axis. Methylation beta-values were transformed to M-values, and CpG-specific smooth functions of biological age were fitted in cross-validation. Local derivatives of these functions were used to project each individuals buccal methylome forward by a small time step. The projected methylome was converted back to beta-values, biological age was recalculated, and the change in biological age per unit time was defined as a pseudo-aging velocity. This raw velocity was transformed to a non-negative PACE-like score centered at 1.0. We then trained cross-fitted models to predict the derived score from buccal CpG methylation profiles. Results: In 151 individuals, the proposed score was reproducibly predicted from buccal methylomes in out-of-fold analysis, with a Pearson correlation of 0.706 and Spearman correlation of 0.710 between observed and predicted PACE-like scores. Sensitivity analyses across CpG selection size and regression models showed broadly consistent performance. In contrast, the proposed buccal PACE-like score showed only modest association with measured DunedinPACE, and alternative attempts to reconstruct DunedinPACE from buccal methylomes, including supervised proxy modeling and buccal-to-blood CpG imputation, showed limited sample-level performance. Conclusions: These results support the feasibility of deriving a tissue-specific PACE-like aging-rate score from cross-sectional buccal methylome data by treating biological age as a pseudo-time axis. The proposed score should not be interpreted as a replacement for blood-derived DunedinPACE, but rather as an exploratory buccal methylome dynamics index that may capture tissue-specific aging-related variation.
]]></description>
<dc:creator><![CDATA[ Shoji, T., Nakaki, R. ]]></dc:creator>
<dc:date>2026-08-09</dc:date>
<dc:identifier>doi:10.64898/2026.08.03.742627</dc:identifier>
<dc:title><![CDATA[A Pseudo-Longitudinal Methylome Projection Framework Defines a Buccal PACE-like Aging-Rate Score from Cross-Sectional DNA Methylation Data]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-08-09</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.03.742604v1?rss=1">
<title>
<![CDATA[
pbcftools: parallel execution of bcftools for large variant call sets 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.03.742604v1?rss=1
</link>
<description><![CDATA[
Summary bcftools is the standard toolkit for handling VCF and BCF variant files, but it processes records on a single core; its --threads option speeds up only compression of the output, not the work done on variant records. Processing large call sets is therefore slow, and users often divide the genome and reassemble the results by hand. We present pbcftools, a Perl wrapper that does this automatically: it splits the genome into chunks, runs an ordinary bcftools command on each in parallel, and reassembles the outputs by a method suited to the data type. Across Linux servers, Windows/WSL2 workstations and Apple laptops, with bcftools 1.21 to 1.24, parallel output was identical to serial output for every command tested. On 1000 Genomes Phase 3 data, operations writing compressed VCF ran 10.8 to 21.1 times faster with 32 cores and up to 35.4 times with 64, those writing text 3.7 to 12.8 times, and merging 100 VCF files 19.2 times. pbcftools also runs on LSF and Slurm clusters. Availability and implementation pbcftools is written in Perl (>= 5.16) and requires bcftools; local parallel execution also requires Perl module Parallel::ForkManager. It is released under the MIT license at https://github.com/zhangge-uc/pbcftools (DOI: 10.5281/zenodo.21780361).
]]></description>
<dc:creator><![CDATA[ Zhang, G. ]]></dc:creator>
<dc:date>2026-08-09</dc:date>
<dc:identifier>doi:10.64898/2026.08.03.742604</dc:identifier>
<dc:title><![CDATA[pbcftools: parallel execution of bcftools for large variant call sets]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-08-09</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.03.742561v1?rss=1">
<title>
<![CDATA[
Assessing Computational Models for Pharmacogenomic Variant Interpretation 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.03.742561v1?rss=1
</link>
<description><![CDATA[
Accurately predicting the effects of pharmacogenomic variants is essential for the development of personalized therapeutic strategies, as genetic variability can influence drug response differently across patients. Here, we assessed several computational approaches using a dataset of pharmacogenomic variants with either clinical annotations or functional characterization by deep mutational scanning, compiled from the literature, with an additional focus on CYP2C9, a clinically relevant drug-metabolizing enzyme. Our results show that, despite recent methodological advances, substantial room for improvement remains. In particular, current methods struggle to distinguish gain-of-function variants associated with increased drug clearance and fast-metabolizer phenotypes from neutral variants, whereas loss-of-function variants that reduce drug clearance are predicted more accurately. The integration of structural and evolutionary information appears to be a key strategy for improving performance, with the coevolution-based StructureDCA method achieving the highest accuracy compared with classical genetic variant-effect predictors and recent deep learning approaches, including the pathogenic-variant predictor AlphaMissense and general protein language model-based methods. Finally, our results indicate that computational models can complement in vitro experiments in clinical variant interpretation, as StructureDCA predictions showed better agreement with clinically annotated phenotypes than large-scale deep mutational scanning data in several cases.
]]></description>
<dc:creator><![CDATA[ Pucci, F., Hermans, P., Tsishyn, M., Cusato, J., Rooman, M. ]]></dc:creator>
<dc:date>2026-08-09</dc:date>
<dc:identifier>doi:10.64898/2026.08.03.742561</dc:identifier>
<dc:title><![CDATA[Assessing Computational Models for Pharmacogenomic Variant Interpretation]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-08-09</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.05.743110v1?rss=1">
<title>
<![CDATA[
Move BeTween modAlities (MBTA) employs flow matching to predict single cell data modalities 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.05.743110v1?rss=1
</link>
<description><![CDATA[
Integrating diverse molecular modalities to obtain a comprehensive view of cellular identity remains a major challenge in single-cell biology. A fundamental but underappreciated obstacle is structural mismatch-the phenomenon in which the neighborhood structure of a cell differs depending on which molecular modality is used to define it. Existing approaches typically embed modalities into a shared latent space, which actively erases the structural differences between modalities that make multimodal measurements scientifically valuable. Here we introduce Move BeTween modAlities (MBTA), the first framework explicitly designed to address structural mismatch. Rather than forcing modalities into a shared representation, MBTA maintains modality-specific latent spaces and connects them via flow matching, preserving the structural integrity of each modality while enabling accurate cross-modal translation. Across extensive benchmarks on multi-modal single-cell datasets, MBTA consistently outperformed existing methods, with the largest gains observed in datasets with pronounced structural mismatch. Applied to joint genomic and transcriptomic profiles of breast cancer patients, MBTA identified transcriptomic lineage relationships corroborated by genomic variation and outperformed state-of-the-art transcriptomics-based copy number inference methods. Extending this framework to mouse embryonic development, we reconstructed temporal trajectories jointly defined by gene expression and seven complementary epigenetic modalities. MBTA can connect any number of molecular readouts without erasing their individual character, serving as the computational foundation for assembling multi-layered portraits of cells.
]]></description>
<dc:creator><![CDATA[ Xu, B., Zhang, Y., Michor, F. ]]></dc:creator>
<dc:date>2026-08-09</dc:date>
<dc:identifier>doi:10.64898/2026.08.05.743110</dc:identifier>
<dc:title><![CDATA[Move BeTween modAlities (MBTA) employs flow matching to predict single cell data modalities]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-08-09</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.04.742537v1?rss=1">
<title>
<![CDATA[
pysigscore: gene signatures scoring across bulk and single-cell transcriptomics 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.04.742537v1?rss=1
</link>
<description><![CDATA[
High-throughput transcriptomics has made gene signatures central to interpreting gene expression data, with applications in diagnosis, prognosis, and prediction. Quantifying signature activity and assessing its robustness remain challenging because scoring methods primarily rely on various assumptions, and no single approach is universally optimal. Here, we present pysigscore, a Python framework for gene set scoring in bulk and single-cell RNA-seq data. pysigscore integrates 18 built-in scoring methods with a fully customisable scorer, allowing users to define and benchmark new scoring functions. It also provides reliability analyses, including p-value estimation and leave-one-out experiments, to assess the significance of scores and gene-level contributions. We validated pysigscore on the CCLE, TCGA, and PBMC datasets, recovering the expected enrichment in liver, hypoxia, inflammatory, and cell-cycle signatures.
]]></description>
<dc:creator><![CDATA[ Giacomello, T., Mazzara, S., Abbruzzese, G., Barberis, A., tangherloni, a., Buffa, F. M. ]]></dc:creator>
<dc:date>2026-08-09</dc:date>
<dc:identifier>doi:10.64898/2026.08.04.742537</dc:identifier>
<dc:title><![CDATA[pysigscore: gene signatures scoring across bulk and single-cell transcriptomics]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-08-09</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.06.743405v1?rss=1">
<title>
<![CDATA[
Scalable Extraction of Information on Protein-Protein Interactions using Topological Data Analysis 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.06.743405v1?rss=1
</link>
<description><![CDATA[
Protein-protein interactions (PPIs) govern a wide range of cellular functions. The ability to predict PPI interfaces from protein molecular surfaces is important for understanding protein function and enabling therapeutic discovery. While recent advances in structure-based learning, particularly molecular-surface geometric deep learning frameworks, have demonstrated that protein surfaces encode rich geometric and physicochemical information, such approaches often remain computationally intensive and data-hungry. Alternatively, topological data analysis (TDA) has emerged as a mathematically rigorous framework for extracting robust, multiscale shape information from complex data. In this work, we introduce a scalable TDA framework for extracting information on PPIs directly from localized protein surface patches. Our approach leverages multiscale topological descriptors, evaluated from patch-wise point cloud representations of protein mesh surfaces, combined with supervised machine learning models for interface prediction. On a full dataset of 3,362 proteins, the proposed approach substantially reduced computational cost relative to an established geometric deep learning method, MaSIF-site, decreasing preprocessing time from approximately 27 s/protein to 5-8 s/protein and total training time from approximately 6 h to 1-1.3 h. Importantly, this computational reduction is achieved while maintaining mean test area under the receiver operating characteristic curve (AUC) values of 0.76 and 0.77 for patch radii of 0.9 nm and 1.2 nm, respectively, thus approaching the MaSIF-site test AUC of 0.84. Our results suggest that topology offers a scalable and computationally efficient approach for high-throughput extraction of information from complex biomolecular interfaces.
]]></description>
<dc:creator><![CDATA[ Mukherjee, A., Park, B., Malmstrom, A., Cisewski-Kehe, J., Van Lehn, R. C., Zavala, V. M. ]]></dc:creator>
<dc:date>2026-08-09</dc:date>
<dc:identifier>doi:10.64898/2026.08.06.743405</dc:identifier>
<dc:title><![CDATA[Scalable Extraction of Information on Protein-Protein Interactions using Topological Data Analysis]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-08-09</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.06.743374v1?rss=1">
<title>
<![CDATA[
Inferring disruption of directed graphs using LIKA reveals altered protein phosphorylation networks in schizophrenia 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.06.743374v1?rss=1
</link>
<description><![CDATA[
Motivation: Kinases regulate a multitude of protein functions, and their dysregulation is pivotal for many human diseases. Direct measurement of kinase activity, however, is often challenging; therefore, inferring activity from the behavior of their substrates is a widely adopted strategy. Nonetheless, traditional methods typically oversimplify the underlying network, ignoring that any particular substrate can be phosphorylated by multiple kinases. Results: We present LIKA, a likelihood-based framework for inferring kinase activity from phosphoproteomic data. By modeling the many-to-many structure of kinase-substrate interactions, LIKA achieves high efficiency, even with limited data, while capturing network complexity. Simulation and cell line analyses confirm the robustness and accuracy of LIKA. Importantly, analysis of a phosphoproteomic dataset from schizophrenia and control subjects reveals novel dysregulated kinases. Availability and Implementation: The implementation code and publicly available data are provided at: https://github.com/lujingz/LIKA.
]]></description>
<dc:creator><![CDATA[ Zhang, L., Demarco, A. G., Ghafari, K., Devlin, B., MacDonald, M. L., Roeder, K. ]]></dc:creator>
<dc:date>2026-08-07</dc:date>
<dc:identifier>doi:10.64898/2026.08.06.743374</dc:identifier>
<dc:title><![CDATA[Inferring disruption of directed graphs using LIKA reveals altered protein phosphorylation networks in schizophrenia]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-08-07</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.03.742620v1?rss=1">
<title>
<![CDATA[
De novo transformer modeling improves recovery of genetic cell types from sparse single-cell RNA sequencing 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.03.742620v1?rss=1
</link>
<description><![CDATA[
Single-cell RNA sequencing (scRNA-seq) simultaneously provides gene-expression profiles and genetic variants from individual cells, creating an opportunity to relate cellular phenotypes to their somatic evolutionary histories. However, delineation of genetic type (GTs) from scRNA-seq remains difficult because most variant positions are unobserved in individual cells and the observed base calls contain substantial false-positive and false-negative errors. We evaluated some existing phylogenetic and imputation methods using one simulated dataset and two empirical tumor datasets. We found that extreme sparsity prevented reliable recovery of known or independently inferred GTs when multiple GTs were present. This led us to adapt the STICI transformer architecture to train a separate model de novo on each sparse cell-variant (CV) matrix. These data-specific models predicted millions of missing bases, greatly reducing matrix sparsity. Phylogenetic analyses of the imputed CV matrices showed substantially improved recovery of GTs in both simulated and empirical datasets. In the empirical dataset, transformer-based analysis also suggested finer-scale genetic structure within some previously reported GTs that was not apparent with the existing methods. These results demonstrate that highly sparse scRNA-seq datasets contain substantially more recoverable lineage information than previously appreciated and that de novo transformer modeling provides an effective approach for recovering much of this hidden information. Nevertheless, sequencing errors persisted, limiting reconstruction of cellular lineage structure and leaving significant room for methodological improvement before expression phenotypes can be examined reliably in the context of their cellular evolutionary relationships.
]]></description>
<dc:creator><![CDATA[ Craig, J. M., Mowlaei, M. E., Miura, S., Shi, X., Kumar, S. ]]></dc:creator>
<dc:date>2026-08-07</dc:date>
<dc:identifier>doi:10.64898/2026.08.03.742620</dc:identifier>
<dc:title><![CDATA[De novo transformer modeling improves recovery of genetic cell types from sparse single-cell RNA sequencing]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-08-07</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.03.742388v1?rss=1">
<title>
<![CDATA[
Distribution-Constrained Optimization for Reliable ML-Guided 5'UTR Sequence Design 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.03.742388v1?rss=1
</link>
<description><![CDATA[
The 5' untranslated region (5' UTR) shapes translation initiation, so its design is central to mRNA therapeutics and to improving protein-production cell lines. Deep-learning models that predict translation efficiency, measured as mean ribosome load (MRL), from the 5' UTR sequence have been combined with genetic algorithms (GAs) for sequence optimization. However, optimizing against a model trained on offline data risks reward hacking that exploits the models estimation error outside the training distribution, yielding sequences that score highly in prediction yet fail to perform in the wet lab. Yet for 5' UTR design, few studies have systematically examined which region should be treated as untrustworthy (the definition of out-of-distribution, OOD) or which constraints keep the search away from it.

We present a constrained optimization that keeps candidates within a trust region where the predictors validated accuracy holds; here "reliable" denotes keeping candidates within the training distribution over which prediction has been validated, not a guarantee of measured performance. As the OOD score, we compare the k-nearest-neighbor (KNN) distance in the predictors embedding space against a pseudo-perplexity (PPPL) from the encoder and LM head, and show that for nucleotide sequences--whose vocabulary is small--PPPL fails to separate in- vs out-of-distribution, whereas the KNN distance is an effective OOD score that can define a trust region even from unlabeled native UTR sequences. Using the KNN distance as a hard GA constraint keeps all candidates inside the trust region while maintaining predicted MRL: under unconstrained optimization most final-generation candidates (72-96% across seeds) left the trust region (self-KNN p95), whereas the hard constraint holds predicted MRL at the unconstrained level and yields about 4.3x more selectable low-risk candidates than post-hoc filtering of the unconstrained output. Comparing an output extrapolation guard, reference-sequence similarity and structural accessibility (RNAplfold), we find that the guard and the similarity constraint also suppress OOD as a side effect, whereas making accessibility a secondary objective broadens the search without suppressing OOD.
]]></description>
<dc:creator><![CDATA[ Yamaguchi, R., Mori, C., Inoue, S. ]]></dc:creator>
<dc:date>2026-08-07</dc:date>
<dc:identifier>doi:10.64898/2026.08.03.742388</dc:identifier>
<dc:title><![CDATA[Distribution-Constrained Optimization for Reliable ML-Guided 5'UTR Sequence Design]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-08-07</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.07.743468v1?rss=1">
<title>
<![CDATA[
DigiAra Computationally Designs Plant Mutants for Resistance to Microbial Infection in Arabidopsis 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.07.743468v1?rss=1
</link>
<description><![CDATA[
Plant breeding is a resource-intensive process that requires repeated cultivation and selection across multiple generations to develop varieties with desirable traits, yet computational tools capable of supporting this process remain limited. Here, we present DigiAra, an AI-based framework for designing Arabidopsis thaliana mutants with targeted traits, particularly enhanced microbial resistance. DigiAra implements an S3 pipeline--simulation, scoring, and screening: it simulates the transcriptional effects of genetic perturbations and microbial infections, scores the predicted responses in terms of relevant traits through biological pathway analysis, and screens candidate perturbations at multiple levels. In doing so, DigiAra enables the computational exploration of the genome-wide effects of genetic perturbations and diverse microbial infections in Arabidopsis. To develop DigiAra, we address two fundamental challenges. Methodologically, we introduce a hybrid architecture that integrates local gene-level interaction modeling with global transcriptional-state modeling to predict perturbation-induced changes in the Arabidopsis transcriptional state. From a data perspective, we establish a standardized pipeline for curating, harmonizing, and processing an integrated Arabidopsis-microbe transcriptional dataset comprising 495 samples from 26 projects. As a result, DigiAra accurately predicts gene-expression changes induced by unobserved genetic perturbations and microbial infections, achieving a Pearson correlation of 0.49. Moreover, it recapitulates the general non-self response (GNSR), a 24-gene program reflecting broad transcriptional reprogramming across bacterial perturbations. In an independent study, the predicted pattern-triggered immunity pathway scores further correlate with bacterial load, with a Pearson correlation of 0.57. Lastly, we deploy DigiAra to identify 27 gene knockouts through genome-wide screening that are predicted to enhance resistance to Pseudomonas syringae pv. tomato DC3000 (Pst DC3000) while limiting growth compromise, 9 of which are supported by published studies. Together, these results establish DigiAra as an effective framework for the computational design of Arabidopsis mutants. We have made our implementation openly available at https://github.com/youlab2025/DigiAra.
]]></description>
<dc:creator><![CDATA[ Bai, T., Cui, S., You, Y. ]]></dc:creator>
<dc:date>2026-08-07</dc:date>
<dc:identifier>doi:10.64898/2026.08.07.743468</dc:identifier>
<dc:title><![CDATA[DigiAra Computationally Designs Plant Mutants for Resistance to Microbial Infection in Arabidopsis]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-08-07</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.04.742344v1?rss=1">
<title>
<![CDATA[
A hierarchical orthology framework reveals viral carbohydrate-active genes across the global virosphere 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.04.742344v1?rss=1
</link>
<description><![CDATA[
Carbohydrate-active enzymes (CAZymes) shape virus-host interactions by modifying virion structures, host surfaces and extracellular glycans. However, the diversity and evolutionary origins of viral carbohydrate-active enzymes remain poorly understood, partly due to limited viral protein annotations. To address this, we present VirGenes, a database of viral orthologous groups constructed from the KEGG viral gene dataset. VirGenes uses a hierarchical framework that integrates sequence similarity, remote homology, and structural similarity to support evolutionary and functional analyses of viral proteins. By screening the sequence space of VirGenes, we identified 558 CAZyme-associated gene clusters spanning 102 CAZyme families, revealing particularly enriched repertoires in dsDNA viral lineages. Two bacteriophage families, Kleczkowskaviridae and Pootjesviridae, encoded more than 10 CAZymes per genome, followed by Mimiviridae, a representative family of eukaryotic giant viruses. Phylogenetic analyses systematically revealed divergent evolutionary histories of viral carbohydrate-active genes, including frequent horizontal transfer of endolysin genes from bacteria, which likely represents a viral strategy in the ongoing evolutionary arms race with their cellular hosts. Within the structural space of VirGenes, a large number of viral genes were found to contain CAZyme-like folds despite more than 85% of them lacking detectable sequence similarity to annotated CAZyme sequences. Notably, numerous hypothetical sequences from giant viruses exhibited glycoside hydrolase-like five-bladed {beta}-propeller folds. Overall, by integrating sequence, structural and functional evidence, we show that viral carbohydrate-active systems exemplify how distributed innovations, constrained by ancient folds, collectively build the functional complexity of the global virosphere. VirGenes is publicly accessible at https://www.genome.jp/vogdb/.
]]></description>
<dc:creator><![CDATA[ Meng, L., Zhang, R., De Castro, C., Uchiyama, I., Kanehisa, M., Ogata, H. ]]></dc:creator>
<dc:date>2026-08-07</dc:date>
<dc:identifier>doi:10.64898/2026.08.04.742344</dc:identifier>
<dc:title><![CDATA[A hierarchical orthology framework reveals viral carbohydrate-active genes across the global virosphere]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-08-07</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.07.743473v1?rss=1">
<title>
<![CDATA[
Phylogeny-aided detection of contamination in nearly 5 million SARS-CoV-2 genomes 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.07.743473v1?rss=1
</link>
<description><![CDATA[
Contamination can occur during genome sequencing when a contaminant genome is accidentally mixed with the intended genome to be sequenced. Contamination can lead to incorrect consensus genome calling, disrupting analyses of pathogen evolution and transmission. To investigate the extent of this issue, we developed PhyCD, a phylogeny-aided computational approach to investigate contamination in SARS-CoV-2 genome sequencing data. PhyCD masks consensus genome positions associated with suspicious sequencing read coverage drops, then leverages pandemic-scale phylogenetic placement techniques to identify putative contamination events. Applying PhyCD to nearly 5 million SARS-CoV-2 genomes, we identified 10,942 putative contamination events under conservative parameters. Across the flagged genomes, PhyCD flagged -- and so permits masking of -- a total of 64,753 substitutions that could cause errors in downstream genome data analyses.
]]></description>
<dc:creator><![CDATA[ Anoufa, O., Ly-Trong, N., Goldman, N., De Maio, N. ]]></dc:creator>
<dc:date>2026-08-07</dc:date>
<dc:identifier>doi:10.64898/2026.08.07.743473</dc:identifier>
<dc:title><![CDATA[Phylogeny-aided detection of contamination in nearly 5 million SARS-CoV-2 genomes]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-08-07</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.03.741541v1?rss=1">
<title>
<![CDATA[
HANSEN: An Integrated Structural and Functional Proteome Resource for Structure-Guided Drug Discovery in Mycobacterium leprae 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.03.741541v1?rss=1
</link>
<description><![CDATA[
Structural biology has advanced antimicrobial discovery by enabling drug-target identification and validation and supporting structure-guided inhibitor design. However, Mycobacterium leprae (M. leprae), the obligate intracellular bacillus that causes leprosy (Hansens disease), remains structurally under-characterised. Only 10 Protein Data Bank (PDB) entries represent seven unique proteins within a proteome encoded by 1,603 protein-coding genes. To address this gap, we present HANSEN (https://hansen-leprosy.medschl.cam.ac.uk/home), an integrated structural and functional resource containing computationally predicted three-dimensional models across the M. leprae proteome. Monomeric and oligomeric models were generated using complementary structure-prediction methods, including AlphaFold 3, Boltz-1, Chai-1, and Boltz-2. Models were annotated with predicted Local Distance Difference Test (pLDDT) scores and predicted aligned error (PAE) values. Ligand-binding pockets were predicted using AF2BIND, P2Rank, and fpocket, and ligands from the best-matching PDB templates were modelled within oligomeric complexes. Residue-level B-cell epitope propensity was estimated using DiscoTope-3.0, and ProteomeLM-derived essentiality scores were calculated for each protein. These features were integrated into a relational web database with interactive visualisation through Mol*. We also ranked all 1,603 proteins using a Target Priority Score ranging from 0 to 100. The score combines ProteomeLM-derived essentiality with pocket and AF2BIND predictions, functional annotations, and Boltz-2-associated measures of model quality and tractability. The essentiality model used a logistic-regression head trained on Mycobacterium tuberculosis (M. tuberculosis) Tn-seq labels. It achieved an AUROC of 0.84 in homology-grouped M. tuberculosis cross-validation and a transfer AUROC of 0.78 against the orthologue-aligned M. leprae reference set. Proteins were assigned to four tiers, ranging from high priority to exploratory candidates. Together, HANSEN provides a practical resource for generating and prioritising experimentally testable hypotheses for M. leprae target discovery and structure-guided drug development.

TeaserPredicted structures and druggability annotations for the whole M. leprae proteome in an open resource.

Key pointsO_LIHANSEN provides a proteome-wide structural resource for M. leprae, integrating predicted monomeric and oligomeric models with confidence metrics and functional annotations.
C_LIO_LIHANSEN integrates UniProt annotations, ligand and cofactor associations, structure-based predictions of small-molecule binding pockets, B-cell epitope propensity, gene-essentiality estimates and multi-parameter target prioritisation within a single, protein-centric interface for the proteome of M. leprae.
C_LIO_LIThe resource further incorporates a dedicated analytical module for Oxford Nanopore MinION amplicon-sequencing data, enabling the identification of mutations within drug-resistance-determining regions that confer antimicrobial resistance (AMR) in M. leprae.
C_LIO_LIBenchmarking against available experimental structures, together with cross-method concordance analyses, supports the use of pLDDT, PAE and agreement between prediction methods as complementary indicators of model reliability.
C_LIO_LIIntegrated target prioritisation produced a ranked set of candidate proteins, including established mycobacterial drug targets, to support experimental hypothesis generation for leprosy drug discovery.
C_LI
]]></description>
<dc:creator><![CDATA[ Vedithi, S. C., Rees, R., Malhotra, S., Munir, A., Matusevicius, M., Alsulami, A. F., Beaudoin, C. A., Sunkara, K. S., Das, M., Blundell, T. L., Floto, R. A. ]]></dc:creator>
<dc:date>2026-08-07</dc:date>
<dc:identifier>doi:10.64898/2026.08.03.741541</dc:identifier>
<dc:title><![CDATA[HANSEN: An Integrated Structural and Functional Proteome Resource for Structure-Guided Drug Discovery in Mycobacterium leprae]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-08-07</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.03.742514v1?rss=1">
<title>
<![CDATA[
AI semantics for biomedical data integration 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.03.742514v1?rss=1
</link>
<description><![CDATA[
Researchers increasingly need to explore hypotheses that span multimodal data across different scales, organisms, and domains. In practice, this requires connecting knowledge across fragmented databases with incompatible APIs and heterogeneous annotation practices. Large language model (LLM) agents can automate this data integration process, but grounding LLM agent outputs in scientifically correct sources of truth remains a significant challenge.

Here we describe our deployment of a novel AI semantics workflow using LLM agents to enable scalable data integration, grounded in biological knowledge in the form of ontologies. Our workflow comprises (1) a multi-agent system curating scientific knowledge across ontologies using the Ontology Lookup Service (OLS) as grounding; (2) an LLM embedding service to enable interoperability between scientific databases by mapping ontology terms; and (3) GrEBI, a knowledge graph and Model Context Protocol (MCP) server enabling LLM agents to conduct cross-cutting, multi-omic biomedical queries.

O_FIG O_LINKSMALLFIG WIDTH=187 HEIGHT=200 SRC="FIGDIR/small/742514v1_ufig1.gif" ALT="Figure 1">
View larger version (39K):
org.highwire.dtl.DTLVardef@379930org.highwire.dtl.DTLVardef@2a2286org.highwire.dtl.DTLVardef@40ca12org.highwire.dtl.DTLVardef@1929729_HPS_FORMAT_FIGEXP  M_FIG C_FIG
]]></description>
<dc:creator><![CDATA[ McLaughlin, J., Puig-Barbe, A., Ibrahim, A., Pava, D., Pendlington, Z. M., Matentzoglu, N., Sollis, E., Foreman, A., Wilson, R., Lopez Gomez, F., Harris, L., Adeleye, Y., Kaur, S., Meldal, B., Smedley, D., Parkinson, H. ]]></dc:creator>
<dc:date>2026-08-07</dc:date>
<dc:identifier>doi:10.64898/2026.08.03.742514</dc:identifier>
<dc:title><![CDATA[AI semantics for biomedical data integration]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-08-07</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.03.741871v1?rss=1">
<title>
<![CDATA[
Resolving the immune response across clonal cancer evolution in situ with whole transcriptome profiling 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.03.741871v1?rss=1
</link>
<description><![CDATA[
Single-cell spatial transcriptomics is now central to studying tumors in their native tissue context. Here we present the first comprehensive, independent evaluation of Atera, a new spatial whole transcriptome platform, compared against Xenium on adjacent sections of human ductal carcinoma in situ (DCIS). We show that Atera enables granular cell-state annotation and resolves rare cell populations, which we experimentally validate by multiplex immunofluorescence (IF). We further show that its transcriptome-wide coverage enables inference of copy-number alterations at single-cell resolution, allowing us to reconstruct the clonal evolution of DCIS. We orthogonally confirm the inferred copy-number alterations by whole-genome sequencing of 16 microdissected tumor regions from a consecutive tissue section. Finally, by mapping the immune microenvironment onto this clonal architecture, we demonstrate the feasibility of tracking the changes in immune response along the clonal tumor evolution in situ. Together, our results establish Atera as a validated platform for tracking clonal evolution and immune adaptation in clinical samples.
]]></description>
<dc:creator><![CDATA[ Wu, S., Mahajan, M., van der Linde, R. M., Zhu, C., van IJzendoorn, D., West, R. B., Matusiak, M. ]]></dc:creator>
<dc:date>2026-08-07</dc:date>
<dc:identifier>doi:10.64898/2026.08.03.741871</dc:identifier>
<dc:title><![CDATA[Resolving the immune response across clonal cancer evolution in situ with whole transcriptome profiling]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-08-07</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.03.742524v1?rss=1">
<title>
<![CDATA[
RiboRep: Replicate-Aware Cross-Modal Transformers for Codon-Resolved Ribosome Density Prediction 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.03.742524v1?rss=1
</link>
<description><![CDATA[
Ribosome profiling enables genome-wide measurement of translation at nucleotide resolution and provides a dynamic view of cellular protein synthesis under diverse biological conditions. Existing computational approaches primarily operate on codon-level representations, potentially losing fine-grained translational signals critical for modeling context-dependent cellular responses. Such predictive translational modeling is increasingly important for emerging biological digital twins, where accurate simulation of molecular-state dynamics is required to characterize cellular adaptation, perturbation response, and phenotype progression. We present RiboRep, a replicate-aware cross-modal transformer for codon-resolved ribosome density prediction. RiboRep jointly models nucleotide-resolution RNA sequences and reference ribosome occupancy signals using dual-stream convolutional encoders, RoPE-based self-attention, asymmetric cross-attention, and replicate-aware conditioning tokens. By explicitly modeling replicate-specific variation and integrating sequence context with experimentally observed translational activity, RiboRep provides a framework for reconstructing and simulating translational states across biological conditions. Across bacterial, yeast, and plant ribosome profiling datasets, RiboRep achieves competitive or improved performance compared with existing baselines, with particularly strong gains on replicate-rich plant datasets. Ablation studies further demonstrate the importance of local codon-aware feature extraction, replicate-aware conditioning, and gated readout. Beyond predictive performance, the proposed framework establishes a foundation for translation-aware molecular digital twins capable of modeling ribosome occupancy landscapes, perturbation-induced translational responses, and condition-specific regulatory programs at codon resolution. 1
]]></description>
<dc:creator><![CDATA[ Kuo, A., Yue, Z., Ku, W.-S., Chen, H. ]]></dc:creator>
<dc:date>2026-08-07</dc:date>
<dc:identifier>doi:10.64898/2026.08.03.742524</dc:identifier>
<dc:title><![CDATA[RiboRep: Replicate-Aware Cross-Modal Transformers for Codon-Resolved Ribosome Density Prediction]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-08-07</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.03.742471v1?rss=1">
<title>
<![CDATA[
PhysioMap: an ontology-grounded causal knowledge graph of human physiology 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.03.742471v1?rss=1
</link>
<description><![CDATA[
Computational physiology needs representations that connect traits across biological scales while distinguishing causal, constitutive, and mathematical dependencies. We present PhysioMap, an ontology-grounded knowledge base of contextualized physiological traits and precisely defined relation types. A versioned projection maps entailed ontology patterns to a typed causal knowledge graph that constrains quantitative structural causal models. Derivative signs provide a separate qualitative abstraction, which the PhysioMap solver uses to analyze steady-state responses in the presence of feedback. A stratified expert review across all relation types supported most sampled relations and isolated a minority for correction or further investigation. In a rare metabolic disease application, nearly all determinate predictions agreed with the HPO-derived reference before post-hoc review; after the discordant reference directions were excluded, all remaining determinate predictions agreed. Shortest signed paths produced directional errors, particularly on cases for which the PhysioMap solver did not determine a direction, indicating that its abstentions concentrated difficult cases. Abduction usually narrowed the candidate set but often did not identify a unique cause. PhysioMap therefore connects ontology-grounded physiological content to interventional prediction and abduction under incomplete quantitative knowledge. Because PhysioMap curation and the HPO-derived reference may share supporting literature, and because abduction used a closed candidate pool, these analyses do not constitute independent clinical validation.
]]></description>
<dc:creator><![CDATA[ Hoehndorf, R., SCHOFIELD, P., Gkoutos, G. V. ]]></dc:creator>
<dc:date>2026-08-07</dc:date>
<dc:identifier>doi:10.64898/2026.08.03.742471</dc:identifier>
<dc:title><![CDATA[PhysioMap: an ontology-grounded causal knowledge graph of human physiology]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-08-07</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.03.742474v1?rss=1">
<title>
<![CDATA[
Synthetic Longitudinal Tabular Data Generation via Copula 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.03.742474v1?rss=1
</link>
<description><![CDATA[
Synthetic data generation is increasingly used to enable data sharing and secondary analysis while protecting participant privacy, particularly for longitudinal tabular health data, where repeated measures per subject create within-subject dependence that most synthetic data methods are not designed to preserve. Existing generative methods, particularly generative adversarial network (GAN)-based approaches, can model complex distributions, but their estimated dependence structures are often difficult to interpret and their performance may be unstable or prone to overfitting in modestly sized datasets. Here we show that eCDF-copula, a statistically rooted approach using the empirical cumulative distribution function (eCDF) and copula modeling, preserves within- and between-visit dependence structure. To handle pervasive missing data, we propose a two-stage strategy combining multiple imputation with copula-based synthesis, enabling a variance decomposition that quantifies replication variability across methods. We benchmarked the proposed approach against four established methods on two longitudinal clinical datasets spanning markedly different sample sizes (n = 120 vs. n = 3, 612). eCDF-copula achieved resemblance and utility exceeding those of state-of-the-art synthetic data methods, while maintaining comparable privacy.
]]></description>
<dc:creator><![CDATA[ Cai, H., Yu, W., Lu, R., Chattopadhyay, I., Zhang, X., Liu, J. ]]></dc:creator>
<dc:date>2026-08-07</dc:date>
<dc:identifier>doi:10.64898/2026.08.03.742474</dc:identifier>
<dc:title><![CDATA[Synthetic Longitudinal Tabular Data Generation via Copula]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-08-07</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.08.03.742406v1?rss=1">
<title>
<![CDATA[
REFCON: Reference-free and robust copy number inference in single-cell tumor transcriptomes 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.08.03.742406v1?rss=1
</link>
<description><![CDATA[
Single-cell RNA sequencing (scRNA-seq) is widely used to infer copy number profiles from tumor cells. Existing methods build on a reference-based normalization paradigm: normalizing each tumor cell against a reference of normal cells, whether supplied, in-sample, or synthesized. This makes them reference-dependent and as a result, sensitive to cohort composition, and prone to false positives. To address these limitations, we introduce REFCON, a deep-learning model that enables reference-free copy number profiling from scRNA-seq data. REFCON estimates local copy-number deviations and jointly optimizes them into a genome-wide per-cell profile. It profiles pure tumors, generalizes to unseen tissues and platforms, and stays robust to cohort composition. Predicted copy number profiles distinguish malignant cells with high specificity, producing far fewer false-positive calls, and improve clonal reconstruction. The model can also benefit from reference cells when available, turning a field requirement into an optional refinement. Hence, REFCON extends reliable per-cell copy number profiling to the scRNA-seq data collected without matched normals.
]]></description>
<dc:creator><![CDATA[ Gencturk, M. M., Cicek, A. E. ]]></dc:creator>
<dc:date>2026-08-07</dc:date>
<dc:identifier>doi:10.64898/2026.08.03.742406</dc:identifier>
<dc:title><![CDATA[REFCON: Reference-free and robust copy number inference in single-cell tumor transcriptomes]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-08-07</prism:publicationDate>
<prism:section></prism:section>
</item>
</rdf:RDF>
