<?xml version="1.0" encoding="UTF-8" ?>
<rdf:RDF xmlns:admin="http://webns.net/mvcb/" xmlns="http://purl.org/rss/1.0/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:prism="http://purl.org/rss/1.0/modules/prism/" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:syn="http://purl.org/rss/1.0/modules/syndication/">
<channel rdf:about="https://biorxiv.org">
<admin:errorReportsTo rdf:resource="mailto:biorxiv@cshlpress.edu"/>
<title>bioRxiv Subject Collection: Genomics Bioinformatics</title>
<link>https://biorxiv.org</link>
<description>
This feed contains articles for bioRxiv Subject Collection "Genomics Bioinformatics"
</description>

<items>
<rdf:Seq>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.07.10.737871v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.07.08.737305v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.07.08.737319v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.07.08.737273v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.07.08.737248v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.07.10.737551v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.07.10.737715v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.07.11.737898v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.07.08.737157v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.07.10.737679v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.07.10.737796v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.07.10.737754v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.07.10.737777v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.07.09.736945v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.07.11.737882v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.07.10.737775v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.07.10.737545v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.07.08.737144v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.07.07.736241v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.07.08.737203v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.07.07.736532v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.07.07.737032v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.07.07.737090v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.07.07.737093v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.07.07.737051v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.07.07.736817v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.07.07.737037v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.07.07.737010v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.07.07.736993v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.07.10.735823v1?rss=1"/>
</rdf:Seq>
</items>
<prism:eIssn/>
<prism:publicationName>bioRxiv</prism:publicationName>
<prism:issn/>

<image rdf:resource=""/>
</channel>
<image rdf:about="">
<title>bioRxiv</title>
<url>https://www.biorxiv.org/sites/default/files/bioRxiv_article.jpg</url>
<link>https://www.biorxiv.org</link>
</image>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.07.10.737871v1?rss=1">
<title>
<![CDATA[
EcoMorph: Universal morphological trait quantification from natural language prompts for ecological research 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.07.10.737871v1?rss=1
</link>
<description><![CDATA[
1. Morphological traits such as floral area and body size are fundamental to ecological research, serving as inputs for studies of pollinator-plant interactions, habitat quality, and biodiversity monitoring. However, accurately measuring these traits from images remains challenging, particularly in complex field conditions where existing tools exhibit reduced accuracy and limited generalizability across taxa. 2. We present EcoMorph, a modular morphological measurement system that leverages the Segment Anything Model 3 (SAM3) to quantify traits across diverse ecological contexts. Unlike task-specific segmentation models requiring domain-specific training data, SAM3's prompt-based architecture enables segmentation of arbitrary biological structures from natural-language prompts, using the same underlying model across flowers, insects, and other targets without retraining. From the resulting segmentations, EcoMorph extracts three classes of measurement: area, linear dimensions, and object counts. 3. We validated EcoMorph across two ecological scales. At the intermediate scale, EcoMorph-derived floral area agreed closely with manual ImageJ measurements (R2 = 0.935, n = 74) under simple-background conditions and (R2 = 0.928, n = 58) under complex-background conditions, with valid predictions for 95% of images. At the fine scale, EcoMorph-derived insect body area was strongly correlated with hand-measured intertegular distance (r = 0.810, n = 349), capturing body-size variation across species from the small Bombus impatiens to the large Xylocopa virginica. Object counts matched manual counts almost exactly for well-separated insects in an insect box (R2 = 0.9997, n = 12). 4. By combining prompt-based segmentation with modular measurement, EcoMorph enables high-throughput quantification of area, size, and abundance from heterogeneous image sources without taxon-specific training. This generality supports a broad range of ecological applications, including pollinator and plant trait research, biodiversity and abundance monitoring, and allometric biomass estimation.
]]></description>
<dc:creator><![CDATA[ Amoah, E. I., Bunch, Z., Thomas, H. M., Patch, H. M., Grozinger, C. ]]></dc:creator>
<dc:date>2026-07-12</dc:date>
<dc:identifier>doi:10.64898/2026.07.10.737871</dc:identifier>
<dc:title><![CDATA[EcoMorph: Universal morphological trait quantification from natural language prompts for ecological research]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-07-12</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.07.08.737305v1?rss=1">
<title>
<![CDATA[
TEscape: Defining the human transposable element transcriptome using multiplatform long-read sequencing 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.07.08.737305v1?rss=1
</link>
<description><![CDATA[
Transposable elements (TEs) not only account for half of the human genome sequence but also generate transcripts that contribute to transcriptomic diversity. Yet, their repetitive nature has hindered accurate quantification of the full TE-derived transcriptome, a challenge that long-read sequencing can overcome. Here, we combined multiplexed arrays isoform sequencing (MAS-ISO-seq) with a dedicated computational framework (TEscape) to perform an in-depth annotation of the human TE transcriptome. To capture the breadth of human transcriptome diversity, we profiled six representative cell types spanning three distinct biological contexts, including metabolism with, primary patient-derived adipogenic cells at two differentiation stages, and iPSC derived hepatic progenitor cells; the nervous system with iPSC-derived neurons, neural progenitor cells (NPCs), and pluripotency using induced pluripotent stem cells (iPSCs). Together, these datasets yielded over 235 million full-length long reads. First, to assess data coverage and transcriptome depth, we quantified protein-coding gene expression, detecting 14,312 genes (73.6% of all annotated protein-coding genes), which is a level consistent with deep and comprehensive transcriptome representation. Second, focusing on TE-derived transcripts, we identified >83,000 previously unannotated isoforms, the vast majority (84%) originating from a complex combination of multi-TEs. We also identified solo TEs, which are predominantly from LINE1 (14%). We confirmed that TE-transcripts are able to be exemplified by signatures detected in Liver Hepatocellular Carcinoma (LICH). Together, MAS-ISO-seq and TEscape establish the first long-read-based, high-resolution atlas of transcribed human TEs, providing a foundational resource for integrative transcriptome analyses and for investigating TE expression and regulation in health and disease.
]]></description>
<dc:creator><![CDATA[ Mercuri, R. L. V., Mombach, D. M., dos Santos, F. R. C., Perez-Schindler, J., Huang, Y., Spealman, P., Pintacuda, G., Al'Khafaji, A., Donnard, E. R., Claussnitzer, M., Galante, P. A. F. ]]></dc:creator>
<dc:date>2026-07-12</dc:date>
<dc:identifier>doi:10.64898/2026.07.08.737305</dc:identifier>
<dc:title><![CDATA[TEscape: Defining the human transposable element transcriptome using multiplatform long-read sequencing]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-07-12</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.07.08.737319v1?rss=1">
<title>
<![CDATA[
A geometric atlas of how ESM3 organizes modalities across depth 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.07.08.737319v1?rss=1
</link>
<description><![CDATA[
Protein language models learn general-purpose representations from large collections of protein sequences and structures, and have advanced the prediction of protein structure and function. ESM3 is a multimodal protein language model that ingests a protein through several channels at once, including amino-acid sequence, three-dimensional structure, secondary structure (SS8), solvent accessibility (SASA), and discrete functional annotations, summing their embeddings into a single residual stream. Little is known about whether these modalities occupy separate subspaces and the depth at which they fuse. The present analysis examines ESM3 (esm3-sm-open-v1; 1.4 billion parameters; 48 transformer layers) once per modality in isolation and applies representational-similarity analysis across all 48 layers. The four physical modalities (sequence, structure, SS8, SASA) begin in distinct subspaces, remain maximally separated through roughly the first half of layers, and then fuse into a shared low-dimensional subspace between layers 25 and 35. The fusion is ordered. The structure-derived modalities (structure, SS8, SASA) are mutually aligned from the input, whereas sequence joins last, after layer 28. The functional-annotation modality never fuses; instead, it remains representationally orthogonal to the physical modalities at every layer, and this orthogonality holds whether the annotation is supplied as whole-protein or per-residue, suggesting that it is content-driven rather than a tokenization artifact. The fusion is a learned property, absent in a randomly initialized model of the same architecture, holds at the residue level below the mean-pool, and reorganizes variance, converting between-condition variance into within-condition variance while the stream never approaches isotropy. Fusion depth is independent of protein length but is delayed by structural disorder. The phenomenon is universal across diverse organisms. Across 5,555 proteins from 12 organisms spanning eukaryota, bacteria, and archaea, every superkingdom (and every individual organism) reaches peak modality fusion at the same network depth (layer 35).
]]></description>
<dc:creator><![CDATA[ Steenwyk, J. L. ]]></dc:creator>
<dc:date>2026-07-12</dc:date>
<dc:identifier>doi:10.64898/2026.07.08.737319</dc:identifier>
<dc:title><![CDATA[A geometric atlas of how ESM3 organizes modalities across depth]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-07-12</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.07.08.737273v1?rss=1">
<title>
<![CDATA[
Genomic Annotation Infrastructure (GAIn): Pipelines and Resource Repositories for Annotating Variants, Positions, and Regions 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.07.08.737273v1?rss=1
</link>
<description><![CDATA[
Interpretation of genomic variants, positions, and regions depends on reliable annotation - adding evidence such as predicted effect, conservation, population frequency, and gene-level context - yet the underlying resources are numerous, versioned, and assembly-specific. We present the Genomic Annotation Infrastructure (GAIn), a platform that generates transparent, reproducible annotations via declarative pipelines that define annotation tasks as ordered lists of components, called annotators, that produce annotation attributes using genomic resources from Genomic Resource Repositories (GRRs). We provide two public GRRs: a main repository containing more than 250 heterogeneous genomic resources, and a separate GRR-ENCODE repository containing resources derived from thousands of ENCODE (Encyclopedia of DNA Elements) project experiments. Users can use the annotation pipelines we made available, author custom annotation pipelines, and execute annotation tasks with these pipelines via GAIn's web and command-line interfaces. The web interface can be used without any setup, but it relies on shared computational infrastructure and imposes limits on the size of annotation tasks. The command-line interface requires setup but supports arbitrarily large annotation tasks through simple-to-use parallelization and offers a broader set of features. For example, command-line GAIn can be extended by using custom GRRs or creating custom annotators via its plugin architecture. In addition, GAIn's re-annotation feature, which updates annotations as they evolve, substantially simplifies maintaining annotations in a large genomics analysis project. GAIn's resource management, explicit versioning, and pipeline abstraction provide an auditable, maintainable, and efficient foundation for modern genomic annotation across reference assemblies and use cases.
]]></description>
<dc:creator><![CDATA[ Cokol, M., Chorbadjiev, L., Lee, Y.-h., Jamsandekar, M., Gergova, I., Todorov, I., Iossifov, I. ]]></dc:creator>
<dc:date>2026-07-12</dc:date>
<dc:identifier>doi:10.64898/2026.07.08.737273</dc:identifier>
<dc:title><![CDATA[Genomic Annotation Infrastructure (GAIn): Pipelines and Resource Repositories for Annotating Variants, Positions, and Regions]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-07-12</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.07.08.737248v1?rss=1">
<title>
<![CDATA[
OCellus: A Language-Model Framework for Single-Cell, Spatial, and Perturbation Biology with Natural-Language Reasoning 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.07.08.737248v1?rss=1
</link>
<description><![CDATA[
Computational modeling of cellular behavior - the virtual cell - has emerged as a stated grand challenge at the intersection of artificial intelligence and biology, yet existing foundation models remain specialized: single-cell models process dissociated transcriptomes only, spatial models require dedicated spatial-aware architectures, and perturbation predictors depend on manually curated knowledge bases that cap generalization. Here we introduce OCellus, a single nine-billion-parameter language model (Qwen3.5-9B) fine-tuned on twenty-two biological tasks that simultaneously addresses all three limitations through three coordinated technical contributions on a shared backbone. First, EvenClock encodes two-dimensional spatial coordinates as eighteen clockface sectors of text, enabling spatial reasoning on a vanilla language model without architectural modification; on ten spatial transcriptomics tasks OCellus attains 77 percent spatial-neighborhood accuracy, 96 percent spatial-cellchat accuracy, and 0.70 proportion-cosine similarity on spatial deconvolution, all without any spatial-aware architectural components. Second, per-gene language-model embeddings replace the Gene Ontology annotations that GEARS depends on, achieving Pearson correlation 0.945 on the Replogle 2022 perturbation benchmark versus 0.84 for GEARS across 457 completely unseen knockout genes. Third, OCellus-Agent provides a Planner-Router-Verifier natural-language interface that achieves 75 percent pipeline accuracy on eighty multi-task queries. Removing language-model embeddings collapses perturbation Pearson to 0.06, confirming that learned functional representations - not graph topology - drive the gain. As a cell-type encoder, OCellus ranks first among fourteen foundation models in linear-probe accuracy at 95.1 percent across four benchmark datasets, and reaches 72.6 percent average across twenty-two evaluated biological tasks - a 57-percentage-point absolute gain over the strongest baseline configuration. As a language model, OCellus uniquely generates natural-language explanations of its predictions, a capability absent from all competing methods. Code, pre-trained model weights, the graph-neural-network module, and the agent system will be made available upon publication.
]]></description>
<dc:creator><![CDATA[ Zhang, C., Sun, J., Xu, Z., Liao, R., Yin, A., Gao, H., Liu, E., Bao, Y., Zhao, L., Wang, G. ]]></dc:creator>
<dc:date>2026-07-12</dc:date>
<dc:identifier>doi:10.64898/2026.07.08.737248</dc:identifier>
<dc:title><![CDATA[OCellus: A Language-Model Framework for Single-Cell, Spatial, and Perturbation Biology with Natural-Language Reasoning]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-07-12</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.07.10.737551v1?rss=1">
<title>
<![CDATA[
ProtBLIP2-SST: Protein Function Prediction via BLIP2 with Sequence, Structure, and Text 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.07.10.737551v1?rss=1
</link>
<description><![CDATA[
Protein function prediction traditionally relies on structured gene ontology (GO) labels or multi-label classifiers. However, these labels or classifiers cannot flexibly describe molecular function, biological process, cellular component, and free-text functional narratives in a single output. In comparison, generation-based approaches offer an intuitive paradigm for flexible free-text protein annotation, with large language models (LLMs) as a representative method for protein-text modeling. Recent efforts on utilizing LLMs for protein semantic understanding and annotation generation have adopted sequence-only encoding or sequence-text contrastive alignment paradigms, yet without explicit consideration of three-dimensional structural information. To address these limitations in current protein function prediction methods, we present ProtBLIP2-SST, a two-stage framework built on the BLIP2 model architecture that bridges protein sequence, structure, and text for open-ended protein functional caption generation. Specifically, we first integrate sequence and structure information through SaProt, a protein language model (PLM) with a structure-aware vocabulary that fuses residue tokens with Foldseek-derived 3Di structural tokens. To empower the LLM to understand protein semantics, we employ a Q-Former (a querying transformer in BLIP2) with learnable query tokens as the cross-modal projector to align protein features from the frozen SaProt encoder and text features from a frozen BiomedBERT via protein--text contrasting, protein--text matching, and protein captioning objectives. After alignment, the protein features are linearly projected and prepended to the prompt embeddings of the LLM for protein captioning fine-tuning with LoRA. Trained on 441k protein--text pairs from Swiss-Prot with corresponding structures from the AlphaFold Database, our ProtBLIP2-SST outperforms sequence-only and sequence-text alignment baselines on protein captioning metrics, with ablation studies demonstrating the effectiveness of integrating structure with sequence information for improved protein understanding. Through a unified two-stage alignment-and-generation pipeline, ProtBLIP2-SST integrates protein sequence and structural information, overcomes the rigidity of traditional GO-centric classification, generating open-ended captions that jointly describe molecular function, subcellular location, and homology context in one single output.
]]></description>
<dc:creator><![CDATA[ Chen, Z., Luo, Q. ]]></dc:creator>
<dc:date>2026-07-12</dc:date>
<dc:identifier>doi:10.64898/2026.07.10.737551</dc:identifier>
<dc:title><![CDATA[ProtBLIP2-SST: Protein Function Prediction via BLIP2 with Sequence, Structure, and Text]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-07-12</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.07.10.737715v1?rss=1">
<title>
<![CDATA[
High resolution Streptococcus pyogenes core genome MLST and LIN coding scheme for outbreak detection 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.07.10.737715v1?rss=1
</link>
<description><![CDATA[
Streptococcus pyogenes is a globally important pathogen responsible for at least 500,000 deaths a year, causing significant burden on healthcare systems. It is the causative agent for ailments such as impetigo and strep throat to septicaemia and necrotizing fasciitis. Assessment of genetic relatedness for the detection of outbreaks within communities or healthcare facilities is vital in decreasing the propagation of S. pyogenes within these settings, alongside epidemiological data. As the volume of isolates being sequenced increases year on year, more scalable and sharable methodologies of assessing genetic relatedness are required by reference laboratories and for international collaboration. LIN codes, applied to core genome MLST (cgMLST) represent a method which is extensible to large scale whole genome sequencing (WGS) while still being sufficiently sensitive to detect outbreak clusters. Here we present a novel cgMLST and LIN code scheme, hosted by PubMLST, enabling international collaboration and global tracking of variants, that is highly scalable and usable for all. The schemes are available at https://pubmlst.org/organisms/streptococcus-pyogenes.
]]></description>
<dc:creator><![CDATA[ Ryan, Y., Jolley, K. A., Hearn, H., Parfitt, K. M., Platt, S., Lamagni, T., Moganeradj, K. ]]></dc:creator>
<dc:date>2026-07-11</dc:date>
<dc:identifier>doi:10.64898/2026.07.10.737715</dc:identifier>
<dc:title><![CDATA[High resolution Streptococcus pyogenes core genome MLST and LIN coding scheme for outbreak detection]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-07-11</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.07.11.737898v1?rss=1">
<title>
<![CDATA[
ArthroVerse: mapping protein family diversity across arthropod-associated microbiomes 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.07.11.737898v1?rss=1
</link>
<description><![CDATA[
Metagenomic studies of arthropod-associated microbiomes have generated vast amounts of sequence data, yet the functional and structural organization of these proteins remains largely unexplored. Here, we present ArthroVerse, the first comprehensive database of protein families derived from arthropod-associated metagenomes. Non-redundant protein families were generated after rigorous filtering, deduplication, and clustering. The protein families were further annotated with microbial taxonomy, host associations, protein structural information, and Carbohydrate-active enzymes (CAZyme) predictions. The resulting dataset integrates both metagenomic and reference genome-derived proteins, enabling systematic exploration of functional diversity, evolutionary relationships, and host-microbe interactions in insect microbiomes. ArthroVerse provides a valuable resource for the study of microbial ecology and arthropod physiology, offering unprecedented insight into the protein landscape of insect-associated microbial communities.
]]></description>
<dc:creator><![CDATA[ Chasapi, I. N., Aplakidou, E., Chasapi, M. N., Lamari, E., Galaras, A., Diplari, S., Iliopoulos, I., Emiris, I. Z., Georgakopoulos-Soares, I., Patalano, S., Stravopodis, D. J., Karatzas, E., Baltoumas, F. A., Kyrpides, N., Pavlopoulos, G. A. ]]></dc:creator>
<dc:date>2026-07-11</dc:date>
<dc:identifier>doi:10.64898/2026.07.11.737898</dc:identifier>
<dc:title><![CDATA[ArthroVerse: mapping protein family diversity across arthropod-associated microbiomes]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-07-11</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.07.08.737157v1?rss=1">
<title>
<![CDATA[
onsite: An Integrated Framework for Phosphosite Localization and False Localization Rate Estimation 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.07.08.737157v1?rss=1
</link>
<description><![CDATA[
With the rapid development of mass spectrometry-based proteomics, the volume of phosphoproteomic data has increased substantially. However, accurate localization of phosphorylation sites and standardized statistical validation remain critical analytical bottlenecks. To address the lack of standardized cross-algorithm evaluation, we introduce onsite, a unified and open-source Python framework. onsite integrates an alanine-decoy strategy to estimate the false localization rate (FLR) across three algorithms: AScore, PhosphoRS, and pyLucXor. This modular architecture efficiently processes large-scale datasets and enables global FLR calculation. Benchmarking on the standard synthetic phosphopeptide dataset PXD000138 highlighted distinct inter-algorithmic variations. Using the same 5% global FLR threshold, pyLucXor localized the most target sites (28,353). It also reached a high accuracy (91.22%) against the known ground truth, resulting in the largest number of correctly localized sites (25,865). Reanalysis of the highly fractionated, large-scale PXD012255 dataset further demonstrated that native integration of onsite into the quantms pipeline enables scalable processing and provides a standardized framework for FLR control in large-scale phosphoproteomics.
]]></description>
<dc:creator><![CDATA[ Yue, Q.-X., Wei, Z., Dai, C., Bai, M., Perez-Riverol, Y., Sachsenberg, T. ]]></dc:creator>
<dc:date>2026-07-11</dc:date>
<dc:identifier>doi:10.64898/2026.07.08.737157</dc:identifier>
<dc:title><![CDATA[onsite: An Integrated Framework for Phosphosite Localization and False Localization Rate Estimation]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-07-11</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.07.10.737679v1?rss=1">
<title>
<![CDATA[
M6AFormer Prioritizes Unannotated Functional m6A Candidate Sites in the Human m6A Epitranscriptome 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.07.10.737679v1?rss=1
</link>
<description><![CDATA[
N6-methyladenosine (m6A) is a pervasive RNA modification with critical roles in post-transcriptional regulation, yet accurate transcriptome-wide identification of functional m6A sites remains challenging. Here, we present M6AFormer, a hybrid deep-learning framework that combines convolutional feature extraction with a lightweight Transformer to capture both local sequence motifs and broader contextual dependencies. M6AFormer consistently outperformed representative m6A predictors, including MST-M6A, CLSM6A and deepSRAMP. Transcriptome-wide scanning revealed a large repertoire of previously unannotated candidate m6A sites that retained hallmark m6A features, including canonical motif enrichment, characteristic spatial distribution and preferential overlap with m6A writer and reader binding regions. Importantly, M6AFormer-predicted sites were broadly associated with genetic and disease-relevant features, including SNPs, sequence variants and GWAS-linked loci, suggesting their potential contribution to human disease mechanisms. Finally, experimental validation confirmed a previously unreported m6A site in NEU4 mRNA and demonstrated its functional impact on cancer cell migration. Together, M6AFormer provides an accurate, interpretable and biologically informative framework for m6A site discovery.
]]></description>
<dc:creator><![CDATA[ Niu, Z., Liu, C., Gu, L. ]]></dc:creator>
<dc:date>2026-07-11</dc:date>
<dc:identifier>doi:10.64898/2026.07.10.737679</dc:identifier>
<dc:title><![CDATA[M6AFormer Prioritizes Unannotated Functional m6A Candidate Sites in the Human m6A Epitranscriptome]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-07-11</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.07.10.737796v1?rss=1">
<title>
<![CDATA[
Panomap: Unbiased Nanopore Signal Mapping with Pangenome Variation Graphs 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.07.10.737796v1?rss=1
</link>
<description><![CDATA[
Motivation: Signal-space nanopore mappers enable real-time mapping and filtering decisions directly from raw nanopore signals. However, existing signal-space mappers are built around linear references, and using a single representative reference can introduce reference bias when the sample diverges from that reference. Pangenome reference collections can reduce this bias by representing diversity across related reference sequences, but linear-reference signal mappers must treat each sequence as a separate target, redundantly storing shared sequences. Pangenome variation graphs provide a more compact representation by storing shared sequences once and encoding variants as alternative paths through the graph. Although sequence-to-graph mapping is well established for basecalled reads, existing signal-space methods do not directly use pangenome variation graphs. Results: We present Panomap, the first signal-space mapper that operates on pangenome variation graphs. Panomap maps raw nanopore signals to graph references, allowing signal-space mapping to use pangenome diversity while representing shared sequences once. We evaluate Panomap in three settings. First, when a single reference already maps the sample well, Panomap preserves mapping accuracy as additional reference sequences are added to the reference collection, while state-of-the-art signal-space tools regress. Second, when the exact sample strain is absent from the reference collection, Panomap benefits from adding related assemblies from the same species to the pangenome reference. Third, using a highly polymorphic locus, we show that Panomap can map reads from alleles not represented in the reference collection by using related alleles in the pangenome, with the largest gains for more divergent alleles and for decisions made from short prefixes of the read signal. In addition, Panomap's graph index scales sublinearly with pangenome collection size. Together, these results show that Panomap brings population-aware reference representation into signal-space mapping. Availability and Implementation: Panomap is open source and available at https://github.com/cornell-brg/panomap.
]]></description>
<dc:creator><![CDATA[ Shih, P. J., Sanghani, Z., Guarracino, A., Gamaarachchi, H., Batten, C. ]]></dc:creator>
<dc:date>2026-07-11</dc:date>
<dc:identifier>doi:10.64898/2026.07.10.737796</dc:identifier>
<dc:title><![CDATA[Panomap: Unbiased Nanopore Signal Mapping with Pangenome Variation Graphs]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-07-11</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.07.10.737754v1?rss=1">
<title>
<![CDATA[
Comparative Analysis of Transposable Elements in Hermetia illucens 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.07.10.737754v1?rss=1
</link>
<description><![CDATA[
Background The black soldier fly (Hermetia illucens) is an emerging model for bioconversion and industrial rearing. Its genome is highly repetitive, yet the contribution of transposable elements (TEs) to population divergence and demographic processes. The sampled populations represent a gradient of demographic histories, including wild and near-wild North American populations, and domesticated European strains with shared industrial origins. Difference in TE composition may influence genome structure, regulatory variation, and evolutionary responses to captive environments. Results A comparative analysis of the repetitive landscape was done for four H. illucens genomes, one of which is a wild-caught specimen. Total repeat content was high across all assemblies (67.6% to 70.8%) and dominated by LINE elements. Class-level TE diversity was nearly identical among genomes, but multiple DNA transposon families showed distinct lineage-specific differences. Large families including Maverick and Academ were generally depleted relative to the wild sample. Divergence profiles revealed patterns consistent with recent turnover in several families. Family level turnover, rather than class level change, accounted for the most difference among the genomes. TE-associated structural variants (TESVs) were also not uniformly distributed. Most chromosomes showed mid-chromosome enrichment, and a pronounced TESV peak on chromosome 5 overlapped a histone rich region containing many unclassified repeats. Use of a repeat library derived from multiple genomes increased the number of detected TESVs and improved classification within complex regions, demonstrating that multi-genome libraries enhance annotation accuracy compared to single reference-based models. Conclusions Multiple DNA transposon families show evidence of recent or lineage-specific amplification in H. illucens, suggesting that TE amplification contributes to genome variation during demography-associated TE turnover. The multi-genome-based library improved TE detection and classification, providing a proof of concept that even a small lineage-inclusive repeat library enhances annotation accuracy and capture TE diversity missed by single-reference approaches. Together, these findings demonstrate that TE family turnover plays a significant role in shaping genome architecture and adaptation in this species. Keywords: Hermetia illucens, transposable elements, demographic processes, structural variants
]]></description>
<dc:creator><![CDATA[ Hector Rosche-Flores, H., Fischer, S., Picard, C. J. ]]></dc:creator>
<dc:date>2026-07-11</dc:date>
<dc:identifier>doi:10.64898/2026.07.10.737754</dc:identifier>
<dc:title><![CDATA[Comparative Analysis of Transposable Elements in Hermetia illucens]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-07-11</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.07.10.737777v1?rss=1">
<title>
<![CDATA[
Characterizing the Small Non-Coding RNA Pathways in the Invasive Zebra Mussel (Dreissena polymorpha) 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.07.10.737777v1?rss=1
</link>
<description><![CDATA[
The zebra mussel (Dreissena polymorpha) is an invasive species that causes extensive economic and ecological damage. Here, we identify and characterize the key components of the small RNA (sRNA) and RNA interference (RNAi) pathways in zebra mussels. Like other mollusks, zebra mussels have extensive microRNA (miRNA) and Piwi-interacting RNA (piRNA) machinery but lack or have modified canonical factors needed to produce small interfering RNA (siRNA). Specifically, the zebra mussel Dicer sequence displays substitutions in the conserved DEAD box motif that is required for substrate processivity, and this organism also lacks some attendant accessory factors such as R2D2. We sequenced the small RNA found in both isolated somatic tissue (adductor muscle) and whole animals (including germline), and identified both conserved and novel miRNA and diverse piRNA sequences, but few endogenous siRNAs. To determine whether their remaining sRNA machinery could still be co-opted to initiate gene silencing, we injected dsRNA targeting several genes into zebra mussel adductor muscle. The injected rpn8-targeting dsRNA reduced rpn8 mRNA levels and was processed into sRNA that resemble endogenous miRNAs and piRNAs. The levels of both sRNA types correlated with mRNA knockdown, suggesting that they may act together to initiate RNAi as seen elsewhere. dsRNA targeting other genes produced variable results suggesting that particular criteria may be needed to trigger an RNAi response in this assay. Our results characterize endogenous sRNA pathways in zebra mussels, establish that dsRNA can induce RNAi, and lay the groundwork for further optimizations to establish RNAi-based genetic manipulation tools for this damaging invasive species.
]]></description>
<dc:creator><![CDATA[ Hernandez Elizarraga, V. H., O'Brien, L. G., Ballantyne, S., Gohl, D. M. ]]></dc:creator>
<dc:date>2026-07-11</dc:date>
<dc:identifier>doi:10.64898/2026.07.10.737777</dc:identifier>
<dc:title><![CDATA[Characterizing the Small Non-Coding RNA Pathways in the Invasive Zebra Mussel (Dreissena polymorpha)]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-07-11</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.07.09.736945v1?rss=1">
<title>
<![CDATA[
PKProbDesign: RNA inverse folding including pseudoknots by optimizing thermodynamic folding probability 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.07.09.736945v1?rss=1
</link>
<description><![CDATA[
Motivation: RNA inverse folding, the design of RNA sequences that fold into specified target structures, is a central problem in RNA design, with applications in functional RNA engineering, synthetic biology, and nucleic-acid therapeutics. This task becomes especially challenging for pseudoknotted target structures because pseudoknots disrupt the nested structure assumed by standard thermodynamic folding models. Existing pseudoknot inverse-folding methods often rely on structure-predictor-based objectives. Direct optimization of the thermodynamic folding probability of a specified pseudoknotted target remains limited. This requires an evaluator that can assign target-specific folding probabilities within a pseudoknot-aware ensemble and can be used as an optimization signal. Results: We present PKProbDesign, a sampling-based inverse-folding framework that directly optimizes a thermodynamic folding-probability objective for pseudoknotted targets. For each target, candidate sequences are scored by combining the folding probability of a pseudoknot-free sca[ff]old with the conditional folding probability of the remaining extension component. On 354 PseudoBase++ targets, PKProbDesign achieved the highest folding probability on 221 targets, compared with 117 for DesiRNA and 16 for MODENA. Conclusions: PKProbDesign demonstrates that pseudoknot inverse folding can be formulated around target folding probabilities rather than structure-prediction agreement alone. By combining sca[ff]old decomposition with HFold/CParty-consistent conditional-ensemble evaluation, the method provides a practical probability-based framework for designing sequences for density-2 pseudoknotted targets. Availability: The source code of PKProbDesign is available at https://github.com/TakumiOtagaki/ PKProbDesign.
]]></description>
<dc:creator><![CDATA[ Otagaki, T., Iwakiri, J., terai, g., Asai, K., Sato, K. ]]></dc:creator>
<dc:date>2026-07-11</dc:date>
<dc:identifier>doi:10.64898/2026.07.09.736945</dc:identifier>
<dc:title><![CDATA[PKProbDesign: RNA inverse folding including pseudoknots by optimizing thermodynamic folding probability]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-07-11</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.07.11.737882v1?rss=1">
<title>
<![CDATA[
ProtPen combines sequence- and structure-based approaches to facilitate protein function predictions on a proteome-wide scale 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.07.11.737882v1?rss=1
</link>
<description><![CDATA[
Proteins of unknown function represent a significant gap in our understanding of biological processes, encompassing large portions of the proteomes of many organisms, especially prokaryotes. Addressing this gap is critical to understanding the biology and pathogenicity of such organisms. We introduce ProtPen, an open-source pipeline that facilitates protein function prediction by combining eggNOG-mapper for sequence-based annotation with Foldseek for rapid structural similarity searches using AlphaFold-predicted protein structures. Annotation results from both tools are merged and enriched with UniProt metadata to produce a comprehensive output suitable for downstream analysis. The pipeline requires only a FASTA input file with UniProt identifiers, and is designed to analyze datasets on the scale of whole proteomes. Benchmarking on a curated dataset of well-characterized Pseudomonas aeruginosa proteins demonstrated an annotation accuracy of >90%, and highlighted the complementarity of sequence- and structure-based methods. Further evaluation of ProtPen included its application to biologically relevant datasets, comprising proteins of unknown function that exhibited significant differential abundances in a proteomics dataset of P. aeruginosa, and uncharacterized glycoproteins from Haloferax volcanii. ProtPen is readily extensible to incorporate additional protein function prediction tools. In summary, this pipeline facilitates the systemwide annotation of proteins of unknown function from proteomic datasets and whole proteomes.
]]></description>
<dc:creator><![CDATA[ Mathai, D., Schulze, S. ]]></dc:creator>
<dc:date>2026-07-11</dc:date>
<dc:identifier>doi:10.64898/2026.07.11.737882</dc:identifier>
<dc:title><![CDATA[ProtPen combines sequence- and structure-based approaches to facilitate protein function predictions on a proteome-wide scale]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-07-11</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.07.10.737775v1?rss=1">
<title>
<![CDATA[
GBZ-base and GAF-base: Indexed pangenome file formats 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.07.10.737775v1?rss=1
</link>
<description><![CDATA[
Motivation: Existing pangenome file formats are designed for batch processing. Graphs must be loaded into memory, and alignment files must be read sequentially. Indexed file formats that can be used directly from disk would be more appropriate for interactive applications. Results: We propose GBZ-base and GAF-base -- SQLite-backed file formats comparable to GBZ and GAF. GBZ-base supports efficient extraction of local subgraphs, and GAF-base lets us extract all alignments to the subgraph. Additionally, GAF-base is smaller than any other file format for sequence-to-graph alignments. Availability and implementation: https://github.com/jltsiren/gbz-base and https://crates.io/crates/gbz-base under the MIT license.
]]></description>
<dc:creator><![CDATA[ Siren, J., Paten, B., the Human Pangenome Reference Consortium ]]></dc:creator>
<dc:date>2026-07-11</dc:date>
<dc:identifier>doi:10.64898/2026.07.10.737775</dc:identifier>
<dc:title><![CDATA[GBZ-base and GAF-base: Indexed pangenome file formats]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-07-11</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.07.10.737545v1?rss=1">
<title>
<![CDATA[
ProtAug: An Empirical Investigation of pLM-Guided Data Augmentation for Protein Sequence Prediction Tasks 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.07.10.737545v1?rss=1
</link>
<description><![CDATA[
Protein language models (pLMs) offer great potential for protein sequence analysis, yet the scarcity of labeled data often limits their effectiveness in fine-tuning. Data augmentation is a promising remedy, but systematic evaluation of augmentation strategies for protein sequences remains limited, and the conditions under which augmentation confers downstream benefits are not well understood. In this paper, we systematically investigate pLM-guided substitution-based augmentation across seven protein prediction tasks. We propose ProtAug, a framework that leverages encoder-based (ESM-2) and autoregressive (ProtGPT2) pLMs to generate augmented sequences with user-controlled variation levels. Our investigation focuses on four questions: (Q1) whether pLM-synthesized sequences preserve more original signals than simpler methods, (Q2) to what extent augmentation improves prediction performance, (Q3) how variation levels affect downstream accuracy across tasks and models, and (Q4) whether biological plausibility is a necessary condition for achieving improvement. Our experimental results show that: (1) ProtAug_Esm generally preserves motifs and structural similarity better than simple substitution, often comparable to homology retrieval; (2) augmentation yields consistent but task-dependent improvements, with ProtAug_Esm achieving the best or second-best performance in 5 out of 7 tasks at 10% variation; (3) low-to-moderate variation levels (2-30%) perform best overall, although high-variation augmentation can benefit certain structure-related tasks; (4) the necessity of biological plausibility is task- and variation-dependent-while semantic preservation correlates with performance at low-to-moderate variation levels, improved generalization at high variation levels suggests that regularization effects, rather than label preservation, can also drive performance gains.
]]></description>
<dc:creator><![CDATA[ Chen, Z., Wang, R., Luo, Q. ]]></dc:creator>
<dc:date>2026-07-11</dc:date>
<dc:identifier>doi:10.64898/2026.07.10.737545</dc:identifier>
<dc:title><![CDATA[ProtAug: An Empirical Investigation of pLM-Guided Data Augmentation for Protein Sequence Prediction Tasks]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-07-11</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.07.08.737144v1?rss=1">
<title>
<![CDATA[
AptViralDB: A Repository of Experimentally Validated Antiviral Aptamers 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.07.08.737144v1?rss=1
</link>
<description><![CDATA[
In an era of increasing drug resistance, exploring alternative molecules is crucial for the efficient management and treatment of viral diseases. Nucleic acid aptamers have emerged as highly promising candidates due to their exceptional target specificity, low immunogenicity, and versatile mechanisms for viral blocking. This manuscript describes AptViralDB, a manually curated database providing comprehensive information on experimentally validated antiviral aptamers. It contains 1,768 entries of antiviral aptamers against 40 viral species and 104 molecular targets, compiled from literature and existing databases. Each entry provides detailed annotations, including sequence, aptamer type, target, chemical modifications, binding affinity, antiviral activity, stability, and cytotoxicity. We also provide predicted secondary structures and their corresponding minimum free energy (MFE) values. Additionally, a knowledge graph created using ArcadeDB/openCypher enables users to seamlessly explore connections among aptamers, viruses, molecular targets, and biological activities. Finally, the platform offers advanced search and browsing tools, BLAST-based sequence similarity searches, GC-content analysis, downloadable datasets, and REST API access to support computational applications. (https://webs.iiitd.edu.in/raghava/aptviraldb/).
]]></description>
<dc:creator><![CDATA[ Bajiya, N., Singh, S., Gahlot, P. S., Raghava, G. P. S. ]]></dc:creator>
<dc:date>2026-07-11</dc:date>
<dc:identifier>doi:10.64898/2026.07.08.737144</dc:identifier>
<dc:title><![CDATA[AptViralDB: A Repository of Experimentally Validated Antiviral Aptamers]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-07-11</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.07.07.736241v1?rss=1">
<title>
<![CDATA[
First-Trimester Non-Invasive Prediction of Preterm Birth Using Cell-Free DNA Fragmentomics 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.07.07.736241v1?rss=1
</link>
<description><![CDATA[
Objective. To develop and validate a cell free DNA (cfDNA) fragmentomic classifier for the early prediction of spontaneous preterm birth (PTB) using routine first trimester non-invasive prenatal testing (NIPT) data. Methods. A nested case-control study was conducted within a prospective multicenter Vietnamese cohort comprising 286 pregnancies, including 82 spontaneous PTB cases and 204 term controls. Maternal plasma cfDNA collected during routine first trimester NIPT (median gestational age, 12 weeks) was sequenced to a depth of approximately 20 million reads per sample. Five fragmentomic feature categories, including copy number alterations, end motif composition, nucleosome distance, fragment length, and joint fragment ength end motif were evaluated for PTB prediction. Machine learning classifiers were developed in a training cohort (n = 228, 65 PTB vs 163TB) and tested in a validation cohort (n = 58, 17 PTB vs 41 TB). Results. Among the five fragmentomic feature classes evaluated, 4 mer end motif (EM) profiles exhibited the most pronounced differences between PTB and term control samples. Consistent with these findings, the EM-based classifier demonstrated the highest discriminative performance in the validation cohort, achieving an AUC of 0.970 (95% CI, 0.912 to 1.000). At a specificity >90%, the model achieved a sensitivity of 94% (95% CI, 78 to 100%). Conclusion. These findings demonstrate that cfDNA EM signatures derived from routine first trimester NIPT can accurately identify pregnancies at risk of spontaneous preterm birth, without additional blood collection or sequencing, thereby extending the clinical utility of existing prenatal screening infrastructure.
]]></description>
<dc:creator><![CDATA[ Pham, M.-D. N., Phan, M.-T. T., Tran, N.-T., Vo, T.-S., Le, H.-T., Nguyen, T.-H. T., Nguyen, Q.-H. V., Ha, M.-T. T., Le, T. M., Hoang, D.-T. T., Huynh, K.-T. N., Nguyen, N. V., Nguyen, C. C., Bui, T. C., Nguyen, X. T., Le, S. V., Tran, V. D., Nguyen, M.-N. B., Nguyen, T. V., Nguyen, T.-A. T., Hoang, B. P., Nguyen, T. V., Nguyen, T.-A. T., Nguyen, T. T., Duong, T. D., Pham, C. H., Luong, K.-O. T., Dao, C. N., Hoang, K. V., Huynh, T.-T. T., Nguyen, K. M., Tran, S.-T. T., Tran, H. T., Nguyen, S. C., Tran, T. D., Nguyen, P. T. L., Pham, T. V., Pham, K. C., Thai, M. D., Do, T.-T. T., Dao, H. T., Va ]]></dc:creator>
<dc:date>2026-07-11</dc:date>
<dc:identifier>doi:10.64898/2026.07.07.736241</dc:identifier>
<dc:title><![CDATA[First-Trimester Non-Invasive Prediction of Preterm Birth Using Cell-Free DNA Fragmentomics]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-07-11</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.07.08.737203v1?rss=1">
<title>
<![CDATA[
Interactome Specialization Predicts Genome-Wide Binding-Site Degeneracy in Drosophila melanogaster Transcription Factors 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.07.08.737203v1?rss=1
</link>
<description><![CDATA[
Transcription factors (TFs) recognize short, degenerate DNA motifs that occur thousands of times throughout the genome, implying that binding specificity depends not only on DNA sequence but also on cellular context, including selective protein-protein interactions. Here, we tested whether a TF's integration into the physical TF interaction network predicts the degeneracy of its DNA-binding motif. Using 279 *Drosophila melanogaster* TFs with matched JASPAR position weight matrices, FlyBase expression profiles, and a physical protein interaction network derived from STRING v12.0 using only experimental and curated-database evidence, we quantified each TF's TF-module fraction and compared it with genome-wide predicted binding-site density across an independently constructed 18 Mb genomic sample. TF module fraction showed a significant positive association with binding-site density (partial r = 0.379, P = 5.6e-11) after controlling for motif information content, network degree, and literature bias. The relationship remained significant after excluding the homeodomain family, adding motif architecture controls, and applying multiple robustness analyses, including family-cluster bootstrapping and outlier-resistant correlation tests. Consistent with these findings, TFs formed a highly interconnected physical interaction network far exceeding degree-matched random expectation. Together, these results support a model in which DNA-recognition specificity and protein-interaction specificity represent complementary components of TF targeting: TFs embedded within TF-rich interaction modules tend to possess more degenerate DNA-binding motifs, whereas broadly acting network-generalist TFs rely on more information-rich sequence recognition. We also identify and correct a motif-length-dependent thresholding artifact that can obscure this relationship in genome-wide motif analyses.
]]></description>
<dc:creator><![CDATA[ Ponnambalam, A., Venkiteswaran Pottore, K. ]]></dc:creator>
<dc:date>2026-07-11</dc:date>
<dc:identifier>doi:10.64898/2026.07.08.737203</dc:identifier>
<dc:title><![CDATA[Interactome Specialization Predicts Genome-Wide Binding-Site Degeneracy in Drosophila melanogaster Transcription Factors]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-07-11</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.07.07.736532v1?rss=1">
<title>
<![CDATA[
SPARC: A Graph-based Optimization Framework for Directional Trajectory Reconstruction Across Ordered Single-Cell Conditions 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.07.07.736532v1?rss=1
</link>
<description><![CDATA[
Single-cell transcriptomics has enabled systematic profiling of cellular states across ordered biological contexts, including developmental stages, treatment phases, disease progression, and anatomical compartments. A central challenge is to reconstruct trajectories that respect the directionality imposed by biology or experimental design. Existing trajectory inference methods reconstruct cell-state progressions from latent-space geometry but do not enforce external biological ordering during graph construction, yielding biologically inadmissible transitions. An emerging paradigm of optimal-transport (OT) approaches partially addresses this limitation by incorporating experimental ordering into probabilistic state-to-state correspondences, yet their pairwise formulation cannot resolve whether a given state is an intermediate state or a terminal state along a multi-step progression. In multi-timepoint settings, OT typically estimates couplings only betweenadjacent timepoints and then chains these locally solved couplings to approximate long-range trajectories without a global optimization across all conditions simultaneously. Here we present SPARC, a graph-based optimization framework that quantifies similarity in a shared high-dimensional latent space and reconstruct directional trajectories under biological constraints. Global shortest-path optimization over this graph yields progression routes, from which SPARC derives path-based pseudotime identifies bottlenecks clusters, and detects gene temporal behavior. SPARC was evaluated across three complementary settings representing distinct trajectory-inference challenges. Its application to paired primary and lung metastatic osteosarcoma samples allows us to be the first to propose a cross-organ bone-like microenvironment hypothesis, in which osteoclastogenic signaling establishes a bone-like remodeling niche within the pulmonary metastatic lesion that promotes osteoclast differentiation and activity. The findings are independently recoverable in human osteosarcoma Visium HD spatial transcriptomics.
]]></description>
<dc:creator><![CDATA[ Wu, S., Walker, W. C., Martin, C., Yustein, J. T., Samee, M. A. H. ]]></dc:creator>
<dc:date>2026-07-11</dc:date>
<dc:identifier>doi:10.64898/2026.07.07.736532</dc:identifier>
<dc:title><![CDATA[SPARC: A Graph-based Optimization Framework for Directional Trajectory Reconstruction Across Ordered Single-Cell Conditions]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-07-11</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.07.07.737032v1?rss=1">
<title>
<![CDATA[
Single-cell RNA-seq reveals conserved and divergent cellular states across wound types and species 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.07.07.737032v1?rss=1
</link>
<description><![CDATA[
Chronic wounds, such as diabetic foot ulcers, fail to progress through the normal healing process and impose a significant burden on healthcare systems. While previous single-cell studies have characterized specific wound conditions, a unified understanding of the shared and distinct cellular landscapes across diverse wound microenvironments has been lacking. Therefore, we integrated over 500,541 cells from patients and mice across multiple wound conditions, including acute wound, diabetic foot ulcer, and venous ulcer as well as their healing outcome. Fibroblast-focused analysis identified a bifurcation in differentiation trajectories and identified STAT3 as a potential regulator of a reparative program in chronic wound. Furthermore, we discovered immune dysfunctions in non-healed chronic wounds, contrasting the quiescent memory-like T cells and TIMP1+ macrophages in healed chronic wounds with the exhausted T cells and foamy SPP1+ macrophage enriched in non-healed chronic wounds. Finally, we translated these results into a clinically applicable three-gene signature (CHI3L1, TIMP1, and SPP1) that accurately predicts chronic wound healing. To support wound biology community, we developed WoundSCAtlas, an interactive web resource for exploring diverse wound pathologies. In conclusion, this work provides a comprehensive and cross-species landscape of chronic wound healing, identifying conversed wound outcome-associated molecular programs, predictive biomarker, and interactive data resource.
]]></description>
<dc:creator><![CDATA[ Choi, D., Bakhtiari, M., Amin, A., Mann, J., Bhasin, S., Bhasin, M. ]]></dc:creator>
<dc:date>2026-07-11</dc:date>
<dc:identifier>doi:10.64898/2026.07.07.737032</dc:identifier>
<dc:title><![CDATA[Single-cell RNA-seq reveals conserved and divergent cellular states across wound types and species]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-07-11</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.07.07.737090v1?rss=1">
<title>
<![CDATA[
Multiscale harmonization and semantic integration of biomedical data enable biological insights through immersive exploration 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.07.07.737090v1?rss=1
</link>
<description><![CDATA[
The Human Reference Atlas (HRA) enables multiscale data exploration and visualization. We present "HRA: Powers of Ten," a virtual reality (VR) application for integrating, harmonizing, and visualizing data within the HRA Organ Gallery. It enables immersive navigation from a whole-body view of 81 organs to datasets across 5 organs, 5 assay types, and 4 spatial scales using a Multiscale Elevator System. The application, data, and code are available open-source.
]]></description>
<dc:creator><![CDATA[ Bueckle, A., Zhu, C., Wong, A. Y. H., Enninful, A., Miao, Y., Farzad, N., Pedersen, M., Mattison, C., Sloan, N., Mares, J., Xing, C., Herr, B. W., Khare, J., Kumar, Y. R., Parekh, K., Chavan, S., Luby, P., Patel, U., Hickey, J. W., Bader, G. D., Phatnani, H., Menon, V., Fan, R., Sorger, P., Snyder, M., Boerner, K. ]]></dc:creator>
<dc:date>2026-07-11</dc:date>
<dc:identifier>doi:10.64898/2026.07.07.737090</dc:identifier>
<dc:title><![CDATA[Multiscale harmonization and semantic integration of biomedical data enable biological insights through immersive exploration]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-07-11</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.07.07.737093v1?rss=1">
<title>
<![CDATA[
Locus-Level Transposable Element Profiling Resolves Division-Coupled Transcriptional Dynamics During Human Endoderm Specification 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.07.07.737093v1?rss=1
</link>
<description><![CDATA[
Background: Transposable elements (TEs) constitute nearly half of the human genome and are now recognized as significant contributors to mammalian gene regulatory networks. Despite this, most transcriptomic studies quantify TE expression at the subfamily level, which may obscure meaningful variation arising from individual insertion sites. Whether resolving TE expression to individual loci can reveal biologically distinct signals during stem cell differentiation has not been systematically characterized. Results: We re-analysed a published RNA-seq time course of FUCCI-h9 human embryonic stem cells differentiating into definitive endoderm (0 to 72 hour, seven time points, three division cycles, three biological replicates), quantifying expression in parallel at two complementary resolutions: TE subfamilies using TEtranscripts and individual TE loci using TElocal. The primary finding is that individual TE loci capture heterogeneous transcriptional responses at cell division boundaries that are entirely absent at the subfamily level. Within-division-state PC1 variance for TE loci was substantially elevated at the first division cycle (6.72; 95% bootstrap CI 0.63 - 8.48) compared with TE subfamilies (0.22; CI 0.04 - 0.34), with non-overlapping confidence intervals providing statistically robust support for the resolution advantage. Differential expression analysis identified over 18,000 dynamic TE loci across the time course, exceeding the 268 differentially expressed subfamilies, with alternating phases of silencing and reactivation resolved only at locus resolution. Differentially expressed TE loci were non-randomly enriched at superenhancers (fold enrichment: 9.1 - 23.5-fold; p < 0.001 by permutation), with peak overlap at 36-48 hours coinciding with the second cell division and endoderm commitment. ERV1-class elements, particularly HERVH-int, were the dominant contributors, and representative loci near the endoderm regulators MIXL1 and ID3 showed differentiation-induced RNA-seq signal within proximal superenhancer domains. Conclusions: TE loci exhibit heterogeneous transcriptional responses at cell division boundaries, a signal with non-overlapping bootstrap confidence intervals relative to TE subfamilies at the first division cycle that is entirely masked by subfamily-level aggregation. This division-boundary resolution advantage, together with permutation-confirmed enrichment of dynamic loci at superenhancers during a discrete 36-48 hour endoderm commitment window, supports broader adoption of locus-resolved TE quantification as a complement to conventional gene expression analysis in developmental genomics.
]]></description>
<dc:creator><![CDATA[ GAIRE, A., Kummerfeld, E., Aliferis, C., Wang, J. ]]></dc:creator>
<dc:date>2026-07-11</dc:date>
<dc:identifier>doi:10.64898/2026.07.07.737093</dc:identifier>
<dc:title><![CDATA[Locus-Level Transposable Element Profiling Resolves Division-Coupled Transcriptional Dynamics During Human Endoderm Specification]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-07-11</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.07.07.737051v1?rss=1">
<title>
<![CDATA[
Gene model for the ortholog of raptor in Drosophila grimshawi 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.07.07.737051v1?rss=1
</link>
<description><![CDATA[
Gene model for the ortholog of raptor in the D. grimshawi May 2011 (Agencourt dgri_caf1/DgriCAF1) Genome Assembly (GenBank Accession: GCA_000005155.1) of Drosophila grimshawi. This ortholog was characterized as part of a developing dataset to study the evolution of the Insulin/insulin-like growth factor signaling pathway (IIS) across the genus Drosophila using the Genomics Education Partnership gene annotation protocol for Course-based Undergraduate Research Experiences.
]]></description>
<dc:creator><![CDATA[ Lieser, B. C., Lose, B., Kiser, C. A., Butterfield, S., Laschober, L., Laskowski, L. F., Nielsen, J., Pulford, J., Thompson, J. S., Rele, C. P., Wittke-Thompson, J. K. ]]></dc:creator>
<dc:date>2026-07-11</dc:date>
<dc:identifier>doi:10.64898/2026.07.07.737051</dc:identifier>
<dc:title><![CDATA[Gene model for the ortholog of raptor in Drosophila grimshawi]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-07-11</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.07.07.736817v1?rss=1">
<title>
<![CDATA[
Structure-guided computational design and mechanistic understanding of the p95HER2-targeting NAZ-mAb antibody and its variants 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.07.07.736817v1?rss=1
</link>
<description><![CDATA[
Human epidermal growth factor receptor 2 (HER2) is an oncogenic receptor tyrosine kinase in breast cancer and other malignancies. A subset of HER2-positive tumours expresses 611-CTF-p95HER2, a tumour-specific, hyperactive truncated isoform associated with metastasis and treatment resistance that lacks most of the extracellular domain targeted by conventional HER2-directed antibodies. We previously developed NAZ-mAb (formerly known as Oslo-2), a monoclonal antibody against 611-CTF-p95HER2. Here, we describe a computational antibody-engineering workflow for designing variants of NAZ-mAb. Starting from the sequence alone, we modeled the NAZ-mAb-611-CTF-p95HER2 complex, generated a combinatorial mutational landscape using FoldX 5.0, and prioritized candidate variants using predicted interaction energy and developability criteria. Two variants representing distinct design strategies were selected for validation: an aromatic double mutant, NAZ-mAb v1 (L:S31W/L:H107W), and a conservative single mutant, NAZ-mAb v2 (L:S31M). Both variants were successfully expressed as recombinant IgGs; NAZ-mAb v2 achieved a five-fold higher recombinant expression yield than parental NAZ-mAb, while both variants retained antigen binding with a higher apparent signal than the parental antibody in indirect ELISA. However, Biacore two-state kinetic analysis revealed weaker affinities than the parental antibody (KD NAZ-mAb v1: 32.6 nM, NAZ-mAb v2: 9.45 nM vs. parental NAZ-mAb: 5.33 nM). These findings show that the computational workflow can generate experimentally tractable, antigen-engaging NAZ-mAb variants, while also highlighting the limitations of fixed-backbone interaction-energy ranking as a predictor of binding affinity and yield. This study provides a practical framework for computationally driven, developability-aware antibody optimization in the absence of experimental structural data.
]]></description>
<dc:creator><![CDATA[ Rawat, P., Kyte, J. A., Greiff, V., Dorraji, E. ]]></dc:creator>
<dc:date>2026-07-11</dc:date>
<dc:identifier>doi:10.64898/2026.07.07.736817</dc:identifier>
<dc:title><![CDATA[Structure-guided computational design and mechanistic understanding of the p95HER2-targeting NAZ-mAb antibody and its variants]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-07-11</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.07.07.737037v1?rss=1">
<title>
<![CDATA[
AGPI: An AI-Powered Genomic Pathogen Intelligence Platform for Integrated Classification, Visualization, and Therapeutic Targeting 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.07.07.737037v1?rss=1
</link>
<description><![CDATA[
Rapid and accurate pathogen detection remains a major challenge in modern bioinformatics, as existing tools are often fragmented and require multiple specialized workflows. We present AGPI (AI-powered Genomic Pathogen Intelligence), an integrated platform that combines genomic sequence classification, biological enrichment, three-dimensional structural visualization, and AI-guided therapeutic prioritization within a single interpretable pipeline. AGPI employs a hybrid convolutional Bidirectional Gated Recurrent Unit (BiGRU) architecture trained on DNA sequences from 40 pathogen classes spanning viruses, bacteria, fungi, and protozoan pathogens. The model achieved 99.61% validation accuracy and 94.90% accuracy on an independent held-out evaluation of 600 pathogen sequences following iterative refinement. As a proof of concept, AGPI correctly classified a Zika virus genome with 96.14% confidence, retrieved curated biological context from 245 peer-reviewed studies, and identified Ribavirin as a leading therapeutic candidate against the Zika NS5 polymerase through AI-guided molecular docking. Multi-metric ligand similarity analysis further differentiated candidate compounds according to their structural and pharmacological properties. These results demonstrate that integrated AI-driven genomic pipelines can accelerate pathogen characterization and therapeutic hypothesis generation while providing an accessible and interpretable framework for infectious disease surveillance and computational drug repurposing.
]]></description>
<dc:creator><![CDATA[ Goel, A., Mishra, P. ]]></dc:creator>
<dc:date>2026-07-11</dc:date>
<dc:identifier>doi:10.64898/2026.07.07.737037</dc:identifier>
<dc:title><![CDATA[AGPI: An AI-Powered Genomic Pathogen Intelligence Platform for Integrated Classification, Visualization, and Therapeutic Targeting]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-07-11</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.07.07.737010v1?rss=1">
<title>
<![CDATA[
Somatic mutation inference from single-cell transcriptomics: A survey in the esophagus 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.07.07.737010v1?rss=1
</link>
<description><![CDATA[
Human somatic tissues accumulate mutations during normal aging. Some of these affect cancer-associated driver genes and confer mutant progenitor cells a competitive advantage that leads to clonal expansions. The human esophageal epithelium exemplifies this phenomenon, becoming a dense mosaic of competing mutant clones by adulthood. However, the phenotypic consequence of those mutations and their possible role in carcinogenesis remains unknown. Novel bioinformatic tools for de novo mutant detection from single-cell transcriptomics (scRNA-seq) could potentially leverage on the wealth of publicly available data to help draw mutant cell phenotypes in vivo. In this study we test SComatic algorithm's ability to identify somatic mutations in the normal, polyclonal esophageal epithelium. We analyze a public scRNA-seq dataset from a human cohort with multiple esophageal samples per donor, and an independent study in mice subjected to experimental mutagenesis where samples have been re-sequenced for validation. These unconventional experimental designs allow us to control unspecificity. We observe scRNA-seq variant calling output is heavily affected by undesired technical artifacts and germline variants, which we are able to reduce following a customized series of rational filters that enrich in somatic mutations. Final candidate mutations are then used to reconstitute clonal lineages and map them to differentiation trajectories in the UMAP embeddings. We find low read depth and sparse cellular sampling favor detection of passenger mutations and hinder driver mutant phenotypic inferences. Altogether, we showcase current limitations of scRNA-seq-derived mutation calling, while we offer methodological indications that should be considered for future studies aimed at investigating mutant clone behavior in normal polyclonal tissues from single-cell transcriptomics.
]]></description>
<dc:creator><![CDATA[ Mendez-Alejandre, A., Gonzalez-Menendez, D., Vidal-Notari, S., Skrupskelyte, G., Rodriguez-Rodriguez, M., Ajith, H., Torralba, A. S., Alcolea, M. P., Piedrafita, G. ]]></dc:creator>
<dc:date>2026-07-11</dc:date>
<dc:identifier>doi:10.64898/2026.07.07.737010</dc:identifier>
<dc:title><![CDATA[Somatic mutation inference from single-cell transcriptomics: A survey in the esophagus]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-07-11</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.07.07.736993v1?rss=1">
<title>
<![CDATA[
Metagenomic contextualization of proteins with state space models 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.07.07.736993v1?rss=1
</link>
<description><![CDATA[
Since the early adoption of metagenomics (the culture-free sequencing of microbial community genomes) in 2011, sequence data has increased over 500-fold across ecosystems. This surge in data has outpaced reliable taxonomic and functional annotation, with over half of sequences lacking confident functional assignment. These unknown sequences limit our understanding of microbial processes central to planetary health and human health. Recent advances in genomic language modeling have made progress in the interpretation of metagenomics datasets. Most state-of-the-art models rely on transformer architectures, which limit the maximum sequence length and therefore capture only a fraction of assembled metagenomic sequences due to the quadratic scaling of attention. This prevents training and inference on sequences with broad context, including multiple coding and non-coding regions. To overcome this limitation, we propose leveraging new model architectures that scale linearly with sequence length, making them more suitable for modeling longer metagenomic sequences. Here, we introduce Nammu, a mixed-modality Mamba-based foundation model with 167M parameters trained on the OpenMetaGenomic (OMG) corpus. Nammu is a bidirectional encoder trained with a 20K context length using a curriculum strategy, first on 64M protein sequences and then on 32M mixed-modality metagenomic contigs. We compared Nammu to gLM2, a mixed-modality transformer also trained on OMG using 37% more tokens, using taxonomy inference on a marine dataset from the Critical Assessment of Metagenome Interpretation (CAMI). Nammu outperforms gLM2 at every taxonomic level. We further assessed function via KEGG Orthology prediction in deep-sea metagenome-assembled genomes, where Nammu outperforms gLM2 (150M). These results demonstrate improved performance.
]]></description>
<dc:creator><![CDATA[ Azbijari, N., Wynne, J. H., David, M., Thurber, A. R. ]]></dc:creator>
<dc:date>2026-07-11</dc:date>
<dc:identifier>doi:10.64898/2026.07.07.736993</dc:identifier>
<dc:title><![CDATA[Metagenomic contextualization of proteins with state space models]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-07-11</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.07.10.735823v1?rss=1">
<title>
<![CDATA[
AtlasLens: Metadata-centric exploration and analysis of single-cell atlases 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.07.10.735823v1?rss=1
</link>
<description><![CDATA[
Motivation: The rapid expansion of single-cell RNA sequencing (scRNAseq) atlases has generated datasets comprising millions of cells annotated with increasingly rich metadata, including tissue, cell type, disease status, sex, age, treatment, and temporal information. Biological questions frequently require simultaneous interrogation of multiple metadata dimensions, such as identifying specific cell populations within defined tissues, disease states, demographic groups, and time points. While existing interactive platforms facilitate visualization and analysis of scRNA-seq data, deep metadata-driven exploration and downstream analysis of atlas-scale datasets remain insufficiently supported. Results: We developed AtlasLens, an open-source R/Shiny application for interactive exploration of scRNA-seq datasets and integrated cellular atlases. AtlasLens enables iterative filtering across arbitrary metadata combinations, allowing users to define biologically meaningful cellular subsets and immediately perform downstream analyses. The platform integrates interactive visualization, differential expression analysis, Gene Ontology enrichment with redundancy reduction, temporal expression analysis, and context-dependent gene function profiling through GeneCOCOA. AtlasLens additionally records analysis history and automatically generates corresponding R code to enhance reproducibility. The application is distributed through Docker for simple local deployment, preserving data privacy and eliminating dependency-management challenges. We demonstrate AtlasLens using the Tabula Muris and a time-resolved whole-lung single-cell atlas of bleomycin-induced lung injury and fibrosis, highlighting its ability to support complex metadata-driven biological investigations. Availability: Source code is available at https://github.com/SchulzLab/AtlasLens
]]></description>
<dc:creator><![CDATA[ Ashrafiyan, S., Kosaretskii, I., Schulz, M. H. ]]></dc:creator>
<dc:date>2026-07-10</dc:date>
<dc:identifier>doi:10.64898/2026.07.10.735823</dc:identifier>
<dc:title><![CDATA[AtlasLens: Metadata-centric exploration and analysis of single-cell atlases]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-07-10</prism:publicationDate>
<prism:section></prism:section>
</item>
</rdf:RDF>
