<?xml version="1.0" encoding="UTF-8" ?>
<rdf:RDF xmlns:admin="http://webns.net/mvcb/" xmlns="http://purl.org/rss/1.0/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:prism="http://purl.org/rss/1.0/modules/prism/" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:syn="http://purl.org/rss/1.0/modules/syndication/">
<channel rdf:about="https://biorxiv.org">
<admin:errorReportsTo rdf:resource="mailto:biorxiv@cshlpress.edu"/>
<title>bioRxiv Subject Collection: Genomics Bioinformatics</title>
<link>https://biorxiv.org</link>
<description>
This feed contains articles for bioRxiv Subject Collection "Genomics Bioinformatics"
</description>

<items>
<rdf:Seq>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.10.03.756294v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.10.03.756391v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.10.03.754609v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.10.02.756312v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.10.02.756019v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.10.02.756304v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.10.05.756038v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.10.02.756305v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.10.03.756430v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.10.08.757792v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.10.04.756577v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.10.07.756157v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.10.07.757202v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.10.02.756336v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.10.04.756527v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.10.02.756356v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.10.03.756406v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.10.06.756962v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.10.08.757737v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.10.08.756856v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.10.08.757705v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.10.08.757135v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.10.08.738042v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.10.08.757329v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.10.07.757445v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.10.08.757597v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.10.07.757366v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.10.02.756377v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.10.02.756289v1?rss=1"/>
<rdf:li rdf:resource="https://www.biorxiv.org/content/10.64898/2026.10.08.757551v1?rss=1"/>
</rdf:Seq>
</items>
<prism:eIssn/>
<prism:publicationName>bioRxiv</prism:publicationName>
<prism:issn/>

<image rdf:resource=""/>
</channel>
<image rdf:about="">
<title>bioRxiv</title>
<url>https://www.biorxiv.org/sites/default/files/bioRxiv_article.jpg</url>
<link>https://www.biorxiv.org</link>
</image>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.10.03.756294v1?rss=1">
<title>
<![CDATA[
Spatiotemporal Heterogeneity of HPV Integration Architectures Drive Oncogenic Features of Oropharyngeal Cancer 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.10.03.756294v1?rss=1
</link>
<description><![CDATA[
Viral integration is common in human papillomavirus-positive oropharyngeal squamous cell carcinoma (HPV+ OPSCC), but its structural and clinical significance remains unclear. We integrated targeted capture sequencing of 197 HPV+ OPSCCs with long-read nanopore sequencing of 17 tumors and models to characterize HPV-host rearrangements. We developed a novel quantitative informatics framework that classifies integration structures as Type1 (simple, single/low-copy) or Type2 (clustered, high-copy, rearranged). HPV integration occurred in 83% of tumors, with recurrent hotspots including TP63, MYC, CD274, TRAF2, and KMT2C. Type2 events were enriched for HPV E2 disruption, focal human amplifications, preferential transcription, and noncanonical HPV-host RNA fusion isoforms. Type2 tumors showed increased HPV oncogene expression and activation of MYC, NF-{kappa}B/TNF, PI3K-AKT-mTOR, Hippo/YAP-TAZ, and epithelial stemness programs. Multi-region and longitudinal analyses revealed intratumoral heterogeneity and clonal selection of Type2 architectures during progression, supporting integration structure as a potential marker of aggressive behavior and risk stratification.
]]></description>
<dc:creator><![CDATA[ Gu, W., Bhangale, A. D., Du, X., Deng, X., Gensterblum-Miller, E. U., Currie, J., Brummel, C. V., Nyabashi, V., Buchakjian, M., Casper, K. A., Chinn, S. E., Forner, D., Malloy, K. M., Mierzwa, M. L., Prince, M. E., Shah, J. L., Shuman, A. G., Stucken, C. L., Swiecicki, P. L., Worden, F. P., Yalamanchi, P., McHugh, J. B., Kuhs, K. A., Boyle, A. P., Jiang, H., Neal, M. E., Spector, M. E., Mills, R. E., Brenner, C. J. ]]></dc:creator>
<dc:date>2026-10-10</dc:date>
<dc:identifier>doi:10.64898/2026.10.03.756294</dc:identifier>
<dc:title><![CDATA[Spatiotemporal Heterogeneity of HPV Integration Architectures Drive Oncogenic Features of Oropharyngeal Cancer]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-10</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.10.03.756391v1?rss=1">
<title>
<![CDATA[
Ab initio CryoEM maps minimizing preferred orientation problems in heterogeneous mixtures: ReconSIREN 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.10.03.756391v1?rss=1
</link>
<description><![CDATA[
Cryo-Electron Microscopy (CryoEM) Single Particle Analysis (SPA) is pivotal for elucidating protein structures and their conformational dynamics. Several critical steps are required to obtain accurate maps. First, the textit{ab initio} relative orientations and shifts between images must be determined. Second, the common issue of specimens presenting a preferred orientation should be adequately addressed. Third, all of the above must account for particle flexibility and sample heterogeneity. Additionally, these steps are interrelated, making the search for all these parameters a real challenge when considering convergence robustness. In this context, we present ReconSIREN, a novel deep learning framework that addresses the key steps outlined before in an textit{ab initio} manner while greatly minimizing one of the most important practical burdens in SPA, namely the errors introduced in the CryoEM maps by the existence of preferred orientations; the latter is achieved by a change of the optimization approach that avoids the so-called ``attraction problem'', and by the introduction of a new type of geometrical priors. It does so fully textit{ab initio}, using deep learning training/inference steps without resorting to more traditional optimization approaches, and generating highly accurate initial reconstructions. We present maps as both homogeneous (consensus) maps and a full conformational landscape for analyzing flexibility and heterogeneity. Thanks to its Mixture-of-Experts (MoE) architecture, ReconSIREN distributes the estimation workload across independent experts, improving optimization efficiency and simplifying the process.
]]></description>
<dc:creator><![CDATA[ Herreros, D., Sorzano, C. O. S., Carazo, J. M. ]]></dc:creator>
<dc:date>2026-10-10</dc:date>
<dc:identifier>doi:10.64898/2026.10.03.756391</dc:identifier>
<dc:title><![CDATA[Ab initio CryoEM maps minimizing preferred orientation problems in heterogeneous mixtures: ReconSIREN]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-10</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.10.03.754609v1?rss=1">
<title>
<![CDATA[
Matrix-Mechanical Genetic Susceptibility and Reduced Inflammatory and Immediate-Early Programs in Keratoconus 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.10.03.754609v1?rss=1
</link>
<description><![CDATA[
Purpose. To determine which proposed keratoconus (KC) mechanisms carry its genetic susceptibility and how it relates to corneal structure, and which describe its corneal tissue. Methods. Six mechanism modules and two attribution sets (transforming growth factor-{beta} [TGF-{beta}] and canonical type-2 atopy) were tested for enrichment in three independent European KC genome-wide association studies (GWAS) and in GWAS of central corneal thickness (CCT) and corneal resistance factor (CRF), with colocalization at matrix loci. Tissue expression was pooled across nine arms from eight studies with control-source meta-regression; activator protein-1 (AP-1)/immediate-early regulatory activity was scored in 107 donors, spatial expression in 21, and serum responses in human and rabbit stromal cells. Results. Extracellular matrix (ECM) was the only module enriched in all three KC sources (P [&le;] 5.46 x 10-) and for CCT, CRF, and thickness-independent CRF (P [&le;] 2.07 x 10-). COL5A1 shared causal variants with CCT in all three sources (posterior probability 0.985-0.991), and FNDC3B and COL1A1 with CRF or its thickness-independent component (0.961-0.986). Pooled, KC corneas had lower general and corneal-local inflammatory and stress/immediate-early expression (log2 fold-change -0.97, -0.98, -0.82; P [&le;] 0.049). In whole-cornea and stromal tissue, ECM expression reversed between postmortem-donor and pathological controls (log2 fold-change +0.36, -2.34), whereas inflammatory and TGF-{beta} programs were lower against both. AP-1/immediate-early regulatory activity was lower in four donor cohorts (Hedges' g -1.50) and stress/immediate-early expression in two spatial cohorts; serum raised immediate-early readouts in human and rabbit cells. Conclusions. KC combines matrix-mechanical genetic susceptibility with a low-inflammatory, low-immediate-early corneal state whose matrix expression direction depends on the control source. Because this susceptibility is genetically shared with corneal thickness and resistance, interventions on the inherited disease are judged by corneal structure and mechanics.
]]></description>
<dc:creator><![CDATA[ Jiang, H., Gao, F., Zhang, W., Jie, Y., Li, Y., Jiang, Y. ]]></dc:creator>
<dc:date>2026-10-10</dc:date>
<dc:identifier>doi:10.64898/2026.10.03.754609</dc:identifier>
<dc:title><![CDATA[Matrix-Mechanical Genetic Susceptibility and Reduced Inflammatory and Immediate-Early Programs in Keratoconus]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-10</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.10.02.756312v1?rss=1">
<title>
<![CDATA[
Winnow-tax: sensitive and precise taxonomic profiling of low-coverage organisms in shotgun metagenomes 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.10.02.756312v1?rss=1
</link>
<description><![CDATA[
Accurate species-level profiling of shotgun metagenomes becomes difficult when an organism is represented by only sparse sequence coverage. Per-read k-mer classifiers can retain sensitivity under these conditions, but short reads from conserved or shared genomic regions can generate false-positive calls among closely related references. More conservative marker-gene and genome-containment approaches generally provide stronger species-level specificity but may lose sensitivity as genomic sampling becomes sparse. We present winnow-tax, a two-stage taxonomic profiling pipeline that separates sensitive candidate nomination from genome-level confirmation. Candidate species are nominated using Sylph, Kraken2, and a branch-rescue procedure, then reads are competitively recruited to a sample-specific reference set. Species presence is evaluated using read support, observed genome breadth, and a Lander-Waterman-based breadth ratio that compares observed breadth with the breadth expected at the measured mean depth. In a controlled synthetic community of 82 genomes with human DNA background, winnow-tax maintained a stronger precision-sensitivity balance than Kraken2/Bracken, Sylph, and MetaPhlAn 4 as genome coverage decreased, with its advantage concentrated near the low-coverage detection boundary. In the CAMI III Toy Longitudinal Human Gut benchmark, winnow-tax had higher sensitivity than Sylph (0.803 versus 0.672), lower precision (0.923 versus 0.970), and a modestly but significantly higher F1 score (0.858 versus 0.793). In a clinical enteric stool cohort with culture/PCR reference testing, winnow-tax achieved the best composite performance for Salmonella, detecting 37 of 48 composite-positive samples with one false-positive call. Together, these results support a profiling strategy in which weak taxonomic signals are retained during candidate generation but require genomic evidence distributed as broadly as expected for their sequencing depth before species presence is accepted.
]]></description>
<dc:creator><![CDATA[ Sun, S., Fodor, A. A. ]]></dc:creator>
<dc:date>2026-10-10</dc:date>
<dc:identifier>doi:10.64898/2026.10.02.756312</dc:identifier>
<dc:title><![CDATA[Winnow-tax: sensitive and precise taxonomic profiling of low-coverage organisms in shotgun metagenomes]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-10</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.10.02.756019v1?rss=1">
<title>
<![CDATA[
Point-Process Modelling of Cell-Type Interaction Structure in CosMx Glioma Data for Hypothesis-Driven Tumor Inference from scRNA-seq 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.10.02.756019v1?rss=1
</link>
<description><![CDATA[
Glioblastoma (GBM) exhibits marked spatial heterogeneity that is lost after tissue dissociation for single-cell RNA sequencing. Here, we developed a spatial-statistical framework to characterize tumor-microenvironment organization and derive molecularly predictable spatial phenotypes from CosMx data. Eight GBM specimens comprising 2,427,362 cells, 15 cell populations, and four malignant states were analyzed, revealing substantial inter-patient differences in cellular composition and density. Fine finite-element meshes best preserved continuous spatial intensity and supported whole-tissue Log-Gaussian Cox Process modelling. Immune populations, including macrophages, microglia, monocytes, cDCs, neutrophils, and T cells, showed spatial attraction toward malignant states. These models generated cell-level measures of conditional intensity and tumor-proximity probability. Transcript abundances and metabolic pathway scores served as the sole predictive features to model these spatial phenotypes. Transcriptomic representations outperformed metabolic pathway scores. Leave-one-patient-out validation reduced predictive performance, and patient-wise harmonization failed to restore cross-patient transportability. Spatial outcome distributions varied between specimens, indicating that molecular-spatial relationships are strongly conditioned by patient-specific tumor architecture. Predicted interaction scores stratified transcriptional subclusters within cell types in external non-spatial single-cell datasets. This framework links spatial point-process phenotypes to molecular states and provides a basis for inferring spatial context from non-spatial single-cell data. Generalizing these models requires larger and more diverse spatial cohorts.
]]></description>
<dc:creator><![CDATA[ Carcanholo, F. P., Cassiano, M. H. A., Panepucci, E. M., Laurini, M. P., Malta, T. M. ]]></dc:creator>
<dc:date>2026-10-10</dc:date>
<dc:identifier>doi:10.64898/2026.10.02.756019</dc:identifier>
<dc:title><![CDATA[Point-Process Modelling of Cell-Type Interaction Structure in CosMx Glioma Data for Hypothesis-Driven Tumor Inference from scRNA-seq]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-10</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.10.02.756304v1?rss=1">
<title>
<![CDATA[
From Sparse Visits to Continuous Molecular Trajectories with Flow Matching 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.10.02.756304v1?rss=1
</link>
<description><![CDATA[
Longitudinal transcriptomic studies often capture only a few molecular snapshots of each patient over time. This makes it challenging to understand how patients are progressing, compare them when progression rates differ, and capture molecular changes between observed visits. We introduce CohortFM, a flow-matching framework that reconstructs continuous patient-specific molecular trajectories from sparse longitudinal observations and organizes them through an integrated mixture model into a compact set of cohort-level progression patterns. CohortFM separates molecular progression from calendar time, places patients on a shared progression scale, uses observed intermediate states to guide the learned trajectories, and maps the resulting patterns back to real patient samples for biological interpretation. We evaluated CohortFM across COVID-19, sepsis, influenza, and tuberculosis. Despite different sampling patterns, the method generalizes well and consistently produces meaningful trajectories. These trajectories capture biological differences between patients that cannot be explained by the observed samples alone. Compared with state-of-the-art trajectory baselines, CohortFM produces more biologically distinct trajectories and a more complete representation of the patient population. Its finer progression scale also reveals biological programs and temporal dynamics that are not captured by visit-level analysis. We further evaluated interpolation and true forward prediction to test whether CohortFM can accurately recover unobserved states. CohortFM remains competitive with strong interpolation baselines and is among the strongest learned predictors in the harder forecasting setting. Ablation experiments further show that flow matching is the core learning mechanism and that our design choices improve continuous trajectory learning. Our results support CohortFM as a practical framework for turning sparse molecular snapshots into continuous, biologically interpretable trajectories. By going beyond visit labels, CohortFM makes it possible to study disease progression and predict future molecular states even when patients are sampled only a few times.
]]></description>
<dc:creator><![CDATA[ Passban, P., Gupta, S., Guan, B., Roosta, T. ]]></dc:creator>
<dc:date>2026-10-10</dc:date>
<dc:identifier>doi:10.64898/2026.10.02.756304</dc:identifier>
<dc:title><![CDATA[From Sparse Visits to Continuous Molecular Trajectories with Flow Matching]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-10</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.10.05.756038v1?rss=1">
<title>
<![CDATA[
RiboFlow v2: a configurable ribosome profiling pipeline reveals measurements sensitive to read-mapping methodology 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.10.05.756038v1?rss=1
</link>
<description><![CDATA[
Ribosome profiling experiments capture a snapshot of translation by sequencing ribosome-protected fragments (RPFs). Alignment to a well-curated transcriptome captures most RPF signal, but can miss ribosome activity when expressed exons are absent from the reference. Genome alignment of RPFs expands the detectable sequence space, but can expose ambiguity between homologous loci. Existing workflows often couple reference choice with the handling of ambiguously mapped reads, making it difficult to isolate the effects of either decision on downstream analyses. Here, we developed RiboFlow v2 by adding genome alignment and configurable handling of multimapping reads to the existing transcriptome workflow. Using 24 human ribosome profiling libraries and their matched RNA-seq data, we found that nucleotide-level coverage at matched coding-sequence positions and translation efficiency (TE) were broadly concordant between alignment strategies under stringent alignment filters. Even under these conditions, individual genes showed substantial differences in coverage and TE estimates. These differences included coverage in alternative exons absent from selected transcripts and ambiguous alignments between protein-coding genes and processed pseudogenes. Clustering genes by read-assignment patterns identified groups associated with pseudogene prevalence and highlighted shared exonic sequence among reference transcripts as a source of alignment discrepancies. RiboFlow v2 extends the analysis beyond selected reference transcripts and enables identification of genes whose ribosome coverage and abundance estimates depend on reference composition and how ambiguously mapped reads are handled.
]]></description>
<dc:creator><![CDATA[ Nguyen, D., Cenik, C. ]]></dc:creator>
<dc:date>2026-10-10</dc:date>
<dc:identifier>doi:10.64898/2026.10.05.756038</dc:identifier>
<dc:title><![CDATA[RiboFlow v2: a configurable ribosome profiling pipeline reveals measurements sensitive to read-mapping methodology]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-10</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.10.02.756305v1?rss=1">
<title>
<![CDATA[
Pooled progeny sequencing effectively estimates polyandry in species with large families 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.10.02.756305v1?rss=1
</link>
<description><![CDATA[
Polyandry, multiple mating by females, both benefits females by increases genetic variation within families and provides an opportunity for sexual selection through sperm competition and cryptic female choice. Estimating polyandry in species with large clutches is challenging because genotyping all progeny individually can be impractical. We present a likelihood-based framework for estimating the number of sires and their relative contributions from the maternal genotype and pooled sequencing of entire progeny pools. The method evaluates the probability of observed pooled offspring read counts under alternative paternal genotype mixtures, accounting for Mendelian segregation and sampling variation from finite sequencing depth. We implement this framework for both biallelic SNPs and short-read mini-haplotype markers constructed from closely linked SNPs, which provide multi-allelic information comparable to traditional microsatellite markers. Simulations show that this approach accurately recovers sire number and sire contributions across a range of conditions. We then applied the method to 275 wild-caught Drosophila melanogaster females and their offspring from a Kansas population using 3,053 biallelic SNPs and 239 four-allele mini-haplotype markers. Estimates from the two marker sets were highly concordant, and combined analyses indicate that 57% of families were sired by a single male. These results demonstrate that pooled progeny sequencing provides an efficient genomic approach for estimating polyandry in species with large families.
]]></description>
<dc:creator><![CDATA[ Farmer, T., Bentz, A., Macdonald, S. J., Unckless, R., Kelly, J. ]]></dc:creator>
<dc:date>2026-10-10</dc:date>
<dc:identifier>doi:10.64898/2026.10.02.756305</dc:identifier>
<dc:title><![CDATA[Pooled progeny sequencing effectively estimates polyandry in species with large families]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-10</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.10.03.756430v1?rss=1">
<title>
<![CDATA[
Systematic benchmark of demultiplexing workflows for direct RNA sequencing 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.10.03.756430v1?rss=1
</link>
<description><![CDATA[
Direct RNA nanopore sequencing offers the capacity to directly profile native RNAs, enabling detection of diverse epitranscriptomic modifications based on current changes as the RNA passes through a pore. Innovations in chemistry, hardware and bioinformatic tools have increased sequencing capacity and accuracy, creating demand for sequencing multiple samples within individual runs. Two recent solutions enable barcoding and demultiplexing of samples, but independent comparisons of their performance is limited. Here we benchmark WarpDemuX and SeqTagger using orthogonal ground truths derived from species identity and matched IVT spike-ins across four RNA004 flow cells. Both achieve high demultiplexing precision but differ in their trade-offs in precision and retention. SeqTagger provided strong performance at its recommended settings, whilst WarpDemuX performance varied by model, barcode and read quality. IVT-derived score cutoffs allowed dataset-specific calibration and increased WarpDeMux assignment yield, although precision achieved on IVTs was not consistently reproduced in species reads. These results support matched IVTs as internal controls for guiding choice of demultiplexing settings, while demonstrating the need to validate demultiplexing precision in biological RNA populations.
]]></description>
<dc:creator><![CDATA[ Bryson, J. W., Johansson, L. B., George, A., Chen, J., Auxillos, J. Y., Rennie, S. ]]></dc:creator>
<dc:date>2026-10-10</dc:date>
<dc:identifier>doi:10.64898/2026.10.03.756430</dc:identifier>
<dc:title><![CDATA[Systematic benchmark of demultiplexing workflows for direct RNA sequencing]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-10</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.10.08.757792v1?rss=1">
<title>
<![CDATA[
A chromosome-level genome assembly of the chilli thrips Scirtothrips dorsalis Hood (Thysanoptera: Thripidae) 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.10.08.757792v1?rss=1
</link>
<description><![CDATA[
Thrips are known for being pests of many agricultural crops, as well as vectors of plant viruses. Yet despite their economic importance, genomic resources for this group remain limited, and no chromosome-level genome has been available for the genus Scirtothrips. One such species is the chilli thrips, Scirtothrips dorsalis Hood (Thysanoptera: Thripidae). This invasive and highly polyphagous pest is established across much of the world and transmits several tospoviruses. Genomic resources would provide a new means to understand the traits that make thrips difficult to manage. Here, we present a chromosome-level genome assembly for S. dorsalis. The assembly is 266.5 Mb, with 96.9% of the sequence anchored in 16 chromosome-scale scaffolds. It contains 20,557 genes, and repetitive elements account for 28.8% of the sequence, which is comparable to those of other sequenced Thripidae. A species tree based on 3,019 single copy orthologs placed S. dorsalis as sister to Scirtothrips hansoni, with Scirtothrips most closely related to Frankliniella, Megalurothrips, and Odontothrips. The mitochondrial genome is a single circular molecule of 15,207 bp in which the sequence of the published mini circle is integrated, in contrast to the bipartite mitogenome previously reported for the same cryptic species.
]]></description>
<dc:creator><![CDATA[ Busuulwa, A., Liesenfelt, T., Mongue, A. J., Lahiri, S. ]]></dc:creator>
<dc:date>2026-10-10</dc:date>
<dc:identifier>doi:10.64898/2026.10.08.757792</dc:identifier>
<dc:title><![CDATA[A chromosome-level genome assembly of the chilli thrips Scirtothrips dorsalis Hood (Thysanoptera: Thripidae)]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-10</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.10.04.756577v1?rss=1">
<title>
<![CDATA[
Characterization of artificial riboswitches for Coxsackievirus B3 detection using machine learning 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.10.04.756577v1?rss=1
</link>
<description><![CDATA[
With the increase of recycled water to offset water demand, the potential possibility to spread contagious RNA viruses, such as Coxsackievirus B3, increases. However, detection of viral particles remains challenging because of low viral concentrations in wastewater and high mutation rates of the RNA virus. Robust monitoring is needed with low cost and low infrastructure technologies to increase accessibility of monitoring technologies worldwide. Novel viral detection methods have been developed, based on synthetic riboswitches that bind to the target virus and trigger a reporter gene, thus amplifying the detection signals. Such monitoring technologies can leverage machine learning to optimize the candidate nucleic acid sequences used for detection. To support the design of effective riboswitches, we present a machine learning model for classifying riboswitch performance, integrating RNA sequence data with secondary structural features based on free energy calculations and parameters that represent single strandedness. This model uses a sparsely gated Mixture of Experts (MoE) architecture to route sequence and thermodynamic features to specialized experts, achieving strong generalization performance across cross-validation and held-out testing. When evaluated against baseline Decision Tree, Random Forest, Gaussian Naive Bayes, and Dense Neural Network classifiers, the MoE model demonstrated near-zero classification errors. Furthermore, post-hoc feature importance and k-mer analyzes reveal that near-perfect predictive performance in silico is strongly driven by sequence constructs in addition to generalizable folding mechanics. An ablation study of numerical feature and sequence inputs showed that while the Dense Neural Network also achieved high accuracy across all ablation conditions, the MoE architecture was retained as the primary model because its specialized subnetworks do not require all model parameters to be activated offering a potential computational advantage.
]]></description>
<dc:creator><![CDATA[ Auyong, J., Long, H. A., Hu, S., Jacob, J., Chan, K., Gupta, B., Mengistu, A., Tokuhara, M., Khatib, L., Andreopoulos, W. B. ]]></dc:creator>
<dc:date>2026-10-09</dc:date>
<dc:identifier>doi:10.64898/2026.10.04.756577</dc:identifier>
<dc:title><![CDATA[Characterization of artificial riboswitches for Coxsackievirus B3 detection using machine learning]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-09</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.10.07.756157v1?rss=1">
<title>
<![CDATA[
DEPICT enables sample-centric therapeutic prioritization across evolving and spatially heterogeneous tumor states 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.10.07.756157v1?rss=1
</link>
<description><![CDATA[
Accurate treatment selection remains a major challenge in precision oncology because tumors with similar genomic alterations can respond differently to therapy. Most drug-response models evaluate individual drugs or drug-sample pairs, whereas clinical decision-making requires therapies to be prioritized within each tumor context. Here we present DEPICT (Drug Efficacy Prediction via Integrated ConText), a sample-centric deep learning framework that reframes drug-response prediction as the relative prioritization of therapies within an individual tumor context. DEPICT integrates transcriptomic and mutational features with prior alteration-therapy relationship and was trained on 231,880 cell line-drug pairs comprising 941 cancer cell lines and 289 compounds. Without retraining, DEPICT retained predictive performance in independently profiled cell lines, experimentally induced drug-resistant states and matched primary lung tumors and patient-derived organoids. Representations learned exclusively from bulk pharmacogenomic data further captured progressive drug-response states at single-cell resolution and outperformed existing single-cell drug-response inference methods. Applied to spatial transcriptomic data, DEPICT resolved clone-associated therapeutic heterogeneity within individual lung tumors, and predicted regional differences were recapitulated in matched multiregional organoid drug-response assays. These findings establish sample-centric therapeutic prioritization as a transferable strategy for modeling dynamic and spatially heterogeneous drug vulnerabilities across biological scales.
]]></description>
<dc:creator><![CDATA[ Ma, S., Yan, S., Wen, X., Yu, C., Ma, H., Teng, W., Zhou, X., Gao, J., Wang, S., Li, M., Han, K., Huang, Y., Wu, N., Liu, B., Zheng, C. ]]></dc:creator>
<dc:date>2026-10-09</dc:date>
<dc:identifier>doi:10.64898/2026.10.07.756157</dc:identifier>
<dc:title><![CDATA[DEPICT enables sample-centric therapeutic prioritization across evolving and spatially heterogeneous tumor states]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-09</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.10.07.757202v1?rss=1">
<title>
<![CDATA[
Error correction of SARS-CoV-2 genomic sequences using a phylogenetic prior 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.10.07.757202v1?rss=1
</link>
<description><![CDATA[
Genotype calling of sequencing data has improved over recent years thanks to technological and statistical breakthroughs. One such advance is the use of imputation-based genotyping methods in human genetics. For haploid and mostly non-recombining organisms, a natural equivalent is to leverage phylogenetic signal for imputation-based genotype calling. SARS-CoV-2 is one such case, and the use of improved genotyping methods in this species is particularly pertinent given the well-documented systematic problems with read alignment and assembly in rapidly evolving regions of its genome, motivating the need for statistically principled methods for genotype correction that complement existing community efforts to curate sequencing data. We present a statistical framework to combine phylogenetic prior probabilities with the information from aligned reads, and demonstrate its utility in correcting consensus calls made on short-read, viral sequencing data. We derive phylogenetically-informed posterior probabilities based on genome placement in a reference phylogeny, and read likelihoods calculated from machine-reported base quality scores. We further modify the prior at each site by estimating a maximum-likelihood scaling parameter jointly with reads at all other sites. We find that the phylogenetically-informed base calling method stays highly accurate across simulated depth and error conditions. In real datasets, the method vastly improves call rate while still maintaining similar accuracy as reference-based assembly methods. In problematic genomic regions, the phylogenetically-informed method confers both accuracy and call rate gains. Accuracy and call rate do not stratify with sequencing center and primer scheme used, suggesting broad applicability across sequencing technologies and potentially diverging sequences. Our results in real and simulated data demonstrate that phylogenetically-informed genotyping correction results in more complete sequences, achieving a higher call rate without compromising accuracy. In cases of low mean depth, sparse coverage, or difficult-to-sequence regions that may still retain a low call rate, the method improves the accuracy of genotype calls.
]]></description>
<dc:creator><![CDATA[ Arniella, M., Pipes, L., Nielsen, R. ]]></dc:creator>
<dc:date>2026-10-09</dc:date>
<dc:identifier>doi:10.64898/2026.10.07.757202</dc:identifier>
<dc:title><![CDATA[Error correction of SARS-CoV-2 genomic sequences using a phylogenetic prior]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-09</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.10.02.756336v1?rss=1">
<title>
<![CDATA[
Performance comparison of probe-capture enrichment sequencing of viruses in wastewater influent and settled solids from the California Central Valley 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.10.02.756336v1?rss=1
</link>
<description><![CDATA[
Wastewater sequencing using probe-capture enrichment of public health-relevant viruses is a powerful, non-intrusive approach to monitor infectious diseases, but performance depends on sample type and processing approach. We compared two processing pipelines from wastewater influent samples collected at three timepoints across three wastewater treatment facilities. Each sample was split for parallel processing, with viruses concentrated (1) directly from whole influent using nano-sized affinity particles and (2) by settling and dewatering the solids fraction. Combined DNA/RNA extracts were prepared into libraries using a pre-commercial Illumina protocol, then probe-capture enrichment was performed with Illumina Viral Surveillance Panel (VSP) v1 (66 targets) and v2 (242 targets). Sequencing depth was increased for VSP v2 to accommodate the expanded target set. On average, 160 +/- 146 viral species were detected per sample, with wastewater fraction (sample type) explaining 19% of the variance in viral abundances. Solids contained a higher abundance of DNA viruses compared to RNA viruses, while influent exhibited the inverse. Metagenomic analysis of the VSP v2 sequencing data detected more viral species in influent (median 358) than in solids (median 184), but probe-capture enrichment produced more comparable target-virus readouts across fractions, supporting the use of either fraction for wastewater-based epidemiology.
]]></description>
<dc:creator><![CDATA[ Lavon, A., Pender, J. M., Olson, R., Chatterjee, T., Naughton, C. C., Prashar, P., Kuersten, S., Eisen, J. A., Whiteson, K., Bischel, H. N. ]]></dc:creator>
<dc:date>2026-10-09</dc:date>
<dc:identifier>doi:10.64898/2026.10.02.756336</dc:identifier>
<dc:title><![CDATA[Performance comparison of probe-capture enrichment sequencing of viruses in wastewater influent and settled solids from the California Central Valley]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-09</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.10.04.756527v1?rss=1">
<title>
<![CDATA[
Genotypic Responses of Rice to Ammonium and Nitrate under Contrasting Water Regimes: Agronomic, Biochemical, and Gene-Expression Analyses 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.10.04.756527v1?rss=1
</link>
<description><![CDATA[
Considering the increasing water and labor constraints in rice production, direct-seeded rice is becoming increasingly important. However, direct-seeded rice is vulnerable to water stress and inefficient nutrient use. Nitrogen use efficiency in rice is generally low, and improving nitrogen acquisition and assimilation is important for sustainable productivity. In this study, 48 recombinant inbred lines were evaluated under ammonium and nitrate nutrition in irrigated and rainfed environments across two cropping seasons. Agronomic, physiological, biochemical and molecular responses were assessed. Significant effects of genotype, nitrogen form, environments and their interactions were observed for several traits. Rainfed conditions reduced grain and biological yield compared with irrigated conditions. Ammonium produced a higher mean grain yield than nitrate by approximately 26.2 g m-2 under irrigated conditions and 154.2 g m-2 under rainfed conditions, although the responses varied among genotypes. Biological yield and SPAD values were generally higher with ammonium, whereas total chlorophyll showed season and environment-dependent responses. Glutamine synthetase activity was highest under ammonium treatment, while glutamate synthase activity was higher under nitrate treatment. Grain yield showed strong positive correlations with biological yield, but correlations between grain yield and GS or GOGAT activity were not significant. Expression of OsGln1.2, OsGln2 and OsGlt1 varied according to genotype, tissue and nitrogen treatment. Overall, the results demonstrate that responses to nitrogen forms are genotype- and environment-dependent, and that ammonium may provide a yield advantage under rainfed conditions. Integrating field performance with biochemical and molecular traits may help identify rice genotypes suited to water-limited systems.
]]></description>
<dc:creator><![CDATA[ Baby Thomas, H., Agarwal, T., Verulkar, S. ]]></dc:creator>
<dc:date>2026-10-09</dc:date>
<dc:identifier>doi:10.64898/2026.10.04.756527</dc:identifier>
<dc:title><![CDATA[Genotypic Responses of Rice to Ammonium and Nitrate under Contrasting Water Regimes: Agronomic, Biochemical, and Gene-Expression Analyses]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-09</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.10.02.756356v1?rss=1">
<title>
<![CDATA[
DECONVersation: Single Cell Foundation Model-Derived Embeddings for Robust Cell Type Deconvolution of Bulk RNA-seq 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.10.02.756356v1?rss=1
</link>
<description><![CDATA[
AI and deep learning have transformed single-cell transcriptomics, empowering unsupervised annotation, trajectory inference, and cross-dataset integration at scale. Here, we propose bulk RNA-seq deconvolution as a novel task for single-cell foundation models. Deconvolution estimates cell type proportions from bulk RNA-seq using single-cell reference data, with applications spanning tumor microenvironment profiling to the detection of disease-associated shifts in tissue composition. However, existing methods remain vulnerable to batch effects, subject-level variation, and marker gene selection sensitivity. We introduce DECONVersation, which leverages embeddings from pretrained and fine-tuned single-cell foundation models, coupled with a non-negative least squares solver, to estimate cell type proportions. Foundation model embeddings are robust to batch effects and implicitly encode complex gene relationships, sidestepping explicit marker gene selection. Benchmarked against MuSiC, DWLS, and BayesPrism, DECONVersation achieved comparable or superior performance across diverse tissues. Performance further improved with tissue-specific fine-tuning.
]]></description>
<dc:creator><![CDATA[ Oku, A., Geiger, H., Robine, N., Fu, R. ]]></dc:creator>
<dc:date>2026-10-09</dc:date>
<dc:identifier>doi:10.64898/2026.10.02.756356</dc:identifier>
<dc:title><![CDATA[DECONVersation: Single Cell Foundation Model-Derived Embeddings for Robust Cell Type Deconvolution of Bulk RNA-seq]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-09</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.10.03.756406v1?rss=1">
<title>
<![CDATA[
'broadSeq', a benchmarking tool to contrast differential gene expression methods and identify when genomic neighbourhood matters 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.10.03.756406v1?rss=1
</link>
<description><![CDATA[
Several tools are available to detect differential gene expression. The predominant approach of different statistical methods is to consider changes in gene expression independent from their genomic context. Previously, DELocal was introduced that takes into consideration gene expression dynamics in the genomic neighbourhood. Whereas DELocal was shown to identify developmentally regulated genes in developing mouse tooth, its performance in other biological conditions remains undetermined. To benchmark DElocal together with other methods, we develop another tool 'broadSeq'. It provides easy to use interface for methods like DELocal , deseq2, limma, edger, EBSeq and many more. Here we report benchmarking results on different data sets that comprise wild type developing mouse organs from different embryonic time points, and also human cancer data. The examined datasets are from the heart, kidney, liver, brain, forebrain, and hindbrain. We found that most methods produce relatively comparable lists of differentially expressed genes. Nevertheless, DELocal identifies partially distinct sets of genes compared to the other methods. Evaluating the identified genes using gene and disease ontology terms show that the uniquely detected genes by DELocal are biologically relevant and also highly expressed compared to the genes identified by the other methods. Because consistently highly expressed genes show relatively subtle fold changes, the neighbourhood concept of DELocal appears to correct for the difficulty in detecting these differential expression of highly expressed genes. Finally, we provide the tools as combined bioconductor packages, allowing researchers to explore and decide the best methods for their data and questions.
]]></description>
<dc:creator><![CDATA[ Das Roy, R., Hallikas, O., Kuure, S., Jernvall, J. ]]></dc:creator>
<dc:date>2026-10-09</dc:date>
<dc:identifier>doi:10.64898/2026.10.03.756406</dc:identifier>
<dc:title><![CDATA['broadSeq', a benchmarking tool to contrast differential gene expression methods and identify when genomic neighbourhood matters]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-09</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.10.06.756962v1?rss=1">
<title>
<![CDATA[
A human-in-the-loop approach for faster cell tracking 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.10.06.756962v1?rss=1
</link>
<description><![CDATA[
Traditional cell tracking methods focus on fully automated tracking. The quality of the generated tracks, however, is heavily influenced by factors such as cell speed and density. In complex scenarios where track correctness is essential, an efficient and reliable tool for manual track correction remains a key component of the cell tracking process. Most tracking tools that allow for such manual correction do so only as a post-processing step, using human feedback to correct one track error at a time -- without improving the quality of other tracks in the dataset. This highlights the need for a more efficient track correction method that iteratively integrates human feedback into the automated tracking process. In this paper we propose a human-in-the-loop approach to cell tracking and investigate the extent to which a tracking algorithm can learn from human input during the tracking process. We combine a graph representation of cell detections, Dijkstra's shortest path algorithm, and a neural network that predicts cell displacement to generate initial track suggestions. We introduce a tool that enables human feedback on individual tracks suggested by the algorithm; this feedback is then propagated to improve subsequent track generation. Compared to the traditional approach of fully automated tracking with correction as a post-processing step, our human-in-the-loop tracking framework reduces the number of corrections required. Finally, we show how this approach can be integrated with existing state-of-the-art tracking algorithms by using pre-generated tracks as a prior, improving the quality of tracks suggested to the human annotator.
]]></description>
<dc:creator><![CDATA[ Mihaylova, M., Ankan, A., Sultan, S., Wortel, I. M., Postat, J., Harris, M., Merino, M., Mandl, J. N., Textor, J. ]]></dc:creator>
<dc:date>2026-10-09</dc:date>
<dc:identifier>doi:10.64898/2026.10.06.756962</dc:identifier>
<dc:title><![CDATA[A human-in-the-loop approach for faster cell tracking]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-09</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.10.08.757737v1?rss=1">
<title>
<![CDATA[
Multi-omic long-read sequencing reveals nucleosome control of transcriptional fidelity at the isoform level 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.10.08.757737v1?rss=1
</link>
<description><![CDATA[
Transcription and RNA processing are tightly coupled to chromatin organization, yet how nucleosome positioning is coordinated across the promoter and gene body to shape alternative isoform expression remains poorly understood. This question has been difficult to address because short-read chromatin profiling cannot simultaneously capture promoter and gene-body chromatin states on the same DNA molecule, and short-read RNA sequencing cannot resolve full-length transcripts without requiring computational assembly. To address this, we deleted two conserved chromatin remodelers, ISW1 and CHD1, in Saccharomyces cerevisiae and profiled chromatin fibers, nascent RNA, and full-length mature RNA using Oxford Nanopore long-read sequencing. We developed EpiFLAIR, a computational framework that assigns discrete long-range chromatin states across entire genes and integrates these states with isoform-level transcriptional outputs. Using EpiFLAIR, we show that single-molecule chromatin states together with sequence motifs are strong predictors of alternative TSS isoform expression. Loss of ISW1 and CHD1 produced divergent transcriptional changes by activating transcriptionally silenced promoters while repressing active ones. Single-molecule chromatin profiling suggested that this redistribution of transcriptional activity arose, in part, from an anticorrelation between gene-body and promoter chromatin accessibility along the same chromatin fiber, a finding uniquely enabled by long-read single-molecule resolution. Furthermore, disruption of regular nucleosome spacing impaired RNAPII processivity and co-transcriptional splicing, consequently altering mature RNA levels and splicing isoform composition. Together, our findings demonstrate that nucleosomes maintain transcriptional fidelity during initiation, elongation, and splicing to regulate isoform expression. More broadly, long-read single-molecule chromatin profiling represents a powerful approach for linking chromatin organization to isoform-level gene regulation.
]]></description>
<dc:creator><![CDATA[ Bai, G., Dhillon, N., Meissner, B., Felton, C., Heath, H. D., Boeger, H., Brooks, A. N. ]]></dc:creator>
<dc:date>2026-10-09</dc:date>
<dc:identifier>doi:10.64898/2026.10.08.757737</dc:identifier>
<dc:title><![CDATA[Multi-omic long-read sequencing reveals nucleosome control of transcriptional fidelity at the isoform level]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-09</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.10.08.756856v1?rss=1">
<title>
<![CDATA[
Systematic mapping of enhancers, silencers, and topological elements with prime editing deletion screens 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.10.08.756856v1?rss=1
</link>
<description><![CDATA[
Noncoding regulatory elements in the human genome can control gene expression through distinct mechanisms, including acting as enhancers, silencers, or topological elements. However, we lack tools to identify all such classes systematically and quantify their effects on gene expression in an endogenous genomic context. Here we develop Swap-seq, a method that combines pooled twin prime editing with flow-sorting to delete hundreds of endogenous regulatory sequences and measure their quantitative effects on the expression of a nearby gene. We apply Swap-seq to dissect the regulatory elements that control PPIF expression in THP-1 monocytes. We quantify the effects of distal and intronic enhancers, characterize CTCF insulators, and discover a silencer that regulates distal genes through histone deacetylase (HDAC)-mediated repression of neighboring enhancers. Hundreds of other genomic elements share its chromatin signatures and transcription factor binding sites, suggesting that this silencer represents a broader class of repressive elements. We leverage Swap-seq data to benchmark sequence-to-function models and show that they often capture the local activity of regulatory elements but fail to link them to distal target genes. Swap-seq thus provides a generalizable tool to discover regulatory elements across mechanistic classes, characterize their effects on gene expression, and generate large-scale datasets for evaluating sequence-to-function models.
]]></description>
<dc:creator><![CDATA[ Cai, X. S., Montgomery, M. T., Nagano, M., Nisal, A., Gu, Z., Chen, Z., Obbad, K., Rastogi, R., Tan, Y., Zeng, T., Sheth, M. U., Xia, C., Chiang, S., Kundaje, A., Hansen, A. S., Bintu, L., Engreitz, J. M. ]]></dc:creator>
<dc:date>2026-10-09</dc:date>
<dc:identifier>doi:10.64898/2026.10.08.756856</dc:identifier>
<dc:title><![CDATA[Systematic mapping of enhancers, silencers, and topological elements with prime editing deletion screens]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-09</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.10.08.757705v1?rss=1">
<title>
<![CDATA[
Random mtDNA mutagenesis maps genotype-phenotype associations across metabolic contexts and lineages 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.10.08.757705v1?rss=1
</link>
<description><![CDATA[
Mitochondrial DNA (mtDNA) mutations cause inherited disease, accumulate with age, and recur in tumors, yet their effects remain unresolved because existing models interrogate one engineered allele at a time. We developed mtDIVE, a random C[-&gt;]T mutagenesis platform for pooled mitochondrial forward genetics in human epithelial cells and zebrafish embryos. mtDIVE generates genotypes ranging from single substitutions to >100 mutations per molecule, impairs respiration, and causes lineage-specific developmental delay. Coupling single-cell transcriptomics with mtDNA genotyping revealed that selection on mutant genomes and their transcriptional consequences depend on metabolic state and developmental lineage. Obligate OXPHOS eliminates highly mutated cells by apoptosis, whereas glycolytic conditions promote their competitive displacement. Individual loci exert modest but non-redundant effects that predict cell state beyond aggregate mutation burden, and cancer-recurrent variants are enriched among mutations linked to reduced mitochondrial transcription. Thus, mtDIVE links mtDNA genotypes to cellular phenotypes, separating the consequences of mutation identity from mutational load.
]]></description>
<dc:creator><![CDATA[ Giladi, A., Baik, R., Chung, C.-Y., Kim, M., Adams, J., Saha, R., Land, M., Cui, R., DeBitetto, E., Kelley, M. E., Thompson, C. B., Reznik, E., Sharma, R., Challinge, R., Isaac, S. R., Peterson, J., van Oudenaarden, A., Sfeir, A. ]]></dc:creator>
<dc:date>2026-10-09</dc:date>
<dc:identifier>doi:10.64898/2026.10.08.757705</dc:identifier>
<dc:title><![CDATA[Random mtDNA mutagenesis maps genotype-phenotype associations across metabolic contexts and lineages]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-09</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.10.08.757135v1?rss=1">
<title>
<![CDATA[
Centromeric islands constrain CENP-A domain size and position 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.10.08.757135v1?rss=1
</link>
<description><![CDATA[
Centromere identity is specified epigenetically by nucleosomes containing the histone H3 variant CENP-A, yet centromeres are almost invariably embedded within repetitive DNA, suggesting that underlying DNA sequences contribute to the organization of centromeric chromatin. In Drosophila melanogaster, CENP-A primarily occupies retroelement-rich islands flanked by simple satellite DNA, but their functional contribution remains unclear. Here, we delete the ~50-kb island Giglio from centromere 3 while preserving the flanking satellites and residual CENP-A chromatin. Despite removal of the DNA sequence associated with most centromeric CENP-A, centromere identity and overall CENP-A levels are preserved at the native locus, indicating that the island itself is not required for centromere identity. Instead, Giglio deletion increases variability in CENP-A chromatin organization, producing highly heterogeneous domain sizes and preferential redistribution of the CENP-A domain into the adjacent dodeca satellite. Together, these findings demonstrate that Giglio is dispensable for centromere identity but constrains the size and position of the CENP-A domain, establishing centromeric islands as key determinants of CENP-A domain architecture.
]]></description>
<dc:creator><![CDATA[ Patel, P., Trasmondi, D., Mellone, B. G. ]]></dc:creator>
<dc:date>2026-10-09</dc:date>
<dc:identifier>doi:10.64898/2026.10.08.757135</dc:identifier>
<dc:title><![CDATA[Centromeric islands constrain CENP-A domain size and position]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-09</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.10.08.738042v1?rss=1">
<title>
<![CDATA[
Scalable localization of clustering instability for spatial transcriptomics 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.10.08.738042v1?rss=1
</link>
<description><![CDATA[
Standard spatial transcriptomics workflows rely on clustering methods to define regions of interest for downstream analyses, such as differential expression and pathway enrichment. Although researchers are presented with a single partition, most clustering methods are stochastic, meaning their output depends on random algorithmic choices, typically controlled by a seed; thus, repeating the analysis changing only the seed can yield a different result. Spatial transcriptomics data may be especially sensitive to algorithmic perturbations due to high dimensionality, sparsity, and mixed expression profiles at tissue boundaries and transitional zones; however, the extent of seed-dependent variation and its downstream impact have yet to be systematically evaluated. To bridge this gap, we present InSTability, a framework that localizes disagreement among partitions at the observation level. InSTability is built on (1) an interpretable, per-spot score that quantifies the variability of a spot's cluster assignment across partitions, and (2) an algorithm with time complexity linear in the number of spots for a specified number of runs for computing the instability score at all spots. Together, these advances enable spatially-resolved clustering instability maps for large, high-resolution datasets. Applying InSTability to repeated runs of seven established methods (BayesSpace, GraphST, Leiden, Louvain, SEDR, STAGATE, and SpiceMix) on 18 benchmark samples reveals that clustering instability varies across tissue sections, with shared spatial patterns emerging among different methods applied to the same sample. Furthermore, clustering instability correlates with poor agreement against reference annotations and propagates downstream into differential expression analyses, resulting in inconsistent gene sets across seeds. Thus, InSTability can strengthen spatial transcriptomics studies by highlighting regions where clustering is seed-sensitive and further validation is warranted. InSTability is open-source and available at https://github.com/molloy-lab/InSTability.
]]></description>
<dc:creator><![CDATA[ Wu, W., Nguyen, J. T., Molloy, E. K. ]]></dc:creator>
<dc:date>2026-10-09</dc:date>
<dc:identifier>doi:10.64898/2026.10.08.738042</dc:identifier>
<dc:title><![CDATA[Scalable localization of clustering instability for spatial transcriptomics]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-09</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.10.08.757329v1?rss=1">
<title>
<![CDATA[
qPyCR: A qPCR Data Analyzer for Template Quantification and the Detection of Performance Outliers 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.10.08.757329v1?rss=1
</link>
<description><![CDATA[
Motivation: Real-time qPCR data are commonly analyzed to obtain a Cq values, which are used to compare the relative abundance of a targeted template. However, improper background subtraction distorts Cq and changing the assigned threshold can alter experimental conclusions. In addition, a Cq values provide no information on the quality of the qPCR for a given sample. Alternatively, global fitting of qPCR data provides agnostic quantification without the use of Cq or thresholds, and the fitting terms provide a useful metric of reaction performance. Current implementations of global fitting are limited to custom in-house programs. Results: qPyCR is comprehensive analysis tool written in Python that performs background signal subtractions, computes a threshold, and reports Cq values in a traditional manner. The program also performs global fitting and reports template abundances along with reaction quality values. The quality values can be used to assign performance standards, which allow researchers to distinguish aberrant reactions from true changes in template abundance. Availability and Implementation: Code and example data are freely available on GitHub (https://github.com/sdmoore-labs/qPyCR) and on Zenodo (https://doi.org/10.5281/zenodo.21115342).
]]></description>
<dc:creator><![CDATA[ Moore, S. D. ]]></dc:creator>
<dc:date>2026-10-09</dc:date>
<dc:identifier>doi:10.64898/2026.10.08.757329</dc:identifier>
<dc:title><![CDATA[qPyCR: A qPCR Data Analyzer for Template Quantification and the Detection of Performance Outliers]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-09</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.10.07.757445v1?rss=1">
<title>
<![CDATA[
High-throughput identification of chromatin-opening transcription factors 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.10.07.757445v1?rss=1
</link>
<description><![CDATA[
Cell fate transitions involve major changes to the chromatin landscape, some of which are thought to deepen the basin around a newly adopted cell state, stabilizing gene expression programs that might otherwise be transient. Transcription factors (TFs) are central to this process, binding a subset of their genome-wide motif occurrences to create accessible regions that can act as cell type-specific enhancers. However, apart from a few well-studied cases, it is not clear which of the ~1,600 mammalian TFs, when newly expressed, can directly or indirectly initiate newly accessible sites, nor why only some motif occurrences gain accessibility. Here, we applied pooled CRISPR activation (CRISPRa) to upregulate endogenous TFs to physiological levels in mouse embryonic stem cells (mESCs). After screening guides targeting 240 developmental TFs for transcriptional response, we selected 60 and jointly profiled gene expression and chromatin accessibility in ~140,000 single cells. Some TFs (e.g. Cdx2, Sox10) appeared to initiate new accessibility on their own, predominantly at distal sites that are closed in mESCs but coincide with cell type-specific enhancers later in development. Others appeared to partner with "co-TFs", either by leveraging preexisting mESC heterogeneity (e.g. {uparrow}Grhl1 in Tfap2c+ mESCs [-&gt;] opening of sites bearing motifs for both), or by hierarchical activation (e.g. {uparrow}Sox7 [-&gt;] {uparrow}Gata6 [-&gt;] opening of sites bearing either or both motifs). Such dependencies shape which sites are used and which fates are reachable, and may explain the differential behavior of TFs with nearly identical consensus motifs. Thus, the capacity of a newly expressed TF to open chromatin strongly depends on the cell type, cell state and gene regulatory network it lands in. Charting these dependencies will inform how cell fate transitions are initiated in development, homeostasis and disease, and how to engineer them in silico, in vitro and in vivo.
]]></description>
<dc:creator><![CDATA[ Keith, A., Jain, S., Daza, R. M., Qiu, C., Germain, P.-L., Shendure, J., Domcke, S. ]]></dc:creator>
<dc:date>2026-10-09</dc:date>
<dc:identifier>doi:10.64898/2026.10.07.757445</dc:identifier>
<dc:title><![CDATA[High-throughput identification of chromatin-opening transcription factors]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-09</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.10.08.757597v1?rss=1">
<title>
<![CDATA[
The adaptive architecture of tRNA dependencies across physiological tumor environments 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.10.08.757597v1?rss=1
</link>
<description><![CDATA[
The genetic code is decoded by a highly redundant transfer RNA (tRNA) repertoire, yet whether this apparent redundancy serves a functional role beyond ensuring robust translation remains unclear. Here, we establish an atlas of tRNA dependencies through large-scale CRISPRi perturbation mapping across matched in vivo tumor growth and in vitro cell culture for 16 cancer cell lines spanning seven tissue types. During in vivo tumor growth, tRNA dependencies exhibited tissue organization at isodecoder resolution and were associated with codon demand at the isoacceptor level. Strikingly, these relationships were reorganized when cells were cultured in vitro, establishing environmental sensitivity of tRNA dependency--in contrast to protein components of the decoding machinery such as aminoacyl-tRNA synthetases (aaRSs), whose dependencies were largely preserved across environments. Despite pronounced context specificity, environmental alterations in tRNA dependency were dominated by a shared response across cancer cell lines, organized across isoacceptor and isodecoder levels. By disentangling stable cellular identity from environment-associated dependency states, we resolved a shared axis of isoacceptor dependency remodeling and identified specific tRNA isoacceptors associated with promoting or restraining adaptation to the tumor environment. We further connect this functional architecture to differential translation of codon-usage-biased programs associated with proliferation and tumor-microenvironmental stress. Together, our study reveals that tRNA redundancy does not imply functional equivalence, but instead forms a structured and environmentally plastic control layer linking coding information to translational output.
]]></description>
<dc:creator><![CDATA[ Chen, S., Charbonneau, T., Yousefi, K., Zirak, B., Lee, S., Joshi, T., Pawluk, A., Ramani, V., Goodarzi, H. ]]></dc:creator>
<dc:date>2026-10-09</dc:date>
<dc:identifier>doi:10.64898/2026.10.08.757597</dc:identifier>
<dc:title><![CDATA[The adaptive architecture of tRNA dependencies across physiological tumor environments]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-09</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.10.07.757366v1?rss=1">
<title>
<![CDATA[
Fine-tuning genomic models on rare splicing events identifies a novel biomarker of TDP-43 pathology 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.10.07.757366v1?rss=1
</link>
<description><![CDATA[
Genomic foundation models (GFMs) are pretrained on large genomic datasets, but their ability to predict rare disease-associated events after fine-tuning remains unclear. We tested this using TDP-43 cryptic exons, well characterized splicing events linked to neurodegenerative disease that are poorly represented in standard annotations. We compared the state-of-the-art GFM AlphaGenome with OpenSpliceAI, a specialized splice site prediction model. Fine-tuned OpenSpliceAI detected held-out cryptic exons in human and mouse and captured allele-specific splicing effects at the disease-associated UNC13A locus. In contrast, a splicing prediction head trained on AlphaGenome's frozen pretrained backbone failed to predict cryptic exons. Unfreezing and fine-tuning all model weights, however, enabled AlphaGenome to generalize to held-out data. To investigate cryptic splicing in cell types that are difficult to study experimentally, we applied fine-tuned OpenSpliceAI across the genome to study cell type-specific genes. This led to the identification of new cryptic splicing events in CLDN11 and ARHGAP23, which were subsequently validated in FTLD-TDP brain tissue. Notably, the cryptic exon in CLDN11 is predicted to generate a neoepitope that could serve as an oligodendrocyte-specific biomarker of TDP-43 pathology. Our results show that even a state-of-the-art GFM can fail to generalize to rare splicing events, even with substantial pretraining. However, both foundation and specialized models that learn to generalize to rare biological events can have broad applications in biological discovery and medicine.
]]></description>
<dc:creator><![CDATA[ Martin-Linares, C. P., Peethambaran Mallika, A., Sandal, P. S., Martin, T. W., Morris, M., Wong, P. C., Ling, J. P. ]]></dc:creator>
<dc:date>2026-10-09</dc:date>
<dc:identifier>doi:10.64898/2026.10.07.757366</dc:identifier>
<dc:title><![CDATA[Fine-tuning genomic models on rare splicing events identifies a novel biomarker of TDP-43 pathology]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-09</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.10.02.756377v1?rss=1">
<title>
<![CDATA[
Cerberus: bidirectional state space blocks improve accuracy and efficiency of regulatory sequence models 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.10.02.756377v1?rss=1
</link>
<description><![CDATA[
Sequence-to-function models predict regulatory activity and variant effects from DNA sequence alone, and leading models such as Borzoi compute long-range interactions with self-attention. The reference genome caps unique training sequences, and new assays add labels, not sequences. This constraint favors architectures that learn more from limited sequence diversity. We compared long-range blocks with the rest of the architecture and the training data held fixed, replacing Borzoi's transformer blocks with Hydra, a bidirectional state space block built on Mamba-2. Hydra blocks ran 5.7-fold faster on 8,192-token inputs, and models built on them predicted binned coverage on held-out sequences and classified fine-mapped eQTLs more accurately. Hydra's linear cost in sequence length let us run the blocks at 32 bp resolution and remove the U-net decoder, leaving a convolution-SSM model smaller and more accurate than the one it replaced. Scaling this architecture up on an augmented track collection produced Cerberus, an ensemble of eight models over 786 kb of input. We evaluated it on GTEx v11 fine-mapped eQTL, sQTL, and paQTL benchmarks with negatives matched on allele frequency, gene expression, and phenotype-specific positional annotations to reduce confounding by these properties. Cerberus is more accurate than Borzoi on all three phenotypes, 0.692 versus 0.668 mean eQTL AUPRC across 48 tissues, and estimates eQTL effect sizes better, 0.379 versus 0.321 Spearman $rho$. Against AlphaGenome, it trades leads in eQTL classification, ahead on variants 3 to 100 kb from the transcription start site and behind in coding sequence and within 3 kb. Interpretation of the trained blocks shows interaction range growing with block depth, from a 0.9 kb median half-life in the first block to 79 kb in the seventh, and the deepest heads anchoring on promoters, enhancers, and CTCF peaks. We release the models, the training data, the training and evaluation code, and the benchmark sets, supporting variant scoring and transfer learning.
]]></description>
<dc:creator><![CDATA[ Kelley, D. R., Yuan, H., Huang, X., Linder, J. ]]></dc:creator>
<dc:date>2026-10-09</dc:date>
<dc:identifier>doi:10.64898/2026.10.02.756377</dc:identifier>
<dc:title><![CDATA[Cerberus: bidirectional state space blocks improve accuracy and efficiency of regulatory sequence models]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-09</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.10.02.756289v1?rss=1">
<title>
<![CDATA[
Probing Genomic Foundation Models with Splice-Variant Perturbations 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.10.02.756289v1?rss=1
</link>
<description><![CDATA[
As benchmarks saturate, traditional validation methods fail to capture the probabilistic representations of genomic foundation models. This study probes self-supervised generalists (NTv3 650M, Evo2 40B, Genos 10B) alongside task-specific specialists (SpliceAI, AlphaGenome, Borzoi) using a mechanistically focused paradigm centered on expert-curated single-nucleotide substitutions within the splicing regions of the exon-rich OPA1 gene. Extending this analysis across the broader corpus of OPA1 mutations, variants of uncertain significance (VUS), and an independent 65-gene dataset of spliceogenic mutations demonstrates that generalists can match task-specific networks. Training data scale, genetic diversity, and evolutionary context are key for capturing RNA splicing syntax, whereas sparse routing may constrain performance. Mechanistically, foundation models exhibit a compute-asymmetry: the Evo2 40B model frontloads compute into massive parameter scale to resolve single-nucleotide variants within a narrow 425-bp context window, whereas the bidirectional NTv3 650M model and the Genos 10B mixture-of-experts (MoE) architecture backload compute to inference, requiring spatial logit aggregation and expansive context windows of up to 94 kb. Furthermore, classification accuracy across these foundation models benefitted from locus-specific thresholding rather than mutation-class calibration, exposing representation boundaries of the models.
]]></description>
<dc:creator><![CDATA[ Alavi, M. V. ]]></dc:creator>
<dc:date>2026-10-09</dc:date>
<dc:identifier>doi:10.64898/2026.10.02.756289</dc:identifier>
<dc:title><![CDATA[Probing Genomic Foundation Models with Splice-Variant Perturbations]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-09</prism:publicationDate>
<prism:section></prism:section>
</item>
<item rdf:about="https://www.biorxiv.org/content/10.64898/2026.10.08.757551v1?rss=1">
<title>
<![CDATA[
STEER-seq: a recombinase based platform for pooled multimodal single-cell perturbation screening 
]]>
</title>
<link>
https://www.biorxiv.org/content/10.64898/2026.10.08.757551v1?rss=1
</link>
<description><![CDATA[
Pooled gain-of-function screening is a powerful method for screening for transcription factors whose overexpression leads to changes in cellular identity. However, variable transgene copy number, integration site and genomic context from random viral integration can confound quantitative comparisons between perturbations. We present STEER-seq (Single-cell Transgene Expression via Exchange Recombination sequencing), a virus-free platform for pooled ORF gain-of-function screening in human induced pluripotent stem cells (hiPSCs). STEER-seq combines irreversible Bxb1-mediated integration at a defined genomic safe-harbour locus with promoterless donor libraries and barcode-enabled perturbation tracking, enabling single-copy expression from a common promoter and genomic context and direct recovery of perturbation identity by single-cell sequencing. In a focused haematopoietic screen, SPI1 expression alone was sufficient to drive differentiation towards HSC/MPP-like progenitors in the absence of exogenous cytokines, with transient induction generating progenitors that retained multilineage haematopoietic differentiation potential. Joint RNA-chromatin accessibility profiling linked transcription factor-induced transcriptional responses to chromatin remodelling and enabled reconstruction of perturbation-induced gene regulatory networks, including a RUNX2-driven osteoblast programme. Scaling STEER-seq to around 2,000 human coding sequences enabled systematic discovery of transcription factors capable of inducing distinct cell-state programmes, including neural progenitor-, melanocyte- and vascular smooth muscle-like states, as well as maintenance of pluripotency. STEER-seq provides a scalable, quantitative framework for single-cell and multimodal gain-of-function screening, enabling functional discovery and engineering of human cell identity.
]]></description>
<dc:creator><![CDATA[ Migliori, V., Suo, C., Dufva, O., Gyulev, I., Polanski, K., Symeonidou, V., Rodschinka, G., Sarropoulos, I., Iwama, S., Steemers, A. S., Gontarczyk, A. M., Powell, H., Cujba, A.-M., Park, J.-E. S., Garnett, M., Vassiliou, G. S., Teichmann, S. A., Bassett, A. R. ]]></dc:creator>
<dc:date>2026-10-09</dc:date>
<dc:identifier>doi:10.64898/2026.10.08.757551</dc:identifier>
<dc:title><![CDATA[STEER-seq: a recombinase based platform for pooled multimodal single-cell perturbation screening]]></dc:title>
<dc:publisher>Cold Spring Harbor Laboratory</dc:publisher>
<prism:publicationDate>2026-10-09</prism:publicationDate>
<prism:section></prism:section>
</item>
</rdf:RDF>
