Preprint · August 6, 2026
Predicting how a genetic perturbation reshapes a cell's transcriptome is a central goal of computational biology. Previous studies report that the mean response across training perturbations rivals specialized models on standard accuracy metrics, e ...
Full textCite
Journal articleNature · July 29, 2026
Proteins have evolved over billions of years through coordinated substitutions, insertions and deletions, yet computational protein design cannot fully replicate nature's ability to engineer new proteins from existing templates. Protein language models1-3 ...
Full textLink to itemCite
Journal articleCell Syst · July 17, 2026
The three-dimensional organization of chromatin into topologically associating domains (TADs) may impact gene regulation by bringing distant genes into contact. However, studies of TADs' function and their influence on transcription have been constrained b ...
Full textLink to itemCite
Preprint · 2026
Virtual screening of billion-scale compound libraries has become feasible through machine learning approaches. In particular, CoNCISE (RECOMB 2025) introduced drug quantization via code-books, achieving highly scalable and accurate binary predictions. Howe ...
Full textCite
Preprint · 2026
Protein language models (PLMs) encode amino acid sequences into residue-level embeddings that must be pooled into fixed-size representations for downstream protein-level prediction tasks. Although these embeddings implicitly reflect evolutionary constraint ...
Full textCite
Preprint · 2026
Non-negative matrix factorization (NMF) is a foundational dimensionality-reduction method in single-cell transcriptomics, valued for its interpretable gene programs. However, in case-control settings common in perturbation and disease research, standard NM ...
Full textCite
Preprint · 2026
Cell extrusion contributes to epithelial homeostasis, but its dysregulation can lead to tumorigenesis or degeneration. A fine balance in this process is therefore essential for tissue integrity. Yet the cell types and states vulnerable to extrusion, and th ...
Full textCite
Journal articleNat Biotechnol · October 24, 2025
Training and deploying large-scale protein language models typically requires deep machine learning expertise-a barrier for researchers outside this field. SaprotHub overcomes this challenge by offering an intuitive platform that facilitates training and p ...
Full textLink to itemCite
Journal articleNat Commun · August 29, 2025
Genome-wide association studies (GWAS) identify numerous disease-linked genetic variants at noncoding genomic loci, yet therapeutic progress is hampered by the challenge of deciphering the regulatory roles of these loci in tissue-specific contexts. Single- ...
Full textLink to itemCite
Journal articlebioRxiv · July 11, 2025
The bipotential gonad is the precursor organ to both the ovary and testis and develops as part of the embryonic urogenital system. In mice, gonadogenesis initiates around embryonic day 9.5 (E9.5), when coelomic epithelial (CE) cells overlaying the mesoneph ...
Full textLink to itemCite
Preprint · May 27, 2025
The syncytiotrophoblast (STB) is a multinucleated cell layer that forms the outer surface of human chorionic villi. Its unusual structure, with billions of nuclei in a single cell, makes it difficult to resolve using conventional single-cell methods. To be ...
Full textLink to itemCite
Journal articleeLife · May 27, 2025
The syncytiotrophoblast (STB) is a multinucleated cell layer that forms the outer surface of human chorionic villi. Its unusual structure, with billions of nuclei in a single cell, makes it difficult to resolve using conventional single-cell method ...
Full textCite
Preprint · April 2, 2025
UNLABELLED: The three-dimensional organization of chromatin into topologically associating domains (TADs) may impact gene regulation by bringing distant genes into contact. However, many questions about TADs' function and their influence on transcription r ...
Full textLink to itemCite
Journal articleJ Biol Chem · March 2025
Protein S-palmitoylation is a reversible lipophilic posttranslational modification regulating diverse signaling pathways. Within transmembrane proteins (TMPs), S-palmitoylation is implicated in conditions from inflammatory disorders to respiratory viral in ...
Full textLink to itemCite
Journal articlebioRxiv · February 19, 2025
The outer surface of chorionic villi in the human placenta consists of a single multinucleated cell called the syncytiotrophoblast (STB). The unique cellular ultrastructure of the STB presents challenges in deciphering its gene expression signature at the ...
Full textLink to itemCite
Journal articleProc Natl Acad Sci U S A · January 7, 2025
Protein language models (PLMs) have demonstrated impressive success in modeling proteins. However, general-purpose "foundational" PLMs have limited performance in modeling antibodies due to the latter's hypervariable regions, which do not conform to the ev ...
Full textLink to itemCite
Preprint · 2025
Rapid advances in deep learning have improved in silico methods for drug-target interaction (DTI) prediction. However, current methods do not scale to the massive catalogs that list millions or billions of commercially-available small molecules. Here, we ...
Full textCite
Journal articleBioinform Adv · 2025
MOTIVATION: Protein language models (PLMs) have emerged as powerful approaches for mapping protein sequences into embeddings suitable for various applications. As protein representation schemes, PLMs generate per-token (i.e. per-residue) representations, r ...
Full textOpen AccessLink to itemCite
Book section · January 1, 2025
Rapid advances in deep learning have improved in silico methods for drug-target interaction (DTI) prediction. However, current methods struggle to scale to catalogs listing billions of commercially-available small molecules. Here, we introduce CoNCISE, a m ...
Full textCite
ConferenceLecture Notes in Computer Science · January 1, 2025
We introduce PHILHARMONIC, a computational framework that couples deep learning de novo network inference with robust unsupervised spectral clustering algorithms to uncover functional relationships and high-level organization in non-model organisms. Our no ...
Full textCite
Preprint · 2025
Spatial omics technologies provide complementary and layered molecular insights that span proteins, transcripts, and metabolites. However, aligning and integrating these modalities across serial tissue sections remains a computational challenge. Existing a ...
Full textCite
Preprint · 2025
Foundation models pre-trained on certain biological data modalities exhibit systematic representational biases when encountering out-of-distribution (OOD) data from new assays. The embedding drift largely arises from instrumentation and protocol-related ar ...
Full textCite
Preprint · September 8, 2024
Protein S-palmitoylation is a reversible lipophilic posttranslational modification regulating a diverse number of signaling pathways. Within transmembrane proteins (TMPs), S-palmitoylation is implicated in conditions from inflammatory disorders to respirat ...
Full textLink to itemCite
Journal articleCell Syst · May 15, 2024
Single-cell expression dynamics, from differentiation trajectories or RNA velocity, have the potential to reveal causal links between transcription factors (TFs) and their target genes in gene regulatory networks (GRNs). However, existing methods either ov ...
Full textLink to itemCite
Preprint · 2024
Proteins have evolved over billions of years through extensive and coordinated substitutions, insertions and deletions (indels). Computational protein design cannot yet fully mimic nature’s ability to engineer new proteins from existing templates. Protein ...
Full textCite
Journal articleBioinformatics · November 1, 2023
MOTIVATION: High-quality computational structural models are now precomputed and available for nearly every protein in UniProt. However, the best way to leverage these models to predict which pairs of proteins interact in a high-throughput manner is not im ...
Full textLink to itemCite
Journal articleProc Natl Acad Sci U S A · June 13, 2023
Sequence-based prediction of drug-target interactions has the potential to accelerate drug discovery by complementing experimental screens. Such computational prediction needs to be generalizable and scalable while remaining sensitive to subtle variations ...
Full textLink to itemCite
Journal articleProc Natl Acad Sci U S A · June 13, 2023
The split-Gal4 system allows for intersectional genetic labeling of highly specific cell types and tissues in Drosophila. However, the existing split-Gal4 system, unlike the standard Gal4 system, cannot be repressed by Gal80, and therefore cannot be contro ...
Full textLink to itemCite
Journal articlePLoS One · 2023
With the ease of gene sequencing and the technology available to study and manipulate non-model organisms, the extension of the methodological toolbox required to translate our understanding of model organisms to non-model organisms has become an urgent pr ...
Full textLink to itemCite
Preprint · 2023
The utility of single-cell RNA sequencing (scRNA-seq) is premised on the notion that transcriptional state can faithfully reflect cell phenotype. However, scRNA-seq measurements are noisy and sparse, with individual transcript counts showing limited correl ...
Full textCite
Journal articleTransactions on Machine Learning Research · January 1, 2023
Graph attention networks estimate the relational importance of node neighbors to aggregate relevant information over local neighborhoods for a prediction task. However, the inferred attentions are vulnerable to spurious correlations and connectivity in the ...
Cite
Journal articleBioinformatics · June 24, 2022
SUMMARY: Computational methods to predict protein-protein interaction (PPI) typically segregate into sequence-based 'bottom-up' methods that infer properties from the characteristics of the individual protein sequences, or global 'top-down' methods that in ...
Full textLink to itemCite
ConferenceIclr 2022 10th International Conference on Learning Representations · January 1, 2022
When a dynamical system can be modeled as a sequence of observations, Granger causality is a powerful approach for detecting predictive interactions between its variables. However, traditional Granger causal inference has limited utility in domains where t ...
Cite
Journal articleCell Syst · October 20, 2021
We combine advances in neural language modeling and structurally motivated design to develop D-SCRIPT, an interpretable and generalizable deep-learning model, which predicts interaction between two proteins using only their sequence and maintains high accu ...
Full textLink to itemCite
Journal articleGenome Biol · May 3, 2021
A complete understanding of biological processes requires synthesizing information across heterogeneous modalities, such as age, disease status, or gene expression. Technological advances in single-cell profiling have enabled researchers to assay multiple ...
Full textLink to itemCite
Preprint · 2021
Summary Chromosome conformation capture technologies such as Hi-C have revealed a rich hierarchical structure of chromatin, with topologically associating domains (TADs) as a key organizational unit, but experimentally reported TAD architectures, ...
Full textCite
Journal articleSci Signal · October 25, 2011
Characterizing the extent and logic of signaling networks is essential to understanding specificity in such physiological and pathophysiological contexts as cell fate decisions and mechanisms of oncogenesis and resistance to chemotherapy. Cell-based RNA in ...
Full textLink to itemCite
Journal articleAlgorithms Mol Biol · April 19, 2011
BACKGROUND: Proteins are dynamic molecules that exhibit a wide range of motions; often these conformational changes are important for protein function. Determining biologically relevant conformational changes, or true variability, efficiently is challengin ...
Full textLink to itemCite
Journal articleNucleic Acids Res · January 2011
We describe IsoBase, a database identifying functionally related proteins, across five major eukaryotic model organisms: Saccharomyces cerevisiae, Drosophila melanogaster, Caenorhabditis elegans, Mus musculus and Homo Sapiens. Nearly all existing algorithm ...
Full textLink to itemCite
ConferenceLecture Notes in Computer Science Including Subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics · November 10, 2010
Proteins are dynamic molecules that exhibit a wide range of motions; often these conformational changes are important for protein function. Determining biologically relevant conformational changes, or true variability, efficiently is challenging due to the ...
Full textCite
Journal articleNucleic Acids Res · July 2010
Struct2Net is a web server for predicting interactions between arbitrary protein pairs using a structure-based approach. Prediction of protein-protein interactions (PPIs) is a central area of interest and successful prediction would provide leads for exper ...
Full textLink to itemCite
ConferenceBioinformatics · June 15, 2009
MOTIVATION: With the increasing availability of large protein-protein interaction networks, the question of protein network alignment is becoming central to systems biology. Network alignment is further delineated into two sub-problems: local alignment, to ...
Full textLink to itemCite
ConferenceProceedings of the Annual ACM SIAM Symposium on Discrete Algorithms · December 1, 2008
The post-genomic era has witnessed an explosion in the quality, quantity and variety of biological data-sequence, structure, and networks. However, when building computational models on these data, some abstractions recur often. In particular, graph-based ...
Cite
Journal articleProc Natl Acad Sci U S A · September 2, 2008
Protein-protein interactions (PPIs) and their networks play a central role in all biological processes. Akin to the complete sequencing of genomes and their comparative analysis, complete descriptions of interactomes and their comparative analysis is funda ...
Full textLink to itemCite
ConferencePac Symp Biocomput · 2008
UNLABELLED: We describe an algorithm for global alignment of multiple protein-protein interaction (PPI) networks, the goal being to maximize the overall match across the input networks. The intuition behind our algorithm is that a protein in one PPI networ ...
Link to itemCite
Journal articleJ Comput Biol · October 2007
We introduce a computational method to predict and annotate the catalytic residues of a protein using only its sequence information, so that we describe both the residues' sequence locations (prediction) and their specific biochemical roles in the catalyze ...
Full textLink to itemCite
ConferenceLecture Notes in Computer Science Including Subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics · January 1, 2007
We describe an algorithm, ISORANK, for global alignment of two protein-protein interaction (PPI) networks. ISORANK aims to maximize the overall match between the two networks; in contrast, much of previous work has focused on the local alignment problem- i ...
Full textCite
ConferencePac Symp Biocomput · 2007
UNLABELLED: We describe a novel probabilistic approach to estimating errors in two-hybrid (2H) experiments. Such experiments are frequently used to elucidate protein-protein interaction networks in a high-throughput fashion; however, a significant challeng ...
Link to itemCite
ConferencePac Symp Biocomput · 2006
UNLABELLED: This paper presents a framework for predicting protein-protein interactions (PPI) that integrates structure-based information with other functional annotations, e.g. GO, co-expression and co-localization, etc., Given two protein sequences, the ...
Link to itemCite
ConferenceIcml 2005 Proceedings of the 22nd International Conference on Machine Learning · December 1, 2005
Many time-series experiments seek to estimate some signal as a continuous function of time. In this paper, we address the sampling problem for such experiments: determining which time-points ought to be sampled in order to minimize the cost of data collect ...
Cite
ConferencePac Symp Biocomput · 2005
When searching for an optimal protein structure, it is often necessary to generate a set of structures similar, e.g., within 4A Root Mean Square Deviation (RMSD), to some base structure. Current methods to do this are designed to produce only small deviati ...
Link to itemCite
Journal articlePac Symp Biocomput · 2003
In biological macromolecules, structural patterns (motifs) are often repeated across different molecules. Detection of these common motifs in a new molecule can provide useful clues to the functional properties of such a molecule. We formulate the problem ...
Link to itemCite