Conference · April 18, 2026
Recent advances in summary evaluation are based on model-based metrics to assess quality dimensions, such as completeness, conciseness, and faithfulness. However, these methods often require large language models, and predicted scores are frequently miscal ...
Link to itemCite
ConferenceLecture Notes in Computer Science · January 1, 2026
Pharmacovigilance and clinical decision support systems utilize structured drug safety data to guide medical practice. However, existing datasets frequently depend on terminologies such as MedDRA, which limits their semantic reasoning capabilities and thei ...
Full textCite
Conference · September 29, 2025
We consider inferring the causal effect of a treatment (intervention) on an outcome of interest in situations where there is potentially an unobserved confounder influencing both the treatment and the outcome. This is achievable by assuming access to two s ...
Link to itemCite
ConferenceJACC Adv · May 2025
BACKGROUND: Hypertrophic cardiomyopathy (HCM) remains underdiagnosed, and artificial intelligence tools for echocardiographic recognition have been hampered by lack of insight into drivers of model predictions and ease of implementation. OBJECTIVES: The pu ...
Full textLink to itemCite
ConferenceProceedings of Machine Learning Research · January 1, 2025
High-resolution spatial transcriptomics (ST) technologies can capture gene expression at the cellular level along with spatial information, but are limited in the number of genes that can be profiled. In contrast, single-cell RNA sequencing (SC) provides m ...
Cite
ConferenceProceedings of Machine Learning Research · January 1, 2025
Synthetically generated data can improve privacy, fairness, and data accessibility; however, it can be challenging in specialized scenarios such as survival analysis. One key challenge in this setting is censoring, i.e., the timing of an event is unknown i ...
Cite
ConferenceIEEE International Workshop on Machine Learning for Signal Processing Mlsp · January 1, 2025
Most existing works on Continual Learning (CL) tend to assume exclusivity or dissimilarity among learning tasks, consequently these methods usually require constantly accumulating task-specific knowledge in memory for each task. This results in the eventua ...
Full textCite
ConferenceIEEE International Workshop on Machine Learning for Signal Processing Mlsp · January 1, 2025
Increasing concerns for data privacy and other difficulties associated with retrieving source data for model training have created the need for source-free transfer learning, in which one only has access to pre-trained models instead of data from the origi ...
Full textCite
ConferenceEmnlp 2025 2025 Conference on Empirical Methods in Natural Language Processing Proceedings of the Conference · January 1, 2025
Subjectivity in NLP tasks, e.g., toxicity classification, has emerged as a critical challenge precipitated by the increased deployment of NLP systems in content-sensitive domains. Conventional approaches aggregate annotator judgements (labels), ignoring mi ...
Full textCite
ConferenceProceedings of Machine Learning Research · January 1, 2025
In-context learning based on attention models is examined for data with categorical outcomes, with inference in such models viewed from the perspective of functional gradient descent (GD). We develop a network composed of attention blocks, with each block ...
Cite
ConferenceProceedings of Machine Learning Research · January 1, 2025
Probabilistic survival analysis models seek to estimate the distribution of the future occurrence (time) of an event given a set of covariates. In recent years, these models have preferred nonparametric specifications that avoid directly estimating surviva ...
Cite
ConferenceProceedings of the 2025 Annual Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics Human Language Technologies Long Papers Naacl Hlt 2025 · January 1, 2025
Smart word substitution aims to enhance sentence quality by improving word choices; however current benchmarks rely on human-labeled data. Since word choices are inherently subjective, ground-truth word substitutions generated by a small group of annotator ...
Full textCite
ConferenceFindings of the Association for Computational Linguistics Naacl 2024 Findings · January 1, 2024
In this paper, we study personalized federated learning for text classification with Pretrained Language Models (PLMs). We identify two challenges in efficiently leveraging PLMs for personalized federated learning: 1) Communication. PLMs are usually large ...
Full textCite
ConferenceProceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining · August 4, 2023
In digital marketing, experimenting with new website content is one of the key levers to improve customer engagement. However, creating successful marketing content is a manual and time-consuming process that lacks clear guiding principles. This paper seek ...
Full textCite
ConferenceCell Mol Gastroenterol Hepatol · 2023
BACKGROUND & AIMS: Nonalcoholic steatohepatitis (NASH), a leading cause of cirrhosis, strongly associates with the metabolic syndrome, an insulin-resistant proinflammatory state that disrupts energy balance and promotes progressive liver degeneration. We a ...
Full textLink to itemCite
ConferenceProceedings 2023 IEEE Winter Conference on Applications of Computer Vision Wacv 2023 · January 1, 2023
Weight pruning is among the most popular approaches for compressing deep convolutional neural networks. Recent work suggests that in a randomly initialized deep neural network, there exist sparse subnetworks that achieve performance comparable to the origi ...
Full textCite
ConferenceProceedings of Machine Learning Research · January 1, 2023
Total correlation (TC) is a fundamental concept in information theory which measures statistical dependency among multiple random variables. Recently, TC has shown noticeable effectiveness as a regularizer in many learning tasks, where the correlation amon ...
Cite
ConferenceProceedings of Machine Learning Research · January 1, 2023
Pretrained language models (PLMs), such as GPT-2, have achieved remarkable empirical performance in text generation tasks. However, pretrained on large-scale natural language corpora, the generated text from PLMs may exhibit social bias against disadvantag ...
Cite
ConferenceProceedings of Machine Learning Research · January 1, 2023
One straightforward metric to evaluate a survival prediction model is based on the Mean Absolute Error (MAE) - the average of the absolute difference between the time predicted by the model and the true event time, over all subjects. Unfortunately, this is ...
Cite
ConferenceProceedings of the Annual Meeting of the Association for Computational Linguistics · January 1, 2023
Federated learning involves collaborative training with private data from multiple platforms, while not violating data privacy. We study the problem of federated domain adaptation for Named Entity Recognition (NER), where we seek to transfer knowledge acro ...
Full textCite
ConferenceProceedings of Machine Learning Research · January 1, 2023
Recently proposed encoder-decoder structures for modeling Hawkes processes use transformer-inspired architectures, which encode the history of events via embeddings and self-attention mechanisms. These models deliver better prediction and goodness-of-fit t ...
Cite
ConferenceAdvances in Neural Information Processing Systems · January 1, 2023
Soft prompt tuning achieves superior performances across a wide range of few-shot tasks. However, the performances of prompt tuning can be highly sensitive to the initialization of the prompts. We have also empirically observed that conventional prompt tun ...
Cite
ConferenceAdvances in Neural Information Processing Systems · January 1, 2023
We address the challenge of generating fair and unbiased image retrieval results given neutral textual queries (with no explicit gender or race connotations), while maintaining the utility (performance) of the underlying vision-language (VL) model. Previou ...
Cite
ConferenceProc Mach Learn Res · August 2022
End-to-end learning of dynamical systems with black-box models, such as neural ordinary differential equations (ODEs), provides a flexible framework for learning dynamics from data without prescribing a mathematical model for the dynamics. Unfortunately, t ...
Link to itemCite
ConferenceInt Conf Learn Represent · April 2022
Though recent works have developed methods that can generate estimates (or imputations) of the missing entries in a dataset to facilitate downstream analysis, most depend on assumptions that may not align with real-world applications and could suffer from ...
Link to itemCite
ConferenceICASSP IEEE International Conference on Acoustics Speech and Signal Processing Proceedings · January 1, 2022
Supervised training of Named Entity Recognition (NER) models generally require large amounts of annotations, which are hardly available for less widely used (low resource) languages, e.g., Armenian and Dutch. Therefore, it will be desirable if we could lev ...
Full textCite
ConferenceProceedings of the Annual Meeting of the Association for Computational Linguistics · January 1, 2022
Previous work of class-incremental learning for Named Entity Recognition (NER) relies on the assumption that there exists abundance of labeled data for the training of new classes. In this work, we study a more challenging but practical problem, i.e., few- ...
Full textCite
ConferenceProceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition · January 1, 2022
An increasing number of applications in computer vision, specially, in medical imaging and remote sensing, become challenging when the goal is to classify very large images with tiny informative objects. Specifically, these classification tasks face two ke ...
Full textCite
ConferenceProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing Emnlp 2022 · January 1, 2022
Open world classification is a task in natural language processing with key practical relevance and impact. Since the open or unknown category data only manifests in the inference phase, finding a model with a suitable decision boundary accommodating for t ...
Full textCite
ConferenceFindings of the Association for Computational Linguistics Emnlp 2022 · January 1, 2022
Supervised training of existing deep learning models for sequence labeling relies on large scale labeled datasets. Such datasets are generally created with crowd-source labeling. However, crowd-source labeling for tasks of sequence labeling can be expensiv ...
Full textCite
ConferenceProceedings of Machine Learning Research · January 1, 2022
End-to-end learning of dynamical systems with black-box models, such as neural ordinary differential equations (ODEs), provides a flexible framework for learning dynamics from data without prescribing a mathematical model for the dynamics. Unfortunately, t ...
Cite
ConferenceAdv Neural Inf Process Syst · December 2021
Dealing with severe class imbalance poses a major challenge for many real-world applications, especially when the accurate classification and generalization of minority classes are of primary interest. In computer vision and NLP, learning from datasets wit ...
Link to itemCite
ConferenceACM Chil 2021 Proceedings of the 2021 ACM Conference on Health Inference and Learning · April 8, 2021
Set classification is the task of predicting a single label from a set comprising multiple instances. The examples we consider are pathology slides represented by sets of patches and medical text data represented by sets of word embeddings. State-of-the-ar ...
Full textCite
ConferenceACM CHIL 2021 (2021) · April 2021
Balanced representation learning methods have been applied successfully to counterfactual inference from observational data. However, approaches that account for survival outcomes are relatively limited. Survival data are frequently encountered across dive ...
Full textLink to itemCite
ConferenceNaacl Hlt 2021 2021 Conference of the North American Chapter of the Association for Computational Linguistics Human Language Technologies Proceedings of the Conference · January 1, 2021
In many natural language processing applications, identifying predictive text can be as important as the predictions themselves. When predicting medical diagnoses, for example, identifying predictive content in clinical notes not only enhances interpretabi ...
Full textCite
ConferenceEmnlp 2021 2021 Conference on Empirical Methods in Natural Language Processing Proceedings · January 1, 2021
Unsupervised consistency training is a way of semi-supervised learning that encourages consistency in model predictions between the original and augmented data. For Named Entity Recognition (NER), existing approaches augment the input sequence with token r ...
Full textCite
ConferenceArXiv · September 17, 2020
Combining the increasing availability and abundance of healthcare data and the current advances in machine learning methods have created renewed opportunities to improve clinical decision support systems. However, in healthcare risk prediction applications ...
Link to itemCite
ConferenceDiabetes · June 1, 2020
T2DM increases risk for advanced fibrosis in NAFLD, however little is known about the impact of glycemic control on fibrosis severity.Objective: To assess whether glycemic control is associated with severity of fibr ...
Full textCite
ConferenceAaai 2020 34th Aaai Conference on Artificial Intelligence · January 1, 2020
Reinforcement learning (RL) has been widely used to aid training in language generation. This is achieved by enhancing standard maximum likelihood objectives with user-specified reward functions that encourage global semantic consistency. We propose a prin ...
Cite
ConferenceFindings of the Association for Computational Linguistics Findings of Acl Emnlp 2020 · January 1, 2020
Pretrained Language Models (PLMs) have improved the performance of natural language understanding in recent years. Such models are pretrained on large corpora, which encode the general prior knowledge of natural languages but are agnostic to information ch ...
Cite
ConferenceProceedings of Machine Learning Research · January 1, 2020
Event time models predict occurrence times of an event of interest based on known features. Recent work has demonstrated that neural networks achieve state-of-the-art event time predictions in biomedical applications, where event time models are frequently ...
Cite
ConferenceCancer Research · July 1, 2019
AbstractProstate cancer (PC) is the most common non-cutaneous malignancy in men. The prediction of outcome based on multicore biopsy specimens is problematic, mainly due to tumor multifocality compounded ...
Full textCite
ConferenceProgress in Biomedical Optics and Imaging Proceedings of SPIE · January 1, 2019
Many researchers in the field of machine learning have addressed the problem of detecting anomalies within Computed Tomography (CT) scans. Training these machine learning algorithms requires a dataset of CT scans with identified anomalies (labels), usually ...
Full textCite
ConferenceProgress in Biomedical Optics and Imaging Proceedings of SPIE · January 1, 2019
Purpose: When conducting machine learning algorithms on classification and detection of abnormalities for medical imaging, many researchers are faced with the problem that it is hard to get enough labeled data. This is especially difficult for modalities s ...
Full textCite
ConferenceProgress in Biomedical Optics and Imaging Proceedings of SPIE · January 1, 2019
Purpose To accurately segment organs from 3D CT image volumes using a 2D, multi-channel SegNet model consisting of a deep Convolutional Neural Network (CNN) encoder-decoder architecture. Method We trained a SegNet model on the extended cardiac-Torso (XCAT) ...
Full textCite
Conference33rd Aaai Conference on Artificial Intelligence Aaai 2019 31st Innovative Applications of Artificial Intelligence Conference Iaai 2019 and the 9th Aaai Symposium on Educational Advances in Artificial Intelligence Eaai 2019 · January 1, 2019
Learning probability distributions on the weights of neural networks has recently proven beneficial in many applications. Bayesian methods such as Stochastic Gradient Markov Chain Monte Carlo (SG-MCMC) offer an elegant framework to reason about model uncer ...
Full textCite
Conference35th International Conference on Machine Learning Icml 2018 · January 1, 2018
To assess the difference between real and synthetic data, Generative Adversarial Networks (GANs) are trained using a distribution discrepancy measure. Three widely employed measures are information-theoretic divergences, integral probability metrics, and H ...
Cite
Conference35th International Conference on Machine Learning Icml 2018 · January 1, 2018
Recent advances on the scalability and flexibility of variational inference have made it successful at unravelling hidden patterns in complex data. In this work we propose a new variational bound formulation, yielding an estimator that extends beyond the c ...
Cite
ConferenceProceedings of Machine Learning Research · January 1, 2018
A new generative adversarial network is developed for joint distribution matching. Distinct from most existing approaches, that only learn conditional distributions, the proposed model aims to learn a joint distribution of multiple random variables (domain ...
Cite
ConferenceProceedings of Machine Learning Research · January 1, 2018
Recent advances on the scalability and flexibility of variational inference have made it successful at unravelling hidden patterns in complex data. In this work we propose a new variational bound formulation, yielding an estimator that extends beyond the c ...
Cite
ConferenceProceedings of Machine Learning Research · January 1, 2018
To assess the difference between real and synthetic data, Generative Adversarial Networks (GANs) are trained using a distribution discrepancy measure. Three widely employed measures are information-theoretic divergences, integral probability metrics, and H ...
Cite
ConferenceEmnlp 2017 Conference on Empirical Methods in Natural Language Processing Proceedings · January 1, 2017
We propose a new encoder-decoder approach to learn distributed sentence representations that are applicable to multiple purposes. The model is learned by using a convolutional neural network as an encoder to map an input sentence into a continuous vector, ...
Full textCite
ConferenceAdvances in Neural Information Processing Systems · January 1, 2017
A new method for learning variational autoencoders (VAEs) is developed, based on Stein variational gradient descent. A key advantage of this approach is that one need not make parametric assumptions about the form of the encoder distribution. Performance i ...
Cite
ConferenceAdvances in Neural Information Processing Systems · January 1, 2017
We investigate the non-identifiability issues associated with bidirectional adversarial training for joint distribution matching. Within a framework of conditional entropy, we propose both adversarial and non-adversarial approaches to learn desirable match ...
Cite
ConferenceProceedings IEEE International Conference on Data Mining Icdm · July 2, 2016
We introduce a novel dynamic model for discrete time-series data, in which the temporal sampling may be nonuniform. The model is specified by constructing a hierarchy of Poisson factor analysis blocks, one for the transitions between latent states and the ...
Full textCite
ConferenceProceedings IEEE International Conference on Data Mining Icdm · July 2, 2016
We propose a non-linear extension to factor analysis with beta process priors for improved data representation ability. This non-linear Beta Process Factor Analysis (nBPFA) allows data to be represented as a non-linear transformation of a standard sparse f ...
Full textCite
ConferenceLecture Notes in Computer Science Including Subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics · January 1, 2016
We proposed a Hamiltonian Monte Carlo (HMC) method with Laplace kinetic energy, and demonstrate the connection between slice sampling and proposed HMC method in one-dimensional cases. Based on this connection, one can perform slice sampling using a numeric ...
Full textCite
ConferenceIjcai International Joint Conference on Artificial Intelligence · January 1, 2016
In dictionary learning for analysis of images, spatial correlation from extracted patches can be leveraged to improve characterization power. We propose a Bayesian framework for dictionary learning, with spatial location dependencies captured by imposing a ...
Cite
ConferenceAdvances in Neural Information Processing Systems · January 1, 2016
We unify slice sampling and Hamiltonian Monte Carlo (HMC) sampling, demonstrating their connection via the Hamiltonian-Jacobi equation from Hamiltonian mechanics. This insight enables extension of HMC and slice sampling to a broader family of samplers, cal ...
Cite
ConferenceAdvances in Neural Information Processing Systems · January 1, 2016
A novel variational autoencoder is developed to model images, as well as associated labels or captions. The Deep Generative Deconvolutional Network (DGDN) is used as a decoder of the latent image features, and a deep Convolutional Neural Network (CNN) is u ...
Cite
Conference32nd International Conference on Machine Learning Icml 2015 · January 1, 2015
Point process data are commonly observed in fields like healthcare and the social sciences. Designing predictive models for such event streams is an under-explored problem, due to often scarce training data. In this work we propose a multitask point proces ...
Cite
ConferenceAdvances in Neural Information Processing Systems · January 1, 2015
We present a scalable Bayesian multi-label learning model based on learning lowdimensional label embeddings. Our model assumes that each label vector is generated as a weighted combination of a set of topics (each topic being a distribution over labels), w ...
Cite
ConferenceAdvances in Neural Information Processing Systems · January 1, 2015
We propose a new deep architecture for topic modeling, based on Poisson Factor Analysis (PFA) modules. The model is composed of a Poisson distribution to model observed vectors of counts, as well as a deep hierarchy of hidden binary units. Rather than usin ...
Cite
ConferenceAdvances in Neural Information Processing Systems · January 1, 2014
A new Bayesian formulation is developed for nonlinear support vector machines (SVMs), based on a Gaussian process and with the SVM hinge loss expressed as a scaled mixture of normals. We then integrate the Bayesian SVM into a factor model, in which feature ...
Cite
ConferenceAdvances in Neural Information Processing Systems 22 Proceedings of the 2009 Conference · January 1, 2009
In this paper we present a novel approach to learn directed acyclic graphs (DAGs) and factor models within the same framework while also allowing for model comparison between them. For this purpose, we exploit the connection between factor models and DAGs ...
Cite
ConferenceLecture Notes in Computer Science Including Subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics · January 1, 2006
This paper introduces a temporal version of Probabilistic Kernel Principal Component Analysis by using a hidden Markov model in order to obtain optimized representations of observed data through time. Recently introduced. Probabilistic Kernel Principal Com ...
Full textCite
Conference2006 28TH ANNUAL INTERNATIONAL CONFERENCE OF THE IEEE ENGINEERING IN MEDICINE AND BIOLOGY SOCIETY, VOLS 1-15 · January 1, 2006Link to itemCite
ConferenceProceedings of the Annual Conference of the International Speech Communication Association Interspeech · January 1, 2004
In this article, it is studied the usefulness of the support vector machines (SVM) algorithm in the active classification of voice records into the sets normal and pathologic. In practice, each one of the samples employed on the classifier training must be ...
Cite