Conference13th International Conference on Learning Representations Iclr 2025 · January 1, 2025
We show theoretically and empirically that the linear Transformer, when applied to graph data, can implement algorithms that solve canonical problems such as electric flow and eigenvector decomposition. The Transformer has access to information on the inpu ...
Cite
ConferenceProceedings of the Annual Meeting of the Association for Computational Linguistics · January 1, 2025
Automatic post-editing (APE) aims to correct errors in machine-translated text, enhancing translation quality, while reducing the need for human intervention. Despite advances in neural machine translation (NMT), the development of effective APE systems ha ...
Full textCite
ConferenceProceedings of Machine Learning Research · January 1, 2025
In-context learning based on attention models is examined for data with categorical outcomes, with inference in such models viewed from the perspective of functional gradient descent (GD). We develop a network composed of attention blocks, with each block ...
Cite
ConferenceProceedings 2024 IEEE Winter Conference on Applications of Computer Vision Wacv 2024 · January 3, 2024
Zero-shot learning (ZSL) is a promising approach to generalizing a model to categories unseen during training by leveraging class attributes, but challenges remain. Recently, methods using generative models to combat bias towards classes seen during traini ...
Full textCite
ConferenceProceedings 2023 IEEE Winter Conference on Applications of Computer Vision Wacv 2023 · January 1, 2023
Weight pruning is among the most popular approaches for compressing deep convolutional neural networks. Recent work suggests that in a randomly initialized deep neural network, there exist sparse subnetworks that achieve performance comparable to the origi ...
Full textCite
ConferenceProceedings of Machine Learning Research · January 1, 2023
Total correlation (TC) is a fundamental concept in information theory which measures statistical dependency among multiple random variables. Recently, TC has shown noticeable effectiveness as a regularizer in many learning tasks, where the correlation amon ...
Cite
ConferenceInternational Conference on Information and Knowledge Management Proceedings · October 17, 2022
Numbers are essential components of text, like any other word tokens, from which natural language processing (NLP) models are built and deployed. Though numbers are typically not accounted for distinctly in most NLP tasks, there is still an underlying amou ...
Full textCite
ConferenceProc Mach Learn Res · August 2022
End-to-end learning of dynamical systems with black-box models, such as neural ordinary differential equations (ODEs), provides a flexible framework for learning dynamics from data without prescribing a mathematical model for the dynamics. Unfortunately, t ...
Link to itemCite
ConferenceInt Conf Learn Represent · April 2022
Though recent works have developed methods that can generate estimates (or imputations) of the missing entries in a dataset to facilitate downstream analysis, most depend on assumptions that may not align with real-world applications and could suffer from ...
Link to itemCite
ConferenceProceedings 2022 IEEE Cvf Winter Conference on Applications of Computer Vision Wacv 2022 · January 1, 2022
In many real-world tasks, a canonical 'big data' problem is created by combining data from several individual groups or domains. Because test data will likely come from a new group of data, we want to utilize the grouped structure of our training data to e ...
Full textCite
ConferenceSpringer Proceedings in Mathematics and Statistics · January 1, 2022
Control variates are a well-established tool to reduce the variance of Monte Carlo estimators. However, for large-scale problems including high-dimensional and large-sample settings, their advantages can be outweighed by a substantial computational cost. T ...
Full textCite
ConferenceDeelio 2022 Deep Learning Inside Out 3rd Workshop on Knowledge Extraction and Integration for Deep Learning Architectures Proceedings of the Workshop · January 1, 2022
GPT-3 has attracted lots of attention due to its superior performance across a wide range of NLP tasks, especially with its in-context learning abilities. Despite its success, we found that the empirical results of GPT-3 depend heavily on the choice of in- ...
Cite
ConferenceProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing Emnlp 2022 · January 1, 2022
Open world classification is a task in natural language processing with key practical relevance and impact. Since the open or unknown category data only manifests in the inference phase, finding a model with a suitable decision boundary accommodating for t ...
Full textCite
ConferenceProceedings of Machine Learning Research · January 1, 2022
End-to-end learning of dynamical systems with black-box models, such as neural ordinary differential equations (ODEs), provides a flexible framework for learning dynamics from data without prescribing a mathematical model for the dynamics. Unfortunately, t ...
Cite
ConferenceAdvances in Neural Information Processing Systems · January 1, 2022
Successful applications of InfoNCE (Information Noise-Contrastive Estimation) and its variants have popularized the use of contrastive variational mutual information (MI) estimators in machine learning. While featuring superior stability, these estimators ...
Cite
ConferenceAdv Neural Inf Process Syst · December 2021
Dealing with severe class imbalance poses a major challenge for many real-world applications, especially when the accurate classification and generalization of minority classes are of primary interest. In computer vision and NLP, learning from datasets wit ...
Link to itemCite
ConferenceProceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining · August 14, 2021
The outbreak of COVID-19 Disease due to the novel coronavirus has caused a shortage of medical resources. To aid and accelerate the diagnosis process, automatic diagnosis of COVID-19 via deep learning models has recently been explored by researchers across ...
Full textCite
ConferenceIEEE Computer Society Conference on Computer Vision and Pattern Recognition Workshops · June 1, 2021
Federated learning has emerged as an important distributed learning paradigm, where a server aggregates a global model from many client-trained models, while having no access to the client data. Although it is recognized that statistical heterogeneity of t ...
Full textCite
ConferenceACM Chil 2021 Proceedings of the 2021 ACM Conference on Health Inference and Learning · April 8, 2021
Set classification is the task of predicting a single label from a set comprising multiple instances. The examples we consider are pathology slides represented by sets of patches and medical text data represented by sets of word embeddings. State-of-the-ar ...
Full textCite
ConferenceACM CHIL 2021 (2021) · April 2021
Balanced representation learning methods have been applied successfully to counterfactual inference from observational data. However, approaches that account for survival outcomes are relatively limited. Survival data are frequently encountered across dive ...
Full textLink to itemCite
ConferenceLecture Notes in Computer Science · January 1, 2021
Recent unsupervised approaches to domain adaptation primarily focus on minimizing the gap between the source and the target domains through refining the feature generator, in order to learn a better alignment between the two domains. This minimization can ...
Full textCite
ConferenceNaacl Hlt 2021 2021 Conference of the North American Chapter of the Association for Computational Linguistics Human Language Technologies Proceedings of the Conference · January 1, 2021
In many natural language processing applications, identifying predictive text can be as important as the predictions themselves. When predicting medical diagnoses, for example, identifying predictive content in clinical notes not only enhances interpretabi ...
Full textCite
ConferenceProceedings of Machine Learning Research · January 1, 2021
Naively trained neural networks tend to experience catastrophic forgetting in sequential task settings, where data from previous tasks are unavailable. A number of methods, using various model expansion strategies, have been proposed recently as possible s ...
Cite
Conference35th Aaai Conference on Artificial Intelligence Aaai 2021 · January 1, 2021
We propose a novel and principled method to learn a nonparametric graph model called graphon, which is defined in an infinite-dimensional space and represents arbitrary-size graphs. Based on the weak regularity lemma from the theory of graphons, we leverag ...
Full textCite
Conference35th Aaai Conference on Artificial Intelligence Aaai 2021 · January 1, 2021
An unbiased low-variance gradient estimator, termed GO gradient, was proposed recently for expectation-based objectives Eqγ (y)[f(y)], where the random variable (RV) y may be drawn from a stochastic computation graph (SCG) with continuous (non-reparameteri ...
Full textCite
ConferenceProceedings 2021 IEEE Winter Conference on Applications of Computer Vision Wacv 2021 · January 1, 2021
We propose an optimal transport (OT) framework for generalized zero-shot learning (GZSL), seeking to distinguish samples for both seen and unseen classes, with the assist of auxiliary attributes. The discrepancy between features and attributes is minimized ...
Full textCite
ConferenceCeur Workshop Proceedings · January 1, 2021
Attention-based deep learning models have demonstrated significant improvement over traditional algorithms in several NLP tasks. The Transformer, for instance, is an illustrative example that generates abstract representations of tokens that are input to a ...
Cite
ConferenceProceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition · January 1, 2021
As neural networks are increasingly being applied to real-world applications, mechanisms to address distributional shift and sequential task learning without forgetting are critical. Methods incorporating network expansion have shown promise by naturally a ...
Full textCite
ConferenceAdvances in Neural Information Processing Systems · January 1, 2021
We present a continual learning approach for generative adversarial networks (GANs), by designing and leveraging parameter-efficient feature map transformations. Our approach is based on learning a set of global and task-specific parameters. The global par ...
Cite
ConferenceFindings of the Association for Computational Linguistics Findings of Acl Emnlp 2021 · January 1, 2021
It has been shown that training multi-task models with auxiliary tasks can improve the target tasks quality through cross-task transfer. However, the importance of each auxiliary task to the primary task is likely not known a priori. While the importance w ...
Full textCite
ConferenceNaacl Hlt 2021 2021 Conference of the North American Chapter of the Association for Computational Linguistics Human Language Technologies Proceedings of the Conference · January 1, 2021
Natural language often exhibits inherent hierarchical structure ingrained with complex syntax and semantics. However, most state-of-the-art deep generative models learn embeddings only in Euclidean vector space, without accounting for this structural prope ...
Full textCite
ConferenceIclr 2021 9th International Conference on Learning Representations · January 1, 2021
Large-scale language models have recently demonstrated impressive empirical performance. Nevertheless, the improved results are attained at the price of bigger models, more power consumption, and slower inference, which hinder their applicability to low-re ...
Cite
ConferenceIclr 2021 9th International Conference on Learning Representations · January 1, 2021
Pretrained text encoders, such as BERT, have been applied increasingly in various natural language processing (NLP) tasks, and have recently demonstrated significant performance gains. However, recent studies have demonstrated the existence of social bias ...
Cite
ConferenceIclr 2021 9th International Conference on Learning Representations · January 1, 2021
Voice style transfer, also called voice conversion, seeks to modify one speaker's voice to generate speech as if it came from another (target) speaker. Previous works have made progress on voice conversion with parallel training data and pre-known speakers ...
Cite
ConferenceAnnu Int Conf IEEE Eng Med Biol Soc · July 2020
Over the last decade, convolutional neural networks (CNNs) have emerged as the leading algorithms in image classification and segmentation. Recent publication of large medical imaging databases have accelerated their use in the biomedical arena. While trai ...
Full textLink to itemCite
ConferenceAcl 2019 57th Annual Meeting of the Association for Computational Linguistics Proceedings of the Conference · January 1, 2020
Vector representations of sentences, trained on massive text corpora, are widely used as generic sentence embeddings across a variety of NLP problems. The learned representations are generally assumed to be continuous and real-valued, giving rise to a larg ...
Cite
ConferenceAcl 2019 57th Annual Meeting of the Association for Computational Linguistics Proceedings of the Conference · January 1, 2020
We present a syntax-infused variational autoencoder (SIVAE), that integrates sentences with their syntactic trees to improve the grammar of generated sentences. Distinct from existing VAE-based text generative models, SIVAE contains two separate latent spa ...
Cite
ConferenceAcl 2019 57th Annual Meeting of the Association for Computational Linguistics Proceedings of the Conference · January 1, 2020
Variational autoencoders (VAEs) have received much attention recently as an end-to-end architecture for text generation with latent variables. However, previous works typically focus on synthesizing relatively short sentences (up to 20 words), and the post ...
Cite
ConferenceAcl 2019 57th Annual Meeting of the Association for Computational Linguistics Proceedings of the Conference · January 1, 2020
Constituting highly informative network embeddings is an important tool for network analysis. It encodes network topology, along with other useful side information, into low-dimensional node-based feature representations that can be exploited by statistica ...
Cite
ConferenceAaai 2020 34th Aaai Conference on Artificial Intelligence · January 1, 2020
Reinforcement learning (RL) has been widely used to aid training in language generation. This is achieved by enhancing standard maximum likelihood objectives with user-specified reward functions that encourage global semantic consistency. We propose a prin ...
Cite
ConferenceProceedings of SPIE the International Society for Optical Engineering · January 1, 2020
Recently, progress has been made in the supervised training of Convolutional Object Detectors (e.g. Faster R-CNN) for threat recognition in carry-on luggage using X-ray images. This is part of the Transportation Security Administration's (TSA's) mission to ...
Full textCite
ConferenceProceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition · January 1, 2020
Neural networks are known to be vulnerable to carefully crafted adversarial examples, and these malicious samples often transfer, i.e., they remain adversarial even against other models. Although significant effort has been devoted to the transferability a ...
Full textCite
ConferenceProceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition · January 1, 2020
Learning to navigate in a visual environment following natural-language instructions is a challenging task, because the multimodal inputs to the agent are highly variable, and the training data on a new task is often limited. We present the first pre-train ...
Full textCite
ConferenceFindings of the Association for Computational Linguistics Findings of Acl Emnlp 2020 · January 1, 2020
Pretrained Language Models (PLMs) have improved the performance of natural language understanding in recent years. Such models are pretrained on large corpora, which encode the general prior knowledge of natural languages but are agnostic to information ch ...
Cite
ConferenceEmnlp 2020 2020 Conference on Empirical Methods in Natural Language Processing Proceedings of the Conference · January 1, 2020
Legislator preferences are typically represented as measures of general ideology estimated from roll call votes on legislation, potentially masking important nuances in legislators' political attitudes. In this paper we introduce a method of measuring more ...
Full textCite
ConferenceProceedings of Machine Learning Research · January 1, 2020
Particle-optimization-based sampling (POS) is a recently developed effective sampling technique that interactively updates a set of particles to approximate a target distribution. A representative algorithm is the Stein variational gradient descent (SVGD). ...
Cite
ConferenceProceedings of Machine Learning Research · January 1, 2020
Reinforcement learning (RL) has been widely studied for improving sequence-generation models. However, the conventional rewards used for RL training typically cannot capture sufficient semantic information and therefore manifest model bias. Further, the sp ...
Cite
ConferenceProceedings of the Annual Meeting of the Association for Computational Linguistics · January 1, 2020
Auto-regressive text generation models usually focus on local fluency, and may cause inconsistent semantic meaning in long text generation. Further, automatically generating words with similar semantics is challenging, and hand-crafted linguistic rules are ...
Cite
ConferenceAaai 2020 34th Aaai Conference on Artificial Intelligence · January 1, 2020
We propose a novel graph-driven generative model, that unifies multiple heterogeneous learning tasks into the same framework. The proposed model is based on the fact that heterogeneous learning tasks, which correspond to different generative processes, oft ...
Cite
Conference37th International Conference on Machine Learning Icml 2020 · January 1, 2020
Stochastic particle-optimization sampling (SPOS) is a recently-developed scalable Bayesian sampling framework unifying stochastic gradient MCMC (SG-MCMC) and Stein variational gradient descent (SVGD) algorithms based on Wasserstein gradient flows. With a r ...
Cite
Conference37th International Conference on Machine Learning Icml 2020 · January 1, 2020
Mutual information (MI) minimization has gained considerable interests in various machine learning tasks. However, estimating and minimizing MI in high-dimensional spaces remains a challenging problem, especially when only samples, rather than distribution ...
Cite
Conference37th International Conference on Machine Learning Icml 2020 · January 1, 2020
Recent work has shown generative adversarial networks (GANs) can generate highly realistic images, that are often indistinguishable (by humans) from real images. Most images so generated are not contained in the training dataset, suggesting potential for a ...
Cite
Conference37th International Conference on Machine Learning Icml 2020 · January 1, 2020
Cross-domain alignment between two sets of entities (e.g., objects in an image, words in a sentence) is fundamental to both computer vision and natural language processing. Existing methods mainly focus on designing advanced attention mechanisms to simulat ...
Cite
ConferenceAaai 2020 34th Aaai Conference on Artificial Intelligence · January 1, 2020
Learning to generate text with a given label is a challenging task because natural language sentences are highly variable and ambiguous. It renders difficulties in trade-off between sentence quality and label fidelity. In this paper, we present CARA to all ...
Cite
ConferenceAaai 2020 34th Aaai Conference on Artificial Intelligence · January 1, 2020
Maximum likelihood (ML) and adversarial learning are two popular approaches for training generative models, and from many perspectives these techniques are complementary. ML learning encourages the capture of all data modes, and it is typically characteriz ...
Cite
ConferenceAdvances in Neural Information Processing Systems · January 1, 2020
As a fundamental issue in lifelong learning, catastrophic forgetting is directly caused by inaccessible historical data; accordingly, if the data (information) were memorized perfectly, no forgetting should be expected. Motivated by that, we propose a GAN ...
Cite
ConferenceAdvances in Neural Information Processing Systems · January 1, 2020
We present an approach for lifelong/continual learning of convolutional neural networks (CNN) that does not suffer from the problem of catastrophic forgetting when moving from one task to the other. We show that the activation maps generated by the CNN tra ...
Cite
ConferenceAdvances in Neural Information Processing Systems · January 1, 2020
Synchronization is a key step in data-parallel distributed machine learning (ML). Different synchronization systems and strategies perform differently, and to achieve optimal parallel training throughput requires synchronization strategies that adapt to mo ...
Cite
ConferenceAdvances in Neural Information Processing Systems · January 1, 2020
We consider the blackbox transfer-based targeted adversarial attack threat model in the realm of deep neural network (DNN) image classifiers. Rather than focusing on crossing decision boundaries at the output layer of the source model, our method perturbs ...
Cite
ConferenceAdvances in Neural Information Processing Systems · January 1, 2020
There has been recent interest in exploring generative goals for counterfactual reasoning, e.g., individualized treatment effect (ITE) estimation. However, existing solutions often fail to address issues that are unique to causal inference, such as covaria ...
Cite
ConferenceFindings of the Association for Computational Linguistics Findings of Acl Emnlp 2020 · January 1, 2020
In sequence-to-sequence models, classical optimal transport (OT) can be applied to semantically match generated sentences with target sentences. However, in non-parallel settings, target sentences are usually unavailable. To tackle this issue without losin ...
Cite
ConferenceEmnlp 2020 2020 Conference on Empirical Methods in Natural Language Processing Proceedings of the Conference · January 1, 2020
Neural language models are often trained with maximum likelihood estimation (MLE), where the next word is generated conditioned on the ground-truth word tokens. During testing, however, the model is instead conditioned on previously generated tokens, resul ...
Cite
ConferenceEmnlp 2020 2020 Conference on Empirical Methods in Natural Language Processing Proceedings of the Conference · January 1, 2020
Word embedding models are typically able to capture the semantics of words via the distributional hypothesis, but fail to capture the numerical properties of numbers that appear in a text. This leads to problems with numerical reasoning involving tasks suc ...
Cite
ConferenceProceedings of Machine Learning Research · January 1, 2020
Cross-domain alignment between two sets of entities (e.g., objects in an image, words in a sentence) is fundamental to both computer vision and natural language processing. Existing methods mainly focus on designing advanced attention mechanisms to simulat ...
Cite
ConferenceProceedings of Machine Learning Research · January 1, 2020
Mutual information (MI) minimization has gained considerable interests in various machine learning tasks. However, estimating and minimizing MI in high-dimensional spaces remains a challenging problem, especially when only samples, rather than distribution ...
Cite
ConferenceProceedings of the Annual Meeting of the Association for Computational Linguistics · January 1, 2020
Learning disentangled representations of natural language is essential for many NLP tasks, e.g., conditional text generation, style transfer, personalized dialogue systems, etc. Similar problems have been studied extensively for other forms of data, such a ...
Cite
Conference8th International Conference on Learning Representations Iclr 2020 · January 1, 2020
We investigate new methods for training collaborative filtering models based on actor-critic reinforcement learning, to more directly maximize ranking-based objective functions. Specifically, we train a critic network to approximate ranking-based metrics, ...
Cite
Conference8th International Conference on Learning Representations Iclr 2020 · January 1, 2020
Almost all current adversarial attacks of CNN classifiers rely on information derived from the output layer of the network. This work presents a new adversarial attack based on the modeling and exploitation of class-wise and layer-wise deep feature distrib ...
Cite
Conference31st British Machine Vision Conference Bmvc 2020 · January 1, 2020
We tackle an unsupervised domain adaptation problem for which the domain discrepancy between labeled source and unlabeled target domains is large, due to many factors of inter- and intra-domain variation. While deep domain adaptation methods have been real ...
Cite
Conference31st British Machine Vision Conference Bmvc 2020 · January 1, 2020
As with other deep learning methods, label quality is important for learning modern convolutional object detectors. However, the potentially large number and wide diversity of object instances that can be found in complex image scenes makes constituting co ...
Cite
ConferenceProceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition · June 1, 2019
In this work, we propose a new task called Story Visualization. Given a multi-sentence paragraph, the story is visualized by generating a sequence of images, one for each sentence. In contrast to video generation, story visualization focuses less on the co ...
Full textCite
Conference33rd Aaai Conference on Artificial Intelligence Aaai 2019 31st Innovative Applications of Artificial Intelligence Conference Iaai 2019 and the 9th Aaai Symposium on Educational Advances in Artificial Intelligence Eaai 2019 · January 1, 2019
Learning probability distributions on the weights of neural networks has recently proven beneficial in many applications. Bayesian methods such as Stochastic Gradient Markov Chain Monte Carlo (SG-MCMC) offer an elegant framework to reason about model uncer ...
Full textCite
Conference36th International Conference on Machine Learning Icml 2019 · January 1, 2019
Stochastic blockmodels (SBM) and their variants, e.g., mixed-membership and overlapping stochastic blockmodels, are latent variable based generative models for graphs. They have proven to be successful for various tasks, such as discovering the community s ...
Cite
Conference36th International Conference on Machine Learning Icml 2019 · January 1, 2019
The impact of softmax on the value function itself in reinforcement learning (RL) is often viewed as problematic because it leads to sub-optimal value (or Q) functions and interferes with the contraction properties of the Bellman operator. Surprisingly, de ...
Cite
Conference36th International Conference on Machine Learning Icml 2019 · January 1, 2019
A novel Gromov-Wasserstein learning framework is proposed to jointly match (align) graphs and learn embedding vectors for the associated graph nodes. Using Gromov-Wasserstein discrepancy, we measure the dissimilarity between two graphs and find their corre ...
Cite
Conference36th International Conference on Machine Learning Icml 2019 · January 1, 2019
The generative adversarial network (GAN) has received considerable attention recently as a model for data synthesis, without an explicit specification of a likelihood function. There has been commensurate interest in leveraging likelihood estimates to impr ...
Cite
Conference36th International Conference on Machine Learning Icml 2019 · January 1, 2019
Particle-based variational inference methods (ParVIs) have gained attention in the Bayesian inference literature, for their capacity to yield flexible and accurate approximations. We explore ParVIs from the perspective of Wasscrstcin gradient flows, and ma ...
Cite
ConferenceAistats 2019 22nd International Conference on Artificial Intelligence and Statistics · January 1, 2019
Concerns related to data security and confidentiality have been raised when applying machine learning to real-world applications. Differential privacy provides a principled and rigorous privacy guarantee for machine learning models. While it is common to i ...
Cite
ConferenceAistats 2019 22nd International Conference on Artificial Intelligence and Statistics · January 1, 2019
We investigate adversarial learning in the case when only an unnormalized form of the density can be accessed, rather than samples. With insights so garnered, adversarial learning is extended to the case for which one has access to an unnormalized form u(x ...
Cite
ConferenceAistats 2019 22nd International Conference on Artificial Intelligence and Statistics · January 1, 2019
Thompson sampling (TS) is a class of algorithms for sequential decision making, in which a posterior distribution is maintained over a reward model. However, calculating exact posterior distributions is intractable for all but the simplest models. Developm ...
Cite
Conference7th International Conference on Learning Representations Iclr 2019 · January 1, 2019
Within many machine learning algorithms, a fundamental problem concerns efficient calculation of an unbiased gradient wrt parameters γ for expectation-based objectives Eqγ(y)[f(y)]. Most existing methods either (i) suffer f ...
Cite
Conference5th International Conference on Learning Representations, ICLR 2017 - Workshop Track Proceedings · January 1, 2019
A new model for video captioning is developed, using a deep three-dimensional Convolutional Neural Network (C3D) as an encoder for videos and a Recurrent Neural Network (RNN) as a decoder for captions. A novel attention mechanism with spatiotemporal alignm ...
Cite
Conference7th International Conference on Learning Representations Iclr 2019 · January 1, 2019
Sequence-to-sequence models are commonly trained via maximum likelihood estimation (MLE). However, standard MLE training considers a word-level objective, predicting the next word given the previous ground-truth partial sentence. This procedure focuses on ...
Cite
ConferenceEmnlp Ijcnlp 2019 2019 Conference on Empirical Methods in Natural Language Processing and 9th International Joint Conference on Natural Language Processing Proceedings of the Conference · January 1, 2019
Generating high-quality paraphrases is a fundamental yet challenging natural language processing task. Despite the effectiveness of previous work based on generative models, there remain problems with exposure bias in recurrent neural networks, and often a ...
Full textCite
ConferenceNaacl Hlt 2019 2019 Conference of the North American Chapter of the Association for Computational Linguistics Human Language Technologies Proceedings of the Conference · January 1, 2019
Variational autoencoders (VAEs) with an auto-regressive decoder have been applied for many natural language processing (NLP) tasks. The VAE objective consists of two terms, (i) reconstruction and (ii) KL regularization, balanced by a weighting hyper-parame ...
Cite
ConferenceNaacl Hlt 2019 2019 Conference of the North American Chapter of the Association for Computational Linguistics Human Language Technologies Proceedings of the Conference · January 1, 2019
We propose a topic-guided variational autoencoder (TGVAE) model for text generation. Distinct from existing variational autoencoder (VAE) based approaches, which assume a simple Gaussian prior for the latent code, our model specifies the prior as a Gaussia ...
Cite
ConferenceAdvances in Neural Information Processing Systems · January 1, 2019
Language models are essential for natural language processing (NLP) tasks, such as machine translation and text summarization. Remarkable performance has been demonstrated recently across many NLP domains via a Transformer-based language model with over a ...
Cite
ConferenceAdvances in Neural Information Processing Systems · January 1, 2019
We propose a scalable Gromov-Wasserstein learning (S-GWL) method and establish a novel and theoretically-supported paradigm for large-scale graph analysis. The proposed method is based on the fact that Gromov-Wasserstein discrepancy is a pseudometric on gr ...
Cite
ConferenceAdvances in Neural Information Processing Systems · January 1, 2019
The existence of adversarial data examples has drawn significant attention in the deep-learning community; such data are seemingly minimally perturbed relative to the original data, but lead to very different outputs from a deep-learning algorithm. Althoug ...
Cite
ConferenceAdvances in Neural Information Processing Systems · January 1, 2019
Inference, estimation, sampling and likelihood evaluation are four primary goals of probabilistic modeling. Practical considerations often force modeling approaches to make compromises between these objectives. We present a novel probabilistic learning fra ...
Cite
ConferenceAdvances in Neural Information Processing Systems · January 1, 2019
Text-based interactive recommendation provides richer user feedback and has demonstrated advantages over traditional interactive recommender systems. However, recommendations can easily violate preferences of users from their past natural-language feedback ...
Cite
ConferenceProceedings of Machine Learning Research · January 1, 2019
Stochastic blockmodels (SBM) and their variants, e.g., mixed-membership and overlapping stochastic blockmodels, are latent variable based generative models for graphs. They have proven to be successful for various tasks, such as discovering the community s ...
Cite
ConferenceProceedings of Machine Learning Research · January 1, 2019
The impact of softmax on the value function itself in reinforcement learning (RL) is often viewed as problematic because it leads to sub-optimal value (or Q) functions and interferes with the contrac-tion properties of the Bellman operator. Surpris-ingly, ...
Cite
ConferenceProceedings of Machine Learning Research · January 1, 2019
Particle-based variational inference methods (ParVIs) have gained attention in the Bayesian inference literature, for their capacity to yield flexible and accurate approximations. We explore ParVIs from the perspective of Wasserstein gradient flows, and ma ...
Cite
ConferenceProceedings of Machine Learning Research · January 1, 2019
The generative adversarial network (GAN) has received considerable attention recently as a model for data synthesis, without an explicit specification of a likelihood function. There has been commensurate interest in leveraging likelihood estimates to impr ...
Cite
ConferenceProceedings of Machine Learning Research · January 1, 2019
A novel Gromov-Wasserstein learning framework is proposed to jointly match (align) graphs and learn embedding vectors for the associated graph nodes. Using Gromov-Wasserstein discrepancy, we measure the dissimilarity between two graphs and find their corre ...
Cite
ConferenceProceedings of Machine Learning Research · January 1, 2019
We investigate adversarial learning in the case when only an unnormalized form of the density can be accessed, rather than samples. With insights so garnered, adversarial learning is extended to the case for which one has access to an unnormalized form u(x ...
Cite
ConferenceProceedings of Machine Learning Research · January 1, 2019
Concerns related to data security and confidentiality have been raised when applying machine learning to real-world applications. Differential privacy provides a principled and rigorous privacy guarantee for machine learning models. While it is common to i ...
Cite
ConferenceProceedings of Machine Learning Research · January 1, 2019
Thompson sampling (TS) is a class of algorithms for sequential decision making, in which a posterior distribution is maintained over a reward model. However, calculating exact posterior distributions is intractable for all but the simplest models. Developm ...
Cite
ConferenceProceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition · December 14, 2018
Low-rank signal modeling has been widely leveraged to capture non-local correlation in image processing applications. We propose a new method that employs low-rank tensor factor analysis for tensors generated by grouped image patches. The low-rank tensors ...
Full textCite
ConferenceProgress in Biomedical Optics and Imaging Proceedings of SPIE · January 1, 2018
Detecting an anomaly such as a malignant tumor or a nodule from medical images including mammogram, CT or PET images is still an ongoing research problem drawing a lot of attention with applications in medical diagnosis. A conventional way to address this ...
Full textCite
ConferenceIjcai International Joint Conference on Artificial Intelligence · January 1, 2018
A continuous-time tensor factorization method is developed for event sequences containing multiple “modalities.” Each data element is a point in a tensor, whose dimensions are associated with the discrete alphabet of the modalities. Each tensor data elemen ...
Full textCite
Conference35th International Conference on Machine Learning Icml 2018 · January 1, 2018
Policy optimization is a core component of reinforcement learning (RL), and most existing RL methods directly optimize parameters of a policy based on maximizing the expected total reward, or its surrogate. Though often achieving encouraging empirical succ ...
Cite
Conference35th International Conference on Machine Learning Icml 2018 · January 1, 2018
Two fundamental problems in unsupervised learning are efficient inference for latent-variable models and robust density estimation based on large amounts of unlabeled data. Algorithms for the two tasks, such as normalizing flows and generative adversarial ...
Cite
Conference35th International Conference on Machine Learning Icml 2018 · January 1, 2018
To assess the difference between real and synthetic data, Generative Adversarial Networks (GANs) are trained using a distribution discrepancy measure. Three widely employed measures are information-theoretic divergences, integral probability metrics, and H ...
Cite
Conference35th International Conference on Machine Learning Icml 2018 · January 1, 2018
Recent advances on the scalability and flexibility of variational inference have made it successful at unravelling hidden patterns in complex data. In this work we propose a new variational bound formulation, yielding an estimator that extends beyond the c ...
Cite
Conference35th International Conference on Machine Learning Icml 2018 · January 1, 2018
A parametric point process model is developed, with modeling based on the assumption that sequential observations often share latent phenomena, while also possessing idiosyncratic effects. An alternating optimization method is proposed to learn a "register ...
Cite
Conference32nd Aaai Conference on Artificial Intelligence Aaai 2018 · January 1, 2018
We present a deep generative model for Zero-Shot Learning (ZSL). Unlike most existing methods for this problem, that represent each class as a point (via a semantic embedding), we represent each seen/unseen class using a class-specific latent-space distrib ...
Cite
Conference32nd Aaai Conference on Artificial Intelligence Aaai 2018 · January 1, 2018
Previous models for video captioning often use the output from a specific layer of a Convolutional Neural Network (CNN) as video features. However, the variable context-dependent semantics in the video may make it more appropriate to adaptively select feat ...
Cite
ConferenceAdvances in Neural Information Processing Systems · January 1, 2018
Textual network embedding leverages rich text information associated with the network to learn low-dimensional vectorial representations of vertices. Rather than using typical natural language processing (NLP) approaches, recent research exploits the relat ...
Cite
ConferenceAdvances in Neural Information Processing Systems · January 1, 2018
We propose a novel Wasserstein method with a distillation mechanism, yielding joint learning of word embeddings and topics. The proposed method is based on the fact that the Euclidean distance between word embeddings may be employed as the underlying dista ...
Cite
ConferenceAdvances in Neural Information Processing Systems · January 1, 2018
Generative adversarial networks (GANs) have achieved significant success in generating real-valued data. However, the discrete nature of text hinders the application of GAN to text-generation tasks. Instead of using the standard GAN objective, we propose t ...
Cite
ConferenceInternational Conference on Artificial Intelligence and Statistics Aistats 2018 · January 1, 2018
A new form of the variational autoencoder (VAE) is proposed, based on the symmetric Kullback-Leibler divergence. It is demonstrated that learning of the resulting symmetric VAE (sVAE) has close connections to previously developed adversarial-learning metho ...
Cite
ConferenceInternational Conference on Artificial Intelligence and Statistics Aistats 2018 · January 1, 2018
Learning probability distributions on the weights of neural networks (NNs) has recently proven beneficial in many applications. Bayesian methods, such as Stein variational gradient descent (SVGD), offer an elegant framework to reason about NN model uncerta ...
Cite
ConferenceInternational Conference on Artificial Intelligence and Statistics Aistats 2018 · January 1, 2018
The superposition of temporal point processes has been studied for many years, although the usefulness of such models for practical applications has not be fully developed. We investigate superposed Hawkes process as an important class of such models, with ...
Cite
ConferenceInternational Conference on Artificial Intelligence and Statistics Aistats 2018 · January 1, 2018
We propose a Topic Compositional Neural Language Model (TCNLM), a novel method designed to simultaneously capture both the global semantic meaning and the local word-ordering structure in a document. The TCNLM learns the global semantic coherence of a docu ...
Cite
ConferenceProceedings of Machine Learning Research · January 1, 2018
Health risks from cigarette smoking - the leading cause of preventable death in the United States - can be substantially reduced by quitting. Although most smokers are motivated to quit, the majority of quit attempts fail. A number of studies have explored ...
Cite
ConferenceProceedings of the 2018 Conference on Empirical Methods in Natural Language Processing Emnlp 2018 · January 1, 2018
Convolutional neural networks (CNNs) have recently emerged as a popular building block for natural language processing (NLP). Despite their success, most existing CNN models employed in NLP share the same learned (and static) set of filters for all input s ...
Cite
ConferenceProceedings of Machine Learning Research · January 1, 2018
We propose a Topic Compositional Neural Language Model (TCNLM), a novel method designed to simultaneously capture both the global semantic meaning and the local wordordering structure in a document. The TCNLM learns the global semantic coherence of a docum ...
Cite
ConferenceProceedings of Machine Learning Research · January 1, 2018
A new form of the variational autoencoder (VAE) is proposed, based on the symmetric KullbackLeibler divergence. It is demonstrated that learning of the resulting symmetric VAE (sVAE) has close connections to previously developed adversarial-learning method ...
Cite
ConferenceProceedings of Machine Learning Research · January 1, 2018
Learning probability distributions on the weights of neural networks (NNs) has recently proven beneficial in many applications. Bayesian methods, such as Stein variational gradient descent (SVGD), offer an elegant framework to reason about NN model uncerta ...
Cite
ConferenceProceedings of Machine Learning Research · January 1, 2018
Policy optimization is a core component of reinforcement learning (RL), and most existing RL methods directly optimize parameters of a policy based on maximizing the expected total reward, or its surrogate. Though often achieving encouraging empirical succ ...
Cite
ConferenceProceedings of Machine Learning Research · January 1, 2018
Two fundamental problems in unsupervised learning are efficient inference for latent-variable models and robust density estimation based on large amounts of unlabeled data. Algorithms for the two tasks, such as normalizing flows and generative adversarial ...
Cite
ConferenceProceedings of Machine Learning Research · January 1, 2018
A new generative adversarial network is developed for joint distribution matching. Distinct from most existing approaches, that only learn conditional distributions, the proposed model aims to learn a joint distribution of multiple random variables (domain ...
Cite
ConferenceProceedings of Machine Learning Research · January 1, 2018
Recent advances on the scalability and flexibility of variational inference have made it successful at unravelling hidden patterns in complex data. In this work we propose a new variational bound formulation, yielding an estimator that extends beyond the c ...
Cite
ConferenceProceedings of Machine Learning Research · January 1, 2018
To assess the difference between real and synthetic data, Generative Adversarial Networks (GANs) are trained using a distribution discrepancy measure. Three widely employed measures are information-theoretic divergences, integral probability metrics, and H ...
Cite
ConferenceProceedings of Machine Learning Research · January 1, 2018
A parametric point process model is developed, with modeling based on the assumption that sequential observations often share latent phenomena, while also possessing idiosyncratic effects. An alternating optimization method is proposed to learn a “register ...
Cite
ConferenceProceedings of Machine Learning Research · January 1, 2018
The superposition of temporal point processes has been studied for many years, although the usefulness of such models for practical applications has not be fully developed. We investigate superposed Hawkes process as an important class of such models, with ...
Cite
ConferenceProceedings 30th IEEE Conference on Computer Vision and Pattern Recognition Cvpr 2017 · November 6, 2017
A Semantic Compositional Network (SCN) is developed for image captioning, in which semantic concepts (i.e., tags) are detected from the image, and the probability of each tag is used to compose the parameters in a long short-term memory (LSTM) network. The ...
Full textCite
ConferenceProceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining · August 13, 2017
Extensive information on 3 million randomly sampled United States citizens is used to construct a statistical model of constituent preferences for each U.S. congressional district. This model is linked to the legislative voting record of the legislator fro ...
Full textCite
Conference31st Aaai Conference on Artificial Intelligence Aaai 2017 · January 1, 2017
Gaussian graphical models (GGMs) are widely used for statistical modeling, because of ease of inference and the ubiquitous use of the normal distribution in practical approximations. However, they are also known for their limited modeling abilities, due to ...
Cite
ConferenceAcl 2017 55th Annual Meeting of the Association for Computational Linguistics Proceedings of the Conference Long Papers · January 1, 2017
Recurrent neural networks (RNNs) have shown promising performance for language modeling. However, traditional training of RNNs using back-propagation through time often suffers from overfitting. One reason for this is that stochastic optimization (used for ...
Full textCite
ConferenceEmnlp 2017 Conference on Empirical Methods in Natural Language Processing Proceedings · January 1, 2017
We propose a new encoder-decoder approach to learn distributed sentence representations that are applicable to multiple purposes. The model is learned by using a convolutional neural network as an encoder to map an input sentence into a continuous vector, ...
Full textCite
ConferenceAdvances in Neural Information Processing Systems · January 1, 2017
We consider the analysis of Electroencephalography (EEG) and Local Field Potential (LFP) datasets, which are "big" in terms of the size of recorded data but rarely have sufficient labels required to train complex models (e.g., conventional deep learning me ...
Cite
ConferenceAdvances in Neural Information Processing Systems · January 1, 2017
In neuropsychiatric disorders such as schizophrenia or depression, there is often a disruption in the way that regions of the brain synchronize with one another. To facilitate understanding of network-level synchronization between brain regions, we introdu ...
Cite
ConferenceAdvances in Neural Information Processing Systems · January 1, 2017
A Triangle Generative Adversarial Network (Δ-GAN) is developed for semi-supervised cross-domain joint distribution matching, where the training data consists of samples from each domain, and supervision of domain correspondence is provided by only a few pa ...
Cite
ConferenceAdvances in Neural Information Processing Systems · January 1, 2017
A new method for learning variational autoencoders (VAEs) is developed, based on Stein variational gradient descent. A key advantage of this approach is that one need not make parametric assumptions about the form of the encoder distribution. Performance i ...
Cite
ConferenceAdvances in Neural Information Processing Systems · January 1, 2017
We investigate the non-identifiability issues associated with bidirectional adversarial training for joint distribution matching. Within a framework of conditional entropy, we propose both adversarial and non-adversarial approaches to learn desirable match ...
Cite
ConferenceAdvances in Neural Information Processing Systems · January 1, 2017
We propose a scalable algorithm for model selection in sigmoid belief networks (SBNs), based on the factorized asymptotic Bayesian (FAB) framework. We derive the corresponding generalized factorized information criterion (gFIC) for the SBN, which is proven ...
Cite
ConferenceAdvances in Neural Information Processing Systems · January 1, 2017
We present a probabilistic framework for nonlinearities, based on doubly truncated Gaussian distributions. By setting the truncation points appropriately, we are able to generate various types of nonlinearities within a unified framework, including sigmoid ...
Cite
ConferenceAdvances in Neural Information Processing Systems · January 1, 2017
We propose a new method that uses deep learning techniques to accelerate the popular alternating direction method of multipliers (ADMM) solution for inverse problems. The ADMM updates consist of a proximity operator, a least squares regression that include ...
Cite
Conference34th International Conference on Machine Learning Icml 2017 · January 1, 2017
We present a probabilistic framework for overlapping community discovery and link prediction for relational data, given as a graph. The proposed framework has: (1) a deep architecture which enables us to infer multiple layers of latent features/communities ...
Cite
ConferenceProceedings of the 20th International Conference on Artificial Intelligence and Statistics, AISTATS 2017 · January 1, 2017
Copyright 2017 by the author(s). A multi-way factor analysis model is introduced for tensor-variate data of any order. Each data item is represented as a (sparse) sum of Kruskal decompositions, a Kruskal-factor analysis (KFA). KFA is nonparametric and can ...
Cite
ConferenceProceedings of the 20th International Conference on Artificial Intelligence and Statistics Aistats 2017 · January 1, 2017
Deep neural networks (DNNs) are increasingly popular in modern machine learning. Bayesian learning affords the opportunity to quantify posterior uncertainty on DNN model parameters. Most existing work adopts independent Gaussian priors on the model weights ...
Cite
ConferenceProceedings of the 20th International Conference on Artificial Intelligence and Statistics Aistats 2017 · January 1, 2017
A multi-way factor analysis model is introduced for tensor-variate data of any order. Each data item is represented as a (sparse) sum of Kruskal decompositions, a Kruskal-factor analysis (KFA). KFA is nonparametric and can infer both the tensor-rank of eac ...
Cite
Conference5th International Conference on Learning Representations Iclr 2017 Workshop Track Proceedings · January 1, 2017
A new model for video captioning is developed, using a deep three-dimensional Convolutional Neural Network (C3D) as an encoder for videos and a Recurrent Neural Network (RNN) as a decoder for captions. A novel attention mechanism with spatiotemporal alignm ...
Cite
ConferenceProceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition · December 9, 2016
Learning the representation of shape cues in 2D & 3D objects for recognition is a fundamental task in computer vision. Deep neural networks (DNNs) have shown promising performance on this task. Due to the large variability of shapes, accurate recognition r ...
Full textCite
ConferenceIEEE Transactions on Information Theory · November 1, 2016
This paper offers a characterization of fundamental limits on the classification and reconstruction of high-dimensional signals from low-dimensional features, in the presence of side information. We consider a scenario where a decoder has access both to li ...
Full textCite
Conference2015 IEEE Nuclear Science Symposium and Medical Imaging Conference NSS Mic 2015 · October 3, 2016
We consider X-ray coherent scatter imaging, where the goal is to reconstruct momentum transfer profiles (spectral distributions) at each spatial location from multiplexed measurements of scatter. Each material is characterized by a unique momentum transfer ...
Full textCite
ConferenceProceedings IEEE International Conference on Data Mining Icdm · July 2, 2016
We introduce a novel dynamic model for discrete time-series data, in which the temporal sampling may be nonuniform. The model is specified by constructing a hierarchy of Poisson factor analysis blocks, one for the transitions between latent states and the ...
Full textCite
ConferenceICASSP IEEE International Conference on Acoustics Speech and Signal Processing Proceedings · May 18, 2016
We develop a general framework for compressive linear-projection measurements with side information. Side information is an additional signal correlated with the signal of interest. We investigate the impact of side information on classification and signal ...
Full textCite
ConferenceProgress in Biomedical Optics and Imaging Proceedings of SPIE · January 1, 2016
Coded aperture X-ray diffraction (coherent scatter spectral) imaging provides fast and dose-efficient measurements of the molecular structure of an object. The information provided is spatially-dependent and material-specific, and can be utilized in medica ...
Full textCite
ConferenceAaai Spring Symposium Technical Report · January 1, 2016
We present a new algorithm called PIEM to approximately solve for the policy of an infinite-horizon decentralized partially observable Markov decision process (DEC-POMDP). The algorithm uses expectation maximization (EM) only in the step of policy improvem ...
Cite
ConferenceProceedings of SPIE the International Society for Optical Engineering · January 1, 2016
Coded aperture X-ray coherent scatter imaging is a novel modality for ascertaining the molecular structure of an object. Measurements from different spatial locations and spectral channels in the object are multiplexed through a radiopaque material (coded ...
Full textCite
ConferenceProceedings of SPIE the International Society for Optical Engineering · January 1, 2016
A long-term goal for checked baggage screening in airports has been to include passenger information, or at least a predetermined passenger risk level, in the screening process. One method for including that information could be treating the checked baggag ...
Full textCite
ConferenceLecture Notes in Computer Science Including Subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics · January 1, 2016
We present Deep Stochastic Neighbor Compression (DSNC), a framework to compress training data for instance-based methods (such as k-nearest neighbors). We accomplish this by inferring a smaller set of pseudo-inputs in a new feature space learned by a deep ...
Full textCite
ConferenceLecture Notes in Computer Science Including Subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics · January 1, 2016
We proposed a Hamiltonian Monte Carlo (HMC) method with Laplace kinetic energy, and demonstrate the connection between slice sampling and proposed HMC method in one-dimensional cases. Based on this connection, one can perform slice sampling using a numeric ...
Full textCite
Conference33rd International Conference on Machine Learning Icml 2016 · January 1, 2016
We introduce the truncated Gaussian graphical model (TGGM) as a novel framework for designing statistical models for nonlinear learning. A TGGM is a Gaussian graphical model (GGM) with a subset of variables truncated to be nonneg- Ative. The truncated vari ...
Cite
Conference33rd International Conference on Machine Learning Icml 2016 · January 1, 2016
Deep conditional generative models are developed to simultaneously learn the temporal dependencies of multiple sequences. The model is designed by introducing a three-way weight tensor to capture the multiplicative interactions between side information and ...
Cite
ConferenceIjcai International Joint Conference on Artificial Intelligence · January 1, 2016
In dictionary learning for analysis of images, spatial correlation from extracted patches can be leveraged to improve characterization power. We propose a Bayesian framework for dictionary learning, with spatial location dependencies captured by imposing a ...
Cite
Conference30th Aaai Conference on Artificial Intelligence Aaai 2016 · January 1, 2016
Learning in deep models using Bayesian methods has generated significant attention recently. This is largely because of the feasibility of modern Bayesian methods to yield scalable learning and inference, while maintaining a measure of uncertainty in the m ...
Cite
ConferenceAdvances in Neural Information Processing Systems · January 1, 2016
We unify slice sampling and Hamiltonian Monte Carlo (HMC) sampling, demonstrating their connection via the Hamiltonian-Jacobi equation from Hamiltonian mechanics. This insight enables extension of HMC and slice sampling to a broader family of samplers, cal ...
Cite
ConferenceAdvances in Neural Information Processing Systems · January 1, 2016
A novel variational autoencoder is developed to model images, as well as associated labels or captions. The Deep Generative Deconvolutional Network (DGDN) is used as a decoder of the latent image features, and a deep Convolutional Neural Network (CNN) is u ...
Cite
ConferenceAdvances in Neural Information Processing Systems · January 1, 2016
Feature construction is of vital importance in reinforcement learning, as the quality of a value function or policy is largely determined by the corresponding features. The recent successes of deep reinforcement learning (RL) only increase the importance o ...
Cite
ConferenceAdvances in Neural Information Processing Systems · January 1, 2016
Stochastic gradient MCMC (SG-MCMC) has played an important role in large-scale Bayesian learning, with well-developed theoretical convergence properties. In such applications of SG-MCMC, it is becoming increasingly popular to employ distributed systems, wh ...
Cite
ConferenceProceedings of the 19th International Conference on Artificial Intelligence and Statistics Aistats 2016 · January 1, 2016
A deep generative model is developed for representation and analysis of images, based on a hierarchical convolutional dictionary-learning framework. Stochastic unpooling is employed to link consecutive layers in the model, yielding top-down image generatio ...
Cite
ConferenceProceedings of the 19th International Conference on Artificial Intelligence and Statistics Aistats 2016 · January 1, 2016
We present a scalable probabilistic framework for learning from multi-relational data, given in form of entity-relation-entity triplets, with a potentially massive number of entities and relations (e.g., in multi-relational networks, knowledge bases, etc.) ...
Cite
ConferenceProceedings of the 19th International Conference on Artificial Intelligence and Statistics Aistats 2016 · January 1, 2016
We utilize copulas to constitute a unified framework for constructing and optimizing variational proposals in hierarchical Bayesian models. For models with continuous and non-Gaussian hidden variables, we propose a semiparametric and automated variational ...
Cite
ConferenceProceedings of the 19th International Conference on Artificial Intelligence and Statistics Aistats 2016 · January 1, 2016
We present a probabilistic framework for efficient non-negative matrix factorization of discrete (count/binary) data with side-information. The side-information is given as a multi-level structure, taxonomy, or ontology, with nodes at each level being cate ...
Cite
ConferenceIEEE International Symposium on Information Theory Proceedings · September 28, 2015
Classical compressive sensing typically assumes a single measurement, and theoretical analysis often relies on corresponding concentration-of-measure results. There are many real-world applications involving multiple compressive measurements, from which th ...
Full textCite
ConferenceIEEE International Symposium on Information Theory Proceedings · September 28, 2015
This paper offers a characterization of performance limits for classification and reconstruction of high-dimensional signals from noisy compressive measurements, in the presence of side information. We assume the signal of interest and the side information ...
Full textCite
ConferenceProceedings of the National Conference on Artificial Intelligence · June 1, 2015
We present a probabilistic model for tensor decomposition where one or more tensor modes may have sideinformation about the mode entities in form of their features and/or their adjacency network. We consider a Bayesian approach based on the Canonical PARAF ...
Cite
ConferenceProceedings of the National Conference on Artificial Intelligence · June 1, 2015
We present a probabilistic framework for learning pairwise similarities between objects belonging to different modalities, such as drugs and proteins, or text and images. Our framework is based on learning a binary code based representation for objects in ...
Cite
ConferenceProceedings of the National Conference on Artificial Intelligence · June 1, 2015
We present a probabilistic framework for learning with heterogeneous multiview data where some views are given as ordinal, binary, or real-valued feature matrices, and some views as similarity matrices. Our framework has the following distinguishing aspect ...
Cite
ConferenceProgress in Biomedical Optics and Imaging Proceedings of SPIE · January 1, 2015
We propose an alternating minimization (AM) algorithm for estimating attenuation functions in X-ray transmission tomography using priors that promote sparsity in the pixel/voxel differences domain. As opposed to standard maximum-a-posteriori (MAP) estimati ...
Full textCite
ConferenceUncertainty in Artificial Intelligence - Proceedings of the 31st Conference, UAI 2015 · January 1, 2015
We present a scalable Bayesian model for lowrank factorization of massive tensors with binary observations. The proposed model has the following key properties: (1) in contrast to the models based on the logistic or probit likelihood, using a zero-truncate ...
Cite
Conference32nd International Conference on Machine Learning Icml 2015 · January 1, 2015
Point process data are commonly observed in fields like healthcare and the social sciences. Designing predictive models for such event streams is an under-explored problem, due to often scarce training data. In this work we propose a multitask point proces ...
Cite
ConferenceLecture Notes in Computer Science · January 1, 2015
We present a Bayesian non-negative tensor factorization model for count-valued tensor data, and develop scalable inference algorithms (both batch and online) for dealing with massive tensors. Our generative model can handle overdispersed counts as well as ...
Full textCite
ConferenceIjcai International Joint Conference on Artificial Intelligence · January 1, 2015
Expectation maximization (EM) has recently been shown to be an efficient algorithm for learning finite-state controllers (FSCs) in large decentralized POMDPs (Dec-POMDPs). However, current methods use fixed-size FSCs and often converge to maxima that are f ...
Cite
ConferenceIjcai International Joint Conference on Artificial Intelligence · January 1, 2015
Tensor factorization methods provide a useful way to extract latent factors from complex multirelational data, and also for predicting missing data. Developing tensor factorization methods for massive tensors, especially when the data are binary- or count- ...
Cite
ConferenceAdvances in Neural Information Processing Systems · January 1, 2015
We present a scalable Bayesian multi-label learning model based on learning lowdimensional label embeddings. Our model assumes that each label vector is generated as a weighted combination of a set of topics (each topic being a distribution over labels), w ...
Cite
ConferenceAdvances in Neural Information Processing Systems · January 1, 2015
We propose a new deep architecture for topic modeling, based on Poisson Factor Analysis (PFA) modules. The model is composed of a Poisson distribution to model observed vectors of counts, as well as a deep hierarchy of hidden binary units. Rather than usin ...
Cite
ConferenceAdvances in Neural Information Processing Systems · January 1, 2015
Recent advances in Bayesian learning with large-scale data have witnessed emergence of stochastic gradient MCMC algorithms (SG-MCMC), such as stochastic gradient Langevin dynamics (SGLD), stochastic gradient Hamiltonian MCMC (SGHMC), and the stochastic gra ...
Cite
ConferenceUncertainty in Artificial Intelligence Proceedings of the 31st Conference Uai 2015 · January 1, 2015
We present a scalable Bayesian model for lowrank factorization of massive tensors with binary observations. The proposed model has the following key properties: (1) in contrast to the models based on the logistic or probit likelihood, using a zero-truncate ...
Cite
Conference · January 1, 2015
Video camera architects must design cameras capable of high-quality, dynamic event capture, while adhering to power and communications constraints. Though modern imagers are capable of both simultaneous spatial and temporal resolutions at micrometer and mi ...
Full textCite
Conference3rd International Conference on Learning Representations Iclr 2015 Workshop Track Proceedings · January 1, 2015
A generative model is developed for deep (multi-layered) convolutional dictionary learning. A novel probabilistic pooling operation is integrated into the deep model, yielding efficient bottom-up (pretraining) and top-down (refinement) probabilistic learni ...
Cite
ConferenceOptics Infobase Conference Papers · January 1, 2015
An information-theoretical adaptive sensing and classification framework is proposed for Quadrupole mass filter systems. Simulation results demonstrate significant reduction in number of measurement and improvement of classification accuracy using the adap ...
Cite
ConferenceOptics Infobase Conference Papers · January 1, 2015
We present a compressive camera that combines mechanical translation and spectral dispersion to compress a multi-spectral, high-speed scene onto a monochrome, video-rate detector. Single-frame reconstructions of 15 spectral channels and 10 temporal frames ...
Cite
ConferenceProceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition · September 24, 2014
A simple and inexpensive (low-power and low-bandwidth) modification is made to a conventional off-the-shelf color video camera, from which we recover multiple color frames for each of the original measured frames, and each of the recovered frames can be fo ...
Full textCite
ConferenceProceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition · September 24, 2014
The capture of multiple images is a simple way to increase the chance of capturing a good photo with a light-weight hand-held camera, for which the camera-shake blur is typically a nuisance problem. The naive approach of selecting the single best captured ...
Full textCite
ConferenceOptics InfoBase Conference Papers · January 1, 2014
This talk will review recent developments in the use of statistical methods for inversion of data that are acquired compressively. A particular focus will be placed on dictionary learning and its connection to mixture models. It will be explained how these ...
Cite
Conference31st International Conference on Machine Learning Icml 2014 · January 1, 2014
We investigate design of general nonlinear functions for mapping high-dimensional data into a lower-dimensional (compressive) space. The nonlinear measurements are assumed contaminated by additive Gaussian noise. Depending on the application, we are either ...
Cite
Conference31st International Conference on Machine Learning Icml 2014 · January 1, 2014
We present a scalable Bayesian framework for low-rank decomposition of multiway tensor data with missing observations. The key issue of pre-specifying the rank of the decomposition is sidestepped in a principled manner using a multiplicative gamma process ...
Cite
Conference31st International Conference on Machine Learning Icml 2014 · January 1, 2014
2014 The analysis of correlated point process data has wide applications, ranging from biomedical research to network analysis. In this work, we model such data as generated by a latent collection of continuous-time binary semi-Markov processes,' correspon ...
Cite
ConferenceAdvances in Neural Information Processing Systems · January 1, 2014
This paper is concerned with compressive sensing of signals drawn from a Gaussian mixture model (GMM) with sparse precision matrices. Previous work has shown: (i) a signal drawn from a given GMM can be perfectly reconstructed from r noise-free measurements ...
Cite
ConferenceAdvances in Neural Information Processing Systems · January 1, 2014
We propose a semi-parametric and dynamic rank factor model for topic modeling, capable of (i) discovering topic prevalence over time, and (ii) learning contemporary multi-scale dependence structures, providing topic and word correlations as a byproduct. Th ...
Cite
ConferenceAdvances in Neural Information Processing Systems · January 1, 2014
A new Bayesian formulation is developed for nonlinear support vector machines (SVMs), based on a Gaussian process and with the SVM hinge loss expressed as a scaled mixture of normals. We then integrate the Bayesian SVM into a factor model, in which feature ...
Cite
ConferenceOptics Infobase Conference Papers · January 1, 2014
This talk will review recent developments in the use of statistical methods for inversion of data that are acquired compressively. A particular focus will be placed on dictionary learning and its connection to mixture models. It will be explained how these ...
Full textCite
ConferenceProceedings of the 6th International Conference on Educational Data Mining Edm 2013 · January 1, 2013
Consider a large database of questions that assess the knowledge of learners on a range of different concepts. In this paper, we study the problem of maximizing the estimation accuracy of each learner’s knowledge about a concept while minimizing the number ...
Cite
ConferenceFrontiers in Optics Fio 2012 · January 1, 2012
Blind compressive sensing (CS) is considered for reconstruction of hyperspectral data imaged by a coded aperture camera. The measurements are manifested as a superposition of the coded wavelengthdependent data, with the ambient three-dimensional hyperspect ...
Full textCite
ConferenceAdvances in Neural Information Processing Systems 24 25th Annual Conference on Neural Information Processing Systems 2011 Nips 2011 · January 1, 2011
A new Lévy process prior is proposed for an uncountable collection of covariate-dependent feature-learning measures; the model is called the kernel beta process (KBP). Available covariates are handled efficiently via the kernel construction, with covariate ...
Cite
ConferenceAdvances in Neural Information Processing Systems 24 25th Annual Conference on Neural Information Processing Systems 2011 Nips 2011 · January 1, 2011
The nested Chinese restaurant process is extended to design a nonparametric topic-model tree for representation of human choices. Each tree path corresponds to a type of person, and each node (topic) has a corresponding probability vector over items that m ...
Cite
ConferenceAdvances in Neural Information Processing Systems 23 24th Annual Conference on Neural Information Processing Systems 2010 Nips 2010 · January 1, 2010
We consider problems for which one has incomplete binary matrices that evolve with time (e:g:, the votes of legislators on particular legislation, with each year characterized by a different such matrix). An objective of such analysis is to infer structure ...
Cite
ConferenceAdvances in Neural Information Processing Systems 20 - Proceedings of the 2007 Conference · January 1, 2008
A semi-supervised multitask learning (MTL) framework is presented, in which M parameterized semi-supervised classifiers, each associated with one of M partially labeled data manifolds, are learned jointly under the constraint of a soft-sharing prior impose ...
Cite
ConferenceAdvances in Neural Information Processing Systems 20 Proceedings of the 2007 Conference · January 1, 2008
A semi-supervised multitask learning (MTL) framework is presented, in which M parameterized semi-supervised classifiers, each associated with one of M partially labeled data manifolds, are learned jointly under the constraint of a soft-sharing prior impose ...
Cite
ConferenceProceedings of SPIE the International Society for Optical Engineering · October 24, 2005
Last year, we reported on a preliminary evaluation of GE's frequency-domain EMI prototype sensor capable of measuring the wideband response of simulant and inert low metal mines at shallow depths over a frequency range from 100 Hz to 150 kHz. Since then, t ...
Full textCite
ConferenceAdvances in Neural Information Processing Systems · January 1, 2005
A graph-based prior is proposed for parametric semi-supervised classification. The prior utilizes both labelled and unlabelled data; it also integrates features from multiple views of a given sample (e.g., multiple sensors), thus implementing a Bayesian fo ...
Cite
ConferenceICASSP IEEE International Conference on Acoustics Speech and Signal Processing Proceedings · September 27, 2004
A method to detect airports in large aerial optical imagery is considered. Combining texture segmentation and shape detection, this method shows advantages in analyzing large aerial imagery. First, large aerial images are segmented and interpreted accordin ...
Cite
ConferenceICASSP IEEE International Conference on Acoustics Speech and Signal Processing Proceedings · September 27, 2004
An information-theoretic approach is developed for target detection, with active selection of training set, directly from the site-specific measured data For the proposed kernel-based algorithm, a set of basis functions are defined first to characterize th ...
Cite
ConferenceProceedings of SPIE the International Society for Optical Engineering · November 26, 2003
Unexploded ordnance (UXO) discrimination is investigated using the wide band electromagnetic induction (EMI) data. The main focus of this paper is on the practical phenomenological modeling for the induced wideband EMI sensor response from different target ...
Full textCite
ConferenceProceedings of SPIE the International Society for Optical Engineering · November 26, 2003
Detection and remediation of unexploded ordnance (UXO) represents a major challenge. The detection problem is exacerbated by the fact that on sites contaminated with UXO, extensive surface and sub-surface clutter and shrapnel is also present. Traditional m ...
Full textCite
ConferenceConference Record of the Asilomar Conference on Signals Systems and Computers · January 1, 2003
In the search for diagnostic and therapeutic strategies for lung cancer, matrix-assisted laser desorption/ionization time-of-flight mass spectrometry (MALDI-TOF MS) has been evinced as a new and promising discovery platform to generate protein expression p ...
Full textCite
ConferenceProceedings of the Annual International Conference on Computational Molecular Biology RECOMB · January 1, 2003
Recent research has demonstrated quite convincingly that accurate cancer diagnosis can be achieved by constructing classifiers that arc designed to compare the gene expression profile of a tissue of unknown cancer status to a database of stored expression ...
Full textCite
ConferenceInternational Geoscience and Remote Sensing Symposium IGARSS · January 1, 2002
Detection and remediation of unexploded ordnance (UXO) represents a major challenge on closed, closing, and transferred military ranges as well as on active installations. The detection problem is exacerbated by the fact that on sites contaminated with UXO ...
Cite
ConferenceProceedings of the IEEE Sensor Array and Multichannel Signal Processing Workshop · January 1, 2002
A new algorithm is developed for independent component analysis (ICA) with or without constraints on the mixing matrix or sources. The algorithm is based on the criterion of Joint Approximate Diagonalization of Eigen-matrices (JADE). We propose a column-wi ...
Full textCite
ConferenceProceedings of SPIE the International Society for Optical Engineering · January 1, 2000
Traditional algorithms for UXO remediation experience severe difficulties distinguishing buried targets from anthropic clutter, and in most cases UXO items are found amongst extensive surface clutter and shrapnel from ordnance operations. These problems re ...
Full textCite
ConferenceProceedings of SPIE the International Society for Optical Engineering · January 1, 2000
Traditionally, field EMI sensors are operated in the time-domain. The time-domain (TD) EMI sensor usually is a pulsed system. It contains both a transmitting coil and a receiving coil. After transmitting an excitation pulse, which generates the primary fie ...
Cite
ConferenceInternational Geoscience and Remote Sensing Symposium IGARSS · December 1, 1999
A study is carried out to investigate sub-optimal detectors that continue to incorporate the physical nature of the wideband frequency-domain electromagnetic induction (EMI) signal, but are less computationally burdensome. In addition, a comparison is made ...
Cite
ConferenceIEEE Antennas and Propagation Society International Symposium Wireless Technologies and Information Networks Aps 1999 Held in Conjunction with Usnc Ursi National Radio Science Meeting · January 1, 1999
We demonstrate the accuracy of the half-space fast multipole method (FMM) by considering two targets: a model unexploded ordnance (UXO) buried under soil and a rectangular box situated above the ground. In both examples, the bistatic radar cross sections ( ...
Full textCite
ConferenceProceedings of SPIE the International Society for Optical Engineering · January 1, 1999
Nuclear quadrupole resonance (NQR) is a technique that discriminates mines from clutter by exploiting unique properties of explosives, rather than the attributes of the mine that exist in many forms of anthropic clutter (e.g., metal content). After excitin ...
Full textCite
ConferenceProceedings of SPIE the International Society for Optical Engineering · December 1, 1998
A principal problem with traditional, narrowband EMI sensors involves target identification. As a consequence, in minefield or unexploded ordinance (UXO) detection, for example, each piece of buried metal must be excavated, causing significant false alarms ...
Full textCite
ConferenceProceedings of SPIE the International Society for Optical Engineering · December 1, 1998
In this paper we model time-domain plane-wave scattering from targets buried under a rough (random) air-ground interface. The properties of the interface are parametrized as a random process with known statistics. Since the fields incident upon a buried ta ...
Full textCite
ConferenceLecture Notes in Computer Science Including Subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics · January 1, 1997
In this paper we propose a neural approach based on the Random Neural Network (RNN) model (Gelenbe 1989, 1990, 1991, 1993 [3, 4, 6, 5]), to detect shaped targets with the help of multiple neural networks whose outputs are combined for making decisions. ...
Full textCite
ConferenceProceedings of SPIE the International Society for Optical Engineering · June 20, 1995
Over the years, many different sensor types have been evaluated in an attempt to satisfy the need to detect and discriminate tactical and strategic targets concealed in foliage or underground. In large measure these early efforts were disappointing because ...
Full textCite
ConferenceProceedings of SPIE the International Society for Optical Engineering · August 1, 1991
Recent developments make it possible to radiate and coherently detect electromagnetic pulses consisting of a few half-cycles of a sine wave having a period on the order of lOps. The antennas involved are compact, typically consisting of conducting films on ...
Full textCite