Journal articleJournal of Privacy and Confidentiality · January 1, 2026
We develop formal privacy mechanisms for releasing statistics from data with many outlying values, such as income data. These mechanisms ensure that a per-record differential privacy guarantee degrades slowly in the protected records’ influence on the stat ...
Full textCite
Journal articleVLDB Journal · March 1, 2025
Differential privacy (DP) is the state-of-the-art and rigorous notion of privacy for answering aggregate database queries while preserving the privacy of sensitive information in the data. In today’s era of data analysis, however, it poses new challenges f ...
Full textCite
Journal articleCommunications of the ACM · March 1, 2025
Answering SPJA queries under differential privacy (DP), including graph-pattern counting under node-DP as an important special case, has received considerable attention in recent years. The dual challenge of foreign-key constraints and self-joins is partic ...
Full textCite
Journal articleACM Transactions on Database Systems · November 8, 2024
Answering SPJA queries under differential privacy (DP), including graph pattern counting under node-DP as an important special case, has received considerable attention in recent years. The dual challenge of foreign-key constraints combined with self-joins ...
Full textCite
Journal articleJournal of Privacy and Confidentiality · August 31, 2023
In this work we describe the High-Dimensional Matrix Mechanism (HDMM), a differentially private algorithm for answering a workload of predicate counting queries. HDMM represents query workloads using a compact implicit matrix representation and exploits th ...
Full textCite
Journal articleSIGMOD Record · June 8, 2023
Answering SPJA queries under differential privacy (DP), including graph pattern counting under node-DP as an important special case, has received considerable attention in recent years. The dual challenge of foreign-key constraints and self-joins is partic ...
Full textCite
Journal articleProceedings of the VLDB Endowment · January 1, 2023
In this work, we propose Longshot, a novel design for secure outsourced database systems that supports ad-hoc queries through the use of secure multi-party computation and differential privacy. By combining these two techniques, we build and maintain data ...
Full textCite
Journal articleProceedings of the VLDB Endowment · January 1, 2023
Synthetic data generation methods, and in particular, private synthetic data generation methods, are gaining popularity as a means to make copies of sensitive databases that can be shared widely for research and data analysis. Some of the fundamental opera ...
Full textCite
Journal articleJournal of Defense Modeling and Simulation · July 1, 2022
This paper describes the collaborative effort between privacy and security researchers at nine different institutions along with researchers at the Naval Information Warfare Center to deploy, test, and demonstrate privacy-preserving technologies in creatin ...
Full textCite
Journal articleProceedings of the VLDB Endowment · January 1, 2022
Most differentially private mechanisms are designed for the use of a single analyst. In reality, however, there are often multiple stake-holders with different and possibly conflicting priorities that must share the same privacy loss budget. This motivates ...
Full textCite
Journal articleProceedings of the ACM SIGMOD International Conference on Management of Data · June 14, 2020
Local sensitivity of a query Q given a database instance D, i.e. how much the output Q(D) changes when a tuple is added to D or deleted from D, has many applications including query analysis, outlier detection, and differential privacy. However, it is NP-h ...
Full textCite
Journal articleACM Transactions on Database Systems · February 1, 2020
The adoption of differential privacy is growing, but the complexity of designing private, efficient, and accurate algorithms is still high. We propose a novel programming framework and system, ϵktelo, for implementing both existing and new privacy algorith ...
Full textCite
Journal articleSIGMOD Record · November 5, 2019
The adoption of differential privacy is growing but the complexity of designing private, efficient and accurate algorithms is still high. We propose a novel programming framework and system, εktelo, for implementing both existing and new privacy algorithms ...
Full textCite
Journal articleJournal of Privacy and Confidentiality · October 23, 2019
The U.S. Census Bureau recently released data on earnings percentiles of grad-uates from post-secondary institutions. This paper describes and evaluates the disclosure avoidance system developed for these statistics. We propose a differentially private alg ...
Full textCite
Journal articleKnowledge and Information Systems · January 1, 2018
Linear and logistic regression are popular statistical techniques for analyzing multi-variate data. Typically, analysts do not simply posit a particular form of the regression model, estimate its parameters, and use the results for inference or prediction. ...
Full textCite
Journal articleCommunications of the ACM · March 1, 2015
Preparing data for public release requires significant attention to fundamental principles of privacy. If a privacy definition is chosen wisely by the data curator, the sensitive information will be protected. Algorithms that satisfy the spec are called pr ...
Full textCite
Journal articleJournal of the American Medical Informatics Association : JAMIA · March 2014
ObjectiveRecord linkage to integrate uncoordinated databases is critical in biomedical research using Big Data. Balancing privacy protection against the need for high quality record linkage requires a human-machine hybrid system to safely manage u ...
Full textCite
Journal articleACM Transactions on Database Systems · January 1, 2014
In this article, we introduce a new and general privacy framework called Pufferfish. The Pufferfish framework can be used to create new privacy definitions that are customized to the needs of a given application. The goal of Pufferfish is to allow experts ...
Full textCite
Journal articleComputer · January 1, 2014
Data-intensive research using distributed, federated, person-level datasets in near real time has the potential to transform social, behavioral, economic, and health sciences - but issues around privacy, confidentiality, access, and data integration have s ...
Full textCite
Journal articleProceedings of the ACM SIGMOD International Conference on Management of Data · January 1, 2014
Privacy definitions provide ways for trading-off the privacy of individuals in a statistical database for the utility of downstream analysis of the data. In this paper, we present Blowfish, a class of privacy definitions inspired by the Pufferfish framewor ...
Full textCite
Journal articleUbicomp 2014 Adjunct Proceedings of the 2014 ACM International Joint Conference on Pervasive and Ubiquitous Computing · January 1, 2014
The increasing popularity of wearable devices that continuously capture video, and the prevalence of third-party applications that utilize these feeds have resulted in a new threat to privacy. In many situations, sensitive objects/regions are maliciously ( ...
Full textCite
Journal articleProceedings International Conference on Data Engineering · August 15, 2013
Given a large graph G = (V, E) with millions of nodes and edges, how do we compute its connected components efficiently? Recent work addresses this problem in map-reduce, where a fundamental trade-off exists between the number of map-reduce rounds and the ...
Full textCite
Journal articleProceedings of the ACM SIGMOD Workshop on Databases and Social Networks Dbsocial 2013 · July 26, 2013
While social networking platforms allow users to control how their private information is shared, recent research has shown that a user's sensitive attribute can be inferred based on friendship links and group memberships, even when the attribute value is ...
Cite
Journal articleProceedings of the VLDB Endowment · January 1, 2013
We present SPARSI, a novel theoretical framework for partitioning sensitive data across multiple non-colluding adversaries. Most work in privacy-aware data sharing has considered disclosing summaries where the aggregate information about the data is preser ...
Full textCite
Journal articleACM International Conference Proceeding Series · December 19, 2012
De-duplication - identification of distinct records referring to the same real-world entity - is a well-known challenge in data integration. Since very large datasets prohibit the comparison of every pair of records, blocking has been identified as a techn ...
Full textCite
Journal articleInternational Conference on Information and Knowledge Management Proceedings · December 10, 2012
Internet users spend billions of minutes per month on so- cial networking sites like Facebook, LinkedIn and Twitter. Not only do they create tons of data everyday in the form of posts, tweets and photos, the connections between users have given rise to new ...
Full textCite
Journal articleProceedings of the ACM SIGACT SIGMOD SIGART Symposium on Principles of Database Systems · May 21, 2012
In this paper we introduce a new and general privacy framework called Pufferfish. The Pufferfish framework can be used to create new privacy definitions that are customized to the needs of a given application. The goal of Pufferfish is to allow experts in ...
Full textCite
Journal articleWww 12 Proceedings of the 21st Annual Conference on World Wide Web · May 16, 2012
Often an interesting true value such as a stock price, sports score, or current temperature is only available via the observations of noisy and potentially conflicting sources. Several techniques have been proposed to reconcile these conflicts by computing ...
Full textCite
Journal articleIEEE Transactions on Knowledge and Data Engineering · February 6, 2012
Search engine companies collect the database of intentions, the histories of their users' search queries. These search logs are a gold mine for researchers. Search engine companies, however, are wary of publishing search logs in order not to disclose sensi ...
Full textCite
Journal articleProceedings of the VLDB Endowment · January 1, 2012
In this paper, we analyze the nature and distribution of structured data on the Web. Web-scale information extraction, or the problem of creating structured tables using extraction from the entire web, is gathering lots of research interest. We perform a s ...
Full textCite
Journal articleProceedings of the VLDB Endowment · January 1, 2012
This tutorial brings together perspectives on ER from a variety of fields, including databases, machine learning, natural language processing and information retrieval, to provide, in one setting, a survey of a large body of work. We discuss both the pract ...
Full textCite
Journal articleProceedings of the 20th International Conference on World Wide Web Www 2011 · December 1, 2011
In this paper, we present a highly scalable algorithm for structurally clustering webpages for extraction. We show that, using only the URLs of the webpages and simple content features, it is possible to cluster webpages effectively and efficiently. At the ...
Full textCite
Journal articleACM Transactions on Internet Technology · March 1, 2011
In peer-to-peer (P2P) systems, computers from around the globe share data and can participate in distributed computation. P2P became famous, and infamous, due to file-sharing systems like Napster. However, the scalability and robustness of these systems ma ...
Full textCite
Journal articleProceedings of the VLDB Endowment · January 1, 2011
With the recent surge of social networks such as Facebook, new forms of recommendations have become possible - recommendationsthat rely on one's social connections in orderto make personalized recommendations of ads, content, products, and people. Since re ...
Full textCite
Journal articleProceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining · January 1, 2011
Given an unknown set of objects embedded in the Euclidean plane and a nearest-neighbor oracle, how to estimate the set size and other properties of the objects? In this paper we address this problem. We propose an efficient method that uses the Voronoi par ...
Full textCite
Journal articleWorkshop on Databases and Social Networks Dbsocial 11 · January 1, 2011
Internet users spend billions of minutes per month on sites like Facebook and Twitter. These sites support feed following, where users "follow" activity streams associated with other users and entities. Followers get personalized feeds that blend streams p ...
Full textCite
Journal articleProceedings of the ACM SIGMOD International Conference on Management of Data · January 1, 2011
Differential privacy is a powerful tool for providing privacy-preserving noisy query answers over statistical databases. It guarantees that the distribution of noisy query answers changes very little with the addition or deletion of any tuple. It is freque ...
Full textCite
Journal articleProceedings of the 4th ACM International Conference on Web Search and Data Mining Wsdm 2011 · January 1, 2011
Automatic extraction of structured records from inconsistently formatted lists on the web is challenging: different lists present disparate sets of attributes with variations in the ordering of attributes; many lists contain additional attributes and noise ...
Full textCite
Journal articleFoundations and Trends in Databases · December 31, 2009
Privacy is an important issue when one wants to make use of data that involves individuals' sensitive information. Research on protecting the privacy of individuals and the confidentiality of data has received contributions from many fields, including comp ...
Full textCite
Journal articleProceedings of the VLDB Endowment · January 1, 2009
Privacy in data publishing has received much attention recently. The key to defining privacy is to model knowledge of the attacker - if the attacker is assumed to know too little, the published data can be easily attacked, if the attacker is assumed to kno ...
Full textCite
Journal articleProceedings International Conference on Data Engineering · October 1, 2008
In this paper, we propose the first formal privacy analysis of a data anonymization process known as the synthetic data generation, a technique becoming popular in the statistics community. The target application for this work is a mapping program that sho ...
Full textCite
Journal articleProceedings of the VLDB Endowment · January 1, 2008
Publish/subscribe (pub/sub) systems are designed to efficiently match incoming events (e.g., stock quotes) against a set of subscriptions (e.g., trader profiles specifying quotes of interest). However, current pub/sub systems only support a simple binary n ...
Full textCite
Journal articleProceedings of the International Conference on Microelectronics, ICM · 2008
As the temperature became a first class design metric due to increased power densities, one of the key challenges in deep sub-micron technologies is to guarantee thermal safety while minimizing the performance impact. This highlights the need for thermal-a ...
Full textCite
Journal articleProceedings of the ACM SIGMOD International Conference on Management of Data · October 30, 2007
Peer-to-peer systems have emerged as a robust, scalable and decentralized way to share and publish data. In this paper, we propose P-Ring, a new P2P index structure that supports both equality and range queries. P-Ring is fault-tolerant, provides logarithm ...
Full textCite
Journal articleProceedings International Conference on Data Engineering · September 24, 2007
Recent work has shown the necessity of considering an attacker's background knowledge when reasoning about privacy in data publishing. However, in practice, the data publisher does not know what background knowledge the attacker possesses. Thus, it is impo ...
Full textCite
Journal articleProceedings International Conference on Data Engineering · October 17, 2006
Publishing data about individuals without revealing sensitive information about them is an important problem. In recent y ears, a new definition of privacy called k-anonymity has gained popularity. In a k-anonymized dataset, each record is indistinguishabl ...
Full textCite
Journal articleProceedings of the ACM SIGACT SIGMOD SIGART Symposium on Principles of Database Systems · June 26, 2006
Privacy-preserving query-answering systems answer queries while provably guaranteeing that sensitive information is kept secret. One very attractive notion of privacy is perfect privacy-a secret is expressed through a query QS, and a query Q
Full textCite
Journal articleThirteenth International World Wide Web Conference Proceedings Www2004 · December 1, 2004
We present a modularized storage and indexing framework that cleanly separates the functional components of a P2P system, enabling us to tailor the P2P infrastructure to the specific needs of various Internet applications. ...
Cite
Journal articleProceedings of the Annual ACM Symposium on Principles of Distributed Computing · July 21, 2002
We study the interplay of network connectivity and perfectly secure message transmission under the corrupting influence of generalized Byzantine adversaries. It is known that in the threshold adversary model, where the Byzantine adversary can corrupt upto ...
Full textCite
Journal articleInformatica Ljubljana · November 1, 2001
Commercial-off-the-shelf (COTS) components are black box software products. The absence of their code precludes them from any kind of inspection of certify that the code is safe. This increases the security risk for safety-sensitive applications. The appli ...
Cite