Skip to main content

A sampling-based approach to information recovery

Publication ,  Journal Article
Xie, J; Yang, J; Chen, Y; Wang, H; Yu, PS
Published in: Proceedings - International Conference on Data Engineering
October 1, 2008

There has been a recent resurgence of interest in research on noisy and incomplete data. Many applications require information to be recovered from such data. Ideally, an approach for information recovery should have the following features. First, it should be able to incorporate prior knowledge about the data, even if such knowledge is in the form of complex distributions and constraints for which no close-form solutions exist. Second, it should be able to capture complex correlations and quantify the degree of uncertainty in the recovered data, and further support queries over such data. The database community has developed a number of approaches for information recovery, but none is general enough to offer all above features. To overcome the limitations, we take a significantly more general approach to information recovery based on sampling. We apply sequential importance sampling, a technique from statistics that works for complex distributions and dramatically outperforms naive sampling when data is constrained. We illustrate the generality and efficiency of this approach in two application scenarios: cleansing RFID data, and recovering information from published data that has been summarized and randomized for privacy. © 2008 IEEE.

Duke Scholars

Published In

Proceedings - International Conference on Data Engineering

DOI

ISSN

1084-4627

Publication Date

October 1, 2008

Start / End Page

476 / 485
 

Citation

APA
Chicago
ICMJE
MLA
NLM
Xie, J., Yang, J., Chen, Y., Wang, H., & Yu, P. S. (2008). A sampling-based approach to information recovery. Proceedings - International Conference on Data Engineering, 476–485. https://doi.org/10.1109/ICDE.2008.4497456
Xie, J., J. Yang, Y. Chen, H. Wang, and P. S. Yu. “A sampling-based approach to information recovery.” Proceedings - International Conference on Data Engineering, October 1, 2008, 476–85. https://doi.org/10.1109/ICDE.2008.4497456.
Xie J, Yang J, Chen Y, Wang H, Yu PS. A sampling-based approach to information recovery. Proceedings - International Conference on Data Engineering. 2008 Oct 1;476–85.
Xie, J., et al. “A sampling-based approach to information recovery.” Proceedings - International Conference on Data Engineering, Oct. 2008, pp. 476–85. Scopus, doi:10.1109/ICDE.2008.4497456.
Xie J, Yang J, Chen Y, Wang H, Yu PS. A sampling-based approach to information recovery. Proceedings - International Conference on Data Engineering. 2008 Oct 1;476–485.

Published In

Proceedings - International Conference on Data Engineering

DOI

ISSN

1084-4627

Publication Date

October 1, 2008

Start / End Page

476 / 485