Skip to main content

SD-NAE: GENERATING NATURAL ADVERSARIAL EXAMPLES WITH STABLE DIFFUSION

Publication ,  Conference
Lin, Y; Zhang, J; Chen, Y; Li, H
Published in: 2nd Tiny Papers Track at Iclr 2024 Tiny Papers @ Iclr 2024
January 1, 2024

Natural Adversarial Examples (NAEs), images arising naturally from the environment and capable of deceiving classifiers, are instrumental in robustly evaluating and identifying vulnerabilities in trained models. In this work, unlike prior works that passively collect NAEs from real images, we propose to actively synthesize NAEs using the state-of-the-art Stable Diffusion. Specifically, our method formulates a controlled optimization process, where we perturb the token embedding that corresponds to a specified class to generate NAEs. This generation process is guided by the gradient of loss from the target classifier, ensuring that the created image closely mimics the ground-truth class yet fools the classifier. Named SD-NAE (Stable Diffusion for Natural Adversarial Examples), our innovative method is effective in producing valid and useful NAEs, which is demonstrated through a meticulously designed experiment. Code is available at https://github.com/linyueqian/SD-NAE.

Duke Scholars

Published In

2nd Tiny Papers Track at Iclr 2024 Tiny Papers @ Iclr 2024

Publication Date

January 1, 2024
 

Citation

APA
Chicago
ICMJE
MLA
NLM
Lin, Y., Zhang, J., Chen, Y., & Li, H. (2024). SD-NAE: GENERATING NATURAL ADVERSARIAL EXAMPLES WITH STABLE DIFFUSION. In 2nd Tiny Papers Track at Iclr 2024 Tiny Papers @ Iclr 2024.
Lin, Y., J. Zhang, Y. Chen, and H. Li. “SD-NAE: GENERATING NATURAL ADVERSARIAL EXAMPLES WITH STABLE DIFFUSION.” In 2nd Tiny Papers Track at Iclr 2024 Tiny Papers @ Iclr 2024, 2024.
Lin Y, Zhang J, Chen Y, Li H. SD-NAE: GENERATING NATURAL ADVERSARIAL EXAMPLES WITH STABLE DIFFUSION. In: 2nd Tiny Papers Track at Iclr 2024 Tiny Papers @ Iclr 2024. 2024.
Lin, Y., et al. “SD-NAE: GENERATING NATURAL ADVERSARIAL EXAMPLES WITH STABLE DIFFUSION.” 2nd Tiny Papers Track at Iclr 2024 Tiny Papers @ Iclr 2024, 2024.
Lin Y, Zhang J, Chen Y, Li H. SD-NAE: GENERATING NATURAL ADVERSARIAL EXAMPLES WITH STABLE DIFFUSION. 2nd Tiny Papers Track at Iclr 2024 Tiny Papers @ Iclr 2024. 2024.

Published In

2nd Tiny Papers Track at Iclr 2024 Tiny Papers @ Iclr 2024

Publication Date

January 1, 2024