Skip to main content

A Data Set and Deep Learning Algorithm for the Detection of Masses and Architectural Distortions in Digital Breast Tomosynthesis Images.

Publication ,  Journal Article
Buda, M; Saha, A; Walsh, R; Ghate, S; Li, N; Swiecicki, A; Lo, JY; Mazurowski, MA
Published in: JAMA Netw Open
August 2, 2021

IMPORTANCE: Breast cancer screening is among the most common radiological tasks, with more than 39 million examinations performed each year. While it has been among the most studied medical imaging applications of artificial intelligence, the development and evaluation of algorithms are hindered by the lack of well-annotated, large-scale publicly available data sets. OBJECTIVES: To curate, annotate, and make publicly available a large-scale data set of digital breast tomosynthesis (DBT) images to facilitate the development and evaluation of artificial intelligence algorithms for breast cancer screening; to develop a baseline deep learning model for breast cancer detection; and to test this model using the data set to serve as a baseline for future research. DESIGN, SETTING, AND PARTICIPANTS: In this diagnostic study, 16 802 DBT examinations with at least 1 reconstruction view available, performed between August 26, 2014, and January 29, 2018, were obtained from Duke Health System and analyzed. From the initial cohort, examinations were divided into 4 groups and split into training and test sets for the development and evaluation of a deep learning model. Images with foreign objects or spot compression views were excluded. Data analysis was conducted from January 2018 to October 2020. EXPOSURES: Screening DBT. MAIN OUTCOMES AND MEASURES: The detection algorithm was evaluated with breast-based free-response receiver operating characteristic curve and sensitivity at 2 false positives per volume. RESULTS: The curated data set contained 22 032 reconstructed DBT volumes that belonged to 5610 studies from 5060 patients with a mean (SD) age of 55 (11) years and 5059 (100.0%) women. This included 4 groups of studies: (1) 5129 (91.4%) normal studies; (2) 280 (5.0%) actionable studies, for which where additional imaging was needed but no biopsy was performed; (3) 112 (2.0%) benign biopsied studies; and (4) 89 studies (1.6%) with cancer. Our data set included masses and architectural distortions that were annotated by 2 experienced radiologists. Our deep learning model reached breast-based sensitivity of 65% (39 of 60; 95% CI, 56%-74%) at 2 false positives per DBT volume on a test set of 460 examinations from 418 patients. CONCLUSIONS AND RELEVANCE: The large, diverse, and curated data set presented in this study could facilitate the development and evaluation of artificial intelligence algorithms for breast cancer screening by providing data for training as well as a common set of cases for model validation. The performance of the model developed in this study showed that the task remains challenging; its performance could serve as a baseline for future model development.

Duke Scholars

Altmetric Attention Stats
Dimensions Citation Stats

Published In

JAMA Netw Open

DOI

EISSN

2574-3805

Publication Date

August 2, 2021

Volume

4

Issue

8

Start / End Page

e2119100

Location

United States

Related Subject Headings

  • Reproducibility of Results
  • ROC Curve
  • Middle Aged
  • Mammography
  • Humans
  • Female
  • False Positive Reactions
  • Early Detection of Cancer
  • Deep Learning
  • Datasets as Topic
 

Citation

APA
Chicago
ICMJE
MLA
NLM
Buda, M., Saha, A., Walsh, R., Ghate, S., Li, N., Swiecicki, A., … Mazurowski, M. A. (2021). A Data Set and Deep Learning Algorithm for the Detection of Masses and Architectural Distortions in Digital Breast Tomosynthesis Images. JAMA Netw Open, 4(8), e2119100. https://doi.org/10.1001/jamanetworkopen.2021.19100
Buda, Mateusz, Ashirbani Saha, Ruth Walsh, Sujata Ghate, Nianyi Li, Albert Swiecicki, Joseph Y. Lo, and Maciej A. Mazurowski. “A Data Set and Deep Learning Algorithm for the Detection of Masses and Architectural Distortions in Digital Breast Tomosynthesis Images.JAMA Netw Open 4, no. 8 (August 2, 2021): e2119100. https://doi.org/10.1001/jamanetworkopen.2021.19100.
Buda M, Saha A, Walsh R, Ghate S, Li N, Swiecicki A, et al. A Data Set and Deep Learning Algorithm for the Detection of Masses and Architectural Distortions in Digital Breast Tomosynthesis Images. JAMA Netw Open. 2021 Aug 2;4(8):e2119100.
Buda, Mateusz, et al. “A Data Set and Deep Learning Algorithm for the Detection of Masses and Architectural Distortions in Digital Breast Tomosynthesis Images.JAMA Netw Open, vol. 4, no. 8, Aug. 2021, p. e2119100. Pubmed, doi:10.1001/jamanetworkopen.2021.19100.
Buda M, Saha A, Walsh R, Ghate S, Li N, Swiecicki A, Lo JY, Mazurowski MA. A Data Set and Deep Learning Algorithm for the Detection of Masses and Architectural Distortions in Digital Breast Tomosynthesis Images. JAMA Netw Open. 2021 Aug 2;4(8):e2119100.

Published In

JAMA Netw Open

DOI

EISSN

2574-3805

Publication Date

August 2, 2021

Volume

4

Issue

8

Start / End Page

e2119100

Location

United States

Related Subject Headings

  • Reproducibility of Results
  • ROC Curve
  • Middle Aged
  • Mammography
  • Humans
  • Female
  • False Positive Reactions
  • Early Detection of Cancer
  • Deep Learning
  • Datasets as Topic