Quality gaps in public pancreas imaging datasets: Implications & challenges for AI applications

Garima Suman; Anurima Patra; Panagiotis Korfiatis; Shounak Majumder; Suresh T. Chari; Mark J. Truty; Joel G. Fletcher; Ajit H. Goenka

doi:10.1016/j.pan.2021.03.016

Quality gaps in public pancreas imaging datasets: Implications & challenges for AI applications

Garima Suman, Anurima Patra, Panagiotis Korfiatis, Shounak Majumder, Suresh T. Chari, Mark J. Truty, Joel G. Fletcher, Ajit H. Goenka

Research output: Contribution to journal › Article › peer-review

Abstract

Objective: Quality gaps in medical imaging datasets lead to profound errors in experiments. Our objective was to characterize such quality gaps in public pancreas imaging datasets (PPIDs), to evaluate their impact on previously published studies, and to provide post-hoc labels and segmentations as a value-add for these PPIDs. Methods: We scored the available PPIDs on the medical imaging data readiness (MIDaR) scale, and evaluated for associated metadata, image quality, acquisition phase, etiology of pancreas lesion, sources of confounders, and biases. Studies utilizing these PPIDs were evaluated for awareness of and any impact of quality gaps on their results. Volumetric pancreatic adenocarcinoma (PDA) segmentations were performed for non-annotated CTs by a junior radiologist (R1) and reviewed by a senior radiologist (R3). Results: We found three PPIDs with 560 CTs and six MRIs. NIH dataset of normal pancreas CTs (PCT) (n = 80 CTs) had optimal image quality and met MIDaR A criteria but parts of pancreas have been excluded in the provided segmentations. TCIA-PDA (n = 60 CTs; 6 MRIs) and MSD(n = 420 CTs) datasets categorized to MIDaR B due to incomplete annotations, limited metadata, and insufficient documentation. Substantial proportion of CTs from TCIA-PDA and MSD datasets were found unsuitable for AI due to biliary stents [TCIA-PDA:10 (17%); MSD:112 (27%)] or other factors (non-portal venous phase, suboptimal image quality, non-PDA etiology, or post-treatment status) [TCIA-PDA:5 (8.5%); MSD:156 (37.1%)]. These quality gaps were not accounted for in any of the 25 studies that have used these PPIDs (NIH-PCT:20; MSD:1; both: 4). PDA segmentations were done by R1 in 91 eligible CTs (TCIA-PDA:42; MSD:49). Of these, corrections were made by R3 in 16 CTs (18%) (TCIA-PDA:4; MSD:12) [mean (standard deviation) Dice: 0.72(0.21) and 0.63(0.23) respectively]. Conclusion: Substantial quality gaps, sources of bias, and high proportion of CTs unsuitable for AI characterize the available limited PPIDs. Published studies on these PPIDs do not account for these quality gaps. We complement these PPIDs through post-hoc labels and segmentations for public release on the TCIA portal. Collaborative efforts leading to large, well-curated PPIDs supported by adequate documentation are critically needed to translate the promise of AI to clinical practice.

Original language	English (US)
Pages (from-to)	1001-1008
Number of pages	8
Journal	Pancreatology
Volume	21
Issue number	5
DOIs	https://doi.org/10.1016/j.pan.2021.03.016
State	Published - Aug 2021

Keywords

Benchmarking
Bias
Deep learning
Metadata
Pancreatic carcinoma

ASJC Scopus subject areas

Endocrinology, Diabetes and Metabolism
Hepatology
Gastroenterology

Access to Document

10.1016/j.pan.2021.03.016

Cite this

@article{5461226d1ada41bfae713d7cffe7c513,

title = "Quality gaps in public pancreas imaging datasets: Implications & challenges for AI applications",

abstract = "Objective: Quality gaps in medical imaging datasets lead to profound errors in experiments. Our objective was to characterize such quality gaps in public pancreas imaging datasets (PPIDs), to evaluate their impact on previously published studies, and to provide post-hoc labels and segmentations as a value-add for these PPIDs. Methods: We scored the available PPIDs on the medical imaging data readiness (MIDaR) scale, and evaluated for associated metadata, image quality, acquisition phase, etiology of pancreas lesion, sources of confounders, and biases. Studies utilizing these PPIDs were evaluated for awareness of and any impact of quality gaps on their results. Volumetric pancreatic adenocarcinoma (PDA) segmentations were performed for non-annotated CTs by a junior radiologist (R1) and reviewed by a senior radiologist (R3). Results: We found three PPIDs with 560 CTs and six MRIs. NIH dataset of normal pancreas CTs (PCT) (n = 80 CTs) had optimal image quality and met MIDaR A criteria but parts of pancreas have been excluded in the provided segmentations. TCIA-PDA (n = 60 CTs; 6 MRIs) and MSD(n = 420 CTs) datasets categorized to MIDaR B due to incomplete annotations, limited metadata, and insufficient documentation. Substantial proportion of CTs from TCIA-PDA and MSD datasets were found unsuitable for AI due to biliary stents [TCIA-PDA:10 (17%); MSD:112 (27%)] or other factors (non-portal venous phase, suboptimal image quality, non-PDA etiology, or post-treatment status) [TCIA-PDA:5 (8.5%); MSD:156 (37.1%)]. These quality gaps were not accounted for in any of the 25 studies that have used these PPIDs (NIH-PCT:20; MSD:1; both: 4). PDA segmentations were done by R1 in 91 eligible CTs (TCIA-PDA:42; MSD:49). Of these, corrections were made by R3 in 16 CTs (18%) (TCIA-PDA:4; MSD:12) [mean (standard deviation) Dice: 0.72(0.21) and 0.63(0.23) respectively]. Conclusion: Substantial quality gaps, sources of bias, and high proportion of CTs unsuitable for AI characterize the available limited PPIDs. Published studies on these PPIDs do not account for these quality gaps. We complement these PPIDs through post-hoc labels and segmentations for public release on the TCIA portal. Collaborative efforts leading to large, well-curated PPIDs supported by adequate documentation are critically needed to translate the promise of AI to clinical practice.",

keywords = "Benchmarking, Bias, Deep learning, Metadata, Pancreatic carcinoma",

author = "Garima Suman and Anurima Patra and Panagiotis Korfiatis and Shounak Majumder and Chari, {Suresh T.} and Truty, {Mark J.} and Fletcher, {Joel G.} and Goenka, {Ajit H.}",

note = "Publisher Copyright: {\textcopyright} 2021 IAP and EPC",

year = "2021",

month = aug,

doi = "10.1016/j.pan.2021.03.016",

language = "English (US)",

volume = "21",

pages = "1001--1008",

journal = "Pancreatology",

issn = "1424-3903",

publisher = "S. Karger AG",

number = "5",

}

TY - JOUR

T1 - Quality gaps in public pancreas imaging datasets

T2 - Implications & challenges for AI applications

AU - Suman, Garima

AU - Patra, Anurima

AU - Korfiatis, Panagiotis

AU - Majumder, Shounak

AU - Chari, Suresh T.

AU - Truty, Mark J.

AU - Fletcher, Joel G.

AU - Goenka, Ajit H.

PY - 2021/8

Y1 - 2021/8

N2 - Objective: Quality gaps in medical imaging datasets lead to profound errors in experiments. Our objective was to characterize such quality gaps in public pancreas imaging datasets (PPIDs), to evaluate their impact on previously published studies, and to provide post-hoc labels and segmentations as a value-add for these PPIDs. Methods: We scored the available PPIDs on the medical imaging data readiness (MIDaR) scale, and evaluated for associated metadata, image quality, acquisition phase, etiology of pancreas lesion, sources of confounders, and biases. Studies utilizing these PPIDs were evaluated for awareness of and any impact of quality gaps on their results. Volumetric pancreatic adenocarcinoma (PDA) segmentations were performed for non-annotated CTs by a junior radiologist (R1) and reviewed by a senior radiologist (R3). Results: We found three PPIDs with 560 CTs and six MRIs. NIH dataset of normal pancreas CTs (PCT) (n = 80 CTs) had optimal image quality and met MIDaR A criteria but parts of pancreas have been excluded in the provided segmentations. TCIA-PDA (n = 60 CTs; 6 MRIs) and MSD(n = 420 CTs) datasets categorized to MIDaR B due to incomplete annotations, limited metadata, and insufficient documentation. Substantial proportion of CTs from TCIA-PDA and MSD datasets were found unsuitable for AI due to biliary stents [TCIA-PDA:10 (17%); MSD:112 (27%)] or other factors (non-portal venous phase, suboptimal image quality, non-PDA etiology, or post-treatment status) [TCIA-PDA:5 (8.5%); MSD:156 (37.1%)]. These quality gaps were not accounted for in any of the 25 studies that have used these PPIDs (NIH-PCT:20; MSD:1; both: 4). PDA segmentations were done by R1 in 91 eligible CTs (TCIA-PDA:42; MSD:49). Of these, corrections were made by R3 in 16 CTs (18%) (TCIA-PDA:4; MSD:12) [mean (standard deviation) Dice: 0.72(0.21) and 0.63(0.23) respectively]. Conclusion: Substantial quality gaps, sources of bias, and high proportion of CTs unsuitable for AI characterize the available limited PPIDs. Published studies on these PPIDs do not account for these quality gaps. We complement these PPIDs through post-hoc labels and segmentations for public release on the TCIA portal. Collaborative efforts leading to large, well-curated PPIDs supported by adequate documentation are critically needed to translate the promise of AI to clinical practice.

AB - Objective: Quality gaps in medical imaging datasets lead to profound errors in experiments. Our objective was to characterize such quality gaps in public pancreas imaging datasets (PPIDs), to evaluate their impact on previously published studies, and to provide post-hoc labels and segmentations as a value-add for these PPIDs. Methods: We scored the available PPIDs on the medical imaging data readiness (MIDaR) scale, and evaluated for associated metadata, image quality, acquisition phase, etiology of pancreas lesion, sources of confounders, and biases. Studies utilizing these PPIDs were evaluated for awareness of and any impact of quality gaps on their results. Volumetric pancreatic adenocarcinoma (PDA) segmentations were performed for non-annotated CTs by a junior radiologist (R1) and reviewed by a senior radiologist (R3). Results: We found three PPIDs with 560 CTs and six MRIs. NIH dataset of normal pancreas CTs (PCT) (n = 80 CTs) had optimal image quality and met MIDaR A criteria but parts of pancreas have been excluded in the provided segmentations. TCIA-PDA (n = 60 CTs; 6 MRIs) and MSD(n = 420 CTs) datasets categorized to MIDaR B due to incomplete annotations, limited metadata, and insufficient documentation. Substantial proportion of CTs from TCIA-PDA and MSD datasets were found unsuitable for AI due to biliary stents [TCIA-PDA:10 (17%); MSD:112 (27%)] or other factors (non-portal venous phase, suboptimal image quality, non-PDA etiology, or post-treatment status) [TCIA-PDA:5 (8.5%); MSD:156 (37.1%)]. These quality gaps were not accounted for in any of the 25 studies that have used these PPIDs (NIH-PCT:20; MSD:1; both: 4). PDA segmentations were done by R1 in 91 eligible CTs (TCIA-PDA:42; MSD:49). Of these, corrections were made by R3 in 16 CTs (18%) (TCIA-PDA:4; MSD:12) [mean (standard deviation) Dice: 0.72(0.21) and 0.63(0.23) respectively]. Conclusion: Substantial quality gaps, sources of bias, and high proportion of CTs unsuitable for AI characterize the available limited PPIDs. Published studies on these PPIDs do not account for these quality gaps. We complement these PPIDs through post-hoc labels and segmentations for public release on the TCIA portal. Collaborative efforts leading to large, well-curated PPIDs supported by adequate documentation are critically needed to translate the promise of AI to clinical practice.

KW - Benchmarking

KW - Bias

KW - Deep learning

KW - Metadata

KW - Pancreatic carcinoma

UR - http://www.scopus.com/inward/record.url?scp=85103939672&partnerID=8YFLogxK

UR - http://www.scopus.com/inward/citedby.url?scp=85103939672&partnerID=8YFLogxK

U2 - 10.1016/j.pan.2021.03.016

DO - 10.1016/j.pan.2021.03.016

M3 - Article

C2 - 33840636

AN - SCOPUS:85103939672

SN - 1424-3903

VL - 21

SP - 1001

EP - 1008

JO - Pancreatology

JF - Pancreatology

IS - 5

ER -

Quality gaps in public pancreas imaging datasets: Implications & challenges for AI applications

Abstract

Keywords

ASJC Scopus subject areas

Access to Document

Other files and links

Fingerprint

Cite this