Please login or create an account. If you do not have access to this content, you will be shown a 30 second preview and licensing options.

  • Presentation

Datasets, Clinical Studies, and External Validation in Dermatology AI

Description

In this talk, Albert Chu discusses the role of datasets and clinical studies in evaluating algorithms used in dermatology, particularly in the context of AI. He highlights the importance of understanding how algorithms perform under real-world conditions compared to their performance on curated retrospective datasets. While algorithms can show promising results in experimental settings, moving to practical applications reveals a different scenario. Chu emphasizes the need for prospective data sets that capture the variability and challenges encountered in clinical practice, as retrospective data may not represent the full spectrum of conditions encountered in patients. The talk presents case studies, illustrating how algorithms can experience significant performance degradation when deployed in real-world settings, particularly when the data used differs from training data. An example includes a melanoma classifier that performed well under controlled testing but faltered when faced with actual patient presentations. The session encourages careful evaluation of algorithm performance across diverse datasets to understand their clinical utility, urging practitioners to remain mindful of how an algorithm's original task definition affects its adaptability. Ultimately, Chu advocates for robust external datasets and local validations to ensure algorithms are clinically relevant and capable of supporting dermatological care effectively.

View more

Conclusions

  • Algorithms in dermatology can achieve performance comparable to board-certified dermatologists in controlled settings, but real-world applicability remains uncertain.
  • Prospective data sets are more useful than retrospective data for assessing algorithm performance in clinical practice.
  • Real-world performance evaluations are critical to understand the strengths and limitations of dermatology algorithms.
  • Performance degradation of algorithms is expected when transitioning from experimental to real-world settings, and must be carefully analyzed.
  • Understanding how and why algorithms fail is essential for determining their clinical utility and safety.
  • Task definition is crucial for the deployment of diagnostic algorithms; altering the task can lead to unpredictable performance changes.
  • The development of inclusive and diverse data sets is necessary to improve algorithm performance across various patient populations.
  • Local validation of algorithms may provide better insights into their usefulness than external validation, which might not reflect real-world situations.
  • Tschandl et al. JAMA Dermatol. 2021;157(11):1271-1273
  • JAMA Dermatol. 2020;156(1):29-37.
  • Sci Rep. 2022 Sep 28;12(1):16260.
  • J Invest Dermatol. 2022 Sep;142(9):2353-2362.e2.
  • Marchetti, Rotemberg et al. NPJ Digit Med. 2023 Jul 12;6(1):127.
  • Nature. 2017 Feb 2; 542(7639): 115-118.
  • Alexey Youssef, Michael Pencina, Anshul Thakur, Tingting Zhu, David Clifton & Nigam H. Shah. Nature Medicine 29, 2686-2687 (2023) | Cite this article.