Big Data and Artificial Intelligence in Cancer Drug Discovery: Promise, Challenges, and Emerging Opportunities

Oncologic drug development is lengthy (~14 years) and expensive (~1.2 billion USD) with low clinical trial success rates (4.1%). Big data and artificial intelligence (AI) are widely proposed as tools to address these challenges. In this review, we examine the current performance and future potential of big data and AI applied to preclinical discovery and development, clinical trials, and the regulatory approval process. We first examine the data foundation required for effective AI, including data harmonization, data commons, and analytical tools. We then assess preclinical applications spanning target identification, compound-library curation, virtual ligand screening, generative chemical design, and high-throughput and high-content screening. In clinical development, we consider the use of big data and AI for outcome prediction, trial design, external and synthetic control arms, adaptive monitoring, and in silico trials. Finally, we discuss how post-approval electronic health records can generate real-world data and real-world evidence to support drug repurposing and improve future oncology drug discovery. Big data is conventionally characterized by a series of “Vs.” In this review, we have used seven “Vs” spanning descriptive and constraining properties of big data and a singular outcome. We have proposed an eighth, Vernacular, a constraint defined as the combined alignment of data semantics and terminology, data representation, data exchange, and data governance across heterogeneous, independently generated datasets to promote interoperability and combined analysis. Although cancer data exhibit substantial Volume, Velocity, and Variety, they remain distributed across fragmented repositories that often cannot be readily integrated. We conclude with a discussion of tabulated resources currently available for the application of big data and AI to oncologic therapeutics.

Authors

Institutions

Publication Details

Journal
Cancers
Published
2026-09-10
DOI
https://doi.org/10.3390/cancers18182939
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Big Data and Artificial Intelligence in Cancer Drug Discovery: Promise, Challenges, and Emerging Opportunities

Charles E. McKenna, Terrence M. Lee, Jonathan G. Katz, Fakhar U. Singhera et al.
Cancers
Artificial Intelligence in Healthcare and Education
article

Big Data and Artificial Intelligence in Cancer Drug Discovery: Promise, Challenges, and Emerging Opportunities

Charles E. McKenna, Terrence M. Lee, Jonathan G. Katz, Fakhar U. Singhera, Justin M. Overhulse, Jerry Lee
article en

Abstract

Oncologic drug development is lengthy (~14 years) and expensive (~1.2 billion USD) with low clinical trial success rates (4.1%). Big data and artificial intelligence (AI) are widely proposed as tools to address these challenges. In this review, we examine the current performance and future potential of big data and AI applied to preclinical discovery and development, clinical trials, and the regulatory approval process. We first examine the data foundation required for effective AI, including data harmonization, data commons, and analytical tools. We then assess preclinical applications spanning target identification, compound-library curation, virtual ligand screening, generative chemical design, and high-throughput and high-content screening. In clinical development, we consider the use of big data and AI for outcome prediction, trial design, external and synthetic control arms, adaptive monitoring, and in silico trials. Finally, we discuss how post-approval electronic health records can generate real-world data and real-world evidence to support drug repurposing and improve future oncology drug discovery. Big data is conventionally characterized by a series of “Vs.” In this review, we have used seven “Vs” spanning descriptive and constraining properties of big data and a singular outcome. We have proposed an eighth, Vernacular, a constraint defined as the combined alignment of data semantics and terminology, data representation, data exchange, and data governance across heterogeneous, independently generated datasets to promote interoperability and combined analysis. Although cancer data exhibit substantial Volume, Velocity, and Variety, they remain distributed across fragmented repositories that often cannot be readily integrated. We conclude with a discussion of tabulated resources currently available for the application of big data and AI to oncologic therapeutics.

CancersVol. 18(18)
University of Southern California (US), Larry Ellison Foundation (US)
Openalex Percentile: Top 14%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.