AI Drug Discovery and Therapeutics

AI Requirements and Challenges

Artificial intelligence (AI) has the potential to revolutionize pharma R&D and pharma digitization. It is adept at analyzing immense complex interactions between a large number of variables in vast datasets to drive actionable insights. The effectiveness of an AI solution depends on the nature of the problem and the required accuracy. Very high accuracy remains the biggest challenge to many aspects of pharma R&D in particular drug discovery.

Quality of Biomedical Data

The quality of the biomedical data is the primary barrier to accuracy for the following reasons:

  • Data is quintessentially messy and needs to be cleaned.
  • Missing values, duplicate records, and inconsistent formats.
  • Integration barriers: Caused by high heterogeneity across all data types. Complementary information about disease and biology exist in different data types. Integration is necessary to improve AI accuracy and build effective machine learning models.
  • Dark Data: More than 50% of unstructured data is trapped in inaccessible silos. (Gartner).
  • Lack of accurate labeling: Undermines the backbone of biomedical data management, directly impacting patient outcomes, research integrity, and legal approval. Pharma companies heavily rely on manual data labeling, a process that is inconsistent, time-consuming, unscalable, and error-prone. To make matters worse, existing labels lose validity as soon as the data updates.

 

The Challenge of Biomedical Data

Used AI Algothms are Inadequate

Widely used AI algorithms for cleansing and integration rely on statistical probability, leading to suboptimal solutions and low AI accuracy. Furthermore, these solutions cannot scale to handle massive biomedical datasets. When using manual solutions only 0.01% of available data is integrated.

AI Success is Constrained

Open-source AI excels at specific R&D tasks like protein-shape prediction, but its performance on heterogeneous biomedical data is heavily constrained by notoriously poor data quality. Automating the cleansing, integrating, labeling, and inclusion of dark data ultimately yields more effective AI models such as regression, generative, discriminative, and foundation models. Effective AI models directly improve predictive accuracy by capturing complex, high-dimensional data patterns and uncovering hidden relationships.

Iteru’s Solution

Iteru built a platform to address biomedical data problems. The platform is scalable to 1PB of data and its functionality is expandable. Its extensible design enables it to accommodates evolving data types and AI algorithms to address broad pharmaceutical R&D and digitization challenges

Iteru provides:

  • Accurate solutions to data cleansing and integration, based on proprietary AI algorithms. It automatically integrates 8 critical data types before AI analysis (clinical trials, digital pathology, genomics, etc.)
  • Provides solutions to missing values, duplicate records, and inconsistent formats.
  • Integrate biomedical data, non-multimodal and multimodal, in one place (a data lake).
  • Tears down data silos to reveal hidden knowledge.
  • Automated accurate data labelling.
  • The platform scales to petabyte to accommodate all critical biomedical data types.

Integration Reduces R&D Cost

According to McKinsey & Company, integrating biomedical data cuts pharmaceutical R&D costs by up to 30% and shrinks time-to-market by 25%

Maximizing Efficiency and Reducing Costs in Generative AI

Iteru Reduces AI Computational Cost

The computational cost of AI escalates rapidly when processing large datasets. Consequently, isolating relevant information before feeding it to AI algorithms is essential. Iteru’s classification and labeling algorithms are used to isolate data related to the objective of analysis. If the objective of analysis is a specific disease, only data related to that disease is processed. This greatly reduces computational cost and removes statistical bias.

Data Cleansing and Labeling Reduces Generative AI Costs by up to 500x

  • Data cleansing significantly reduces the overall cost of generative AI by shrinking token usage, preventing redundant fine-tuning, and stopping expensive model hallucinations.
  • Automated data labeling reduces generative AI costs by shifting human work from generation to review, cutting annotation expenses by up to 70% to 500x depending on the task complexity.
  • Removing duplicate data lowers generative AI costs by shrinking storage requirements, accelerating model training times, and reducing token consumption during retrieval-augmented generation.

Getting Pharma Scientists Involved

It is a textbook notion that to attain very high AI accuracy, understanding of the data is a MUST. Iteru automates data mining and AI analysis to provide an out of the box platform usable by bio scientists, who understand the data, and have no experience in AI or data science. There is lots of entanglement and ambiguity in biomedical data. For instance, cGMP is involved in heart failure, regulating blood pressure, cardiovascular health and prevention and treatment of breast cancer. A bio scantiest, because of his/her understanding of the data is best suited to refine the objective of analysis to be used by AI to provide desired results. In the objective of analysis, he/she can specify whether the analysis pertains to heart failure, blood pressure, cardiovascular health or cancer. He/she can use the platform to interrogate the data to gain more understanding and add more refinement to the objective of analysis by including oncogenes, pathways, receptors, etc. Refinements remove statistical bias and increase accuracy. Software engineer and data scientists use the initial objective of analysis provided by a bio scientist, but they cannot interrogate the data or effectively refine it.