The Challenge of Data Integration

The Big Challenge of Data Integration

Drug discovery is based on analysis of complementary association between drugs and diseases existing in many data types. Integration of all data types in a single repository (data lake) is an important prerequisite for accuracy. The main challenges to integration, are the heterogeneity of a large number of data types and some data is stored in inaccessible silos. Iteru tackles data heterogeneity with a proprietary extraction solution, unlocking critical data stored in hidden silos.

Iteru Data Sources

The Iteru platform ingests eight critical data types: publications, FDA data, lab results, clinical trials, patient data, and drug banks, alongside the analytical results of two multimodal datasets, genomics and digital pathology. In comparison, many competitors use a small portion of only 2 or 3 data types and struggle to scale..

Iteru completed a prototype including:

  • Data ingestion of 8 critical data types.
  • Automated data integration and labeling, scalable up to 1 PB
  • Neural network determines the association of disease with biological entities
  • GUI and visualization software

In the future, Iteru will seek partnerships with biomedical companies and publishers. These organizations can provide data in exchange for Iteru’s services, such as data labeling.

Data and Cost Reduction Without Sacrificing Accuracy

AI for one petabyte of data is challenging even for an expensive computer with a performance of 32 petaflops, for example Nvidia DGX H100 server. A high-end AI server costs between $200K -$300K. A company developing AI for petabyte of data needs multiple servers. In the cloud, AI processing for 1 PB of data can reach $1.5M per year. To reduce computational cost, Iteru implemented a data reduction scheme without sacrificing accuracy.

Iteru's Data Reduction

Pronounced Reduction of computational cost

Out of 1 petabyte relevant data is extracted, causing data reduction by a factor of 1% – 2%, with an average of 1.5%. For instance, in PubMed out of its 37 million documents only 1.4% cover cancer (525,117). The reduction in storage, using the average of 1.5%, is from 1 petabyte to 15 terabytes. Reducing data to 15 terabytes significantly lowers computational cost and improves accuracy by removing statistical bias from irrelevant data. After computational reduction, there is no need for ultra-high-end infrastructure; standard enterprise servers priced at $20K–$40K now satisfy our requirements.