Preprint
Machine Learning

IMPPAT: A curated database of Indian Medicinal Plants, Phytochemistry And Therapeutics

Karthikeyan Mohanraj(Homi Bhabha National Institute), Bagavathy Shanmugam Karthikeyan(Homi Bhabha National Institute), R.P. Vivek-Ananth(Homi Bhabha National Institute), Ramesh Chand(Homi Bhabha National Institute), S.R. Aparna, Pattulingam Mangalapandi(Homi Bhabha National Institute), Areejit Samal(Homi Bhabha National Institute)
March 6, 2018Scientific Reports722 citations

722

Citations

50

Influential Citations

Scientific Reports

Venue

2018

Year

Abstract

Phytochemicals of medicinal plants encompass a diverse chemical space for drug discovery. India is rich with a flora of indigenous medicinal plants that have been used for centuries in traditional Indian medicine to treat human maladies. A comprehensive online database on the phytochemistry of Indian medicinal plants will enable computational approaches towards natural product based drug discovery. In this direction, we present, IMPPAT, a manually curated database of 1742 Indian Medicinal Plants, 9596 Phytochemicals, And 1124 Therapeutic uses spanning 27074 plant-phytochemical associations and 11514 plant-therapeutic associations. Notably, the curation effort led to a non-redundant in silico library of 9596 phytochemicals with standard chemical identifiers and structure information. Using cheminformatic approaches, we have computed the physicochemical, ADMET (absorption, distribution, metabolism, excretion, toxicity) and drug-likeliness properties of the IMPPAT phytochemicals. We show that the stereochemical complexity and shape complexity of IMPPAT phytochemicals differ from libraries of commercial compounds or diversity-oriented synthesis compounds while being similar to other libraries of natural products. Within IMPPAT, we have filtered a subset of 960 potential druggable phytochemicals, of which majority have no significant similarity to existing FDA approved drugs, and thus, rendering them as good candidates for prospective drugs. IMPPAT database is openly accessible at: https://cb.imsc.res.in/imppat .

Analysis

Why This Paper Matters

This paper addresses a critical gap in computational drug discovery by providing a curated, open-access database of Indian medicinal plants and their phytochemicals. Traditional medicine systems, particularly in India, have a long history of use, but their chemical diversity has been underexplored in modern drug discovery pipelines. By systematically compiling 1742 plants, 9596 phytochemicals, and 1124 therapeutic uses, IMPPAT enables researchers to apply machine learning and cheminformatic approaches to identify novel drug candidates. The database's focus on Indian flora is particularly significant given the region's rich biodiversity and the growing interest in natural product-based therapeutics.

Technical Contributions

  • Manual Curation: The authors manually curated 1742 Indian medicinal plants, 9596 phytochemicals, and 1124 therapeutic uses, resulting in 27074 plant-phytochemical and 11514 plant-therapeutic associations. This level of curation ensures high-quality, reliable data for downstream analysis.
  • Cheminformatic Analysis: They computed physicochemical properties, ADMET profiles, and drug-likeliness for all phytochemicals, providing a standardized chemical library with identifiers and structure information.
  • Druggable Subset Identification: Using cheminformatic filters, they identified 960 potential druggable phytochemicals, most of which have no significant similarity to existing FDA-approved drugs, highlighting their novelty.
  • Comparative Analysis: They showed that IMPPAT phytochemicals have distinct stereochemical and shape complexity compared to commercial compound libraries and diversity-oriented synthesis compounds, while being similar to other natural product libraries.

Results

The database includes 1742 plants, 9596 phytochemicals, and 1124 therapeutic uses. The cheminformatic analysis revealed that the stereochemical complexity and shape complexity of IMPPAT phytochemicals differ from libraries of commercial compounds or diversity-oriented synthesis compounds. A subset of 960 potential druggable phytochemicals was identified, with the majority having no significant similarity to existing FDA-approved drugs, making them promising candidates for drug discovery.

Significance

IMPPAT provides a foundational resource for AI-driven drug discovery from natural products. By making the database openly accessible, it enables researchers to apply machine learning models for virtual screening, target prediction, and lead optimization. The identification of novel druggable phytochemicals expands the chemical space for drug development, particularly for diseases prevalent in India. This work bridges traditional medicine and modern computational approaches, potentially accelerating the discovery of new therapeutics from India's rich botanical heritage.