Preprint
Machine Learning

DrugBank: a comprehensive resource for in silico drug discovery and exploration

David S. Wishart(University of Alberta)
December 28, 2005Nucleic Acids Research4,044 citations

4.0k

Citations

338

Influential Citations

Nucleic Acids Research

Venue

2005

Year

Abstract

DrugBank is a unique bioinformatics/cheminformatics resource that combines detailed drug (i.e. chemical) data with comprehensive drug target (i.e. protein) information. The database contains >4100 drug entries including >800 FDA approved small molecule and biotech drugs as well as >3200 experimental drugs. Additionally, >14,000 protein or drug target sequences are linked to these drug entries. Each DrugCard entry contains >80 data fields with half of the information being devoted to drug/chemical data and the other half devoted to drug target or protein data. Many data fields are hyperlinked to other databases (KEGG, PubChem, ChEBI, PDB, Swiss-Prot and GenBank) and a variety of structure viewing applets. The database is fully searchable supporting extensive text, sequence, chemical structure and relational query searches. Potential applications of DrugBank include in silico drug target discovery, drug design, drug docking or screening, drug metabolism prediction, drug interaction prediction and general pharmaceutical education. DrugBank is available at http://redpoll.pharmacy.ualberta.ca/drugbank/.

Analysis

Why This Paper Matters

DrugBank addresses a critical gap in drug discovery by integrating chemical and biological data into a single, freely accessible database. Before DrugBank, researchers had to consult multiple disparate sources for drug properties and target information, slowing down in silico screening and target identification. By combining >4100 drug entries with >14,000 protein sequences, DrugBank enables rapid hypothesis generation and validation, accelerating the early stages of drug development.

The database's design, with over 80 data fields per entry and hyperlinks to major resources like KEGG, PubChem, and Swiss-Prot, makes it a central hub for pharmaceutical research. Its support for text, sequence, chemical structure, and relational queries allows flexible exploration, catering to both chemists and biologists. This integration has made DrugBank a standard reference in the field, evidenced by its high citation count (4044).

Technical Contributions

  • Comprehensive Data Integration: Merges drug chemical data (e.g., structure, properties) with target protein data (e.g., sequence, function) in a single record.
  • Rich Annotation: Each DrugCard contains >80 fields, covering drug nomenclature, pharmacology, pharmacokinetics, and target details.
  • External Linking: Hyperlinks to KEGG, PubChem, ChEBI, PDB, Swiss-Prot, and GenBank enable cross-referencing and data enrichment.
  • Advanced Search: Supports text, sequence (BLAST), chemical structure (substructure/similarity), and relational queries, making it versatile for various in silico tasks.
  • Broad Applicability: Designed for drug target discovery, docking, metabolism prediction, interaction prediction, and education.

Results

The paper reports that DrugBank contains >4100 drug entries, including >800 FDA-approved small molecule and biotech drugs, and >3200 experimental drugs. It links these to >14,000 protein or drug target sequences. The database is fully searchable and freely available online. While no specific performance metrics are provided, the database's utility is demonstrated by its widespread adoption and 4044 citations.

Significance

DrugBank has had a profound impact on computational drug discovery by providing a unified, high-quality resource that bridges chemistry and biology. It has enabled researchers to perform in silico target identification, drug repurposing, and interaction prediction without needing to manually integrate data from multiple sources. The database's design has influenced subsequent bioinformatics resources and remains a cornerstone for pharmaceutical AI and machine learning applications. Its continued use underscores the importance of curated, integrated databases in accelerating drug development.