Preprint
Machine Learning

PubChem 2023 update

Sunghwan Kim(National Institutes of Health), Jie Chen(National Institutes of Health), Tiejun Cheng(National Institutes of Health), Asta Gindulytė(National Institutes of Health), Jia He(National Institutes of Health), Siqian He(National Institutes of Health), Qingliang Li(National Institutes of Health), Benjamin A. Shoemaker(National Institutes of Health), Paul Thiessen(National Institutes of Health), Bo Yu(National Institutes of Health), Leonid Zaslavsky(National Institutes of Health), Jian Zhang(National Institutes of Health), Evan Bolton(National Institutes of Health)
October 28, 2022Nucleic Acids Research3,249 citations

3.2k

Citations

128

Influential Citations

Nucleic Acids Research

Venue

2022

Year

Abstract

PubChem (https://pubchem.ncbi.nlm.nih.gov) is a popular chemical information resource that serves a wide range of use cases. In the past two years, a number of changes were made to PubChem. Data from more than 120 data sources was added to PubChem. Some major highlights include: the integration of Google Patents data into PubChem, which greatly expanded the coverage of the PubChem Patent data collection; the creation of the Cell Line and Taxonomy data collections, which provide quick and easy access to chemical information for a given cell line and taxon, respectively; and the update of the bioassay data model. In addition, new functionalities were added to the PubChem programmatic access protocols, PUG-REST and PUG-View, including support for target-centric data download for a given protein, gene, pathway, cell line, and taxon and the addition of the 'standardize' option to PUG-REST, which returns the standardized form of an input chemical structure. A significant update was also made to PubChemRDF. The present paper provides an overview of these changes.

Analysis

Why This Paper Matters

PubChem is one of the most widely used chemical information resources, serving researchers in drug discovery, toxicology, and chemical biology. This 2023 update is significant because it addresses key user needs: expanding patent coverage through Google Patents integration, enabling cell-line- and taxonomy-specific queries, and modernizing the bioassay data model. These changes make PubChem more relevant for AI practitioners who rely on large-scale chemical data for training models, such as in molecular property prediction or drug-target interaction. The enhanced programmatic access (PUG-REST, PUG-View) lowers barriers for automated data retrieval, which is essential for building reproducible AI pipelines.

Technical Contributions

  • Google Patents Integration: Adds millions of patent-chemical associations, enriching the patent data collection and enabling cross-referencing of chemical compounds with patent literature.
  • Cell Line and Taxonomy Collections: Provide pre-filtered views of chemical data for specific cell lines or taxa, simplifying data extraction for biological context.
  • Bioassay Data Model Update: Improves the structure and consistency of bioassay data, which is critical for training predictive models on assay outcomes.
  • PUG-REST Enhancements: New 'standardize' option returns canonical forms of input structures; target-centric download allows retrieval of data for a given protein, gene, pathway, cell line, or taxon.
  • PUG-View Updates: Supports similar target-centric data access for visualization.
  • PubChemRDF Update: Improves semantic web representation, enabling more sophisticated SPARQL queries and integration with linked data resources.

Results

The paper reports that data from over 120 new sources were added, but does not provide quantitative metrics (e.g., number of new compounds, patents, or assays). The Google Patents integration is highlighted as a major expansion, but no specific numbers are given. The bioassay data model update is described but not benchmarked. The new programmatic features are listed without performance comparisons. Overall, the results are qualitative, focusing on new capabilities rather than empirical evaluation.

Significance

For the AI community, PubChem's updates are important because they improve the quality and accessibility of chemical data used in machine learning. The ability to programmatically retrieve standardized structures and target-specific data reduces preprocessing overhead. The expanded patent coverage may enable new applications in patent analysis and prior art search. However, the lack of quantitative benchmarks means that practitioners must evaluate the new features themselves. The paper reinforces PubChem's role as a foundational resource for cheminformatics and AI-driven drug discovery.