ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
3.2k
Citations
128
Influential Citations
Nucleic Acids Research
Venue
2022
Year
PubChem (https://pubchem.ncbi.nlm.nih.gov) is a popular chemical information resource that serves a wide range of use cases. In the past two years, a number of changes were made to PubChem. Data from more than 120 data sources was added to PubChem. Some major highlights include: the integration of Google Patents data into PubChem, which greatly expanded the coverage of the PubChem Patent data collection; the creation of the Cell Line and Taxonomy data collections, which provide quick and easy access to chemical information for a given cell line and taxon, respectively; and the update of the bioassay data model. In addition, new functionalities were added to the PubChem programmatic access protocols, PUG-REST and PUG-View, including support for target-centric data download for a given protein, gene, pathway, cell line, and taxon and the addition of the 'standardize' option to PUG-REST, which returns the standardized form of an input chemical structure. A significant update was also made to PubChemRDF. The present paper provides an overview of these changes.
PubChem is one of the most widely used chemical information resources, serving researchers in drug discovery, toxicology, and chemical biology. This 2023 update is significant because it addresses key user needs: expanding patent coverage through Google Patents integration, enabling cell-line- and taxonomy-specific queries, and modernizing the bioassay data model. These changes make PubChem more relevant for AI practitioners who rely on large-scale chemical data for training models, such as in molecular property prediction or drug-target interaction. The enhanced programmatic access (PUG-REST, PUG-View) lowers barriers for automated data retrieval, which is essential for building reproducible AI pipelines.
The paper reports that data from over 120 new sources were added, but does not provide quantitative metrics (e.g., number of new compounds, patents, or assays). The Google Patents integration is highlighted as a major expansion, but no specific numbers are given. The bioassay data model update is described but not benchmarked. The new programmatic features are listed without performance comparisons. Overall, the results are qualitative, focusing on new capabilities rather than empirical evaluation.
For the AI community, PubChem's updates are important because they improve the quality and accessibility of chemical data used in machine learning. The ability to programmatically retrieve standardized structures and target-specific data reduces preprocessing overhead. The expanded patent coverage may enable new applications in patent analysis and prior art search. However, the lack of quantitative benchmarks means that practitioners must evaluate the new features themselves. The paper reinforces PubChem's role as a foundational resource for cheminformatics and AI-driven drug discovery.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba