Preprint
Large Language Models

The Power of Generative AI: A Review of Requirements, Models, Input–Output Formats, Evaluation Metrics, and Challenges

Ajay Bandi(Northwest Missouri State University), Pydi Venkata Satya Ramesh Adapa(Northwest Missouri State University), Yudu Eswar Vinay Pratap Kumar Kuchi(Northwest Missouri State University)
July 31, 2023Future Internet587 citations

587

Citations

22

Influential Citations

Future Internet

Venue

2023

Year

Abstract

Generative artificial intelligence (AI) has emerged as a powerful technology with numerous applications in various domains. There is a need to identify the requirements and evaluation metrics for generative AI models designed for specific tasks. The purpose of the research aims to investigate the fundamental aspects of generative AI systems, including their requirements, models, input–output formats, and evaluation metrics. The study addresses key research questions and presents comprehensive insights to guide researchers, developers, and practitioners in the field. Firstly, the requirements necessary for implementing generative AI systems are examined and categorized into three distinct categories: hardware, software, and user experience. Furthermore, the study explores the different types of generative AI models described in the literature by presenting a taxonomy based on architectural characteristics, such as variational autoencoders (VAEs), generative adversarial networks (GANs), diffusion models, transformers, language models, normalizing flow models, and hybrid models. A comprehensive classification of input and output formats used in generative AI systems is also provided. Moreover, the research proposes a classification system based on output types and discusses commonly used evaluation metrics in generative AI. The findings contribute to advancements in the field, enabling researchers, developers, and practitioners to effectively implement and evaluate generative AI models for various applications. The significance of the research lies in understanding that generative AI system requirements are crucial for effective planning, design, and optimal performance. A taxonomy of models aids in selecting suitable options and driving advancements. Classifying input–output formats enables leveraging diverse formats for customized systems, while evaluation metrics establish standardized methods to assess model quality and performance.

Analysis

Why This Paper Matters

This survey paper addresses a critical gap in the rapidly evolving field of generative AI: the lack of a unified framework for understanding and comparing the diverse models, requirements, and evaluation methods. As generative AI applications proliferate across industries, practitioners need clear guidance on selecting appropriate architectures and metrics. The paper's structured taxonomy and categorization provide a much-needed roadmap for both newcomers and experienced researchers.

The paper's significance lies in its comprehensive scope, covering everything from hardware requirements to user experience considerations. By organizing the field into distinct categories, it enables more informed decision-making and helps standardize evaluation practices, which is essential for reproducible research and practical deployment.

Technical Contributions

  • Requirements Categorization: Divides generative AI system requirements into three clear categories: hardware (compute, memory), software (frameworks, libraries), and user experience (interface, latency).
  • Model Taxonomy: Presents a detailed taxonomy of generative AI models including variational autoencoders (VAEs), generative adversarial networks (GANs), diffusion models, transformers, language models, normalizing flow models, and hybrid models.
  • Input-Output Format Classification: Provides a comprehensive classification of input and output formats used in generative AI systems, enabling practitioners to match data types to appropriate models.
  • Evaluation Metrics Survey: Discusses commonly used evaluation metrics such as Inception Score (IS), Fréchet Inception Distance (FID), BLEU, ROUGE, and others, with guidance on their appropriate use cases.

Results

As a review paper, the primary results are the structured frameworks and taxonomies presented. The paper does not provide quantitative experimental results but offers qualitative insights into the strengths and weaknesses of different model architectures. The taxonomy and classification systems serve as actionable tools for researchers and developers to navigate the generative AI landscape.

Significance

This paper serves as a valuable reference for the AI community by consolidating disparate knowledge about generative AI into a coherent framework. It helps standardize terminology and evaluation practices, which is crucial for advancing the field. The practical guidance on requirements and model selection can accelerate development cycles and improve the quality of generative AI applications across domains such as text, image, audio, and video generation.