prompt logo

prompt

Free

Design LLM evaluation benchmarks like an architect

FreeFree tier
Inputs: textOutputs: text
Type
Open Source

About prompt

A system prompt designed to guide an AI to act as an Evaluation Architect for LLM systems. It includes detailed expertise areas such as benchmark design methodology, evaluation metrics, test strategies, quality gates, bias detection, scalability, cost-effectiveness, and failure mode analysis. The prompt provides a structured analysis process with steps for evaluation objective definition, benchmark design, metric design, evaluation rubric creation, failure mode analysis, and reporting.

Key Features

Benchmark design methodology (task selection, difficulty calibration, dataset construction)
Evaluation metrics and scoring rubrics (automated, manual, hybrid)
Test strategy for LLM systems (unit, integration, behavior, regression)
Quality gates and passing criteria definition
Bias detection and fairness evaluation
Scalability and reproducibility assessment
Cost-effectiveness analysis (compute budgets, batch vs. online evaluation)
Failure mode analysis and edge case discovery
Structured analysis process with 6 steps from objective definition to reporting

Pros & Cons

Pros
  • Comprehensive framework covering all aspects of LLM evaluation design
  • Structured process ensures thorough and reproducible analysis
  • Includes bias detection and fairness evaluation capabilities
  • Supports scalability and cost-effectiveness analysis
  • Practical focus with defined evaluation constraints and stakeholder requirements
Cons
  • Requires deep understanding of LLM evaluation concepts to use effectively
  • May be too granular for simple or small-scale projects
  • Designed specifically for LLM systems, not general AI evaluation

Best For

Designing evaluation benchmarks for LLM applications and systemsCreating quality frameworks and scoring rubrics for AI outputsConducting failure mode analysis and edge case discovery on LLM behaviorEstablishing objective success metrics and stakeholder requirements for LLM evaluationDefining bias detection and fairness evaluation strategies for LLMs