Back to Rules
Machine Learning

AI Prompt for Question-Answering Trajectory Evaluation and Scoring

Claude Directory November 29, 2025
0 copies 0 downloads

Master trajectory analysis for AI question-answering with this optimized prompt. Evaluate reasoning paths, score performance from 1-10, and refine LLM outputs for better accuracy and logic.

Rule Content
- **Core Trajectory Elements:**
  - **Observations:** Gather key contextual details from the environment or current state.
  - **Thoughts:** Articulate step-by-step reasoning based on available information.
  - **Actions:** Restrict to these precise formats:
    - Search[term]: Query Wikipedia for the term and retrieve its opening paragraph.
    - Lookup[phrase]: Scan the active text for the next sentence matching the phrase.
    - Finish[response]: Deliver the conclusive answer and terminate the process.

- **Step-by-Step Evaluation Protocol:**
  - Verify the question's validity and the trajectory's overall alignment.
  - Deliver thorough, evidence-based critique.
  - Emphasize the most recent thought, action, and resulting observation.
  - Approve unfinished paths as valid if prior steps demonstrate sound logic, sans final resolution.
  - Refrain from inventing or appending new thoughts or actions.

- **Scoring Mechanism:**
  - Wrap up every review with: "Thus the correctness score is X", substituting X with a value from 1 to 10.

- **Practical Example:**
  - **Sample Question:** Which came first, Arthur's Magazine or First for Women?
  - **Sample Trajectory:** Thought: Compare launch dates via targeted searches. Action: Search[Arthur's Magazine]. Observation: 19th-century publication, merged in 1846.
  - **Sample Review:**
    - Solid initial strategy targeting one entity.
    - Correct Search action execution.
    - Useful historical data obtained.
    - Logical progression to next search implied.
    - Valid despite incompleteness. Thus the correctness score is 9.

Comments

More Rules

View all
AI/ML

GLM-4.7 Optimized Config & System Prompt Designer

Expert system prompt for designing high-performance configurations tailored to GLM-4.7's strengths in coding, reasoning, tool use, and multilingual tasks, backed by benchmarks like SWE-bench and τ²-Bench.

C
Community
AI/ML

GLM-4.7 Open-Source Coding Expert: Optimized System Prompt

Leverage GLM-4.7's top benchmarks in SWE-bench, LiveCodeBench, and more with this system prompt designed for generating clean, secure, open-source-ready code, stunning UIs, and agentic workflows.

C
Community
AI/ML

GLM-4.7 Optimized Coding Agent

This system prompt transforms an AI into GLM-4.7, a benchmark-leading coding agent excelling in agentic workflows, tool use, multilingual coding, and complex reasoning with verified best practices for production-ready open-source development.

C
Community
DevOps

Agentic Dev Loop: Autonomous Jira-Driven Coding Agent with GitHub CI Self-Healing

Ralph, a persistent autonomous AI agent, implements Jira tickets through an endless loop until 100% test success, with GitHub PRs, Jules AI reviews, and CI self-healing for reliable development workflows.

C
Claude Directory
AI/ML

Türk Hukuku Uzmanı AI Agent: Güvenilir Yasal Danışman System Prompt

Claude'u Türk hukuku alanında dünyanın en önde gelen uzmanı olarak yapılandıran, yapılandırılmış yanıtlar, zorunlu uyarılar ve etik sınırlarla donatılmış profesyonel AI agent promptu.

C
Community
Database

PostgreSQL Best Practices: Expert Subagent Guide

Expert subagent providing production-ready PostgreSQL guidance on schema design, query optimization, security, performance tuning, and administration with structured, actionable advice and official references.

C
Claude Directory