GOLDEN_SET.md
Gold-standard reference outputs for evaluation
14 documents availableCategories
Coding & Development
0 docsBusiness & Operations
0 docsMarketing & Content
0 docsData & Analytics
0 docsDesign & Creative
0 docsCustomer Support
0 docsSales & Revenue
0 docsHR & Recruiting
0 docsLegal & Compliance
0 docsFinance & Accounting
0 docsEducation & Training
0 docsResearch & Science
0 docsDevOps & Infrastructure
0 docsSecurity & Privacy
0 docsProduct Management
0 docsAutomation & Efficiency
0 docsDecision Making
0 docsContent Generation
0 docsQuality Assurance
0 docsTeam Collaboration
0 docsKnowledge Management
0 docsProcess Optimization
0 docsRisk Mitigation
0 docsRecent GOLDEN_SET.md Documents
View allGolden Dataset Guidelines
A golden dataset is a curated collection of examples with known-correct answers that you use to:
Project Memory
**Last Updated:** 2026-01-29 22:00
Understanding the Sources of Uncertainty - and Why Our Evals are Biased
Part 4 of *Iterating in the Dark:
Understanding the Sources of Uncertainty - and Why Our Evals are Biased
Part 4 of *Optimizing in the Dark:
Evaluation Scripts
The `create_test_set.py` script helps you interactively build a golden test dataset for evaluating the retrieval system.
IR-Copilot — Incident Response AI Assistant

📈 Trading RAG Mentor
> **Personal AI Trading Mentor** — A custom Retrieval-Augmented Generation (RAG) system built on momentum & price action video transcripts. Ask questions and get answers grounded exclusively in your own trading knowledge base.
AWS Certified Generative AI Developer – Professional (AIP-C01)
These are my personal study notes for the **AWS Certified Generative AI Developer – Professional (AIP-C01)** exam.