Recently Added

3,528 documents

EVALS.md

GenAI Benchmarks & Evaluation — Product-Based Companies

Understanding how to **benchmark, evaluate, and compare LLMs** is essential for roles at Google, OpenAI, Anthropic, Cohere, and AI research teams. This file covers the most important benchmarks, evaluation methodologies, and how to build custom evaluation harnesses.

aiagentllm
0
2
CodeWithDhruvX
RAG.md

RAG Deep Dive Part 7: Evaluation and Debugging RAG Systems

**Series:** RAG (Retrieval-Augmented Generation) A Developer's Deep Dive from Scratch to Production

aillmrag
0
2
Sachinchaurasiya360
EVALS.md

Data-Driven RAG Evaluation: Testing Qdrant Apps with Relari AI

url: "https://qdrant.tech/blog/qdrant-relari/"

aillmrag
0
0
Kohnnn
EVALS.md

After LangGraph node execution, convert messages

**RAGAS** (Retrieval-Augmented Generation Assessment) is a specialized evaluation framework designed to measure RAG pipeline performance through reference-free metrics, making it ideal for production systems. **LangGraph** is a state-based orchestration framework that structures AI workflows as directed graphs. Integrating these two creates a powerful system for building and evaluating complex RAG pipelines systematically.

aiagentllm
0
0
lowkaihon
GOLDEN_SET.md

Evaluation Scripts

The `create_test_set.py` script helps you interactively build a golden test dataset for evaluating the retrieval system.

aieval
0
0
bennettck
GOLDEN_SET.md

IR-Copilot — Incident Response AI Assistant

![Static Badge](https://img.shields.io/badge/automated%20tests-135-blue)

aiagentllm
0
0
giladresisi
EVALS.md

RAG System Testing Methodologies: A Comprehensive Guide

**Document Version:** 1.0

aillmrag
0
2
destefani
EVALS.md

Day 20: Evaluation & Benchmarks 📏

root((Day 20: Evaluation & Benchmarks 📏))

aillmrag
0
3
Ravikiran-Bhonagiri
GOLDEN_SET.md

AWS Certified Generative AI Developer – Professional (AIP-C01)

These are my personal study notes for the **AWS Certified Generative AI Developer – Professional (AIP-C01)** exam.

ai
0
2
vicsz
GOLDEN_SET.md

📈 Trading RAG Mentor

> **Personal AI Trading Mentor** — A custom Retrieval-Augmented Generation (RAG) system built on momentum & price action video transcripts. Ask questions and get answers grounded exclusively in your own trading knowledge base.

aillmrag
0
3
sudhakarbadugu
EVALS.md

Using Performance Metrics to Evaluate RAG Systems

title: "Data-Driven RAG Evaluation: Testing Qdrant Apps with Relari AI"

aillmrag
0
0
AlexisBalayre
GOLDEN_SET.md

Understanding the Sources of Uncertainty - and Why Our Evals are Biased

Part 4 of *Iterating in the Dark:

aiagentrag
0
0
reliableai
EVALS.md

Domain 5: Testing, Validation, and Troubleshooting

**AIP-C01 Study Guide — Dr. Priya Ramanathan**

aiagentllm
0
1
rahulbhavani-il
GOLDEN_SET.md

AI Tester Interview Preparation Guide

1. [Core Competencies Overview](#core-competencies)

aillmprompt
0
2
k21academyuk
EVALS.md

[BEE-30004] Evaluating and Testing LLM Applications

title: Evaluating and Testing LLM Applications

aillmrag
0
1
alivedise
SKILL.md

LLM Evaluation

LLM output evaluation — automated metrics, LLM-as-judge, A/B testing, regression testing. Use when measuring LLM output quality, comparing prompt or model versions, building an automated eval pipeline, setting up regression tests for prompt changes, or evaluating RAG systems and bias/safety.

aillmrag
0
1
projectious-work
EVALS.md

Instructions for Claude Code: n8n Meal Feedback LLM Evaluation Workflow

Create a plan to build an n8n workflow that evaluates multiple LLM prompts for generating meal feedback using a **thinking model to generate ground truth** for comparison.

aillmrag
0
1
B-vR
GLOSSARY.md

Glossary

*[Deutsche Version](GLOSSARY_DE.md)*

aillmrag
0
0
hanasobi
GOLDEN_SET.md

Thesis Falsifier

A tool to aid researchers in assessing whether research papers adhere to scientific best practices. This application uses AI to automatically generate falsification forms, helping researchers verify the scientific robustness of their work across disciplines including social sciences and natural sciences.

aillmrag
0
0
fobert789
RUBRIC.md

QualRubric

The goal of a Qualifying Exam ("qual") is for a student to *effectively demonstrate that they have the knowledge and skills that will be needed to conduct meaningful research in their chosen subfield*. There are a number of key phrases in this sentence:

aieval
0
0
adsarwate
EVALS.md

Agent Evaluation Reference Guide

Complete documentation for the `agent-eval` CLI, metrics, data formats, and customization.

aiagenteval
0
8
danielazamorah
GOLDEN_SET.md

[[Retrieval Augmented Generation|RAG]]: Trade-offs & Evaluation Strategy

* **Rapid Time to Market:** Easier to implement than fine-tuning a model from scratch.

aillmrag
0
0
TanKaizokuO
RUBRIC.md

! Project 3: Web APIs & NLP

In week four we've learned about a few different classifiers. In week five we learned about webscraping, APIs, and Natural Language Processing (NLP). This project will put those skills to the test.

ai
0
0
dmartorano
RUBRIC.md

Sprint 0 Marking Scheme

**Team Name:** BC Hub

ai
0
0
UTSCCSCC01
Page 13 of 147