Recently Added

3,528 documents

PII.md

Document Purpose and Scope

<!-- ISMS-CORE:CTX:ISMS-CTX-A.8.11-data-masking-technical-reference:framework:CTX:a.8.11 -->

ai
0
0
isms-core-project
PII.md

Code Quality Report: Egnyte-LangChain Connector

The Egnyte-LangChain connector demonstrates **enterprise-grade code quality** with comprehensive testing, high coverage, and adherence to industry best practices. This report provides detailed evidence of code quality suitable for partnership evaluation and production deployment.

airageval
0
0
egnyte
PII.md

Tags

- .NET Worker Service

0
0
OneUptime
PII.md

PDF Scanner Python Server

A FastAPI-based Python server for PDF processing, PII detection, redaction, and analytics with ClickHouse integration.

airag
0
0
udaiveerS
COMPLIANCE.md

ThemisDB Admin Tools Overview

ThemisDB provides a comprehensive suite of administrative tools designed to help database administrators, security officers, and compliance teams manage their ThemisDB deployments effectively. This document provides an overview of all admin tools with focus on privacy, security, and compliance features.

airag
0
0
makr-code
ARCHITECTURE.md

🚀 Domain-First Autonomous Data Architecture

**Revolutionary Solution for GenAI Data Problems**

aiagent
0
0
data-riot
SKILL.md

Skill Schema

name: defense-implementation

aiagentllm
0
0
majiayu000
RAG.md

Glossar

Begriffe und Konzepte, die in den Experiment-Dokumenten verwendet werden.

aillmrag
0
0
hanasobi
GOLDEN_SET.md

Golden Dataset Guidelines

A golden dataset is a curated collection of examples with known-correct answers that you use to:

aieval
0
0
natnew
EVALS.md

Evaluation of RAG Systems + Presentation Outline

How do we evaluate our RAG system?

aillmrag
0
2
ayanahye
CLAUDE.md

PROJECT PLAN STARTER PACK - Complete Index

**Last Updated:** January 20, 2026

airag
0
0
owenlim225
SKILL.md

LLM Evaluation

LLM output evaluation — automated metrics, LLM-as-judge, A/B testing, regression testing. Use when measuring LLM output quality, comparing prompt or model versions, building an automated eval pipeline, setting up regression tests for prompt changes, or evaluating RAG systems and bias/safety.

aillmrag
0
1
projectious-work
DEPLOYMENT.md

SunCube AI - Comprehensive Documentation

- [Overview](#overview)

aiworkflowsafety
0
4
SunenaB3504
REVIEW_GUIDE.md

Cassandra - AI Review Agent

> *The truth about your code. Seeing bugs before the fall of production. Ignore at your own peril.*

aiagentllm
0
25
menny
GOLDEN_SET.md

Project Memory

**Last Updated:** 2026-01-29 22:00

aillmrag
0
0
shah-data-scientist
EVALS.md

GenAI Benchmarks & Evaluation — Product-Based Companies

Understanding how to **benchmark, evaluate, and compare LLMs** is essential for roles at Google, OpenAI, Anthropic, Cohere, and AI research teams. This file covers the most important benchmarks, evaluation methodologies, and how to build custom evaluation harnesses.

aiagentllm
0
2
CodeWithDhruvX
EVALS.md

Evaluation Framework

This document describes how Agent Invest measures quality, detects regressions, and ensures safety. The system uses three evaluation layers: online scoring (every production run), offline evaluation (golden dataset), and guardrails (real-time safety checks).

aiagentllm
0
0
yussaaa
GOLDEN_SET.md

Understanding the Sources of Uncertainty - and Why Our Evals are Biased

Part 4 of *Optimizing in the Dark:

aiagentrag
0
0
reliableai
RAG.md

Metrics

This document outlines the evaluation metrics available for assessing the performance of Retrieval Augmented Generation (RAG) systems, particularly focusing on the retrieval and generation components. The implementations can be found in `datapizza/evaluation/metrics.py`.

airageval
0
0
alessiogandelli
RAG.md

Project Memory

**Last Updated:** 2026-01-29 22:00

aillmrag
0
0
shah-data-scientist
GOLDEN_SET.md

Evaluation Scripts

The `create_test_set.py` script helps you interactively build a golden test dataset for evaluating the retrieval system.

aieval
0
0
bennettck
EVALS.md

Data-Driven RAG Evaluation: Testing Qdrant Apps with Relari AI

url: "https://qdrant.tech/blog/qdrant-relari/"

aillmrag
0
0
Kohnnn
SKILL.md

LLM Evaluation

LLM output evaluation — automated metrics, LLM-as-judge, A/B testing, regression testing. Use when measuring LLM output quality, comparing prompt or model versions, building an automated eval pipeline, setting up regression tests for prompt changes, or evaluating RAG systems and bias/safety.

aillmrag
0
4
projectious-work
FINE_TUNING.md

Agents and LLMs

* how to measure hallucinations?

aiagentllm
0
0
btcoal
Page 12 of 147