Recently Added
3,528 documents
Document Purpose and Scope
<!-- ISMS-CORE:CTX:ISMS-CTX-A.8.11-data-masking-technical-reference:framework:CTX:a.8.11 -->
Code Quality Report: Egnyte-LangChain Connector
The Egnyte-LangChain connector demonstrates **enterprise-grade code quality** with comprehensive testing, high coverage, and adherence to industry best practices. This report provides detailed evidence of code quality suitable for partnership evaluation and production deployment.
Tags
- .NET Worker Service
PDF Scanner Python Server
A FastAPI-based Python server for PDF processing, PII detection, redaction, and analytics with ClickHouse integration.
ThemisDB Admin Tools Overview
ThemisDB provides a comprehensive suite of administrative tools designed to help database administrators, security officers, and compliance teams manage their ThemisDB deployments effectively. This document provides an overview of all admin tools with focus on privacy, security, and compliance features.
🚀 Domain-First Autonomous Data Architecture
**Revolutionary Solution for GenAI Data Problems**
Skill Schema
name: defense-implementation
Glossar
Begriffe und Konzepte, die in den Experiment-Dokumenten verwendet werden.
Golden Dataset Guidelines
A golden dataset is a curated collection of examples with known-correct answers that you use to:
Evaluation of RAG Systems + Presentation Outline
How do we evaluate our RAG system?
PROJECT PLAN STARTER PACK - Complete Index
**Last Updated:** January 20, 2026
LLM Evaluation
LLM output evaluation — automated metrics, LLM-as-judge, A/B testing, regression testing. Use when measuring LLM output quality, comparing prompt or model versions, building an automated eval pipeline, setting up regression tests for prompt changes, or evaluating RAG systems and bias/safety.
SunCube AI - Comprehensive Documentation
- [Overview](#overview)
Cassandra - AI Review Agent
> *The truth about your code. Seeing bugs before the fall of production. Ignore at your own peril.*
Project Memory
**Last Updated:** 2026-01-29 22:00
GenAI Benchmarks & Evaluation — Product-Based Companies
Understanding how to **benchmark, evaluate, and compare LLMs** is essential for roles at Google, OpenAI, Anthropic, Cohere, and AI research teams. This file covers the most important benchmarks, evaluation methodologies, and how to build custom evaluation harnesses.
Evaluation Framework
This document describes how Agent Invest measures quality, detects regressions, and ensures safety. The system uses three evaluation layers: online scoring (every production run), offline evaluation (golden dataset), and guardrails (real-time safety checks).
Understanding the Sources of Uncertainty - and Why Our Evals are Biased
Part 4 of *Optimizing in the Dark:
Metrics
This document outlines the evaluation metrics available for assessing the performance of Retrieval Augmented Generation (RAG) systems, particularly focusing on the retrieval and generation components. The implementations can be found in `datapizza/evaluation/metrics.py`.
Project Memory
**Last Updated:** 2026-01-29 22:00
Evaluation Scripts
The `create_test_set.py` script helps you interactively build a golden test dataset for evaluating the retrieval system.
Data-Driven RAG Evaluation: Testing Qdrant Apps with Relari AI
url: "https://qdrant.tech/blog/qdrant-relari/"
LLM Evaluation
LLM output evaluation — automated metrics, LLM-as-judge, A/B testing, regression testing. Use when measuring LLM output quality, comparing prompt or model versions, building an automated eval pipeline, setting up regression tests for prompt changes, or evaluating RAG systems and bias/safety.
Agents and LLMs
* how to measure hallucinations?