All Documents
3,528 documents available
rawrubric
This rubric defines a **standardised metric** for evaluating how well a software repository implements core **kernel** and **operating‑system (OS)** primitives. It is based on the function manifest and status report from the Echo.Kern project and draws on general operating‑system principles ([Wikipedia: Kernel](https://en.wikipedia.org/wiki/Kernel_(operating_system)#:~:text=operating%20system%20%20that%20always,for%20the%20central%20processing%20unit)). The goal is to provide a repeatable method
The Evals Gap
It doesn't matter how beautiful your theory is, <br>
RAG Evaluation Guide
This guide explains how to evaluate the RAG (Retrieval-Augmented Generation) performance of the Clarity and Rigor agents using different retriever configurations.
Code Span Semantic Chunking Executive Summary
LLMC’s retrieval system must balance context relevance with token limitations, especially for large code
Module 6: Synthesis
[← Back: Cost Model](05_cost_model.md) | [Back to Project →](README.md)
Prompt Testing Skill
description: Comprehensive prompt testing and LLM output evaluation skill covering hallucination detection, response quality scoring, regression testing for prompts, A/B testing, and building evaluation pipelines for AI-powered applications.
EGG Rubric: Corporate Sustainability Evaluation Framework
The **EGG (Environmental, Governance & Goals) Rubric** is a comprehensive evaluation framework for assessing corporate sustainability performance across five critical sustainability themes. This rubric employs a multi-dimensional scoring approach that evaluates both the **quantity** and **quality** of corporate commitments, as well as their **specificity** and **temporal evolution**.
PRD-010 — Evaluation Framework
title: Evaluation Framework
Judging Rubric
**AI for Social Good Hackathon – SUST 2026**
RAG Evaluation
title: RAG Evaluation
Decodable Story Quality Rubric
Use this checklist when writing or reviewing decodable readers. The phonics constraints are hard enough—don't let the story suffer too.
Reddit Virality Grading Rubric
This document defines the scoring criteria for evaluating Reddit post/rumour virality potential. Each attribute is scored from **0.0 (no presence)** to **1.0 (very strong)**. These scores are used by the LLM to grade injected rumours before simulation.
! Project 3: Web APIs & NLP
In week four we've learned about a few different classifiers. In week five we learned about webscraping, APIs, and Natural Language Processing (NLP). This project will put those skills to the test.
Hook Rubric (score out of 10)
1. **Curiosity Gap (0–2)**
Criteria 1: Quality of Exploratory Data Analysis (20%) [20]
module_title: Data Science and Machine Learning
Sprint 0 Marking Scheme
**Team Name:** BC Hub
Project 2: Full-Stack Application
**Due:** Friday by 2:59am
RE-Bench Formal Scoring Rubric
RE-Bench evaluates reverse engineering LLMs across seven orthogonal axes:
Sprint 0 Marking Scheme
**Team Name:** [Byte-Peeps]
proj2rubric
Repository link: https://github.com/agupta15k/ncsu_se_fall22_22_pr_2
TA Grading Rubric
> **Core Principle:** Focus on whether the student has demonstrated mastery of the
Evaluation Rubric for LLM-Generated Outputs
title: Evaluation Rubric for LLM Outputs
Evaluation Metrics for Language Model Comparative Analysis
This document defines the metrics used to evaluate the performance of different language models in generating Python game scripts. The metrics focus on three key areas: Accuracy, Bug Frequency, and Feature Completeness.
Project 2 (P2) grading rubric
- A TA / grader will be reviewing your code after the deadline.