Recently Added
3,528 documents
Thesis Falsifier
A tool to aid researchers in assessing whether research papers adhere to scientific best practices. This application uses AI to automatically generate falsification forms, helping researchers verify the scientific robustness of their work across disciplines including social sciences and natural sciences.
RAG Evaluation Guide
This guide explains how to evaluate the RAG (Retrieval-Augmented Generation) performance of the Clarity and Rigor agents using different retriever configurations.
Code Span Semantic Chunking Executive Summary
LLMC’s retrieval system must balance context relevance with token limitations, especially for large code
The Evals Gap
It doesn't matter how beautiful your theory is, <br>
rawrubric
This rubric defines a **standardised metric** for evaluating how well a software repository implements core **kernel** and **operating‑system (OS)** primitives. It is based on the function manifest and status report from the Echo.Kern project and draws on general operating‑system principles ([Wikipedia: Kernel](https://en.wikipedia.org/wiki/Kernel_(operating_system)#:~:text=operating%20system%20%20that%20always,for%20the%20central%20processing%20unit)). The goal is to provide a repeatable method
Module 6: Synthesis
[← Back: Cost Model](05_cost_model.md) | [Back to Project →](README.md)
Judging Rubric
**AI for Social Good Hackathon – SUST 2026**
Prompt Testing Skill
description: Comprehensive prompt testing and LLM output evaluation skill covering hallucination detection, response quality scoring, regression testing for prompts, A/B testing, and building evaluation pipelines for AI-powered applications.
RAG Evaluation
title: RAG Evaluation
EGG Rubric: Corporate Sustainability Evaluation Framework
The **EGG (Environmental, Governance & Goals) Rubric** is a comprehensive evaluation framework for assessing corporate sustainability performance across five critical sustainability themes. This rubric employs a multi-dimensional scoring approach that evaluates both the **quantity** and **quality** of corporate commitments, as well as their **specificity** and **temporal evolution**.
PRD-010 — Evaluation Framework
title: Evaluation Framework
! Project 3: Web APIs & NLP
In week four we've learned about a few different classifiers. In week five we learned about webscraping, APIs, and Natural Language Processing (NLP). This project will put those skills to the test.
Hook Rubric (score out of 10)
1. **Curiosity Gap (0–2)**
Reddit Virality Grading Rubric
This document defines the scoring criteria for evaluating Reddit post/rumour virality potential. Each attribute is scored from **0.0 (no presence)** to **1.0 (very strong)**. These scores are used by the LLM to grade injected rumours before simulation.
Project 2: Full-Stack Application
**Due:** Friday by 2:59am
Sprint 0 Marking Scheme
**Team Name:** BC Hub
Criteria 1: Quality of Exploratory Data Analysis (20%) [20]
module_title: Data Science and Machine Learning
TA Grading Rubric
> **Core Principle:** Focus on whether the student has demonstrated mastery of the
RE-Bench Formal Scoring Rubric
RE-Bench evaluates reverse engineering LLMs across seven orthogonal axes:
Sprint 0 Marking Scheme
**Team Name:** [Byte-Peeps]
Evaluation Metrics for Language Model Comparative Analysis
This document defines the metrics used to evaluate the performance of different language models in generating Python game scripts. The metrics focus on three key areas: Accuracy, Bug Frequency, and Feature Completeness.
proj2rubric
Repository link: https://github.com/agupta15k/ncsu_se_fall22_22_pr_2
Evaluation Rubric for LLM-Generated Outputs
title: Evaluation Rubric for LLM Outputs
proj3rubric
|Score|Notes| Evidence|Self-Assessment|