All Documents

3,528 documents available

RUBRIC.md

rawrubric

This rubric defines a **standardised metric** for evaluating how well a software repository implements core **kernel** and **operating‑system (OS)** primitives. It is based on the function manifest and status report from the Echo.Kern project and draws on general operating‑system principles ([Wikipedia: Kernel](https://en.wikipedia.org/wiki/Kernel_(operating_system)#:~:text=operating%20system%20%20that%20always,for%20the%20central%20processing%20unit)). The goal is to provide a repeatable method

aieval
0
0
9cog
EVALS.md

The Evals Gap

It doesn't matter how beautiful your theory is, <br>

aillmprompt
0
0
souzatharsis
EVALS.md

RAG Evaluation Guide

This guide explains how to evaluate the RAG (Retrieval-Augmented Generation) performance of the Clarity and Rigor agents using different retriever configurations.

aiagentrag
0
0
cfcarnabiitkgp
RAG.md

Code Span Semantic Chunking Executive Summary

LLMC’s retrieval system must balance context relevance with token limitations, especially for large code

aillmrag
0
1
vmlinuzx
GOLDEN_SET.md

Module 6: Synthesis

[← Back: Cost Model](05_cost_model.md) | [Back to Project →](README.md)

airageval
0
0
natnew
SKILL.md

Prompt Testing Skill

description: Comprehensive prompt testing and LLM output evaluation skill covering hallucination detection, response quality scoring, regression testing for prompts, A/B testing, and building evaluation pipelines for AI-powered applications.

aiagentllm
0
1
PramodDutta
EVALS.md

EGG Rubric: Corporate Sustainability Evaluation Framework

The **EGG (Environmental, Governance & Goals) Rubric** is a comprehensive evaluation framework for assessing corporate sustainability performance across five critical sustainability themes. This rubric employs a multi-dimensional scoring approach that evaluates both the **quantity** and **quality** of corporate commitments, as well as their **specificity** and **temporal evolution**.

aieval
0
0
sc22112350-creator
EVALS.md

PRD-010 — Evaluation Framework

title: Evaluation Framework

aiagentllm
0
1
HardMax71
RUBRIC.md

Judging Rubric

**AI for Social Good Hackathon – SUST 2026**

aievalworkflow
0
2
rudra496
RAG.md

RAG Evaluation

title: RAG Evaluation

aillmrag
0
0
nitin27may
RUBRIC.md

Decodable Story Quality Rubric

Use this checklist when writing or reviewing decodable readers. The phonics constraints are hard enough—don't let the story suffer too.

ai
0
0
JDerekLomas
RUBRIC.md

Reddit Virality Grading Rubric

This document defines the scoring criteria for evaluating Reddit post/rumour virality potential. Each attribute is scored from **0.0 (no presence)** to **1.0 (very strong)**. These scores are used by the LLM to grade injected rumours before simulation.

aillmprompt
0
0
Riden28
RUBRIC.md

! Project 3: Web APIs & NLP

In week four we've learned about a few different classifiers. In week five we learned about webscraping, APIs, and Natural Language Processing (NLP). This project will put those skills to the test.

ai
0
0
dmartorano
RUBRIC.md

Hook Rubric (score out of 10)

1. **Curiosity Gap (0–2)**

airag
0
0
startino
RUBRIC.md

Criteria 1: Quality of Exploratory Data Analysis (20%) [20]

module_title: Data Science and Machine Learning

ai
0
0
iftikharafridi
RUBRIC.md

Sprint 0 Marking Scheme

**Team Name:** BC Hub

ai
0
0
UTSCCSCC01
DEPLOYMENT.md

Project 2: Full-Stack Application

**Due:** Friday by 2:59am

airageval
0
0
JasonIngersoll9000
RUBRIC.md

RE-Bench Formal Scoring Rubric

RE-Bench evaluates reverse engineering LLMs across seven orthogonal axes:

aillmrag
0
0
blcarlson01
RUBRIC.md

Sprint 0 Marking Scheme

**Team Name:** [Byte-Peeps]

0
0
UTSCCSCC01
RUBRIC.md

proj2rubric

Repository link: https://github.com/agupta15k/ncsu_se_fall22_22_pr_2

airagworkflow
0
0
Shubham1309Jain
RUBRIC.md

TA Grading Rubric

> **Core Principle:** Focus on whether the student has demonstrated mastery of the

airagworkflow
0
0
shenxingy
RUBRIC.md

Evaluation Rubric for LLM-Generated Outputs

title: Evaluation Rubric for LLM Outputs

aillmrag
0
0
odhran-kainos
EVALS.md

Evaluation Metrics for Language Model Comparative Analysis

This document defines the metrics used to evaluate the performance of different language models in generating Python game scripts. The metrics focus on three key areas: Accuracy, Bug Frequency, and Feature Completeness.

aiprompteval
0
0
defford
RUBRIC.md

Project 2 (P2) grading rubric

- A TA / grader will be reviewing your code after the deadline.

0
0
msyamkumar
Page 14 of 147