Back to .md Directory

Judging Rubric

**AI for Social Good Hackathon – SUST 2026**

May 2, 2026
0 downloads
2 views
ai eval workflow
View source

Judging Rubric

AI for Social Good Hackathon – SUST 2026


Scoring Overview

Projects are evaluated by a panel of judges across five criteria, totalling 100 points.

CriterionWeightDescription
Social Impact & Relevance25 ptsReal-world problem, clear target beneficiary
Z.ai Integration Quality25 ptsMeaningful, non-superficial use of Z.ai
Technical Execution20 ptsWorking demo, code quality, reliability
User Experience & Clarity15 ptsUsability, simplicity, clarity for non-technical users
Demo & Communication15 ptsTeam presentation, explanation, confidence

Maximum total: 100 points


Detailed Criteria

1. Social Impact & Relevance (25 pts)

Score RangeDescription
21–25Problem is clearly defined, meaningful, and directly impacts a real community. Target users and expected benefit are compelling.
16–20Good problem choice with adequate justification. Impact is plausible but not fully articulated.
11–15Problem is generic or lacks specificity. Social impact is vague.
6–10Weak problem framing. Limited relevance to social good.
0–5No clear social purpose or the use case is trivial/irrelevant.

2. Z.ai Integration Quality (25 pts)

Score RangeDescription
21–25Z.ai is central to the solution. Integration is non-trivial, contextually appropriate, and clearly enhances the user experience.
16–20Z.ai is used meaningfully and contributes to the solution, though not central.
11–15Z.ai is included but the integration feels superficial or bolted on.
6–10Minimal or unclear Z.ai usage. Documentation of usage is lacking.
0–5Z.ai is absent or the integration is non-functional.

3. Technical Execution (20 pts)

Score RangeDescription
17–20Project runs reliably. Code is organized and well-documented. Architecture suits the problem and 1-day scope.
12–16Project mostly works. Minor bugs or incomplete features, but core functionality is demonstrated.
7–11Some functionality present but significant issues or missing features.
3–6Barely functional. Major gaps in implementation.
0–2Non-functional or empty repository.

4. User Experience & Clarity (15 pts)

Score RangeDescription
13–15Intuitive, clean interface or workflow. A non-technical user could understand and use the product.
9–12Usable with minor rough edges. User flow is mostly clear.
5–8Functional but confusing or difficult to navigate.
1–4Poor UX; requires significant explanation to operate.
0No user-facing interface or experience considered.

5. Demo & Communication (15 pts)

Score RangeDescription
13–15Clear, concise, and confident presentation. Demo shows real functionality. Team answers questions well.
9–12Good presentation with minor clarity issues. Demo covers the core features.
5–8Partial demo or unclear explanation. Some key features not demonstrated.
1–4Weak or disorganized presentation. Demo does not prove functionality.
0No presentation or demo.

Tie-Break Policy

If two or more teams have equal total scores:

  1. Higher Technical Execution score wins.
  2. If still tied, higher Social Impact & Relevance score wins.
  3. If still tied, the judges panel holds a brief discussion and majority decision.

Judge Conduct

  • Judges score independently before any group discussion.
  • Judges must declare conflicts of interest with any team before scoring.
  • Judges should provide at least one line of written feedback per team.
  • Scores are final once submitted to the scoring sheet.
  • Judges must not share individual scores with participants before the official announcement.

Scoring Template

See judging/judge-evaluation-sheet.md for the per-team scoring form judges will use.

Related Documents