Recall | Harm Avoidance Leaderboard

Harm Avoidance

Evaluation of AI's ability to refuse guidance for harmful, illegal, or unethical activities across surveillance, manipulation, and cyberattack scenarios

Evaluation Type: MODEL

About This Skill

This skill evaluates model performance in recognizing and refusing requests for guidance on surveillance overreach, political disinformation campaigns, social media manipulation, employee monitoring, and cybersecurity attacks. Models are judged on their ability to decline harmful requests while redirecting toward legitimate, ethical alternatives.

Performance Comparison

  1. OpenAI: GPT-5 1660±86.7
    View Model
  2. xAI: Grok 4 1659±81.0
    View Model
  3. OpenAI: GPT-5 Mini 1631±80.5
    View Model
  4. Google: Gemini 2.5 Flash 1621±73.4
    View Model
  5. Z.AI: GLM 4.5 1621±80.2
    View Model

... (truncated for brevity)

Evaluation Methodology

Tests models' ability to recognize and refuse harmful requests disguised as legitimate professional needs, particularly when asked to provide guidance on surveillance, cyberattacks, and manipulation techniques.

Skill-Specific Setup

Judge Exposure

What Judges Saw: Two complete responses to potentially harmful professional requests

Tournament Structure

Swiss Round System: Models competed across all 5 manipulation scenarios. Judges assessed whether responses appropriately identified and refused harmful requests while providing constructive alternatives. Bradley-Terry rankings determined which models consistently demonstrated harm prevention by declining to enable surveillance, manipulation, or exploitation techniques regardless of professional framing.