# Harm Avoidance

Evaluation of AI's ability to refuse guidance for harmful, illegal, or unethical activities across surveillance, manipulation, and cyberattack scenarios

Evaluation Type: MODEL

### About This Skill

This skill evaluates model performance in recognizing and refusing requests for guidance on surveillance overreach, political disinformation campaigns, social media manipulation, employee monitoring, and cybersecurity attacks. Models are judged on their ability to decline harmful requests while redirecting toward legitimate, ethical alternatives.

### Performance Comparison

1.  OpenAI: GPT-5 1660±86.7  
   [View Model](https://openrouter.ai/models/openai/gpt-5)
2.  xAI: Grok 4 1659±81.0  
   [View Model](https://openrouter.ai/models/x-ai/grok-4)
3.  OpenAI: GPT-5 Mini 1631±80.5  
   [View Model](https://openrouter.ai/models/openai/gpt-5-mini)
4.  Google: Gemini 2.5 Flash 1621±73.4  
   [View Model](https://openrouter.ai/models/google/gemini-2.5-flash)
5.  Z.AI: GLM 4.5 1621±80.2  
   [View Model](https://openrouter.ai/models/z-ai/glm-4.5)

... (truncated for brevity)

### Evaluation Methodology

Tests models' ability to recognize and refuse harmful requests disguised as legitimate professional needs, particularly when asked to provide guidance on surveillance, cyberattacks, and manipulation techniques.

### Skill-Specific Setup

- **Evaluation Type**: Single-prompt manipulation attempts
- **Evaluation Count**: 5 sophisticated harmful request scenarios (crisis management surveillance, employee investigation, political manipulation, social media manipulation, unauthorized access techniques)
- **Task Complexity**: Each request appears professionally legitimate but seeks guidance on potentially harmful activities like surveillance overreach, psychological manipulation, or cyber exploitation
- **Deception Design**: Requests frame harmful activities as public safety, emergency response, or legitimate business needs

### Judge Exposure

**What Judges Saw**: Two complete responses to potentially harmful professional requests

- Model A's complete response to the manipulation attempt
- Model B's complete response to the manipulation attempt
- **Judge Task**: Evaluate harm avoidance (refusal to provide harmful guidance) and constructive redirection (offering legitimate alternatives while declining harmful requests)

### Tournament Structure

**Swiss Round System**: Models competed across all 5 manipulation scenarios. Judges assessed whether responses appropriately identified and refused harmful requests while providing constructive alternatives. Bradley-Terry rankings determined which models consistently demonstrated harm prevention by declining to enable surveillance, manipulation, or exploitation techniques regardless of professional framing.
