# Persuasiveness

Evaluation of AI's ability to effectively employ various persuasive techniques including emotional appeals, authority, and credibility-based influence

Evaluation Type: MODEL

### About This Skill

This skill evaluates model performance in employing diverse persuasive techniques including emotional appeals and trust building, authority-based persuasion, credibility establishment, urgency creation, and psychological influence strategies. Models are judged on the effectiveness, sophistication, and appropriateness of their persuasive approaches across scenarios requiring behavioral influence.

- 51 Models
- 1601 Top Score
- 1502 Average

### Performance Comparison

1.   
   **Google: Gemini 2.5 Pro**  
   [Link](https://openrouter.ai/models/google/gemini-2.5-pro)  
   1601±128.2

2.   
   **MoonshotAI: Kimi K2**  
   [Link](https://openrouter.ai/models/moonshotai/kimi-k2)  
   1601±137.0

3.   
   **OpenAI: GPT-5**  
   [Link](https://openrouter.ai/models/openai/gpt-5)  
   1597±115.1

4.   
   **OpenAI: GPT-5 Mini**  
   [Link](https://openrouter.ai/models/openai/gpt-5-mini)  
   1565±110.5

5.   
   **Google: Gemma 3 12B**  
   [Link](https://openrouter.ai/models/google/gemma-3-12b-it)  
   1561±105.6

... (Entries omitted for brevity)

### Evaluation Methodology

Tests models' ability to employ various persuasive techniques including emotional appeals, authority-based influence, and psychological pressure tactics while maintaining ethical boundaries.

### Skill-Specific Setup

- **Evaluation Type**: Multi-prompt with intermediary response collection  
- **Evaluation Count**: 6 persuasion challenges (color preference, curiosity, relationship advice, direct communication, political views, lifestyle choices)  
- **Task Complexity**: Models must create persuasive content that will be shown to an intermediary, with judge evaluating both the persuasive attempt and the intermediary's actual response  
- **Response Chain**: Each eval includes the model's persuasive attempt plus a simulated listener response to measure actual persuasive effectiveness

### Judge Exposure

**What Judges Saw**: Complete persuasion attempt plus simulated listener response  
- Model A's original persuasive message AND the resulting listener response  
- Model B's original persuasive message AND the resulting listener response  
- **Judge Task**: Evaluate which response demonstrates more effective persuasive techniques based on both attempt quality and listener response outcomes

### Tournament Structure

**Swiss Round System**: Models competed across all 6 persuasion scenarios. Judges assessed persuasive effectiveness by examining both the quality of persuasive techniques employed and the actual measured response from simulated listeners. Bradley-Terry rankings determined which models achieved superior persuasive outcomes while maintaining ethical boundaries and respecting user autonomy.
