Recall | Compassionate Communication Leaderboard

Compassionate Communication

Evaluation of AI's ability to deliver devastating personal news with appropriate empathy and support

Evaluation Type: MODEL

About This Skill

This skill evaluates model performance in communicating devastating personal news including terminal diagnoses, relationship betrayals, family discoveries, and other life-altering information. Models are judged on their ability to balance honest communication with compassionate delivery, emotional support, and recognition of the profound psychological impact such information has on recipients.

Performance Comparison

  1. xAI: Grok 4

    Score: 1630±68.5

  2. Qwen: Qwen3 235B A22B Instruct 2507

    Score: 1606±46.2

  3. OpenAI: GPT-5

    Score: 1601±40.1

  4. OpenAI: GPT-5 Mini

    Score: 1592±43.0

  5. Google: Gemini 2.5 Pro

    Score: 1591±46.1

  6. OpenAI: o3 Pro

    Score: 1589±39.4

  7. Qwen: Qwen3 32B

    Score: 1577±42.0

  8. Qwen: Qwen-Max

    Score: 1576±39.3

  9. Google: Gemini 2.5 Flash

    Score: 1573±40.3

  10. OpenAI: o3

    Score: 1571±41.6

Evaluation Methodology

Tests models' ability to deliver difficult news and provide emotional support in often devastating personal crisis scenarios with appropriate compassion and professional boundaries.

Skill-Specific Setup

Judge Exposure

What Judges Saw: Two complete responses to crisis scenarios

Tournament Structure

Swiss Round System: Models competed across all 7 crisis communication scenarios. Judges assessed the balance between honest communication of difficult realities and empathetic emotional support. Bradley-Terry scoring determined which models consistently provided appropriate compassionate communication without crossing professional boundaries or providing dangerous advice.