Recall | Respect No Em Dashes Leaderboard

Respect No Em Dashes

Evaluation of AI's ability to follow the specific instruction to avoid using em dashes in extended writing

About This Skill

This skill evaluates model performance in following the specific formatting constraint of avoiding em dashes (—) in extended writing tasks while maintaining content quality, flow, and readability. Models are judged on both perfect compliance with the punctuation restriction and preservation of writing effectiveness despite the constraint.

Performance Comparison

Rank Model Score (± SD) Link
1 Meta: Llama 3.1 405B (base) 1568 (±109.8)
2 ReMM SLERP 13B 1548 (±107.4)
3 OpenAI: o1 1536 (±105.7)
4 Anthropic: Claude 3.7 Sonnet (thinking) 1534 (±119.9)
5 OpenAI: GPT-5 1533 (±113.9)
6 Anthropic: Claude Opus 4 1529 (±106.6)
7 DeepSeek: DeepSeek V3 1528 (±113.7)
8 Perplexity: Sonar 1528 (±101.6)
9 AI21: Jamba Large 1.7 1523 (±99.6)
10 DeepSeek: R1 1523 (±103.7)

Evaluation Methodology

Tests models' ability to follow explicit formatting constraints (avoiding em dashes) while maintaining content quality and coherence in extended writing tasks.

Skill-Specific Setup

Judge Exposure

What Judges Saw: Complete expanded blog posts from both models

Tournament Structure

Swiss Round System: Models competed across all 3 writing topics through the two-stage expansion process. Judges focused exclusively on instruction-following compliance by checking for em dash usage violations. Bradley-Terry rankings identified models that consistently followed explicit formatting constraints without debate, justification, or violation of user-specified requirements.