Recall | Document Summarization Leaderboard

Document Summarization

Evaluation of AI's ability to create concise, accurate summaries of ArXiv research papers

Evaluation Type: MODEL

About This Skill

This skill evaluates model performance in summarizing academic research papers from ArXiv across diverse scientific domains. Models are judged on content coverage of main contributions, technical accuracy, clarity of structure, appropriate length, and fidelity to source material compared to human-generated reference summaries.

51 Models
1558 Top Score
1503 Average Score

Performance Comparison

  1. MoonshotAI: Kimi K2 - moonshotai
    Score: 1558±63.5

  2. OpenAI: o3 - openai
    Score: 1558±63.9

  3. DeepSeek: R1 - deepseek
    Score: 1550±61.6

  4. Z.AI: GLM 4.5 - z-ai
    Score: 1544±68.1

  5. Qwen: Qwen3 235B A22B Instruct 2507 - qwen
    Score: 1544±57.0

  6. Perplexity: Sonar Pro - perplexity
    Score: 1543±57.3

  7. Perplexity: Sonar - perplexity
    Score: 1542±55.9

  8. OpenAI: Codex Mini - openai
    Score: 1540±56.6

  9. OpenAI: GPT-4.1 Mini - openai
    Score: 1536±59.5

  10. OpenAI: GPT-5 - openai
    Score: 1535±57.8

Evaluation Methodology

Tests models' ability to create concise, accurate summaries of academic research papers from ArXiv, judged against human reference summaries.

Skill-Specific Setup

Judge Exposure

What Judges Saw: Two candidate summaries compared against a human reference summary

Tournament Structure

Swiss Round System: Models competed across all 5 research papers. Judges used the standardized scoring rubric to evaluate technical accuracy, comprehensiveness, and clarity. Bradley-Terry rankings incorporated the structured scoring to determine relative performance across the academic summarization domain.