Recall | Document Summarization Leaderboard
Document Summarization
Evaluation of AI's ability to create concise, accurate summaries of ArXiv research papers
Evaluation Type: MODEL
About This Skill
This skill evaluates model performance in summarizing academic research papers from ArXiv across diverse scientific domains. Models are judged on content coverage of main contributions, technical accuracy, clarity of structure, appropriate length, and fidelity to source material compared to human-generated reference summaries.
51 Models
1558 Top Score
1503 Average Score
Performance Comparison
MoonshotAI: Kimi K2 - moonshotai
Score: 1558±63.5OpenAI: o3 - openai
Score: 1558±63.9DeepSeek: R1 - deepseek
Score: 1550±61.6Z.AI: GLM 4.5 - z-ai
Score: 1544±68.1Qwen: Qwen3 235B A22B Instruct 2507 - qwen
Score: 1544±57.0Perplexity: Sonar Pro - perplexity
Score: 1543±57.3Perplexity: Sonar - perplexity
Score: 1542±55.9OpenAI: Codex Mini - openai
Score: 1540±56.6OpenAI: GPT-4.1 Mini - openai
Score: 1536±59.5OpenAI: GPT-5 - openai
Score: 1535±57.8
Evaluation Methodology
Tests models' ability to create concise, accurate summaries of academic research papers from ArXiv, judged against human reference summaries.
Skill-Specific Setup
- Evaluation Type: Single-prompt with document input
- Evaluation Count: 5 ArXiv research papers across different domains
- Task Complexity: Models must read full academic papers and extract main contributions, methodology, and key findings
- Input Materials: Complete research papers provided as document text, with human-authored reference abstracts for comparison
Judge Exposure
What Judges Saw: Two candidate summaries compared against a human reference summary
- Model A's complete summary
- Model B's complete summary
- Human reference summary (gold standard)
- Judge Task: Score using structured rubric (Content Coverage 0-4, Accuracy 0-3, Clarity & Structure 0-2, Conciseness 0-1) and return JSON evaluation
Tournament Structure
Swiss Round System: Models competed across all 5 research papers. Judges used the standardized scoring rubric to evaluate technical accuracy, comprehensiveness, and clarity. Bradley-Terry rankings incorporated the structured scoring to determine relative performance across the academic summarization domain.