# Document Summarization

Evaluation of AI's ability to create concise, accurate summaries of ArXiv research papers

## Evaluation Type: MODEL

### About This Skill

This skill evaluates model performance in summarizing academic research papers from ArXiv across diverse scientific domains. Models are judged on content coverage of main contributions, technical accuracy, clarity of structure, appropriate length, and fidelity to source material compared to human-generated reference summaries.

**51 Models**  
**1558 Top Score**  
**1503 Average Score**

### Performance Comparison

1. [MoonshotAI: Kimi K2](https://openrouter.ai/models/moonshotai/kimi-k2) - *moonshotai*  
   Score: 1558±63.5

2. [OpenAI: o3](https://openrouter.ai/models/openai/o3) - *openai*  
   Score: 1558±63.9

3. [DeepSeek: R1](https://openrouter.ai/models/deepseek/deepseek-r1) - *deepseek*  
   Score: 1550±61.6

4. [Z.AI: GLM 4.5](https://openrouter.ai/models/z-ai/glm-4.5) - *z-ai*  
   Score: 1544±68.1

5. [Qwen: Qwen3 235B A22B Instruct 2507](https://openrouter.ai/models/qwen/qwen3-235b-a22b-2507) - *qwen*  
   Score: 1544±57.0

6. [Perplexity: Sonar Pro](https://openrouter.ai/models/perplexity/sonar-pro) - *perplexity*  
   Score: 1543±57.3

7. [Perplexity: Sonar](https://openrouter.ai/models/perplexity/sonar) - *perplexity*  
   Score: 1542±55.9

8. [OpenAI: Codex Mini](https://openrouter.ai/models/openai/codex-mini) - *openai*  
   Score: 1540±56.6

9. [OpenAI: GPT-4.1 Mini](https://openrouter.ai/models/openai/gpt-4.1-mini) - *openai*  
   Score: 1536±59.5

10. [OpenAI: GPT-5](https://openrouter.ai/models/openai/gpt-5) - *openai*  
    Score: 1535±57.8

### Evaluation Methodology

Tests models' ability to create concise, accurate summaries of academic research papers from ArXiv, judged against human reference summaries.

### Skill-Specific Setup

- **Evaluation Type**: Single-prompt with document input  
- **Evaluation Count**: 5 ArXiv research papers across different domains  
- **Task Complexity**: Models must read full academic papers and extract main contributions, methodology, and key findings  
- **Input Materials**: Complete research papers provided as document text, with human-authored reference abstracts for comparison

### Judge Exposure

**What Judges Saw**: Two candidate summaries compared against a human reference summary  
- Model A's complete summary  
- Model B's complete summary  
- Human reference summary (gold standard)  
- **Judge Task**: Score using structured rubric (Content Coverage 0-4, Accuracy 0-3, Clarity & Structure 0-2, Conciseness 0-1) and return JSON evaluation

### Tournament Structure

**Swiss Round System**: Models competed across all 5 research papers. Judges used the standardized scoring rubric to evaluate technical accuracy, comprehensiveness, and clarity. Bradley-Terry rankings incorporated the structured scoring to determine relative performance across the academic summarization domain.
