# JavaScript Coding

Evaluation of AI's ability to create interactive browser-based games using JavaScript

**Evaluation Type:** MODEL

### About This Skill

This skill evaluates model short-horizon coding performance in creating complete, functional JavaScript games that render in HTML5 Canvas with real-time interaction, animations, and game logic. Models are judged on code quality, visual presentation, and successfully delivering the requirements.

**Total Models:** 51  
**Top Score:** 1671  
**Average Score:** 1502

### Performance Comparison

1.   
   **Google:** Gemini 2.5 Pro  
   **Score:** 1671±65.3  
   [Link](https://openrouter.ai/models/google/gemini-2.5-pro)

2.   
   **OpenAI:** GPT-5 Mini  
   **Score:** 1648±67.2  
   [Link](https://openrouter.ai/models/openai/gpt-5-mini)

3.   
   **OpenAI:** GPT-5  
   **Score:** 1643±77.2  
   [Link](https://openrouter.ai/models/openai/gpt-5)

4.   
   **Anthropic:** Claude 3.7 Sonnet (thinking)  
   **Score:** 1607±45.1  
   [Link](https://openrouter.ai/models/anthropic/claude-3.7-sonnet:thinking)

5.   
   **Anthropic:** Claude Opus 4  
   **Score:** 1586±41.8  
   [Link](https://openrouter.ai/models/anthropic/claude-opus-4)

6.   
   **Anthropic:** Claude Opus 4.1  
   **Score:** 1578±44.0  
   [Link](https://openrouter.ai/models/anthropic/claude-opus-4.1)

7.   
   **Anthropic:** Claude Sonnet 4  
   **Score:** 1574±40.8  
   [Link](https://openrouter.ai/models/anthropic/claude-sonnet-4)

8.   
   **Qwen:** Qwen3 235B A22B Instruct 2507  
   **Score:** 1574±39.2  
   [Link](https://openrouter.ai/models/qwen/qwen3-235b-a22b-2507)

9.   
   **Google:** Gemini 2.5 Flash  
   **Score:** 1564±42.0  
   [Link](https://openrouter.ai/models/google/gemini-2.5-flash)

10.   
    **OpenAI:** GPT-4.1 Mini  
    **Score:** 1561±36.6  
    [Link](https://openrouter.ai/models/openai/gpt-4.1-mini)

... (more data omitted for brevity)

### Evaluation Methodology

Tests models' ability to create complete, functional JavaScript games with advanced features using simple canvas rendering.

### Skill-Specific Setup
- **Evaluation Type:** Single-prompt technical implementation  
- **Evaluation Count:** 7 distinct game challenges (Conway's Game of Life, Flight Simulator, Maze Generator, Procedural FPS Map, Racing Game, Sorting Visualizer, Space Invaders)  
- **Task Complexity:** Each evaluation requires implementing a complete interactive game with multiple features, UI controls, and smooth performance.  
- **Expected Output:** Valid JavaScript code only, no explanatory text or markdown formatting.

### Judge Exposure
**What Judges Saw:** Two complete JavaScript implementations side-by-side
- Model A's complete code solution  
- Model B's complete code solution  
- **Judge Task:** Evaluate functionality, correctness, and feature completeness without penalizing different coding styles.

### Tournament Structure
**Swiss Round System:** Each model competed against others across all 7 game development challenges. Judges compared pairs of implementations, selecting the superior solution based on technical merit, feature completeness, and code quality. Rankings were determined using Bradley-Terry scoring across all comparisons.
