Recall | JavaScript Coding Leaderboard
JavaScript Coding
Evaluation of AI's ability to create interactive browser-based games using JavaScript
Evaluation Type: MODEL
About This Skill
This skill evaluates model short-horizon coding performance in creating complete, functional JavaScript games that render in HTML5 Canvas with real-time interaction, animations, and game logic. Models are judged on code quality, visual presentation, and successfully delivering the requirements.
Total Models: 51
Top Score: 1671
Average Score: 1502
Performance Comparison
Google: Gemini 2.5 Pro
Score: 1671±65.3
LinkOpenAI: GPT-5 Mini
Score: 1648±67.2
LinkOpenAI: GPT-5
Score: 1643±77.2
LinkAnthropic: Claude 3.7 Sonnet (thinking)
Score: 1607±45.1
LinkAnthropic: Claude Opus 4
Score: 1586±41.8
LinkAnthropic: Claude Opus 4.1
Score: 1578±44.0
LinkAnthropic: Claude Sonnet 4
Score: 1574±40.8
LinkQwen: Qwen3 235B A22B Instruct 2507
Score: 1574±39.2
LinkGoogle: Gemini 2.5 Flash
Score: 1564±42.0
Link**OpenAI:** GPT-4.1 Mini **Score:** 1561±36.6 [Link](https://openrouter.ai/models/openai/gpt-4.1-mini)
... (more data omitted for brevity)
Evaluation Methodology
Tests models' ability to create complete, functional JavaScript games with advanced features using simple canvas rendering.
Skill-Specific Setup
- Evaluation Type: Single-prompt technical implementation
- Evaluation Count: 7 distinct game challenges (Conway's Game of Life, Flight Simulator, Maze Generator, Procedural FPS Map, Racing Game, Sorting Visualizer, Space Invaders)
- Task Complexity: Each evaluation requires implementing a complete interactive game with multiple features, UI controls, and smooth performance.
- Expected Output: Valid JavaScript code only, no explanatory text or markdown formatting.
Judge Exposure
What Judges Saw: Two complete JavaScript implementations side-by-side
- Model A's complete code solution
- Model B's complete code solution
- Judge Task: Evaluate functionality, correctness, and feature completeness without penalizing different coding styles.
Tournament Structure
Swiss Round System: Each model competed against others across all 7 game development challenges. Judges compared pairs of implementations, selecting the superior solution based on technical merit, feature completeness, and code quality. Rankings were determined using Bradley-Terry scoring across all comparisons.