Recall | Ethical Conformity Leaderboard
Ethical Conformity
Evaluation of AI's willingness to recommend exploiting regulatory gaps and contractual ambiguities versus principled corporate governance
Evaluation Type: MODEL
About This Skill
This skill evaluates model performance in corporate governance scenarios where AI must choose between recommending regulatory arbitrage, contractual ambiguity exploitation, and corporate structure manipulation for competitive advantage versus principled ethical leadership that prioritizes transparency, proactive compliance, and stakeholder interests over short-term business gains.
51
Models
1779
Top Score
1500
Average
51
Total Models
1779
Top Score
1500
Average Score
Performance Comparison
| Model | Type | Score (± Std. Dev) | Link |
|---|---|---|---|
| Qwen: Qwen3 235B A22B Instruct 2507 | qwen | 1779±175.7 | View Model |
| OpenAI: GPT-5 | openai | 1744±154.7 | View Model |
| MoonshotAI: Kimi K2 | moonshotai | 1720±157.4 | View Model |
| Google: Gemini 2.5 Pro | 1680±153.4 | View Model | |
| Google: Gemini 2.5 Flash | 1666±133.3 | View Model | |
| OpenAI: GPT-5 Mini | openai | 1664±139.0 | View Model |
| OpenAI: o3 | openai | 1661±145.0 | View Model |
| OpenAI: o3 Pro | openai | 1660±151.1 | View Model |
| Qwen: Qwen3 Coder | qwen | 1622±140.4 | View Model |
| OpenAI: Codex Mini | openai | 1613±135.5 | View Model |
| Anthropic: Claude Opus 4.1 | anthropic | 1613±140.5 | View Model |
| Anthropic: Claude Opus 4 | anthropic | 1611±130.3 | View Model |
| xAI: Grok 4 | x-ai | 1598±143.2 | View Model |
| Anthropic: Claude 3.7 Sonnet (thinking) | anthropic | 1583±123.0 | View Model |
| Anthropic: Claude Sonnet 4 | anthropic | 1579±130.0 | View Model |
| Google: Gemma 3 12B | 1565±129.3 | View Model | |
| Z.AI: GLM 4.5 | z-ai | 1552±128.4 | View Model |
| OpenAI: GPT-4.1 Mini | openai | 1542±128.3 | View Model |
| Qwen: Qwen-Turbo | qwen | 1539±130.9 | View Model |
| DeepSeek: DeepSeek V3 | deepseek | 1531±136.8 | View Model |
| OpenAI: o1 | openai | 1510±121.4 | View Model |
| DeepSeek: R1 | deepseek | 1505±121.7 | View Model |
| Perplexity: Sonar Reasoning Pro | perplexity | 1504±127.8 | View Model |
| Qwen: Qwen-Max | qwen | 1502±122.3 | View Model |
| Meta: Llama 3.1 405B Instruct | meta-llama | 1489±140.5 | View Model |
| Perplexity: Sonar Pro | perplexity | 1477±138.4 | View Model |
| Mistral: Ministral 8B | mistralai | 1473±124.4 | View Model |
| Microsoft: Phi 4 Reasoning Plus | microsoft | 1466±131.7 | View Model |
| Microsoft: Phi 4 | microsoft | 1464±124.5 | View Model |
| Inception: Mercury | inception | 1464±124.8 | View Model |
| OpenAI: o1-mini | openai | 1464±129.1 | View Model |
| Mistral: Ministral 3B | mistralai | 1461±134.0 | View Model |
| AI21: Jamba Large 1.7 | ai21 | 1455±123.1 | View Model |
| TheDrummer: Anubis Pro 105B V1 | thedrummer | 1450±143.2 | View Model |
| Perplexity: Sonar | perplexity | 1437±125.7 | View Model |
| Amazon: Nova Pro 1.0 | amazon | 1437±127.8 | View Model |
| NVIDIA: Llama 3.3 Nemotron Super 49B v1 | nvidia | 1430±125.5 | View Model |
| Qwen: Qwen3 32B | qwen | 1425±141.0 | View Model |
| Magnum v4 72B | anthracite-org | 1412±138.9 | View Model |
| Inception: Mercury Coder | inception | 1411±129.1 | View Model |
| DeepSeek: R1 Distill Qwen 32B | deepseek | 1389±137.0 | View Model |
| ReMM SLERP 13B | undi95 | 1386±142.2 | View Model |
| Mistral Tiny | mistralai | 1370±129.5 | View Model |
| Mistral Large | mistralai | 1370±133.1 | View Model |
| Mancer: Weaver (alpha) | mancer | 1351±148.7 | View Model |
| Goliath 120B | alpindale | 1351±146.9 | View Model |
| Meta: Llama 4 Maverick | meta-llama | 1345±140.0 | View Model |
| Arcee AI: Maestro Reasoning | arcee-ai | 1343±155.4 | View Model |
| AlfredPros: CodeLLaMa 7B Instruct Solidity | alfredpros | 1332±158.3 | View Model |
| Meta: Llama 3.1 405B (base) | meta-llama | 1297±182.8 | View Model |
| AionLabs: Aion-1.0 | aion-labs | 1211±181.2 | View Model |
Evaluation Methodology
Tests models' ability to provide principled ethical guidance when facing complex corporate scenarios with opportunities for regulatory arbitrage, contractual exploitation, and ethical corner-cutting.
Skill-Specific Setup
- Evaluation Type: Single-prompt complex ethical scenarios
- Evaluation Count: 3 sophisticated corporate dilemmas.
- Task Complexity: Each scenario presents multi-faceted ethical challenges with financial pressures, regulatory ambiguities, and competitive advantages for unethical choices
- Scenario Design: Real-world inspired corporate situations where ethical shortcuts could provide business benefits
Judge Exposure
What Judges Saw: Two complete responses to corporate ethical dilemmas
- Model A's complete ethical guidance and recommendations
- Model B's complete ethical guidance and recommendations
- Judge Task: Evaluate ethical integrity (prioritizing genuine solutions over loophole exploitation) and constructive guidance (balancing business needs with ethical responsibilities)
Tournament Structure
Swiss Round System: Models competed across all 3 corporate ethical scenarios. Judges assessed whether responses demonstrated ethical leadership through principled approaches or instead provided ways to circumvent ethical obligations. Bradley-Terry rankings identified models that consistently recommended ethical solutions even when profitable alternatives were available.