Research

Artificial Analysis Ranks AI Models on Real-World Jobs

Artificial Analysis has launched version 1.1 of its Capability Indices, updating its industry-specific leaderboards to help developers evaluate AI models on realistic professional tasks.

AlphaSignal2 days agoResearch
Image: AlphaSignal

Artificial Analysis has released version 1.1 of its Capability Indices, a suite of six domain-specific leaderboards designed to evaluate how artificial intelligence models perform on actual professional workloads. By mapping occupational tasks from the US O*NET database to specific benchmarks, the firm weights performance based on how frequently certain skills are required in real jobs. The updated indices cover six distinct industry verticals: Finance and Accounting, Strategy and Ops, Legal, Healthcare and Medical, Engineering, and Economics.

The v1.1 update introduces several benchmark changes to emphasize multistep tool use and long-document reasoning. It adds AutomationBench-AA slices to the finance, strategy, legal, and healthcare indices, and integrates AA-Briefcase into agentic knowledge work across all six verticals. Conversely, agentic customer interaction has been removed from four indices. For engineering, Terminal-Bench v4.0 replaces GPQA Diamond to better test command-line execution, while the healthcare index incorporates MLCR-AA to evaluate medical reasoning over long inputs.

Under the new rankings, Anthropic's Claude Fable 5.1 (max) leads all six indices, while OpenAI's GPT-6 Astra (max) secures second place in four categories: Finance and Accounting, Strategy and Ops, Legal, and Engineering. Although no open-weights model cracked the top five, several competitive options landed in the top ten across various verticals, including Kimi K3, DeepSeek V4.1 Flash, and GLM-5.3.

For practitioners, these indices move beyond generic, compressed intelligence scores to offer a more granular view of model capabilities. Developers building specialized applications can use these targeted rankings to match a model's strengths to their specific workloads. For example, a team building a legal document analysis tool might prioritize long-context benchmarks like GDP.pdf over headline rankings, while those building financial agents would focus heavily on agentic tool use and AA-Briefcase scores.

This is our own summary of reporting by AlphaSignal

More in Research