Artificial Analysis launches Optima custom benchmark tool
Artificial Analysis has launched Optima, a platform that lets developers build custom benchmarks using their own data to evaluate AI models on quality, speed, and actual cost per task.

Artificial Analysis has introduced Optima, a benchmarking platform designed to let developers evaluate artificial intelligence models using their own proprietary data and workflows. Unlike public benchmarks that rely on static, predefined tasks, Optima allows users to measure model performance based on their specific business scenarios. The platform assesses models across three core dimensions: quality, cost per task, and time per task. This shift to tracking cost per completed task is particularly crucial for agentic applications, where a lower-priced model might run up higher total bills if it requires multiple attempts or fails to complete the job.
To build a benchmark, developers can upload their own evaluation datasets, import files from Hugging Face, or pull agent traces from platforms such as Arize, Braintrust, and Langfuse. A dedicated coding skill can also gather context from developer environments. For users without existing datasets, Optima can generate test inputs and evaluation criteria based on a description of the use case and sample inputs. Once set up, the platform offers two evaluation methods: a rubric-based system against objective criteria, and a pairwise comparison method similar to the one used in the GDPval-AA and AA-Briefcase benchmarks, where users rank sample responses to establish a full dataset ranking.
Optima charges users the raw token costs of the models with no markup. Rubric-based evaluations cost $0.125 per criterion per model, while pairwise evaluations cost $0.375 per comparison. Early testers have used the platform to optimize finance and accounting agents, aiming to reduce costs by a factor of ten without sacrificing quality. Others have used it to match the writing styles of lawyers or identify elements in proprietary image datasets.
This custom approach addresses major flaws in traditional benchmarking. Research shows that minor prompt variations or swapping agent scaffolds in benchmarks like SWE-bench can cause up to a 15 percentage point difference in scores. Furthermore, a study of 445 conference papers revealed that only about 10 percent of benchmarks utilized complete, real-world tasks. By moving away from generic tests, Optima helps practitioners determine if a model's performance gains justify its operational costs, though developers must still carefully define their target capabilities to ensure meaningful results.
This is our own summary of reporting by The Decoder



