Domain testing suites
Each domain is a curated evaluation template covering a distinct set of capability categories. Scores reflect finalized runs only. Click any model row to highlight it across all domains at once.
Core Evaluation
Default evaluation template
Planning Evaluation
This battery tests a models ability to create a plan and then create instructions for a smaller AI Model to execute the plan successfully.
Response Consistency
Response Consistency — does the model approach the same situation the same way every time, or drift depending on how you ask? Always milk for skunk stink, the same database recommendation however the question is framed. Not whether it's right — whether it's dependable enough to build on.
Verbosity
Verbosity measures whether a model says the right amount — length and register calibrated to the task, not minimized. Every other benchmark asks if the answer is correct; this one asks if it was the right size to be useful, and what it costs you when it isn't. Models are trained toward over-explaining, and Verbosity exposes that directly — the padding reflex, the failure to compress, and the meltdown under nonsense — across everyday conversation, copywriting, functional deliverables, explicit length constraints, and prompts with no precedent at all. Higher scores represent fit, not token count; "appropriate + relevant answer" is the target - correctness of answer is not judged here, only whether it is addressing the question itself and what it does with that question.
