Nautilus Deep
79 models · updated July 2026
Evaluation domains

Domain testing suites

Each domain is a curated evaluation template covering a distinct set of capability categories. Scores reflect finalized runs only. Click any model row to highlight it across all domains at once.

Core Evaluation

Default evaluation template

AgenticCodingCreative WritingEngineeringFactual HallucinationInstruction FollowingKnowledgeLong Context ReasoningMathRAG HallucinationVoice
Top models
1
GLM 5.1
Z.ai
161,548
2
OPUS 4.7
Unknown
85,944
3
STEP 3.7 FLASH
StepFun
84,587

Planning Evaluation

This battery tests a models ability to create a plan and then create instructions for a smaller AI Model to execute the plan successfully.

System & Architecture PlanningData & API DesignLibrary / Dependency SelectionSecurity & Threat PlanningTest Strategy PlanningUX / Interaction PlanningVisual Design PlanningAccessibility PlanningPolicy-Adherence PlanningDevOps PlanningProject ScopingImplementation InstructionsWorkload Chunking & Context/Session PlanningDocumentation DesignHandoff / Diary AuthoringHarness-Fit: Autonomous-AgentHarness-Fit: Async / BackgroundHarness-Fit: Interactive / IDEHarness-Fit: Always-on OrchestratorHarness-Fit: Generative PlatformDuct Tape
Top models
1
STEP 3.7 FLASH
StepFun
121,247
2
QWEN 3.6 27B FABLE-REASONING
Qwen
120,904
3
MINIMAX M2.7
MiniMax
120,372

Response Consistency

Response Consistency — does the model approach the same situation the same way every time, or drift depending on how you ask? Always milk for skunk stink, the same database recommendation however the question is framed. Not whether it's right — whether it's dependable enough to build on.

Stance ConsistencyRecommendation ConsistencyEstimate ConsistencyRanking ConsistencyJudgment-Call ConsistencyOpen-Approach Consistency
Top models
1
QWEN 3.6 27B FABLE-REASONING
Qwen
24,220
2
GEMINI 3.1 PRO
Google
23,831
3
GEMINI 3.1 FLASH LITE
Google
23,804

Verbosity

Verbosity measures whether a model says the right amount — length and register calibrated to the task, not minimized. Every other benchmark asks if the answer is correct; this one asks if it was the right size to be useful, and what it costs you when it isn't. Models are trained toward over-explaining, and Verbosity exposes that directly — the padding reflex, the failure to compress, and the meltdown under nonsense — across everyday conversation, copywriting, functional deliverables, explicit length constraints, and prompts with no precedent at all. Higher scores represent fit, not token count; "appropriate + relevant answer" is the target - correctness of answer is not judged here, only whether it is addressing the question itself and what it does with that question.

Conversational: ShortConversational: ActivityConversational: ReflectiveConversational: Long FormConversational: Chat ResponseCopywriting: ProseCopywriting: NovelaCopywriting: NovelCopywriting: AcademicCopywriting: ScientificFunctional: Email ResponseFunctional: Email OriginationFunctional: Report OutlineFunctional: Report Rough DraftFunctional: Report Final DraftFunctional: SummarizationFunctional: Explain to a 6 Year OldFunctional: Explain to a 12 Year OldFunctional: Explain to a High School GradFunctional: Explain to an ExpertConstraint: ShortConstraint: MediumConstraint: LongBloat TrapsAbsurd
Top models
1
DEEPSEEK v4 PRO
DeepSeek
53,484
2
QWEN 3.6 27B FABLE-REASONING
Qwen
49,312
3
GEMINI 3.1 FLASH LITE
Google
46,720