# StackPilot Model Role Benchmark

Use this worksheet to choose an AI model for a real job—not from a launch headline.

## 1. Freeze the job

- Business task:
- Representative input:
- Required output:
- Acceptance checks:
- Maximum useful time:
- Maximum acceptable cost:

## 2. Freeze the conditions

- Date and time:
- Product surface: ChatGPT / Work / Codex / API / other
- Plan or access tier:
- Model and reasoning level:
- Tools or connectors enabled:
- Same prompt and attachments used for every run: yes / no

## 3. Run receipt

| Model route | Usable result (0–5) | Corrections | Minutes | Estimated total cost | Important failure |
|---|---:|---:|---:|---:|---|
| Route A |  |  |  |  |  |
| Route B |  |  |  |  |  |
| Route C |  |  |  |  |  |

## 4. Decide

- Best route for this job:
- Why it won:
- What still needs human review:
- When to rerun this benchmark:

## Approval wall

A benchmark result may guide private, reversible work. It does not authorize publishing, messaging, spending, deploying, changing an account, or using private identity or payment data.

StackPilot Guides · `stackpilotguides.com/pages/guide-solo-stack-planner.html`
