← All posts

Opus, Sonnet, Haiku: How to Actually Pick One

May 15, 2026Updated #Model Selection561 words · 3 min readExperience notes阅读中文原文 ↗

The short version

Use model families to form a shortlist. Let your own task set determine the default and the conditions for upgrading.

Model selectionEvaluation methodsCost analysis
Table of contents5 sections

Define the task first

This is an experience-based selection method. It does not include model-version records, raw outputs, or measured costs, so it does not claim a benchmark ranking or treat a family name as a quality guarantee.

These combinations can help organize an initial comparison:

Task Candidates to start with What to check
Classification, extraction, simple routing Haiku and Sonnet Field correctness, failure handling, cost per task
Everyday code, documents, conversation Sonnet and a second candidate Tests passed, editing needed, response time
Complex planning and multi-step work Sonnet and Opus Completion rate, missed constraints, total workflow cost

Record the full model ID, date, prompt, and settings. Results from different versions should not be mixed without keeping that distinction.

Turn “good” into acceptance criteria

For an order classifier, require a valid label, an explicit rule for empty input, and a priority rule when a request mentions both shipping and a refund.

A long input can still be difficult. Contract extraction and financial analysis need checks for omissions, cited evidence, and human review effort. A model upgrade alone does not settle those questions.

Measure quality, latency, and cost separately. Include errors and retries in the workflow totals.

Start with the same examples

The following is a small Chinese-language classification exercise. Its reference labels come from the task rules; they are not measured model outputs. It checks the evaluation process and does not represent a full production distribution.

Download the JSONL example task set

Give every candidate the same instruction:

Classify the request as order, refund, or other.
refund: an explicit request for a refund or return; this takes priority over shipping questions.
order: a question about order status or shipping without a refund or return request.
other: any other request, including empty input.
Return only the label.
Input Reference label
我的订单到哪了? — Where is my order? order
包裹迟迟没到,我要退款。 — The parcel is late; I want a refund. refund
我要取消邮件订阅。 — I want to unsubscribe from emails. other

Add empty inputs, ambiguous words, multiple intents, and failures from your own traffic. Keep prompt-tuning examples separate from the final acceptance set.

Keep a record someone can check

For each call, save the case ID, full model ID, input, raw output, start and end times, usage, errors, and retries. Calculate costs using the prices in effect on the test date and save that source.

This minimal check reads your collected outputs from results.jsonl, where each line looks like {"id":"order-status","output":"order"}. Missing results stay in the denominator.

import json

def read_jsonl(path):
    with open(path, encoding="utf-8") as file:
        return [json.loads(line) for line in file if line.strip()]

cases = read_jsonl("model-selection-cases.jsonl")
results = {row["id"]: row["output"] for row in read_jsonl("results.jsonl")}
correct = sum(results.get(row["id"], "").strip() == row["expected"] for row in cases)
missing = sum(row["id"] not in results for row in cases)
print({"correct": correct, "total": len(cases), "missing": missing})

This only checks label agreement. Run tests for code tasks; define a rubric and human review process for open-ended tasks. Keep repeat-run variation instead of selecting the best output.

Upgrade and recheck when the evidence calls for it

Choose a candidate that meets your acceptance criteria, then decide whether any quality improvement is worth the additional time and cost. Add an upgrade route only after identifying failures that another candidate consistently improves.

Recheck when the model version, prompt, tools, data distribution, or acceptance criteria change. A publication date alone cannot make a selection table remain valid.

Was this useful?

If this post helped, you can buy me a coffee ☕

The next post, in your reader

Get new posts via RSS in Feedly, Reeder, or your preferred feed reader.

Subscribe via RSS