Cost per completed task, not cost per token

Leaderboards measure somebody else's tasks. Compare measures yours, in your environment, and reports the number that decides it.

A cheap model that needs three attempts costs more than an expensive one that succeeds first time. Only one of those numbers appears on a leaderboard.

Cost per token is an input. Cost per completed task is the output you are actually buying, and it is the one figure that changes a decision in a steering meeting.

The arithmetic illustrative shape, not a result

Two candidates compared by cost per completed task
Measure Candidate A Candidate B
Cost per million tokens lower higher
Attempts per completed task more fewer
Cost per completed task higher lower

What a run measures

The same tasks, one candidate at a time, scored by an executable check. One run can put a local model beside two hosted ones on the same provider; hosted provider and frontier API routes are in build.

The three-row cross-environment report: a small model on your own hardware beside two hosted open models on one named provider, on one task set. Both hosted rows are real calls to that provider through the enforcement hook, each under the provider's own route name; neither served byte-set was checked by us, so both carry the label provenance by host attestation, not verified. Each row carries its environment, its provenance label and its cost basis with the date the price was captured, and the local row carries a utilisation-sensitivity table.
Three rows, one task set, one report.
The same cross-environment report at 390px.

One task set on one laptop. The two hosted rows are real calls to the named provider through the hook, but what that provider serves is not byte-checked, and the local row's latency is this laptop's, not a hardware claim. Provider routing is in build.

Metrics captured per task, per model
Metric How it is produced
MetricTask pass rate How it is producedA test suite, build, lint or type check run in a sandbox. Pass or fail, not a score out of ten.
MetricTime to first token How it is producedMeasured on the stream, because that is what a person waits for.
MetricLatency percentiles How it is producedPer model across the run, so the tail is visible rather than averaged away.
MetricTokens per task How it is producedFrom the gateway's own ledger rows, not from an estimate.
MetricMemory high-water How it is producedPeak accelerator memory: whether the model fits at all.
MetricCost per completed task How it is producedTotal cost of the run divided by the tasks that actually passed.
MetricEnvironment How it is producedYour hardware, your private cloud, a hosted provider or a frontier API; the hosted and frontier routes are in build.
MetricCost basis How it is producedStated on every row: hardware amortisation and utilisation for a local row, the captured token price and date for a hosted one. Never presented as equivalent.
MetricProvenance label How it is producedVerified for weights we byte-checked; provenance by host attestation, not verified for a hosted route, and hosted routes are in build. A row missing either label is not rendered.

A human review sample covers the subjective remainder, recorded as a sample rather than a measurement.

A local row also shows cost per completed task across utilisation levels, so you can see the case for hosted first and hardware later.

The dataset stays yours

The golden dataset, tasks with an executable expected outcome, lives in your environment. With no external route enabled, no prompt leaves; with one, it crosses only the route your policy names.

Results are versioned, so the next model is measured against the same bar.

Reading the results

  • Results rank the candidates you admitted, on the tasks you ran.
  • Re-run when the workload, prompts or settings change.
  • Pilot figures describe the evaluation machine; production sizing is part of the pilot's recommendation.

Deciding which model costs less on your own tasks? Heliast runs Compare inside your environment, on your policy.

Book a discovery session