Know which AI models you can run. Keep the evidence.

Every model checked on its licence and provenance before it's used, tested on your own work, and kept on a record you can verify yourself. Built in New Zealand for Australian and New Zealand organisations.

  • Runs inside your environment
  • No kill switch
  • Records you can verify yourself

Latest verdict

Qwen / Qwen3-32B

revision 9216db57

Admitted

No findings: every admission gate passed for this revision

  • Licence passed: Apache-2.0 licence file in the repository
  • Provenance passed: Traced to the original publisher
  • Files passed: Safetensors, digests recorded
  • Export screening not assessed: Not yet assessed
  • Admit

    Every model checked on licence, provenance and files before anyone can use it.

    How the gate decides
  • Compare

    Candidate models tested on your own tasks, with the cost of each completed task.

    How Compare works
  • Prove

    A tamper-evident record of every request and the rule that allowed it.

    How the record works

The number that decides it

A model with cheaper tokens that needs three attempts costs more than a dearer one that succeeds first time.

Compare reports cost per completed task on your workload, per environment. No public leaderboard publishes it, and it answers the question most teams start with: at your volume, is running it privately cheaper than paying per token?

Read how a run is scored

Compare report two candidates, one workload

A Compare report line, illustrative
Measure Candidate A Candidate B
Task outcome pass or fail pass or fail
Cost per completed task reported reported
Time to first token measured measured
Tokens per task measured measured

Straight answers

Govern whatever you run, wherever it runs. Everyone promises private AI; we let you prove it.

  • Cost. Private AI can be dramatically cheaper at sustained utilisation when the model, workload and hardware line up. Compare shows whether yours do.
  • Compliance. We supply the evidence your auditors and risk team ask for, built into every request. We do not deliver compliance; we supply the evidence for yours.
  • Model choice. For the hardest reasoning, frontier models still lead; the pilot measures the difference on your workload.
  • Data. Nothing leaves by default: the egress allow-list is one identity-provider entry plus one entry for each provider you enable, and you generate it.
  • Provenance. Weights we byte-checked are labelled verified; a hosted route, in build, is labelled provenance by host attestation, not verified.

Is this for you?

Start with one workload

A free discovery session covers your highest-value AI workload, what data cannot leave your environment, and what a pilot would need to prove.

You leave with a clear view of whether a pilot is worth doing, and if it is, exactly what it would cover and what hardware you would need.