All Insights

Product and engineering note · reviewed 10 September 2026

Wave evaluation: what the evidence must show

A qualified research note on evaluating privacy-preserving context without claiming a universal benchmark.

Current evidence status

Earlier material reported a small quality change in an internal evaluation. The public site does not provide the complete task set, sample size, model versions, scoring rubric or reproducible run artefacts. We are not presenting that result as an independently verified benchmark or a general performance promise.

Compare the same task

A meaningful evaluation should compare the same task with original and transformed context, using the same model settings and a stated rubric. Report task failures and exceptions as well as aggregate scores. Assess disclosure risks separately: a useful answer does not by itself demonstrate privacy.

What to request

For an evaluation, ask for the input categories, provider/model versions, transformation scope, number of examples, human-review method and known failure cases. Quality and disclosure can vary by task and deployment. Choose an integration on the evidence for your intended workflow, not on a single cross-product percentage.