01

Invisible work takes time

Some model configurations perform reasoning not emitted as ordinary answer text. A monitor may receive metadata or reasoning-related events before the first visible response. Counting those as the first readable token shortens the reported wait without improving what the reader actually sees.

02

Settings define the baseline

Reasoning configuration belongs in experiment identity. Low-effort and high-effort runs should not be pooled as identical tasks. If a provider changes a model alias or its supported settings, treat the transition as a revision boundary and re-evaluate whether old baselines remain comparable.

03

Billing differs from display

Usage records can describe billed tokens, reasoning tokens or other categories. These help accounting but should not automatically become visible-output counts. Store the source and method with every derived metric. If the relationship is unclear, withhold throughput instead of subtracting fields based on an assumption.

04

Interpret the whole task

A longer initial wait is not enough to judge answer quality; a faster stream is not enough to judge usefulness. This service describes measured responsiveness under controlled settings. It does not grade reasoning quality or claim every extra second is productive reasoning. Consider completion and reliability alongside your workflow needs.

05

Sources and further reading

Our methodology · More guides