Speed after the first text
Output throughput describes generation after a response starts. It answers a different question from how long you waited for the first text. A response can start quickly and then arrive slowly; another can pause before producing an answer in a short burst. Keep both phases visible when evaluating responsiveness.
A delta is not a token
A stream message may contain several tokens or part of a string. Counting events and calling the result tokens per second creates a misleading comparison. Tokenization also differs across model families and languages. Character counts are not interchangeable with the token count used by a provider for billing.
Reasoning changes the interpretation
Billed output can include work that never appears as answer text. Dividing that count by the time between visible text events can inflate apparent reading speed. We retain billed and reasoning counts separately when supplied and withhold throughput when a reliable visible count is absent.
How to compare responsibly
Use a fixed task and output budget, document the count source and consider completion time alongside TTFT. One-chunk responses do not support a meaningful visible-generation interval. A zero-length interval must never become an infinite rate. No throughput ranking is published before its counting method is explained and tested.