Measurements we run ourselves, on a schedule, so the same question can be asked again next quarter and the numbers can be compared. Every edition ships with its prompt set or instrument, its scoring rule, its limitations, and the raw data as CSV and JSON. If the method is wrong, the data is there to prove it.
Which vendors do AI assistants actually name — and how stable is the answer?
When a buyer asks a general-purpose AI assistant which product to use, which names come back, how often, and do they hold still when the same question is asked again?
Datasets are published under CC BY 4.0. Use them, quote them, re-run them — cite the edition and its version so readers can tell which snapshot a number came from.