The Voice Agent Benchmark Landscape is a dataset we publish: a structured map of public benchmarks for voice agents, spoken assistants, speech-enabled tool use, computer action, speech recognition and meeting understanding. This page is its record of publication.
What it is
- 13 benchmarks, one row each, in a machine-readable file (data/benchmarks.jsonl) with a documented field contract (data/schema.json).
- Each row records the benchmark's primary focus; whether it covers spoken input, spoken output, multi-turn interaction, tool use, goal completion, computer or browser action, meeting or long-form material, and real-time operation; whether the data is public; the code and data licences, recorded separately; the repository, dataset card and paper addresses; the date last verified; and evidence notes.
- Values are yes, no, partial or unclear. Partial means the capability appears in part of the suite or is evaluated indirectly; it does not describe quality.
- Every row has at least one first-party source, listed in the source audit. Facts in the current release were verified on 1 September 2026 from the maintainers' public repositories, dataset cards and papers.
What it is not
- Not a leaderboard. It contains no scores and ranks nothing. Inclusion does not imply endorsement.
- Not an evaluation framework. There is nothing to run against a model; the repository ships the data, its schema, a validator and a query tool.
- Not Cue's benchmark. The private 227-sample production benchmark described in Google DeepMind's case study is not part of it.
Where it is published
| Where | Address | Role |
|---|---|---|
| GitHub | Sophon-LLC/voice-agent-benchmark-landscape | Source of record: data, schema, source audit, contribution rules, tests |
| Zenodo | doi.org/10.5281/zenodo.22227265 | Archived releases with DOIs; the concept DOI always resolves to the latest version |
| Hugging Face | chatjesus/voice-agent-benchmark-landscape | Mirror of the dataset |
Versions
| Version | Date | Records | Licence | Version DOI |
|---|---|---|---|---|
| v1.0.2 | 1 September 2026 | 13 | MIT | 10.5281/zenodo.22227266 |
The Zenodo record for v1.0.2 carries the dataset's SHA-256 (1d36c7b3f73565147fb7afff076c611c5fddc331e3ec47dc867625b75b70bbf4) so a downloaded copy can be checked against the archived one.
How to cite
The repository's CITATION.cff names the author as Sophon LLC. Cite the version DOI when exact reproducibility matters and the concept DOI for the evolving project:
Sophon LLC. Voice Agent Benchmark Landscape, version 1.0.2. 1 September 2026. https://doi.org/10.5281/zenodo.22227266
Corrections
A correction names the row, the field, the proposed value and a first-party source, following the repository's CONTRIBUTING.md. Changes ship as a new version with a new DOI.