GenUI Model Benchmarks
Speed is only one dimension of a good experience. How well a model builds content and catalog items that address the user's needs is crucial, and these benchmarks cannot measure that.
Average over all timed round trips, sorted fastest-first by total round-trip time. Latency stats exclude failed turns; the error rate counts them.
| Model | Avg round-trip | Avg TTFC † | p95 TTFC † | Avg per pass | Error rate | Avg chunks † | Avg output tokens | Turns × iters |
|---|---|---|---|---|---|---|---|---|
| gemini-3.1-flash-lite-no-thinking | 1,866 ms | 673 ms | 792 ms | 5,599 ms | 0.0% | 20.7 | 482.0 | 3 × 5 |
| gemini-3.1-flash-lite | 1,915 ms | 699 ms | 896 ms | 5,744 ms | 0.0% | 20.7 | 479.6 | 3 × 5 |
| gpt-5.4-mini | 3,089 ms | 808 ms | 993 ms | 9,268 ms | 0.0% | 364.7 | 367.7 | 3 × 5 |
| gemini-2.5-flash-no-thinking | 3,267 ms | 914 ms | 1,297 ms | 9,801 ms | 0.0% | 14.2 | 584.9 | 3 × 5 |
| gpt-5.4-mini-no-thinking | 3,370 ms | 1,041 ms | 1,551 ms | 10,111 ms | 0.0% | 444.5 | 472.1 | 3 × 5 |
| gemini-2.5-flash | 3,462 ms | 1,623 ms | 2,134 ms | 10,386 ms | 60.0% | 12.0 | 534.8 | 3 × 5 |
| gemini-3-flash-preview-no-thinking | 3,492 ms | 1,198 ms | 1,668 ms | 10,477 ms | 0.0% | 24.4 | 565.1 | 3 × 5 |
| gpt-5.4 | 3,678 ms | 1,008 ms | 1,766 ms | 11,033 ms | 0.0% | 486.7 | 489.7 | 3 × 5 |
| gpt-5.6-luna | 3,721 ms | 1,348 ms | 2,142 ms | 11,163 ms | 0.0% | 528.4 | 585.2 | 3 × 5 |
| gpt-5.6-luna-no-thinking | 3,835 ms | 1,668 ms | 4,148 ms | 11,506 ms | 0.0% | 544.5 | 584.9 | 3 × 5 |
| kimi-k2.7-code-highspeed | 4,090 ms | 2,452 ms | 4,053 ms | 12,269 ms | 0.0% | 486.9 | 610.4 | 3 × 5 |
| gpt-5.4-no-thinking | 4,407 ms | 1,338 ms | 2,435 ms | 13,222 ms | 0.0% | 505.0 | 541.7 | 3 × 5 |
| gpt-5.6-terra-no-thinking | 4,447 ms | 1,864 ms | 5,265 ms | 13,341 ms | 0.0% | 493.9 | 534.8 | 3 × 5 |
| gpt-5.6-terra | 4,499 ms | 1,663 ms | 2,411 ms | 13,497 ms | 0.0% | 547.4 | 617.7 | 3 × 5 |
| gemini-3.5-flash-no-thinking | 4,526 ms | 1,752 ms | 2,498 ms | 13,578 ms | 0.0% | 24.9 | 584.5 | 3 × 5 |
| deepseek-v4-flash-no-thinking | 5,261 ms | 1,280 ms | 1,599 ms | 15,783 ms | 0.0% | 556.0 | 556.7 | 3 × 5 |
| claude-haiku-4-5 | 5,279 ms | 2,423 ms | 4,734 ms | 15,837 ms | 0.0% | 14.1 | 518.5 | 3 × 5 |
| claude-sonnet-5 | 5,665 ms | 1,816 ms | 2,521 ms | 16,995 ms | 0.0% | 11.9 | 561.9 | 3 × 5 |
| deepseek-v4-flash | 5,971 ms | 2,436 ms | 3,582 ms | 17,913 ms | 0.0% | 526.9 | 612.1 | 3 × 5 |
| gemini-3-flash-preview | 6,322 ms | 4,120 ms | 7,264 ms | 18,967 ms | 0.0% | 23.9 | 568.3 | 3 × 5 |
| gpt-5.5-no-thinking | 6,396 ms | 1,800 ms | 4,294 ms | 17,907 ms | 6.7% | 493.8 | 501.2 | 3 × 5 |
| gemini-3.5-flash | 7,788 ms | 5,192 ms | 7,006 ms | 23,365 ms | 0.0% | 22.7 | 538.0 | 3 × 5 |
| claude-opus-4-8 | 7,818 ms | 1,752 ms | 3,936 ms | 23,453 ms | 0.0% | 13.5 | 616.3 | 3 × 5 |
| gpt-5.5 | 7,976 ms | 2,998 ms | 6,424 ms | 23,929 ms | 0.0% | 519.7 | 608.5 | 3 × 5 |
| claude-sonnet-4-6 | 8,193 ms | 1,487 ms | 2,128 ms | 24,578 ms | 0.0% | 16.5 | 553.4 | 3 × 5 |
| mercury-2 | 8,375 ms | 7,895 ms | 34,970 ms | 25,124 ms | 0.0% | 5.9 | 545.2 | 3 × 5 |
| kimi-k2.7-code | 17,056 ms | 6,305 ms | 15,354 ms | 51,167 ms | 0.0% | 485.5 | 666.5 | 3 × 5 |
Column glossary (hover any header for a tooltip)
- Model — the model id. A -no-thinking row is the same model with thinking disabled. Rows are sorted fastest-first.
- Avg round-trip — mean wall-clock time per turn, from request sent to stream complete, over successful turns. Comparable across providers; this is the headline metric.
- Avg TTFC † — time to first chunk: request sent until the first non-empty streamed chunk.
- p95 TTFC † — 95th-percentile time to first chunk, i.e. tail latency.
- Avg per pass — mean time to complete one full pass (all turns in an iteration summed). Comparable across providers.
- Error rate — share of turns that errored, produced no valid GenUI surface, or whose components didn't conform to the catalog item JSON schemas (props or component type). Hover the value for the breakdown.
- Avg chunks † — mean number of non-empty streamed chunks per turn.
- Avg output tokens — mean completion tokens per turn (excludes the prompt). Includes reasoning tokens for thinking models, so it drops on -no-thinking rows.
- Turns × iters — turns per iteration × iteration count behind each average.
† TTFC and chunk counts reflect each provider's streaming granularity — OpenAI streams token-by-token, Google and others send fewer, larger chunks — so compare them only within a provider, not across. Latency and token stats are shown only when the provider reports them.