Compare
Compare models
Plot cost against writing quality, speed or context for up to six models, then share the exact view as a link.
Models · 4/6
01Chart
Filters and reference
02Side by side
| Model | 1 request | 100 requests | vs. reference | Writing Elo | First token | Speed | Context |
|---|---|---|---|---|---|---|---|
Best marks the strongest value in each column. Estimates exclude top-up fees and cache writes. Speed and first-token figures are external benchmarks.
Sources and calculation details
100 requests repeats the same workload; it does not predict a growing conversation. Changing token counts updates cost only. The cost reference stays fixed when filters change.
Writing Elo assesses story generation. EQ-Bench 4 assesses emotional and social intelligence. Both are useful signals for roleplay, with different test designs. Neither measures every aspect of character chat. First-token latency is not time to a visible answer for reasoning models.
Ranking methodology