Available models.
The multiplier is how heavily a model counts toward your token quota, not a price. A 2× model spends two tokens of quota per token used.
Latency, throughput, and success rate are measured through the live gateway: 3 streaming 24-token completions per model on 2026-09-26. Click a column header to sort; real results vary with prompt size and load.
Showing 1–10 of 37
| Model | Status | Latency | Throughput | Success | Multiplier | Vision |
|---|---|---|---|---|---|---|
| Greg/deepseek-v4-flash | Operational | 1019 ms | 260.3 tok/s | 100% | 2.3× | No |
| Greg/deepseek-v4-flash-0731 | Operational | 1005 ms | 173.4 tok/s | 100% | 2.5× | No |
| Greg/deepseek-v4-flash-vision-exp | Operational | 1252 ms | — | 100% | 3× | Yes |
| Greg/deepseek-v4-mod | Unavailable | 1293 ms | — | 100% | 3.5× | No |
| Greg/deepseek-v4-pro | Unavailable | 1880 ms | 657.9 tok/s | 100% | 2× | No |
| Greg/deepseek-v4-pro-0813 | Operational | 1694 ms | 465.1 tok/s | 100% | 2.5× | No |
| Greg/deepseek-v4.1-flash | Operational | 1054 ms | 163.8 tok/s | 100% | 3× | Yes |
| Greg/deepseek-v4.1-mod | Unavailable | 1303 ms | — | 100% | 4× | Yes |
| Greg/glm-5.1 | Unavailable | — | — | 0% | 2× | No |
| Greg/glm-5.2 | Unavailable | — | — | 0% | 2× | No |