nbgateLog in

Available models.

The multiplier is how heavily a model counts toward your token quota, not a price. A 2× model spends two tokens of quota per token used.

Latency, throughput, and success rate are measured through the live gateway: 3 streaming 24-token completions per model on 2026-09-26. Click a column header to sort; real results vary with prompt size and load.

Showing 1–10 of 37

ModelStatusLatencyThroughputSuccessMultiplierVision
DeepSeekGreg/deepseek-v4-flashOperational1019 ms260.3 tok/s100%2.3×No
DeepSeekGreg/deepseek-v4-flash-0731Operational1005 ms173.4 tok/s100%2.5×No
DeepSeekGreg/deepseek-v4-flash-vision-expOperational1252 ms—100%3×Yes
DeepSeekGreg/deepseek-v4-modUnavailable1293 ms—100%3.5×No
DeepSeekGreg/deepseek-v4-proUnavailable1880 ms657.9 tok/s100%2×No
DeepSeekGreg/deepseek-v4-pro-0813Operational1694 ms465.1 tok/s100%2.5×No
DeepSeekGreg/deepseek-v4.1-flashOperational1054 ms163.8 tok/s100%3×Yes
DeepSeekGreg/deepseek-v4.1-modUnavailable1303 ms—100%4×Yes
ZhipuGreg/glm-5.1Unavailable——0%2×No
ZhipuGreg/glm-5.2Unavailable——0%2×No