Skip to main content

Understand throughput metrics

  • RPM (requests per minute) The number of requests that can be processed per minute.
  • IPM (images per minute) The number of images that can be processed per minute.
  • TPM (tokens per minute) The number of tokens that can be processed per minute.

Handle rate-limit responses

Control concurrency before sending requests. If you receive 429, retry with exponential backoff and jitter, then reduce burst traffic if the responses continue. See Handle rate limits and concurrency for examples, and use the model catalog to select model IDs.
Last modified on October 9, 2026