AI Token Rate Limits and Quotas
Enforces token-per-minute limits and token quotas by API consumer at the API Management runtime boundary so one application, team, or workload cannot exhaust shared model capacity.
Enforces token-per-minute limits and token quotas by API consumer at the API Management runtime boundary so one application, team, or workload cannot exhaust shared model capacity.