AI Token Rate Limits and Quotas

Enforces token-per-minute limits and token quotas by API consumer at the API Management runtime boundary so one application, team, or workload cannot exhaust shared model capacity.