Skip to main content

Limits and Quotas

To ensure ultra-low latency (sub-500ms voice turn-around) and protect against infrastructure denial-of-service, Sayvy AI applies the following payload, concurrency, and operational limits.

Concurrency Quotas

Simultaneous call limits are evaluated across your entire workspace in real-time. If a dispatch exceeds your concurrency cap, the API returns a 429 concurrency_limit_reached error.

Payload Size Limits


Operational Timeouts


Requesting Limit Increases

Enterprise customers can request increased concurrency quotas, custom TTS model instances, and private SIP trunking:
  1. Visit the Billing & Plan tab in your Sayvy AI Console.
  2. Click Request Quota Increase or reach out directly to your account manager at support@sayvy.ai.