A banking application programming interface can receive thousands of legitimate requests and abusive requests through the same technical doorway. Rate limiting controls how quickly a defined user, application or service may consume that doorway's resources.
Every API request consumes something
A balance request may use computing, database and network capacity, while a password-recovery request can also trigger an outside message and cost. If clients can repeat an expensive operation without an appropriate bound, one faulty integration or attacker can degrade service for everyone else.
Rate limiting protects availability and cost by restricting request frequency over a defined period. It complements capacity planning; it does not excuse an architecture that cannot support the legitimate transaction volume and peak demand the bank has agreed to serve.
The limit follows the operation and identity
Teams decide what they are limiting: an authenticated customer, device, software client, access token, network address or combination of signals. They also distinguish low-cost read requests from sensitive or resource-intensive actions such as login attempts, statement generation, payment initiation or one-time-code delivery.
A single global ceiling is rarely sufficient. Limits may apply per second and over longer windows, and high-impact operations can have tighter transaction or value controls. Authentication and authorization still decide who may act; rate limits decide how often an allowed request may be attempted.
The service needs a controlled response
When a client reaches a limit, the API rejects or delays additional requests in a predictable way and may indicate when the client can try again. Good client design uses controlled retries and avoids sending the same instruction in a rapid loop that makes the original problem worse.
Payment and account-changing APIs also use unique request identifiers or idempotency controls so a retry does not create an unintended duplicate. Queues can smooth bursts where delay is acceptable, but time-sensitive requests need clear expiration and customer-status handling.
Resource protection needs several layers
Request counts alone do not control oversized files, unusually complex queries, long-running calls or a batch that contains hundreds of operations. Banks pair rate limits with payload and batch limits, timeouts, spending thresholds, resource isolation and validation of request parameters.
Controls also need to account for distributed attacks and shared infrastructure, where many sources can stay below individual thresholds while overwhelming a dependency. Upstream gateways, application services, databases and third parties should have compatible protections rather than relying on one perimeter control.
Monitoring keeps protection aligned with customer use
Teams monitor rejection rates, latency, resource consumption, client behavior and customer impact. Testing covers ordinary peaks, failover traffic, major payment dates and a provider slowdown so limits do not unexpectedly block a valid critical service when capacity is already strained.
A limit set too high may not prevent exhaustion, while one set too low can deny customers or partners legitimate access. Changes therefore move through testing and approval, with documented emergency adjustments and post-event review rather than permanent, untracked bypasses.
Read the primary material
Banking Explained prioritizes regulators, official publications and first-party announcements.
