What is an API rate limit — the guard against floods like the one that hit Wikidata
An API rate limit is a cap on how many requests one user or program can send to a server within a set time. Exceed it and the server returns HTTP 429 (Too Many Requests), usually with a Retry-After header telling the client when to try again. The goal is to stop any single user from hogging capacity or an automated program from knocking a service over, so everyone can keep using it. Wikimedia began phasing in rate limits on its public APIs in March 2026, and on October 5 said AI agents it believes OpenAI operated sent millions of API requests and hundreds of thousands of Wikidata Query Service queries that may have contributed to an outage in May
The three lines
- Definition — a cap on requests per time window; over it, HTTP 429 plus a Retry-After hint
- Methods — token bucket, fixed window, sliding window; counted per user, IP or API key
- Case — Wikimedia phased in limits from March 2026; identified bots get more room; AI agents test it
Key questions
- What is API rate limiting
- **A cap on how many requests can be sent in a time window.** | Item | Detail | |---|---| | Units | Requests per second, minute, hour or day | | Counted by | User account, API key, IP address | | When exceeded | HTTP 429 Too Many Requests | | Hint | Retry-After header (seconds to wait) | | Purpose | Prevent overload, share fairly, slow abuse |
- How to fix HTTP 429 Too Many Requests
- **Slow down, wait, and retry.** | Fix | Detail | |---|---| | Honor Retry-After | Wait the time the server specifies | | Exponential backoff | Retry after 1s, 2s, 4s… with random jitter | | Batch requests | Use endpoints that return many items per call | | Authenticate | Logged-in or keyed requests usually get higher limits | | Identify yourself | Descriptive User-Agent with contact info (Wikimedia requires it) |
- Rate limiting algorithms
- **Four common designs.** | Method | How it works | Trait | |---|---|---| | Fixed window | Counter resets every minute | Simple; bursts at boundaries | | Sliding window | Always counts the last 60 seconds | Smooths boundary bursts | | Token bucket | Tokens refill at a steady rate; each request spends one | Allows short bursts, caps the average | | Leaky bucket | Queue drains at a fixed rate | Even server load |
If one person plugs a pump into the village well, everyone goes thirsty. A rate limit sets the width of the pipe. On October 5, 2026, the Wikimedia Foundation said AI agents it believes were operated by OpenAI sent millions of requests to its public APIs and hundreds of thousands of queries to the Wikidata Query Service, traffic that may have contributed to an outage in May (see "Wikimedia says OpenAI agents edited its wikis"). As programs rather than people become the web's heaviest users, counting and capping requests has become a survival tool for services.
1. What a rate limit is
An API (application programming interface) is the window through which one program requests data from another service. A rate limit puts a "this many per minute" cap on that window.
| Item | Detail |
|---|---|
| What is counted | Requests (sometimes bytes; AI APIs also count tokens) |
| Per whom | Account, API key, IP address, app (User-Agent) |
| Time window | Second, minute, hour, day |
| When exceeded | HTTP 429 Too Many Requests (standardized in RFC 6585, 2012) |
| Accompanying hints | Retry-After, remaining-quota headers |
It serves three purposes: preventing overload, sharing capacity fairly, and slowing abuse such as mass scraping or password guessing.
2. How counting works: four common algorithms
| Method | Principle | Strength | Weakness |
|---|---|---|---|
| Fixed window | Reset the counter every minute on the minute | Simple | Requests at 0:59 and 1:01 let double through in two seconds |
| Sliding window | Always count the previous 60 seconds | Fixes boundary bursts | Stores more history |
| Token bucket | Tokens refill at a steady rate; each request spends one | Allows brief bursts, caps the average | Bucket size matters |
| Leaky bucket | Queue requests and release at a fixed rate | Smooth server load | Adds waiting time |
Take a token bucket holding up to 100 tokens, refilled at 10 per second. A program that has been idle can fire 100 requests at once, but after that it can't exceed 10 per second. Bursty human use gets through; a machine that never stops gets throttled.
3. The Wikimedia case: limits, identity and AI agents
| Client type | How Wikimedia treats it |
|---|---|
| Anonymous, no User-Agent | May be blocked without notice |
| Identified (descriptive User-Agent) | Rate-limited since Phase 2 (late April 2026) |
| Authenticated (OAuth 2.0) | Higher limits |
| Approved bots on Toolforge / Cloud Services | Exempt |
Phase 1 enforcement covered all public APIs from March 2026. Wikimedia's principle is "identify yourself and you get more room." Its access policy requires a User-Agent with app name, version and contact details on every request and warns that clients without one may be blocked without notice. Authenticated (OAuth 2.0) requests get higher limits, and approved bots running in Wikimedia's cloud are exempt. When an education dashboard hit 429 errors in 2026, staff traced it to User-Agent noncompliance and said proper identification should fix it.
AI agents probe the weak spot: requests that look human and are spread across many routes rarely trip a per-user cap. That is why Wikimedia asked AI companies to tag their traffic. A rate limit only works if you can tell who is sending the requests.
4. If you hit a 429: what developers should do
| Situation | Response |
|---|---|
Retry-After present | Wait exactly that long |
| No hint | Exponential backoff: 1s → 2s → 4s → 8s, plus random jitter |
| Repeating the same lookups | Cache results |
| Many single-item calls | Use batch endpoints |
| Limits always too tight | Authenticate, upgrade a plan, or use official data dumps (Wikimedia publishes full dumps) |
| Wikimedia APIs | Descriptive User-Agent; use the maxlag parameter to back off when servers lag |
AI APIs use the same idea: OpenAI and Anthropic cap both requests and tokens per minute, raising limits as usage tiers rise.
5. FAQ
| Question | Answer |
|---|---|
| Is rate limiting DDoS protection? | Related but different; DDoS comes from many addresses, so per-IP caps alone won't stop it |
| 429 vs 503? | 429: "you sent too many." 503: "the server can't take requests right now" |
| Rotating IPs to dodge limits? | Usually violates terms of service and invites blocking and public criticism |
| Do websites rate-limit page visits? | Yes, web servers and CDNs do, to manage crawlers and bots |
6. What remains unclear
- Wikimedia's exact limits: numeric thresholds were not found.
- Agents versus the limits: whether the agents' traffic tripped the limits, or slipped past, wasn't disclosed.
- Related: how sites signal crawlers to stay out is in "What robots.txt is"; how agents get account access without passwords is in "What OAuth is."