Skip to content
TEN Brief Ten verified stories a day 2026.10.09 KO

이 기사는 한국어로도 읽을 수 있습니다 →

Tech · 3 min read · Explainer

What is an API rate limit — the guard against floods like the one that hit Wikidata

An API rate limit is a cap on how many requests one user or program can send to a server within a set time. Exceed it and the server returns HTTP 429 (Too Many Requests), usually with a Retry-After header telling the client when to try again. The goal is to stop any single user from hogging capacity or an automated program from knocking a service over, so everyone can keep using it. Wikimedia began phasing in rate limits on its public APIs in March 2026, and on October 5 said AI agents it believes OpenAI operated sent millions of API requests and hundreds of thousands of Wikidata Query Service queries that may have contributed to an outage in May

Cars passing one by one through highway toll gates on a sunny day

The three lines

  • Definition — a cap on requests per time window; over it, HTTP 429 plus a Retry-After hint
  • Methods — token bucket, fixed window, sliding window; counted per user, IP or API key
  • Case — Wikimedia phased in limits from March 2026; identified bots get more room; AI agents test it

Key questions

What is API rate limiting
**A cap on how many requests can be sent in a time window.** | Item | Detail | |---|---| | Units | Requests per second, minute, hour or day | | Counted by | User account, API key, IP address | | When exceeded | HTTP 429 Too Many Requests | | Hint | Retry-After header (seconds to wait) | | Purpose | Prevent overload, share fairly, slow abuse |
How to fix HTTP 429 Too Many Requests
**Slow down, wait, and retry.** | Fix | Detail | |---|---| | Honor Retry-After | Wait the time the server specifies | | Exponential backoff | Retry after 1s, 2s, 4s… with random jitter | | Batch requests | Use endpoints that return many items per call | | Authenticate | Logged-in or keyed requests usually get higher limits | | Identify yourself | Descriptive User-Agent with contact info (Wikimedia requires it) |
Rate limiting algorithms
**Four common designs.** | Method | How it works | Trait | |---|---|---| | Fixed window | Counter resets every minute | Simple; bursts at boundaries | | Sliding window | Always counts the last 60 seconds | Smooths boundary bursts | | Token bucket | Tokens refill at a steady rate; each request spends one | Allows short bursts, caps the average | | Leaky bucket | Queue drains at a fixed rate | Even server load |

If one person plugs a pump into the village well, everyone goes thirsty. A rate limit sets the width of the pipe. On October 5, 2026, the Wikimedia Foundation said AI agents it believes were operated by OpenAI sent millions of requests to its public APIs and hundreds of thousands of queries to the Wikidata Query Service, traffic that may have contributed to an outage in May (see "Wikimedia says OpenAI agents edited its wikis"). As programs rather than people become the web's heaviest users, counting and capping requests has become a survival tool for services.

1. What a rate limit is

An API (application programming interface) is the window through which one program requests data from another service. A rate limit puts a "this many per minute" cap on that window.

ItemDetail
What is countedRequests (sometimes bytes; AI APIs also count tokens)
Per whomAccount, API key, IP address, app (User-Agent)
Time windowSecond, minute, hour, day
When exceededHTTP 429 Too Many Requests (standardized in RFC 6585, 2012)
Accompanying hintsRetry-After, remaining-quota headers

It serves three purposes: preventing overload, sharing capacity fairly, and slowing abuse such as mass scraping or password guessing.

2. How counting works: four common algorithms

MethodPrincipleStrengthWeakness
Fixed windowReset the counter every minute on the minuteSimpleRequests at 0:59 and 1:01 let double through in two seconds
Sliding windowAlways count the previous 60 secondsFixes boundary burstsStores more history
Token bucketTokens refill at a steady rate; each request spends oneAllows brief bursts, caps the averageBucket size matters
Leaky bucketQueue requests and release at a fixed rateSmooth server loadAdds waiting time

Take a token bucket holding up to 100 tokens, refilled at 10 per second. A program that has been idle can fire 100 requests at once, but after that it can't exceed 10 per second. Bursty human use gets through; a machine that never stops gets throttled.

3. The Wikimedia case: limits, identity and AI agents

Client typeHow Wikimedia treats it
Anonymous, no User-AgentMay be blocked without notice
Identified (descriptive User-Agent)Rate-limited since Phase 2 (late April 2026)
Authenticated (OAuth 2.0)Higher limits
Approved bots on Toolforge / Cloud ServicesExempt

Phase 1 enforcement covered all public APIs from March 2026. Wikimedia's principle is "identify yourself and you get more room." Its access policy requires a User-Agent with app name, version and contact details on every request and warns that clients without one may be blocked without notice. Authenticated (OAuth 2.0) requests get higher limits, and approved bots running in Wikimedia's cloud are exempt. When an education dashboard hit 429 errors in 2026, staff traced it to User-Agent noncompliance and said proper identification should fix it.

AI agents probe the weak spot: requests that look human and are spread across many routes rarely trip a per-user cap. That is why Wikimedia asked AI companies to tag their traffic. A rate limit only works if you can tell who is sending the requests.

4. If you hit a 429: what developers should do

SituationResponse
Retry-After presentWait exactly that long
No hintExponential backoff: 1s → 2s → 4s → 8s, plus random jitter
Repeating the same lookupsCache results
Many single-item callsUse batch endpoints
Limits always too tightAuthenticate, upgrade a plan, or use official data dumps (Wikimedia publishes full dumps)
Wikimedia APIsDescriptive User-Agent; use the maxlag parameter to back off when servers lag

AI APIs use the same idea: OpenAI and Anthropic cap both requests and tokens per minute, raising limits as usage tiers rise.

5. FAQ

QuestionAnswer
Is rate limiting DDoS protection?Related but different; DDoS comes from many addresses, so per-IP caps alone won't stop it
429 vs 503?429: "you sent too many." 503: "the server can't take requests right now"
Rotating IPs to dodge limits?Usually violates terms of service and invites blocking and public criticism
Do websites rate-limit page visits?Yes, web servers and CDNs do, to manage crawlers and bots

6. What remains unclear

  • Wikimedia's exact limits: numeric thresholds were not found.
  • Agents versus the limits: whether the agents' traffic tripped the limits, or slipped past, wasn't disclosed.
  • Related: how sites signal crawlers to stay out is in "What robots.txt is"; how agents get account access without passwords is in "What OAuth is."

Sources

  1. IETF — RFC 6585: Additional HTTP Status Codes (429 Too Many Requests)
  2. MDN — 429 Too Many Requests
  3. MediaWiki — Wikimedia APIs/Changelog
  4. MediaWiki — Wikimedia APIs/Access policy
  5. The Register — Wikimedia Foundation comes forward as latest OpenAI agent assault victim

Verification

Published
Last modified
Cross-check
Checked against 5 independent sources.
Unverified
  • Wikimedia's specific numeric limits (requests per minute or hour) were not found in public material.
  • Whether the agents' requests hit Wikimedia's rate limits, or how they avoided them, was not disclosed.
  • Algorithm descriptions and numbers in examples are illustrative, not any service's real settings.
Authoring
Reviewed by a person before publication. The full process is described in the Editorial.

Ten stories, once each morning

We send the three-line summaries only; the full pieces stay on the site. One-click unsubscribe, any time.

Related