‹ API Foundations Lesson 9 of 16
Contents Lesson 9 of 16

5 min read · foundations

Why does one script die at noon and another die in ten seconds?

Because there are two completely different limits, enforced by different mechanisms, with different symptoms and different fixes. Confusing them means applying the wrong fix, which is worse than applying none — you will change something, the problem will persist, and you will conclude the API is broken.

Wall one: the daily quota

How many calls you may make in a day. A counter that goes up as you spend and resets on a date boundary.

The published figure is 100,000 calls per day on paid plans, and 20 calls per day on the Free tier. You can read your own position from GET /user, which on a test account on 2026-07-28 returned apiRequests: 1632 against dailyRateLimit: 100000.

Symptom: everything works normally, then at some point everything stops, and it stays stopped. No amount of waiting a minute helps. Tomorrow it works again.

Fix: make fewer calls. Bulk endpoints instead of loops, date filters instead of full histories you throw away, and a cache so you never fetch the same closed trading day twice.

Wall two: the per-minute rate limit

How fast you may make those calls. A speed cap, measured over a rolling minute.

The published figure is 1,000 requests per minute, and the documentation shows X-RateLimit-Limit: 1000 in its own example. Read the header rather than the docs, though: live responses currently carry X-RateLimit-Limit: 1200. Pacing to the published number stays safe; hard-coding either number is how you find out the hard way that it moved.

The headers arrive on every response, not only on a throttled one — the documentation says so explicitly and the server agrees. So X-RateLimit-Remaining is a gauge you can watch continuously, for free, alongside Retry-After in seconds on the throttled case.

Symptom: a burst of 429 Too Many Requests, which clears by itself within a minute. Your daily quota is nowhere near spent.

Fix: slow the burst down. A concurrency cap, a queue, a small sleep between batches. Nothing about your total volume needs to change.

Say "per minute", and never "per second"

Nothing in this stack publishes a per-second limit — the budget you are given is per minute, and you should plan against that. If you see a per-second figure quoted anywhere, it is an average somebody derived by dividing an internal per-minute setting by sixty — and the fix that framing suggests (a fixed delay between every call) is not the fix an actual per-minute budget needs. A per-minute budget lets you burst and then rest; a per-second framing forbids the burst and hides the rest.

This is not a pedantic distinction. It changes what you build.

Both limits belong to the token, not to the machine sending the request. Two workers that each pace themselves to the published minute ceiling are together sending twice it, and the 429s land on whichever request arrives second. Run one pacer per token: a shared counter, a queue, or a single process that owns the key. The daily quota works the same way; every script, notebook and colleague using one token draws on one apiRequests count, which is why /user is worth polling from a job that spends nothing itself. The counter resets at midnight GMT (API limits page, read 2026-09-07); write that boundary into the scheduler rather than assuming your own time zone.

The arithmetic that ties the two walls together

One unit conversion first, because it is the whole trick. The documentation's own heading reads "Minute Limit (requests) — Minute request limit (not API calls)", and the header confirms it: X-RateLimit-Remaining falls by exactly one per request, whatever that request cost. Wall one counts calls; wall two counts requests. They are different currencies and the next lesson but one prices them.

Run flat out at the per-minute ceiling on a 1-call endpoint and you spend a full daily quota in:

100,000 ÷ 1,000 = 100 minutes.

An hour and forty minutes, and your day is over. But that is the best case, because it assumes every request costs one call. Run the same thousand requests a minute against /fundamentals at 10 calls each and you are spending 10,000 calls a minute:

100,000 ÷ 10,000 = 10 minutes. On /eod-bulk-last-day at 100 calls a request, 60 seconds.

So the two limits are not two versions of the same idea, and the gap between them is up to two orders of magnitude wider than the headline sum suggests — the rate limit is fast enough to destroy your quota before you have finished reading this paragraph.

Turn it around. Spread 100,000 calls evenly across 24 hours:

100,000 ÷ 1,440 minutes = about 69 calls per minute.

Sixty-nine, against a ceiling of a thousand. A job that paces itself never meets the rate limit at all. Which is the practical conclusion of the whole lesson: the per-minute wall is almost always a symptom of a burst you did not need to make.

Try it now

  1. /user answers about the key that calls it, so it cannot be rendered here as a live table. Below is the shape it takes for an account on a 100,000-call plan with 50,000 calls of purchased extra capacity, trimmed to the quota fields; the figures are illustrative. Read apiRequests against dailyRateLimit and say what share of the day's base quota had gone. Then do the same for the test account quoted above.
{"apiRequests": 41250,
 "apiRequestsDate": "2026-09-28",
 "dailyRateLimit": 100000,
 "extraLimit": 50000}
  1. Divide a daily allowance by 1,440: the 100,000 of the plan, then the 150,000 it becomes with the extra capacity in the response above. That is the even pace in calls per minute — compare it to 1,000 and see how much headroom a well-behaved job has.
  2. Take your slowest scheduled job and count the calls it makes. If the answer is more than a few thousand, the next two lessons are about why.