‹ EODHD: The Product Lesson 4 of 17
Contents Lesson 4 of 17

3 min read · practitioner

Where our data actually comes from

Why this matters. "Is your data reliable?" is asked of every source by every serious client, and the honest answer is a description of a process, not a slogan. Foundation: How do professionals sanity-check data?

The supply chain

Data arrives from more than fifty sources: exchange feeds we license directly, established aggregators, and specialists for particular corners of the market. It is reconciled centrally before it reaches an endpoint, which is why one ticker looks the same whichever way you ask for it.

Direct licensing matters and is worth stating precisely. Some feeds come straight from the exchange, and some come through an aggregator or a specialist vendor. "We have direct exchange contracts" is true; "all our data is direct from exchanges" is not, and the difference is exactly the kind of thing a technical client checks.

The watch

Reconciled feeds still produce nonsense: an impossible price jump, a missing bar, a feed that quietly drifts. An automated anomaly watch runs continuously over the data looking for exactly those shapes, and what it flags enters a documented data-quality process: triage by how much is affected, severity-tiered targets, a trace back through the pipeline to the cause, a fix, and a backfill of the history behind it.

What the claim actually is

Read that process for what it promises. It commits to finding and fixing, on a measured clock. It does not promise the data is never wrong, and no honest data vendor promises that. Fifty sources reconciled continuously will produce anomalies. The claim is vigilance, and vigilance that is measured.

That is a stronger position than perfection, because it survives contact with the first bad number a client finds.

Three things are not defects, and ruling them out comes before a report. First, convention: close is the traded price and adjusted_close folds in splits and dividends, so a mismatch against another source is usually two conventions. Second, rewriting: the adjustment is computed backward from the latest row, so every new dividend or split rescales every earlier adjusted value, and a client who cached rows years ago and finds them changed is seeing the series work as documented; check /div/{ticker} and /splits/{ticker} for an action after their cache date. Third, timing: the row may not be due yet, as the prices lesson explains. What survives those three checks is a report.

What a useful report looks like

Triage runs on three fields: a ticker, a date, and the expected value beside the actual one. A report missing any of them cannot enter the process at all. Knowing that turns a vague "the prices look wrong" into something that gets fixed.

Try it now

  1. Answer "is your data reliable?" in two sentences, naming the sources, the watch, and the measured process, without ever claiming the data cannot be wrong.
  2. Now file a report that could actually enter triage. A single day of one ticker is below: /eod/AAPL.US?from=2020-08-28&to=2020-08-28&fmt=json. That single row gives you two of the three required fields — the ticker and the date — and the close is the actual value. Write the third yourself: the expected value, and where you got it.
Live API response: apple one daily bar
  1. Say out loud what a report missing any of the three does. It does not enter the process at all, which is why "the prices look wrong" is not a report.
  2. Practise the honest boundary once: name one feed that comes straight from an exchange and one that comes through an aggregator. "We have direct exchange contracts" is true; "all our data is direct from exchanges" is not, and a technical client checks exactly that difference.