Contents Lesson 4 of 16

4 min read · practitioner

Which field decides whether your study is honest?

The date field. Not the sentiment, not the tags, not the text. If the timestamp handling is wrong, everything downstream is a time machine, and it will produce results that look wonderful.

Three real timestamps

Three articles returned by /news for s=AAPL.US, all of them "January 6 news":

2026-01-06T21:31:05+00:00   Michael Burry / Tesla valuation piece
2026-01-06T22:00:06+00:00   Best First Jobs ... Lucrative Stock
2026-01-06T23:15:00+00:00   BitMEX lanza Equity Perps

US equities close at 16:00 New York time. In January, New York is five hours behind UTC, so that close is 21:00 UTC. Convert the three stamps: 16:31, 17:00 and 18:15 New York time. Every one of them was published after the January 6 close.

So "January 6 news" and "the January 6 close" are not contemporaneous, and joining them on a date column silently pairs each price with information that did not exist when it printed.

Which direction is the crime

Be precise about this, because the same join is fine in one direction and fatal in the other.

  • Using these three items to explain the January 6 close is a look-ahead. You have explained a price with text published after it.
  • Using them to study the January 7 close is legitimate — provided you have checked that every item in the window genuinely predates the price, and not just the three you looked at.

The rule that survives contact with real data is not a date join at all. It is: an item is available to a decision at time T only if item.date <= T, and every decision then maps to the next bar that could actually have been traded. Implement that once, as a function, and never write merge(news, prices, on="date") again.

Two mechanical details make it work. The timestamps are offset-aware — +00:00 is part of the string — so parse them as timezone-aware objects and never strip the offset to "simplify". And from and to on /news take dates, not timestamps, so a window ending to=2026-01-06 includes items published at 23:15 UTC that day. Filter again on the timestamp after you fetch.

The two clocks the API does not show you

There is a third trap that no amount of careful joining fixes. The date field records when the source stamped the item. It does not record:

  • Event time. A press release describing something that happened on Friday is stamped Monday. An earnings story is stamped when the story was written, not when the filing hit.
  • Ingestion time. When the crawler actually saw the item and it became fetchable. A backtest that uses publication timestamps assumes you had every article at the instant its publisher did, which no real system does.

Neither is available in the response, so the honest move is to record your own ingestion time when you store items, and to build in a deliberate lag — treat an item as usable some minutes after its stated publication time — rather than pretending the gap is zero.

This is the same failure mode as the look-ahead material in survivorship-and-biases, and a close cousin of the staleness question in stale-and-missing. It is described here as mechanics: how to avoid accidentally testing a strategy on information nobody had.

Try it now

  1. Here is a day of news for AAPL.US, 25 September 2026, the newest nine of its 33 items. New York was on daylight time, so the close was 20:00 UTC. Count the items stamped after it; every item not shown is earlier than the ninth. That fraction of 33 is the size of the trap for that symbol on that day.
Live API response: mda12 apple news sept 25 latest
  1. Write the availability rule as a function of one argument, T, and use it everywhere. Then delete every date-level join in your codebase.
  2. Add an ingested_at column to your own news store today. In six months it will be the only way you can answer "when could I actually have known this?"