Contents Lesson 3 of 16

4 min read · practitioner

Why does the same story arrive five times?

Anyone who has counted news items has met this. You ask "how many articles mentioned this company yesterday", get 40, read them, and find you are looking at maybe eight actual stories. Duplication in a news feed comes from two separate mechanisms, and they need two separate fixes.

Mechanism one: one article, many symbols

Here is a real item from /news, fetched with s=AAPL.US for 2026-01-06. A single consumer-finance article titled "Best First Jobs Where You Can Earn Potentially Lucrative Stock" arrived with 22 entries in symbols:

1AAPL.MI, AAPL.BA, AAPL.MX, AAPL.US, AAPL34.SA, APC.BE, APC.DU,
APC.F, APC.HA, APC.HM, APC.MU, APC8.F,
1TSLA.MI, TSLA.MX, TSLA.US, TSLA34.SA, TL0.BE, TL0.F, TL01.F,
COST.US, NVDA.US, SBUX.US

Twenty-two identifiers; five companies. Twelve of those rows are Apple — its US line plus Milan, Buenos Aires, Mexico, São Paulo, Berlin, Düsseldorf, Hanover, Hamburg, Munich and two separate Frankfurt lines. Seven are Tesla. The remaining three are Costco, NVIDIA and Starbucks.

If your loading code explodes symbols into one row per symbol — which is the obvious relational shape — this one article becomes 22 rows. Ask "how many articles mentioned Apple" against that table and the answer is inflated twelvefold by a single piece of text. The fix is structural: keep articles and article-symbol links as two tables, count distinct articles, and decide explicitly whether a cross-listing counts as a mention.

Mechanism two: one story, many publishers

The second mechanism is syndication. A press release goes out on a wire and is republished by a dozen outlets. Each copy has its own URL, so each copy has a distinct link. Titles are near-identical but not identical — outlets append their own house style, translate, or trim.

That gives you a dedup ladder, cheapest first:

  1. Exact link. Free, and only catches literal re-fetches of the same page. Always do it.
  2. Normalised title plus publication day. Lower-case, strip punctuation and outlet suffixes, group. Catches most wire syndication.
  3. First N characters of content. Catches copies whose headlines were rewritten but whose body was not.
  4. Near-duplicate similarity over shingles or embeddings, with a threshold. Catches translations and rewrites, costs real compute, and is the only rung that can produce false merges.

Say the honest thing out loud: no deduplication is exact. Every rung is a choice about which error you prefer — counting one story twice, or merging two genuinely different stories into one. Write down which you chose and why.

Why this is not a tidiness problem

It changes numbers you will later treat as evidence.

A "news volume" series is a count, so duplicates go straight into it. A day on which one wire release was picked up twelve times looks like a day of intense coverage. And, as Unit 2 will show, aggregated sentiment is a mean over the day's items — so a single strongly-worded release, republished twelve times, votes twelve times on the day's tone. The signal you thought you were measuring is partly a measurement of how syndicated a publisher happens to be.

Try it now

  1. One day of /news for s=AAPL.US, 25 September 2026, counted on 28 September: 33 rows, 33 distinct link values, 33 distinct titles after lower-casing and stripping punctuation. The newest items are below. Three numbers, three definitions of "how much news there was", and on this day they agree: say what that tells you about syndication on that day, and why you still run all three every day.
Live API response: mda12 apple news sept 25 latest
  1. The longest symbols array that day, on "Trump-Xi Meeting Puts These 5 Chinese AI Stocks in Focus", held 50 identifiers. Group them by company. The ratio of identifiers to companies is your cross-listing inflation factor.
{"symbols": ["0700.HK", "1AAPL.MI", "1AMD.MI", "1TSLA.MI", "2RR.F", "3896.HK", "9888.HK", "9988.HK",
 "A1MD34.SA", "AAPL.BA", "AAPL.MX", "AAPL.US", "AAPL34.SA", "AHLA.F", "AMD.DU", "AMD.F",
 "AMD.HM", "AMD.MU", "AMD.MX", "AMD.US", "AMD0.F", "APC.DU", "APC.F", "APC.HA",
 "APC.HM", "APC.MU", "APC8.F", "B1C.F", "B1CB.F", "BABA.BA", "BABA.US", "BABA34.SA",
 "BABAF.US", "BABAN.MX", "BAIDF.US", "BIDU.US", "BIDU34.SA", "BIDUN.MX", "HSAI.US", "KC.US",
 "KS7.F", "KS70.F", "NNN1.DU", "NNN1.F", "NNN1.MU", "NNND.F", "NNND.MU", "NVDA.US",
 "PONY.US", "SPCX.US"]}
  1. The two most similar titles of that day that are not identical, by a character-similarity ratio of 0.46, were "What Musk's Tesla and SpaceX Stocks Are Doing After That White House State Dinner" and "Here's who is attending the Trump-Xi state dinner". Decide by hand whether they are the same story. That single judgement is the threshold your dedup rule has to encode.