Can a sentiment score tell you what happens next?
The previous three lessons make sentiment look easy. One request, one number per symbol per day, an obvious axis from negative to positive. It is the most inviting field in any market-data API, and the one most likely to produce a confident conclusion that is not true.
Start from what the number is. A sentiment score is a model's output. It is a measurement of text, and only of text.
Five things it is not
Not a measurement of the world. The model reads what was written. What was written is chosen by editors chasing attention and by press offices managing a message. A quiet week for a company can be a quiet week for the company or a quiet week for the desk that covers it, and the score cannot distinguish them.
Not a validated number. You have not seen the training data, the label definitions, the language handling or the error rate. Two concrete cracks show through even in the small samples in this unit: a Spanish-language press release carried a full per-article sentiment block, while EODHD documents the aggregated /sentiments series as built on an English-language news base. Whether those are the same model over the same population is not something the response tells you.
Not a stationary scale. Across four consecutive articles, polarity read 1.000, 0.997, 0.997 and −0.273. A distribution pressed against its own ceiling has almost no room to say "even more positive", so the field's information is concentrated in rare low values — and a threshold you tuned on one period can select a completely different fraction of days in another.
Not independent of volume. A daily aggregate mixes tone with how much was written and how often it was syndicated. Apple's 422 articles over five days and Bitcoin's 20 produce numbers of the same type and nothing like the same reliability.
Not a forecast. Nothing in the response refers to any future price. Any relationship between a score and a subsequent return is a hypothesis, and it has to be tested the hard way: out of sample, on a universe you defined before you looked, with the timestamp discipline from Unit 1 — because most impressive sentiment results are look-ahead in disguise, built on a date join that paired an evening article with that afternoon's close.
What it is genuinely good for
All of that leaves a real job, and it is worth naming precisely.
A sentiment score is an excellent index into the text. It compresses thousands of articles into one number per symbol per day so that you can find the twenty days out of a thousand worth reading properly. Used that way — as a filter or as a description of coverage, with a human reading the articles the filter surfaced — it earns its place, because the score never gets the last word. Feed it into a model as one feature among many and that safeguard is gone; the score is then standing in for the text, which is the line the next paragraph draws.
It stops being safe at exactly the moment it becomes a substitute for the text rather than a route to it.
The test that keeps you honest
Before you use a score for anything, write down two things: the decision it will feed, and the error rate at which that decision becomes wrong. If you cannot state the second, the number has no job and you are collecting it because it is available.
And note the trap in evaluating it at all: if the only evidence a score works is the returns it appears to predict, and you measure that on the same data that suggested the idea, you have measured nothing at all. That is not a sentiment problem — it is the general problem — but sentiment data is unusually good at hiding it, because there are so many defensible ways to build the daily aggregate.
This is an educational description of how a published number is produced and what it can support. It is not a suggestion to trade on sentiment, and nothing here is a recommendation about any instrument, dataset or strategy.
Try it now
- Over the 60 days of
/sentimentsforAAPL.USto 25 September 2026, the lowest-scoring day was Sunday 23 August:normalized0.524 on acountof 4. Below are that day's actual/newsitems, all of them, with their per-article polarity. Decide whether you would have described the day the same way, and how much of the day's low score one article carries. The first table is the same series for an earlier August window, for comparison.
- Over the 42 trading days from 28 July to 24 September 2026, the correlation between Apple's daily
normalizedand the next session's return onadjusted_closewas 0.066 (computed 28 September 2026). State in one sentence exactly what that number does and does not establish. The sentence is the exercise. - Write down the decision and the tolerable error rate before you compute anything else. If you cannot finish that sentence, stop — you have saved yourself a month.