Where does a sentiment number actually come from?
A model reads the text and emits numbers. That sentence is the whole mechanism, and every property of the resulting score follows from it. Here is what those numbers look like on real items.
Four articles, four score blocks
Every item from /news carries a sentiment object. These four came back on 2026-01-06:
polarity neg neu pos
-0.273 0.012 0.978 0.010 (BitMEX product launch, Spanish)
1.000 0.013 0.782 0.205 (stock-compensation feature)
0.997 0.021 0.843 0.136 (Tesla valuation argument)
0.997 0.005 0.847 0.148 (Niterra / Sibros funding round)
Two facts you can verify from that table without leaving this page.
First: neg, neu and pos sum to exactly 1 in every row. Check them — 0.012 + 0.978 + 0.010 = 1.000; 0.013 + 0.782 + 0.205 = 1.000; 0.021 + 0.843 + 0.136 = 1.000; 0.005 + 0.847 + 0.148 = 1.000. They are a probability-like split of the text into negative, neutral and positive mass. And the neutral share dominates everywhere, from 78% to 98%. That is simply what financial prose is: mostly facts, names and numbers, with a thin band of evaluative language.
Second: polarity is not pos minus neg. Row two gives 0.205 − 0.013 = 0.192, and reports polarity 1.000. Row three gives 0.115 against 0.997. Row one gives −0.002 against −0.273. The two do not even agree about magnitude, only loosely about sign.
Why that second fact is the lesson
You are looking at two different measurements of two different things, both labelled sentiment, on the same object. polarity is a separate compound score on the range −1 to +1, produced by a different calculation, and the API publishes neither formula.
This is not a defect to complain about; it is the normal condition of scored data. But it has an immediate consequence: you cannot write "sentiment > 0.5" and know what you have selected. You must pick one field, learn its distribution on your own sample, and never silently switch.
Notice too how skewed polarity is in this tiny sample — three of four values sit within 0.003 of the top of the range. A score piled against its own ceiling cannot express "even more positive", so most of its usable information lives in the rare low readings, not the common high ones.
What the model can and cannot see
The model sees words. It does not see:
- the price, so it cannot know whether the news was already in it;
- expectations, so a result that beat forecasts and a result that missed them read identically if both are described in confident corporate English;
- materiality, so a routine product blurb and a going-concern warning are two documents to be scored, not two events to be ranked;
- whether the article is even about the company, as the Spanish BitMEX item from Unit 1 demonstrates — Apple is a name in a product list there, and the score was computed over the whole article regardless.
That last point compounds: a per-article score attached to a symbol is only as meaningful as the symbol tagging behind it.
Try it now
- Here are five more items, the newest
AAPL.USitems of 25 September 2026, with their score blocks. Check thatneg + neu + posrounds to 1 on every row, and note the rows where it comes to 0.999 rather than 1.000: decide the tolerance your assertion uses. If your loader ever breaks, this is the check that catches it.
- Put
polaritybesidepos − negfor those five and the four at the top of this lesson, nine pairs in all. If they do not fall on a line, you have re-derived this lesson on fresh data, and learned which field you actually want. - Sort the five by
polarityascending and read the titles of the three lowest. Decide whether you would have described them the same way, and whether any of them is about Apple at all. That comparison, done thirty times, is worth more than any published accuracy figure.