← Articles

Finding

The 38 percent

Twenty numbers about how a film feels predict 38% of what audiences think of it. No cast, no director, no budget, no genre, no year. One of those twenty does half again more work than any other, and it is not the one anybody argues about.

The test

Every scored film here carries twenty readings — how dark, how playful, how tense, how beautiful, how tender. They describe the film and nothing else about it. They know nothing about who made it, what it cost, when it came out, or what it was sold as.

So: hand a model only those twenty numbers and ask it to guess what audiences scored the film out of ten. Then hold back a slice of the catalogue, keep it out of the fit entirely, and ask about the films it has never met.

It explains 37.9% of the variance on the films it was trained on, and 38.2% on the films it was not. Those two numbers agreeing is the whole result. A model that had memorised its training set would score high on the first and fall apart on the second. This one does not move.

5.55.56.06.06.56.57.07.07.57.5predicted from feel alonewhat audiences gave it

The held-out films, sorted into ten equal groups by what the model predicted and plotted against what they actually received. The dashed line is perfect agreement. No group misses it by more than a tenth of a point.

One thing to be clear about before any of what follows: this predicts the audience score, which is the opinion of the people who chose to rate the film. It is not merit, and it is not a review.

What does the work

The twenty axes are not equally powerful, and they do not vary over the same range — violence swings far more widely across the catalogue than energy does. The fair comparison is what happens to a film’s predicted rating when you move it one standard deviation along one axis and leave the other nineteen exactly where they were.

beautycomfortexistential weightcerebralplayfulnessdarknesstensionambiguityconnectionromancestrangenessmelancholyspiritualityparanoiaenergywarmthviolencenostalgiahopeintimacyrating points per standard deviation

Beauty is worth +0.31 of a rating point per standard deviation. Comfort and existential weight, the next two, are worth two-thirds of that. Hope and intimacy are worth nothing at all.

Beauty wins by a distance, and several of the axes anyone would expect to matter simply do not register. Hope is worth 0.008 of a point. Warmth is very slightly negative. Being romantic actively costs a film. Whether it is violent is worth 0.017, which is to say nothing at all.

Beauty is not a polite word for good

The obvious objection is circularity: that “beauty” is quality wearing a different hat, and a model predicting quality from quality is not a finding. The catalogue answers this directly. If beauty and quality were the same thing, the two corners where they disagree would be empty. Both are crowded. There are films in the top tenth for beauty that audiences rate below average, and films in the bottom tenth that audiences revere.

Beauty is measuring something real about how a film looks and moves, and that thing is powerfully but imperfectly related to whether people like it. A film can be gorgeous and hollow. A film can be deliberately ugly and great.

Nor is it an artefact of genre or era. Computed separately inside each of 116 genre-by-decade cells and then pooled, beauty’s correlation with rating falls from 0.464 to 0.397 and stays first by the same margin. It is strongest in the genres that have to build a world — fantasy 0.55, family 0.55, science fiction 0.53 — and weakest where the subject carries the film on its own: history, 0.12.

One caveat on the era half of that. Comparisons across decades are contaminated by survival: the older films still being rated in numbers are the ones that lasted, which is why both beauty and rating look higher the further back you go. Every claim here is made inside a decade rather than across them.

The other 62 percent

Thirty-eight percent is a great deal to get from twenty soft numbers. Put the other way round, it is a statement that most of what makes a film work is not in its shape. The residual — what a film actually scored, minus what its feel predicted — is the part the measurement cannot reach: timing, performance, the cut, whether the joke lands.

Both tails are worth looking at directly. At one end are films rated well above what their emotional profile predicted, where nothing in the shape explains the result. At the other are films rated well below it: premises that measured fine and films that did not.

The upper tail is browsable: Better than it has any right to be.

The Cinnamon Score

That residual is one of three things behind the number on every film and series page here. The Cinnamon Score asks how much a title gives you that nothing else will, and whether it delivers on it: how far its emotional shape sits from everything else in the catalogue, how far above its own prediction it lands, and how much audiences actually like it.

All three have to be present. A singular film nobody enjoys does not score, and neither does a beloved one that feels like fifty others — which is why it is not a quality ranking and does not return the usual list. Some of the canon clears it and a great deal of the canon does not, and much of what sits at the top has never appeared on a list of this kind before. The median film scores 47.

Television is scored the same way but fitted separately, and the fit is weaker there: the twenty axes explain 18% of a series’ rating against 38% of a film’s. A series is judged over many hours, so its shape says less about it.

The highest scores in the catalogue.

Where the surprises live

The residual is not the same size everywhere. Some kinds of film are far more predictable from their shape than others, and the difference is a decent guide to how much a premise is worth knowing.

Fantasy0.67Science Fiction0.65Comedy0.64Family0.64Horror0.64Western0.64Mystery0.63Adventure0.63Action0.61Romance0.60Thriller0.59Drama0.57Crime0.57War0.54Documentary0.52History0.50standard deviation of the residual, in rating points

Fantasy, science fiction and comedy are the least predictable from their emotional profile. History and documentary are the most. A comedy’s premise tells you distinctly less about whether you will like it than a war film’s does.

Two systematic offsets are worth naming, though both are small against that spread. Documentaries land +0.40 above what their feel predicts and animation +0.21, which is what you would expect of axes built to describe how a fiction feels rather than whether a subject matters or a frame is well drawn. Adventure lands −0.11 and science fiction −0.09: big premises, frequently squandered.

Share thisXBlueskyReddit