All articles

How Wren decides a pattern is real

The statistical test behind the What's working summary — why Wren refuses to report most gaps, and what it takes for a pattern to earn a claim.

At the top of What's working, Wren writes a few sentences about what your post analytics show. Before anything appears there, it has to pass a test — and most things fail it. This article explains the test, because a summary you can't interrogate is a summary you shouldn't trust.

The problem: small numbers lie confidently

Group a few dozen posts by any label — post type, length, first line — and some group will always be "winning." With a handful of posts per group, one strong post landing in one group creates the entire gap. On a real corpus we tested, one post type showed nearly four times the median reach of another — and the gap turned out to be indistinguishable from luck.

So the question Wren asks is never "is there a gap?" There always is. It's "is this gap bigger than luck produces?"

The shuffle test

Here's the whole idea, and it needs no formulas.

If a label genuinely doesn't matter, then the labels are just decoration — you could peel them off your posts, stick them back on at random, and the groups should separate about as well as the real ones do.

So that's literally what Wren does: it shuffles your labels about 1,500 times and measures, each time, how well the shuffled groups separate. Then it asks: out of 1,500 fake labelings, how many separated your posts as well as the real one did? If random labeling matches your real gap all the time, the gap is noise. If almost no shuffle can reproduce it, the pattern is probably real.

Asking many questions raises the bar

Wren doesn't test one thing — it tests every label against every measure, a couple of dozen questions at once. Ask that many questions of pure noise and one or two will look "significant" by luck alone. That's not a flaw in the data; it's arithmetic.

So the bar rises with the number of questions asked. A pattern that would just barely pass as the only question fails when it's one of twenty-four — because at that point, something barely passing is exactly what chance predicts. This single rule is what separates the summary from most analytics dashboards, which happily report the luckiest of many comparisons as a finding.

What else the test guards against

Runaway posts. Reach is lottery-shaped: one post can outdo the rest of a corpus combined. The test works on a scale where "ten times bigger" counts as one big step rather than swamping everything, so a single outlier can't manufacture a pattern — though when one post does carry a group, the summary says exactly that, because "one post carried this" is often the truest available claim.

Labels that secretly measure something else. If a label only exists on part of your corpus — say, visual styles that were only captured for your recent posts — then comparing by that label really compares eras, not content. Wren detects this and refuses to make claims on such an axis, no matter how strong its numbers look.

Mood swings. The shuffles are run the same way every time, so the same data always produces the same verdicts. The summary doesn't reword itself day to day unless a finding actually changed.

What "nothing yet" means

At a typical corpus size — fifty or sixty posts — the honest answer for most people is that no pattern clears the bar yet. Wren says so plainly, tells you what is true about your corpus, and names the closest near-miss as something worth watching rather than something proven.

That's deliberate. A pattern reported before the data supports it isn't insight, it's horoscope — and acting on it costs you real posts. As your corpus grows, real patterns accumulate evidence and cross the bar; noise doesn't. The summary is designed to be worth trusting on the day it finally does make a claim.

What Wren still doesn't claim

Even a pattern that survives the test is a description, not an explanation. Wren reports what your groups did; it doesn't claim to know why, and it won't tell you what to write next. You wrote those posts about particular subjects, at particular moments, to an audience that was changing the whole time — no test on this data can fully separate those.

Still need help?

Can't find what you're looking for? Email us at support@writewithwren.com and we'll get back to you.

What to include in your message — and which address to use for what.