Retrics

RESULTS — THE PROOF STANDARD

Measured,
not promised.

Every number Retrics publishes must survive a holdout. Today the honest sample is small — so this page is a program, not a trophy case. Ten design partners, measured in the open.
How proof works ↓

ONE BACKTESTED FINDING TODAY · TEN PARTNER SLOTS OPEN

WHAT A NUMBER MUST SURVIVENO EXCEPTIONS

The number gets smaller. And truer.

FROZEN

The model is frozen at a cutoff date and graded only on the orders that really followed.

LEFT ALONE

Where an audience is large enough, a holdout slice gets no outreach at all. Whatever it spends sets the bar.

NETTED OUT

A controlled result is read against the holdout’s own observed return rate. A send too small to hold one back is read against the fixed ~2.7% floor instead.

PROVENANCE TRAVELS WITH THE NUMBER · ALWAYS

SCORED OR UNPUBLISHEDEVERY RESULT NETS OUT A BASELINEPROVENANCE TRAVELS WITH THE NUMBERTEN PARTNER SLOTSCASE STUDIES EARNED, NOT WRITTEN

ACT 01 — HOW PROOF WORKS

A causal figure faces a holdout first.

Where an audience is large enough, Retrics leaves a slice alone on purpose. Whatever those customers spend anyway is the baseline the treated group is read against. An audience too small to split carries no control of its own and is reported observationally instead — the workspace-wide control is still left alone. A number that fails the gate stays unpublished.

~2.7%

natural-return floor

the share of lapsed customers who return with no outreach — the baseline for an uncontrolled result; a controlled send is read against the observed return rate of its own withheld holdout

3.2×

second-order window flag lift

12-month backtest of real Shopify orders, measuring flag precision rather than campaign uplift

PUBLISH GATEEVERY FIGURE, NO EXCEPTIONS

MODEL FROZEN AT CUTOFF

Graded only on the orders that really followed.

HOLDOUT LEFT ALONE

No outreach to the control slice. Ever.

BASELINE NETTED OUT

A controlled result nets out its holdout’s own return rate; an uncontrolled one nets out the fixed ~2.7% floor.

WHAT SURVIVED

3.2×second-order window flag lift

12-month backtest of real Shopify orders, measuring flag precision rather than campaign uplift.

FIGURE WITHHELD · REBUILDING THE COMPARISON

The holdout is carved on every eligible send and never messaged, so the evidence accumulates. No recovered dollar is published until the comparison counts only the customers a message reached.

PROVENANCE TRAVELS WITH THE NUMBER · ALWAYS

THE ONE FINDING WE PUBLISH TODAY · AND THE ONE WE WITHHOLD

THE PROOF, IN MOTION

The recovery is visible.
The claim is earned.

This is the loop the product runs — drift scored across the base, a win-back audience frozen with a control, and the message you wrote held for one tap. Every figure on this surface is a concept preview, the recovered total included: it is withheld in the product too, while the comparison is rebuilt to count only the customers a message reached. The holdout itself is real, carved on every eligible send and never messaged.

A CONTROL WITHHELD ON EVERY ELIGIBLE SEND · THE RECOVERED FIGURE STAYS WITHHELD

Drift radarSCANNING 14,208 CUSTOMERS
C-4821

8 orders · due in 6 days

HELD
C-1177

3 orders · 71 days quiet

WIN-BACK RANKED
C-9204

5 orders · 44 days quiet

WIN-BACK RANKED
C-3358

12 orders · steady

HELD
WOULD BE CLAIMED · WITHHELD TODAY$0

CONCEPT PREVIEW · NUMBERS ILLUSTRATIVE · YOU WRITE THE MESSAGE

THE METHOD, IN FULL

Six checks between a result and your eyes.

The gate is not a slogan — it is a sequence, run the same way every time. Each step makes the surviving number smaller and harder to argue with. Nothing skips a step, and a figure that fails one never reaches this page.

01

FREEZE

The model is locked at a cutoff date before a single order is graded. From then on it can only be judged on the orders that actually landed after — never tuned to flatter a past it already saw.

02

HOLD OUT

Where an audience is large enough to withhold one, a randomly chosen slice is set aside and gets no outreach at all. It is the control the method rests on: your own store’s behaviour, left untouched, as the thing a causal claim is measured against. Audiences too small to split carry no control of their own and are reported observationally instead — the workspace-wide control is still left alone, at any audience size.

03

ACT — HELD FOR APPROVAL

Everyone outside the holdout gets the move you approved: Retrics names who and when and freezes the audience, you write the message, and it waits for a human tap before the audience is exported. Autonomy is earned: suggest first, approve next, and only then act inside limits with a kill switch.

04

GRADE ON REAL ORDERS

When the window closes, both groups are scored on orders that truly happened in Shopify — not opens, not clicks, not intent. Revenue is the only outcome that counts.

05

NET OUT THE BASELINE

A controlled result subtracts the HOLDOUT’S OWN observed return rate from the treated group — not a fixed figure. The fixed ~2.7% natural-return floor is what uncontrolled sends are credited above instead, because they have no holdout to read. Neither lane publishes a dollar today: the comparison is being rebuilt to count only the customers a message reached, and the figure waits for it.

06

PUBLISH — OR DON’T

Only a figure that survives all of the above, with its provenance attached, is ever shown. A number that fails the gate stays off the page. Today that leaves exactly one — a backtest of the prediction. Every recovered dollar and rate is withheld.

That is why this page is short. An honest result has a smaller surface area than a marketing one — and provenance travels with it, always.

WHAT WE DON’T PUBLISH

Customer logos we haven’t earned.

ROI multiples we awarded ourselves.

Aggregates, while the honest sample is one.

Projections dressed as actuals.

One backtested finding, provenance attached. Until partners are measured, that is the page.

The Honest Metrics Manifesto →

ACT 02 — THE PROGRAM

Ten brands. Three months. Measured in the open.

Design partners run Retention free for three months on their real order history. We meet once a week. What the holdout verifies becomes the first case studies on this page — that is the whole trade.

The fit: repeat-purchase Shopify brands with a thousand or more customers. The kind of store where a missed reorder window is real money.

DESIGN-PARTNER PROGRAMPHASE 1

THE FIT

  • REPEAT-PURCHASE CATEGORY
  • SHOPIFY STORE
  • 1,000+ CUSTOMERS

THE TERMS

  • THREE MONTHS FREE
  • ONE CALL A WEEK
  • CASE-STUDY RIGHTS

SLOTS

10 · OPEN

PROGRAM TERMS AS WRITTEN · NOTHING BILLED FOR THREE MONTHS

ACT 03 — THE EXCHANGE

Plain terms, both directions.

YOU GET

  • Retention, live on your store

    The full product on your real order history from day one — not a sandbox.

  • Three months free

    No invoices while we prove it. If the holdout says it worked, you decide what happens next.

  • A weekly working session

    One call with the people building Retrics. Your store sets the agenda.

  • A say in what ships next

    The weekly call is where the order of the queue gets argued. Chat is already being built; what partners press on decides what lands beside it.

WE GET

  • One hour a week

    The weekly call. That is the time commitment.

  • Your order history, measured

    Retention runs against your real cohorts, holdout included.

  • Case-study rights

    The write-up that becomes this page — holdout-measured numbers only, provenance attached, and nothing published until the comparison counts only who was messaged.

TEN SLOTS · THE EXCHANGE IN WRITING BEFORE ANYTHING RUNS

QUESTIONS

The method, answered.

What is a holdout, exactly?

A holdout is a randomly chosen slice of an audience that receives no outreach at all. It is the control a causal claim rests on: whatever those customers spend on their own is the baseline. Audiences below the minimum size are deliberately not split — splitting a tiny list produces a control too small to read — so those sends are reported observationally rather than causally.

What is the natural-return floor?

Some lapsed customers come back with no prompting whatsoever. The share that does is the natural-return floor — currently ~2.7%. It is the baseline for UNCONTROLLED sends, which have no holdout to read. A controlled send does not use it: that comparison subtracts its own holdout’s observed return rate instead.

Why does Retrics freeze the model at a cutoff?

The model is locked at a cutoff date and then graded only on the orders that actually landed after it. Freezing removes the temptation — and the possibility — of tuning a prediction to fit a past it has already seen. A frozen model is judged on the future it did not get to look at.

Why are there no customer case studies yet?

Because the honest, scored sample is still one finding. We refuse to publish aggregates while the real sample is one, and we will not write a case study a holdout has not verified. The design-partner program exists precisely to earn those case studies in the open rather than manufacture them.

What is the one number you publish today?

A 3.2× second-order window flag lift, from a 12-month backtest of real Shopify orders, measuring flag precision rather than campaign uplift. It grades the reorder-window prediction against the orders that really followed a frozen cutoff, and its provenance travels with it wherever it is shown. It is not a recovered-revenue figure. The recovered-revenue figure is withheld today. The holdout is still carved and never messaged on every eligible send, so the evidence accumulates — but the comparison is being rebuilt to count only the customers a message actually reached, and Retrics won’t publish a dollar measured over anything looser.

Are the numbers in the previews on this page real results?

No. Any figure inside a concept preview — the drift radar, the recovered total, the illustrative cohorts — is exactly that: illustrative, and labelled as such. The only published result on this page is the single backtested finding. Everything else is a picture of how the product works, not a record of what it earned.

What will Retrics never publish?

Customer logos we have not earned, ROI multiples we awarded ourselves, aggregates while the honest sample is one, and projections dressed up as actuals. If a number cannot survive the publish gate with its provenance attached, it does not go on the page.

How do I get a controlled result for my own store?

Join the design-partner program: ten slots, three months free, one call a week, and case-study rights on holdout-measured numbers only. Retention runs against your real order history, holdout included — and whatever it measures becomes one of the first case studies published here, once the comparison counts only the customers a message actually reached.

The first case study could be yours.

Ten slots. Three months free. A holdout withheld on every eligible send, and no number carries your name until it can be told apart from what would have happened anyway.

See what gets measured →