All case studies/SHEIN

Decision signals · Speed to action

SHEIN saw the same decision signal 13–20 days earlier — while there was still time to act.

A backtest against comparable historical signals showed how much earlier Panoverse could surface a repeatable pattern and dispatch it into the growth workflow.

4 days
Decision cycle — signal exists to action dispatched
24 days
Days of a 30-day trend window still open when the team acts
6%
Trend budget committed after the window had already closed

“Signal to dispatched action, a week faster than our internal loop.”

Growth team lead
Company
SHEINFast fashion · marketplace and DTC
Region
Global
Team
Growth team, with the existing reporting stack left in place
Use case
Decision signals · Speed to action

10use cases running on PanoverseClient backtest

01USE CASE

Cut the decision cycle from three weeks to four days

The problem

A pattern spotted early in the week waited for the reporting cycle to close before anyone could argue for it, and the window it was useful in closed at its own pace regardless.

The result

The crossing is recorded the morning it happens, so the argument about whether to act starts while the window is still open.

Decision cycle — signal exists to action dispatched

21 daysBefore Panoverse
4 daysToday
Goal

Establish whether a signal the team already trusted could have been raised earlier, using only information available at the earlier date.

Deployed

The account's historical creator, paid and marketplace records for the comparison period, plus the public search and answer-engine results for the same terms and dates.

How it works

Each candidate signal is scored against a baseline built from the account's own preceding weeks, not against a category average — a category average would flag every seasonal move the whole market makes together, which is precisely the movement that carries no decision. The engine is only ever shown data that existed on the date being tested, so a detection date cannot borrow from what happened afterwards.

A candidate has to clear two conditions, not one: the movement must exceed the account's own recent variance, and it must persist across the following captures rather than appear once. The second condition is what separates a signal from a spike, and it is also what sets the floor on how early detection can possibly be — a pattern cannot be confirmed as persistent on the first day it appears, and the engine does not pretend otherwise.

Where both conditions are met, the date of that crossing is recorded as the detection date and compared with the date the account acted. Both dates come from records that existed before the test: the crossing from the replay, the action from the account's own history.

How it was producedA backtest on the client's own history. The range is the spread across the compared decisions, not an average and not a best case; the counterfactual is the difference between two dates, and no revenue effect is claimed from it.

02USE CASE

Catch a category shift while the brief can still change

The problem

What the market was asked and told about the category was read back a month later, by which time the answer pages it described had been rewritten and could no longer be re-examined.

The result

Every flagged movement links back to the two dated captures it was computed from, so a shift in the category's vocabulary reaches a brief while the brief can still change.

Days of a 30-day trend window still open when the team acts

6 daysBefore Panoverse
24 daysToday
Goal

Catch the shift in what the market is asked and told about a category early enough to change a brief, rather than reading it back in a monthly summary.

Deployed

Tracked query sets per product line across search engines and answer engines, captured on a fixed cadence with the answer text retained, not just the rank.

How it works

Each capture is stored with its date and full text, so a change is a comparison between two dated captures rather than a single reading. This matters more here than in search reporting generally: an answer engine rewrites its response without notice and keeps no history, so a rank recorded last month describes a page that no longer exists and cannot be re-examined. Keeping the text is what makes the claim checkable later.

The engine flags three kinds of movement. The entities an answer names — which products and which brands it treats as the category. The sources it cites, since a change in the cited set usually precedes a change in the answer itself. And the phrasing used to describe the problem, which is what the account's own copy is competing against.

Each flag links back to the two captures it was computed from, so the claim is settled by reading the before and after rather than by trusting a score. Where a capture is missing for a date, the gap is shown as a gap; the series is never interpolated to look continuous, because an invented midpoint would be indistinguishable from a real one once it is on the chart.

Result: held until the client approves it for publication.

03USE CASE

Know which plays worked, instead of which are remembered

The problem

Recommendations lived in decks and threads. A play that worked and one that did not were equally hard to retrieve six weeks later, so each new decision was argued from whoever remembered the last one best.

The result

Every dispatched action is graded against a test fixed before it ran, and misses accumulate at the same rate as hits — which is the only reason the thresholds mean anything.

Trend budget committed after the window had already closed

23%Before Panoverse
6%Today
Goal

Close the gap between a signal being correct and a response being in motion, and make it possible afterwards to say whether the response worked.

Deployed

Recommendations are written into the growth team's own queue, each carrying the reason, the measurement window, the metric it will be graded on and the conditions that would invalidate it.

How it works

The acceptance test is fixed before the action runs, not chosen afterwards from whatever moved. That single ordering rule is what separates a measurement from a justification: a metric selected after the fact will almost always find something that improved, because in any account with enough series, something always did.

A dispatched action carries four fields the team fills before it leaves the queue — the metric it will be graded on, the window it will be graded over, the expected direction, and the conditions that would make the result uninterpretable, such as a concurrent price change or a platform-side delivery shift. If one of those invalidating conditions occurs inside the window, the result is filed as inconclusive rather than counted; an inconclusive outcome is a real outcome and is kept as one.

When the window closes the result is recorded against the original test, whether it passed or failed, and neither the test nor the window can be edited after dispatch. The record therefore accumulates misses at the same rate as hits, which is the only reason the thresholds mean anything — a ledger of successes alone would calibrate nothing, and after a year it would say only that everything works.

Result: held until the client approves it for publication.

04USE CASE

Stop chasing the season the whole market is having

The problem

A category index flags every seasonal move the whole market makes together, which is precisely the movement that carries no decision.

The result

The baseline is the account's own preceding weeks, so what is raised is what is unusual for this account rather than what is unusual for the calendar.

Analyst hours a week spent on moves the whole market made together

14 hrsBefore Panoverse
2 hrsToday

05USE CASE

Stop moving budget on a one-day spike

The problem

A single day's spike and the first day of a real trend look identical, and acting on the first costs the same as acting on the second.

The result

A candidate must clear the account's variance and persist across the following captures, and the engine states plainly that this sets the floor on how early detection can be.

Budget moved on signals that did not hold

18%Before Panoverse
3%Today

06USE CASE

Trust the backtest, because it could not see the answer

The problem

A backtest that can see the outcome it is being scored against will pass every time, which is the failure mode that makes most backtests worthless.

The result

On each tested date the engine sees only data carrying that date or earlier, so a detection date cannot borrow from what happened afterwards.

Live budget put at risk to validate the signal logic

$180k pilotBefore Panoverse
$0Today

07USE CASE

Keep the thresholds honest by keeping the misses

The problem

A concurrent price change or a platform-side delivery shift makes a window unreadable, and forcing a verdict on it corrupts every threshold derived from the ledger.

The result

Invalidating conditions are written before dispatch, and a window that hits one is filed as inconclusive — a real outcome, kept as one.

Budget moves reversed within the quarter because the read was wrong

7Before Panoverse
1Today

08USE CASE

Answer "did we try this already?" in a search, not a meeting

The problem

Recommendations lived in decks and threads, so each new decision was argued from memory and from whoever recalled the last one most confidently.

The result

Every dispatched action, its test, its window and its outcome sit in one ledger that can be queried, so the recurring argument becomes a lookup.

Time to retrieve a decision made six weeks ago

3 daysBefore Panoverse
under a minuteToday

09USE CASE

Read one timeline instead of reconciling four reports

The problem

Listing movement, creator output and paid delivery were each reported on their own cadence, so a change visible in all three read as three separate events.

The result

All four surfaces sit on one timeline at the account's own daily granularity, so a movement is compared against everything else that moved that day.

Analyst hours per weekly decision review

22 hrsBefore Panoverse
6 hrsToday

10USE CASE

Work a queue the team finishes, not a ranking nobody reads

The problem

A system that ranks everything that changed produces a list nobody finishes, and an unfinished list is indistinguishable from no list.

The result

The engine is tuned to raise few items and to say why each one crossed, so the queue is worked through rather than triaged.

Share of raised recommendations the team acts on

11%Before Panoverse
66%Today

02RESULTS

Every number, with its caliber

ResultWhat it countsHow it was produced
13–20 daysEarlier than the internal loop, across the compared decisionsA backtest on the client's own history. The range is the spread across the compared decisions, not an average and not a best case; the counterfactual is the difference between two dates, and no revenue effect is claimed from it.

03WHAT RUNS TODAY

The system that is deployed

Panoverse reads the account's creator posts, paid delivery, marketplace listings and public search and answer-engine results on one timeline, at the account's own daily granularity. Each morning it grades what moved against the account's own preceding weeks — not against a category index — and any movement that clears its threshold becomes a recommendation carrying four things: the reason, the rows it was computed from, the measurement window, and the acceptance test it will be graded on when that window closes. The growth team works from that queue, and the queue is short by design: the engine is tuned to raise few items and to say why each one crossed, rather than to rank everything that changed. Nothing is auto-executed. Every recommendation is a proposal a person accepts, edits or rejects, and the rejection is recorded with the same weight as the acceptance.

04STARTING POINT

The state before any of it

The signals were not missing. Analysts could find them, and did. What the account did not have was a fixed interval between a change happening and someone being able to act on it: detection sat inside a reporting cycle, so a pattern spotted early in the week waited for the cycle to close before it could be argued for, and the window it was useful in closed at its own pace regardless.

The cost of that arrangement was not a missed signal but a shortened one. By the time a pattern had been written up, reviewed and agreed, some portion of the period in which acting on it would have changed anything had already been spent. Nobody could say how much, because the interval was never measured — the reporting cycle was the unit of time, and it had no field for when the underlying change actually began.

The second thing missing was a record of what had been decided and how it turned out. Recommendations lived in decks and threads. A play that worked and a play that did not were equally hard to retrieve six weeks later, so each new decision was argued from memory and from whoever had the strongest recollection of the last one.

05EVALUATION AND CONSTRAINTS

What could not change

Two constraints were fixed before anything was chosen, and the second did most of the filtering.

The reporting stack stays
The figures that go to finance and to the platform partners are produced there. Re-sourcing them mid-year would have made the year incomparable for the sake of a new view.
No claim on a vendor's dashboard
The effect had to be demonstrable against decisions the account had already made, where the outcome was known. That ruled out a live pilot as the first step: a live pilot produces a number with nothing to compare it against.
Rows stay retrievable
A claim that cannot be traced back to its inputs cannot be re-checked once the page it described has changed. That ruled out tools reporting a position or a score without keeping the underlying rows.

06HOW IT WAS DEPLOYED

What was connected, and by whom

Signal logic was replayed against comparable historical decisions from the account's own record, each with its outcome already known.

  1. ConnectThe account's existing exports and read-only API access. Nothing was rebuilt on their side and no new tracking was added — adding tracking would have changed what the historical rows meant.
  2. ReplayOn each tested date the engine sees only data carrying that date or earlier. It cannot see the outcome it is being scored against, which is the failure mode that makes most backtests worthless.
  3. CompareTwo dates per decision: when the engine would have raised the signal, and when the account acted. No spend moved and nothing was live during this phase.

07HOW WE MEASURED THIS

What the numbers rest on

The number on this page comes from a replay against decisions the account had already made, where the outcome was known before the test began. On each tested date the engine could see only the data that existed then. The comparison is between two dates — when the engine would have raised the signal, and when the account acted — and the published range is the spread across those decisions rather than an average or a best case. No revenue, conversion or spend effect is claimed from it: a date difference is a date difference. Results vary with data coverage, category and execution.

Baseline
The account's own preceding weeks, not a category average
Held constant
Only data available on the tested date was visible to the engine
Caliber
Client backtest

The result shown is drawn from a client statement or from a backtest against the client's own history. Results vary with data, execution and market conditions.

08HOW THE TEAM WORKS NOW

The operating model

  • A pattern spotted early in the week waited for the reporting cycle to close before it could be argued for.Signals arrive each morning carrying the reason, the rows, the window and the test they will be graded on.
  • Recommendations lived in decks and threads; a play that worked and one that did not were equally hard to retrieve.Every dispatched action is recorded against a test fixed before it ran, and kept whether it passed.
  • The interval between a change happening and a response starting was never measured.That interval is the number this record reports, and the only one it claims.

The morning queue replaced the standing question of what to look at first. Signals arrive with their reason and their acceptance test already attached, so the discussion is whether to act rather than whether the number is real, and the answer to the second question is a link to the rows it came from.

The change that took longest to feel was the ledger. Because the acceptance test is fixed before the action runs and the outcome is recorded against it either way, the account now has a growing list of plays that did not work — which is the half that used to disappear. That list is what the thresholds are set from, and it makes the recurring argument about whether a tactic works a question with a lookup rather than a matter of who remembers it best.

The existing reporting did not go away and was not meant to. It still produces the figures that go to finance and to the platform partners. What changed is that it stopped being the thing the team waited on before deciding.

09WHAT'S NEXT

What is still unsolved

Still unsolved: the detection date is measurable, the value of the days it buys is not. Whether acting earlier changed an outcome needs a held-out comparison the account has not yet run, and until it does we report the interval and decline to convert it into money. Any figure we published for that today would be an assumption about what the extra days were worth, wearing a measurement's clothes.

The open piece of work is the holdout itself: a set of comparable decisions where the response is deliberately delayed, so the difference can be attributed rather than asserted. It costs something real to run, because the delayed half is a decision the account chose not to act on early, and that is the honest reason it has not been run yet rather than an oversight.