The anatomy of a store that runs itself — one real decision traced end to end, the prediction engine underneath it, every seat at the table, and the economics on top.
Every stream is a stream of decisions — and every decision routes through one person, in real time, during business hours only.
The bottleneck isn’t effort. It’s architecture.
Filling the empty corner takes something genuinely new: every decision becomes a record — with its own forecast, its own safety check, its own receipt, and its own grade.
Not “an AI answers questions.” A system that picks the best allowed move — and keeps the paper trail to prove it. That difference is the rest of this talk.
A customer messages a boutique at 23:42. We follow that single message through every mechanism it touches — council, guard, gate, checkout, receipt, and grade — with the real thresholds and the real endpoints.
The eleven roles are copied from the code (lib/external-commerce/types.ts), not invented for this slide. Each specialist files a structured commitment — not chat text. Margin figures live on the blackboard and never reach anything the customer sees.
The rules behind this gate are the same ones the owner edits in plain words — the sentence you read is the rule that runs. Across 203 live offers: 185 pass · 17 need approval · 1 blocked.
A chatbot has one path: reply. This system has three: reply, ask you, or refuse — and the third is what makes the first trustworthy.
Everything the system does in the outside world leaves a receipt tied to one decision and one store. That single habit powers all the proof to come: no receipt, no claim.
A chat answer evaporates. This object accumulates — evidence, receipts, actuals, and finally a grade.
Where does “$1,850, 62% sure” actually come from? One button press fans out into eleven stages — gathering evidence, simulating thousands of similar store-weeks, racing five forecasts against each other, policing time-travel, and passing nine proof gates. We walk them one by one.
source: bazaar/v2/portal/v2_portal/workrun_full_stack_orchestrator.py · scorer_service.py · prediction_leakage_guard.py
Don’t read every box — each numbered stage gets its own slide next. Just notice two things: evidence flows left to right, and nothing skips a gate.
Stages 1–4 run top-left to top-right; 5–8 bottom-left to bottom-right. Canonical spine, stamped on every run: StudioTwin → prediction_kernel → Bazaar receipts → Autodune proof → WorkRun packet.
| run type | twin ladder | CI width target |
|---|---|---|
| micro | 256 · 512 · 1,024 · 2,048 | 0.22 |
| team | 512 → 4,096 | 0.18 |
| full | 1,024 → 8,192 | 0.13 |
| diligence | 2,048 → 16,384 | 0.10 |
Run 016 ran 2,200 twins over 3,049 source rows. That’s what one “simple” $1,850 forecast costs — and why it can be trusted.
This is the answer to “isn’t this just GPT with a nice UI?” — the language model is a bounded advisor inside a calibrated instrument.
| public benchmark | metric | predicted band | actual | baseline |
|---|---|---|---|---|
| M5 / Walmart | WRMSSE ↓ | 0.62 – 0.70 | 0.65 ✓ | 1.00 naive |
| Criteo uplift | Qini | 0.078 – 0.090 | 0.084 ✓ | 0.000 random |
| X5 / RetailHero | AUUC | 0.595 – 0.625 | 0.608 ✓ | 0.500 coin-flip |
| Hillstrom email | +pp vs control | 5.1 – 5.9 | 5.7 ✓ | 0.0 no-mail |
| HIGGS | AUROC | 0.866 – 0.882 | 0.879 ✓ | 0.733 logistic |
| FreshRetailNet | wMAPE ↓ | 0.30 – 0.36 | 0.32 ✓ | 0.58 naive |
↓ = lower is better. Bands were published before evaluation; every score is against held-out targets — 60,288 rows (M5) up to 450,000 (HIGGS). Six for six inside the band.
Three brains, separated by design: StudioTwin thinks, Bazaar acts and signs, Autodune grades. None may grade its own homework.
The same machine serves seven different humans — and rebuilds itself around each one at login. Not themes or role badges: different nouns, different front doors, different physics of what each seat may see and touch.
Default = Simple mode: one plain answer on top, at most 3 act-now items (enforced mechanically: stats.slice(0,3) · SIMPLE_ACT_NOW = 3 · simpleCap = 3), an honest “N more in the full view,” and Advanced one tap away.
The design law behind both: the primary number and the primary action live in the resting view — never one click behind a fold.
Approvals become a portfolio with a P&L; idle hours become escrowed, dual-proof trades.
The owner never posts a job, never exports a CSV. The boundary does the trust work that NDAs used to fake.
Seven seats, one invariant: every seat sees the truth it’s entitled to, and nothing it isn’t.
Once decisions are recorded objects with verified outcomes, they become sellable, priceable, and settleable. A proven playbook stops dying as tribal knowledge — and money learns to move only on proof.
Sellers are paid when a buyer’s run actually resolves — not when a dashboard feels optimistic.
| law | the enforcing code path |
|---|---|
| benchmark points never print as $ | formatter requires an ISO currency to emit “$” |
| ◆ modeled can never wear ● verified | verified requires latest_result_proof w/ real actual |
| self-rated confidence never wears green | Badge “verified” is contractually actual/paid |
| loading is never zero | in-flight reads render “—”, not 0 |
| “Sent on WhatsApp” needs a receipt | receipt bound to run_id + tenant + external_write |
| peer stats can’t identify a store | benchmarks render only at k ≥ 12 peers |
| test data can never pose as proof | force-stamped production_eligible: false |
Most products promise honesty in a values deck. Here it’s a set of compile-time properties — which is why every number in this talk could be read straight out of the system.
Put the mechanisms back together and read the delta, job by job — what the owner could not buy at any price last year, running overnight now.
| the job | the old ceiling | the mechanism that raised it |
|---|---|---|
| Sell after close | A chatbot replies and overpromises; the message waits for morning. | Council + write guard close checked sales: €142 paid at 23:51; 606 offers, 21 paid, €6,720 tracked. |
| Hold the margin | Discounts slip through whoever answers; rules live in a prompt. | Plain-word rules enforced at offer-build time — previewed by replaying your own history before shipping. |
| Run the back office | Dashboards report; you remain the integration layer. | Decisions are objects: 37 ran in a week, you approved 4 — autonomy earned at ≥5 runs / ≥80% beat rate. |
| Get expert help | Agencies need raw exports, bill monthly, never re-check. | Blind queries on sealed shapes, paid on verified results — the $650 job clears only when the fix works. Raw releases: 0. |
| Prove it worked | “AI-powered” claims with no baseline and no misses shown. | Forecasts sealed pre-actual, graded per run — “predicted 74 → landed 76, beat doing nothing by 6” — misses left in. |
| Trade what works | Playbooks die as tribal knowledge; spare capacity rots. | Two clicks list a certified run (+$2,310); van-hours clear as an escrowed $224 dual-proof match. |
Every tool before either thought or did. This one checks, acts, proves — and learns.
Loading the canvas…
Signal → council → gate → receipt → grade → lesson. Around that loop: seven seats, an economy, and an honesty constitution enforced in code.