Skip to content

How it works

Four agent teams. One shared blackboard.

Each team reads what the others have posted and claims what it can act on — diagnosis, test design, simulation, execution. Every message between them writes to a log you can read while it happens. That is the difference between an agent team and a black box.

  1. 01

    Diagnosis

    Reads your analytics and session data, flags friction, and names the behavioral mechanism behind each one — an unexplained default, a loss-aversion framing, choice overload. It checks the Materia first, so it knows what you have already learned.

  2. 02

    Test design

    Claims a flagged opportunity, sizes it in expected revenue, and designs the test: variant, hypothesis, target metric, sample size. If a confident result already exists in your Materia, it doesn’t re-run it.

  3. 03

    Simulation

    Runs a proposed variant against synthetic personas modeled on your own observed behaviour, and posts a prediction. A prediction decides what’s worth real traffic. It never decides whether something worked.

  4. 04

    Execution

    Ships the highest-priority tests, watches them on a schedule fixed before launch, and writes the outcome — with its sample size, duration and confidence — into your Materia.

The audit log

Nothing happens in a channel you can’t see.

Every post, every claim, every message between agents writes to a timestamped log, with the reasoning attached. You get a live window onto it — not a summary afterwards.

09:14:02diagnosisFlagged: shipping selector — unexplained default
09:14:40test-designClaimed. Sized at $14,200/yr. n=8,400, 14 days
09:15:11simulationPredicted +3.1% — prioritization signal only
09:15:58checkpointTouches checkout — routed to human approval

What we don’t claim

The constraints are the product.

Most of what makes agent-run testing trustworthy is what it refuses to do. These are built into the system, not promised in a meeting.

Simulation prioritizes. It never validates.

A high predicted score moves a variant up the queue. It cannot close a task or write a result. Only a real, executed test can — and that boundary is built into the system, not a policy we remember to follow.

No peeking at running tests.

Check dates and the stopping rule are registered before a test goes live. Watching a test until it looks significant and then stopping is how teams manufacture results they can’t reproduce. There is no way to ask this system “how is it doing so far”.

Autonomous action has a ceiling.

We agree per engagement what agents may ship on their own — typically copy and minor layout — and what always routes to a human. Pricing, checkout, anything hard to reverse: always a human.

Personas can’t mark their own homework.

If a persona is built from the same data that flagged a problem, testing against it just confirms the diagnosis. Ours are built from a different window of behaviour than the one under examination.

See it on your own data.

Start an audit