Trust & Autonomy
Authority grows per action, not per store.
On price, budget, refunds and critical brand calls, control stays with a human; authority only grows with evidence.
- A0 ObserveNow
- A1 CopilotNow
- A2 ApprovedPilot
- A3 BoundedTarget
- A4 CoordinatedSoon
Authority is per action, and grows with evidence
Four working modes
Observe and explain. Prepare. Approve and apply. Run within limits.
Effective autonomy is the lowest of four things: workflow maturity, risk ceiling, merchant policy and current incident state.
A0
Observe and explain
Reads the data and shows the anomaly with its evidence.
A1
Prepare
Drafts the content, the plan or the exact action.
A2
Approve and apply
Brings critical work for approval, then runs the approved command unchanged.
A3
Run within limits
Handles low-risk, reversible work inside your policy and budget.
effective_autonomy = min(proven_workflow_maturity, action_risk_ceiling, merchant_policy, current_incident_state)
The autonomy ladder
A0 to A4: the same action, more authority, more evidence.
Levels apply to actions, not to stores. A0–A1 are active today; A2 is in pilot; A3 and A4 open on evidence.
Autonomy ladder
Sample action · Stock risk and transfer
Stockout risk detected
Estimated cover for Linen Pillow / Pearl is 14 days against a 21-day target. The signal comes from sell-through rate and current stock.
- Estimated cover
- 14 days
- Target
- 21 days
- Source
- Sell-through + stock
- System
- Gathers data and shows the anomaly and its context with the source.
- Merchant
- Verifies the data.
- Evidence required
- A correct, sourced diagnosis.
- Control
- Read-only · source shown
Full comparison table for all five levels
| Level | System behaviour | Merchant role | Evidence required | Status |
|---|---|---|---|---|
| A0 Observe | Gathers data and shows the anomaly and its context with the source. | Verifies the data. | A correct, sourced diagnosis. | Now |
| A1 Copilot | Explains, recommends and drafts. | Decides and applies. | A useful recommendation that gets accepted. | Now |
| A2 Approved executor | Asks for approval with an exact diff, impact and risk; runs a typed command once approved. | Approves in one click, or with a second pair of eyes. | Reliable execution and readback. | Pilot / Early access |
| A3 Bounded autonomy | Performs low-risk, reversible actions automatically within policy and budget. | Sets the policy and the limits. | Low incident rate, working rollback and measured value. | Target experience |
| A4 Coordinated agents | Agents share the same state and resolve conflicting goals. | Sets the goals and the thresholds. | Coordination beating isolated optimisation. | Coming soon |
Risk classes
Every action has a risk ceiling.
Read-only work can be automatic; anything that moves money, changes price or alters policy asks for approval or stays with a human.
R0
Read-only
Example actions
- KPI explanation
- Readiness
- Anomaly
- Root cause
Automatic
Teorik maksimum
A5
R1
Low / reversible
Example actions
- Internal tag
- Draft
- Retry
- Safe sync
- Approved template notification
Automatic within policy; fully logged
Teorik maksimum
A4 / A5
With a full log
R2
Medium
Example actions
- Publishing content
- Small campaign
- Low-value goodwill
- Order reroute
Threshold or one-click approval
Teorik maksimum
A3
After evidence
R3
High
Example actions
- Price
- Large budget
- Refund
- Inventory transfer
- PO
Simulation + mandatory approval; two people where needed
Teorik maksimum
A2
Approval required
R4
Restricted
Example actions
- Permission/policy change
- Data deletion
- Contracts
- Recall
- Legal/medical claims
Human-owned; the agent only prepares
Teorik maksimum
A1
The decision is human
Promotion gates
When does a workflow earn more autonomy?
A separate gate for every tool and action; no authority is ever granted on "general trust".
- Bounded autonomy
- Approval
- Copilot
- Disabled
If any one of these criteria breaks, the workflow automatically drops to the mode below.
EvidenceIs it correct first?
- 01Enough shadow runs
- 02Proven, correct input grounding
- 03Policy compliance
- 04Approval accept / reject reasons
ExecutionThen, is it reliable?
- 05Idempotent execution and authoritative readback
- 06Timeout, duplicate and unknown-outcome recovery
- 07Successful rollback or compensation
Outcome and ownershipAnd finally, is it valuable and owned?
- 08A measurable commercial or operational outcome after the action
- 09Low incident rate and low operator correction
- 10Tenant / Shop isolation and PII safety
- 11A named owner, a runbook and a kill switch
Principles
Trust is built with mechanisms, not promises.
Evidence first, authority second
A workflow only gets more autonomy once it has earned shadow-run, reliable-execution, outcome-evaluation and rollback evidence.
Exact plan, applied unchanged
The diff you approve and the command that runs are the same. If the plan changes after approval, it needs approval again.
Authoritative readback
The result is never assumed; it's re-read from the source system. Timeouts, duplicates and unknown outcomes are handled separately.
Everything on the record
Agent proposal, human decision, tool call, result, failure and rollback all go into the same action ledger.
Kill switch and demotion
Every workflow has an owner, a runbook and an off switch. If a criterion breaks, authority narrows automatically.
Tenant and PII boundaries
Agents only read the Shop context they're authorised for; personal data is subject to policy and masking rules.
- Price and large budget decisions
- Purchase orders and high-value refunds
- Changes to permissions, policy and approval thresholds
- Domain cutover, PSP and risk policy
- Critical brand messaging and regulatory claims
- Data deletion, contracts and recalls
Control stays with you
Let's define your policy and thresholds together, and let the system work inside them.
In the first pilot we set the approval thresholds, the human-owned decisions and the kill switch together.