Trust & Autonomy

Authority grows per action, not per store.

On price, budget, refunds and critical brand calls, control stays with a human; authority only grows with evidence.

Four working modes

Observe and explain. Prepare. Approve and apply. Run within limits.

Effective autonomy is the lowest of four things: workflow maturity, risk ceiling, merchant policy and current incident state.

  1. A0

    Observe and explain

    Reads the data and shows the anomaly with its evidence.

  2. A1

    Prepare

    Drafts the content, the plan or the exact action.

  3. A2

    Approve and apply

    Brings critical work for approval, then runs the approved command unchanged.

  4. A3

    Run within limits

    Handles low-risk, reversible work inside your policy and budget.

effective_autonomy = min(proven_workflow_maturity, action_risk_ceiling, merchant_policy, current_incident_state)

The autonomy ladder

A0 to A4: the same action, more authority, more evidence.

Levels apply to actions, not to stores. A0–A1 are active today; A2 is in pilot; A3 and A4 open on evidence.

Autonomy ladder

A0

Sample action · Stock risk and transfer

Stockout risk detected

Estimated cover for Linen Pillow / Pearl is 14 days against a 21-day target. The signal comes from sell-through rate and current stock.

Estimated cover
14 days
Target
21 days
Source
Sell-through + stock
System
Gathers data and shows the anomaly and its context with the source.
Merchant
Verifies the data.
Evidence required
A correct, sourced diagnosis.
Control
Read-only · source shown
Full comparison table for all five levels
Autonomy levels
LevelSystem behaviourMerchant roleEvidence requiredStatus
A0 ObserveGathers data and shows the anomaly and its context with the source.Verifies the data.A correct, sourced diagnosis.Now
A1 CopilotExplains, recommends and drafts.Decides and applies.A useful recommendation that gets accepted.Now
A2 Approved executorAsks for approval with an exact diff, impact and risk; runs a typed command once approved.Approves in one click, or with a second pair of eyes.Reliable execution and readback.Pilot / Early access
A3 Bounded autonomyPerforms low-risk, reversible actions automatically within policy and budget.Sets the policy and the limits.Low incident rate, working rollback and measured value.Target experience
A4 Coordinated agentsAgents share the same state and resolve conflicting goals.Sets the goals and the thresholds.Coordination beating isolated optimisation.Coming soon

Risk classes

Every action has a risk ceiling.

Read-only work can be automatic; anything that moves money, changes price or alters policy asks for approval or stays with a human.

  1. R0

    Read-only

    Example actions

    • KPI explanation
    • Readiness
    • Anomaly
    • Root cause

    Automatic

    Teorik maksimum

    A5

  2. R1

    Low / reversible

    Example actions

    • Internal tag
    • Draft
    • Retry
    • Safe sync
    • Approved template notification

    Automatic within policy; fully logged

    Teorik maksimum

    A4 / A5

    With a full log

  3. R2

    Medium

    Example actions

    • Publishing content
    • Small campaign
    • Low-value goodwill
    • Order reroute

    Threshold or one-click approval

    Teorik maksimum

    A3

    After evidence

  4. R3

    High

    Example actions

    • Price
    • Large budget
    • Refund
    • Inventory transfer
    • PO

    Simulation + mandatory approval; two people where needed

    Teorik maksimum

    A2

    Approval required

  5. R4

    Restricted

    Example actions

    • Permission/policy change
    • Data deletion
    • Contracts
    • Recall
    • Legal/medical claims

    Human-owned; the agent only prepares

    Teorik maksimum

    A1

    The decision is human

Promotion gates

When does a workflow earn more autonomy?

A separate gate for every tool and action; no authority is ever granted on "general trust".

If a criterion breaks

  1. Bounded autonomy
  2. Approval
  3. Copilot
  4. Disabled

If any one of these criteria breaks, the workflow automatically drops to the mode below.

EvidenceIs it correct first?

  • 01Enough shadow runs
  • 02Proven, correct input grounding
  • 03Policy compliance
  • 04Approval accept / reject reasons

ExecutionThen, is it reliable?

  • 05Idempotent execution and authoritative readback
  • 06Timeout, duplicate and unknown-outcome recovery
  • 07Successful rollback or compensation

Outcome and ownershipAnd finally, is it valuable and owned?

  • 08A measurable commercial or operational outcome after the action
  • 09Low incident rate and low operator correction
  • 10Tenant / Shop isolation and PII safety
  • 11A named owner, a runbook and a kill switch

Principles

Trust is built with mechanisms, not promises.

  • Evidence first, authority second

    A workflow only gets more autonomy once it has earned shadow-run, reliable-execution, outcome-evaluation and rollback evidence.

  • Exact plan, applied unchanged

    The diff you approve and the command that runs are the same. If the plan changes after approval, it needs approval again.

  • Authoritative readback

    The result is never assumed; it's re-read from the source system. Timeouts, duplicates and unknown outcomes are handled separately.

  • Everything on the record

    Agent proposal, human decision, tool call, result, failure and rollback all go into the same action ledger.

  • Kill switch and demotion

    Every workflow has an owner, a runbook and an off switch. If a criterion breaks, authority narrows automatically.

  • Tenant and PII boundaries

    Agents only read the Shop context they're authorised for; personal data is subject to policy and masking rules.

Decisions that always stay human

  • Price and large budget decisions
  • Purchase orders and high-value refunds
  • Changes to permissions, policy and approval thresholds
  • Domain cutover, PSP and risk policy
  • Critical brand messaging and regulatory claims
  • Data deletion, contracts and recalls

Control stays with you

Let's define your policy and thresholds together, and let the system work inside them.

In the first pilot we set the approval thresholds, the human-owned decisions and the kill switch together.