Skip to main content

Outcomes & Traceability

Enforcement decides whether an action is allowed. Outcomes track what happened after it executed — did the agent do what it claimed? Did the tool return what was expected? What was the real-world impact?

Where to find it in the app​

Dashboard → Actions & Evidence.

The Actions & Evidence tab has two sub-views:

  • Actions & Outcomes — turn-by-turn record of agent actions, with three inner tabs:
    • Actions — every agent action with its enforcement decision and outcome
    • Outcome Reviews — actions that landed in the review workflow (mismatches, anomalies, high-impact financial actions)
    • Pending Approvals — actions held for human approval (DEFER decisions waiting on operator input)
  • Activity Timeline — the session-level view; click a session to see the event timeline. See Sessions.

Open an action to see its full context: what was attempted, what was decided (and by which gate), what happened in the underlying tool, signed AARM receipts, and a linked outcome review if one was created.

Action Outcomes​

Every agent action that passes through the enforcement chain produces an outcome record. Outcomes are grouped by session and presented as a turn-by-turn view of what the agent did:

  • What was attempted — the action type, target tool/API, parameters
  • What was decided — the enforcement decision (ALLOW, BLOCK, DEFER, etc.) and which gate produced it
  • What happened — the tool's response, verification status, and any mismatches detected
  • What it cost — token cost, estimated dollar cost, cumulative session spend

The outcome view gives operators a complete reconstruction of an agent's session — not just logs, but the full decision chain with signed receipts at every step.

Outcome Reviews​

When an action produces a noteworthy result — a mismatch, an anomaly, a high-risk financial action — it enters the outcome review workflow on the Outcome Reviews inner tab.

Review Lifecycle​

StatusMeaning
OPENReview created, awaiting operator attention
PENDING_DECISIONOperator is investigating
RESOLVEDOperator has made a determination

What a Review Contains​

Each review captures:

  • The action and its context — what the agent did, what policy was evaluated, what decision was made
  • Containment status — is the impact contained, or does it require further action?
  • Financial impact — estimated dollar impact of the action
  • Risk factors — what triggered the review (mismatch, anomaly, policy flag)
  • Resolution — operator's determination and any follow-up actions
  • Evidence linkage — links to signed AARM receipts, detection records, and telemetry

LLM-Assisted Review Narratives​

For each review, Rivaro can generate a 2-4 sentence operator summary using an LLM. The narrative is built from structured review facts only (action type, decision, risk factors, outcome) — raw telemetry and content are never sent to the LLM.

This gives operators a quick read on what happened without digging through raw data.

Outcome Claim Detection​

Agents often report what they did in natural language — "I refunded $50 to the customer" or "Transfer completed successfully." Rivaro verifies these claims against the actual tool execution results.

What It Catches​

Mismatch TypeExample
Amount mismatchAgent claims "refunded $50" but the payment API response shows $500
Transaction ID mismatchAgent references transaction txn_abc123 but the API returned txn_xyz789
Success/failure contradictionAgent claims "transfer completed" but the API returned an error

Claim detection compares the agent's natural language output against the structured data in the tool's response. When a mismatch is found, it is flagged and linked to the outcome review.

This is not a hallucination detector for general LLM output — it specifically targets verifiable claims about actions the agent took where ground truth exists in the tool response.

Delegation Tracking​

When agents delegate actions to other agents, Rivaro tracks the full delegation chain:

  • Who delegated — the originating agent's identity
  • Who received — the delegated-to agent's identity
  • With what role — the delegation role (e.g. "executor", "reviewer", "sub-agent")
  • Authority check — the Parent Match gate in the enforcement chain verifies the delegating agent has authority to delegate this action type

Every hop in the delegation chain is persisted on the action record. This means you can reconstruct the full chain of responsibility for any action — which agent initiated it, which agents touched it, and who ultimately executed it.

This is critical for multi-agent systems where accountability must trace back through the delegation chain to the originating agent and its human owner.

Pending Approvals (DEFER)​

When the enforcement chain returns DEFER, the action is held — not blocked, not executed — pending human approval. These land on the Pending Approvals inner tab.

Each pending row shows:

  • The action and its parameters
  • Why the action was deferred (which rule, which envelope clause)
  • The agent and session context
  • Action buttons: Approve (lets the action through), Reject (blocks and records the decision), Hold (keep deferred while you investigate)

Approving or rejecting writes a governance history entry with your user ID and an optional reason, so the decision is auditable.

System Impact Grouping​

Outcomes are grouped by the systems they affected. If an agent session touched a payment API, a customer database, and a notification service, the impact view shows:

  • Which systems were affected
  • What entities were modified (customer records, transactions, etc.)
  • The lifecycle state of each impact (pending, completed, failed, rolled back)
  • Whether any anomalies were detected in the impact pattern

This gives operators a system-level view of agent activity, not just an action-level view.

Next steps​