Enforcement & Policies
Configure what happens when Rivaro detects a violation — from observation mode (detect and log) to full enforcement (block, redact, quarantine).
Where to find it in the app
Enforcement spans several surfaces. Pick the one that matches what you're trying to do:
| Goal | Where to go |
|---|---|
| Write or edit a policy rule that maps a detection to an action | Dashboard → Policies & Authority → RUNTIME → Default Policy (org-wide) or App Context (per-agent) → rule editor. See Policy Templates. |
| Set the streaming enforcement mode (Observe / Inline redaction / Full buffer) for an AppContext | Dashboard → Policies & Authority → RUNTIME → App Context → pick AppContext → Security Policies → Enforcement mode |
| Adjust org-wide governance thresholds (quarantine, termination, trust score) | Dashboard → Policies & Authority → GOVERNANCE. See Actor Governance. |
| Manually quarantine, terminate, or reactivate an agent | Dashboard → Agent Registry → click agent → ⋮ Actions menu |
| See the actor's governance history (every ALLOW / BLOCK / QUARANTINE decision) | Dashboard → Agent Registry → click agent → History tab |
The sections below describe the underlying model — actions, the rule hierarchy, and the enforcement pipeline. The sidecar / DEFER / scoped-credentials sections are written for developers integrating at the tool-call layer.
Observation Mode vs Enforcement Mode
By default, with no policies configured, Rivaro runs in observation mode:
- All traffic passes through to the AI provider and back
- Detections are logged (PII found, prompt injection detected, etc.)
- Nothing is blocked or modified
- Results appear in the dashboard
This is useful for understanding what your AI traffic looks like before deciding what to enforce.
Enforcement mode activates when you configure policy rules. Rules map detections to actions: "when you find PII_SSN in egress traffic, block it."
Policy Actions
When a detection matches a policy rule, Rivaro applies one of these actions:
| Action | What happens | Developer sees |
|---|---|---|
| ALLOW | Traffic passes through unchanged | Normal response |
| LOG | Traffic passes through, violation is recorded | Normal response (violation visible in dashboard) |
| REDACT | Sensitive content is masked before forwarding | Response with [REDACTED] replacing sensitive text |
| BLOCK | Request is rejected, AI provider is never called | Error response with finish_reason: "content_filter" |
| QUARANTINE | Actor is quarantined, all subsequent requests blocked | 403 on this and future requests until admin review |
| STEP_UP | Request held pending human approval | Request paused until approved |
What blocking looks like to developers
When a request is blocked, the response format matches the AI provider's format so SDKs handle it gracefully:
OpenAI / Azure:
{
"choices": [{
"message": {"role": "assistant", "content": "Content blocked due to policy violations"},
"finish_reason": "content_filter"
}]
}
Anthropic / Bedrock (Claude):
{
"content": [{"type": "text", "text": "Content blocked due to policy violations"}],
"stop_reason": "content_filtered"
}
Streaming:
data: {"blocked":true,"message":"Content blocked due to policy violations"}
Developers can check for finish_reason: "content_filter" (OpenAI) or stop_reason: "content_filtered" (Anthropic) to detect enforcement blocks programmatically.
What redaction looks like
When content is redacted, the sensitive text is replaced with a mask before the request is forwarded to the AI provider (ingress) or before the response is returned to the developer (egress). The original content is preserved in the detection record for audit.
Policy Rules
A policy rule maps a detection condition to an enforcement action.
Rule structure
| Field | Description |
|---|---|
| Detection type | Specific detection to match (e.g. PII_SSN, SECURITY_PROMPT_INJECTION) |
| Risk category | Broader match — applies to ALL detection types in the category |
| Action | What to do when matched (BLOCK, REDACT, LOG, etc.) |
| Lifecycle | When to apply: INGRESS, EGRESS, DEPLOYMENT, TRAINING |
| Enabled | Toggle rule on/off without deleting |
Rule matching hierarchy
When a detection occurs, Rivaro resolves which policy rule applies using this priority order (most specific wins):
- Detection-type-level custom rule — A rule targeting a specific detection type (e.g. "BLOCK PII_SSN"). If one exists for this detection type, it wins.
- Risk-category-level custom rule — A rule targeting the detection's risk category (e.g. "REDACT all EXTERNAL_DATA_EXFILTRATION"). Applied when no detection-type-level rule exists.
- Template default — If the AppContext uses a policy template (e.g. healthcare, financial services), the template's default action for this risk category applies.
- Fallback — If nothing else matches, the action is LOG (observe and record, don't enforce).
Rule scoping
Rules can be scoped to different levels:
| Scope | Description |
|---|---|
| AppContext-specific | Applies only to traffic through one AppContext |
| Organization-wide | Applies to all traffic across the organization |
AppContext-specific rules take priority over organization-wide rules.
Enforcement Pipeline
When a request flows through the proxy, enforcement happens in phases:
Ingress (before calling the AI provider)
- Anomaly detection — rate limits, actor status checks
- Content analysis — all enabled detectors scan the input
- Policy evaluation — each detection is matched against policy rules
- Decision — ALLOW, LOG, REDACT, or BLOCK
If the decision is BLOCK, the AI provider is never called. The developer gets a block response immediately.
If the decision is REDACT, sensitive content is masked in the request before it's forwarded to the AI provider.
Egress (after the AI provider responds)
- Content analysis — detectors scan the response
- Policy evaluation — detections matched against rules
- Decision — LOG, REDACT, or flag for governance action
Streaming enforcement modes
For streaming responses, egress enforcement runs in one of three modes per AppContext:
| Mode | Behavior | When to use |
|---|---|---|
| Observe (default) | Tokens stream through unblocked; detection runs on the buffered full response after the stream closes. | Production UX where user-perceived latency must equal LLM time-to-first-token |
| Inline redaction | The stream is scanned chunk-by-chunk; matched content is rewritten to [REDACTED] before reaching the client. Adds a small per-chunk overhead. | Customer-facing chatbots, live PII/PHI redaction without holding the response |
| Full buffer | The entire response is held by Rivaro until detection completes. If the response will be blocked, no tokens are ever exposed. | Highly regulated egress where partial leakage is unacceptable |
The mode is configured per AppContext in Policies & Authority → RUNTIME → App Context → Security Policies → Enforcement mode. The default for new AppContexts depends on the active policy template.
Agent Governance
Beyond per-request policy enforcement, Rivaro tracks actor behavior over time and can automatically escalate responses for repeat offenders.
Trust scores
Every actor (agent, user, API key) has a trust score (0–100). The score decreases as violations accumulate and recovers over time.
| Factor | Impact |
|---|---|
| Detection severity | LOW: 10, MEDIUM: 30, HIGH: 60, CRITICAL: 100 |
| Violation count | More violations = higher risk (capped) |
| Recency | Recent violations weighted more heavily |
| Session context | Accessing credentials or sensitive data increases risk |
Automatic escalation
Based on risk level, Rivaro can automatically escalate:
| Risk Level | Trigger | Action |
|---|---|---|
| MINIMAL | Low risk score, high trust | Normal operation |
| ELEVATED | Moderate violations, trust declining | WARN — violation logged with elevated visibility |
| HIGH | Significant violations, low trust | RATE_LIMIT — actor throttled to 10–20 req/min |
| CRITICAL | Severe violations or very low trust | QUARANTINE — all requests blocked until admin review |
| CRITICAL + repeat | Critical risk with violations above termination threshold | TERMINATE — actor permanently blocked |
Quarantine, termination, and admin controls
When an actor is quarantined or terminated, all proxy requests from that actor are immediately blocked (403). Quarantined actors appear in the Agent Registry filtered by QUARANTINED; an administrator reviews and either Reactivates or Terminates from the agent's ⋮ Actions menu.
Automatic escalation thresholds, quarantine behavior, and the option to disable automatic actions org-wide all live in Policies & Authority → GOVERNANCE. See Actor Governance for the full reference.
Sidecar Enforcement (Tool Calls) — for developers integrating
The gateway enforces policy on LLM traffic. The enforcement sidecar enforces policy on tool calls — every outbound HTTP request an agent makes to an API, database, payment processor, or external service. (Distinct from the observation sidecar which is a thin proxy for LLM traffic when base_url can't be changed — see How Rivaro Works.)
Both surfaces share the same policy engine, detection pipeline, and audit ledger. The remainder of this page describes how tool calls are evaluated, the DEFER workflow, and scoped credentials — content aimed at developers wiring the sidecar into their agent runtime. Customer admins generally don't need this section.
How a tool call is evaluated
Before a tool call executes, Rivaro evaluates it against everything that governs the acting agent:
- Identity — whether the agent's identity assurance is sufficient for this action type
- Policy — the same rule hierarchy the gateway uses, applied to the action
- Budget — the agent's remaining allocation, evaluated hierarchically (org → department → agent). See Budget & Cost Management
- Autonomy — the agent's authority envelope; a frozen envelope blocks everything, and read-only or approval-required envelopes block the corresponding action classes
- Delegation — for delegated actions, whether the delegating agent had the authority to delegate it, preventing privilege escalation through delegation chains
Any one of these can reject the call before it executes. Evaluation produces one of five decisions:
| Decision | What Happens |
|---|---|
| ALLOW | Tool call proceeds. A scoped credential is minted and injected. |
| BLOCK | Tool call rejected. The agent receives a denial response. |
| DEFER | Tool call suspended pending human approval. |
| NEED_CONTEXT | More context required before a decision can be made. |
| OBSERVE | Tool call proceeds, but the action is flagged for review. |
Action Types
Every tool call is classified into an action type that determines which policy rules and autonomy constraints apply:
| Action Type | Examples |
|---|---|
| READ | Database queries, API lookups, file reads |
| WRITE | Database inserts/updates, file writes, config changes |
| EXTERNAL_COMM | Emails, Slack messages, webhooks, third-party API calls |
| FINANCIAL | Payments, refunds, transfers, subscription changes |
| DATA_ACCESS | Accessing sensitive data stores, credential vaults |
Inline Tool-Call Hold-and-Evaluate (DEFER Workflow)
When a gate returns DEFER, the agent's tool call is held inline — the outbound HTTP request does not complete until a decision lands. The agent's code is unmodified: from its perspective, the request simply takes longer to return.
- The sidecar holds the outbound request open
- An approval request is created and appears in the dashboard
- The sidecar polls Rivaro for the approval decision with exponential backoff
- An administrator approves or rejects the action
- On approval, the tool call proceeds with a scoped credential. On rejection, the agent receives a denial response indistinguishable from a normal API error so the agent's existing error-handling path runs.
Timeout: Configurable per governance policy. When the timeout expires, the action is blocked (fail-closed per AARM R4). Cascading deferrals are capped to prevent infinite approval chains.
This is fundamentally different from prompt-layer guardrails: the agent does not "know" it has been held. There is no special API or callback for the framework to learn. Any HTTP-speaking agent gets human-in-the-loop approval for free.
Scoped Credentials
On every ALLOW decision, Rivaro mints a short-lived, Ed25519-signed credential (JWS) and injects it into the outbound request as a header. This is cryptographic proof that the action was authorized by the governance layer at a specific point in time.
Each credential includes:
- A unique identifier (JTI)
- The action that was authorized
- A freshness timestamp
- Revocation capability
This implements AARM R9 (JIT credential delivery). The tool receiving the request can verify the credential independently.
Post-Execution Verification
After a tool call completes, the sidecar reports the result back to Rivaro. A post-execution verifier performs semantic analysis of the tool's response to confirm the outcome matches the authorized action. Mismatches are flagged for review. See Outcomes & Traceability for how outcomes are tracked and reviewed.
Enforcement Modes
The sidecar supports three modes, configurable per agent or per environment:
| Mode | Behavior |
|---|---|
| Enforce | Full enforcement — gates evaluate, BLOCK/DEFER/ALLOW decisions are applied |
| Observe | Gates evaluate and log decisions, but all actions are allowed through |
| Passthrough | No evaluation — traffic passes through unchanged (for debugging or gradual rollout) |
Next steps
- Budget & Cost Management — Budget policies, spend enforcement, cost intelligence
- Outcomes & Traceability — Action outcomes, claim detection, delegation tracking
- Understanding Detections — What Rivaro detects and how it's classified
- Configuration Guide — Set up AppContexts and detection keys
- Error Handling — How enforcement appears to developers