Skip to main content
Published February 2026 10 min read

Building Auditable Agents: Receipts, Rankings, and Runtime Monitoring

The final piece of the security puzzle: proving what happened after your agent ran, detecting anomalies in real-time, and building trust signals that compound over time.

The Visibility Problem

OWASP A10 identifies a critical gap: most AI systems have no way to prove what they actually did. After an agent runs, you're left with:

• No logs — or logs that can be easily modified or deleted
• No attribution — which agent version produced this output?
• No anomaly detection — was this behavior normal or suspicious?
• No kill switch — if something goes wrong, how do you stop it?

Without visibility, you can't audit, you can't debug, and you can't build trust.

Usage Receipts

A usage receipt is a cryptographically signed record of what an agent did. It's generated at runtime, immutable, and independently verifiable.

Receipt Structure

{

"agent_id": "cdsa-v1",

"agent_commit": "abc123...",

"timestamp": "2026-02-05T14:32:00Z",

"action": "query_database",

"input_hash": "sha256:...",

"output_hash": "sha256:...",

"tokens_used": 1247,

"signature": "Ed25519..."

}

Tamper-evident: The Ed25519 signature proves the receipt wasn't modified after creation
Attributable: Links the action to a specific agent version at a specific commit
Auditable: Anyone with the public key can verify the receipt's authenticity

Rankings and Quality Signals

Receipts enable usage-proof rankings—trust signals based on verified execution history, not just claims.

Execution Count

How many times has this agent been invoked successfully? More usage = more confidence.

Error Rate

What percentage of invocations resulted in errors? Lower = better.

Token Efficiency

How many tokens does this agent typically consume? Helps predict costs.

Consistency Score

How stable are the outputs for similar inputs? Measures reliability.

These metrics are derived from verified receipts—not self-reported claims. The registry aggregates them to produce trust scores that help orchestrators choose between agent versions.

Runtime Monitoring

Receipts are post-hoc. For real-time protection, you need runtime monitoring:

Guardrails

Define boundaries for acceptable behavior: token limits, allowed tool calls, output patterns. The runtime enforces these before actions complete.

Anomaly Detection

Compare current behavior to historical baselines. Flag unusual patterns: sudden token spikes, unexpected tool calls, output distributions that don't match.

🛑 Kill Switches

Registry-controlled disable flags. If an agent is compromised or misbehaving, the registry can revoke it—and all running instances stop accepting new requests.

Alert Channels

Route anomalies to the right responders. Critical alerts to on-call, warnings to Slack, info to dashboards. Different severity, different response.

The Audit Trail

When you combine receipts with Git-verified agents, you get a complete chain of custody:

User Action
↓ triggers
Receipt Generated
↓ links to
Verified Agent ID
↓ resolves to
Git Commit + File Path
↓ signed by
Author Identity

From any action, you can trace back to: what code ran, when it was approved, who wrote it, and whether it behaved normally. This is the foundation of accountability.

Building for Accountability

The goal isn't to prevent all failures—it's to make failures visible, attributable, and recoverable. Design principles:

Log everything: Every agent invocation should produce a receipt
Sign everything: Cryptographic signatures make forgery impossible
Version everything: Pin to specific commits, never to mutable references
Monitor in real-time: Don't wait for audits to discover problems
Build kill switches: You need the ability to stop agents immediately

Series Complete

You've reached the end of the Agentic Security Series

We've covered the threat landscape (Moltbot, OWASP Top 10), the attacks (prompt injection), and the defenses (privilege separation, Git verification, auditability).

Building secure agentic systems isn't about finding a silver bullet—it's about layering defenses, assuming breaches will happen, and designing for visibility. The tools exist. The patterns exist. It's time to build.

Back to all posts