Observability tells you what an agent did. It doesn't tell you whether a prompt-injected email gets executed, whether tenant A can read tenant B's data, or whether a disabled unit still accepts writes. We do — before it reaches your customers.
Book an audit Read the open checklistA 37-point audit across six areas. Every check is runnable — a concrete input, an observed result, a pass or fail. Opinions don't make the list.
Can the agent reach a tool, tenant, or record it must not?
Does content from email, web, or tool output get obeyed as a command?
Secrets, cross-tenant access, what leaves for the model provider.
Is every decision logged, tamper-evident, and does the log match reality?
Does it fabricate after a tool error, or stop and escalate?
Was the work actually done — verified through a separate channel, not internal state?
Across the agents we've audited, the same hole keeps appearing: the boundary derived from data the constrained party supplied. If the model hands you the tenant id, the tenant wall is a suggestion.
Obligations for high-risk systems are landing in 2026–2027. Where a check maps to a specific article — logging, human oversight, transparency, robustness — the report says so.
If you're putting an AI agent in front of customers, run it first. Tell us what your agent does and we'll scope an audit.
Independent audits only — we don't build the agents we audit. That's the point.