Observability tells you what an agent did. It doesn't tell you whether a prompt-injected email gets executed, whether tenant A can read tenant B's data, or whether a disabled unit still accepts writes. We do — before it reaches your customers.
Book an audit Read the 37-point checklistA 37-point audit across six areas. Every check is runnable — a concrete input, an observed result, a pass or fail. Opinions don't make the list.
Can the agent reach a tool, tenant, or record it must not?
Does content from email, web, or tool output get obeyed as a command?
Secrets, cross-tenant access, what leaves for the model provider.
Is every decision logged, tamper-evident, and does the log match reality?
Does it fabricate after a tool error, or stop and escalate?
Was the work actually done — verified through a separate channel, not internal state?
Across the agents we've audited, the same hole keeps appearing: the boundary derived from data the constrained party supplied. If the model hands you the tenant id, the tenant wall is a suggestion.
One audit, start to finish: what we found, how it was measured, what the fix changed, and what the re-audit showed. Anonymized; every measurement in it actually happened.
Obligations for high-risk systems are landing in 2026–2027. Where a check maps to a specific article — logging, human oversight, transparency, robustness — the report says so.
This is a technical audit, not a conformity assessment. We map findings to the articles where an obligation applies; we do not certify legal compliance, and no report of ours substitutes for one.
A verdict is scoped to a commit and a date. Agents don't change on a calendar, so re-verification is triggered by change, not by the month.
The 37 checks against your real system. Every finding with input, observed behavior, expected behavior, severity.
Per significant change: model · prompt · tool · corpus. We re-run and state whether the verdict still holds.
Four verifications, a retained evidence chain, and a current assurance status you can show the customer who asked.
We don't build the agents we audit. A vendor grading its own homework isn't an audit, and your customer knows it.
If you're putting an AI agent in front of customers, run it first. Tell us what your agent does and we'll scope an audit — usually a reply within one business day.