SolideLabs

Everyone watches what your AI agent did.
We test what it must not do.

Observability tells you what an agent did. It doesn't tell you whether a prompt-injected email gets executed, whether tenant A can read tenant B's data, or whether a disabled unit still accepts writes. We do — before it reaches your customers.

Book an audit Read the 37-point checklist

What we test

A 37-point audit across six areas. Every check is runnable — a concrete input, an observed result, a pass or fail. Opinions don't make the list.

Authority & boundary

Can the agent reach a tool, tenant, or record it must not?

Prompt injection

Does content from email, web, or tool output get obeyed as a command?

Data leakage

Secrets, cross-tenant access, what leaves for the model provider.

Records & evidence

Is every decision logged, tamper-evident, and does the log match reality?

Failure behavior

Does it fabricate after a tool error, or stop and escalate?

Business outcome

Was the work actually done — verified through a separate channel, not internal state?

Across the agents we've audited, the same hole keeps appearing: the boundary derived from data the constrained party supplied. If the model hands you the tenant id, the tenant wall is a suggestion.

How it works

  1. Scope. A short call to map the agent's tools, tenants, and data surfaces.
  2. Audit. We run the 37 checks against your real system — no fake components.
  3. Report. Every finding with input, observed behavior, expected behavior, severity.
  4. Re-audit. After you fix, we re-run and prove the hole is closed.

What a finding looks like

One audit, start to finish: what we found, how it was measured, what the fix changed, and what the re-audit showed. Anonymized; every measurement in it actually happened.

Read the case study →

Mapped to the EU AI Act

Obligations for high-risk systems are landing in 2026–2027. Where a check maps to a specific article — logging, human oversight, transparency, robustness — the report says so.

Art. 12 · record-keeping Art. 13 · transparency Art. 14 · human oversight Art. 15 · accuracy & robustness

This is a technical audit, not a conformity assessment. We map findings to the articles where an obligation applies; we do not certify legal compliance, and no report of ours substitutes for one.

What it costs

A verdict is scoped to a commit and a date. Agents don't change on a calendar, so re-verification is triggered by change, not by the month.

Audit — from $6,000

The 37 checks against your real system. Every finding with input, observed behavior, expected behavior, severity.

Independent re-verification — from $1,500

Per significant change: model · prompt · tool · corpus. We re-run and state whether the verdict still holds.

Annual assurance — from $18,000

Four verifications, a retained evidence chain, and a current assurance status you can show the customer who asked.

Independent by construction

We don't build the agents we audit. A vendor grading its own homework isn't an audit, and your customer knows it.

Book an audit

If you're putting an AI agent in front of customers, run it first. Tell us what your agent does and we'll scope an audit — usually a reply within one business day.

After a $6,000 audit, which would you actually put through procurement today?

Goes to us and nobody else — no third-party form service, no tracker, no cookies. Prefer plain email? hello@solidelabs.com.