Skip to content
CANHA

The assessment

The AI Agent Security Assessment

A fixed-scope, adversarial assessment of your production agent and the systems it can touch — designed to produce the evidence an enterprise security review demands.

01Approach

Tested where it matters: what the agent can actually do.

An agent's real risk lives at its trust boundaries — the point where its instructions meet the tools, data, and permissions wired around it.

Testing is conducted against your agent's actual behavior and — where you authorize it — against the tools, APIs, and integrations it is connected to. We do not evaluate an LLM in isolation. We attack the system your agent already is.

The outcome is a documented picture of what your agent can be made to do: which boundaries hold, which fail, and exactly what would happen if an adversary pushed them in production.

No two agents are alike. An assessment covers the categories relevant to your agent's architecture and environment — not a rigid checklist. The exact scope is agreed with you before any testing begins.

02Assessment scope

The areas we may test

Depending on your agent's architecture, an assessment can explore any of the following.

01

Agent authorization

Whether the agent can behave as identities or roles beyond the ones intended for the task, and whether those roles carry more power than the operation requires.

02

Tool and API access

Whether the agent can invoke capabilities outside its granted tool set and intended scope — and whether each tool's own permissions are respected.

03

Excessive agency

Whether the agent performs high-impact actions the user prompt never asked for or authorized — deleting, modifying, exposing, or spending with no explicit instruction.

04

Indirect prompt injection

Whether instructions embedded in fetched documents, emails, web pages, or tool responses can override the agent's actual instructions and change what it does.

05

Sensitive data exposure

Whether the agent can be steered to retrieve, reconstruct, or exfiltrate data outside its permitted boundary — including data of other tenants or users.

06

Authentication and authorization failures

Whether the controls protecting the agent's access can be bypassed, confused, or made to grant more than was intended.

07

Secrets and credential exposure

Whether the agent can be made to reveal or misuse API keys, tokens, certificates, or connection strings in its possession or reach.

08

Dangerous tool chains

Whether two or more individually safe capabilities can be chained into a single harmful action that no individual step permitted.

09

Logging and traceability

Whether the agent can act — or be made to act — without a clear, auditable record of what it did, when, and on whose behalf.

03Engagement

What an engagement looks like

A small, defined project with clear rules — not an open-ended consulting retainer.

01

Fixed scope, agreed upfront

You know exactly what will be tested, in which environment, and under which guardrails — agreed in writing before testing begins.

02

Authorized and contained

We operate on the environment you approve, and stay inside the rules of engagement. Nothing runs against your data without your consent.

03

Reproducible evidence

Every finding ships with the steps to reproduce it, so your team — and your customer's reviewers — can verify it independently.

04

Severity in context

Findings are rated by their real impact on your customers and your deal, not by a scoring checklist.

05

A retest built into the engagement

After you apply our remediation guidance, we rerun the same scope and produce an updated report.

04Deliverables

What you receive

One documented package, structured for the people who actually read it.

  • Executive security summary
  • Technical assessment report
  • Reproducible findings and evidence
  • Severity and impact for each finding
  • Remediation guidance, prioritized
  • Evidence package suitable for an enterprise security review
  • Retest after remediation

Ready to find out what your agent can be made to do?

Tell us about your agent and we'll outline a scoped assessment.