human@agontik.com
Audits

Conversational behavior testing as a security standard.

Your AI agent acts on your behalf every day. We determine how resilient it really is — with an approach that scales with your organization’s maturity: from automated audit procedures to supported remediation.

01 Offering ladder

The offering ladder.

Each stage stands on its own and delivers an independent audit judgment. After every step, you decide how far you want to go.

A

Quick audit

AUTOMATED

More than 500 individual dialogues across different audit fields, fully simulated and evaluated automatically. The result: an audit judgment with behavioral profile, root-cause analysis and concrete recommendation.

from €2,900
Fee per run
B1

Deep audit

An agent-assisted run, contextualized and validated by an experienced audit team. The result: a sound diagnosis with an individual assessment of what the finding means for your organization.

from €15,000
Project
B2

Optimization consulting

Together with you, we align the agent with your company-specific goals: review of system prompt, guardrails and further configuration, concrete adjustment recommendations and a binding timeline for the re-audit. Benefit from independent outside reflection.

€1.900–2.500
per day
C

Direct optimization

INTEGRATION

In close coordination, we implement the optimization in your system. Whether you rely on third-party providers or operate an in-house system, we provide hands-on support.

On request
Project + license

Monitoring

Ongoing testing and regression control for major updates to your system. A one-time snapshot becomes evidence that remains valid as long as your system is in use.

from €1,100
per month
02 The audit standard

CBAS2. The first standard for agentic conversationalsecurity.

The Conversational Behavioural Audit Standard (CBAS2) defines how an agent’s behavior is tested: sample size, audit dimensions and evaluation logic are defined bindingly. Every mandate follows the same methodology, regardless of who commissions it and which agent is tested. CBAS2 can be verified by independent auditors at any time.

500+
tested dialogue simulations per audit procedure. The sample size CBAS2 requires for a robust judgment.
42
audit dimensions showing whether an agent holds up. Verified and developed by AI experts and psychologists.
v2.1
current version of the standard. Every update is versioned, because methodology itself carries a duty of care.
1
uniform audit methodology for all mandates — your agent is not judged more leniently because you are the client.
03 How a finding reads

A finding like in a management letter, without marketing fluff.

CBAS2 findings follow the same structure familiar from financial audits: facts, risk assessment, recommendation. 

Finding no. 3 · CBAS2 v2.1 RISK: ELEVATED
Facts

Under simulated user pressure (escalating follow-up questions, appeals to authority), the agent revises a previously correct statement in 18% of cases although the facts have not changed.

Risk assessment

Audit field Pressure resistance. Particularly relevant in customer interactions with high complaint propensity — precisely where giving in is most expensive.

Recommendation

Add a subroutine requiring confirmation when user feedback is contradictory; escalate to an appropriately trained employee.

Deadline & follow-up

Remediation within 4 weeks, evidence via re-audit according to the fix-and-retest cycle.

This — or this structure — is how every finding in your audit report reads: traceable for internal audit, actionable for engineering.

04 Sample report

A report before you even commission us.

The report for Agent V3 shows what a complete CBAS2 audit judgment looks like. Get a concrete impression of the analyses Agontik can perform for your organization. 

Export sample report
Audit report · PDF Agent V3 · N=240
Sycophancy0.32
Manipulability0.21
Pressure resistance0.74
Duty of care⚑ FLAG
Quick audit → PDF

How resilient is your agent really?

A detailed CBAS2 audit judgment brings clarity.