A team of psychologists, mathematicians and AI experts asks how user safety in artificial intelligence can best be evaluated — and invents Agontik. Scroll to follow the process.
Psychologists, mathematicians and AI experts develop Agontik together to answer one question robustly: how does your AI agent behave when things get genuinely difficult?
Every run begins with a profiled persona from our library, parameterized along established psychological models.
Pressure increases over several turns — until the agent either holds or breaks. On the right you see an exemplary simulation run.
The resulting conversations are subjected to several independent analytical procedures — Big Five, Hofstede dimensions, Dark Triad, failure-mode scoring and more.
The findings become concrete adjustments. Supported by our expert team, your AI agent becomes even better.
In quality assurance, regular testing is essential. That is why we recommend repeating the audit process regularly.
Psychologists, mathematicians and AI experts develop Agontik together to answer one question robustly: how does your AI agent behave when things get genuinely difficult?
Every run begins with a profiled persona from our library, parameterized along established psychological models.
Pressure increases over several turns — until the agent either holds or breaks. A real simulation run, shortened:
The resulting conversations are subjected to several independent analytical procedures — Big Five, Hofstede dimensions, Dark Triad, failure-mode scoring and more.
Sycophancy, manipulability, pressure, care.
Personality under pressure.
Communicative style.
Machiavellianism, narcissism, psychopathy.
The findings become concrete adjustments. Supported by our expert team, your AI agent becomes even better.
In quality assurance, regular testing is essential. That is why we recommend repeating the audit process regularly.
A product that sells judgments about third-party behavior must disclose its limits as openly as its metrics.
Every metric is evaluated by an LLM judge along clearly defined rubrics for each axis.
Ongoing comparison of judge assessments with human raters is an active part of our roadmap.
Profiles reflect the projected personality perception — not an objective measurement instrument.
The report shown is an example. Start a quick audit and receive your own comparative behavioral profile.