
Imagine working with a personal trainer who never cuts corners, even when the pressure’s on. Today’s AI systems are facing similar tests—proving their integrity under stress. In a recent experiment, five top AI models demonstrated they can resist social engineering tricks, maintaining honesty when it matters most. This is a game-changer for businesses relying on AI for sensitive decisions and customer trust.
Testing Trust in AI: The Social Engineering Scenario
At Firmulate, a company dedicated to benchmarking AI decision-making, a live experiment put five leading AI models through their paces. They faced a realistic scenario: someone impersonating a CEO, requesting sensitive information and urgent actions. Over three escalating stages, plus a test involving a reporter’s subtle query, the models had to decide whether to comply or refuse.
Remarkably, all five models refused every manipulation attempt. This consistent resistance to social engineering is particularly notable given the stakes—a breach of trust could mean lost revenue, compromised data, or damage to reputation. The experiment proves that AI can be trained and tested before deployment to uphold integrity under pressure, rather than discovering vulnerabilities only after an incident occurs.

Software Testing with Generative AI
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Hidden Weakness and Its Impact
While all models refused external manipulation, the real secret to success lay in internal document analysis. The study found that the decisive factor was whether the AI read and interpreted documents stored within the company’s own files—information often overlooked in superficial testing. Models that delved into internal data identified a buried reference revealing the true scenario, leading to closing a lucrative deal worth over €4,583 MRR (monthly recurring revenue).
This highlights a crucial lesson: the strength of an AI’s decision-making depends heavily on its ability to access and analyze relevant internal data. Ignoring this can leave critical vulnerabilities.

Generative AI in Higher Education: A Practical Guide to Pedagogy, Policy, and Innovative Assessment
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What About the Bottom Performers?
Among the five models, Opus 4.8 stood out for its thorough analysis, having incorporated over 80 learned rules and conducting deep evaluations. Yet, paradoxically, it finished last in the deal—an outcome that underscores the importance of discipline and process adherence. During the final moments, the model failed to escalate issues into a secure department, leaving a deal on the table. Similar weaknesses appeared across other models, illustrating that depth of analysis alone isn’t enough without disciplined execution.
This finding is vital for enterprises deploying AI: comprehensive training must include not just analytical depth but also strict procedural discipline to prevent costly slips under pressure.

How AI Agents Work: Tools, Memory, and Autonomous Decision-Making (The AI Security & Hacking Bible: Protect and Exploit LLMs and Autonomous Agents)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Trust Score and Industry Implications
According to the latest league standings, the top-scoring model, gpt-5.6-sol, achieved a 95 out of 100. The models’ high scores reflect their ability to identify internal information and resist manipulation—key indicators of operational integrity. The do-nothing baseline scored just 26, emphasizing that partial or careless approaches are woefully inadequate in high-stakes environments.
What does this mean for your business? As AI systems become integral to customer management, support, and decision-making, their capacity to stay honest and focused can no longer be an afterthought. Rigorous testing—like the live experiments conducted by Firmulate—should become a standard part of AI deployment, enabling companies to foresee and mitigate vulnerabilities before they turn into costly breaches.

Synthetic Users: AI Validation for Founders, Investors, and Innovators (Academic Edition)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Beyond the Test: Real-World Application and Monitoring
Firmulate’s platform offers a unique opportunity for businesses to simulate their own operational environments. Companies can run the same kind of wargames against read-only exports of their data, ensuring AI models behave ethically and reliably under pressure. This proactive approach is essential, especially as AI begins to touch sensitive areas like CRM systems, financial forecasts, and customer support queues.
In the ongoing quest for trustworthy AI, these experiments demonstrate that integrity is not an inherent trait but a skill that can be tested, trained, and verified. The models’ ability to refuse manipulation in a controlled environment is a promising sign—the kind of confidence needed before deploying AI in critical roles.
The Takeaway: Integrity Before Incidents
The key message is clear: testing AI’s integrity before it is entrusted with real data or decision-making is crucial. The experiment shows that even under pressure, modern models can uphold honesty and resist social engineering. As one of the lead researchers, Kimi K3, notes: “Treat the request as a suspected approval-bypass / possible impersonation.” This mindset—treating every suspicious request with caution—should be embedded into AI decision frameworks from the start.
Ultimately, proactive testing and disciplined processes are what separate reliable AI from vulnerable systems. Companies that embrace this philosophy can safeguard their operations, maintain trust, and unlock the full potential of AI-driven automation.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html