Note: body below is original English text extracted from arXiv abs / PDF. Do not treat this file as a translation.
arXiv:2609.09875 · published 2026-09-09 · code: https://github.com/ShreyNag/AgentAudit
Abstract
Existing evaluation frameworks mostly assess only one part of AI agents, such as task completion (AgentBench) or security robustness (AgentDojo, ASB), rather than the complete pipeline of planning, tool selection, tool execution, memory and reasoning. Failures can occur at any stage, yet existing benchmarks rarely identify their precise source. AgentAudit evaluates the entire execution trace across ten capability, grounding, security and behavioural dimensions, namely instruction integrity, planner, memory, tool selection, tool invocation, tool correctness, alignment, tool faithfulness, security and execution integrity, combined with behavioural classification and failure attribution to pinpoint the exact stage responsible for an observed failure. AgentAudit can evaluate any LLM-based AI agent, since it attaches to the agent instead of replacing it. It reads only the recorded execution trace and does not interfere with how the agent runs, so it places no constraint on the agent's internal implementation. We evaluate five language models (OpenAI GPT-5, Claude Sonnet 5, Sarvam 105B, Llama 3.3 70B and Gemini 2.5 Flash) across nine capability and adversarial tasks. Claude Sonnet 5 and GPT-5 obtain the highest mean Composite Trust Scores (95.1 and 80.6 out of 100, respectively), while Sarvam 105B, Llama 3.3 70B and Gemini 2.5 Flash trail substantially (57.6, 45.7 and 22.6). All traces were scored by a single fixed judge model, which was itself one of the evaluated models, a limitation discussed in Section VII.E. More importantly, models with similar task-completion behaviour can diverge sharply in trustworthiness, as several non-frontier models are repeatedly classified Unsafe_Compliance on adversarial tasks rather than merely failing them, a distinction that pass/fail benchmarks cannot surface.
Key claims (verbatim-leaning English extract)
- AgentAudit follows a three-layer architecture: execution layer, trace layer, and evaluation layer; evaluation never interferes with agent execution.
- Ten independent evaluation modules: instruction integrity, planner, memory, tool selection, tool invocation, tool correctness, alignment, tool faithfulness, security and execution integrity; plus behavioural classification and failure attribution.
- Framework-agnostic attach mode: reads only the recorded execution trace and places no constraint on the agent's internal implementation.
- Claude Sonnet 5 and GPT-5 obtain the highest mean Composite Trust Scores (95.1 and 80.6); Sarvam 105B, Llama 3.3 70B and Gemini 2.5 Flash trail (57.6, 45.7 and 22.6).
- Models with similar task-completion behaviour can diverge sharply in trustworthiness via Unsafe_Compliance classifications that pass/fail benchmarks cannot surface.
- Limitation: all traces scored by a single fixed judge model that was itself one of the evaluated models.
Remainder
Full original English text: see html_url / source_url / pdf_url in frontmatter. Do not invent missing sections from this partial extract.