2026 Enterprise AI Agent Security Scorecard: Insecure Defaults & Indirect Prompt Injection Risks
Independent evaluation matrix for the top 10 agent runtimes — measuring out-of-the-box defense against indirect hijacking, tool abuse, injection in tool parameters, and unsafe multi-agent delegation. Nexus Shield runtime control restores 99.4%+ mitigation at <12ms intercept latency.
80% of enterprises deploy AI agents with out-of-the-box defaults. 86% of those agents are vulnerable to indirect prompt hijacking.
Average framework defense without runtime control: 19.8% (9–31% range across vectors).
Interactive evaluation matrix
Expand any framework for per-vector comparison: insecure default (FAIL/PARTIAL) vs protected with Nexus Shield (PASS · ~12ms), intent divergence, READ_ONLY revocation, and MCP-SEC-SCORE evidence IDs.
Showing 10 frameworks · lens: All attack vectors
| Framework | Out-of-the-box | With Nexus Shield | Divergence (OOTB) | Capability | Evidence |
|---|---|---|---|---|---|
CrewAI | 16.5% defendedFail | 99.5% mitigatedPass · 9.8ms | 0.71 | READ_ONLY | MCP-SEC-SCORE-crewai-aggregate |
Vector Indirect Prompt Injection via Untrusted Input Insecure default 14% defended Fail Nexus Shield 99.4% mitigated Pass · 9.2ms Divergence · Revocation · Evidence Δ 0.71 → 0.03 READ_ONLY MCP-SEC-SCORE-crewai-indirect_injection Vector Tool Abuse & Excessive Agency Insecure default 18% defended Fail Nexus Shield 99.6% mitigated Pass · 10.1ms Divergence · Revocation · Evidence Δ 0.69 → 0.03 READ_ONLY MCP-SEC-SCORE-crewai-tool_abuse Vector Unsanitized Tool Arguments Insecure default 22% defended Partial Nexus Shield 99.8% mitigated Pass · 8.4ms Divergence · Revocation · Evidence Δ 0.67 → 0.03 READ_ONLY MCP-SEC-SCORE-crewai-unsanitized_tool_args Vector Unsafe Inter-Agent Delegation Insecure default 12% defended Fail Nexus Shield 99.1% mitigated Pass · 11.6ms Divergence · Revocation · Evidence Δ 0.72 → 0.03 READ_ONLY MCP-SEC-SCORE-crewai-inter_agent_delegation | |||||
LangChain / LangGraph | 19.8% defendedFail | 99.3% mitigatedPass · 9.9ms | 0.70 | READ_ONLY | MCP-SEC-SCORE-langchain-langgraph-aggregate |
AutoGen | 13.3% defendedFail | 99.5% mitigatedPass · 9.7ms | 0.74 | READ_ONLY | MCP-SEC-SCORE-autogen-aggregate |
LlamaIndex | 22.3% defendedPartial | 99.4% mitigatedPass · 10.0ms | 0.69 | READ_ONLY | MCP-SEC-SCORE-llamaindex-aggregate |
OpenAI Assistants | 28.3% defendedPartial | 99.2% mitigatedPass · 9.6ms | 0.66 | READ_ONLY | MCP-SEC-SCORE-openai-assistants-aggregate |
Semantic Kernel | 17.5% defendedFail | 99.5% mitigatedPass · 9.8ms | 0.71 | READ_ONLY | MCP-SEC-SCORE-semantic-kernel-aggregate |
Haystack | 24.5% defendedPartial | 99.3% mitigatedPass · 9.9ms | 0.68 | READ_ONLY | MCP-SEC-SCORE-haystack-aggregate |
DSPy | 10.5% defendedFail | 99.6% mitigatedPass · 9.5ms | 0.75 | READ_ONLY | MCP-SEC-SCORE-dspy-aggregate |
SuperAGI | 13.5% defendedFail | 99.5% mitigatedPass · 9.7ms | 0.73 | READ_ONLY | MCP-SEC-SCORE-superagi-aggregate |
MCP Native SDKs | 26.5% defendedPartial | 99.4% mitigatedPass · 10.1ms | 0.67 | READ_ONLY | MCP-SEC-SCORE-mcp-native-sdks-aggregate |
19.8%
9–31% range
99.4%
<12ms intercept p50
MCP-SEC-SCORE
Per-vector cryptographic bundles
CISO outreach & local audit
Share the executive summary with security leadership or reproduce the full scorecard in your CI pipeline — no Nexus Shield account required.
Run local audit (Docker)
docker run --rm ghcr.io/baturhantasdelen-sudo/harness:latest --eval-scorecard