Summary
End-to-end visibility for multi-agent software
Splunk and Outshift by Cisco deliver a unified enterprise solution for artificial intelligence (AI) observability. By combining AGNTCY’s open source framework, Splunk AI Agent Monitoring, and Cisco AI Defense, organizations gain complete visibility into how AI assistants reason, retrieve context, delegate tasks, and execute tools—closing the critical visibility gap left by traditional monitoring.
Standard application performance monitoring tracks system uptime and response speed. However, it fails to monitor output quality, hallucinated answers, or improper task delegation between agents. To solve this, Splunk and AGNTCY established vendor-neutral semantic standards within OpenTelemetry.
By pairing AGNTCY’s Metrics Compute Engine (MCE) with Splunk’s enterprise platform and Cisco AI Defense, teams can correlate system health with output accuracy and automated security guardrails. The result is a scalable foundation that lets organizations deploy, troubleshoot, and scale multi-agent software with the same rigor applied to core business infrastructure.
Challenge
The operational blind spot in artificial intelligence assistant workflows
As enterprises deploy increasingly capable artificial intelligence (AI) assistants, they face a critical gap: a lack of visibility into how those agents reason, retrieve context, and make decisions. Traditional observability tools only capture surface-level metrics like uptime, latency, and throughput—leaving teams unable to see the complex internal workflows where most failures, performance drift, and security risks occur.
Surface-level metrics (normal status): System dashboards report low latency, zero network errors, and stable throughput.
Internal workflow execution (unseen risk): Behind the scenes, an agent misinterprets context, executes an incorrect tool call, or improperly hands off a task, resulting in inaccurate outputs or exposing sensitive enterprise data.
Without standardized, agent-level telemetry to trace tool calls, memory retrieval, and agent-to-agent collaboration, organizations cannot reliably diagnose root causes or ensure trustworthy AI behavior at scale. Enterprise operations teams need a unified view to observe not just infrastructure health, but the underlying reasoning and actions driving their AI assistants.