← Back to all stories

Tracing the Swarm: How OpenTelemetry and Semantic Spans Conquered Multi-Agent Observability

Imagine an airline passenger whose luggage was lost during a journey that involved five connecting flights across three different airline alliances. If each airport only writes unindexed notes on paper napkins that get thrown away at midnight, finding the lost bag is impossible. But with a centralized digital baggage tracking system scanning barcodes at every conveyor belt, you can pinpoint the exact second the suitcase was routed to the wrong carousel in Zurich. This is Distributed Observability for AI Swarms.

The Chaos of Multi-Agent Communication

In a modern multi-agent system, an orchestration supervisor delegates tasks to specialized sub-agents: a Search Agent, a Code Writer, a Linter Agent, and a Database Query Agent. These agents exchange messages, execute tools in parallel, and recursively invoke one another.

When the final output fails after 20 intermediate steps, inspecting flat stdout terminal logs is like reading an unformatted phonebook: you cannot determine which agent hallucinated, which tool call timed out, or how token costs were distributed across the hierarchy.

[Distributed Trace Graph for Multi-Agent Workflow]

Root Trace: `UserRequest: Refactor Payment Service` (Total: 4.2s, $0.08)
├── Span 1: `SupervisorAgent.plan()` ────────────────► (600ms, 1,200 tokens)
├── Span 2: `SearchAgent.grep_search("process_payment")` (120ms, 3 files found)
├── Span 3: `CoderAgent.draft_patch()` ───────────────► (1,800ms, 4,000 tokens)
│     └── Child Span: `ToolCall: replace_file_content` ─► (15ms, SUCCESS)
└── Span 4: `TesterAgent.run_pytest()` ───────────────► (1,600ms, FAIL: SyntaxError)
      └── Child Span: `SupervisorAgent.retry_node()` ──► (Self-Correction Loop!)

The OpenTelemetry Semantic Conventions for GenAI

The industry solved this crisis by adopting OpenTelemetry Distributed Tracing extended with standardized GenAI semantic conventions:

  • Hierarchical Trace Spans: Every LLM invocation, vector query, and tool execution is wrapped in a discrete OpenTelemetry span with unique trace_id and parent_span_id attributes.
  • Standardized GenAI Attributes: Spans record standardized telemetry: gen_ai.system (e.g. Ollama, Anthropic), gen_ai.request.model, gen_ai.usage.prompt_tokens, gen_ai.usage.completion_tokens, and tool execution payloads.
  • Time-Series Latency & Cost Flamegraphs: Platforms like Arize Phoenix, Langfuse, and Jaeger render full interactive flamegraphs, allowing engineers to spot token bottlenecks, recursive loops, and hallucination propagation in seconds.

Engineering Takeaway

You cannot improve what you cannot measure. Treat multi-agent architectures like mission-critical distributed microservices: instrument every prompt, tool call, and state transition with standardized OpenTelemetry spans.

Reference Paper / Context: OpenTelemetry Semantic Conventions for Generative AI & Arize Phoenix — Read source ↗
👨‍💻
About the Author

I am Vikram Samal, an AI systems architect exploring how intelligent systems reason, adapt, and act—and how to make them reliable at scale. I connect emerging AI capabilities with the architectural decisions that shape performance, trust, and practical value. Through this blog, I share insights into the ideas and engineering choices shaping AI’s next chapter. As a proud father of two, I believe curiosity, human judgment, and continuous learning are essential in a world being transformed by AI.

Previous
← The Blast Chamber: Securing Autonomous Agent Terminal Access with MicroVMs and Ephemeral Sandboxes
Next
Grading the Journey: How Process Reward Models (PRMs) Replaced Flawed Outcome-Based Evaluation →