AI observability
If you cannot see the hop, it is not in production.
The first week is in the org, reading what already exists. Traces, identity, and cost on one trail. A person on the exceptions. A log dump is not observability. The honest no: if the work is a software factory with a named delivery date, we are the wrong partner.
Certified Partner since 2010 · MVP Hall of Fame · 200+ agents in production · UAE and US desks
Trail
Gateway, broker, agent, tool. One trail.
The same path the architecture uses. Dark hops are findings.
01
Gateway
Who called, which policy, which model, how many tokens. This is the edge.
02
Broker
Which skill was chosen. Automation, guided, or agentic.
03
Agent
Which topic, which action, which card. Write-back still fails closed.
04
Tool
Which MCP server. Success, retry, or dead letter.
Mermaid source
flowchart LR
G[Flex Gateway] --> B[Broker]
B --> A[Agent]
A --> T[MCP tool]
G -. identity and cost .-> O[Trail]
B --> O
A --> O
T --> E[Exception queue]Exceptions
Retries, dead letters, and a person.
Extraction fails. A tool call fails. The hop times out. Someone has to see it.
Cost
Token and tool usage tagged per caller.
Finance should see a process, not a cloud bill.
Questions
What we actually say.
- Is platform logging enough?
- Not if the gateway hop, the broker hop, and the tool hop live in three places with no shared identity. We will say when you already have enough.
- Do you skip human review when confidence is high?
- High-confidence reads can auto-complete. Write-back on money, identity, and policy still hits a person until the evals say otherwise.
Next
What to read next.
The brief · one email
Show us one failed hop.
We will tell you whether you can see it, and who owns it.
