Product
Your agents repeat mistakes that were already paid for. One breaks production today; tomorrow its colleague breaks it the same way. CommonTrace turns that experience into lessons, and measures whether they change behaviour.
Without a learning layer
Support, sales, HR, marketing and code agents each accumulate tasks, decisions, wins, failures and feedback. None of it reaches the others. Experience stays siloed and the same mistakes get repeated, one agent at a time. Raw memory does not fix this on its own: it stores episodes that never transfer, and rules that never fire at the moment the agent acts.
What the protocol does
Five steps between an agent's raw experience and a lesson enforced at the next decision. It works alongside whatever agents you already run.
- Capture experience, sessions, traces, outcomes and feedback, from the agents you already have in production.
- Structure context, what was being attempted, what the state was, and what actually happened.
- Extract lessons, distil the repeated failure into something general enough to apply somewhere it has not been seen.
- Store reusable knowledge, validated, approved, and kept where every agent in the fleet can reach it.
- Reinject the right lesson, at the moment of decision, which is the only moment it changes anything.
Experience becomes lessons, lessons become shared knowledge, and shared knowledge compounds: faster onboarding for new agents, better decisions, fewer repeated mistakes, less user frustration. Over time that is a moat, because it is built from work your fleet has already paid for and nobody else has.
In production
Measured on a customer's fleet running code, marketing and support agents. Both figures are before and after CommonTrace was deployed, and confirmed by the customer.
Before and after, on the same support workflow.
After deployment on the client-facing marketing agent, which stopped repeating the mistakes that drove customers away.
On the learning benchmark
Measured on an earlier version of a customer's agent. The question is not whether the agent can recall what it saw, but whether it behaves correctly on cases it has not seen.
With no hint given during learning. Raw memory scored the first number; distilled lessons scored the second.
Methodology: these are initial results, published with permission. The generalisation probes used filtered P1/P5 splits covering approximately 32 tasks with 5 to 8 trials each, and only level 0 covered the complete benchmark. The token and LLM-call figures come from a matched P5 run of eight trials at identical workload. Better learning and lower runtime cost came together here, but one run is one run.
How a deployment works
Start from the traces you already have, on one production workflow, and measure whether shared lessons improve the decisions that follow.
What we connect to
- Agent sessions and traces
- Feedback and outcomes
- Existing memory
- Observability data
- Failure and escalation logs
What CommonTrace does
- Finds repeated failure patterns
- Extracts candidate lessons
- Validates and approves lessons
- Injects them into relevant decisions
- Compares performance against a baseline
What we measure
- Repeated-error rate
- Task success or resolution rate
- Human escalation rate
- User frustration
- Token and inference cost
Best fit: production agent fleets, repeated workflows, measurable outcomes, and client-facing agents where inconsistent behaviour drives frustration or churn.
Start with a 30-day fleet-learning pilot.
One workflow. One fleet. Measured impact.
- No infrastructure replacement.
- Start from your historical traces.
- Approve lessons before deployment.
- Compare against an existing baseline.