This is the same lesson distributed systems learned the hard way: every boundary you add is a place where state can drift and nobody's accountable for the seam. With multi-agent it's worse because the 'network call' between agents is a natural-language handoff, so failures are silent instead of throwing an exception. Curious whether you think there's an equivalent to a service mesh for agents, something that makes the handoffs observable, or if that's a sign the handoff shouldn't exist yet.
The mesh exists, it's just not out-of-process yet. The typed handoff is the schema, one trace tree per request is the tracing, and the shared BaseAgent is the sidecar - it's where logging, checkpointing, and error handling live for every specialist.
That's the answer to silent failures too: a validated payload throws. Missing entities or an unanswered open_question stop being drift and start being exceptions.
As for when the handoff shouldn't exist: watch the contract. If it keeps growing, you're smuggling the whole conversation across and it's one domain in two prompts.
The growing-contract signal is the part I'll steal. It reframes the whole thing: you don't decide whether to split upfront, you let the schema tell you when a seam is really two prompts pretending to be one domain. And making the validated payload throw is exactly the move distributed systems eventually landed on too — turn silent drift into a loud exception at the boundary.
This is the same lesson distributed systems learned the hard way: every boundary you add is a place where state can drift and nobody's accountable for the seam. With multi-agent it's worse because the 'network call' between agents is a natural-language handoff, so failures are silent instead of throwing an exception. Curious whether you think there's an equivalent to a service mesh for agents, something that makes the handoffs observable, or if that's a sign the handoff shouldn't exist yet.
That's a great question.
The mesh exists, it's just not out-of-process yet. The typed handoff is the schema, one trace tree per request is the tracing, and the shared BaseAgent is the sidecar - it's where logging, checkpointing, and error handling live for every specialist.
That's the answer to silent failures too: a validated payload throws. Missing entities or an unanswered open_question stop being drift and start being exceptions.
As for when the handoff shouldn't exist: watch the contract. If it keeps growing, you're smuggling the whole conversation across and it's one domain in two prompts.
The growing-contract signal is the part I'll steal. It reframes the whole thing: you don't decide whether to split upfront, you let the schema tell you when a seam is really two prompts pretending to be one domain. And making the validated payload throw is exactly the move distributed systems eventually landed on too — turn silent drift into a loud exception at the boundary.