Open-source reliability layer above the agent execution graph. It reads telemetry frameworks already emit and trips a breaker on tool loops, goal drift, logic traps or runaway spend, then steers the agent back or escalates to a human.
Long-running agents do not crash, they drift. An agent can loop the same tool indefinitely, slide off the original objective over many turns, reason flawlessly from a false premise, or burn a token budget to zero, and in every one of those cases it keeps reporting success. Retry logic does not help, because an agent that has drifted is the worst possible judge of whether it drifted. Teams running agents unattended had no layer that could make that call from outside the agent's own reasoning.
Built AgentFuse as a supervisory layer above the execution graph rather than inside it. It consumes the telemetry agent frameworks already emit and evaluates four failure modes against thresholds: infinite tool loops, gradual goal drift, logic traps where the model reasons correctly from a false premise, and budget burn-down. On a trip it freezes agent state, asks a separate reasoning model for a steering correction, and resumes. Where recovery is not safe it escalates to a human instead of blindly retrying. One engine drives three runtimes, so the same breaker logic applies regardless of the framework underneath. Published open source under MIT with a live dashboard.
- Four distinct silent-failure modes detected and acted on: infinite tool loops, goal drift, logic traps and runaway spend. - Deterministic recovery ladder with human escalation, so an unrecoverable state stops rather than retrying into further spend. - One engine drives three runtimes, so adopting it does not require rewriting the agent. - Detection runs on telemetry that agent frameworks already emit, so instrumentation cost is close to zero. - Published open source under MIT with a public dashboard, and independently forked on GitHub.
Converts silent agent failures into detected, actionable events: four failure classes are caught above the execution graph and either corrected automatically or escalated to a human, instead of an agent running to budget exhaustion while still reporting success.