- Home
- Skills
- Agents & Orchestration
- multi agent failure mode audit
Works with the AI tools you already use
multi agent failure mode audit
Static pre-deploy audit for LangGraph, CrewAI, and AutoGen to catch deadlocks and unbounded loops.
$14
multi agent failure mode audit
Example session with this skill installed
Audit this LangGraph orchestration for potential deadlocks or unbounded loops. It handles customer support tickets with a retry loop on the search tool and a conditional edge back to the triage node.
- Read your context and instructions
- Compiled the multi agent failure
- Generated the document
I have completed the static audit of your LangGraph orchestration. I identified one high-severity structural cycle and one instance of context contamination.
Multi-Agent Failure Mode Audit Results
| Severity | Category | Location | Finding (MAST Code) |
|---|---|---|---|
| High | Deadlock | triage_node -> search_tool | FM-1.5 Unaware of Termination Conditions: The conditional edge back to triage on tool failure lacks a hard recursion_limit override or an "attempt_counter" in the state. Relying on the default limit of 25 may mask a no-progress cycle that burns tokens before failing. |
| Medium | Contamination | search_tool -> responder | FM-2.4 Information Withholding/Contamination: The responder node receives the FullState object. It only requires the search_results and original_query, but is currently inheriting the entire raw JSON trace from the search tool. |
| Low | Bounded Loop | search_retry block | FM-1.3 Step Repetition: The retry logic uses a fixed sleep with no exponential backoff. While capped, it risks hitting the same rate-limit failure 3 times in rapid succession. |
Next steps
- Implement a
max_retriesinteger in yourGraphStateand increment it at thesearch_toolnode. - Update the
respondernode to use aselectorTypedDictsubset of the state to reduce context bloat. - Set an explicit
recursion_limitin thecompile()call to 15 to catch cycles earlier.
multi-agent-failure-mode-audit.pdf
PDF · document
Example file from a real run - the skill writes it into your workspace.
Connects securely to your tools. The creator never sees your data.
What you get
About this skill
Multi-agent systems often fail in production through "Step Repetition" or deadlock cycles that burn tokens without finishing. Observability tools catch these failures after they cost you money. This skill provides a static, pre-deploy audit of your orchestration source code to identify structural flaws in graph topology, delegation logic, and state-passing design.
What it does
- Detects deadlocks by tracing conditional edges and delegation paths for reachable cycles without exit conditions.
- Identifies unbounded loops in retry logic or multi-agent handoffs that can bypass framework-level iteration caps.
- Audits termination conditions in AutoGen or LangGraph systems to ensure every possible end-state (text, tool-call, handoff) is covered.
- Flags context contamination where subagents inherit irrelevant or stale history that degrades performance or bloats costs.
How it works
- Topology Mapping: The skill reads your LangGraph, CrewAI, or AutoGen code to map every node, agent, and state-passing boundary.
- Cycle Analysis: It checks for framework-specific risks, such as
allow_delegation=Trueloops in CrewAI or unreachable exit nodes in LangGraph. - Context Review: It evaluates the scope of data passed between agents, identifying where full histories should be replaced with scoped state.
Frameworks & tools
Supports LangGraph, CrewAI, AutoGen, and custom orchestration patterns using OpenAI Agents SDK or LlamaIndex.
Why this beats prompting it yourself
Generic prompts often miss framework-specific quirks, like how CrewAI's max_iter fails to bound cross-agent delegation cycles. This skill applies a specific failure taxonomy (MAST) to find structural bugs that a general LLM would overlook as "valid" code.
Use cases
- Auditing a complex LangGraph topology before moving from staging to production.
- Reviewing a CrewAI team for potential delegation ping-pongs between peer agents.
- Validating AutoGen termination logic after adding new tool-calling capabilities.
Known limitations
This is a static code audit, not a runtime tracer. It cannot detect live data drift, spend-cap violations, or adversarial prompt injection attacks.
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean 1 day ago
- Passed all security checks, Safe to install