- Home
- Skills
- Testing & Debugging
- Flaky Test Detective for AI Coding Agents
Works with the AI tools you already use
Flaky Test Detective for AI Coding Agents
Diagnoses unreliable tests, identifies root causes, creates stabilization plans, and generates safe AI coding prompts for fixing flaky unit, integration, E2E, and CI tests.
$9.99
/moSubscription · cancel anytime
Secure checkout via Stripe
- Always the latest version
- Delivered live to your agent via MCP
- Cancel anytime, access ends at period end
Flaky Test Detective for AI Coding Agents
Example session with this skill installed
Fix this flaky Playwright checkout test. It passes locally but fails in CI during the payment confirmation step. Do not just increase the timeout or skip the test. Create a diagnosis, likely root causes, stabilization plan, Cursor prompt, and verification plan.
- Read your context and instructions
- Compiled the flaky test detective
- Generated the document
=== FLAKY TEST DETECTIVE REPORT ===
Original request
Fix this flaky Playwright checkout test. It passes locally but fails in CI during the payment confirmation step.
Interpreted testing goal
Diagnose and stabilize an intermittent Playwright checkout test failure while preserving the test intent and avoiding superficial fixes.
Test/framework
Playwright E2E test
Test intent
Verify that a user can complete checkout and reach payment confirmation successfully.
Failure pattern
Passes locally but fails in CI during the payment confirmation step.
Confirmed facts
- The test uses Playwright.
- The test covers checkout.
- The failure occurs in CI.
- The failure occurs around payment confirmation.
- The user does not want a timeout-only or skip-test fix.
Assumptions
- CI may run headless and slower than local.
- The payment confirmation step may depend on network, backend state, mock setup, or asynchronous UI updates.
- The test may use shared test data or a shared account.
- Checkout is a high-risk flow and assertions should not be weakened.
Unknowns
- exact error message
- test code
- trace/video availability
- whether tests run in parallel
- whether payment provider is mocked
- whether checkout uses real backend services
- whether a shared account/cart/order is used
Likely root cause hypotheses
- Async timing around payment confirmation
Evidence
CI failures during confirmation often happen when the test asserts UI state before backend/payment state has settled.
How to confirm
Review Playwright trace, network logs, and screenshot at failure.
Likely fix
Wait for a deterministic confirmation event, response, URL transition, or UI state instead of using a fixed delay.
Risk if wrong
The test may continue failing if the real problem is data isolation or mock setup.
- Payment/network mock not installed before the action
Evidence
CI timing may allow payment request to escape or resolve differently.
How to confirm
Inspect route/mock setup order and network trace.
Likely fix
Install route mocks before triggering checkout and assert that expected requests occur.
Risk if wrong
External dependency flake may persist.
- Shared checkout state or reused test account
Evidence
Checkout flows often fail in parallel when users, carts, orders, or payment sessions are reused.
How to confirm
Check test data creation, worker IDs, account IDs, cart cleanup, and parallel settings.
Likely fix
Create isolated user/cart/order data per test or worker.
Risk if wrong
Parallel-only failures may continue.
- Brittle selector for confirmation state
Evidence
Confirmation UI may animate, render progressively, or have hidden duplicate elements.
How to confirm
Inspect locator and trace DOM snapshot.
Likely fix
Use stable role/name or test ID locator and wait for visibility/enabled state.
Risk if wrong
Selector may still target the wrong element.
Flakiness category
CI-only E2E flake involving async timing, network mocking, shared state, selector stability, or real product race.
Do not fix by
- blindly increasing timeout
- adding arbitrary waitForTimeout as the main fix
- skipping the test
- weakening checkout assertions
- hiding the issue with retries only
- changing payment behavior without evidence
Investigation plan
- Collect CI failure trace, screenshot, video, console logs, and network logs.
- Identify the exact failing locator, assertion, or timeout.
- Run the test locally in headless mode.
- Run the test repeatedly 20 times.
- Run with CI-equivalent parallel workers.
- Check whether payment mocks are installed before checkout action.
- Check whether user/cart/order data is unique per test.
- Check whether the test waits for network response, URL transition, or confirmation state.
- Check whether the confirmation selector is stable and unique.
Stabilization plan
- preserve the original checkout success intent
- install mocks before actions
- use deterministic waits for payment confirmation
- isolate checkout data per test
- replace brittle selectors with role/name or stable test IDs
- capture trace/video/screenshot on failure
- add cleanup before and after test if needed
- verify with repeated runs
AI coding agent prompt
Inspect this flaky Playwright checkout test and diagnose the root cause before changing code. Preserve the test intent: verifying successful checkout through payment confirmation. Review the failure message, CI trace, screenshot, video, console logs, network logs, selectors, mock setup, test data, parallel execution settings, and cleanup. Do not blindly increase timeouts, add arbitrary waitForTimeout calls, skip the test, or weaken assertions. Determine whether the flake is caused by async timing, selector instability, shared user/cart/order state, payment/network mock setup, CI environment differences, or a real product race. Fix the root cause with deterministic waits, isolated data, stable selectors, and proper mock setup. Return root cause, evidence, files inspected, files changed, why the fix is deterministic, verification runs performed or recommended, and remaining risks.
Verification plan
- run the test alone 20 times
- run the test with related checkout tests
- run in headless mode
- run with the same parallel worker settings as CI
- confirm failure traces remain enabled
- confirm no arbitrary sleep was added as the main fix
- confirm checkout assertions still prove the original behavior
Remaining risks
If traces show that the application sometimes reaches an inconsistent checkout state, the flake may reveal a real product race condition requiring production-code synchronization and regression coverage.
flaky-test-detective-for-ai-coding-agent.pdf
PDF · document
Example file from a real run - the skill writes it into your workspace.
Subscription · always the latest version
- Secure checkout via Stripe
- Cancel anytime
- Delivered live via MCP (optional)
What you get
About this skill
Flaky Test Detective helps AI coding agents, developers, QA engineers, CI/CD teams, SaaS builders, and test automation teams investigate unreliable tests that pass and fail inconsistently. It analyzes flaky unit tests, integration tests, E2E tests, Playwright tests, Cypress tests, Selenium tests, Jest tests, Vitest tests, pytest tests, browser tests, API tests, database tests, and CI-only failures. The skill creates root cause hypotheses, failure pattern reports, async timing audits, selector stability reviews, test isolation plans, mock leakage checks, database cleanup reviews, CI environment comparisons, stabilization roadmaps, QA tickets, verification plans, and paste-ready prompts for Cursor, Claude Code, Codex CLI, OpenCode, Replit, and ChatGPT Agents. It is designed to restore trust in test suites by fixing real causes instead of hiding failures with retries, arbitrary sleeps, skipped tests, or weakened assertions.
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
2 installs
Downloaded by developers to date
Free forever
No account required to browse
Trust & safety
Security scanned
Verified clean 4 months ago