- Home
- Skills
- Testing & Debugging
- flaky-test-detector
More screenshots
Works with the AI tools you already use
Flaky Test Detector
Stop wasting 30 minutes per CI rerun on tests that pass locally but fail randomly in CI.
Free
flaky-test-detector
Example session with this skill installed
Our test_payment_retry fails in CI about 1 in 5 runs but always passes locally. Here's the test file.
- Read your context and instructions
- Compiled the flaky-test-detector
Detected 3 potential root causes from code scan: (1) time.sleep(2) on line 47 — timing-dependent assertion, (2) @pytest.fixture(scope='class') creates shared state between test_retry_success and test_retry_exhausted, (3) unmocked HTTP call to payments.stripe.com on line 62. Ran isolation test: passes alone, fails when paired with test_payment_webhook. Root cause:
shared state — the class-level fixture retains payment state across tests. Fix: convert to function-scoped fixture. Applied fix, ran 50-iteration validation: 0/50 failures. Quarantined original with tracking issue #247 while verifying.
Connects securely to your tools. The creator never sees your data.
About this skill
Stop wasting 30 minutes per CI rerun on tests that pass locally but fail randomly in CI.
The Problem
Your CI pipeline red-lights every third build on a test that passes 90% of the time. You re-run it and it passes. Your team starts ignoring CI failures because "it's probably just flaky." Meanwhile, a real regression sneaks through because nobody trusts the test suite anymore. The root cause is always non-determinism — timing races, shared state, unmocked APIs, or environment drift — but finding which one takes tedious manual investigation every single time.
What You Get
- Automated flakiness confirmation — rerun loop scripts (50 iterations) that measure exact failure rates for pytest, Jest, Go, and Maven test suites
- Root cause classification engine — maps symptoms to 6 diagnosed categories: timing/race conditions, shared state, random data, external dependencies, environment differences, and teardown failures
- Isolation testing workflow — scripts to run tests alone vs. with neighbors, in shuffled order, and under parallel stress to pinpoint whether failures are order-dependent or resource-contended
- Framework-specific fix patterns — concrete before/after code for each root cause: replacing
sleep()with poll-with-timeout, replacing class-level fixtures with per-test fixtures, seeding random generators, mocking external calls - Quarantine system — generates
@pytest.mark.flakymarkers,describe.skipblocks, or@Disabledannotations with tracking issues, plus aFLAKY_TESTS.mdregistry with owner and quarantine date - CI prevention pipeline — GitHub Actions YAML with
pytest-rerunfailures,pytest-flakefinder, and flaky-detection steps that parse JUnit XML to flag tests with mixed pass/fail outcomes
Why This Beats Prompting It Yourself
When you ask an AI agent "fix my flaky test," it typically adds a longer sleep() timeout or wraps the assertion in a retry — treating symptoms instead of diagnosing the actual source of non-determinism. This skill forces a systematic 5-phase workflow: confirm → diagnose → fix → quarantine → prevent. Without it, agents gravitate toward the easiest-looking fix (more sleep, bigger timeout) rather than the correct one (eliminating the shared state or mocking the external dependency).
Example Session
You: Our
test_payment_retryfails in CI about 1 in 5 runs but always passes locally. Here's the test file.
Agent: Detected 3 potential root causes from code scan: (1)
time.sleep(2)on line 47 — timing-dependent assertion, (2)@pytest.fixture(scope='class')creates shared state betweentest_retry_successandtest_retry_exhausted, (3) unmocked HTTP call topayments.stripe.comon line 62. Ran isolation test: passes alone, fails when paired withtest_payment_webhook. Root cause: shared state — the class-level fixture retains payment state across tests. Fix: convert to function-scoped fixture. Applied fix, ran 50-iteration validation: 0/50 failures. Quarantined original with tracking issue #247 while verifying.
Use Cases
- Diagnosing why a test fails only on GitHub Actions Ubuntu runners but passes on macOS locally
- Systematically identifying which tests in a 500-test suite are flaky before a major release
- Adding a nightly flaky-test detection job that surfaces newly introduced non-determinism
- Setting up CI quarantine rules that block merging tests with >5% flakiness rate
- Investigating a test that fails when run with
pytest-xdist -n autobut passes single-threaded
Known Limitations
Flaky tests caused by actual hardware differences (SSD vs. HDD timing, real network latency) cannot be fixed by code changes alone — they require environment standardization. This skill focuses on code-level non-determinism, not infrastructure-level flakiness.
Tags: testing ci-cd flaky-tests pytest jest debugging code-quality
Version: 1.0.0
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
2 installs
Downloaded by developers to date
Free forever
No account required to browse
Trust & safety
Security scanned
Verified clean 4 months ago
- Free to download with an account