More screenshots

    Works with the AI tools you already use

    Claude CodeClaude CodeCursorCursorCodex CLICodex CLIGitHub CopilotGitHub CopilotGemini CLIGemini CLI+20 more

    Agent Skill Regression Tester

    2

    You update your `code-review-checklist` SKILL.md to add a new check. It passes your manual test — "looks good.

    Free

    13 installsSecurity scanned
    agent-skill-regression-tester

    agent-skill-regression-tester

    Example session with this skill installed

    "Build regression tests for our commit-message-generator skill before we update it"

    • Read your context and instructions
    • Compiled the agent--regression-tester

    === Running tests for commit-message-generator ===

    --- Deterministic Checks ---
    Scenario CS-001: Single-file feat change → PASS
    Format: type(scope): description ✓
    Type: feat ✓
    Length: 48 chars (10-72) ✓
    No trailing period ✓
    Scenario CS-002: Multi-file refactor → PASS
    Scenario CS-003: DELETE endpoint → PASS (no retry logic added)
    Scenario CS-004: Breaking change → FAIL
    Expected: type includes "!" or footer added
    Got: Standard feat commit with no breaking indicator

    --- Behavioral Snapshot Tests ---
    CS-001: unchanged ✓
    CS-002: equivalent (different wording, same meaning) ✓
    CS-005: improved (better scope detection) ✓

    === Results ===
    Deterministic: 3/4 PASS
    Behavioral: 3/3 PASS
    OVERALL: FAIL — fix CS-003 before shipping

    Connects securely to your tools. The creator never sees your data.

    About this skill

    The Problem

    You update your code-review-checklist SKILL.md to add a new check. It passes your manual test — "looks good." Two weeks later, a teammate reports that the agent now rejects valid code because the new instruction conflicts with an existing constraint you forgot about. SKILL.md files are prompts, not code. They can't be unit-tested, type-checked, or linted. Every edit is a bet that the new instructions produce better behavior — with no way to prove it until it breaks in production.

    What You Get

    • Extract behavioral contracts from SKILL.md files — identify explicit constraints, output format requirements, trigger conditions, and decision rules, then convert them into testable assertions
    • Design test scenarios covering happy path, edge cases, error cases, and regression cases — each with explicit input, expected behavioral contract, and pass/fail criteria
    • Build deterministic validation scripts — regex pattern checks, format validators, and constraint enforcers that run in under 1 second without LLM involvement (80% of regressions caught here)
    • Create behavioral snapshot tests — run the skill against known inputs with both old and new versions, compare outputs using structured LLM-as-judge evaluation, and classify differences as unchanged/improved/equivalent/regressed
    • Assemble a CI-ready test harness — tests/ directory with scenarios.yaml, deterministic_checks.py, snapshot_baselines/, run_tests.sh, and GitHub Actions integration
    • Detect semantic drift across skill versions — flag when a single word change in SKILL.md cascades into different behavior on edge cases you never manually tested

    Why This Beats Prompting It Yourself

    You can't test a SKILL.md by asking an agent "does this still work?" — the agent doesn't know what "work" meant in the previous version. This skill builds a regression harness that defines exactly what the skill should do (behavioral contract), tests it against known inputs (scenarios), and compares outputs between versions (snapshots). The deterministic checks catch format violations instantly; the behavioral snapshots catch semantic drift. You get a pass/fail report just like any other code test.

    Example Session

    You: "Build regression tests for our commit-message-generator skill before we update it"

    Agent:

    === Running tests for commit-message-generator ===
    
    --- Deterministic Checks ---
    Scenario CS-001: Single-file feat change → PASS
      Format: type(scope): description ✓
      Type: feat ✓
      Length: 48 chars (10-72) ✓
      No trailing period ✓
    Scenario CS-002: Multi-file refactor → PASS
    Scenario CS-003: DELETE endpoint → PASS (no retry logic added)
    Scenario CS-004: Breaking change → FAIL
      Expected: type includes "!" or footer added
      Got: Standard feat commit with no breaking indicator
    
    --- Behavioral Snapshot Tests ---
    CS-001: unchanged ✓
    CS-002: equivalent (different wording, same meaning) ✓
    CS-005: improved (better scope detection) ✓
    
    === Results ===
    Deterministic: 3/4 PASS
    Behavioral: 3/3 PASS
    OVERALL: FAIL — fix CS-003 before shipping
    

    Use Cases

    • Pre-update regression testing for shared SKILL.md libraries used by multiple team members
    • Quality assurance before publishing skills to Agensi marketplace — verify cross-harness compatibility
    • CI integration for teams that treat agent skills as versioned, testable artifacts
    • Onboarding verification — confirm new team members are using the correct skill version
    • Evaluating third-party skills — run their behavioral contract against your expected outcomes
    • Capturing regression evidence for contest submissions or compliance documentation

    Known Limitations

    Behavioral snapshot tests have inherent LLM variance — use "equivalent" classification liberally for wording differences that don't affect outcomes. Test scenarios age as codebases evolve and need quarterly rotation. Model behavior changes (e.g., Claude 3.5 → 4.0) can cause failures that aren't skill regressions — flag these separately. The harness is read-only analysis; it doesn't execute the skill against live systems.

    How to install

    Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.

    ~30 seconds
    1. 1

      Download the ZIP

      Free skills download straight away. Paid skills unlock right after purchase.

    2. 2

      Unzip into your skills folder

      Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.

    3. 3

      Ask your agent to use it

      Restart the agent if it was already running. It picks the skill up automatically - no config needed.

    Skills folder by agent

    Click the path to copy it. Create the folder if it does not exist yet.

    Reviews

    No reviews yet

    Be one of the first to try it. Every listed skill passes our trust checks below.

    Security scanned

    Passed our 8-point scan before listing

    13 installs

    Downloaded by developers to date

    Free forever

    No account required to browse

    Trust & safety

    Security scanned

    Verified clean 2 months ago

    • Free to download with an account

    Listed2 months ago
    Updated9 days ago

    What's inside

    Frequently Asked Questions