- Home
- Skills
- Testing & Debugging
- agent-skill-regression-tester
More screenshots
Works with the AI tools you already use
Agent Skill Regression Tester
You update your `code-review-checklist` SKILL.md to add a new check. It passes your manual test — "looks good.
Free
agent-skill-regression-tester
Example session with this skill installed
"Build regression tests for our commit-message-generator skill before we update it"
- Read your context and instructions
- Compiled the agent--regression-tester
=== Running tests for commit-message-generator ===
--- Deterministic Checks ---
Scenario CS-001: Single-file feat change → PASS
Format: type(scope): description ✓
Type: feat ✓
Length: 48 chars (10-72) ✓
No trailing period ✓
Scenario CS-002: Multi-file refactor → PASS
Scenario CS-003: DELETE endpoint → PASS (no retry logic added)
Scenario CS-004: Breaking change → FAIL
Expected: type includes "!" or footer added
Got: Standard feat commit with no breaking indicator
--- Behavioral Snapshot Tests ---
CS-001: unchanged ✓
CS-002: equivalent (different wording, same meaning) ✓
CS-005: improved (better scope detection) ✓
=== Results ===
Deterministic: 3/4 PASS
Behavioral: 3/3 PASS
OVERALL: FAIL — fix CS-003 before shipping
Connects securely to your tools. The creator never sees your data.
About this skill
The Problem
You update your code-review-checklist SKILL.md to add a new check. It passes your manual test — "looks good." Two weeks later, a teammate reports that the agent now rejects valid code because the new instruction conflicts with an existing constraint you forgot about. SKILL.md files are prompts, not code. They can't be unit-tested, type-checked, or linted. Every edit is a bet that the new instructions produce better behavior — with no way to prove it until it breaks in production.
What You Get
- Extract behavioral contracts from SKILL.md files — identify explicit constraints, output format requirements, trigger conditions, and decision rules, then convert them into testable assertions
- Design test scenarios covering happy path, edge cases, error cases, and regression cases — each with explicit input, expected behavioral contract, and pass/fail criteria
- Build deterministic validation scripts — regex pattern checks, format validators, and constraint enforcers that run in under 1 second without LLM involvement (80% of regressions caught here)
- Create behavioral snapshot tests — run the skill against known inputs with both old and new versions, compare outputs using structured LLM-as-judge evaluation, and classify differences as unchanged/improved/equivalent/regressed
- Assemble a CI-ready test harness —
tests/directory withscenarios.yaml,deterministic_checks.py,snapshot_baselines/,run_tests.sh, and GitHub Actions integration - Detect semantic drift across skill versions — flag when a single word change in SKILL.md cascades into different behavior on edge cases you never manually tested
Why This Beats Prompting It Yourself
You can't test a SKILL.md by asking an agent "does this still work?" — the agent doesn't know what "work" meant in the previous version. This skill builds a regression harness that defines exactly what the skill should do (behavioral contract), tests it against known inputs (scenarios), and compares outputs between versions (snapshots). The deterministic checks catch format violations instantly; the behavioral snapshots catch semantic drift. You get a pass/fail report just like any other code test.
Example Session
You: "Build regression tests for our commit-message-generator skill before we update it"
Agent:
=== Running tests for commit-message-generator ===
--- Deterministic Checks ---
Scenario CS-001: Single-file feat change → PASS
Format: type(scope): description ✓
Type: feat ✓
Length: 48 chars (10-72) ✓
No trailing period ✓
Scenario CS-002: Multi-file refactor → PASS
Scenario CS-003: DELETE endpoint → PASS (no retry logic added)
Scenario CS-004: Breaking change → FAIL
Expected: type includes "!" or footer added
Got: Standard feat commit with no breaking indicator
--- Behavioral Snapshot Tests ---
CS-001: unchanged ✓
CS-002: equivalent (different wording, same meaning) ✓
CS-005: improved (better scope detection) ✓
=== Results ===
Deterministic: 3/4 PASS
Behavioral: 3/3 PASS
OVERALL: FAIL — fix CS-003 before shipping
Use Cases
- Pre-update regression testing for shared SKILL.md libraries used by multiple team members
- Quality assurance before publishing skills to Agensi marketplace — verify cross-harness compatibility
- CI integration for teams that treat agent skills as versioned, testable artifacts
- Onboarding verification — confirm new team members are using the correct skill version
- Evaluating third-party skills — run their behavioral contract against your expected outcomes
- Capturing regression evidence for contest submissions or compliance documentation
Known Limitations
Behavioral snapshot tests have inherent LLM variance — use "equivalent" classification liberally for wording differences that don't affect outcomes. Test scenarios age as codebases evolve and need quarterly rotation. Model behavior changes (e.g., Claude 3.5 → 4.0) can cause failures that aren't skill regressions — flag these separately. The harness is read-only analysis; it doesn't execute the skill against live systems.
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
13 installs
Downloaded by developers to date
Free forever
No account required to browse
Trust & safety
Security scanned
Verified clean 2 months ago
- Free to download with an account