Works with the AI tools you already use

    Claude CodeClaude CodeCursorCursorCodex CLICodex CLIGitHub CopilotGitHub CopilotGemini CLIGemini CLI+20 more

    skill regression harness

    1

    Generate a cost-tiered regression testing harness to catch silent behavioral breakages after editing SKILL.md files.

    $16

    Secure checkout via Stripe

    30-day refund guarantee

    Converts to your local currency at checkout

    0 installsSecurity scanned
    skill regression harness

    skill regression harness

    Example session with this skill installed

    I just updated the Iron Rules in skills/billing-agent.md to be stricter about data privacy. Generate a regression harness to make sure I didn't break the existing refund trigger.

    • Read your context and instructions
    • Compiled the regression harness

    I have analyzed skills/billing-agent.md and designed a 35-case regression battery. This harness is configured to use Claude Haiku 4.5 for bulk testing and Claude Opus 5 for escalation and confirmation sampling.

    Regression Harness for Billing Agent

    Estimated Run Cost:

    • Bulk Pass (Haiku 4.5): $0.07
    • Judging & Escalation (Opus 5): $0.12
    • Total Projected Cost: ~$0.19 per run

    The following files have been generated in tests/regression/billing-agent/:

    1. config.yaml: Defines model routing, API environment variables, and a $0.50 safety ceiling.
    2. battery.yaml: 15 trigger-positive cases (refunds), 10 trigger-negative cases (general billing), and 10 compliance cases mapped to your new privacy rules.
    3. run_battery.py: Script to generate responses from the skill under test.
    4. score_responses.py: Rubric-based grader that performs cost-tiered escalation.
    5. README.md: Setup instructions and usage guide.
    
    # config.yaml snippet
    models:
      bulk: "claude-3-haiku-20240307"
      frontier: "claude-3-opus-20240229"
    routing:
      confirmation_sample_rate: 0.1
      escalate_all_failures: true
    costs:
      max_projected_cost_usd: 0.50
    

    Next steps

    1. Review the test cases in battery.yaml to ensure they reflect your expected "strict" privacy behaviors.
    2. Set your ANTHROPIC_API_KEY environment variable.
    3. Run python run_battery.py --estimate-only to verify the setup before making live API calls.

    Connects securely to your tools. The creator never sees your data.

    What you get

    Automate behavior verification for agent skill updates in CI pipelines.Reduce API testing costs using tiered model routing and selective escalation.Identify if a skill failure is a regression or an intentional instruction change.Enforce rubric-based compliance scoring for mission-critical agent rules.

    About this skill

    Editing a shipped SKILL.md is like editing a live program without a compiler. A single wording change can silently break trigger accuracy or cause the agent to ignore its own Iron Rules. This skill solves the problem of expensive, manual re-testing by generating a portable, automated regression harness for your specific skill files. It provides the scripts you need to catch behavior drift before merging edits, using a cost-efficient model routing strategy that reserves frontier models only for critical verifications.

    What it does

    • Battery generation creates a structured set of 25 to 40 test cases covering triggers, compliance, and adversarial inputs.
    • Cost-tiered routing runs bulk tests on cheaper models like Haiku 4.5 and escalates failures to Opus 5 for confirmation.
    • Rubric-based scoring extracts specific rules from your skill text to judge outputs with traceable citations instead of vague scores.
    • Regression triage compares failures against your recent edits to distinguish between silent breakages and intended behavior changes.
    • Budget enforcement calculates and enforces API cost ceilings for both response generation and grading passes.

    How it works

    1. Analyze target skill to identify all trigger surfaces and compliance rules within your SKILL.md.
    2. Design prompt battery including positive/negative triggers and rubric-linked adversarial prompts.
    3. Generate harness code including config.yaml, test scripts, and model routing logic for your local environment.
    4. Execute local scripts to run the battery, score responses, and generate a detailed scorecard.md report.

    Frameworks & tools

    This skill generates Python-based scripts (run_battery.py, score_responses.py) that interface with the Anthropic API. It uses YAML for configuration and Markdown for reporting.

    Why this beats prompting it yourself

    Manual testing is inconsistent and expensive if you use frontier models for every check. This skill automates the "bulk-test, frontier-escalate" workflow, ensuring you only pay for high-end reasoning when a cheaper model flags a potential issue.

    Use cases

    • Validating a PR that modifies core agent instructions in a production skill.
    • Setting up CI/CD gates to prevent behavioral regressions in a skill portfolio.
    • Auditing third-party skills for rule compliance and adversarial robustness.
    • Reducing API spend during the maintenance phase of an agent's lifecycle.

    Known limitations

    This tool is for post-ship regression testing, not interactive first-draft authoring. It requires local Python execution and valid Anthropic API credentials.

    How to install

    Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.

    ~30 seconds
    1. 1

      Download the ZIP

      Free skills download straight away. Paid skills unlock right after purchase.

    2. 2

      Unzip into your skills folder

      Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.

    3. 3

      Ask your agent to use it

      Restart the agent if it was already running. It picks the skill up automatically - no config needed.

    Skills folder by agent

    Click the path to copy it. Create the folder if it does not exist yet.

    Reviews

    No reviews yet

    Be one of the first to try it. Every listed skill passes our trust checks below.

    Security scanned

    Passed our 8-point scan before listing

    Fresh listing

    Recently published to Agensi

    30-day refund

    Not a fit? Get your money back

    Trust & safety

    Security scanned

    Verified clean 1 day ago

    • Passed all security checks, Safe to install

    Listed1 day ago

    What's inside

    Frequently Asked Questions