- Home
- Skills
- Code Quality & Review
- agent-pr-verifier
More screenshots
Works with the AI tools you already use
Agent pr Verifier
Your agent spent 40 minutes on a feature and closed with "All 23 tests passing, lint clean, PR ready." You look at the diff — it looks right. You merg
Free
agent-pr-verifier
Example session with this skill installed
** "Agent just finished the refund-retry feature and says all tests pass. Verify before I merge PR #482.
- Read your context and instructions
- Compiled the agent-pr-verifier
- Generated the document
**
EVIDENCE REPORT — PR #482 (refund-retry)
Claims found: 4 | Re-runnable: 3 | Unverifiable: 1
=== VERIFIED (re-ran from clean state, HEAD a91c3f2) ===
pytest tests/ -q → 23 passed (exit 0)
ruff check . → 0 findings (exit 0)
npm run build → success (exit 0)
=== VACUOUS (checks that cannot fail) ===
test_refund_retry_calls_gateway
Deletion probe: stubbed RefundGateway.send() → test STILL PASSES
Reason: asserts on mock return value, never touches implementation
Proof probe: git stash && pytest tests/test_refund.py::test_refund_retry_calls_gateway
→ 1 of 23 tests verified nothing
=== GAPS (requirement clauses with no test) ===
Requirement: "retry with exponential backoff, max 3 attempts"
Tests exercise retry ONCE. No test covers attempt 2 or 3, or the
give-up path after max retries.
R
agent-pr-verifier.pdf
PDF · document
Example file from a real run - the skill writes it into your workspace.
Connects securely to your tools. The creator never sees your data.
About this skill
"Tests pass" from an AI agent is a report of intent, not an observation. This skill turns the reviewer's side into a mechanical process — and finds the checks that cannot fail.
The Problem
Your agent spent 40 minutes on a feature and closed with "All 23 tests passing, lint clean, PR ready." You look at the diff — it looks right. You merge. Three days later production breaks in exactly the way the tests were supposedly covering. The post-mortem finds the trap: the agent wrote the implementation AND the tests from the same wrong reading of the requirement, so they agree perfectly — on the wrong behavior. Or worse: the "coverage" test passes whether or not the feature exists, because it asserts a mock, not the code. You had 23 green checks and zero evidence.
What You Get
- Claim inventory extraction — pulls every checkable claim from the agent's final report ("tests pass", "lint clean", "verified in browser") and classifies each as re-runnable, stale, or unverifiable.
- Independent re-run — executes each claim's check from a clean state against current HEAD, with exact commands and exit codes recorded — never the agent's terminal history.
- The deletion probe — temporarily deletes or stubs the code each test claims to cover and confirms the test FAILS. A test that still passes when the feature is deleted is vacuous, and gets flagged with a one-command proof probe.
- The zero-findings probe — plants a known-bad artifact into scanners that reported "0 issues" and confirms the scanner catches it. A scanner with a matching problem reads as "clean" forever.
- Shared-misunderstanding detection — flags tests written in the same PR as the implementation whose assertions restate the implementation instead of the original requirement, and requirement clauses no test exercises.
- One-screen evidence report — VERIFIED / VACUOUS / GAPS sections with a merge recommendation (SAFE / CONDITIONAL / BLOCKED) and the single next command that provides the strongest remaining evidence.
Why This Beats Prompting It Yourself
Ask your agent to "double-check your work" and it re-runs its own test command and says "all good" — the same instrument that reported success the first time, re-run by the party with the incentive to see green. It will never delete the code under test to see if the test notices, because that requires suspecting its own work. This skill mechanically distrusts the report: re-runs from clean state, probes for vacuity, and compares tests against the requirement text instead of the implementation. It is the reviewer-side complement to agent self-verification — designed for the case where the "author" is a session that no longer exists.
Example Session
You: "Agent just finished the refund-retry feature and says all tests pass. Verify before I merge PR #482."
Agent:
EVIDENCE REPORT — PR #482 (refund-retry)
Claims found: 4 | Re-runnable: 3 | Unverifiable: 1
=== VERIFIED (re-ran from clean state, HEAD a91c3f2) ===
pytest tests/ -q → 23 passed (exit 0)
ruff check . → 0 findings (exit 0)
npm run build → success (exit 0)
=== VACUOUS (checks that cannot fail) ===
test_refund_retry_calls_gateway
Deletion probe: stubbed RefundGateway.send() → test STILL PASSES
Reason: asserts on mock return value, never touches implementation
Proof probe: git stash && pytest tests/test_refund.py::test_refund_retry_calls_gateway
→ 1 of 23 tests verified nothing
=== GAPS (requirement clauses with no test) ===
Requirement: "retry with exponential backoff, max 3 attempts"
Tests exercise retry ONCE. No test covers attempt 2 or 3, or the
give-up path after max retries.
Requirement: "idempotent by event ID"
No test submits the same event ID twice.
Claim "no breaking changes": UNVERIFIABLE — no API contract check exists
Strongest next command: npm run api-diff --Against=main
=== SHARED-MISUNDERSTANDING CHECK ===
test_refund_amount_matches_request
Asserts: refund.amount == request.amount
Requirement says: "refund amount EXCLUDING original transaction fee"
Implementation and test agree with each other — and both disagree
with the requirement. HIGH severity.
MERGE RECOMMENDATION: BLOCKED
1 vacuous test, 2 untested requirement clauses, 1 requirement divergence
Fix path: assert on gateway side effects (not mock), add attempt-3 and
duplicate-event-ID tests, resolve fee-handling against the requirement.
Use Cases
- Verifying an agent-authored feature PR before merging to a shared branch.
- Orchestrator sessions validating executor session claims ("tests pass" → prove it).
- Auditing a large agent-generated codebase to find which parts have real evidence and which are Ferrari shells.
- CI augmentation: run the deletion probe on the 10 most critical tests as a scheduled job.
- Solo developers using agents as their only implementer — this is the second pair of eyes.
Known Limitations
The deletion probe mutates the working tree and must run on stashed/clean state — it is not safe on uncommitted work-in-progress. It verifies evidence quality, not code style or architecture. Detecting shared misunderstanding requires the original requirement text to be available and specific; vague tickets weaken the check. Runtime/concurrency bugs that no static probe exposes remain the reviewer's judgment call.
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
1 install
Downloaded by developers to date
Free forever
No account required to browse
Trust & safety
Security scanned
Verified clean 11 days ago
- Free to download with an account