- Home
- Skills
- Data & Databases
- A/B Test Clinic
A/B Test Clinic
Analyze A/B tests rigorously: SRM check, significance, confidence intervals, guardrails and a ship/no-ship call.
$12
Works with the AI tools you already use
A/B Test Clinic
Example session with this skill installed
Our checkout test finished. Control: 5,000 users, 480 conversions. Treatment: 5,100 users, 530 conversions. Planned 50/50 split. Should we ship?
- Read your context and instructions
- Compiled the a b test
- Generated the document
Illustrative example on synthetic data. It follows this skill's ab_test_report_template.md.
A/B Test Report: New checkout flow
Hypothesis
The simplified checkout increases purchase conversion.
Design
Randomised by user, 50/50 split. Primary metric: purchase conversion. Two-sided α = 0.05.
Results
Sample Ratio Mismatch (SRM) check
Expected 5,050 / 5,050; observed 5,000 / 5,100. χ² = 0.99, p = 0.32. No SRM. Randomisation looks healthy.
Primary metric
| Variant | Users | Conversions | Rate |
|---|---|---|---|
| Control | 5,000 | 480 | 9.60% |
| Treatment | 5,100 | 530 | 10.39% |
- Absolute lift: +0.79 pp (relative +8.3%)
- Two-proportion z-test: z = 1.33, p = 0.18
- 95% CI for the difference: −0.38 pp to +1.96 pp (includes zero)
Guardrail metrics
No guardrail data was provided. Add refund rate and average order value before any final call.
Decision
Do not ship yet. Extend the test. The direction is positive, but the result is not statistically significant, and the confidence interval still allows a small negative effect. To detect an 8% relative lift with 80% power, you need about
23,900 users per arm. At the current traffic, that is roughly 4–5 more weeks. Do not stop early when the p-value happens to dip below 0.05 (peeking).
a-b-test-clinic.pdf
PDF · document
Example file from a real run - the skill writes it into your workspace.
Connects securely to your tools. The creator never sees your data.
What you get
About this skill
A/B Test Clinic turns experiment results into a defensible decision. Your agent confirms the test design, checks for sample ratio mismatch (SRM) before looking at results, computes per-variant rates with confidence intervals, and runs the right significance test (a two-proportion z-test for rates, Welch's t-test for means). It checks guardrail metrics and finishes with a ship, no-ship or extend recommendation. It also helps with sample-size planning before launch. The bundled analyzer implements the statistics directly and needs no stats package.
Use it when
- An experiment ended and the team needs a ship decision
- Results "look positive" but nobody is sure they are significant
- A test has run for weeks without a clear winner
- You are planning the sample size for a new test
What's inside (installs as one folder, ab-test-analysis/):
ab-test-analysis/LICENSEab-test-analysis/SKILL.mdab-test-analysis/assets/ab_test_report_template.mdab-test-analysis/references/ab_test_design_guide.mdab-test-analysis/scripts/ab_test_analyzer.py
Please read before buying: this is packaged from my free, MIT-licensed open-source library (https://github.com/nimrodfisher/data-analytics-skills). The same content is available there for free. You are paying for a ready-to-install package. It is a one-time purchase, sold as-is, with no support, updates or maintenance included. The MIT license is included.
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
Fresh listing
Recently published to Agensi
30-day refund
Not a fit? Get your money back
Trust & safety
Security scanned
Verified clean 1 day ago
- Passed all security checks, Safe to install