Works with the AI tools you already use

    Claude CodeClaude CodeCursorCursorCodex CLICodex CLIGitHub CopilotGitHub CopilotGemini CLIGemini CLI+20 more

    trading agent evaluation suite

    1

    A deterministic evaluation suite for scoring trading agent logic, math, and policy adherence against reviewed baselines.

    $15

    Secure checkout via Stripe

    30-day refund guarantee

    Converts to your local currency at checkout

    0 installsSecurity scanned
    trading agent evaluation suite

    trading agent evaluation suite

    Example session with this skill installed

    Analyze the results in evaluation_logs.json for the 'aggressive-growth' agent version 2.1. Compare it against the gold-standard fixture for position sizing and news-drift scenarios.

    • Read your context and instructions
    • Compiled the trading agent evaluation

    Analysis complete. I have processed 50 scenarios for agent 'aggressive-growth' v2.1. The evaluation detected a 12% regression in numeric precision for arithmetic-heavy tasks compared to v2.0, while policy adherence for unauthorized tools remained at 100%.

    Trading Agent Evaluation Scorecard

    Metric CategoryPass RateMean Absolute ErrorCritical FailuresNotes
    Position Sizing82%0.045 BTC4Errors in Kelly Criterion calc
    News Drift90%N/A22 instances of 'misleading news' bias
    Tool Policy100%N/A0No unauthorized tool calls detected
    Arithmetic76%12.40 USD7Regressed vs v2.0 baseline
    Latency (Input)N/A450ms0As reported in source logs

    Failure Taxonomy

    • Precision Error (7): Agent rounded down prematurely on high-frequency sizing prompts.
    • Sentiment Bias (2): Agent ignored 'held-out' bearish data in favor of 'misleading bullish' news injection.
    • Logic Regress (4): Version 2.1 failed cases that Version 2.0 passed in the 'synthetic-math-fixture'.

    Next steps

    1. Review the precision_error_logs to adjust the absolute tolerance settings in analyze.py.
    2. Re-run the 'News Drift' scenarios using the updated adversarial_v3 test suite.
    3. Investigate the arithmetic regression before promoting v2.1 to the paper-trading environment.

    Connects securely to your tools. The creator never sees your data.

    What you get

    Score agent tool choices against strict policy expectationsValidate numeric position sizing logic with defined error tolerancesDetect regressions between model versions using fixed test fixturesIdentify agent susceptibility to adversarial or misleading market data

    About this skill

    The problem

    Developers lack a structured way to verify if their trading agents make sound financial decisions or follow strict policy constraints before risking capital. Manual review of execution logs is slow and prone to overlooking subtle numeric errors or unauthorized tool usage.

    What it does

    • Evaluates trading agent outputs against human-reviewed exact or numeric expectations.
    • Identifies failures in arithmetic, position sizing, and response to misleading market news.
    • Categorizes failures into a specific taxonomy for regression testing and debugging.
    • Flags unauthorized tool calls that bypass isolation boundaries.
    • Produces a comparative scorecard between different model versions or configurations.

    Why this beats prompting it yourself

    General LLM prompting often suffers from 'fluency bias,' where an agent explains its reasoning well despite getting the math or policy wrong. This skill enforces strict, deterministic scoring of structured data, ensuring that citations and numeric tolerances are validated rather than just sounding plausible.

    Use cases

    • Regression testing agent performance after updating underlying LLM versions.
    • Benchmarking position sizing logic against a set of verified arithmetic fixtures.
    • Detecting policy violations in agents exposed to adversarial 'misleading news' prompts.
    • Generating a failure taxonomy for audits before moving from paper trading to production.

    Known limitations

    Does not include a model runner, broker adapter, or semantic grading engine. Requires manual mapping of external logs to the specific input contract.

    How to install

    Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.

    ~30 seconds
    1. 1

      Download the ZIP

      Free skills download straight away. Paid skills unlock right after purchase.

    2. 2

      Unzip into your skills folder

      Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.

    3. 3

      Ask your agent to use it

      Restart the agent if it was already running. It picks the skill up automatically - no config needed.

    Skills folder by agent

    Click the path to copy it. Create the folder if it does not exist yet.

    Reviews

    No reviews yet

    Be one of the first to try it. Every listed skill passes our trust checks below.

    Security scanned

    Passed our 8-point scan before listing

    Fresh listing

    Recently published to Agensi

    30-day refund

    Not a fit? Get your money back

    Trust & safety

    Security scanned

    Verified clean 13 days ago

    • Passed all security checks, Safe to install

    Listed13 days ago

    What's inside

    Frequently Asked Questions