More screenshots

    Works with the AI tools you already use

    Claude CodeClaude CodeCursorCursorCodex CLICodex CLIGitHub CopilotGitHub CopilotGemini CLIGemini CLI+20 more

    Agent Token Budget Manager

    2

    You tell your agent to update one API endpoint. Five minutes later, the token counter shows 180K consumed.

    Free

    6 installsSecurity scanned
    agent-token-budget-manager

    agent-token-budget-manager

    Example session with this skill installed

    Fix the login button alignment with a 5K token budget.

    • Read your context and instructions
    • Compiled the agent-token-budget-manager
    • Generated the document

    Budget allocated: 5K tokens (complexity 1)

    • Research: 1K (read the login component)
    • Planning: 0.5K (identify CSS fix)
    • Implementation: 2.5K (edit the component)
    • Verification: 1K (quick visual check)

    [Agent reads the file — 800 tokens. Identifies missing flexbox property.
    Edits file — 1,200 tokens. Verifies — 600 tokens.]

    📊 Token Budget Report: "Fix login button alignment"
    Budget: 5K | Actual: 2,600 tokens (52% used)
    Research: 800 tokens (17%)
    Planning: 400 tokens (11%)
    Implementation: 1,200 tokens (50%)
    Verification: 600 tokens (22%)

    Compression: None needed
    Efficiency: 94% (high — focused on the task)
    Completed: Full fix within budget

    
    Without budget: Agent reads 3 related files (5K), explores the CSS framework (3K), checks browser compatibility (2K), writes a detailed explanation (4K), implements the fix (1.5K), refactors t
    

    agent-token-budget-manager.pdf

    PDF · document

    Generated

    Example file from a real run - the skill writes it into your workspace.

    Connects securely to your tools. The creator never sees your data.

    About this skill

    The Problem

    You tell your agent to update one API endpoint. Five minutes later, the token counter shows 180K consumed. The agent read the entire codebase (not just the endpoint file), explored three unrelated modules "for context," generated a detailed plan with four alternative approaches, wrote the implementation, then refactored two files that weren't part of the task "while it was in there."

    This is token bleed — the silent cost of agents that don't budget. Unlike a crash or API error (which you notice immediately), token bleed just looks like the agent is working hard. It is. But 80% of its tokens went to work that didn't move the needle. On paid API plans ($20+/month per agent), this translates directly to wasted money. On a team of 10 agents, one wasted 150K-token task per day costs $450/month in API fees.

    What You Get

    • Three-layer budget model — Session Budget (default 100K tokens, hard ceiling), Task Budget (default 30K tokens per task), Tool Call Budget (default 5K per call output)
    • Four-phase task allocation — Research (20%), Planning (10%), Implementation (50%), Verification (20%) with hard ceilings per phase
    • Progressive compression at three thresholds — Level 1 (60-75%): trim verbose output, skip reasoning traces. Level 2 (75-90%): narrow scope, drop edge cases. Level 3 (90-100%): emergency mode, deliver minimum viable implementation, flag deferred work with TODO markers
    • Budget escalation protocol — when the task is genuinely underestimated, the agent reports current state, remaining work, requested additional budget, and justification — instead of silently overspending
    • Token attribution reports — after each task: budget vs. actual, per-phase breakdown, compression levels activated, deferred items, efficiency score (productive tokens ÷ total tokens)
    • Cross-session learning — .agent-token-budget.json stores budget patterns per task type, automatically adjusting allocations based on historical accuracy

    Why This Beats Prompting It Yourself

    "Be concise" is a vague instruction the agent may or may not follow. The budget system creates hard ceilings with automatic escalation: when a task hits 60% of its allocation, the agent switches from verbose explanations to concise summaries. At 90%, it delivers the minimum viable implementation and flags deferred work. At 100%, it stops and asks for more budget instead of silently burning tokens. The attribution report shows exactly where tokens went so you can tune the budget model over time.

    Example Session

    User: "Fix the button alignment on the login page."

    Agent (with budget enabled):

    Budget allocated: 5K tokens (complexity 1)
      - Research: 1K (read the login component)
      - Planning: 0.5K (identify CSS fix)
      - Implementation: 2.5K (edit the component)
      - Verification: 1K (quick visual check)
    
    [Agent reads the file — 800 tokens. Identifies missing flexbox property.
     Edits file — 1,200 tokens. Verifies — 600 tokens.]
    
    📊 Token Budget Report: "Fix login button alignment"
       Budget: 5K | Actual: 2,600 tokens (52% used)
       Research:       800 tokens (17%)
       Planning:       400 tokens (11%)
       Implementation: 1,200 tokens (50%)
       Verification:   600 tokens (22%)
       
       Compression: None needed
       Efficiency: 94% (high — focused on the task)
       Completed: Full fix within budget
    

    Without budget: Agent reads 3 related files (5K), explores the CSS framework (3K), checks browser compatibility (2K), writes a detailed explanation (4K), implements the fix (1.5K), refactors two other files "while at it" (6K) = 21.5K tokens for a 5K-token fix.

    Use Cases

    • Cost-conscious daily development — keep per-task token costs predictable on paid API plans
    • Team budget management — provide per-agent token visibility and enforcement across engineering teams
    • Long-running session control — prevent token burn during multi-hour debugging or refactoring sessions
    • Agent output auditing — use attribution reports to identify inefficiency patterns and tune the budget model
    • Budget-constrained projects — enforce hard ceilings when working with fixed API credits

    Known Limitations

    Budget allocation is an estimate based on task complexity assessment — some tasks will be overestimated (wasting budget headroom) and others underestimated (triggering escalation). Progressive compression improves efficiency but may reduce output detail in some cases. The skill manages internal budget discipline, not external token compression tools (RTK, Caveman, Lean CTX) — those are complementary.

    How to install

    Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.

    ~30 seconds
    1. 1

      Download the ZIP

      Free skills download straight away. Paid skills unlock right after purchase.

    2. 2

      Unzip into your skills folder

      Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.

    3. 3

      Ask your agent to use it

      Restart the agent if it was already running. It picks the skill up automatically - no config needed.

    Skills folder by agent

    Click the path to copy it. Create the folder if it does not exist yet.

    Reviews

    No reviews yet

    Be one of the first to try it. Every listed skill passes our trust checks below.

    Security scanned

    Passed our 8-point scan before listing

    6 installs

    Downloaded by developers to date

    Free forever

    No account required to browse

    Trust & safety

    Security scanned

    Verified clean 2 months ago

    • Free to download with an account

    Listed2 months ago
    Updated9 days ago

    What's inside

    Frequently Asked Questions