LG

    llm gpu fit

    2

    Stop OOM crashes with architecture-aware VRAM estimation and 3-way agent compatibility checks for local LLMs.

    Free

    5 installs5.0 (1 review)

    Works with the AI tools you already use

    Claude CodeClaude CodeCursorCursorCodex CLICodex CLIGitHub CopilotGitHub CopilotGemini CLIGemini CLI+20 more
    LG

    llm gpu fit

    Example session with this skill installed

    I have an 11GB RTX 2080 Ti. Can I run Llama 3.1 8B or Qwen2.5 7B if I need a 128K context window for my agent? Show me the math for the KV cache.

    • Read your context and instructions
    • Compiled the llm gpu fit

    I have analyzed the resource requirements for Llama 3.1 8B and Qwen2.5 7B against an 11GB VRAM budget with a 128K context requirement.

    GPU Fit Analysis (11GB VRAM)

    ModelArchWeights (Q4)KV Cache (128K)Total Est.Status
    Qwen2.5 7BGQA4.7 GB~2.1 GB7.8 GB✅ FIT
    Llama 3.1 8BMHA4.9 GB~16.0 GB21.9 GB❌ OOM

    Technical Breakdown

    Architecture Impact: Llama 3.1 uses Multi-Head Attention (MHA), where KV cache grows linearly and heavily (1x). At 128K context, the cache alone exceeds your total VRAM.
    *

    GQA Advantage: Qwen2.5 uses Grouped-Query Attention (GQA), reducing the KV cache footprint by approximately 8x. This allows the model + full context + 1GB system headroom to fit comfortably within 11GB.
    *

    Agent Compatibility: Both models use standard content fields for responses, making them suitable for agentic workflows, unlike newer "thinking" models (e.g., Llama 3.2 3B-Instruct or Qwen2.5-7B-Reasoning).

    Next steps

    1. Pull the GQA model: ollama run qwen2.5:7b
    2. Create a 128K context variant: curl http://localhost:11434/api/create -d '{"name":"qwen-128k","from":"qwen2.5:7b","parameters":{"num_ctx":131072}}'
    3. Verify the agent endpoint: curl http://localhost:11434/v1/chat/completions -d '{"model":"qwen-128k","messages":[{"role":"user","content":"test"}]}'

    Connects securely to your tools. The creator never sees your data.

    About this skill

    The problem

    Running local LLMs often leads to OOM crashes or empty agent responses because standard VRAM math ignores KV cache growth and architectural differences. You waste hours downloading 10GB models only to find they lack the context window your agent requires or hide reasoning in non-standard fields.

    What it does

    • Performs architecture-aware VRAM estimation by calculating KV cache costs for MHA, GQA, and MQA models.
    • Conducts a 3-way fit check verifying VRAM headroom, context window minimums, and agent field compatibility.
    • Identifies models that break OpenAI-compatible tools by routing output to "thinking" fields instead of the standard content field.
    • Provides precise curl commands to verify model loading and create custom context length variants without re-downloading weights.

    Frameworks & tools

    Ollama, NVIDIA SMI, CUDA, and OpenAI-compatible inference servers.

    Why this beats prompting it yourself

    Generic LLMs often use "paper math" that underestimates memory usage by ignoring the tokeniser and CUDA context overhead. This skill uses empirical data from real hardware testing to account for the specific multipliers in modern architectures like Qwen and Llama.

    Use cases

    • Pre-flighting a model selection for a local agent requiring a 64K+ context window.
    • Debugging why a model "fits" on paper but fails to return content to your application.
    • Optimizing GPU utilization by switching from MHA to GQA models for long-context tasks.
    • Automating the creation of model variants with specific num_ctx parameters.

    Known limitations

    Estimates are based on Q4_K_M quantization as the baseline. Real VRAM usage may vary slightly based on specific driver versions and system background tasks.

    How to install

    Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.

    ~30 seconds
    1. 1

      Download the ZIP

      Free skills download straight away. Paid skills unlock right after purchase.

    2. 2

      Unzip into your skills folder

      Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.

    3. 3

      Ask your agent to use it

      Restart the agent if it was already running. It picks the skill up automatically - no config needed.

    Skills folder by agent

    Click the path to copy it. Create the folder if it does not exist yet.

    Reviews

    5.0
    1 review
    5
    1
    4
    0
    3
    0
    2
    0
    1
    0

    5 people have installed this skill.

    Trust & safety

    Security scanned

    Verified clean 29 days ago

    • Free to download with an account

    Listed1 month ago
    Updated29 days ago

    What's inside

    Frequently Asked Questions