llm gpu fit
Stop OOM crashes with architecture-aware VRAM estimation and 3-way agent compatibility checks for local LLMs.
Free
Works with the AI tools you already use
llm gpu fit
Example session with this skill installed
I have an 11GB RTX 2080 Ti. Can I run Llama 3.1 8B or Qwen2.5 7B if I need a 128K context window for my agent? Show me the math for the KV cache.
- Read your context and instructions
- Compiled the llm gpu fit
I have analyzed the resource requirements for Llama 3.1 8B and Qwen2.5 7B against an 11GB VRAM budget with a 128K context requirement.
GPU Fit Analysis (11GB VRAM)
| Model | Arch | Weights (Q4) | KV Cache (128K) | Total Est. | Status |
|---|---|---|---|---|---|
| Qwen2.5 7B | GQA | 4.7 GB | ~2.1 GB | 7.8 GB | ✅ FIT |
| Llama 3.1 8B | MHA | 4.9 GB | ~16.0 GB | 21.9 GB | ❌ OOM |
Technical Breakdown
Architecture Impact: Llama 3.1 uses Multi-Head Attention (MHA), where KV cache grows linearly and heavily (1x). At 128K context, the cache alone exceeds your total VRAM.
*
GQA Advantage: Qwen2.5 uses Grouped-Query Attention (GQA), reducing the KV cache footprint by approximately 8x. This allows the model + full context + 1GB system headroom to fit comfortably within 11GB.
*
Agent Compatibility: Both models use standard content fields for responses, making them suitable for agentic workflows, unlike newer "thinking" models (e.g., Llama 3.2 3B-Instruct or Qwen2.5-7B-Reasoning).
Next steps
- Pull the GQA model:
ollama run qwen2.5:7b - Create a 128K context variant:
curl http://localhost:11434/api/create -d '{"name":"qwen-128k","from":"qwen2.5:7b","parameters":{"num_ctx":131072}}' - Verify the agent endpoint:
curl http://localhost:11434/v1/chat/completions -d '{"model":"qwen-128k","messages":[{"role":"user","content":"test"}]}'
Connects securely to your tools. The creator never sees your data.
About this skill
The problem
Running local LLMs often leads to OOM crashes or empty agent responses because standard VRAM math ignores KV cache growth and architectural differences. You waste hours downloading 10GB models only to find they lack the context window your agent requires or hide reasoning in non-standard fields.
What it does
- Performs architecture-aware VRAM estimation by calculating KV cache costs for MHA, GQA, and MQA models.
- Conducts a 3-way fit check verifying VRAM headroom, context window minimums, and agent field compatibility.
- Identifies models that break OpenAI-compatible tools by routing output to "thinking" fields instead of the standard content field.
- Provides precise curl commands to verify model loading and create custom context length variants without re-downloading weights.
Frameworks & tools
Ollama, NVIDIA SMI, CUDA, and OpenAI-compatible inference servers.
Why this beats prompting it yourself
Generic LLMs often use "paper math" that underestimates memory usage by ignoring the tokeniser and CUDA context overhead. This skill uses empirical data from real hardware testing to account for the specific multipliers in modern architectures like Qwen and Llama.
Use cases
- Pre-flighting a model selection for a local agent requiring a 64K+ context window.
- Debugging why a model "fits" on paper but fails to return content to your application.
- Optimizing GPU utilization by switching from MHA to GQA models for long-context tasks.
- Automating the creation of model variants with specific
num_ctxparameters.
Known limitations
Estimates are based on Q4_K_M quantization as the baseline. Real VRAM usage may vary slightly based on specific driver versions and system background tasks.
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
5 people have installed this skill.
Trust & safety
Security scanned
Verified clean 29 days ago
- Free to download with an account