> BROMANDER_LABS
CAN YOURUN IT?
Honest, math-based estimates for running local LLMs on your hardware. VRAM fit, KV cache, and tokens-per-second — calculated the way llama.cpp actually runs them. Stack up to eight cards and see what layer split really buys you versus tensor parallelism.
All math runs in your browser. We never see your specs.
01 — Hardware
GPU
How many
02 — Model
System report
Verdict
S-TIER
38.4
tokens / sec
Fits comfortably in VRAM. You're running Llama 3.1 8B at full GPU speed.
Best fit: Q8_0auto-selected
Memory budget
Memory required9.98 GB / 11.0 GB VRAM usable
╳ 11.0 GB
Weights8.53 GB
KV cache0.54 GB
Overhead0.91 GB
VRAM
12 GB
Bandwidth
504 GB/s
Context
4K
→ Digital wellness
Running models is easy. Running yourself is harder.
Shinery tracks how AI fits into your day so it stays a tool, not a crutch.
─── Your Report Card ───
This is the actual image that shows on X, LinkedIn, and Facebook when you share the link.
Live preview
Built by
Bromander Studios — Hire Us