BROMANDERLABS
Ten free, no-signup tools we build for the AI engineering community. Real math, total privacy, zero signups.
Each tool ships with our reasoning open and our assumptions documented. If we say your 4090 runs Qwen3.8 27B at 34 tok/sec, you can see exactly why.
Local AI Scanner
Can your rig run that model?
Math-based VRAM, KV cache, and tokens/sec estimates for 140+ GPUs and 50+ open-weight LLMs — Llama 4, Gemma 4, Qwen3, DeepSeek R1 included. Multi-GPU rigs too, layer split or tensor parallel. The honest answer to "Can I Run It?".
Prompt Cost Estimator
Stop guessing your API bill.
Per-call, daily, monthly, and yearly cost across 60+ models and 12 providers — Anthropic, OpenAI, Google, DeepSeek, Z.ai GLM, Qwen, Kimi, MiniMax, Mistral, xAI, Cohere, and hosted Llama. Cache + batch math included.
Context Window Visualizer
See what fits in 128K.
Paste your codebase, docs, or chat history. Live token count + a fit-or-truncate verdict across 28 models from 32K all the way to Llama 4 Scout's 10M.
Inference Latency Map
Who's actually fast?
TTFT + throughput leaderboard across 27 provider+model combos — from Cerebras' 2,600 tok/s wafer-scale to OpenAI's reasoning-mode marathons. Sources cited, reasoning models flagged.
Fine-Tuning VRAM Calculator
Can your rig train it?
Peak VRAM for full, LoRA, and QLoRA across 50+ models — weights, gradients, optimizer state, and activations broken out. The honest answer to "Can I fine-tune this?".
Self-Host vs API Break-Even
Build or buy?
Pit pay-per-token APIs against renting and owning GPUs. Find the exact monthly volume where a 4090 pays for itself versus your favorite frontier model.
Quantization Explorer
Is Q4 good enough?
Size versus quality across 11 GGUF quants, with llama.cpp perplexity data and the best quant that still fits your card. Stop guessing at Q4_K_M.
Energy & Carbon Calculator
What does your AI burn?
Tokens to kilowatt-hours to CO₂, across 20 GPUs and 9 grid regions. See your inference footprint in miles driven, phone charges, and trees.
Agent Loop Cost Simulator
Turn 30 re-sends turn 1.
Agent cost is quadratic in the turn count, and everyone budgets it as linear. See the real token total, what prompt caching claws back, and the case where caching costs you more than it saves.
Serving Capacity Planner
How many users, really?
KV cache runs out long before the weights do. Max concurrency, aggregate throughput, and dollars per million tokens served — for one card or a rack of eight.
Math, not marketing
Every number on these tools traces back to a published formula or a benchmark we can defend. No hand-waving.
Your data, your machine
Calculators run entirely in your browser. Your specs, prompts, and inputs are never sent to us or logged. Privacy is the default.
Always free
We build software, we don't sell your data. Every tool here is funded entirely by our own apps.
Want a tool like this built for you?
Hire Bromander Studios