> BROMANDER_LABS

BROMANDERLABS

Ten free, no-signup tools we build for the AI engineering community. Real math, total privacy, zero signups.

Each tool ships with our reasoning open and our assumptions documented. If we say your 4090 runs Qwen3.8 27B at 34 tok/sec, you can see exactly why.

Built in-houseUpdated continuouslyOpen methodology
─── The Tools ───
01

Local AI Scanner

Can your rig run that model?

Math-based VRAM, KV cache, and tokens/sec estimates for 140+ GPUs and 50+ open-weight LLMs — Llama 4, Gemma 4, Qwen3, DeepSeek R1 included. Multi-GPU rigs too, layer split or tensor parallel. The honest answer to "Can I Run It?".

02

Prompt Cost Estimator

Stop guessing your API bill.

Per-call, daily, monthly, and yearly cost across 60+ models and 12 providers — Anthropic, OpenAI, Google, DeepSeek, Z.ai GLM, Qwen, Kimi, MiniMax, Mistral, xAI, Cohere, and hosted Llama. Cache + batch math included.

03

Context Window Visualizer

See what fits in 128K.

Paste your codebase, docs, or chat history. Live token count + a fit-or-truncate verdict across 28 models from 32K all the way to Llama 4 Scout's 10M.

04

Inference Latency Map

Who's actually fast?

TTFT + throughput leaderboard across 27 provider+model combos — from Cerebras' 2,600 tok/s wafer-scale to OpenAI's reasoning-mode marathons. Sources cited, reasoning models flagged.

05

Fine-Tuning VRAM Calculator

Can your rig train it?

Peak VRAM for full, LoRA, and QLoRA across 50+ models — weights, gradients, optimizer state, and activations broken out. The honest answer to "Can I fine-tune this?".

06

Self-Host vs API Break-Even

Build or buy?

Pit pay-per-token APIs against renting and owning GPUs. Find the exact monthly volume where a 4090 pays for itself versus your favorite frontier model.

07

Quantization Explorer

Is Q4 good enough?

Size versus quality across 11 GGUF quants, with llama.cpp perplexity data and the best quant that still fits your card. Stop guessing at Q4_K_M.

08

Energy & Carbon Calculator

What does your AI burn?

Tokens to kilowatt-hours to CO₂, across 20 GPUs and 9 grid regions. See your inference footprint in miles driven, phone charges, and trees.

09

Agent Loop Cost Simulator

Turn 30 re-sends turn 1.

Agent cost is quadratic in the turn count, and everyone budgets it as linear. See the real token total, what prompt caching claws back, and the case where caching costs you more than it saves.

10

Serving Capacity Planner

How many users, really?

KV cache runs out long before the weights do. Max concurrency, aggregate throughput, and dollars per million tokens served — for one card or a rack of eight.

Math, not marketing

Every number on these tools traces back to a published formula or a benchmark we can defend. No hand-waving.

Your data, your machine

Calculators run entirely in your browser. Your specs, prompts, and inputs are never sent to us or logged. Privacy is the default.

Always free

We build software, we don't sell your data. Every tool here is funded entirely by our own apps.

Want a tool like this built for you?

Hire Bromander Studios