FitLLM is a free, auditable local-LLM fit calculator for NVIDIA and AMD GPUs and Apple Silicon Macs. Choose a model, weight quant, KV-cache quant and context length to see whether it fits your GPU's VRAM or your Mac's unified memory and how much is left. It models the sliding-window and hybrid attention of Gemma 4 and Qwen 3.6 and the MLA compressed KV cache of GLM-5.2, GLM-4.7-Flash and the DeepSeek family — cases that naive uniform-KV formulas miss — and lets you paste a Hugging Face config from a modeled architecture family. Architecture inputs are pinned to official configs; runtime and OS reserves remain documented estimates in the open-source MIT engine. 이 LLM, 내 GPU·맥에 들어갈까?