Skip to content

Add llm-vram-calculator to Local / On-device Inference - #790

Merged
alvinunreal merged 1 commit into
alvinreal:mainfrom
159753a52:add-modelvram-vram-calculator
Sep 30, 2026
Merged

alvinunreal merged 1 commit into
alvinreal:mainfrom
159753a52:add-modelvram-vram-calculator

Conversation

@159753a52

Copy link
Copy Markdown

Project

Why it belongs

It helps people running models locally decide what fits before downloading: it estimates inference memory (weights, KV cache, runtime overhead) for any Hugging Face model from its config.json and safetensors metadata, rather than from a fixed model table. It handles cases generic calculators get wrong: all MoE experts counted toward weights, MLA's compressed latent cache, sliding-window layers, linear-attention / state-space layers, and MXFP4 / 4-bit expert checkpoints.

Quality signals

  • License: MIT
  • Maintenance status: actively maintained (last push 2026-09-29)
  • Documentation/examples: README documents each architecture case it models; tests pin results to real checkpoint sizes
  • Distinction from similar projects: llmfit scores models against your detected hardware from a curated list; this computes memory per model from the Hub's own files and runs in the browser without install

python3 tools/validate_awesome.py --skip-remote passes with 0 errors (the 10 warnings are pre-existing and unrelated).

Disclosure: I built and maintain this project. GitHub stars are low since it's new; happy to adjust or drop it if it doesn't meet the bar.

@alvinunreal
alvinunreal merged commit 4b86e6e into alvinreal:main Sep 30, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants