Skip to content
Merged
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -287,6 +287,7 @@ Good entries should have a clear reason to exist. They should help people build,
- [Claude Code Local](https://github.com/nicedreamzapp/claude-code-local) - MLX-native server that speaks the Anthropic Messages API so the unmodified Claude Code CLI runs against local models on Apple Silicon, with parsing for local models' tool-call formats. MIT licensed. ![GitHub stars](https://img.shields.io/github/stars/nicedreamzapp/claude-code-local?style=social)
- [nemotron-omni-mlx](https://github.com/nicedreamzapp/nemotron-omni-mlx) - Pure MLX runtime for the vision and audio towers of NVIDIA Nemotron 3 Nano Omni on Apple Silicon, tested for parity against NVIDIA's PyTorch reference. MIT licensed. ![GitHub stars](https://img.shields.io/github/stars/nicedreamzapp/nemotron-omni-mlx?style=social)
- [Magnitude](https://github.com/magnitudedev/magnitude) - Hardware-aware local inference engine that profiles host hardware, recommends suitable open models, and tunes execution across Apple Silicon, NVIDIA, AMD, and CPU. Apache 2.0 licensed. ![GitHub stars](https://img.shields.io/github/stars/magnitudedev/magnitude?style=social)
- [llm-vram-calculator](https://github.com/159753a52/llm-vram-calculator) - Dependency-free TypeScript core that estimates inference memory (weights, KV cache, overhead) of any Hugging Face model from its config.json and safetensors metadata, covering MoE, MLA, sliding-window and linear-attention layers; runs in the browser at [modelvram.com](https://modelvram.com/llm-vram-calculator/). MIT licensed. ![GitHub stars](https://img.shields.io/github/stars/159753a52/llm-vram-calculator?style=social)

#### High-performance Serving & API Servers

Expand Down