sys1rust runs Laya's System 1 models on your Mac's GPU, in Rust with its own Metal kernels and no Python. Send it some text and a few questions, and it answers each one with a choice, a score or a yes probability. Its server, sys1rust serve, speaks the /v1/systemone API of laya serve, so Jev and Laya clients send it the same requests.
Paste this line into Claude Code, Codex, Cursor or another coding agent:
Set up sys1rust on my Mac from https://github.com/krishhgg/sys1rust
The agent clones this repository and follows AGENTS.md, which has it check your Mac, run install.sh, download the model, start the server and send it a test request.
To install it yourself on an Apple silicon Mac with macOS 14 or later, run:
curl -fsSL https://raw.githubusercontent.com/krishhgg/sys1rust/main/install.sh | sh
sys1rust serveThe script installs the latest release into ~/.local/share/sys1rust and links ~/.local/bin/sys1rust, without sudo. If ~/.local/bin is not on your PATH, it prints the line to add to ~/.zshrc. Until you add it, run ~/.local/bin/sys1rust serve instead.
On macOS 26.2 or later the script installs the bundle with MLX's build for that version, which is about 3x faster on the M5 than the bundle for macOS 14 to 26.1. The speed numbers in this README use the macOS 26 bundle. The Releases page has both tarballs, and the script checks each download against the release's SHA256SUMS.
The first serve downloads the typed-decisions model (846 MB) into the Hugging Face cache, then prints listening on http://127.0.0.1:8000 when it's ready. sys1rust pull [model] downloads a model ahead of time and sys1rust models lists what is downloaded.
To run it in the background now and at every login, first stop the foreground server with Ctrl-C to free port 8000, then install the service:
curl -fsSL https://raw.githubusercontent.com/krishhgg/sys1rust/main/install.sh | sh -s -- --serviceThe service is a LaunchAgent that runs sys1rust serve at login, restarts it if it crashes and logs to ~/Library/Logs/sys1rust.log. From a clone of this repository, ./install.sh --service does the same.
To update, run the install command again, and it restarts a running service on the new version. To uninstall, run it with sh -s -- --uninstall in place of sh, which removes the service, the install, the link and the log, keeps the downloaded models and prints how to delete them.
With the server from Install running, in the foreground or as the service, send it a request from another terminal:
curl -s localhost:8000/v1/systemone -H 'content-type: application/json' -d '{
"state": "Customer: I was charged twice for my subscription this month and I need it fixed today.",
"questions": {
"team": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"billing": "payments, charges and refunds", "tech": "bugs and outages", "sales": "new purchases"}},
"urgency": {"type": "score", "instructions": "How urgent is this?",
"criteria": ["none", "low", "high", "now"]},
"money": {"type": "noul", "instructions": "Is money involved?"}
}
}'On the base M5, inference for this request took 13.0 to 13.3 ms over 18 runs, 6 in each of 3 fresh processes after 2 warm-up requests. That is the X-Inference-Time-Ms header, which curl -si shows. It covers tokenizing, running the model and decoding. It leaves out reading and checking the request, waiting for the GPU and writing the reply.
Each question gets one of 3 answer types:
| type | you give | you get back |
|---|---|---|
choice |
labels, each with a description | choice, the most likely label, and probabilities for every label |
score |
an ordered list of levels | score, the expected level (it can fall between levels), probabilities per level, and a legend |
noul |
only the instructions | noul, the probability that the answer is yes |
Every answer also has 2 confidence numbers, which measure different things:
answer_confidenceis the probability of the most likely option. Forchoicethat is the chosen label, forscorethe most likely level, which can differ from the expectedscore, and fornoulthe larger of yes and no. It is the one to use for deciding whether to trust an answer.- For
choiceandscore,confidenceis 1 minus the normalized entropy of the probabilities. It is 0 for an even split and 1 when everything is on one option. Fornoul, it equalsanswer_confidence, which is 0.5 for an even split.
Don't compare them against the same threshold. action.act_probability is the model's probability, from a second output, for acting on the answer rather than escalating. Ctrl-C stops a foreground server, and launchctl bootout gui/$(id -u)/io.github.krishhgg.sys1rust stops the service until the next login. When the server stops, the model is out of memory.
- It checks requests the way
laya servedoes, with the same limits, error codes anddetailstrings. It checks the API key first, then takes a slot, then reads and checks the body. - The GPU runs one request at a time.
sys1rustholds up to 16 requests (LAYA_MAX_CONCURRENT), one running and the rest waiting their turn. A request that arrives while all 16 slots are in use gets503withRetry-After: 1at once, so clients must retry it. More clients don't get more throughput, since one request already fills the GPU. - All the questions in a request go through the model together, one row per question, in fp16.
- It beats Python MLX by doing less work and tuning its matmuls, not by being Rust. It skips computing on padding, runs the decision head's last layer only at the positions the answer reads, uses dense attention where that is cheaper, and from 512 tokens computes local attention only over the keys each window reaches. Its matmuls run MLX's own kernels for the M5's matrix units, with tile sizes tuned on this M5. The answers don't change.
- It downloads a model once, at a pinned revision. The first
sys1rust servefetches the 5 files the model needs (846 MB for typed-decisions) into the Hugging Face cache and checks each one's hash. Later starts read them from there.--offlineturns downloads off.
Median time per request for the typed-decisions model on a base M5, in ms. These are the numbers in the chart at the top.
| median ms | 1 question, 128 tokens | 1 question, 512 tokens | 10 questions, 512 tokens |
|---|---|---|---|
| sys1rust | 15.5 | 37.9 | 379 |
| laya-mlx, Python MLX at its fastest (compiled, buffer cache capped) | 19.0 | 47.6 | 430 |
laya serve, stock, over HTTP |
53.9 | 157.6 | 708 |
- Against stock
laya serve, it is 1.9 to 4.1x faster per request, countingsys1rust's 0.3 ms of HTTP. Thelaya servetimes come from the bake-off (REPORT.md), an earlier session than the round 3 sys1rust times, and separate runs on this laptop vary by about 5%. In earlier runs,sys1rustsustained 7.9 requests/s over 5 minutes, andlaya serve3.54 over 2 minutes.laya serveruns fp32 for requests with fewer than 5 questions, but even upstream in fp16, which it doesn't ship, is 1.9 to 2.9x slower on these 3 sizes, also against times from an earlier session. - Against Python laya-mlx at its fastest, it is 1.25x faster, and faster on all 12 benchmark sizes, by 13 to 55%.
- It gives the same answers. On the 1,500-answer correctness workload, 1,498 agree with the upstream PyTorch fp32 reference (99.9%, above the 99% gate). The 2 that differ are near ties.
- Its tail stays close to the median. In the timing runs, no request took more than twice the median for its size. Python MLX without a capped cache had 9% of requests over that line.
- It starts fast and is small. From process start to the first answer takes 243 to 375 ms, with the model files already in the OS file cache.
sys1rustis a 7 MB binary, and its HTTP layer adds 0.3 ms per request. macOS caches the compiled GPU kernels by the binary's path and the kernel source, so the first start from a new path or after a kernel change takes about 1.7 s longer.
All these measurements come from one Mac, a MacBook Pro 14 with a base M5, on macOS 26.2 and wall power. Other chips are untested. The write-ups are in results/. They run from the bake-off of 13 existing runtimes (REPORT.md) to the speed round (SPEED.md).
--model |
Hugging Face repo | encoder | agreement with upstream fp32 answers |
|---|---|---|---|
typed-decisions (default) |
convaiinnovations/laya-typed-decisions |
ModernBERT-large | 1,498 of 1,500 on correctness |
multilingual |
convaiinnovations/laya-multilingual |
mmBERT-base | 100% on correctness, smoke and short |
english |
convaiinnovations/laya |
ModernBERT-large | 100% on smoke, short and cold; no upstream correctness reference exists |
The first serve of a model downloads it. --model also takes a local checkpoint directory, which sys1rust reads without downloading anything; the Hugging Face CLI's hf download <repo> fetches another hub checkpoint and prints its directory.
You need an Apple silicon Mac, Rust 1.89 or newer, CMake, the Xcode command line tools, and Python 3.10 or newer. Python only fetches MLX. sys1rust doesn't run it.
# A prebuilt MLX 0.32.2 (the Python wheel ships libmlx and its CMake files).
python3 -m venv .mlx && .mlx/bin/pip install mlx==0.32.2
export MLX_SYS_PREBUILT_DIR="$(.mlx/bin/python -c 'import mlx.core, os; print(os.path.dirname(mlx.core.__file__))')"
# The binaries go to runtime/target/release/.
cargo build --release --manifest-path runtime/Cargo.tomlThen run runtime/target/release/sys1rust serve. Where this README runs sys1rust, use that path, or add runtime/target/release to your PATH. The binary loads MLX from that venv by its absolute path, so keep .mlx/ where it is. To change the code and run the tests, follow .agents/skills/sys1rust-develop/SKILL.md.
Options
Each flag falls back to an environment variable. The LAYA_* ones are the same as laya serve's.
| flag | environment variable | default | |
|---|---|---|---|
--model |
SYS1_MODEL |
typed-decisions |
a name from Models, its repo id, or a checkpoint directory |
--revision |
SYS1_REVISION |
the pinned one | serve snapshots/<sha> of the cached repo instead. The pin's first 7 characters, as sys1rust models shows them, also mean the pin |
--host |
LAYA_HOST |
127.0.0.1 |
bind address; laya serve binds 0.0.0.0 |
--port |
LAYA_PORT |
8000 |
0 picks a free port, printed on the ready line |
--api-key |
LAYA_API_KEY |
none | when set, /v1/systemone needs Authorization: Bearer <key> |
--max-concurrent |
LAYA_MAX_CONCURRENT |
16 |
requests held at once; the next one gets 503 |
--tuning |
SYS1_MLX_TUNING |
the measured default | engine settings, see Knobs in runtime/crates/laya-mlx |
--f32 |
SYS1_F32 |
off | run the transformer in f32 instead of the checkpoint's f16 |
--offline |
HF_HUB_OFFLINE |
off | never download; a model that is not in the cache is an error. HF_HUB_OFFLINE turns it on when set to 1, ON, YES or TRUE, in any letter case |
Downloads honor HF_ENDPOINT and HF_TOKEN. sys1rust sends the token only to the endpoint, never after a redirect.
sys1rust sets MLX's MLX_MAX_MB_PER_BUFFER to 10, measured faster on the M5, unless it is already set. At load it checks its M5 matmul kernels against MLX's, bit for bit, and prints the result on stderr. If the check fails, it uses MLX's matmuls and prints why.
When it's ready, sys1rust prints one JSON line on stdout with the address, model, revision, load time and warm-up time. GET /health reports the model, revision and engine. SIGINT or SIGTERM lets requests in flight finish before it exits. sys1rust serve --help lists everything.
Serving other machines
sys1rust listens on 127.0.0.1 by default, so only programs on the same Mac can reach it. Keep that default. sys1rust speaks plain HTTP without TLS, and it has no header read timeout and no connection limit. Only the request body has a deadline, and a body that takes over 10 s gets 408.
An API key alone does not make remote serving safe. With LAYA_API_KEY set, /v1/systemone answers only requests that send Authorization: Bearer <key>, but over plain HTTP anyone on the network path can read that key. sys1rust checks the key only once the headers have arrived, so it does nothing against connections that never finish sending them. Each such connection holds a socket for as long as the client keeps it open, and nothing caps how many there are.
To serve clients on other machines, run a reverse proxy on the same Mac in front of sys1rust. The proxy should terminate TLS, time out slow headers, limit connections and forward to 127.0.0.1. Set an API key as well:
export LAYA_API_KEY=replace-with-a-secret
sys1rust serve --port 8000 # the proxy forwards to 127.0.0.1:8000--host changes the bind address, but any client that can reach a wider address talks to sys1rust directly, with none of the proxy's protections.
Differences from laya serve
- Bind address.
sys1rustlistens on 127.0.0.1 by default.laya servelistens on 0.0.0.0. - Downloads.
sys1rustdownloads only the 3 Laya models, only at the revisions pinned inbench/models.lock.json, and only the 5 files each one needs.laya servedownloads any model repo at its newest revision. - Duplicate choice keys. When 2 of a choice's labels print as the same JSON key, such as
"1"and1, upstream writes that key twice inprobabilities, andsys1rustwrites it once with the second label's probability. Python'sjson, JavaScript'sJSON.parseand serde_json'sValuekeep the last duplicate, so a client parsing with one of them gets the same object from both servers. A parser that rejects duplicate keys or keeps the first one reads the 2 responses differently. - JSON.
sys1rustaccepts standard JSON in UTF-8. Upstream'sjson.loadsalso acceptsNaN,Infinityand-Infinity,\uescapes of unpaired surrogates, and bodies in UTF-16 or UTF-32 or with a byte order mark.sys1rustanswers those with 400request body must be valid JSON.
Build and run inside the benchmark setup
Use this instead of Build from source when working on the benchmark. bench/env.sh takes MLX from the laya-mlx contender's venv but does not create it, so the first step sets that venv up once, as in bench/contenders/laya-mlx/NOTES.md. That step needs uv. The script also keeps the Cargo output and the Hugging Face cache under bench/, so the binary is at $CARGO_TARGET_DIR/release/sys1rust. The commands download and serve the typed-decisions revision pinned in bench/models.lock.json, the one the results used.
source bench/env.sh
# Once: the laya-mlx contender's venv, with MLX 0.32.2 and the hf CLI.
git clone https://github.com/mizorewww/laya-mlx bench/contenders/laya-mlx/src
(cd bench/contenders/laya-mlx/src && git checkout -q 0a859518634112655cb97c745dbf04f5191aaf13 &&
UV_PROJECT_ENVIRONMENT=../.venv uv sync --frozen --python 3.12 --managed-python)
cargo build --release --manifest-path runtime/Cargo.toml
REV=1a793eb568e6718f15941d08f85432581df534e3 # typed-decisions sha in bench/models.lock.json
bench/contenders/laya-mlx/.venv/bin/hf download convaiinnovations/laya-typed-decisions --revision $REV
$CARGO_TARGET_DIR/release/sys1rust serve --model typed-decisions --port 8000Repository layout
install.sh: the installer for prebuilt releases.AGENTS.md,CLAUDE.mdand.agents/skills/: instructions for coding agents that set up, call or change sys1rust.runtime/: the Rust workspace.laya-core: request parsing, tokenization, sequence layout and answer decoding, with no GPU code.laya-mlx: the forward pass on MLX through mlx-rs.sys1rust: the command and its HTTP server.sys1-bench: the benchmark adapter andsys1-probe.vendor/mlx-sys: mlx-sys 0.6.0 with a build that can link a prebuilt MLX.
bench/: the benchmark harness, workloads and upstream reference answers (bench/PLAN.md).packaging/: the release bundle builds, the smoke test and the release checklist (packaging/RELEASE.md).results/: measured write-ups, from the bake-off of existing runtimes to the speed round.research/: sourced reports on the models, runtimes and hardware.docs/assets/: the diagrams in this README.
Apache-2.0. laya-core and laya-mlx started as a fork of tjameswilliams/laya-r-mlx, and vendor/mlx-sys is from oxiglade/mlx-rs. See NOTICE.