trinity_coordinator ships with a small set of named runtime
profiles that bundle a backend choice, a default coordinator SLM,
and a set of validation expectations into one keyword. Profiles are
defined in TrinityCoordinator.RuntimeProfile and resolved by name
or struct.
A profile answers three questions in one place:
- Which Nx backend? (
:nx_backend— e.g.{EXLA.Backend, client: :cuda},{EMLX.Backend, device: :gpu},Nx.BinaryBackend.) - Is CUDA required? (
:require_cuda?— gatesRuntime.put_cuda_backend!/0so callers can opt out cleanly.) - What is the default SLM coordinator profile? (
:default_slm_profile.)
Plus metadata for downstream callers (:export_svd?, :large_svd?,
:qwen_runtime?, :artifact_runtime?) and operator-facing
:notes / :warnings.
Linux/CUDA happy path. Used for everything the project does today.
nx_backend: {EXLA.Backend, client: :cuda}require_cuda?: truedefault_slm_profile: :qwen_coordinator
Requires XLA_TARGET=cuda12 and a working EXLA CUDA stack.
Apple Silicon (MLX-backed) lane. Brings up EMLX.Backend, device: :gpu
when the optional :emlx dep is present.
nx_backend: {EMLX.Backend, device: :gpu}require_cuda?: falsedefault_slm_profile: :qwen_coordinator
EMLX is an optional dependency. To use this profile, add it to your
parent application's mix.exs:
{:emlx, "~> 0.3"}Then:
mix deps.get
mix run examples/qwen_router_prompt_eval.exs --runtime-profile emlx \
--snapshot examples/fixtures/qwen_router_prompt_eval_logits.json \
--determinism-runs 2- Thin SVD memory footprint. Nx main as of commit
6424c89(Paulo Valente, PR #1753) refactoredNx.LinAlg.svd/2withfull_matrices?: falseso it does not materialise the fullm × mU on the Qwen3-0.6B embedder (wherem = 151_936, i.e. (92 GB of U under the old path). This fix is in the Nx version thattrinity_coordinatorpins to. Both EMLX and EXLA benefit from this change. --svd-compute-type f32. Recommended on Apple. The thin-SVD path uses aneighdecomposition under the hood; doing that work in f32 keeps the small-σ tail precise.- Backend label. When the exporter validates per-tensor backend
during the SVD reconstruction step, it accepts the
"EMLX.Backend"label as well as"EXLA.Backend<cuda:". No code changes needed for the user. - Bumblebee Qwen3 support. Bumblebee is git-pinned to a Qwen3-
supporting commit (post-v0.7.0 main). EMLXAxon
(github.com/elixir-nx/emlx) has
independently validated Qwen3-0.6B loading through the EMLX backend.
Paulo Valente confirmed on 2026-05-21 that running with the bare
EMLX backend (no
EMLXAxon.rewrite/1) successfully exports and passes 37/37 on the prompt eval. - bf16 round-trip. The bundle is bf16 safetensors. EMLX accepts
bf16 natively (
{:bf, 16}↔ MLXbfloat16). No quantisation or type cast required.
Apple Silicon (MLX-backed) research/validation lane via the
Emily backend. Same Apple-shaped flags
as :emlx but routes to Emily.Backend and ships
ausimian's empirically-derived per-profile margin floors
(agent: 0.33, role: 0.82) so a clean Emily run does not mark
the escalate_to_human case as a near-miss against the canonical CUDA
role floor of 1.06.
nx_backend: {Emily.Backend, []}require_cuda?: falsedefault_slm_profile: :qwen_coordinatordefault_min_agent_margin: 0.33default_min_role_margin: 0.82
Emily is an optional dependency. To use this profile, add it to your
parent application's mix.exs — do NOT add it to
trinity_coordinator's own mix.exs:
{:emily, "~> 0.4", only: [:dev, :test]}Then:
mix deps.get
XLA_TARGET=cuda12 mix trinity.sakana.export_adapted \
--force \
--svd-compute-type f32 \
--runtime-profile emily \
--out tmp/emily_adapted_qwen3_0_6b_layer26
mix run examples/qwen_router_prompt_eval.exs \
--runtime-profile emily \
--artifact-dir tmp/emily_adapted_qwen3_0_6b_layer26 \
--determinism-runs 2Note: the --min-agent-margin / --min-role-margin flags are no
longer required — the :emily profile seeds its own floors via
RuntimeProfile.default_margins/1 (see "Per-Profile Snapshot Fixtures
And Margin Floors" below). Pass them explicitly only if you want to
override the seeded values for a one-off run.
The canonical Apple lane for production-shaped workloads remains
:emlx; :emily is the research/validation lane. They are both
Apple-shaped and they both pass the prompt eval — pick :emily when
you want Paulo Valente's thin-SVD path under MLX, and :emlx when you
want the EMLX runtime that EMLXAxon was built against.
:emlx and :emily are both Apple-Silicon (MLX-family) but differ at
the Nx-backend layer:
:emlx→EMLX.Backend, device: :gpu. Canonical Apple lane; EMLXAxon has independently validated Qwen3-0.6B through it.:emily→Emily.Backend. Research/validation backend; ships theGram-matrix thin-SVD path that adapted on top of Nx PR #1753 ("better memory footprint for thin SVD") and was the lane on which the Apple-side end-to-end run was first proven (ausimian, 2026-05-21, 37/37 decisions match CUDA; one role-margin near-miss absorbed by the per-profile floor seeded above).
Both lanes pass the same prompt-eval suite and are decision-stable
against the CUDA snapshot. route_hash will drift on every case for
both lanes — that's expected on a different kernel stack and is exactly
what the per-profile snapshot fixture lane below is designed for.
Pure-Elixir CPU fallback. Useful for unit tests and for quick sanity-checks on machines without any GPU.
nx_backend: Nx.BinaryBackendrequire_cuda?: falsedefault_slm_profile: :qwen_coordinator
Expect order-of-magnitude slower latencies than CUDA or EMLX. Not intended for production use.
Synthetic profile used in tests; not for real workloads. Skip unless you're writing tests.
Tuple-shaped profile for anyone wiring up a backend that does not have
a built-in name. The runtime calls Nx.global_default_backend({BackendMod, opts})
when this profile is selected.
After the Phase D refactor:
mix trinity.sakana.export_adapted --runtime-profile <name>— pick the backend used to run the SVD/SVF pipeline.mix trinity.sakana.router_trace --runtime-profile <name>— trace a routing call with the selected backend.mix trinity.sakana.large_tensor_chunks --runtime-profile <name>— chunked tensor work (default: CUDA for back-compat).mix trinity.sakana.parity_sample --runtime-profile <name>— parity sampling.mix trinity.hitl.adapted --runtime-profile <name>, alsotrinity.hitl.base_qwen,trinity.hitl.gpu,trinity.hitl.head_route.mix run examples/qwen_router_prompt_eval.exs --runtime-profile <name>.mix run examples/local_coordinator_route.exs --runtime-profile <name>.mix run examples/mock_orchestration_trace.exs --runtime-profile <name>.
--runtime-profile cuda_exla is the default for every task, so no
existing CUDA workflow needs a flag.
TrinityCoordinator.Sakana.Coordinator.load/1 accepts:
TrinityCoordinator.Sakana.Coordinator.load(
runtime_profile: :emlx,
artifact_dir: "priv/sakana_trinity/adapted_qwen3_0_6b_layer26"
)For finer-grained overrides — for example, picking a non-named backend
without writing a {:custom, ...} profile — you can pass the
backend tuple directly:
TrinityCoordinator.Sakana.Coordinator.load(
runtime_profile: :emlx, # for require_cuda? = false
backend: {EMLX.Backend, device: :cpu} # but use CPU device
)The :backend and :require_cuda keys are compatibility overrides
that pre-date the profile system; they remain supported.
| You have… | Use |
|---|---|
| NVIDIA GPU + CUDA-12 toolchain + Linux | :cuda_exla (default) |
| Apple Silicon (M-series), production-shaped | :emlx + add {:emlx, "~> 0.3"} to your deps |
| Apple Silicon (M-series), research / Emily MLX | :emily + add {:emily, "~> 0.4"} to your deps |
| No GPU; want to run unit tests / quick sanity checks | :binary |
| Some other backend (e.g. Torchx, custom NIF) | {:custom, BackendMod, opts} |
mix trinity.env.checkreports the current XLA_TARGET and any artifact-directory issues
without loading EXLA. For richer per-profile validation, the
RuntimeProfile.compatibility_probe/1 family of functions returns a
structured report indicating whether the profile's expected backend is
loadable in this process, whether the artifact path exists, and so on.
examples/qwen_router_prompt_eval.exs supports a per-profile snapshot
fixture lane and per-profile margin defaults so a non-CUDA backend can
land its own empirical floors without rewriting or lowering the
canonical CUDA snapshot.
Resolution order for the --snapshot flag (Phase 5):
- Explicit
--snapshot path— wins unconditionally; existing CI flows pinning the canonical CUDA fixture keep working without reinterpretation. examples/fixtures/runtime_profiles/<profile>/qwen_router_prompt_eval_logits.json— picked up automatically when the per-profile file is present.nil— no snapshot drift check (the same default behaviour as before Phase 5). To pin against the CUDA snapshot, pass--snapshot examples/fixtures/qwen_router_prompt_eval_logits.jsonexplicitly. We deliberately do not fall through to the legacy fixture path automatically: that would silently enable a strict 6dp logits byte-equivalence check for operators who did not opt in.
Margin floor resolution (--min-agent-margin / --min-role-margin):
- Explicit CLI flag — wins.
RuntimeProfile.default_margins(profile)— every built-in profile inherits the canonical CUDA defaults (agent: 0.24,role: 1.06) unless overridden viaRuntimeProfile.override_default_margins/2(e.g. for a future:emilyprofile that wantsagent: 0.33,role: 0.82).- Module-level defaults (legacy fallback in the eval script).
Emily is a first-class profile — see the :emily section
above for the full recipe. The short version:
- Add
{:emily, "~> 0.4", only: [:dev, :test]}to your parent application'smix.exs. Do NOT add it totrinity_coordinator's ownmix.exs. mix deps.get.- Pass
--runtime-profile emilytomix trinity.sakana.export_adaptedandmix run examples/qwen_router_prompt_eval.exs.
The profile's default_min_agent_margin / default_min_role_margin
fields are pre-seeded with the empirical Emily floors (0.33 / 0.82)
from ausimian's 2026-05-21 validation pass, so a clean run does not
require any explicit --min-*-margin overrides.
If you would rather keep your run shaped exactly like the prior
{:custom, Emily.Backend, []} recipe — for example to pin a different
backend module — the custom-tuple form still works:
profile =
TrinityCoordinator.RuntimeProfile.resolve({:custom, Emily.Backend, []})
|> TrinityCoordinator.RuntimeProfile.override_default_margins(
agent: 0.33,
role: 0.82
)- 0/37 drift on the decision-stable fields (
agent_id,role_id,token_count,transcript_hash). - 37/37 differ on
route_hash(6dp logit drift — expected on a different kernel stack). - Empirical worst margins were
agent: 0.417(two_assistant_turns) androle: 1.029(escalate_to_human); the 80% floors are therefore0.33/0.82. These are exactly the values the built-in:emilyprofile now ships. - Phase 1 (lazy-backend timing sync) makes
decompose_elapsed_msreport real GPU wall time on Emily / EMLX instead of the host-side dispatch cost of an unmaterialised future.
route_hash drifts on every case under Emily because float aggregation
order differs from CUDA. To pin Emily-stable snapshots, drop a
examples/fixtures/runtime_profiles/emily/qwen_router_prompt_eval_logits.json
file next to the legacy CUDA fixture; the eval entry point's
SnapshotResolver will pick it up automatically when
--runtime-profile emily is passed without --snapshot. See "Per-Profile
Snapshot Fixtures And Margin Floors" above for the resolution order.
The seed snapshot can be generated by running the eval once with
--snapshot-out examples/fixtures/runtime_profiles/emily/qwen_router_prompt_eval_logits.json
on Apple Silicon.
TrinityCoordinator.RuntimeProfile— the module that defines and resolves profiles.TrinityCoordinator.Sakana.Coordinator.load/1— the canonical load entry point.- Artifact Distribution — how to fetch / publish the bundle.
- Troubleshooting — common failure modes by symptom.