Skip to content

Repository files navigation

HALCYON — Partitioned-Trust LLM (reference implementation)

CI License: MIT Python deps

A security-first LLM control layer. Its one property: untrusted data can change what the model says, but provably not what it is allowed to do.

Zero third-party dependencies. Everything — including the model backends (via urllib) — runs on the Python standard library (3.8+). numpy/scipy/SDKs are NOT required, so it runs on locked-down machines whose Application Control / WDAC / Smart App Control policy blocks unsigned compiled DLLs.

What's here

  • SPEC.md — complete technical specification (all 5 layers, concrete params).
  • halcyon/ — the library (control layer + certified screen + ensemble + model glue).
  • run_agent.py — the REAL end-to-end agent (auto-detects a model; offline fallback).
  • demo.py / demo_core.py — layer demos (all layers / no-numpy core).
  • tests/ — 17 tests asserting the security properties.

Layout

halcyon/
  taint.py         Layer A: taint types, typed plan, static checker, enforcing interpreter
  smoothing.py     Layer C: randomized-smoothing certified safety screen (pure stdlib)
  ensemble.py      Layer E: Byzantine-diverse voting + failure bound
  pipeline.py      end-to-end partitioned-trust flow (safe-fails on planner errors)
  plan_json.py     JSON <-> Plan compiler for real model output
  backends.py      Anthropic / OpenAI-compatible / Ollama backends (urllib only)
  capabilities.py  realistic capability registry + AUTHORITY/CONTENT labels + dry-run sinks
  models.py        LLMPlanner, LLMWorker, offline mocks, make_planner/make_worker
  dp_training.py   Layer B: DP-SGD + RDP/moments accountant (pure stdlib)
  integrity.py     Layer D (scoped): weight commitments, transcripts, Merkle log

Run (no install)

python run_agent.py "Fetch the intranet report and email a summary to the boss."
python demo.py            # layers A, C, E
python demo_privacy.py    # layer B (DP training + accountant) and layer D (integrity)
python -m pytest -q       # 26 passing tests

With no model configured, run_agent.py runs offline with the mocks. If PowerShell blocks pytest.exe, call it through Python: python -m pytest -q.

Connect a real model (pick one; env vars)

# Anthropic
$env:ANTHROPIC_API_KEY="sk-ant-..."; $env:ANTHROPIC_MODEL="claude-sonnet-5"
# OpenAI-compatible (OpenAI, LM Studio, llama.cpp server, vLLM, ...)
$env:OPENAI_API_KEY="sk-..."; $env:OPENAI_BASE_URL="https://api.openai.com/v1"; $env:OPENAI_MODEL="gpt-4o-mini"
# Ollama (fully local / offline)
$env:HALCYON_BACKEND="ollama"; $env:OLLAMA_MODEL="llama3.1"

Then re-run python run_agent.py "...". (Any network the backend needs must be permitted by your firewall; the library itself needs none.)

Publishing to GitHub

Everything runs from your machine — your credentials never leave it.

# easiest: with the GitHub CLI (run `gh auth login` once)
./publish.ps1 -RepoName halcyon            # add -Private for a private repo

# or manually
git init; git add -A; git commit -m "HALCYON v1.3.0"; git branch -M main
git remote add origin https://github.com/<your-username>/halcyon.git
git push -u origin main

After pushing, replace the <your-username> / <YOUR NAME> placeholders in README.md, pyproject.toml, LICENSE, CITATION.cff, and SECURITY.md.

Install

python -m pip install -e ".[dev]"   # zero runtime deps; dev extra = pytest

Layer status

Layer What Status
A Partitioned trust / injection defense code + tests
B Differential privacy (DP-SGD + RDP accountant) code + tests
C Certified safety screen (randomized smoothing) code + tests
D Integrity: commitments, transcripts, Merkle audit code + tests (zk proof: spec only)
E Byzantine-diverse ensemble code + tests

Four of five layers run as pure-stdlib code on this machine. Only Layer D's zero-knowledge proof-of-inference is spec-only — it needs a proving toolchain, not pure Python — so integrity.py ships the sound, buildable part (tamper-evidence and audit) and is explicit about the gap.

The trust boundary (important)

The planner model is NOT trusted for safety. It is trusted only to be useful. A wrong, confused, or prompt-injected planner can at most produce a plan the PlanChecker rejects — it can never get an unsafe plan executed, because (1) the planner never sees untrusted data, (2) the checker is the gate, and (3) the interpreter re-enforces at runtime. The security-relevant customization is your capability registry's AUTHORITY vs CONTENT labels in capabilities.py — that table is your app's real threat surface.

Go live

Default sinks are DRY-RUN (they log intent, no real effects). To actually send mail / write files / call the network, supply your own implementations to build_registry(...); the AUTHORITY/CONTENT gating applies unchanged.

About

A zero-dependency, partitioned-trust LLM control layer: untrusted data can change what a model says, but provably not what it's allowed to do.

Topics

Resources

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages