A prompt-injection firewall for any LLM — with a private, streaming, cryptographically-signed chat app on top. No paid API key required.
Try it in 30 seconds (no key, offline demo model):
pip install cryptography
python secure_chat.py # opens a chat UI at http://localhost:8000Pick a real model with no API key from the dropdown: Local — Ollama (private, on your machine) or Free online (llm7.io). Every message is firewalled and Ed25519-signed either way.
A channel firewall that sits in front of any LLM, plus a ChatGPT/Claude-style chat app built on top of it. It is the deployable, orchestration-layer realization of CIPHER's pillar I (non-interference): keep trusted instructions and untrusted data in structurally separate channels, and never let untrusted data act as an instruction.
You can't split a hosted model's residual stream — but before a single token reaches the model you can type every input by trust, neutralize known injection vectors, seal untrusted content behind an unguessable fence, and sign the result. That is what this does.
pip install cryptography
python secure_chat.py # opens a chat UI at http://localhost:8000
# optional real model (else an offline demo model is used):
# set ANTHROPIC_API_KEY=... or set OPENAI_API_KEY=...secure_chat.py is one self-contained file: a browser chat interface with
multi-turn memory, a system-prompt editor, an "attach untrusted context" box for
testing injection defense, and a live security panel under every reply
showing what the firewall caught, the grounding support, and the Ed25519
provenance signature. Replies stream in token-by-token (real streaming from
Claude/GPT, chunked for the demo model) and render as markdown (code blocks,
lists, headings). With no API key it runs an offline demo model; set a key to
chat with a real Claude or GPT model — every turn firewalled and signed either
way.
Each turn is signed, so you can download a signed transcript and verify it independently — nothing has to trust the server:
python secure_chat.py verify transcript.json # re-checks every Ed25519 signatureThe same file also runs redteam, demo, test, and selfcheck.
The chat app can run on several backends, chosen from a dropdown in the UI (or
FB_BACKEND=free|local|offline at launch):
| Backend | Real model? | API key? | Notes |
|---|---|---|---|
| Offline demo (default) | no (rule-based) | none | fully local & private; there to exercise the firewall |
| Free online — llm7.io | yes | none | a real model over the internet with no key, via an OpenAI-compatible keyless endpoint |
| Local — Ollama / LM Studio | yes | none | fully private; install Ollama, ollama run llama3.2 |
| Claude / GPT | yes | ANTHROPIC_API_KEY / OPENAI_API_KEY |
your own account |
So yes — there is an online option that needs no API key: select "Free online",
and it streams from https://api.llm7.io/v1 (OpenAI-compatible, keyless). Because
these free catalogs rotate model IDs, the app uses the stable default alias and
probes /v1/models on switch to lock onto a model that is actually being served,
so a rotated-out model name doesn't 400. Honest caveat, which the UI states plainly: a free public endpoint is rate-limited,
rotates models, and may log your prompts. The firewall still neutralizes
injection and signs every turn, but prompt confidentiality is not guaranteed
on a free third-party service — for private real-model use, run Local (Ollama),
which keeps everything on your machine with no key either.
All of this rides on one generic OpenAICompatProvider (standard-library HTTP,
no openai package), so you can point it at any OpenAI-compatible endpoint.
pip install cryptography # only hard dependency
python demo.py # offline end-to-end demo
pytest -q # 16 checks, all green
python -m firewall_bridge # smoke self-checkfrom firewall_bridge import ChannelFirewall, MockProvider
fw = ChannelFirewall(MockProvider()) # swap in a real provider below
r = fw.run(
system="You are a careful assistant.", # trusted
user="Summarize this document.", # the principal
retrieved=[poisoned_web_page], # UNTRUSTED - injection lives here
)
r.output # the model's answer (injection did not take effect)
r.findings # audit: which injection vectors were caught, and where
r.grounding # per-sentence support against the sources
r.verify() # True -> the signed transcript is authentic and unaltered
r.messages # exactly what was sent to the model (safe, channel-separated)Adapters import their SDK lazily, so the package installs without them.
from firewall_bridge import ChannelFirewall, AnthropicProvider, OpenAIProvider
fw = ChannelFirewall(AnthropicProvider(model_id="claude-sonnet-4-5"))
# or: ChannelFirewall(OpenAIProvider(model_id="gpt-4o-mini"))
r = fw.run(system=SYSTEM, user=user_msg, retrieved=rag_chunks, sources=rag_chunks)Keys are read from ANTHROPIC_API_KEY / OPENAI_API_KEY. The Anthropic adapter
passes trusted instructions through the dedicated system parameter, keeping the
trusted channel separate at the API boundary as well.
Injection defense here is defense-in-depth, layered on purpose so neither layer has to be perfect:
A — Sanitize (sanitize.py). Pattern-detects and neutralizes high-signal
vectors in untrusted content: instruction overrides, persona/role changes,
system-prompt exfiltration, role-header spoofing, chat-markup breakouts,
tool-hijack phrasing. Findings go to the audit log. Policy chooses QUARANTINE
(neutralize, keep as data), STRIP (delete), or BLOCK (refuse).
B — Spotlight (compose.py). Wraps untrusted content in a fence keyed by a
per-request random nonce, datamarks every untrusted line with that nonce, and
gives the system channel an explicit contract: text inside the fence is data,
never instructions. Because the nonce is unpredictable, untrusted content
cannot forge a fence-close to break out — and this layer needs no list of
known attacks, which is exactly why it complements layer A.
Then the output is grounding-checked against the sources and the whole transcript is signed.
Pattern matching alone is trivially bypassed, so with FirewallPolicy(harden=True)
the sanitizer first canonicalizes (unicode NFKC, strips zero-width/bidi
characters, folds Cyrillic/Greek homoglyphs) and then decodes and re-scans
base64 / hex / rot13 / de-spaced / l33t variants, redacting encoded blobs that
decode to instructions. A canary token can be planted in the system prompt;
if it appears in the output, the system prompt was exfiltrated and the result is
flagged (r.canary_leaked).
redteam.py fires a categorized battery of 20 injection/jailbreak payloads
(plain, exec-marker, role-spoof, chat-markup, homoglyph, zero-width, base64, hex,
leetspeak, de-spacing, persona, audit-pretext, and a novel semantic attack)
through a naive pipeline and the hardened firewall, using a capable offline
model that decodes obfuscations the way a real model would:
python redteam.py
...
indirect naive lands 10/12 firewall lands 0/12 protection 100.0%
direct naive lands 8/8 firewall lands 1/8 protection 87.5%
OVERALL naive lands 18/20 firewall lands 1/20 protection 94.4%
The reading is deliberately honest: indirect injection is stopped structurally (the nonce fence doesn't depend on recognizing the attack, so obfuscation doesn't help). Direct attacks with any detectable trigger are neutralized by canonicalization + deep scan. The one payload that survives is a purely semantic jailbreak with no encoding and no keyword — which is model-alignment territory, not something an orchestration layer can catch. A firewall that claimed to stop it would be lying.
| Property | Test |
|---|---|
| Injection in retrieved content does not hijack the model | test_injection_in_retrieved_does_not_hijack |
| The same attack does hijack a naive pipeline (control) | test_same_attack_hijacks_a_naive_pipeline |
| Untrusted data may still inform the answer | test_untrusted_data_still_informs_the_answer |
| The fence stops imperatives even with the sanitizer bypassed | test_fence_isolates_even_unsanitized_imperatives |
| Untrusted content cannot forge the nonce fence | test_nonce_is_unforgeable_and_untrusted_cannot_break_out |
| Known vectors are flagged; raw instruction removed | test_sanitizer_*, test_neutralize_* |
BLOCK refuses; benign requests pass through |
test_block_policy_*, test_benign_* |
| Provenance verifies and detects tampering; keys persist | test_provenance_* |
| Grounding separates supported from fabricated claims | test_grounding_* |
firewall_bridge/
channels.py trust levels + trust-typed Segment/Envelope
policy.py FirewallPolicy (quarantine/strip/block, harden, canary, limits)
sanitize.py layer A: injection detection + neutralization (+ deep/obfuscation-aware)
canonicalize.py layer A+: unicode/homoglyph/zero-width fold + base64/hex/rot13 decode
compose.py layer B: nonce-fence spotlighting + datamarking
providers.py Mock/Capability/Echo (offline) + lazy OpenAI/Anthropic adapters
grounding.py attribution gate (lexical proxy for an NLI model)
provenance.py Merkle + Ed25519 signing, key save/load, verify
firewall.py ChannelFirewall orchestrator + FirewallResult
__main__.py python -m firewall_bridge self-check
redteam.py categorized attack battery + scored naive-vs-firewall benchmark
demo.py offline end-to-end demo
- The sanitizer's pattern list is not complete; treat it as one layer, not a guarantee. The nonce spotlight is the layer that doesn't depend on enumerating attacks.
- Against a real model, non-interference is empirical: a model can still be coaxed into treating fenced data as instructions. This raises the bar substantially and makes every attempt auditable and signed — it does not make injection impossible. That honesty is the point; a firewall that claims perfection is lying.
- The grounding score is a lexical proxy. Swap in a trained entailment model for production strength.
MockProvideris a stand-in that lets the guarantees be tested offline; it models an agent that honors the fence contract so the pipeline's effect is visible without keys.
firewall_bridge/ the package (channels, sanitize, canonicalize, compose,
providers, grounding, provenance, firewall, chat, verify)
tests/ pytest suite (also proves the guarantees at runtime)
secure_chat.py single-file build of the whole app (just `python secure_chat.py`)
chat_app.py the web app (uses the package)
redteam.py categorized attack battery + scored benchmark
demo.py non-interactive attack/defense demo
docs/OLLAMA_SETUP.md run a real model locally with no key
docs/THEORY-CIPHER.md the mathematical foundation (non-interference, etc.)
theory/cipher_all.py runnable reference implementation of the CIPHER theory
.github/workflows/ CI (pytest on every push)
This firewall is the deployable half of a two-part project. The other half,
CIPHER (docs/THEORY-CIPHER.md, theory/cipher_all.py), is a from-scratch
model architecture whose "pillar I" proves non-interference bit-exactly —
untrusted input provably cannot alter control flow. You can't do that to a hosted
model you don't own, so this repo brings the same idea to the orchestration
layer, where it becomes strong-but-empirical instead of provable. Run the theory
with python theory/cipher_all.py.
Please do not commit secrets. API keys and tokens are read from environment
variables (ANTHROPIC_API_KEY, OPENAI_API_KEY, FB_FREE_KEY) and never stored
in code; .gitignore excludes .env, *.pem, and saved transcripts. If you
find a vulnerability, open an issue (or a private advisory) rather than a public
PR with a working exploit.
PRs welcome. Run pytest -q before submitting; new firewall behavior should come
with a test that turns the property into an assertion (see tests/test_firewall.py
for the pattern). Keep the "honest limits" spirit — document what a change does
not protect against, too.
MIT — see LICENSE. Replace <YOUR NAME> in the license file with your
name before publishing.