Skip to content

About

An orchestration-layer firewall that neutralizes LLM prompt injection (94.4% on a 20-attack red-team battery) and Ed25519-signs every turn — plus a ChatGPT-style chat app that runs locally with no API key.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

firewall-bridge

A prompt-injection firewall for any LLM — with a private, streaming, cryptographically-signed chat app on top. No paid API key required.

CI python license

Try it in 30 seconds (no key, offline demo model):

pip install cryptography
python secure_chat.py        # opens a chat UI at http://localhost:8000

Pick a real model with no API key from the dropdown: Local — Ollama (private, on your machine) or Free online (llm7.io). Every message is firewalled and Ed25519-signed either way.


A channel firewall that sits in front of any LLM, plus a ChatGPT/Claude-style chat app built on top of it. It is the deployable, orchestration-layer realization of CIPHER's pillar I (non-interference): keep trusted instructions and untrusted data in structurally separate channels, and never let untrusted data act as an instruction.

You can't split a hosted model's residual stream — but before a single token reaches the model you can type every input by trust, neutralize known injection vectors, seal untrusted content behind an unguessable fence, and sign the result. That is what this does.

The chat app (start here)

pip install cryptography
python secure_chat.py                 # opens a chat UI at http://localhost:8000
# optional real model (else an offline demo model is used):
#   set ANTHROPIC_API_KEY=...   or   set OPENAI_API_KEY=...

secure_chat.py is one self-contained file: a browser chat interface with multi-turn memory, a system-prompt editor, an "attach untrusted context" box for testing injection defense, and a live security panel under every reply showing what the firewall caught, the grounding support, and the Ed25519 provenance signature. Replies stream in token-by-token (real streaming from Claude/GPT, chunked for the demo model) and render as markdown (code blocks, lists, headings). With no API key it runs an offline demo model; set a key to chat with a real Claude or GPT model — every turn firewalled and signed either way.

Each turn is signed, so you can download a signed transcript and verify it independently — nothing has to trust the server:

python secure_chat.py verify transcript.json   # re-checks every Ed25519 signature

The same file also runs redteam, demo, test, and selfcheck.

Model backends (including keyless ones)

The chat app can run on several backends, chosen from a dropdown in the UI (or FB_BACKEND=free|local|offline at launch):

Backend Real model? API key? Notes
Offline demo (default) no (rule-based) none fully local & private; there to exercise the firewall
Free online — llm7.io yes none a real model over the internet with no key, via an OpenAI-compatible keyless endpoint
Local — Ollama / LM Studio yes none fully private; install Ollama, ollama run llama3.2
Claude / GPT yes ANTHROPIC_API_KEY / OPENAI_API_KEY your own account

So yes — there is an online option that needs no API key: select "Free online", and it streams from https://api.llm7.io/v1 (OpenAI-compatible, keyless). Because these free catalogs rotate model IDs, the app uses the stable default alias and probes /v1/models on switch to lock onto a model that is actually being served, so a rotated-out model name doesn't 400. Honest caveat, which the UI states plainly: a free public endpoint is rate-limited, rotates models, and may log your prompts. The firewall still neutralizes injection and signs every turn, but prompt confidentiality is not guaranteed on a free third-party service — for private real-model use, run Local (Ollama), which keeps everything on your machine with no key either.

All of this rides on one generic OpenAICompatProvider (standard-library HTTP, no openai package), so you can point it at any OpenAI-compatible endpoint.

pip install cryptography          # only hard dependency
python demo.py                    # offline end-to-end demo
pytest -q                         # 16 checks, all green
python -m firewall_bridge         # smoke self-check

30-second use

from firewall_bridge import ChannelFirewall, MockProvider

fw = ChannelFirewall(MockProvider())          # swap in a real provider below
r = fw.run(
    system="You are a careful assistant.",    # trusted
    user="Summarize this document.",          # the principal
    retrieved=[poisoned_web_page],            # UNTRUSTED - injection lives here
)

r.output        # the model's answer (injection did not take effect)
r.findings      # audit: which injection vectors were caught, and where
r.grounding     # per-sentence support against the sources
r.verify()      # True -> the signed transcript is authentic and unaltered
r.messages      # exactly what was sent to the model (safe, channel-separated)

Against a real model

Adapters import their SDK lazily, so the package installs without them.

from firewall_bridge import ChannelFirewall, AnthropicProvider, OpenAIProvider

fw = ChannelFirewall(AnthropicProvider(model_id="claude-sonnet-4-5"))
# or: ChannelFirewall(OpenAIProvider(model_id="gpt-4o-mini"))
r = fw.run(system=SYSTEM, user=user_msg, retrieved=rag_chunks, sources=rag_chunks)

Keys are read from ANTHROPIC_API_KEY / OPENAI_API_KEY. The Anthropic adapter passes trusted instructions through the dedicated system parameter, keeping the trusted channel separate at the API boundary as well.

The two enforcement layers

Injection defense here is defense-in-depth, layered on purpose so neither layer has to be perfect:

A — Sanitize (sanitize.py). Pattern-detects and neutralizes high-signal vectors in untrusted content: instruction overrides, persona/role changes, system-prompt exfiltration, role-header spoofing, chat-markup breakouts, tool-hijack phrasing. Findings go to the audit log. Policy chooses QUARANTINE (neutralize, keep as data), STRIP (delete), or BLOCK (refuse).

B — Spotlight (compose.py). Wraps untrusted content in a fence keyed by a per-request random nonce, datamarks every untrusted line with that nonce, and gives the system channel an explicit contract: text inside the fence is data, never instructions. Because the nonce is unpredictable, untrusted content cannot forge a fence-close to break out — and this layer needs no list of known attacks, which is exactly why it complements layer A.

Then the output is grounding-checked against the sources and the whole transcript is signed.

Hardening (obfuscation-aware) and the red-team benchmark

Pattern matching alone is trivially bypassed, so with FirewallPolicy(harden=True) the sanitizer first canonicalizes (unicode NFKC, strips zero-width/bidi characters, folds Cyrillic/Greek homoglyphs) and then decodes and re-scans base64 / hex / rot13 / de-spaced / l33t variants, redacting encoded blobs that decode to instructions. A canary token can be planted in the system prompt; if it appears in the output, the system prompt was exfiltrated and the result is flagged (r.canary_leaked).

redteam.py fires a categorized battery of 20 injection/jailbreak payloads (plain, exec-marker, role-spoof, chat-markup, homoglyph, zero-width, base64, hex, leetspeak, de-spacing, persona, audit-pretext, and a novel semantic attack) through a naive pipeline and the hardened firewall, using a capable offline model that decodes obfuscations the way a real model would:

python redteam.py
...
indirect   naive lands 10/12  firewall lands  0/12  protection 100.0%
direct     naive lands  8/8   firewall lands  1/8   protection  87.5%
OVERALL    naive lands 18/20  firewall lands  1/20  protection  94.4%

The reading is deliberately honest: indirect injection is stopped structurally (the nonce fence doesn't depend on recognizing the attack, so obfuscation doesn't help). Direct attacks with any detectable trigger are neutralized by canonicalization + deep scan. The one payload that survives is a purely semantic jailbreak with no encoding and no keyword — which is model-alignment territory, not something an orchestration layer can catch. A firewall that claimed to stop it would be lying.

What the tests prove (pytest -q)

Property Test
Injection in retrieved content does not hijack the model test_injection_in_retrieved_does_not_hijack
The same attack does hijack a naive pipeline (control) test_same_attack_hijacks_a_naive_pipeline
Untrusted data may still inform the answer test_untrusted_data_still_informs_the_answer
The fence stops imperatives even with the sanitizer bypassed test_fence_isolates_even_unsanitized_imperatives
Untrusted content cannot forge the nonce fence test_nonce_is_unforgeable_and_untrusted_cannot_break_out
Known vectors are flagged; raw instruction removed test_sanitizer_*, test_neutralize_*
BLOCK refuses; benign requests pass through test_block_policy_*, test_benign_*
Provenance verifies and detects tampering; keys persist test_provenance_*
Grounding separates supported from fabricated claims test_grounding_*

Package layout

firewall_bridge/
  channels.py     trust levels + trust-typed Segment/Envelope
  policy.py       FirewallPolicy (quarantine/strip/block, harden, canary, limits)
  sanitize.py     layer A: injection detection + neutralization (+ deep/obfuscation-aware)
  canonicalize.py layer A+: unicode/homoglyph/zero-width fold + base64/hex/rot13 decode
  compose.py      layer B: nonce-fence spotlighting + datamarking
  providers.py    Mock/Capability/Echo (offline) + lazy OpenAI/Anthropic adapters
  grounding.py    attribution gate (lexical proxy for an NLI model)
  provenance.py   Merkle + Ed25519 signing, key save/load, verify
  firewall.py     ChannelFirewall orchestrator + FirewallResult
  __main__.py     python -m firewall_bridge self-check
redteam.py        categorized attack battery + scored naive-vs-firewall benchmark
demo.py           offline end-to-end demo

Honest limits

  • The sanitizer's pattern list is not complete; treat it as one layer, not a guarantee. The nonce spotlight is the layer that doesn't depend on enumerating attacks.
  • Against a real model, non-interference is empirical: a model can still be coaxed into treating fenced data as instructions. This raises the bar substantially and makes every attempt auditable and signed — it does not make injection impossible. That honesty is the point; a firewall that claims perfection is lying.
  • The grounding score is a lexical proxy. Swap in a trained entailment model for production strength.
  • MockProvider is a stand-in that lets the guarantees be tested offline; it models an agent that honors the fence contract so the pipeline's effect is visible without keys.

Repository layout

firewall_bridge/     the package (channels, sanitize, canonicalize, compose,
                     providers, grounding, provenance, firewall, chat, verify)
tests/               pytest suite (also proves the guarantees at runtime)
secure_chat.py       single-file build of the whole app (just `python secure_chat.py`)
chat_app.py          the web app (uses the package)
redteam.py           categorized attack battery + scored benchmark
demo.py              non-interactive attack/defense demo
docs/OLLAMA_SETUP.md  run a real model locally with no key
docs/THEORY-CIPHER.md the mathematical foundation (non-interference, etc.)
theory/cipher_all.py  runnable reference implementation of the CIPHER theory
.github/workflows/    CI (pytest on every push)

The theory behind it (optional reading)

This firewall is the deployable half of a two-part project. The other half, CIPHER (docs/THEORY-CIPHER.md, theory/cipher_all.py), is a from-scratch model architecture whose "pillar I" proves non-interference bit-exactly — untrusted input provably cannot alter control flow. You can't do that to a hosted model you don't own, so this repo brings the same idea to the orchestration layer, where it becomes strong-but-empirical instead of provable. Run the theory with python theory/cipher_all.py.

Security

Please do not commit secrets. API keys and tokens are read from environment variables (ANTHROPIC_API_KEY, OPENAI_API_KEY, FB_FREE_KEY) and never stored in code; .gitignore excludes .env, *.pem, and saved transcripts. If you find a vulnerability, open an issue (or a private advisory) rather than a public PR with a working exploit.

Contributing

PRs welcome. Run pytest -q before submitting; new firewall behavior should come with a test that turns the property into an assertion (see tests/test_firewall.py for the pattern). Keep the "honest limits" spirit — document what a change does not protect against, too.

License

MIT — see LICENSE. Replace <YOUR NAME> in the license file with your name before publishing.

About

An orchestration-layer firewall that neutralizes LLM prompt injection (94.4% on a 20-attack red-team battery) and Ed25519-signs every turn — plus a ChatGPT-style chat app that runs locally with no API key.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages