Status: accepted; running-instance SSH implemented behind production activation and live verification (2026-07-14); in-dashboard Browser Web Shell over that same running instance added (w2/m55, 2026-07-17) and production-activated 2026-09-08 (w2/m90). Amended by ADR054 (w2/m65): a second target kind — the agent-session sandbox (ags-…@ssh.bex.co) — resolves through this same gateway with a per-connection multi-channel exception and an sftp-server subsystem bridge, both scoped to sandbox targets only; srv-… App targets keep the single-channel / no-subsystem / no-forwarding contract unchanged. The inline corrections below fold that amendment in. Regression + fix (w6/m132, 2026-08-28): the public handshake silently stopped sending KEXINIT after w4/m82 enabled Traefik's PROXY protocol without the gateway's matching BEX_SSH_PROXY_PROTOCOL_TRUSTED_CIDRS; see Regression. It is now guarded by a manifest lockstep check and an always-on KEXINIT liveness synthetic.
Render supports public-key SSH into paid web services, private services, and background workers. Its current CLI reads serviceDetails.sshAddress, verifies a live deploy, optionally lists GET /v1/services/{id}/instances, and invokes the user's local OpenSSH binary. This is narrower than an ephemeral shell product: it attaches to one container that is already serving the App.
bex previously grouped every exec-shaped feature under the hosted-execution non-goal. The user reopened running-instance SSH (2026-07-14) and, later, the in-dashboard Browser Web Shell over that same already-running instance (2026-07-17 — see Browser Web Shell). Ephemeral instances, one-off jobs, and hosted sandboxes remain excluded in .pm/DO_NOT_DO.md.
The security constraint is stronger than ordinary API reads: Kubernetes pods/exec can read the complete runtime environment and act with the workload's identity. bex-api must not inherit that permission.
bex runs /ssh-gateway as a third backend entrypoint in the shared lego image and as its own Deployment and ServiceAccount. It accepts SSH on container port 2222. Traefik's dedicated ssh TCP entrypoint exposes it on public port 22 in production.
The connection path is:
- OpenSSH connects to
<service-id>@ssh.bex.co, or<instance-id>@ssh.bex.cofor a selected replica. - The gateway accepts public-key authentication only. It hashes the presented key with OpenSSH SHA-256 and resolves that globally unique fingerprint in
ssh_keys. - After the client proves key possession, the gateway attaches the stored subject as a
core.Identitywith methodssh. - A composite resolver routes by id kind (ADR054 D1). A
srv-…username goes toapps.Service.ResolveSSHSession, which resolves the App through its stable id and calls the resource-scopedAuthorizeApp(can_view_sensitive)seam — the SINK's relation, since a shell reads the pod's env vars and mounted secrets (codex round-4 #8; the weakercan_operatean earlier draft named would let a contributorprintenvaround the boundary). Anags-…username goes to the agent-session resolver, which authorizescan_view_sensitiveon theagent_session:<id>object and derives the session's sandbox pod. Authorization therefore targets the resource's own workspace, not whichever workspace happens to be the caller's default. - For an App target the resolver selects a Running, Ready pod whose revision label and
appcontainer image equal the App's active status. A bare service id selects a random eligible replica, matching Render. A specific instance id selects exactly the pod returned by the instances endpoint. For an agent-session target it derives the live sandbox pod (<sandbox_id>-0in<workspace_id>-sandbox, containersandbox) — refused unless the session's phase implies a running sandbox. - One SSH
sessionchannel maps to one Kubernetespods/execstream in the target container. EXCEPTION (ADR054 D3): an agent-session sandbox target may open several concurrentsessionchannels over one connection — what Zed's ControlMaster remoting multiplexes — each its ownpods/execstream, bounded byBEX_SSH_MAX_CHANNELS_PER_CONN; App targets stay single-channel. No sshd or sidecar is injected into tenant images.
Keys are identity-scoped. One person can use the same key in every workspace where that identity has can_operate; a key is not copied into workspace membership or Kubernetes Secrets.
- REST:
GET/POST /v1/ssh-keys,DELETE /v1/ssh-keys/{id} - GraphQL:
sshKeys,createSSHKey,deleteSSHKey - MCP:
list_ssh_keys,add_ssh_key,delete_ssh_key - Dashboard: Account Settings → SSH Public Keys
Eligible service pages expose the command through a Connect → SSH menu and a left-nav Shell page, matching Render's information architecture. The Shell page shows the copy-ready running-instance SSH command and hosts an in-browser terminal (see Browser Web Shell); unavailable services explain the missing address without inventing one. Private keys never enter the browser — the in-browser terminal authenticates with the caller's dashboard session, not an SSH key.
The store contains typed ssk-… id, subject, display name, canonical public key, SHA-256 fingerprint, and creation time. It never accepts or stores private material. Supported key types match Render's documented set: Ed25519, RSA (minimum 2048 bits, authenticated only with SHA-2 signatures), ECDSA P-256/P-384/P-521, and OpenSSH security-key Ed25519/ECDSA. Comments are discarded. Multiple records, authorized_keys options, trailing payloads, oversized input, malformed text, and duplicate key material are rejected. Options are rejected rather than silently stripped because a user might otherwise believe an option such as command= still restricts the registered account key.
Render documents RSA support but no minimum size. The 2048-bit registration floor is bex's explicit security policy; it rejects obsolete weak RSA material instead of advertising a key that the gateway should not trust.
Any workspace member may manage their own public keys through can_manage_ssh_keys; only contributor/developer/admin identities can open sessions because target authorization separately requires can_operate. Deleting a key prevents every new handshake using it. Existing sessions are not retroactively killed by key deletion.
When BEX_SSH_HOST is set and the SSH-key store is available, an eligible service returns:
{
"serviceDetails": {
"sshAddress": "srv-d5t5d4v8g3c73f5m9peg@ssh.bex.co"
}
}The field is omitted for free, suspended, cron, static, or unsupported service types and when no public host is configured. Web, private, and background-worker services on a paid plan are eligible to advertise an address; the gateway still requires Running phase, an active revision, and a Ready current-image pod before authenticating.
GET /v1/services/{id}/instances returns Render's array of {id,createdAt} objects. Public instance ids are <service-id>-<opaque-suffix>, stably derived from the Pod name so live listing, metrics, and logs agree without a live UID lookup (historical Prom/Loki retain only names). Pre-m87 SSH selectors hashed the Pod UID; the resolver still accepts those legacy ids for the same Ready pod and fails closed on ambiguity. Raw pod names, UIDs, or Kubernetes credentials are never exposed on the wire. The general instances surface follows Render's rollout-observability contract: it includes non-terminating Pending, Running, and Unknown Deployment pods and excludes terminating, terminal, cron, and static pods. The SSH resolver then re-lists the service pods and narrows a requested id to a Running+Ready pod whose app.bex.co/revision label equals App.status.activeRevision and whose app-container image equals App.status.image. Build, pre-deploy, stale-revision, stale-image, and non-Ready pods therefore remain visible where appropriate for instance inspection but fail closed as SSH targets. The operator stamps the revision label on every app Deployment pod template, so a same-image restart cannot make the old ReplicaSet look current.
The server permits only:
- one
sessionchannel per SSH connection (EXCEPTION: an agent-session sandbox target may open up toBEX_SSH_MAX_CHANNELS_PER_CONNconcurrentsessionchannels, each its own exec stream — ADR054 D3); pty-reqbefore execution;window-changewhile a PTY session runs;- one
shellrequest, executed as/bin/sh; or - one
execrequest, executed as/bin/sh -lc <client-command>; or - the
sftpsubsystem, executed as the fixedsftp-serverbinary — ONLY on an agent-session sandbox target, so Zed can upload its remote-server when the sandbox's closed egress blocks a direct download (ADR054 D4). App targets reject every subsystem.
The command is one argv value, not interpolated into a Kubernetes URL or host command. It intentionally receives normal shell semantics inside the app container. The gateway does not log or persist it.
stdin, stdout, stderr, TTY state, and resize events bridge through client-go remotecommand. Kubernetes exit codes become SSH exit-status. Client disconnect, gateway shutdown, session timeout, restart, redeploy, or pod deletion cancels the stream. An image without executable /bin/sh fails with exit 126 and a bounded shell-unavailable message; bex never installs a shell into the image.
The following never reach Kubernetes: direct-tcpip, forwarding, agent forwarding, X11, SCP protocol handling, and environment requests — on every target. Subsystems and extra session channels are likewise banned EXCEPT the two ADR054 D3/D4 carve-outs above (multiple session channels and the sftp subsystem), which apply to agent-session sandbox targets only.
The gateway ServiceAccount has one Role in the configured tenant App namespace: get,list on apps.app.bex.co and pods, plus create on pods/exec. It has no ClusterRole or ClusterRoleBinding.
It cannot read Secrets, logs, Jobs, platform/auth namespaces, or arbitrary tenant namespaces. bex-api's separate Role remains read-only for pods and has no pods/exec rule. A structural manifest test guards this split and rejects any future cluster-wide gateway grant.
The gateway runs non-root with all capabilities dropped, a read-only root filesystem, resource limits, two replicas, a PodDisruptionBudget, and SSH ingress restricted to the Traefik namespace. A separate internal HTTP listener supplies health probes and Prometheus metrics; only the monitoring namespace may reach it. Metrics contain bounded authentication/session outcome and limit-scope labels—never subjects, service or instance ids, addresses, commands, environment values, or terminal content. Global and per-identity session caps default to 100 and 5. Handshakes default to 10 seconds and sessions to four hours.
The bex_ssh_gateway_authentications_total{result} counter is also the server-side proxy for the dashboard SSH-key onboarding funnel (w2/m66's RequiresSshKey gate): result="rejected_key" is an offered key not found in ssh_keys (the caller has not registered it — the dominant first-run failure the gate exists to pre-empt), versus result="accepted". The gate reduces the former, so "the gate is working" reads as rejected_key's share falling as accepted rises — e.g. sum(rate(bex_ssh_gateway_authentications_total{result="rejected_key"}[1h])) / sum(rate(bex_ssh_gateway_authentications_total[1h])) trending down. This exact metric root-caused the w2/m65 "Open in Zed" dead-end (a rejected_key-only spike with zero rejected_target).
The gateway host key is unrelated to the ~/.ssh/bex node-admin key described in ADR019. It is a dedicated, stable Ed25519 key installed out of band as bex-system/bex-ssh-host-key by scripts/ssh-host-key-secret.sh. The Deployment refuses to start when the Secret is absent or malformed.
The local hostNetwork Traefik baseline binds its SSH entrypoint on 2222, avoiding node sshd on port 22. The production Hetzner LoadBalancer exposes that entrypoint on 22. The ssh.bex.co A/AAAA records must be DNS-only and point directly at the Traefik LoadBalancer; an ordinary Cloudflare-proxied record cannot carry this raw SSH endpoint. scripts/ssh-dns-cloudflare.sh reconciles that exact DNS-only A/AAAA set and removes stale or duplicate address records. Its Cloudflare token needs zone-specific Zone / DNS / Edit; zone discovery additionally needs Zone / Zone / Read, which can be avoided by supplying CLOUDFLARE_ZONE_ID. The token is an ephemeral operator credential and must not be written to .env or a checked-in file. --check is read-only and rejects proxied, missing, stale, or duplicate address records.
scripts/ssh-activate.sh --check is a separate read-only preflight that requires the complete public A/AAAA set to equal the Terraform-owned Hetzner edge's public address set and verifies TCP/22 presents the mounted key's published fingerprint. Both SSH edge scripts query that exact named object with HCLOUD_TOKEN; they no longer depend on Kubernetes LoadBalancer status because the production Traefik Service is intentionally a NodePort. Running activation without --check performs the same gates before creating the activation ConfigMap and restarting bex-api. BEX_SSH_HOST must not be set before those checks pass.
Rotation is deliberate: create a new key, publish its fingerprint and maintenance window, replace the Secret with scripts/ssh-host-key-secret.sh (which rolls and waits for both gateway replicas), and tell clients to replace the known-host entry. Routine deploys never regenerate the key.
AuthorizeApp(can_view_sensitive) (App target) or AuthorizeOn(can_view_sensitive, agent_session:<id>) (sandbox target, ADR054 D2) records the allowed or denied connection authorization through the shared audit seam. A successful handshake additionally creates an ssn-… row in ssh_sessions, then closes it with end time and result. The table contains only subject, workspace, service id, instance id, remote address, start/end times, and result. The daily audit sweep deletes these rows after BEX_AUDIT_RETENTION_DAYS (default 90), the same boundary as audit_events.
The audit type and schema have no command, argv, environment, stdin, stdout, stderr, or terminal-content field. This structural omission is the privacy boundary; operators must not add ad hoc stream logging around it.
Authentication fails without revealing whether the cause was an unknown/deleted key, malformed username, missing/foreign service, insufficient role, unsupported/free/suspended service, missing live revision, or missing/non-Ready/stale instance. No session channel opens and no pods/exec request occurs.
After authorization, Kubernetes errors remain bounded. A pod disappearing during restart/redeploy closes the stream. The gateway never retries against another replica, because silently switching a specific session target would violate instance selection and could cross a deploy boundary.
The paragraph above governs failures at and after authentication, and its non-disclosure is a deliberate security boundary. A pre-authentication fault — the connection never completing the version exchange and key exchange, so the client never receives a KEXINIT — is a different case, and it is not a security boundary: it leaks nothing about a key, a service, or a role because it happens before any of those are consulted. Conflating the two is what let the w6/m132 regression (a Traefik PROXY header the gateway was misconfigured to not strip, fed into the version exchange) look like every other quiet refusal for weeks.
The honest-failure decision is therefore: the two are kept structurally distinguishable, and the pipeline fix plus observability — not a new in-band SSH message — carries it. A client that reaches authentication and is refused sees the transport succeed and the auth attempt fail (Permission denied (publickey)); a client hitting an infrastructure fault never gets a KEXINIT and its handshake dies before authentication. There is no protocol-legitimate way to say more before KEXINIT: SSH forbids application messages (a SSH_MSG_DISCONNECT reason, a banner) until the transport is established, and inventing one would mean forking the handshake. Server-side, the bex_ssh_gateway_handshakes_total{result} counter separates established from failed with no subject, address, or cause — so a dead edge is loud in metrics and the always-on liveness synthetic below fails, while authentications_total (which ADR035:106 governs) stays untouched and its opacity unchanged.
- Cross-workspace target guessing: stable service ids resolve through
AuthorizeApp; the resource's tenant label determines the OpenFGA object. - Key ambiguity: fingerprint uniqueness permits exactly one subject per public key.
- Deleted key races: every new handshake reads the database before accepting the offered key and rechecks it after signature verification, before target authorization; no key cache survives deletion.
- Pod-name confusion: clients receive a derived instance id, and resolution re-lists label-selected Ready pods rather than accepting a raw pod name.
- Stale rollout target: pod container image must match observed active image and the pod must be Ready and non-terminating.
- Command injection into gateway/Kubernetes: the SSH command is passed as one
/bin/sh -lcargv element; it is never evaluated by the gateway host or embedded into a URL. - Session exhaustion: bounded handshake/session durations and global/per-identity counters reject excess connections.
- Compromised gateway: namespaced pod/exec RBAC limits the Kubernetes blast radius to the configured App namespace; it cannot read platform Secrets. This remains a powerful tenant-workload privilege, which is why the binary and ServiceAccount are isolated. Database blast radius (w7/m56): the gateway no longer shares bex-api's full-privilege
bex-db-appcredential. It connects as the scopedbex_ssh_gatewayPostgres role —SELECTonssh_keys/tenants/tenant_members(authenticate the caller + resolve their workspace/membership),INSERT/UPDATEon its ownssh_sessions,SELECT/INSERT/DELETEonshell_ticket_nonces, andINSERTonaudit_events— and nothing else. A stolen gateway credential therefore cannot read billing/usage (stripe_billing_events,usage_hourly), tenant credentials (registry_credentials,git_connections,sandbox_tenant_keys), or app/domain/deploy rows: those tables are never granted, and a fresh role has no default table access. This is enforced by Postgres, not by the narrow GoStoreinterface — verified negatively (aSELECTon the sensitive tables is refused with SQLSTATE42501) by the CI-runTestGatewayScopedRoleAllowsOwnSurfaceDeniesTheRest(internal/sshgateway/dbrole/dbrole_integration_test.go), which provisions the role from the sameinternal/sshgateway/dbrole/dbrole.sqlthe productionscripts/ssh-gateway-db-role.shapplies (one DDL source, so the tested and shipped boundaries cannot drift). The gateway is a least-privilege consumer of the control-plane store: it never runs migrations or the ownership check (those stay with bex-api onbex-db-app), which is what lets it run under a role with no DDL privilege. Note that resolving SSH authorization legitimately needs tenant/membership reads, so those are granted (not denied) — the boundary is money + credentials + resource data, which the gateway never touches. - Host-key substitution: stable out-of-band custody and published fingerprints make unexpected replacement visible to OpenSSH.
Render's Shell page hosts an in-browser terminal alongside the copy-ready ssh command. bex matches this (w2/m55) without weakening the isolation above: the browser terminal rides the same isolated gateway, KubeExecutor, Render-compatible instance targeting, AuthorizeApp(can_operate), session caps, and content-free audit as native SSH. It adds only a browser transport and a session-authenticated ticket path. It does not give bex-api pods/exec, and it never puts an SSH private key in the browser.
- In the dashboard Shell page the authenticated caller (a Kratos session, HttpOnly cookie) requests a short-lived exec ticket from bex-api:
POST /v1/services/{id}/shell-ticket(GraphQLcreateShellSession). bex-api runsAuthorizeApp(can_operate)and the same eligibility gate assshAddress(paid, non-suspended web/private/worker, Running with a live revision), then returns{ticket, url, expiresAt}. - The ticket is an HMAC-SHA256-signed token over
{subject, serviceId, instanceId?, issuedAt, expiresAt, nonce}, signed withBEX_SHELL_TICKET_SECRET— a secret shared only between bex-api and the gateway. bex-api can mint it; it cannot exec. The TTL is short (seconds) and the nonce makes it single-use across every gateway replica (w1/042 L7): each redemption atomically claims the nonce in the shared control-plane store (shell_ticket_nonces,INSERT … ON CONFLICT— exactly one replica wins; expired rows are pruned by the next claim), backed by a per-process map as the cheap same-replica first line. A store error refuses the ticket — the session-audit write requires the same database, so this adds no availability dependency. - The browser opens a WebSocket to the gateway at
url(wss://…/shell, Traefik-terminated to the gateway's plain-HTTPBEX_SHELL_WS_ADDR), carrying the ticket exclusively in aSec-WebSocket-Protocolentry (bex.ticket.<ticket>alongside thebex.shellmarker, which is the only protocol the gateway ever selects) — never in the URL, whoseRequestPathTraefik's access log keeps while headers are dropped (w1/042 L8). Query-string tickets are rejected. The gateway verifies the ticket signature, expiry, and single-use, attachescore.Identity{Subject}, and calls the sameResolveSSHSessionused by native SSH — so the gateway re-authorizescan_operateagainst the App's workspace and re-selects a Ready, current-image, current-revision pod. bex-api's authorization is not trusted transitively. - One WebSocket maps to one
pods/execstream in theappcontainer (/bin/sh, TTY). Browser stdin/stdout are binary WebSocket frames; JSON text frames carry terminal resize (mapped to the sameTerminalSizeQueue) and the terminal exit/error. The gateway never logs or persists the stream.
The re-drawn boundary is: terminal bytes reach the browser, but only over an authenticated, can_operate-gated gateway WebSocket, and only for an already-running instance. pods/exec stays confined to the gateway ServiceAccount; bex-api gains no exec permission (the structural RBAC test still guards this). A browser session counts against the same global/per-identity caps and writes the same ssn-… ssh_sessions audit row (subject, workspace, service, instance, remote address, start/end, result) — with no command, argv, or terminal-content field. Metrics stay bounded exactly as for SSH.
Unset BEX_SHELL_TICKET_SECRET (on bex-api or the gateway) disables the Web Shell: bex-api returns 503 for the ticket verb and the gateway does not start its WebSocket listener. That is the byte-identical default; native ssh is unaffected.
The Web Shell rides the same isolated gateway as native SSH, so it inherits that fail-closed sequence and adds one edge route and one shared secret. In order: deploy the image and manifests (they already wire BEX_SHELL_TICKET_SECRET into both bex-api and the gateway as an optional Secret ref, set the gateway's BEX_SHELL_WS_ADDR to :8080, and set bex-api's BEX_SHELL_WS_URL to wss://ssh.bex.co/shell); confirm the gateway Service exposes the plain-HTTP shell port (8080), the ssh-shell Ingress publishes wss://ssh.bex.co/shell with a valid cert-manager TLS cert, and the gateway NetworkPolicy admits Traefik → 8080; and only then install the shared HMAC key with scripts/shell-ticket-secret.sh (the same bex-shell-ticket Secret both Deployments mount). Installing that Secret is the single activation flip: before it, createShellSession honestly 503s ("web shell transport not configured") and the gateway starts no WebSocket listener; after it, the script rolls both Deployments and the dashboard terminal connects. BEX_SHELL_WS_URL is static — the Secret, not the URL, gates the feature — so the URL may be present before the edge is reachable without ever minting a redeemable-nowhere ticket.
Activated 2026-09-08 (w2/m90). Production already had the Secret, the :8080 listener, and Ingress bex-ssh-shell; this milestone verified the public refusal shape, recorded a live dashboard.bex.co Web Shell session plus the fail-closed matrix, and added the always-on ticketless WS probe (scripts/shell-ws-probe.sh on ssh-edge-liveness.yml) plus a gitops-validate.sh lockstep check so the edge cannot die silently again. Sanitized evidence: .pm/w2/done/m90/evidence/2026-09-08-production-activation.md.
This ADR does not authorize ephemeral shell instances, --ephemeral, one-off jobs, cron shell access, hosted sandboxes, SFTP/SCP, forwarding, agent forwarding, direct TCP/Unix-socket channels, shell installation, or an sshd sidecar. The in-dashboard Browser Web Shell attaches to an already-running instance and is in scope; ephemeral and hosted execution remain excluded.
Deterministic tests cover parsing, canonical fingerprints, duplicate/foreign ownership, REST/GraphQL/MCP parity, service eligibility, stale-revision/non-Ready filtering, any/specific-instance resolution, a real in-process SSH handshake, PTY resize, exit status, deleted-key and forbidden-target denial, forwarding rejection, session caps, authenticated-idle timeout/release, and content-free audit writes.
An opt-in real-cluster test (TestGatewayRealKubernetesExec) covers the boundary those fakes cannot: SSH protocol → client-go SPDY → a disposable pod's app container. Against the local CAPD app cluster on 2026-07-15 it read a runtime environment value, propagated exit 37, and closed the attached stream when pod deletion completed its 30-second termination grace. Its shell-less sibling drove a real traefik/whoami container and received the bounded exit-126 response. The tests require an explicit kubeconfig and disposable pod names; the deletion check additionally requires BEX_TEST_SSH_DELETE_POD=1, so ordinary go test ./... never mutates a cluster.
Production activation and the complete public-edge matrix passed on 2026-07-17. ssh.bex.co resolved directly to the production IPv4/IPv6 edge, presented the published stable fingerprint, and the 2/2 Ready gateway completed scripts/ssh-verify.sh with raw OpenSSH and the current unmodified Render CLI. The run covered any/exact instance targeting, runtime environment, PTY resize, exit status, restart/redeploy closure, and every lifecycle/authorization/type denial below. Sanitized identifiers and pass/fail evidence are recorded in the production acceptance.
The production sequence is intentionally fail-closed: deploy the image and manifests; install the dedicated stable host-key Secret; expose Traefik's ssh entrypoint; reconcile and check the direct, DNS-only A/AAAA records with scripts/ssh-dns-cloudflare.sh; verify the public fingerprint; and only then run scripts/ssh-activate.sh to advertise sshAddress. If any earlier step is absent, bex-api continues omitting the address.
That 2026-07-17/18 acceptance was real, and then nothing re-checked it — so a regression six weeks later went unnoticed until the 71st /qa-find-bugs run (2026-08-28) found ssh.bex.co:22 writing its SSH-2.0-bex banner and then never sending KEXINIT, for every target alike. The cause was a half-completed change: w4/m82 (2026-08-16) turned on proxyProtocol: version: 2 on the ssh IngressRouteTCP so Traefik forwards the real client address, but shipped the gateway Deployment without the matching BEX_SSH_PROXY_PROTOCOL_TRUSTED_CIDRS. With no trusted CIDR, proxyproto.Wrap is a no-op, so the un-stripped binary PROXY v2 header is fed straight into the SSH version exchange; Go's readVersion discards it as non-SSH- preamble and then blocks waiting for a valid version line until the 10 s BEX_SSH_HANDSHAKE_TIMEOUT closes the socket. KEXINIT is only sent after a completed version exchange, so it never arrives — the exact reported symptom. The sibling pg-sni-proxy/kv-sni-proxy never regressed because they set their BEX_PROXY_PROTOCOL_TRUSTED_CIDRS; only the SSH gateway's half was missing.
The fix sets BEX_SSH_PROXY_PROTOCOL_TRUSTED_CIDRS to the cluster pod CIDR on the gateway Deployment (Traefik is its only admitted :2222 peer per networkpolicy.yaml, so trusting the pod range is safe), and a scripts/gitops-validate.sh check now fails CI whenever the route sends a PROXY header without the Deployment setting the matching trust — the two halves can no longer drift silently. Two guards make the class detectable rather than silent: the manifest lockstep check above, and an always-on liveness synthetic — the credential-free KEXINIT probe (internal/sshgateway/nativessh.ProbeKEXINIT, driven by scripts/ssh-kexinit-probe.sh on the scheduled ssh-edge-liveness.yml) — that fails within hours of a dead edge instead of on the next human's manual check. Reproduced live under that probe on 2026-08-28: github.com and gitlab.com sent KEXINIT, ssh.bex.co sent its banner then timed out. Live re-verification after the fix's production rollout (the raw probe and the full scripts/ssh-verify.sh matrix) is owed on deploy.
The verifier requires the API URL/token, a disposable private-key file, and the published host-key fingerprint. By default it creates and cleans up a paid, shell-capable two-replica service; an explicitly supplied existing service must be disposable because the matrix restarts, suspends, resizes, and temporarily rolls it to a shell-less image. The default fixture uses one healthy port for both its shell-capable BusyBox server and the shell-less whoami rollout, so a real Ready redeploy must replace the attached pod before the exit-126 check. Readiness also requires a live deploy and retries transient edge errors within its deadline. The matrix registers only the derived public key, pins the observed TCP/22 host key, checks runtime environment, PTY resize, any- and specific-instance raw SSH, exit status, stale-instance denial, separate restart- and redeploy-driven stream closure, suspended/free/shell-less behavior, and unknown/deleted keys. BEX_RENDER_CLI_VERIFY=1 adds interactive CLI checks by service name, service id with the Any instance picker option, and direct complete instance id; the three-selector supplement passed against production on 2026-07-18. BEX_RENDER_CLI_BIN can pin an exact downloaded release binary instead of relying on PATH. The script supplies RENDER_HOST, RENDER_API_KEY, RENDER_WORKSPACE, and an isolated RENDER_CLI_CONFIG_PATH to the unmodified binary without persisting their values, pins the OpenSSH host key/private key explicitly, and asserts the destination the CLI selected. The current official release, render-oss/cli v2.21.0 at c398207, still has an upstream defect in both instance-picker callbacks (they drop the selected instance id), so exact-replica acceptance uses the CLI's supported full-instance-id argument instead of claiming that broken menu path works. BEX_SSH_VERIFY_FULL_MATRIX=1 additionally requires out-of-band viewer, foreign-workspace, static, and cron fixtures, because the primary acceptance identity must not be able to manufacture its own weaker role or foreign workspace.