A spatial computing runtime for the browser. Ten holographic modules orbit you in a volumetric environment; you rotate, grab, throw and expand them with your hands.
Phase 01 — the surface, the gesture engine and the physics. No AI yet.
npm install
npm run dev # http://localhost:3000predev / prebuild vendor MediaPipe's wasm runtime and the hand-landmark model
into public/, so nothing is fetched from a CDN at runtime. The model download
needs network access once; if it fails the app boots in pointer mode and you can
retry with npm run assets.
Hand tracking is the primary input. The pointer is a first-class fallback, not a degraded one — see Pointer parity below.
| Gesture | Action |
|---|---|
| Swipe ◀ / ▶ (open hand) | Rotate the ring one slot |
| Tap (one finger) | Select what you're pointing at |
| Double tap | Open a module — the card flies at you and dissolves into its page. Double tap again to close |
| Pinch | Grab the card under the cursor |
| Release while moving | Throw it into the physics world |
| Pull toward you (open palm) | Also opens the module |
| Push away (open palm) | Close and return to the ring |
| Flat palm, held still | Freeze the world |
| Circle (index finger) | Reserved for Phase 02 |
Pointer: move to aim · left drag to pinch · fling to swipe · wheel for push/pull · right/Shift for the pointing pose · double-click to open a module · Space (held still) to freeze.
Keys: ←/→ rotate the ring · H help · C camera preview · V toggle tracking ·
M mute · K skeleton · R recall thrown cards · 1–4 pin quality · Esc collapse.
DOM chrome and the 3D world are separate interaction surfaces: the pointer source only acts on events whose target is the canvas, so clicking a panel never pinches a card behind it and scrolling a panel never drives the depth gesture.
app/ Next routes, global tokens
src/
core/ Framework-free foundation
config/ modules · theme · tuning · quality ← every magic number
math/ springs · filters · noise · damping
time/ the world clock
events/ typed bus
types/
gesture/ Input → intent
sources/ MediaPipeSource · PointerSource (same contract)
detectors/ pinch · swipe · pushPull · palmHold · circle
HandFrame.ts landmark conditioning
GestureEngine.ts orchestration + 60Hz interpolation
rendering/ The picture
environment/ particles · beams · fog · grid
materials/ holo glass · procedural env map · shaders
postfx/ adaptive/ geometry/
scene/ The content
Card/ slab · painted face · animated frame · face painters
Carousel.tsx rotation, hit-testing, gesture routing
HandCursor.tsx
physics/ Rapier, for released cards only
audio/ procedural synth — no samples
hud/ DOM overlay
stores/ zustand (session) · valtio (telemetry)
fallback/ no-WebGL surface
One clock. Everything animated reads worldClock.time rather than elapsed
time. The palm-freeze gesture scales that single number toward zero, and the whole
world — shaders, springs, physics, camera — stops coherently. No component knows
the gesture exists.
Springs, never lerps. core/math/spring.ts is the only way things move.
Frame-rate independent, substepped, allocation-free. Cards lag the ring when it
turns and lag your hand when you drag them because the spring is allowed to fall
behind — that lag is the feel, not a bug to tune out.
Two rates, one engine. Hand detection runs at ~32Hz (inference costs 4–7ms;
running it every frame would eat the render budget). GestureEngine.tick() runs
every frame and springs the cursor between detections, so the interaction reads as
60fps from half the samples.
Pointer parity. PointerSource doesn't emit synthetic gesture events. It poses
a real 21-landmark hand around the cursor and feeds it through the same
HandFrameBuilder the camera uses. Flinging the mouse produces a genuine swipe via
the same detector; the wheel changes the hand's apparent size, which the depth
model reads as push/pull. The fallback exercises the real code path, so it can't
rot separately.
React stays out of the frame loop. Per-frame state lives in plain mutable
objects (worldClock, the gesture runtime, cardRegistry) read inside useFrame.
Zustand holds session state that changes at human speed; valtio holds HUD telemetry
written at ~10Hz. Nothing re-renders at 60Hz.
Physics only where it earns its place. Cards are spring-driven in the ring. A card becomes a Rapier rigid body only when thrown, and leaves the world when recalled. The body is invisible — it produces a transform the card copies — so the card's visual hierarchy never remounts mid-throw.
No lights. Every material is unlit or lit analytically from a procedural environment map. No shadow pass, no light uniforms, no per-light shader permutations. The glass look is fresnel plus a reflection probe.
Append a ModuleDefinition to core/config/modules.ts and a painter to
scene/Card/faces/painters.ts. Nothing else changes — layout, interaction and
physics are all derived from the registry.
Adaptive quality walks a four-tier ladder from the 90th-percentile frame time (mean frame time hides exactly the stutter people notice). Cheapest things go first: post-processing extras, then particle count, then resolution. Stalls over 120ms are excluded from sampling so a backgrounded tab can't downgrade quality.
Measured on Apple Silicon (ANGLE/Metal) at 1440×810: 60–72 fps, ~15ms frame,
~68 draw calls, ~15k triangles, tier balanced.
Card faces are canvas textures repainted at 5Hz when idle and 20Hz when active, phase-staggered by slot so the ring's repaints never land on the same frame.
Needs WebGL2 (WebGL1 falls back to a lower tier). Without WebGL the app renders a flat module registry rather than an error page. Hand tracking additionally needs a camera and a secure context; every failure path — no API, permission denied, model fetch failed, repeated inference errors — falls back to pointer input silently.
Video never leaves the device. Inference runs entirely in the browser.
Additive. Phase 1's scene, gesture engine and physics are untouched; the assistant
mounts through two new lines (<AssistantLayer /> in SceneRoot, useAssistant()
in Hud) and can be removed by deleting them.
# .env.local (gitignored, never sent to the browser)
GEMINI_API_KEY=your-key-hereGet a key at https://aistudio.google.com/apikey, then restart the dev server. Without one the assistant still mounts and says exactly what is missing.
Wake with "Orion", a circle gesture (index finger, traced in the air), or
N. The environment lifts, a wave crosses the floor, and the microphone
appears. Esc dismisses it.
Speak or type. Responses stream token by token, speech starts as soon as the first sentence completes, and talking over Orion stops it mid-word.
It can drive the interface — "open stocks", "rotate left", "close that" — through the same store actions the gestures use, so a spoken command and a double tap converge on one code path. Ask "explain this" with a module open and it resolves the referent from context.
The key never reaches the browser. All calls go through
app/api/orion/chat/route.ts, which streams NDJSON back. A NEXT_PUBLIC_ key
would ship to every visitor.
The assistant must be deaf while it talks. Two separate paths make it interrupt itself, and both had to be closed:
- The recogniser transcribes the assistant's own voice and submits it as the
next command, which cancels the answer still being spoken.
Recognitionis muted for the duration of each utterance (muted, not stopped — restarting a session costs a few hundred ms and loses the start of the user's reply). - The acoustic detector hears the assistant and reads it as a barge-in. Browser
echo cancellation does not help here: AEC cancels what the browser renders
through the WebRTC path, and
speechSynthesisdoes not go through it, so the canceller has no reference signal.BargeIntherefore measures the echo per utterance — it is deaf for the first 650ms and records what it hears, which at that moment is the assistant — and requires the user to be meaningfully louder than that.
Both are covered by npm run check:voice, which fails if either guard is removed.
Speech is chunked by sentence. Tokens are too short to synthesise without audible seams; whole responses would mean waiting for the model. Sentences are the unit that lets speech begin before the response finishes.
Holographic text is two layers. Particles fly in and settle into the words; a crisp text plane fades in behind them as they arrive. Particle text alone is beautiful in a still frame and genuinely hard to read as prose — and this text is the actual answer.
The model writes for the ear. The system prompt forbids markdown, because a TTS engine reads asterisks aloud, and caps length, because the text has to fit in the air in front of you.
- Speech recognition is Chromium-only. Safari and Firefox get the text field, which is why it is always present rather than a fallback.
- The voice cannot be spatialised. Web Speech output cannot be routed through WebAudio, so the environment reacts to the voice but the voice itself is not positioned. True spatial audio needs a cloud TTS returning audio buffers.
- No live data. The model is told to say so rather than invent a stock price.