Skip to content

Latest commit

 

History

History
223 lines (171 loc) · 9.88 KB

File metadata and controls

223 lines (171 loc) · 9.88 KB

Orion

A spatial computing runtime for the browser. Ten holographic modules orbit you in a volumetric environment; you rotate, grab, throw and expand them with your hands.

Phase 01 — the surface, the gesture engine and the physics. No AI yet.

npm install
npm run dev      # http://localhost:3000

predev / prebuild vendor MediaPipe's wasm runtime and the hand-landmark model into public/, so nothing is fetched from a CDN at runtime. The model download needs network access once; if it fails the app boots in pointer mode and you can retry with npm run assets.


Interaction

Hand tracking is the primary input. The pointer is a first-class fallback, not a degraded one — see Pointer parity below.

Gesture Action
Swipe ◀ / ▶ (open hand) Rotate the ring one slot
Tap (one finger) Select what you're pointing at
Double tap Open a module — the card flies at you and dissolves into its page. Double tap again to close
Pinch Grab the card under the cursor
Release while moving Throw it into the physics world
Pull toward you (open palm) Also opens the module
Push away (open palm) Close and return to the ring
Flat palm, held still Freeze the world
Circle (index finger) Reserved for Phase 02

Pointer: move to aim · left drag to pinch · fling to swipe · wheel for push/pull · right/Shift for the pointing pose · double-click to open a module · Space (held still) to freeze.

Keys: ←/→ rotate the ring · H help · C camera preview · V toggle tracking · M mute · K skeleton · R recall thrown cards · 1–4 pin quality · Esc collapse.

DOM chrome and the 3D world are separate interaction surfaces: the pointer source only acts on events whose target is the canvas, so clicking a panel never pinches a card behind it and scrolling a panel never drives the depth gesture.


Architecture

app/                     Next routes, global tokens
src/
  core/                  Framework-free foundation
    config/              modules · theme · tuning · quality   ← every magic number
    math/                springs · filters · noise · damping
    time/                the world clock
    events/              typed bus
    types/
  gesture/               Input → intent
    sources/             MediaPipeSource · PointerSource (same contract)
    detectors/           pinch · swipe · pushPull · palmHold · circle
    HandFrame.ts         landmark conditioning
    GestureEngine.ts     orchestration + 60Hz interpolation
  rendering/             The picture
    environment/         particles · beams · fog · grid
    materials/           holo glass · procedural env map · shaders
    postfx/ adaptive/ geometry/
  scene/                 The content
    Card/                slab · painted face · animated frame · face painters
    Carousel.tsx         rotation, hit-testing, gesture routing
    HandCursor.tsx
  physics/               Rapier, for released cards only
  audio/                 procedural synth — no samples
  hud/                   DOM overlay
  stores/                zustand (session) · valtio (telemetry)
  fallback/              no-WebGL surface

Ideas the code is built on

One clock. Everything animated reads worldClock.time rather than elapsed time. The palm-freeze gesture scales that single number toward zero, and the whole world — shaders, springs, physics, camera — stops coherently. No component knows the gesture exists.

Springs, never lerps. core/math/spring.ts is the only way things move. Frame-rate independent, substepped, allocation-free. Cards lag the ring when it turns and lag your hand when you drag them because the spring is allowed to fall behind — that lag is the feel, not a bug to tune out.

Two rates, one engine. Hand detection runs at ~32Hz (inference costs 4–7ms; running it every frame would eat the render budget). GestureEngine.tick() runs every frame and springs the cursor between detections, so the interaction reads as 60fps from half the samples.

Pointer parity. PointerSource doesn't emit synthetic gesture events. It poses a real 21-landmark hand around the cursor and feeds it through the same HandFrameBuilder the camera uses. Flinging the mouse produces a genuine swipe via the same detector; the wheel changes the hand's apparent size, which the depth model reads as push/pull. The fallback exercises the real code path, so it can't rot separately.

React stays out of the frame loop. Per-frame state lives in plain mutable objects (worldClock, the gesture runtime, cardRegistry) read inside useFrame. Zustand holds session state that changes at human speed; valtio holds HUD telemetry written at ~10Hz. Nothing re-renders at 60Hz.

Physics only where it earns its place. Cards are spring-driven in the ring. A card becomes a Rapier rigid body only when thrown, and leaves the world when recalled. The body is invisible — it produces a transform the card copies — so the card's visual hierarchy never remounts mid-throw.

No lights. Every material is unlit or lit analytically from a procedural environment map. No shadow pass, no light uniforms, no per-light shader permutations. The glass look is fresnel plus a reflection probe.

Adding a module

Append a ModuleDefinition to core/config/modules.ts and a painter to scene/Card/faces/painters.ts. Nothing else changes — layout, interaction and physics are all derived from the registry.


Performance

Adaptive quality walks a four-tier ladder from the 90th-percentile frame time (mean frame time hides exactly the stutter people notice). Cheapest things go first: post-processing extras, then particle count, then resolution. Stalls over 120ms are excluded from sampling so a backgrounded tab can't downgrade quality.

Measured on Apple Silicon (ANGLE/Metal) at 1440×810: 60–72 fps, ~15ms frame, ~68 draw calls, ~15k triangles, tier balanced.

Card faces are canvas textures repainted at 5Hz when idle and 20Hz when active, phase-staggered by slot so the ring's repaints never land on the same frame.

Browser support

Needs WebGL2 (WebGL1 falls back to a lower tier). Without WebGL the app renders a flat module registry rather than an error page. Hand tracking additionally needs a camera and a secure context; every failure path — no API, permission denied, model fetch failed, repeated inference errors — falls back to pointer input silently.

Video never leaves the device. Inference runs entirely in the browser.


Phase 2 — Orion assistant

Additive. Phase 1's scene, gesture engine and physics are untouched; the assistant mounts through two new lines (<AssistantLayer /> in SceneRoot, useAssistant() in Hud) and can be removed by deleting them.

Setup

# .env.local  (gitignored, never sent to the browser)
GEMINI_API_KEY=your-key-here

Get a key at https://aistudio.google.com/apikey, then restart the dev server. Without one the assistant still mounts and says exactly what is missing.

Using it

Wake with "Orion", a circle gesture (index finger, traced in the air), or N. The environment lifts, a wave crosses the floor, and the microphone appears. Esc dismisses it.

Speak or type. Responses stream token by token, speech starts as soon as the first sentence completes, and talking over Orion stops it mid-word.

It can drive the interface — "open stocks", "rotate left", "close that" — through the same store actions the gestures use, so a spoken command and a double tap converge on one code path. Ask "explain this" with a module open and it resolves the referent from context.

Design notes

The key never reaches the browser. All calls go through app/api/orion/chat/route.ts, which streams NDJSON back. A NEXT_PUBLIC_ key would ship to every visitor.

The assistant must be deaf while it talks. Two separate paths make it interrupt itself, and both had to be closed:

  1. The recogniser transcribes the assistant's own voice and submits it as the next command, which cancels the answer still being spoken. Recognition is muted for the duration of each utterance (muted, not stopped — restarting a session costs a few hundred ms and loses the start of the user's reply).
  2. The acoustic detector hears the assistant and reads it as a barge-in. Browser echo cancellation does not help here: AEC cancels what the browser renders through the WebRTC path, and speechSynthesis does not go through it, so the canceller has no reference signal. BargeIn therefore measures the echo per utterance — it is deaf for the first 650ms and records what it hears, which at that moment is the assistant — and requires the user to be meaningfully louder than that.

Both are covered by npm run check:voice, which fails if either guard is removed.

Speech is chunked by sentence. Tokens are too short to synthesise without audible seams; whole responses would mean waiting for the model. Sentences are the unit that lets speech begin before the response finishes.

Holographic text is two layers. Particles fly in and settle into the words; a crisp text plane fades in behind them as they arrive. Particle text alone is beautiful in a still frame and genuinely hard to read as prose — and this text is the actual answer.

The model writes for the ear. The system prompt forbids markdown, because a TTS engine reads asterisks aloud, and caps length, because the text has to fit in the air in front of you.

Known limits

  • Speech recognition is Chromium-only. Safari and Firefox get the text field, which is why it is always present rather than a fallback.
  • The voice cannot be spatialised. Web Speech output cannot be routed through WebAudio, so the environment reacts to the voice but the voice itself is not positioned. True spatial audio needs a cloud TTS returning audio buffers.
  • No live data. The model is told to say so rather than invent a stock price.