Skip to content

Latest commit

 

History

History

README.md

JavaScript/TypeScript Voice Assistant Samples

Reference documentation | Package (npm)

This folder contains JavaScript samples demonstrating how to build real-time voice assistants using Azure AI Speech VoiceLive service. Each sample is self-contained for easy understanding and deployment.

Available Samples

A Node.js quickstart demonstrating the Voice Live + Foundry Agent v2 flow, including agent creation and voice assistant runtime samples.

Key Features:

  • Agent creation utility with Voice Live metadata
  • Voice Live session targetting Foundry agents
  • Proactive greeting (LLM-generated or pre-defined)
  • Explicit microphone device selection
  • Barge-in handling and conversation logging

A Node.js quickstart demonstrating MCP (Model Context Protocol) server integration with Voice Live, enabling the assistant to use remote tools (DeepWiki, Azure Docs) during voice conversations.

Key Features:

  • MCP server definitions with configurable approval policies
  • MCP tool discovery, execution, and failure event handling
  • Interactive console-based approval flow for sensitive tools
  • Barge-in handling and conversation logging

A Node.js quickstart demonstrating direct Voice Live model integration without Foundry agent orchestration.

Key Features:

  • Direct model-mode session (gpt-realtime by default)
  • Custom instructions and voice configuration
  • API key or Azure credential authentication
  • Proactive greeting (LLM-generated or pre-defined)
  • Explicit microphone device selection
  • Barge-in handling and conversation logging

Shared PowerShell scripts for setting up and verifying Windows development prerequisites (Node.js, SoX, VS Build Tools).

A browser-based voice assistant demonstrating Azure Voice Live SDK integration in a web application using TypeScript and the Web Audio API.

Key Features:

  • Client/Session architecture with type-safe handler-based events
  • Real-time bi-directional audio streaming (PCM16)
  • Live transcription and streaming text responses
  • Barge-in support for natural conversation interruption
  • Audio level visualization
  • Support for OpenAI and Azure Neural voices

A browser sample showing how to stream stereo audio (microphone + a tap of your app's speaker playback) to Azure Voice Live so the service cancels echo using the exact signal the speaker plays. This feature is called Live Reference AEC. Available at api-version 2026-07-15 and later.

Key Features:

  • Stereo PCM16 streaming with mic on channel 0 and speaker reference on channel 1
  • AudioWorklet interleaver and a Web Audio playback bus (speakers + reference)
  • Session config via inputAudioEchoCancellation (referenceSource: "client", channels: 2)
  • Mic capture with browser echo cancellation, noise suppression, and AGC disabled
  • Barge-in support and streaming transcription

A Dockerized sample demonstrating Azure Voice Live API with avatar integration, enabling visual avatar representation during voice conversations.

Key Features:

  • Avatar-enabled voice conversations
  • Prebuilt, custom, and photo avatar character support
  • WebRTC and WebSocket avatar output modes
  • Live scene settings adjustment for photo avatars
  • Proactive greeting support
  • Barge-in support for natural conversation interruption
  • Docker-based deployment
  • Azure Container Apps deployment guide
  • Developer mode for debugging

Also available in Python with a server-side SDK architecture (FastAPI backend).

A React + Vite demo showcasing a Voice-Enabled Car Assistant powered by Azure OpenAI Realtime API.

Live Demo: https://novaaidesigner.github.io/azure-voice-live-for-car/

Key Features:

  • Vehicle Control (lights, windows, temp)
  • Status Monitoring (speed, battery)
  • Media & Navigation simulation
  • Real-time EV driving cycle simulation
  • Latency and token usage benchmarking

A browser-based English pronunciation coaching demo that pairs the Azure Voice Live SDK with the Azure Speech SDK pronunciation assessment (PA) for real-time, targeted coaching feedback.

Key Features:

  • Per-word pronunciation assessment with color-coded visualization
  • Multiple PA scenarios (Conversation, Concise, Read Along)
  • With-reference-text and streaming PA modes
  • Real-time streaming text/audio with barge-in support
  • Optional latency metrics and live event panel

A minimal Vite + React + TypeScript demo that uses Azure Voice Live for real-time speech translation.

Live Demo: https://novaaidesigner.github.io/azure-voice-live-interpreter/

Key Features:

  • Configurable endpoint, apiKey, model (defaults to gpt-5), and target language.
  • “同声传译专家” system prompt (sentence-by-sentence, context-aware translation).
  • Session window logs: ASR, translations, and event logs.
  • Benchmarks per turn: latency + token usage (also keeps totals).
  • One-click export to the Azure Voice Live Calculator via URL params.

A Web App based on Azure Speech Voice Live for real-time trading simulation.

Live Demo: https://novaaidesigner.github.io/voice-live-trader/

Key Features:

  • Real-time trading assistant
  • Simulated matching engine (client-side)
  • Usage statistics (tokens/audio/network)
  • Multi-turn conversation support

Prerequisites

All samples require:

Sample-specific requirements:

Sample Requirements
Agents New Quickstart Node.js 18+ with npm
MCP Quickstart Node.js 18+ with npm
Model Quickstart Node.js 18+ with npm
Basic Web Voice Assistant Node.js 18+ with npm
Live Reference AEC Node.js 22+ with npm
Voice Live Avatar Docker
Voice Live Car Demo Node.js 18+ with npm
Voice Live Education Demo Node.js 18+ with npm
Voice Live Interpreter Demo Node.js 18+ with npm
Voice Live Trader Demo Node.js 18+ with npm

Getting Started

See individual sample READMEs for detailed setup instructions.

Resources

See Also