Gemini Live for NVDA – Documentation
Version 1.0
Gemini Live brings real-time, natural voice conversations with Google's most advanced AI directly into your NVDA screen reader. Speak naturally, hear human-like responses, and interact with Gemini using your voice or typed text — all without leaving your assistive technology workflow. This add-on turns NVDA into a hands-free AI companion for research, writing, learning, or casual conversation.
Table of Contents
- Overview
- Feature List
- Installation
- Getting Started
- Settings and Configuration
- Voice Library (30+ Natural Voices)
- Live Conversation Interface
- Managing API Keys (Rotation)
- Google Search and Code Execution
- Keyboard Shortcuts
- Troubleshooting
- System Requirements
- Support and Community
- Privacy and Security
Overview
Gemini Live for NVDA is a professional-grade add-on that connects you to Google's Gemini real-time audio models. Unlike traditional text-only chatbots, this add-on supports full-duplex voice conversations: you speak, the AI listens and responds with natural speech, and you can interrupt or type at any time. The add-on supports multiple API keys with automatic rotation, 30 built-in ultra-realistic voices (male, female, various styles), and advanced features like Google Search grounding and secure code execution.
All settings — including your system instructions, voice preference, temperature, API keys, and tool toggles — are saved locally and persist across NVDA sessions. Whether you need a brainstorming partner, a coding assistant, a research aide, or just someone to talk to, Gemini Live fits seamlessly into your workflow.
Feature List
- Real-time voice conversation – speak to Gemini and hear spoken responses with ultra-low latency using Google's Live API.
- 30 premium voices – choose from 30 expressive voices (Zephyr, Puck, Charon, Kore, Fenrir, Leda, and more) with unique styles: Bright, Upbeat, Informative, Firm, Gentle, Smooth, and others.
- Text input fallback – type your messages when speaking isn't convenient; ideal for quiet environments or complex prompts.
- API key rotation – add multiple API keys (one per line). If one fails or hits limits, the add-on automatically tries the next key.
- System instructions – define Gemini's personality, role, and constraints (e.g., "You are a patient coding tutor" or "You speak only in haikus").
- Adjustable temperature (creativity) – slider from 0.0 (deterministic, factual) to 2.0 (creative, surprising). Default 1.0.
- Google Search grounding – enable real-time web search to get current information, news, and facts.
- Secure code execution – run Python code in a sandboxed environment; perfect for data analysis, math, or testing snippets.
- Microphone mute/unmute – pause voice input without ending the session.
- Session persistence – keep conversations active until you explicitly end them; no timeouts.
- Audio feedback – system sounds confirm connection and disconnection.
- Two supported models – gemini-3.1-flash-live-preview (fast) and gemini-2.5-flash-native-audio-preview-12-2025 (high quality).
- Standalone start shortcut – launch a live conversation directly with NVDA+S (bypassing the main menu).
- Full NVDA integration – accessible dialogs, progress indicators, and tray menu entry under Tools.
Installation
- Download the
.nvda-addon file for Gemini Live.
- Ensure NVDA is running (press Control+Alt+N if needed).
- Open the NVDA menu with NVDA+N, go to Tools, then Manage Add-ons.
- Press the Install button, browse to the downloaded add-on file, and select it.
- Confirm the installation and restart NVDA when prompted.
- After restart, access Gemini Live via NVDA+Alt+G or from the NVDA Tools menu.
Note: This add-on requires valid Google Gemini API keys with access to live audio models. API keys are not included and must be obtained from
Google AI Studio or Google Cloud Console.
Getting Started
Press NVDA+Alt+G to open the main Gemini Live dialog. From here you can:
- Open the Settings dialog to configure your API keys, voice, system instructions, and tools.
- Start Conversation – begin a live voice session using your saved settings.
- About – view add-on version, author information, and a link to the Telegram support channel.
Alternatively, use the direct shortcut NVDA+S anywhere to immediately start a live conversation using your most recent settings (if configured).
When you start a conversation, a connection dialog appears while the add-on attempts each API key in sequence. Once connected, you will hear a confirmation sound (SystemAsterisk), and the main conversation window opens.
Settings and Configuration
Press the Settings button in the main dialog or access it via the Tools menu. All settings are saved automatically to %APPDATA%\gemini-live-application\settings.ini.
General Settings Tab
- Select AI Model – choose between
gemini-3.1-flash-live-preview (faster, lower latency) and gemini-2.5-flash-native-audio-preview-12-2025 (richer audio quality).
- Select Voice – choose from 30 voices (see Voice Library section below). Each voice includes style, gender, and description.
- System Instructions – enter a multi-line prompt that defines Gemini's behavior. Example: "You are a patient and encouraging math tutor for blind students. Explain concepts step by step."
- Temperature (creativity) – slider from 0.0 to 2.0. Lower values produce more focused, factual answers. Higher values produce more creative, unexpected responses.
- Enable Google Search – allows Gemini to fetch real-time information from the web. Responses may include citations.
- Enable Code Execution – allows Gemini to write and run Python code in a sandbox. Useful for calculations, data manipulation, or testing logic.
- API Key(s) – enter one or more Google API keys, one per line. The add-on will try keys in order; if one fails, it automatically moves to the next.
Voice Library (30+ Natural Voices)
Gemini Live offers a diverse range of high-fidelity neural voices. Each voice has a unique ID, style, gender, and short description. Below is the complete selection:
Zephyr – Bright, Female
Puck – Upbeat, Male
Charon – Informative, Male
Kore – Firm, Female
Fenrir – Excitable, Male
Leda – Youthful, Female
Orus – Firm, Male
Aoede – Breezy, Female
Callirrhoe – Easy-going, Female
Autonoe – Bright, Female
Enceladus – Breathy, Male
Iapetus – Clear, Male
Umbriel – Easy-going, Male
Algieba – Smooth, Male
Despina – Smooth, Female
Erinome – Clear, Female
Algenib – Gravelly, Male
Rasalgethi – Informative, Male
Laomedeia – Upbeat, Female
Achernar – Soft, Female
Alnilam – Firm, Male
Schedar – Even, Male
Gacrux – Mature, Female
Pulcherrima – Forward, Female
Achird – Friendly, Male
Zubenelgenubi – Casual, Male
Vindemiatrix – Gentle, Female
Sadachbia – Lively, Male
Sadaltager – Knowledgeable, Male
Sulafat – Warm, Female
All voices are server-side and require no local downloads. The add-on sends the voice ID with each session request.
Live Conversation Interface
When a session starts, the main conversation window opens with the following controls:
- Your message (multi-line text box) – type a message if you prefer text input. Press the Send button or simply speak into your microphone.
- Send button – sends typed text to Gemini. Voice input is sent automatically.
- Mute Mic button – temporarily disables your microphone. The AI can still respond to typed messages. Press again to unmute.
- End Session button – gracefully terminates the live conversation, closes audio streams, and returns to the main dialog.
During the conversation:
- Speak naturally. There is no "push-to-talk" – the microphone streams continuously while unmuted.
- Gemini's responses play through your default speakers or headphones.
- You can interrupt the AI by speaking again or sending a typed message.
- All interactions are real-time; there is no manual page turning or waiting for processing.
When you end the session, a SystemExit sound plays, and the window closes automatically. No conversation history is saved locally for privacy reasons.
Managing API Keys (Rotation)
The add-on supports multiple API keys for redundancy and rate limit handling. To configure:
- Open Settings from the main dialog.
- In the API Key(s) text box, enter one key per line, for example:
AIzaSyA1b2c3d4e5f6g7h8i9j0k1l2m3n4o5p6q7r8s9t
AIzaSyB2c3d4e5f6g7h8i9j0k1l2m3n4o5p6q7r8s9u
- Save the settings.
When starting a conversation, the add-on tries keys in order. If a key fails (invalid, expired, quota exceeded, or model not accessible), it automatically moves to the next key after a 0.5-second delay. A progress dialog shows which key is being attempted. If all keys fail, an error message appears.
Tip: Obtain API keys from
Google AI Studio. Ensure the keys have access to the "Live API" and the specific model you selected. Free tier quotas apply.
Gemini Live can optionally use two powerful tools:
Google Search Grounding
When enabled, Gemini can search the web in real time to answer questions about current events, recent data, or factual information. This is useful for news, weather, sports scores, stock prices, or any query that requires up-to-date knowledge. The AI will typically cite sources in its response.
Code Execution
When enabled, Gemini can write and execute Python code in a secure, sandboxed environment. This is ideal for:
- Mathematical calculations and formula solving
- Data analysis or CSV manipulation
- Testing algorithms or debugging logic
- Generating charts or visual data (text described)
Code runs server-side on Google's infrastructure and does not access your local files. Results are returned as text or audio descriptions. Enable this feature only if you need computational assistance; it may increase response latency.
Keyboard Shortcuts
| Shortcut | Action |
NVDA+Alt+G | Open Gemini Live main dialog (settings, start, about) |
NVDA+S | Start live conversation directly (uses saved settings) |
Alt+F4 (in conversation window) | End session and close window |
Enter (in text box) | Send typed message (multi-line: Shift+Enter for new line) |
Escape (in any dialog) | Cancel or close without applying |
Ctrl+Tab / Ctrl+Shift+Tab | Navigate between dialog tabs (Settings dialog) |
All buttons are fully accessible via NVDA object navigation and standard Tab key focus.
Troubleshooting
No sound or microphone not working: Ensure your default input and output devices are correctly set in Windows Sound settings. The add-on uses PyAudio and accesses the system's default microphone and speaker.
Connection fails with all API keys: Verify that each API key is valid, has billing enabled (or free tier quota), and has access to the selected live model. Some keys may require enabling the Live API in Google Cloud Console.
High latency or choppy audio: This is often network-related. Use a stable broadband connection. Select the "flash-live-preview" model for lower latency.
Voice doesn't match selection: Some voice IDs may be deprecated or region-restricted. Try another voice from the library.
NVDA becomes unresponsive during session start: Wait a few seconds; the add-on is attempting API keys. If it persists, restart NVDA and ensure your API keys are correct.
Google Search or Code Execution not working: Ensure you have enabled these features in Settings before starting the conversation. They cannot be toggled mid-session.
Note: All audio processing (microphone capture and speaker playback) uses 16kHz for input and 24kHz for output. If you experience echo or feedback, use headphones instead of speakers.
System Requirements
- NVDA 2023.1 or later
- Windows 10 or Windows 11 (64-bit recommended)
- Working microphone and speakers/headphones
- Stable internet connection (broadband, at least 1 Mbps up/down)
- Valid Google Gemini API key with Live API access
- Python environment not required – the add-on bundles all dependencies (pyaudio, google-genai, etc.)
Support and Community
For assistance, feature requests, or to report issues:
- Join the Telegram channel: https://t.me/blindtechvisionary
- Contact the author directly via Telegram: @blindtechvisionary
- Check the About dialog inside the add-on (main dialog -> About button) for version information and quick links.
Your feedback directly shapes future updates. Suggestions for new voices, additional tools, or improved workflows are always welcome.
Privacy and Security
Gemini Live for NVDA is designed with privacy as a core principle:
- No local storage of conversations – Audio and text data are streamed directly to Google's API and are not saved on your computer after the session ends.
- API keys stored locally – Keys are saved in plain text in
%APPDATA%\gemini-live-application\settings.ini. Protect your computer with a strong password if you are concerned about key exposure.
- No usage telemetry – The add-on does not collect any analytics, crash reports, or usage statistics.
- Google's data handling – Audio and text sent to Google's API are subject to Google's privacy policy. Review Google Privacy Policy for details.
- Offline operation – The add-on itself contains no network services beyond connecting to Google's API. No third-party servers are involved.
For maximum privacy, avoid sharing sensitive personal or financial information during conversations, as with any cloud-based AI service.
Important: This add-on is not affiliated with Google. "Gemini" and "Google" are trademarks of Google LLC. The add-on uses official Google Generative AI client libraries.
Gemini Live for NVDA – Version 1.0 – Built for natural, voice-driven AI assistance.
Author: Sujan Rai | Telegram: @blindtechvisionary