Vision Assistant Pro Documentation

Total Downloads: 70,107

Vision Assistant Pro is an advanced, multi-modal AI assistant for NVDA. It leverages world-class AI engines to provide intelligent screen reading, translation, voice dictation, and document analysis.

This add-on was released to the community in honor of the International Day of Persons with Disabilities.

1. Setup & Configuration

Go to NVDA Menu > Preferences > Settings > Vision Assistant Pro. The settings dialog is organized into 9 accessible tabs: Connection, Live Assistant, AI Behavior, Translation Languages, Document Reader, Video, CAPTCHA, Prompts, and Advanced.

1.1 Connection Tab

1.2 Live Assistant Tab

Note: This tab appears only when Google Gemini (or a Gemini-compatible Custom provider) is your active provider.

1.3 AI Behavior Tab

1.4 Translation Languages Tab

1.5 Document Reader Tab

1.6 Video Tab

1.7 CAPTCHA Tab

1.8 Prompts Tab

1.9 Advanced Tab & Global Logging

Navigate to the Advanced tab to configure global add-on logging:
- Enable dedicated log file: Toggles logging of all operational events, API traffic, and errors across all add-on modules into a separate file (vision_assistant.log).
- Log Level: Select verbosity between Debug (All Details), Info (General Information), Warning (Warnings Only), and Error (Errors Only).
- Keep Logs For: Set automatic retention periods to automatically clean up older log entries (ranging from 1 hour to 90 days).
- Log Management Controls: Use Open Log File, Open Log Folder, or Clear Log File to inspect or clear log data directly without restarting NVDA or interfering with standard NVDA logs.

1.10 Settings Backup & Restore

The Advanced tab also includes a Backup and Restore section:
- Backup: Saves your configuration into a single JSON file. When you click it, you choose what to include: Everything (settings, custom labels, OCR progress, and history) or Settings Only.
- Restore: Loads a previously saved backup to restore your configuration and data at any time, on any machine, or after reinstalling NVDA. You will be asked to confirm first, since restoring replaces all of your current settings and data.

2. Command Layer & Shortcuts

To prevent keyboard conflicts, this add-on uses a Command Layer.
1. Press NVDA + Shift + V (Master Key) to activate the layer (you will hear a beep).
2. Release keys, then press one of the following single keys:

Key Function Description
Shift + A AI Operator Autonomous Operation: Tell the AI to perform a task on your screen. Pressing it again instantly aborts active operations.
E UI Explorer Interactive Click: Identifies and clicks UI elements in any app.
T Smart Translator Translates text under navigator cursor or selection.
Shift + T Clipboard Translator Translates content currently in the clipboard.
R Text Refiner Summarize, Fix Grammar, Explain, or run Custom Prompts.
V Object Vision Describes the current navigator object.
O Full Screen Vision Analyzes the entire screen layout and content.
Shift + V Video Analysis Analyze local video files or online YouTube, Instagram, TikTok, or Twitter (X) videos.
Control + V Local Video Recording Records a silent video of your screen and analyzes the actions and layout.
D Document Reader Advanced reader for PDF, images, and plain text/HTML files with page range selection.
F Smart File Action Context-aware recognition from selected image, PDF, or TIFF files.
M Media Transcription & Dubbing Transcribe or Dub audio/video files (MP3, WAV, MP4, etc.) into your target language.
C CAPTCHA Solver Captures and solves CAPTCHAs.
Shift + C Direct Chat Opens a direct text-based chat interface with the AI.
S Smart Dictation Converts speech to text. Press to start recording, again to stop/type.
Control+T Voice Translation Transcribes, translates, and types the result based on your language settings.
Control+L Live Assistant Real-time Copilot (Gemini only): Starts or ends a live voice and screen conversation with the AI assistant.
I Status Reporting Announces current progress (e.g., "Scanning...", "Idle").
L Label Object Semantic AI Labeling: Permanently labels the current focused element/icon.
Shift + L Manage/Scan Labels Opens Label Manager (if labels exist) or scans the app for unnamed elements.
U Update Check Manually check GitHub for the latest version of the add-on.
Space Recall Last Result Shows the last AI response in a chat dialog for review or follow-up.
H Commands Help Displays a list of all available shortcuts.
Control + H History Opens the History dialog listing your past chats and documents, with type filters and Delete/Clear options.
Alt + S Settings Opens the Vision Assistant Pro settings dialog.
Alt + Q Quota Exhausted Keys Report Reports the number of Gemini API keys that have exceeded their daily quota and their reset time.
Alt + M Routing Audit Reports the AI models currently selected in advanced routing.
Up / Down Quick Settings Nav Navigates between quick settings categories (Provider, Model, etc.) in the layer.
Left / Right Change Quick Setting Changes the value of the currently selected quick setting.

3. Chat & History

Chat windows and the History dialog work across all features, so you can review conversations and pick up right where you left off.

3.1 Chat Window Shortcuts

When a chat window is open (Direct Chat, document chat, refine, and similar), you can review the conversation with:
- Alt + Down: Read the next message.
- Alt + Up: Read the previous message.
- Alt + C: Copy the current message.

3.2 History (Control + H)

Press Control + H in the Command Layer to open the History dialog with your past chats and documents, filterable by type (All / Chats / Documents). Open a chat to continue the conversation — including its attached files, which are re-attached automatically — or open a document and keep reading. Press Delete on any item to remove it, or Clear All to empty the list.

4. AI Operator - Autonomous Computer Control

The AI Operator turns Vision Assistant Pro from a passive reader into an active assistant that can interact with your computer on your behalf. You can ask it to describe the screen, answer questions about what it sees, or even take control—clicking buttons, dragging items, typing text, and navigating through applications using natural language commands.

The biggest advantage? It works perfectly in completely inaccessible software. If you are stuck in a custom app, a remote desktop, or a website where your screen reader goes totally silent, the operator doesn't mind. Because it "sees" the screen visually, it can find, read, and interact with elements that have zero accessibility labels.

How It Works

  1. Press NVDA + Shift + V, then press Shift + A (or use the direct shortcut) to open the AI Operator dialog.
  2. Type what you want to do in plain language (e.g., "Click the Save button", "What does the error message say?", or "Rename the file to final.pdf").
  3. The AI will analyze your screen, identify the relevant elements, and carry out the action or provide the answer. If a task requires multiple steps, the operator will continue working until it's complete.
  4. Press Shift + A again at any time to instantly abort an ongoing operation.

Supported Actions

The operator understands a wide range of commands:
- Describe & Answer: "Describe the screen layout" or "What does the error message say?"
- Click: "Click the Save button"
- Right Click: "Right-click the file"
- Double Click: "Double-click the document"
- Drag & Drop: "Drag the document to the Archive folder"
- Type: "Type 'Hello World' in the search box"
- Scroll: "Scroll down three times"
- Keypress: "Press Enter", "Press Tab", "Press Escape"
- Multi-step Tasks: "Open File Explorer, find the report, and rename it to final.pdf"

Important Notes

5. Video Analysis & Audio Description

Note: The Video Analysis and Audio Description features are strictly powered by the Google Gemini provider. Ensure that your active provider in the add-on settings is set to Google Gemini.

Vision Assistant Pro introduces powerful video processing capabilities designed specifically for blind users. It can analyze both online videos and local screen recordings to provide highly detailed visual descriptions and generate professional Audio Description scripts (SRT).

5.1 Local Screen Recording (Control + V)

If you encounter a silent video, an animation, or a tutorial on your screen, you can capture it directly:
1. Press NVDA + Shift + V to enter the Command Layer, then press Control + V.
2. The add-on will silently record your screen in the background.
3. Press Control + V again to stop recording.
4. The AI will then analyze the recorded video segment and provide a highly detailed description of the scene, characters, and actions.

5.2 Video Analysis (Shift + V)

You can analyze both local video files and online videos. Simply select a local video file in Windows Explorer, or copy an online video link to your clipboard. You can also press Shift + V anywhere (like inside a media player) to open a dialog where you can browse for a video file or paste a URL manually.
- Supported Online Platforms: YouTube, Instagram, TikTok, and Twitter (X).
- The AI will automatically detect the local file or the URL, process the video, and provide a comprehensive visual description and audio summary.

5.3 Audio Description Generation (SRT)

For a more structured experience, the add-on can generate professional Audio Description scripts in standard SubRip (SRT) format.
- Smart Gap-Timing: The AI listens to the audio track and specifically anchors its visual descriptions to natural pauses and silent gaps to intelligently minimize dialogue overlap.
- Character Tracking: The engine performs a pre-pass to extract distinct characters based on immutable facial features. It builds a global dictionary to accurately track and label characters across different scenes without confusion.
- Verbatim Text OCR: Any text appearing on the screen (signs, phones, credits) is strictly quoted verbatim.
- How to Use: To listen to the generated subtitle, simply place the .srt file in the same folder as your video file and give it the exact same name. Then, configure your media player (e.g., VLC or PotPlayer) to route the subtitle text directly to your screen reader or TTS engine during playback.

5.4 Synchronized Audio Narration (MP3 Export)

Beyond just creating text-based SRT files, the add-on functions as a complete Audio Description production tool by synthesizing the descriptions into speech and mixing them with the video. You can now choose Gemini Live TTS as the voice engine, which utilizes the Gemini Live API to generate highly realistic, unlimited voice narration. When generating an MP3 for local video files, you have multiple mixing modes:
- Standard AD (Mix Voice): The narration is overlaid directly on top of the video's audio. You will be prompted if you want to apply Audio Ducking (lowering the background volume during descriptions) to ensure the narration is clear.
- Extended AD (Pause Audio): The engine pauses the original video audio during descriptions, ensuring you never miss a single word of the original dialogue or the AI narration.
- YouTube Videos: For YouTube sources (which are not downloaded locally), the MP3 export will strictly contain the synchronized AI voice track without the background video audio.

6. Media Transcription & Dubbing (M)

The Audio Transcriber has been completely rebuilt to support both audio and video files (MP3, WAV, MP4, MKV, etc.). Press M in the Command Layer to select a media file and choose one of 3 distinct operation modes:
1. Transcribe (Original Language): Accurately transcribes the spoken speech in its original language.
2. Transcribe and Translate (Target Language): Transcribes the speech and translates it into your configured target language.
3. Dub and Translate (Target Language) (Gemini Only): A powerful new feature that transcribes the speech, translates it into your target language, and synthesizes a spoken audio dub using the add-on's TTS engine.

7. Advanced Document & Image Reader

The Document Reader turns your documents into clean, readable text — so you can read, translate, and listen to anything from a scanned book to a stack of photos. It handles multi-page PDFs, complex images, iPhone HEIC formats, and even plain text (.txt) and HTML (.html, .htm) files, which are opened instantly with no OCR or AI processing. Select several files at once and they are merged into a single continuous document in page order. Three OCR engines are available — Chrome (Fast), AI (Advanced) for superior layout preservation, and None (Extract Text Layer) for searchable PDFs — selected in Settings → Document Reader.

How It Works

  1. Press NVDA + Shift + V, then D to open the Document Reader — or highlight a file in File Explorer first and press D / F to skip the file dialog entirely.
  2. Pick one or more PDFs or images. The add-on scans them and announces the total page count.
  3. In the Options dialog, choose the page range (From/To). You can also check Translate Output and pick the target language, or toggle Describe images inline during OCR.
  4. Text extraction starts in the background in batches. You can close the window at any time and continue later — nothing is lost.
  5. Once pages are ready, read them in the viewer: move between pages, jump to any page, ask the AI questions, save the text, or generate an audio narration.

7.1 Batch Processing & Resume

You don't need to read a massive document all at once. Choose a page range (e.g., 1-20) or keep the defaults to process everything, and the AI extracts all pages in the background. If NVDA crashes or you interrupt the scan, the add-on remembers your progress and offers to Resume exactly where it left off — even across restarts. Completed documents are also cached, so reopening them (from Recent Documents or via D) loads the text instantly without re-running OCR, unless the source files have changed.

7.2 Smart File Action

You don't always need to open the document first. In Windows File Explorer, simply highlight a PDF, image, or text/HTML file and press D (Document Reader) — or highlight a PDF or image and press F (Smart File Action) — inside the Command Layer. The add-on instantly bypasses the file dialog and begins processing the highlighted file. Selecting several files at once processes them together as one document.

7.3 Document Viewer Controls & Shortcuts

When the Document Reader window is open, you can use the following:

Keyboard Shortcuts

Buttons & Controls

7.4 Recent Documents (D)

Pressing D in the Command Layer lists your recently read documents first. Choose one to continue from the page you were on — even if the OCR already finished — or press Open File... (Ctrl + O) to browse for a file as usual.

8. Semantic AI Labeling & UI Explorer

Stuck in an application with "unlabeled button" everywhere? The Semantic AI Labeling engine solves this permanently.

8.1 Permanent Object Labeling (L)

Focus your screen reader on an unlabeled graphic or button and press L in the Command Layer. The AI will look at the button visually, determine its function, and apply a permanent label.
Unlike older screen reader labeling tools, this add-on uses an advanced hybrid "Object Signature" system (AutomationId/ControlID). Your custom labels will survive window resizing, monitor switching, and application updates!

8.2 Full Application Scan (Shift + L)

Press Shift + L to scan the entire active window at once. The AI will find all unlabeled elements and intelligently name them in one go. You can later manage, rename, or batch-delete these labels from the built-in Label Manager.

8.3 UI Explorer (E)

Need to interact with an element without navigating to it manually? Press E to activate the UI Explorer. The AI will scan the screen and generate an accessible list of every clickable element (ignoring system noise like taskbars). Pick an item from the list, and the add-on will instantly click it for you.

9. Live Voice Assistant

The Live Assistant turns Vision Assistant Pro into a real-time, interactive copilot.
(Note: This feature is exclusive to Google Gemini and Gemini-compatible Custom providers).

10. Custom Prompts & Variables

You can manage prompts in Settings > Prompts > Manage Prompts....

Custom Prompt Shortcuts

Give any custom prompt its own shortcut key directly in the Prompt Manager, and run it instantly with your current selection or context:
- Single key (e.g., 1, p, or F3): Works inside the Command Layer, and also globally as NVDA + Shift + key.
- Key combination (e.g., Control + Shift + 1, Alt + P, or Insert + 1): Works globally on its own.

Supported Variables

11. Real-World Use Cases (Which feature should I use?)

Vision Assistant Pro is packed with advanced tools. Here are some common scenarios to help you choose the right one:


Note: An active internet connection is required for all AI features. Multi-page documents are processed automatically.

12. Support & Community

Stay updated with the latest news, features, and releases:
- Telegram Channel: t.me/VisionAssistantPro
- GitHub Issues: For bug reports and feature requests.

Reporting Bugs & Logs

When opening a GitHub issue or asking for support, please include details about your active AI provider, model, and NVDA version. If you are experiencing connection issues or unexpected crashes, enable the dedicated log file in Settings > Advanced, recreate the issue, and attach your vision_assistant.log file to help us resolve the problem faster.

13. Project Supporters

A heartfelt thank you to our community members who support the continuous development and maintenance of this project through their generous financial contributions:

If you wish to support the project financially and see your name here, you can find the Donate option in the NVDA Tools menu (Vision Assistant submenu) or during the setup process after installation.


Changes for 2026.09.01

Changes for 2026.08.06

Changes for 2026.07.15

Changes for 7.0.0

Changes for 6.5.0

Changes for 6.1.2

Changes for 6.1.1

Changes for 6.1.0

Changes for 6.0

Changes for 5.6

Changes for 5.5.2

Changes for 5.5 (The Automation Update)

Changes for 5.0

Changes for 4.6

Changes for 4.5

Changes for 4.0.3

Changes for 4.0.1

Changes for 3.6.0

Changes for 3.5.0

* **Command Layer:** Introduced a Command Layer system (default: NVDA+Shift+V) to group shortcuts under a single master key. For example, instead of pressing NVDA+Control+Shift+T for translation, you now press NVDA+Shift+V followed by T.
* **Online Video Analysis:** Added a new feature to analyze YouTube and Instagram videos directly by providing a URL.

Changes for 3.1.0

Changes for 3.0

Changes for 2.9

Changes for 2.8

Changes for 2.7

Changes for 2.6

Changes for 2.5

Changes for 2.1.1

Changes for 2.1

Changes for 2.0

Changes for 1.5

Changes for 1.0