Vision Assistant Pro Documentation

Total Downloads: 64,213

Vision Assistant Pro is an advanced, multi-modal AI assistant for NVDA. It leverages world-class AI engines to provide intelligent screen reading, translation, voice dictation, and document analysis.

This add-on was released to the community in honor of the International Day of Persons with Disabilities.

1. Setup & Configuration

Go to NVDA Menu > Preferences > Settings > Vision Assistant Pro. The settings dialog is organized into 8 accessible tabs: Connection, AI Behavior, Translation Languages, Document Reader, Video, CAPTCHA, Prompts, and Advanced.

1.1 Connection Tab

1.2 AI Behavior Tab

1.3 Translation Languages Tab

1.4 Document Reader Tab

1.5 Video Tab

1.6 CAPTCHA Tab

1.7 Prompts Tab

1.8 Advanced Tab & Global Logging

Navigate to the Advanced tab to configure global add-on logging:
- Enable dedicated log file: Toggles logging of all operational events, API traffic, and errors across all add-on modules into a separate file (vision_assistant.log).
- Log Level: Select verbosity between Debug (All Details), Info (General Information), Warning (Warnings Only), and Error (Errors Only).
- Keep Logs For: Set automatic retention periods to automatically clean up older log entries (ranging from 1 hour to 90 days).
- Log Management Controls: Use Open Log File, Open Log Folder, or Clear Log File to inspect or clear log data directly without restarting NVDA or interfering with standard NVDA logs.

2. Command Layer & Shortcuts

To prevent keyboard conflicts, this add-on uses a Command Layer.
1. Press NVDA + Shift + V (Master Key) to activate the layer (you will hear a beep).
2. Release keys, then press one of the following single keys:

Key Function Description
Shift + A AI Operator Autonomous Operation: Tell the AI to perform a task on your screen. Pressing it again instantly aborts active operations.
E UI Explorer Interactive Click: Identifies and clicks UI elements in any app.
T Smart Translator Translates text under navigator cursor or selection.
Shift + T Clipboard Translator Translates content currently in the clipboard.
R Text Refiner Summarize, Fix Grammar, Explain, or run Custom Prompts.
V Object Vision Describes the current navigator object.
O Full Screen Vision Analyzes the entire screen layout and content.
Shift + V Video Analysis Analyze local video files or online YouTube, Instagram, TikTok, or Twitter (X) videos.
Control + V Local Video Recording Records a silent video of your screen and analyzes the actions and layout.
D Document Reader Advanced reader for PDF and images with page range selection.
F Smart File Action Context-aware recognition from selected image, PDF, or TIFF files.
M Media Transcription & Dubbing Transcribe or Dub audio/video files (MP3, WAV, MP4, etc.) into your target language.
C CAPTCHA Solver Captures and solves CAPTCHAs.
Shift + C Direct Chat Opens a direct text-based chat interface with the AI.
S Smart Dictation Converts speech to text. Press to start recording, again to stop/type.
Control+T Voice Translation Transcribes, translates, and types the result based on your language settings.
Control+L Live Assistant Real-time Copilot (Gemini only): Starts or ends a live voice and screen conversation with the AI assistant.
I Status Reporting Announces current progress (e.g., "Scanning...", "Idle").
L Label Object Semantic AI Labeling: Permanently labels the current focused element/icon.
Shift + L Manage/Scan Labels Opens Label Manager (if labels exist) or scans the app for unnamed elements.
U Update Check Manually check GitHub for the latest version of the add-on.
Space Recall Last Result Shows the last AI response in a chat dialog for review or follow-up.
H Commands Help Displays a list of all available shortcuts.
Alt + S Settings Opens the Vision Assistant Pro settings dialog.
Alt + Q Quota Exhausted Keys Report Reports the number of Gemini API keys that have exceeded their daily quota and their reset time.
Alt + M Routing Audit Reports the AI models currently selected in advanced routing.
Up / Down Quick Settings Nav Navigates between quick settings categories (Provider, Model, etc.) in the layer.
Left / Right Change Quick Setting Changes the value of the currently selected quick setting.

3. AI Operator - Autonomous Computer Control

The AI Operator turns Vision Assistant Pro from a passive reader into an active assistant that can interact with your computer on your behalf. You can ask it to describe the screen, answer questions about what it sees, or even take control—clicking buttons, dragging items, typing text, and navigating through applications using natural language commands.

The biggest advantage? It works perfectly in completely inaccessible software. If you are stuck in a custom app, a remote desktop, or a website where your screen reader goes totally silent, the operator doesn't mind. Because it "sees" the screen visually, it can find, read, and interact with elements that have zero accessibility labels.

How It Works

  1. Press NVDA + Shift + V, then press Shift + A (or use the direct shortcut) to open the AI Operator dialog.
  2. Type what you want to do in plain language (e.g., "Click the Save button", "What does the error message say?", or "Rename the file to final.pdf").
  3. The AI will analyze your screen, identify the relevant elements, and carry out the action or provide the answer. If a task requires multiple steps, the operator will continue working until it's complete.
  4. Press Shift + A again at any time to instantly abort an ongoing operation.

Supported Actions

The operator understands a wide range of commands:
- Describe & Answer: "Describe the screen layout" or "What does the error message say?"
- Click: "Click the Save button"
- Right Click: "Right-click the file"
- Double Click: "Double-click the document"
- Drag & Drop: "Drag the document to the Archive folder"
- Type: "Type 'Hello World' in the search box"
- Scroll: "Scroll down three times"
- Keypress: "Press Enter", "Press Tab", "Press Escape"
- Multi-step Tasks: "Open File Explorer, find the report, and rename it to final.pdf"

Important Notes

4. Video Analysis & Audio Description

Note: The Video Analysis and Audio Description features are strictly powered by the Google Gemini provider. Ensure that your active provider in the add-on settings is set to Google Gemini.

Vision Assistant Pro introduces powerful video processing capabilities designed specifically for blind users. It can analyze both online videos and local screen recordings to provide highly detailed visual descriptions and generate professional Audio Description scripts (SRT).

4.1 Local Screen Recording (Control + V)

If you encounter a silent video, an animation, or a tutorial on your screen, you can capture it directly:
1. Press NVDA + Shift + V to enter the Command Layer, then press Control + V.
2. The add-on will silently record your screen in the background.
3. Press Control + V again to stop recording.
4. The AI will then analyze the recorded video segment and provide a highly detailed description of the scene, characters, and actions.

4.2 Video Analysis (Shift + V)

You can analyze both local video files and online videos. Simply select a local video file in Windows Explorer, or copy an online video link to your clipboard. You can also press Shift + V anywhere (like inside a media player) to open a dialog where you can browse for a video file or paste a URL manually.
- Supported Online Platforms: YouTube, Instagram, TikTok, and Twitter (X).
- The AI will automatically detect the local file or the URL, process the video, and provide a comprehensive visual description and audio summary.

4.3 Audio Description Generation (SRT)

For a more structured experience, the add-on can generate professional Audio Description scripts in standard SubRip (SRT) format.
- Smart Gap-Timing: The AI listens to the audio track and specifically anchors its visual descriptions to natural pauses and silent gaps to intelligently minimize dialogue overlap.
- Character Tracking: The engine performs a pre-pass to extract distinct characters based on immutable facial features. It builds a global dictionary to accurately track and label characters across different scenes without confusion.
- Verbatim Text OCR: Any text appearing on the screen (signs, phones, credits) is strictly quoted verbatim.
- How to Use: To listen to the generated subtitle, simply place the .srt file in the same folder as your video file and give it the exact same name. Then, configure your media player (e.g., VLC or PotPlayer) to route the subtitle text directly to your screen reader or TTS engine during playback.

4.4 Synchronized Audio Narration (MP3 Export)

Beyond just creating text-based SRT files, the add-on functions as a complete Audio Description production tool by synthesizing the descriptions into speech and mixing them with the video. You can now choose Gemini Live TTS as the voice engine, which utilizes the Gemini Live API to generate highly realistic, unlimited voice narration. When generating an MP3 for local video files, you have multiple mixing modes:
- Standard AD (Mix Voice): The narration is overlaid directly on top of the video's audio. You will be prompted if you want to apply Audio Ducking (lowering the background volume during descriptions) to ensure the narration is clear.
- Extended AD (Pause Audio): The engine pauses the original video audio during descriptions, ensuring you never miss a single word of the original dialogue or the AI narration.
- YouTube Videos: For YouTube sources (which are not downloaded locally), the MP3 export will strictly contain the synchronized AI voice track without the background video audio.

5. Media Transcription & Dubbing (M)

The Audio Transcriber has been completely rebuilt to support both audio and video files (MP3, WAV, MP4, MKV, etc.). Press M in the Command Layer to select a media file and choose one of 3 distinct operation modes:
1. Transcribe (Original Language): Accurately transcribes the spoken speech in its original language.
2. Transcribe and Translate (Target Language): Transcribes the speech and translates it into your configured target language.
3. Dub and Translate (Target Language) (Gemini Only): A powerful new feature that transcribes the speech, translates it into your target language, and synthesizes a spoken audio dub using the add-on's TTS engine.

6. Advanced Document & Image Reader

Vision Assistant Pro includes a highly optimized Document Reader designed for multi-page PDFs, complex images, and even iPhone HEIC formats.

6.1 Batch Processing & Resume

You don't need to read a massive document all at once. Enter a page range (e.g., 1-20), and the AI will process all pages in the background. If NVDA crashes or you interrupt the scan, the add-on will remember your progress and offer to Resume exactly where it left off!

6.2 Smart File Action

You don't always need to open the document first. In Windows File Explorer, simply highlight a PDF or image and press D (Document Reader) or F (Smart File Action) inside the Command Layer. The add-on will instantly bypass the file dialog and begin processing the highlighted file.

6.3 Document Viewer Shortcuts

When the Document Reader window is open, you can use the following shortcuts:
- Ctrl + PageDown: Move to the next page.
- Ctrl + PageUp: Move to the previous page.
- Alt + A: Open a chat dialog to ask questions about the document.
- Alt + R: Force a Re-scan with AI using your active provider.
- Alt + G: Generate and save a high-quality audio file (WAV/MP3). (Hidden if provider doesn't support TTS).
- Alt + S / Ctrl + S: Save the extracted text as a TXT or HTML file.

7. Semantic AI Labeling & UI Explorer

Stuck in an application with "unlabeled button" everywhere? The Semantic AI Labeling engine solves this permanently.

7.1 Permanent Object Labeling (L)

Focus your screen reader on an unlabeled graphic or button and press L in the Command Layer. The AI will look at the button visually, determine its function, and apply a permanent label.
Unlike older screen reader labeling tools, this add-on uses an advanced hybrid "Object Signature" system (AutomationId/ControlID). Your custom labels will survive window resizing, monitor switching, and application updates!

7.2 Full Application Scan (Shift + L)

Press Shift + L to scan the entire active window at once. The AI will find all unlabeled elements and intelligently name them in one go. You can later manage, rename, or batch-delete these labels from the built-in Label Manager.

7.3 UI Explorer (E)

Need to interact with an element without navigating to it manually? Press E to activate the UI Explorer. The AI will scan the screen and generate an accessible list of every clickable element (ignoring system noise like taskbars). Pick an item from the list, and the add-on will instantly click it for you.

8. Live Voice Assistant

The Live Assistant turns Vision Assistant Pro into a real-time, interactive copilot.
(Note: This feature is exclusive to Google Gemini and Gemini-compatible Custom providers).

9. Custom Prompts & Variables

You can manage prompts in Settings > Prompts > Manage Prompts....

Supported Variables

10. Real-World Use Cases (Which feature should I use?)

Vision Assistant Pro is packed with advanced tools. Here are some common scenarios to help you choose the right one:


Note: An active internet connection is required for all AI features. Multi-page documents are processed automatically.

11. Support & Community

Stay updated with the latest news, features, and releases:
- Telegram Channel: t.me/VisionAssistantPro
- GitHub Issues: For bug reports and feature requests.

Reporting Bugs & Logs

When opening a GitHub issue or asking for support, please include details about your active AI provider, model, and NVDA version. If you are experiencing connection issues or unexpected crashes, enable the dedicated log file in Settings > Advanced, recreate the issue, and attach your vision_assistant.log file to help us resolve the problem faster.

12. Project Supporters

A heartfelt thank you to our community members who support the continuous development and maintenance of this project through their generous financial contributions:

If you wish to support the project financially and see your name here, you can find the Donate option in the NVDA Tools menu (Vision Assistant submenu) or during the setup process after installation.


Changes for 2026.08.06

Changes for 2026.07.15

Changes for 7.0.0

Changes for 6.5.0

Changes for 6.1.2

Changes for 6.1.1

Changes for 6.1.0

Changes for 6.0

Changes for 5.6

Changes for 5.5.2

Changes for 5.5 (The Automation Update)

Changes for 5.0

Changes for 4.6

Changes for 4.5

Changes for 4.0.3

Changes for 4.0.1

Changes for 3.6.0

Changes for 3.5.0

* **Command Layer:** Introduced a Command Layer system (default: NVDA+Shift+V) to group shortcuts under a single master key. For example, instead of pressing NVDA+Control+Shift+T for translation, you now press NVDA+Shift+V followed by T.
* **Online Video Analysis:** Added a new feature to analyze YouTube and Instagram videos directly by providing a URL.

Changes for 3.1.0

Changes for 3.0

Changes for 2.9

Changes for 2.8

Changes for 2.7

Changes for 2.6

Changes for 2.5

Changes for 2.1.1

Changes for 2.1

Changes for 2.0

Changes for 1.5

Changes for 1.0