Vis Aware is an NVDA add-on for OCR, AI-powered image description, automatic recognition, and AI-assisted computer control.
Vis Aware requires NVDA 2026.1 or later. Most hosted engines require an Internet connection and service credentials. Ollama is used for local self-hosting (for example, Google Gemma).
NVDA menu > Preferences > Vis Aware settings.... In General,
you can sign in to your NVDA Chinese Community (www.nvdacn.com) account to
use the related free community services. Then select and configure other
engines as needed in the corresponding OCR, Image description, or
AI Agent panel.NVDA+alt+1 to select a mode (capability type): OCR, Image
description, or AI Agent.NVDA+alt+2 to select an engine supported by that mode. Choose an
engine that you configured in step 1.NVDA+alt+3 to select another recognition source (for example, an image or
image file on the clipboard).Press NVDA+alt+space once to perform recognition using the selected
capability or start or stop the Agent. For regular recognition:
Press once to present the result in an NVDA recognition result document.
For OCR results, you can use Enter or Space to click the coordinates
corresponding to the text; this is useful when recognizing a window.
Vis Aware provides three modes:
OCR and image description can use these recognition sources:
AI Agent mode captures the full screen directly and does not use the selected recognition source.
The following commands have default gestures:
| Command | Default gesture |
|---|---|
| Perform recognition using the current settings | NVDA+alt+space |
| Ask a follow-up question about the previous image description | NVDA+alt+q |
| Cycle through recognition modes | NVDA+alt+1 |
| Cycle through engines for the current mode | NVDA+alt+2 |
| Cycle through recognition sources | NVDA+alt+3 |
The following commands have no default gesture:
After assigning a gesture to Describes images on the clipboard, press it once to present the result in an NVDA recognition result document. Press it twice in quick succession to show the result in a browsable message or have NVDA announce it, depending on the Show text results in a browsable message setting.
When OCR results include coordinates, pressing enter or space on text in
the recognition result document clicks the corresponding location in the
recognized screen area.
Recognition results can be displayed in an NVDA virtual document or, depending on the setting, in a browse mode dialog. Browse mode renders supported Markdown and mathematical formulas. Results can also be copied to the clipboard.
Only the previous recognition result from the current NVDA session is retained.
For image description engines that support follow-up questions, the follow-up
question command opens a multi-turn dialog and streams spoken answers when
supported. In that dialog, control+enter sends a question and escape
cancels the request or closes the dialog. Use View formatted content to view
the initial description or latest answer in browse mode.
Follow-up questions are supported by Google Gemini, Google Gemma, Kimi, VIVO BlueLLM Vision, and Ollama Vision. Outside the conversation dialog, streaming output is available only when browse mode is not in use.
Automatic recognition is off by default. In the Automatic recognition panel, choose OCR or image description, then choose the current engine or a specific enabled engine. The prompt and model fields override that engine only for automatic recognition. Leave the prompt blank to use Vis Aware’s concise default prompt (20 to 30 characters, without Markdown), and choose Use the model selected in engine settings to follow its regular model. When supported, use Fetch models to load available model names.
Automatic recognition runs in the background when the system focus, browse mode cursor, or navigator object moves to a supported graphic control. The result is announced automatically and saved as the previous result; follow-up questions are available when supported by the engine.
For web graphics, Vis Aware normally uses the image’s URL and uses an object screenshot instead when no URL is available. Prefer screenshots for web images recognizes the content shown on the screen first and falls back to the URL if a screenshot cannot be captured.
The AI Agent asks for a task, starts from the foreground window, and can operate across windows. It can click, type, press keys, scroll, drag and drop, navigate, wait, and ask you for information when needed. Actions are not confirmed one by one, so monitor the session and stop it when necessary.
Each AI Agent step sends the selected service a full-screen screenshot, which can include content from other windows. Avoid running the AI Agent while sensitive information is visible on the screen.
The AI Agent cannot start while Screen Curtain is enabled.
Open NVDA menu > Preferences > Vis Aware settings....
The settings dialog contains these panels:
Disabled engines are skipped during normal use and engine cycling. At least one engine must remain enabled in each mode.
For PaddleOCR / PaddleOCR-VL, the OCR settings support an AI Studio hosted task API, an AI Studio deployed service, or a self-hosted PaddleOCR service.
For Ollama engines, enter a full API URL or a host and port such as
localhost:11434. The default API root is http://localhost:11434/api. Use
Fetch models to load model names and then choose a model. If no model is
selected, the first model returned by Ollama is used. The optional API key is
sent as an Authorization: Bearer token. Ollama engines require a
vision-capable model, such as Gemma 4; Ollama OCR provides screen coordinates
only when the model returns valid structured OCR data.
Kimi engines default to the official Kimi Code OpenAI-compatible Base URL
https://api.kimi.com/coding/v1. Enter an API key issued for that endpoint. To use
the public Kimi API instead, set the base URL to https://api.moonshot.ai/v1
and use a public API key. Kimi K3 is the default model; model and thinking
options vary by endpoint and model family.
OCR engines:
Image description engines:
AI Agent engines:
Engine availability depends on its service and configuration.
Apple Vision (OCR Server) is a local-network OCR engine. Vis Aware sends the selected image to OCR Server, an open-source iOS app that recognizes text with Apple Vision and returns text coordinates. No cloud OCR account is required, but the computer and iPhone must be able to reach each other on the same local network.
The current app requires iOS 18.4 or later. The oldest supported iPhones are iPhone XS, iPhone XS Max, iPhone XR, and iPhone SE (2nd generation).
To configure the engine:
NVDA menu > Preferences > Vis Aware settings... > OCR, enable and
select Apple Vision (OCR Server), then enter the displayed address, such
as 192.168.1.10:8000.OCR Server connections use unauthenticated, unencrypted HTTP. Use this engine only on a trusted local network; on public or untrusted networks, transmitted images and returned OCR results can be intercepted or modified.
Keep OCR Server open and the iPhone screen on while using the engine. For continuous operation, follow the upstream project’s instructions for iOS Guided Access. The upstream project also documents app usage and the server API.
For an offline, self-hosted setup, we recommend running a vision-capable Gemma 4 model locally with Ollama. After the model is downloaded and the API URL points to the local service, Vis Aware sends images to that local service. If you use a remote Ollama address hosted by someone else, data is sent to that remote service.
Recognition sends the selected image and, when applicable, a prompt to the service configured for the selected engine. Each AI Agent step sends a full-screen screenshot to the selected service; the screenshot can include other windows. Review the service’s data policy and avoid sending sensitive content.
Manual recognition from a source other than the clipboard requires Screen Curtain to be disabled. The NVDACN password is protected with Windows DPAPI. Other saved API keys are stored unencrypted in the NVDA configuration. Be mindful of data security before creating or sharing a portable copy of NVDA.
This project is licensed under the GNU General Public License version 2. See
COPYING.txt for details.