Language extraction
Extract visible text and language content from images.
Vision language workflow
VXLT turns photographed documents and images into readable, translated language using a local vision model.
Windows 10/11 x64 · CPU supported · GPU recommended
What it does
A narrow, understandable workflow for images that contain text or need a plain-language description.
Extract visible text and language content from images.
Translate foreign-language content from photographed documents or images.
Select one picture and send it for analysis.
Supports common picture formats.
CPU and NVIDIA GPU supported. CPU works, but GPU is recommended.
Unzip the downloaded file and run VXLT.exe.
Quick start
VXLT uses Ollama to download and run the vision model on your computer.
Download it from ollama.com. If CLEAR already uses Ollama, skip this step.
Unzip the download and run VXLT.exe.
Select “Check Ollama Connection.” A check mark confirms Ollama is running.
Select a model size and download it the first time you run the app.
Select a picture and a task, then choose Send.
Click the screenshot to enlarge it and read the details.
Current release
The first public release of the vision extraction and translation workflow.
v0.1Initial release.
VXLT 0.1 MD5ab05d3273b0980e3cb8b2502ec6a417e
Good to know
Vision eXtract Language Translate.
VXLT uses a vision language model to extract content from images and translate it when needed.
Windows 10 and Windows 11 on x64 systems.
No. A CPU works, though a compatible GPU will make processing faster.
Ready to try it?