Vision language workflow
VXLT
Vision eXtract Language Translate. Translate and describe content from images, especially documents that are photographed instead of scanned.
VXLT features
A vision-language workflow for images and photographed documents.
- Language extractionExtract visible text and language content from images.
- TranslationTranslate foreign-language content from photographed documents or images.
- Single file processingSelect one picture and send it for analysis.
- Picture supportSupports common picture formats.
- Offline capabilityCPU and NVIDIA GPU supported. CPU works, but GPU is recommended.
- No install neededUnzip the downloaded file and run VXLT.exe.
Quick Start
VXLT uses Ollama to manage models, process files, and provide responses.
Setup and first run
- Install Ollama from ollama.com. If CLEAR already uses Ollama, skip this.
- Unzip the downloaded file and run
VXLT.exe. - Click “Check Ollama Connection.” A check mark confirms Ollama is running.
- Select a model size and click Download if this is your first run.
- Select a picture and task, then click Send.
Changelog
Recent VXLT releases.
- v0.1Initial release.
Hash
MD5 hash for verifying the VXLT download.
VXLT 0.1:
ab05d3273b0980e3cb8b2502ec6a417e
ab05d3273b0980e3cb8b2502ec6a417e
FAQ
Common platform and setup notes.
What does VXLT stand for?
Vision eXtract Language Translate.
What does it do?
VXLT uses a vision language model to extract content from images and translate it when needed.
What platform is supported?
Windows 10/11 x64.
Do I have to have a GPU?
CPU will do, but GPU is recommended.