Vision language workflow

VXLT

Vision eXtract Language Translate. Translate and describe content from images, especially documents that are photographed instead of scanned.

VXLT features

A vision-language workflow for images and photographed documents.

  • Language extractionExtract visible text and language content from images.
  • TranslationTranslate foreign-language content from photographed documents or images.
  • Single file processingSelect one picture and send it for analysis.
  • Picture supportSupports common picture formats.
  • Offline capabilityCPU and NVIDIA GPU supported. CPU works, but GPU is recommended.
  • No install neededUnzip the downloaded file and run VXLT.exe.

Quick Start

VXLT uses Ollama to manage models, process files, and provide responses.

Setup and first run
  1. Install Ollama from ollama.com. If CLEAR already uses Ollama, skip this.
  2. Unzip the downloaded file and run VXLT.exe.
  3. Click “Check Ollama Connection.” A check mark confirms Ollama is running.
  4. Select a model size and click Download if this is your first run.
  5. Select a picture and task, then click Send.

Changelog

Recent VXLT releases.

  • v0.1Initial release.

Hash

MD5 hash for verifying the VXLT download.

VXLT 0.1:
ab05d3273b0980e3cb8b2502ec6a417e

FAQ

Common platform and setup notes.

What does VXLT stand for?

Vision eXtract Language Translate.

What does it do?

VXLT uses a vision language model to extract content from images and translate it when needed.

What platform is supported?

Windows 10/11 x64.

Do I have to have a GPU?

CPU will do, but GPU is recommended.