All Tools
L
OtherFreeOpen Source
LLAVA-ONEVISION
Open multimodal training for image, video, and spatial tasks
Apache-2.0
ABOUT
Most open VLMs stay on single images and hide data or training logs. LLaVA-OneVision ships codec-aligned encoders, long-video and spatial datasets, and full configs so labs can reproduce an 8B model that handles images, long clips, OCR, and 3D layout from one stack.
INTEGRATION GUIDE
1. Train LLaVA-OneVision-2 on a single 8-GPU node with the published quick-start Docker bundle
2. Run native-resolution image, long-video, and spatial reasoning from one 8B instruct checkpoint
3. Evaluate OneVision models with the lmms-eval llava-onevision2 branch and codec backend
4. Fine-tune on LLaVA-OneVision-2-VideoCaption or Spatial datasets without closed data mixes
TAGS
pythonmultimodalvision-languagevideollavatrainingapache