Jun 13Vibe with Hermes Agent — Bengaluru · RSVP
ToolsMCPBlogResearchCommunityStar on GitHub
All Tools
L
OtherFreeOpen Source

LLAVA-ONEVISION

Open multimodal training for image, video, and spatial tasks

Apache-2.0

ABOUT

Most open VLMs stay on single images and hide data or training logs. LLaVA-OneVision ships codec-aligned encoders, long-video and spatial datasets, and full configs so labs can reproduce an 8B model that handles images, long clips, OCR, and 3D layout from one stack.

INTEGRATION GUIDE

1. Train LLaVA-OneVision-2 on a single 8-GPU node with the published quick-start Docker bundle 2. Run native-resolution image, long-video, and spatial reasoning from one 8B instruct checkpoint 3. Evaluate OneVision models with the lmms-eval llava-onevision2 branch and codec backend 4. Fine-tune on LLaVA-OneVision-2-VideoCaption or Spatial datasets without closed data mixes

TAGS

pythonmultimodalvision-languagevideollavatrainingapache
LLaVA-OneVision — AI Tool | Agentic AI For Good