LLaVA

LLaVA connects a CLIP vision encoder to the Vicuna language model so you can chat about images, get detailed descriptions, and run complex visual reasonin

LLaVA