VisualBERT

VisualBERT is a vision-and-language research model that feeds text and detected image regions into one shared Transformer stack. It extends BERT with visu