A vision-language model with 4 billion parameters for depth-aware visual reasoning and question answering.