A monocular depth estimation model using a large Vision Transformer backbone with DPT head, developed by Google.