A base Vision Transformer-based monocular depth estimation model from the original Depth Anything family.