A monocular depth estimation model combining DINOv3 vision transformer features with a DPT decoder head.