A monocular depth estimation model based on BEiT transformer backbone, producing dense depth maps from single images.