A Vision Transformer base model with patch size 16, pretrained on ImageNet-21k with AugReg for general-purpose image representation.