A small Vision Transformer with patch size 16, pretrained on ImageNet-21k with AugReg and fine-tuned on ImageNet-1k for classification.