A lightweight MaxViT variant combining convolution and attention, trained on ImageNet-1k with RandAugment.