A small You Only Look at One Sequence (YOLOS) vision transformer for object detection, trained on COCO.