You Only Look at One Sequence (YOLOS) tiny model for object detection, using a vision transformer approach without convolutions.