Models / google / ViT Base (Patch 16, 384 resolution)

ViT Base (Patch 16, 384 resolution)

JSON →
googlevision

A base Vision Transformer model with 16x16 patch size and 384x384 input resolution for image classification.

imagevision
Specs
Context & limits
context window
max output
Lifecycle
Dates

No lifecycle dates recorded.

Resources

No resource links recorded.