Imagine an art restorer who is asked to inspect an authentic Renaissance masterpiece for tiny micro-cracks in the brushwork. If the restorer first takes a photograph of the painting, shrinks the resolution down to a tiny 1-inch blurry postage stamp, and then attempts to analyze the cracks through a magnifying glass, they will see nothing but a soup of gray pixels. For years, early Vision Transformers treated high-resolution images with that exact blindness.
The Fixed-Grid Bottleneck of Early Vision Transformers
Early Vision-Language models (such as original CLIP and early LLaVA) were constrained by fixed-resolution Vision Transformers. Regardless of whether an uploaded image was a 4K blueprint, a wide panorama, or a vertical infographic, the system aggressively downscaled and warped the image into a standard square grid (e.g. $224 \times 224$ or $336 \times 336$ pixels).
This aggressive downsampling caused catastrophic failures:
- Unreadable Tiny Text: Fine print in legal contracts, table footnotes, and chart legends were smeared into unreadable digital noise.
- Distorted Aspect Ratios: Tall mobile screenshots and wide architectural schematics were squashed into squares, destroying spatial geometry.
- Small Object Disappearance: Small components in electronic circuit boards vanished entirely from the vision token stream.
[Fixed Downsampling vs. Dynamic Patch Slicing (AnyRes)]
Fixed Downsampling (Lossy & Distorted):
4K Blueprint (3840×2160) ──► Forced Squash to 336×336 ──► Unreadable blurry text & lost lines!
Dynamic Patch Slicing (AnyRes / LLaVA-NeXT):
4K Blueprint (3840×2160)
├── Slice 1: [Top-Left Patch (336×336)] ──► Vision Tokens 1
├── Slice 2: [Top-Right Patch (336×336)] ──► Vision Tokens 2
├── Slice 3: [Bottom-Left Patch (336×336)] ──► Vision Tokens 3
├── Slice 4: [Bottom-Right Patch (336×336)] ──► Vision Tokens 4
└── Overview: [Global Downscaled (336×336)] ──► Global Context Tokens
│
▼
[Preserves 100% Native Resolution!]
The Mechanics of Dynamic Patch Slicing (AnyRes)
Modern Vision-Language architectures (including LLaVA-NeXT, Qwen-2-VL, and Claude 3.5 Sonnet) solve high-resolution perception using Dynamic Aspect-Ratio Patch Slicing:
- Native Aspect Ratio Partitioning: The engine analyzes the original image dimensions and dynamically partitions the canvas into a grid of high-resolution sub-patches (e.g. $2 \times 2$, $1 \times 3$, or $3 \times 2$ tiles) without warping aspect ratios.
- Dual-Level Feature Extraction: Each sub-patch is processed independently by the Vision Transformer to extract high-frequency local details (individual letters, tiny circuit traces). Simultaneously, a downscaled overview image is processed to provide global contextual awareness.
- Spatial 2D Positional Encoding: Special newline and spatial boundary tokens are injected into the visual token sequence, allowing the language model backbone to maintain exact 2D coordinate geometry during spatial reasoning.
Engineering Takeaway
Never squash or warp image inputs before feeding them to multimodal foundation models. Adopt dynamic patch slicing architectures that preserve native resolution and aspect ratios to unlock pinpoint spatial accuracy.