Atlas · GenAI 2026
Multimodal AI
Multimodal Foundation Models (VLMs)
conceptPeak: 2024Multimodal ArchitecturesAI consensus: 3/3
Prerequisites
VLMs combine vision encoders (often ViT = Vision Transformer) with LLM decoders via projection layers — both sides are Transformer-based
Many VLMs use CNN-based vision backbones or their concepts (feature maps, pooling) even when the main architecture is a ViT
Recommended reference
Liu et al. (2024) 'LLaVA: Visual Instruction Tuning' — NeurIPS; foundational VLM architecture paper; plus OpenAI GPT-4V system card
Notes from AI deep research
Anthropic Opus
GPT-4o, Gemini, Claude — vision+text+audio. Vision encoder → projector → LLM backbone
OpenAI Deep Research
Inne metryki i pipeline danych [OA#19]
Google Deep Think
Natywna korelacja tekst/wideo/dźwięk/obraz [G#27]
Related skills
- ← is subcategory of: Vision-Language Models(3/3)
- ← is subcategory of: Audio AI(1/3)
- → is subcategory of: Deep Learning(1/3)