Atlas · GenAI 2026

Multimodal AI

Multimodal Foundation Models (VLMs)

conceptPeak: 2024Multimodal ArchitecturesAI consensus: 3/3

Prerequisites

  • VLMs combine vision encoders (often ViT = Vision Transformer) with LLM decoders via projection layers — both sides are Transformer-based

  • Many VLMs use CNN-based vision backbones or their concepts (feature maps, pooling) even when the main architecture is a ViT

Recommended reference

Liu et al. (2024) 'LLaVA: Visual Instruction Tuning' — NeurIPS; foundational VLM architecture paper; plus OpenAI GPT-4V system card

Notes from AI deep research

Anthropic Opus

GPT-4o, Gemini, Claude — vision+text+audio. Vision encoder → projector → LLM backbone

OpenAI Deep Research

Inne metryki i pipeline danych [OA#19]

Google Deep Think

Natywna korelacja tekst/wideo/dźwięk/obraz [G#27]

Related skills