Video summary
Workshop #5: Vision-Language Model (VLM) và AI đa phương thức
Main summary
Key takeaways
Overview: Vision-Language Models (VLM/VLM) and Multimodal AI Workshop
The speaker introduces a progression of multimodal models that align images and text in a shared embedding space. This is extended to stronger VLMs that support captioning, image-text matching, and image-conditioned language modeling. The talk focuses mainly on CLIP, BLIP, and BLIP-2.
1) CLIP: Contrastive Vision-Language Alignment (baseline idea)
Goal
- Train separate image and text encoders so that:
- a correct image–text pair produces embeddings close in vector space,
- incorrect pairs are far apart.
- The core emphasis is that CLIP is built around contrastive learning.
Architecture & training flow
- Image encoder (e.g., CNN or Vision Transformer):
- Converts each image into an image embedding vector.
- Text encoder:
- Uses a Transformer-like sentence encoder (contrasted with older RNN/LSTM-style approaches).
- Training uses batches of N images and N matching captions:
- Build an N×N cosine similarity matrix between image embeddings and text embeddings.
- Diagonal entries are correct pairs; off-diagonal are incorrect.
Similarity metric and loss intuition
- Uses cosine similarity (range -1 to 1).
- Temperature scaling: multiply cosine scores by a factor (e.g., 100) to sharpen differences.
- Training does not force:
- diagonal similarities to equal 1,
- off-diagonal similarities to equal 0.
- Instead, it ensures diagonal > off-diagonal.
- The loss is computed in two directions:
- Image→Text (row-wise logic)
- Text→Image (column-wise logic)
- Final CLIP loss is the average of both directions.
Inference use cases
- Zero-shot classification:
- encode an image and compare it to embeddings of candidate prompts (e.g., “a dog”, “a car”).
- Text-based image retrieval:
- query text → closest matching image embedding.
- Mentioned limitations:
- high compute cost (both encoders must be trained),
- stronger emphasis on retrieval/embedding similarity, not rich generative Q&A or captioning.
2) BLIP: Upgrade from CLIP to Multitask (contrastive + matching + generation)
The speaker states that BLIP builds on CLIP but adds richer training objectives.
Components
- Vision Transformer as the image backbone.
- Text processing with Transformer-style mechanics using:
- token embeddings and a CLS token.
- Attention blocks include:
- self-attention (conceptually),
- cross-attention between image tokens and text tokens.
Three training objectives (major expansion vs CLIP)
-
ITC (Image-Text Contrastive)
- Similar to CLIP-style contrastive alignment.
- Uses CLS vectors from both modalities to compute similarity.
-
ITM (Image-Text Matching)
- Predicts whether a caption matches the image (binary match score).
- Uses stronger image-text interaction than ITC.
- Trained with cross-entropy loss (score passed through sigmoid).
-
LM (Language Modeling) (image-conditioned text generation)
- Uses causal attention (autoregressive) to predict the next token.
- Conditioning comes from the image through cross-attention, enabling image-grounded generation.
- Example idea: generating a caption continuation like “kitten next to blue fence” based on the image.
Limitations mentioned for BLIP
- The language “decoder” is described as not large enough / not fully leveraging large pretrained LMs, motivating BLIP-2.
3) BLIP-2: Stronger by reusing pretrained components + training a small “bridge”
BLIP-2 is presented as a more practical and powerful variant that reduces training cost and improves capabilities.
Key idea: reuse pretrained models, train only what’s necessary
Rather than training large parts from scratch, BLIP-2:
- keeps the Vision Transformer image encoder (pretrained),
- keeps a large pretrained language model (LLM),
- trains a smaller Query Transformer / “bridge transformer” that converts image features into tokens the LLM can use.
Why not feed image embeddings directly to the LLM?
- The LLM expects inputs in its own token/representation space.
- Image encoder outputs live in a different embedding space.
- A learned projection/bridge is required to translate image information into LLM-compatible “visual tokens”.
The Query Transformer (“Q-Former” / “Wifmer” in subtitles)
- Uses learnable query vectors (e.g., 32 query tokens).
- Training concept:
- start with random query vectors,
- learn them to extract 32 information units from the image.
- Bridge mechanism:
- cross-attention where queries attend to image tokens,
- producing 32 output visual tokens.
Projection + prompt-based generation
- Query outputs are passed through a projection layer/adapters to match the LLM hidden size.
- The LLM receives:
- prompt tokens (instruction, question, etc.),
- plus projected visual tokens as additional context.
- Enables tasks like:
- image captioning
- image question answering (VQA) without retraining (as described).
Two-stage training process (important in the talk)
- Stage 1: train Query Transformer + visual tokens with multimodal losses:
- contrastive objective,
- image-text matching,
- language modeling (next-word prediction).
- Stage 2: keep the Q-Former output fixed, then train projection/adapter and integrate with the LLM to improve generation.
Practical customization described
- To adapt to new datasets or focus on different objects:
- modify/retrain mainly the Q-Former or the projection/adapter,
- rather than rebuilding the vision encoder or language model.
- Mentioned related models:
- LLaVA and other “vision vs language-only” style variants (speaker notes differences depend on whether vision is included).
Limitations mentioned
- BLIP-2 diagrams are described as more complex/confusing, requiring careful review.
- The model structure is reduced conceptually to:
- Q-Former (bridge)
- projection layer
- large language model
Key tutorial-like takeaways (concept checklist)
CLIP
- Build an N×N cosine similarity matrix
- Diagonal pairs high, off-diagonal low
- Use temperature scaling
- Optimize contrastive learning in both directions
BLIP
- Add ITM (matching head) and LM (autoregressive generation with image conditioning)
BLIP-2
- Replace end-to-end retraining with:
- learnable query tokens + Q-Former bridge
- Reuse pretrained vision encoder and pretrained LLM
- Enable captioning and VQA efficiently
Main speakers / sources
- Main speaker: the workshop lecturer (no specific name provided in subtitles)
- Primary referenced model families / papers: CLIP, BLIP, BLIP-2, and related multimodal variants (e.g., LLaVA-like)