Video summary

Workshop #5: Vision-Language Model (VLM) và AI đa phương thức

Main summary

Key takeaways

Technology

Overview: Vision-Language Models (VLM/VLM) and Multimodal AI Workshop

The speaker introduces a progression of multimodal models that align images and text in a shared embedding space. This is extended to stronger VLMs that support captioning, image-text matching, and image-conditioned language modeling. The talk focuses mainly on CLIP, BLIP, and BLIP-2.


1) CLIP: Contrastive Vision-Language Alignment (baseline idea)

Goal

  • Train separate image and text encoders so that:
    • a correct image–text pair produces embeddings close in vector space,
    • incorrect pairs are far apart.
  • The core emphasis is that CLIP is built around contrastive learning.

Architecture & training flow

  1. Image encoder (e.g., CNN or Vision Transformer):
    • Converts each image into an image embedding vector.
  2. Text encoder:
    • Uses a Transformer-like sentence encoder (contrasted with older RNN/LSTM-style approaches).
  3. Training uses batches of N images and N matching captions:
    • Build an N×N cosine similarity matrix between image embeddings and text embeddings.
    • Diagonal entries are correct pairs; off-diagonal are incorrect.

Similarity metric and loss intuition

  • Uses cosine similarity (range -1 to 1).
  • Temperature scaling: multiply cosine scores by a factor (e.g., 100) to sharpen differences.
  • Training does not force:
    • diagonal similarities to equal 1,
    • off-diagonal similarities to equal 0.
    • Instead, it ensures diagonal > off-diagonal.
  • The loss is computed in two directions:
    • Image→Text (row-wise logic)
    • Text→Image (column-wise logic)
  • Final CLIP loss is the average of both directions.

Inference use cases

  • Zero-shot classification:
    • encode an image and compare it to embeddings of candidate prompts (e.g., “a dog”, “a car”).
  • Text-based image retrieval:
    • query text → closest matching image embedding.
  • Mentioned limitations:
    • high compute cost (both encoders must be trained),
    • stronger emphasis on retrieval/embedding similarity, not rich generative Q&A or captioning.

2) BLIP: Upgrade from CLIP to Multitask (contrastive + matching + generation)

The speaker states that BLIP builds on CLIP but adds richer training objectives.

Components

  • Vision Transformer as the image backbone.
  • Text processing with Transformer-style mechanics using:
    • token embeddings and a CLS token.
  • Attention blocks include:
    • self-attention (conceptually),
    • cross-attention between image tokens and text tokens.

Three training objectives (major expansion vs CLIP)

  1. ITC (Image-Text Contrastive)

    • Similar to CLIP-style contrastive alignment.
    • Uses CLS vectors from both modalities to compute similarity.
  2. ITM (Image-Text Matching)

    • Predicts whether a caption matches the image (binary match score).
    • Uses stronger image-text interaction than ITC.
    • Trained with cross-entropy loss (score passed through sigmoid).
  3. LM (Language Modeling) (image-conditioned text generation)

    • Uses causal attention (autoregressive) to predict the next token.
    • Conditioning comes from the image through cross-attention, enabling image-grounded generation.
    • Example idea: generating a caption continuation like “kitten next to blue fence” based on the image.

Limitations mentioned for BLIP

  • The language “decoder” is described as not large enough / not fully leveraging large pretrained LMs, motivating BLIP-2.

3) BLIP-2: Stronger by reusing pretrained components + training a small “bridge”

BLIP-2 is presented as a more practical and powerful variant that reduces training cost and improves capabilities.

Key idea: reuse pretrained models, train only what’s necessary

Rather than training large parts from scratch, BLIP-2:

  • keeps the Vision Transformer image encoder (pretrained),
  • keeps a large pretrained language model (LLM),
  • trains a smaller Query Transformer / “bridge transformer” that converts image features into tokens the LLM can use.

Why not feed image embeddings directly to the LLM?

  • The LLM expects inputs in its own token/representation space.
  • Image encoder outputs live in a different embedding space.
  • A learned projection/bridge is required to translate image information into LLM-compatible “visual tokens”.

The Query Transformer (“Q-Former” / “Wifmer” in subtitles)

  • Uses learnable query vectors (e.g., 32 query tokens).
  • Training concept:
    • start with random query vectors,
    • learn them to extract 32 information units from the image.
  • Bridge mechanism:
    • cross-attention where queries attend to image tokens,
    • producing 32 output visual tokens.

Projection + prompt-based generation

  • Query outputs are passed through a projection layer/adapters to match the LLM hidden size.
  • The LLM receives:
    • prompt tokens (instruction, question, etc.),
    • plus projected visual tokens as additional context.
  • Enables tasks like:
    • image captioning
    • image question answering (VQA) without retraining (as described).

Two-stage training process (important in the talk)

  1. Stage 1: train Query Transformer + visual tokens with multimodal losses:
    • contrastive objective,
    • image-text matching,
    • language modeling (next-word prediction).
  2. Stage 2: keep the Q-Former output fixed, then train projection/adapter and integrate with the LLM to improve generation.

Practical customization described

  • To adapt to new datasets or focus on different objects:
    • modify/retrain mainly the Q-Former or the projection/adapter,
    • rather than rebuilding the vision encoder or language model.
  • Mentioned related models:
    • LLaVA and other “vision vs language-only” style variants (speaker notes differences depend on whether vision is included).

Limitations mentioned

  • BLIP-2 diagrams are described as more complex/confusing, requiring careful review.
  • The model structure is reduced conceptually to:
    1. Q-Former (bridge)
    2. projection layer
    3. large language model

Key tutorial-like takeaways (concept checklist)

CLIP

  • Build an N×N cosine similarity matrix
  • Diagonal pairs high, off-diagonal low
  • Use temperature scaling
  • Optimize contrastive learning in both directions

BLIP

  • Add ITM (matching head) and LM (autoregressive generation with image conditioning)

BLIP-2

  • Replace end-to-end retraining with:
    • learnable query tokens + Q-Former bridge
  • Reuse pretrained vision encoder and pretrained LLM
  • Enable captioning and VQA efficiently

Main speakers / sources

  • Main speaker: the workshop lecturer (no specific name provided in subtitles)
  • Primary referenced model families / papers: CLIP, BLIP, BLIP-2, and related multimodal variants (e.g., LLaVA-like)

Original video