AI Mode conversation: Multi-Head Attention in Transformer architectures.

 

Images

AI Mode conversation: Multi-Head Attention in Transformer architectures.

Multi-Head Attention in Transformer architectures.
Searching

AI Mode reply for Multi-Head Attention in Transformer architectures.

Multi-Head Attention (MHA) is the core building block of the Transformer architecture, originally introduced in the groundbreaking 2017 paper "Attention Is All You Need". It extends basic self-attention by allowing a model to simultaneously process information from different representation subspaces, providing a richer contextual understanding of sequences.

Why Use Multiple Heads?
In a single-head self-attention layer, a token can only focus on one specific type of relationship at a time. However, natural language requires understanding multiple overlapping contextual nuances simultaneously.
MHA resolves this limitation by running several attention mechanisms in parallel:
  • Diverse Perspectives: One head might focus on local syntactic links (e.g., subject-verb agreement), another tracks long-range semantic context, while a third may pay attention to structural positioning.
  • Efficiency Through Splitting: Rather than duplicating full-sized layers, MHA splits the model's total feature dimension (
    𝑑model
    ) across h heads. Each individual head operates in a lower-dimensional subspace (
    ), keeping the total computational cost comparable to a single full-width head while massively increasing parallelizability.

Step-by-Step Mathematical Workflow
Given an input sequence matrix X, the workflow for calculating Multi-Head Attention proceeds as follows:
1. Linear Projection
The input sequence is projected into three distinct spaces—Queries (Q), Keys (K), and Values (V). For each head i (out of h total heads), distinct, learnable weight matrices are used:
𝑄𝑖=𝑋𝑊𝑖𝑄,𝐾𝑖=𝑋𝑊𝑖𝐾,𝑉𝑖=𝑋𝑊𝑖𝑉
2. Scaled Dot-Product Attention
Each head independently computes its own attention weights using a scaled dot product. This measures how strongly each token relates to every other token in the sequence:
Head𝑖=Attention(𝑄𝑖,𝐾𝑖,𝑉𝑖)=softmax(𝑄𝑖𝐾𝑖𝑇𝑑𝑘)𝑉𝑖
  • Note: Dividing by
    𝑑𝑘
    scales the dot products down to prevent gradients from vanishing during training.
3. Concatenation and Final Projection
The output embeddings from all parallel heads are concatenated side by side, restoring the original hidden dimension size (
). A final learned linear layer (
𝑊𝑂
) blends the combined insights from all heads:
MultiHead(𝑄,𝐾,𝑉)=Concat(Head1,…,Headℎ)𝑊𝑂

Where It's Applied in Transformers
As detailed by structural guides on GeeksforGeeks, MHA is deployed in three different configurations depending on the objective:
Attention TypeContext & Data FlowPurpose
Encoder Self-AttentionEvery token attends to every other token in the input sequence.Generates rich, fully contextualised input representations.
Decoder Masked Self-AttentionTokens only attend to themselves and prior tokens; future tokens are hidden by a causal mask.Ensures autoregressive token generation behaves correctly without "cheating".
Decoder Cross-AttentionQ comes from the decoder, while K and V come from the encoder's output.Bridges the encoder and decoder, mapping source information to target tokens during tasks like translation.

No comments:

Post a Comment

Welcome.

Welcome You find here a selection of A I topics and concepts. Some at least could be adopted by local U3A/AI discussion groups. An important...