Classification
2- Complexity
- High
- Impact area
- Technical
A neural architecture paradigm based on self-attention for sequential and multimodal data. Common foundation for large language, vision and multimodal models.
360° overview
Six perspectives place the building block in context. The numbers show where each perspective continues in the reading path.
The building block at a glance
Transformer
360°
Integrations
PyTorch for model implementation
+2
Use self-attention as the central representation mechanism.
Value stream stage: Build