Transformers are a deep-learning architecture based on self-attention that enables efficient processing of sequential data. They replaced recurrence in NLP and power large-scale models for language, vision, and multimodal tasks. Transformers enable parallelization and long-range context modeling but require significant compute and large datasets.
Use this profile to understand the building block briefly, place it in the model, and switch to the 360° assessment when needed.
Theoretical construct: explains a term, principle, or mental model.
What you need to understand to reason about a domain.
A Transformer is a neural network architecture that uses self-attention to process relationships in sequences in parallel.
Researchers at Google and the University of Toronto introduced Transformers in the 2017 paper Attention Is All You Need.
Each token weights other tokens and forms context-dependent representations. The architecture gains parallelism but needs substantial data and compute.
Tokens weight other tokens.
Positions are processed in parallel.
Information is linked within the input window.
For modern AI models, the concept explains performance gains, resource needs, and context limits.
Where this building block is located in the topic model.
No structure path available.
Explore how this building block connects to concepts, methods, technologies, and tools.
These sources establish the term and its professional meaning.
All direct connections of the current building block in a compact text view.
This classification shows where the building block typically matters, how demanding it is, and what kind of impact it has in the model.
The level within the organization (enterprise, domain, team) at which the AssetBlock is applied.