MapReduce is a distributed programming model for parallel processing of large datasets across clusters; it abstracts map and reduce phases and enables horizontal scaling. It simplifies fault tolerance and data partitioning, making it suitable for batch analytics, index construction and large-scale aggregations. Implementations optimize locality, scheduling a…
Use this profile to understand the building block briefly, place it in the model, and switch to the 360° assessment when needed.
Theoretical construct: explains a term, principle, or mental model.
What you need to understand to reason about a domain.
MapReduce is a programming model that splits large-scale data processing into parallel map steps followed by a reduce step.
In 2004, Google's Jeffrey Dean and Sanjay Ghemawat presented MapReduce to address the difficulty of parallelizing large data-processing jobs. The model shaped cluster computing, and Apache Hadoop later carried the idea into a widely used open-source system.
Many workers read separate slices and emit intermediate values with a key. Other workers then gather equal keys and compute one result per group.
Input records are processed independently into intermediate key-value pairs.
Intermediate values are grouped by key and assigned to matching reduce tasks.
Grouped values are combined into a condensed result.
MapReduce scales batch processing of large datasets when work can be split into independent transformations and groupings.
Where this building block is located in the topic model.
No structure path available.
Explore how this building block connects to concepts, methods, technologies, and tools.
These sources establish the term and its professional meaning.
All direct connections of the current building block in a compact text view.
This classification shows where the building block typically matters, how demanding it is, and what kind of impact it has in the model.
The level within the organization (enterprise, domain, team) at which the AssetBlock is applied.