Kafka enables the publishing, storing, and processing of data streams in real-time. It is commonly used in modern architectures to integrate and process data between various systems.
Use this profile to understand the building block briefly, place it in the model, and switch to the 360° assessment when needed.
Technical building block: can be automated, integrated, or operated.
Concrete cog in the system that works inside larger relationships.
Apache Kafka is a distributed event streaming platform. Applications write events to durable data streams, while other applications can read and process those events independently.
Kafka was developed at LinkedIn by Jay Kreps, Neha Narkhede, and Jun Rao. The team needed a scalable way to move the company's growing event streams reliably and in real time between many data producers and consumers. That internal data infrastructure later became the open-source Apache project.
Think of Kafka as a shared, continuously growing logbook. Producers append events to topics. Each topic is split into partitions that are stored and replicated across brokers; order is preserved within a partition. Consumers read events at their own pace and track their position as an offset. Multiple consumers in one consumer group can share the partitions, while other groups independently read the same stream again.
An immutable record that something happened in a business or technical system.
A named stream to which producers write events and from which consumers read them.
An ordered part of a topic and the unit Kafka uses to distribute data and enable parallel processing.
A server in the Kafka cluster that stores partitions and handles requests from producer and consumer clients.
Producers publish events; consumers read and process them. Both sides remain temporally and technically decoupled.
A group distributes partitions among its consumers. The offset describes how far that group has read.
Kafka is relevant when many systems continuously exchange events, data streams need multiple uses, or traffic spikes need buffering. Typical uses include event-driven architecture, data integration, change data capture, telemetry, and streaming analytics. Its value comes from decoupling and reuse; partitioning, delivery guarantees, operations, and schema evolution still require deliberate design.
Where this building block is located in the topic model.
Explore how this building block connects to concepts, methods, technologies, and tools.
These sources establish the term and its professional meaning.
All direct connections of the current building block in a compact text view.
This classification shows where the building block typically matters, how demanding it is, and what kind of impact it has in the model.
The level within the organization (enterprise, domain, team) at which the AssetBlock is applied.