Kafka Exactly-Once Delivery: Surviving Broker Failure

In the previous article in this series, we explored Kafka At-Least-Once delivery. We saw that even when the leader of the partition receiving our messages was stopped, Kafka was able to recover and our messages remained intact. But At-Least-Once has an important trade-off, because the producer retries when delivery is uncertain, duplicate messages are possible.

What happens when duplicates are not acceptable? That brings us to the third delivery guarantee in this series: Exactly-Once delivery.

In this article, we will look at how Kafka combines idempotence and transactions to prevent duplicate writes, how a consumer can process a record and commit its offset within a transaction, and finally test the setup by deliberately stopping a Kafka broker while messages are flowing through the system.

This follows the same experiment demonstrated in my Exactly-Once Delivery video. If you prefer watching the practical demonstration rather than reading through the article, you can watch it here.

The experiment starts from the problem left by At-Least-Once: the messages survived the broker failure, but producer retries mean duplication remains possible.

What Does Exactly-Once Mean?

The goal of Exactly-Once delivery is straightforward:

A message should not be lost, and the same message should not be processed more than once.

Kafka achieves this by combining several mechanisms rather than relying on a single configuration property.

On the producer side, Kafka can identify a producer and associate sequence numbers with the records it sends. On the processing side, Kafka transactions allow several related operations to either succeed together or fail together.

The important idea is all or nothing.

If the transaction succeeds, its operations are committed. If it does not succeed, consumers configured to read only committed data will not treat the aborted transactional records as successfully committed work. The video introduces Exactly-Once using precisely this transactional, all-or-nothing model.

A Delivery Analogy

To make this easier to understand, consider the delivery-company example from the video. Imagine you order an expensive smartphone online.

For an expensive delivery, the delivery company may tell you when the package will arrive. When the courier reaches your house and hands over the phone, they ask you to sign for it.

Your signature provides confirmation that the delivery was completed.

Kafka’s Exactly-Once mechanism is more technical, of course, but the analogy introduces an important concept: the system needs a reliable way of identifying and tracking what has already been successfully delivered.

Producer IDs and Sequence Numbers

When a Kafka producer uses idempotence, Kafka has a mechanism for identifying the producer. The producer also maintains sequence numbers per partition. Conceptually, suppose a producer is writing to partition 0:

Producer ID: 42

Partition 0:
    Record A -> sequence 0
    Record B -> sequence 1
    Record C -> sequence 2

The broker tracks the producer and sequence information associated with records appended to the partition.

This becomes particularly useful when retries happen. Imagine that the producer sends record B. Kafka successfully appends it, but something goes wrong before the producer knows that the operation succeeded.

The producer may retry B.

Without idempotence, that retry could result in another copy being appended. With producer idempotence enabled, Kafka can recognize the retry and prevent it from becoming another successfully appended copy.

This is the mechanism discussed early in the video: the producer and broker use the producer identity and monotonically increasing sequence numbers to detect duplicate writes, with each partition maintaining its own sequence progression.

Adding Kafka Transactions

The producer configuration used for the Exactly-Once example keeps the reliability settings we used previously and adds the properties required for idempotence and transactions.

Conceptually, the important configuration is:

props.put(ProducerConfig.ACKS_CONFIG, "all");
props.put(ProducerConfig.RETRIES_CONFIG, Integer.MAX_VALUE);
props.put(ProducerConfig.ENABLE_IDEMPOTENCE_CONFIG, true);
props.put(ProducerConfig.TRANSACTIONAL_ID_CONFIG, "your-txn-id");

The important change from our At-Least-Once experiment is:

props.put(ProducerConfig.ENABLE_IDEMPOTENCE_CONFIG, true);

and the introduction of a transactional ID:

props.put(ProducerConfig.TRANSACTIONAL_ID_CONFIG, "your-txn-id");

The Consumer Must Read Committed Data

Transactions also affect how we consume records. In the consumer configuration, we add:

props.put(
    ConsumerConfig.ISOLATION_LEVEL_CONFIG,
    "read_committed"
);

This is important because records can be written as part of a transaction before that transaction reaches its final committed state.

A consumer configured with read_committed reads records from successfully committed transactions rather than treating transactional records that are not committed as completed data.

The Consume → Process → Produce Transaction

This is where the Exactly-Once example becomes particularly interesting.

Our application uses two topics: payments and payment-analytics A producer writes payment events into payments. A consumer reads those events and performs the processing required for the analytics flow. So the logical pipeline is: Payments Producer → payments → Transactional Consumer/Processor → payment-analytics.

Outside the transaction boundary: payments topic → transaction → payment-analytics topic.

Why We Don’t Manually Acknowledge the Message

There is an interesting difference between this example and the previous At-Least-Once implementation. The consumer configuration still uses manual acknowledgement configuration, but inside our transactional processing flow we do not manually call acknowledgment.acknowledge();Instead, the consumed offset is sent into the transaction. The relevant operation demonstrated in the video is sendOffsetsToTransaction(...) the idea is that the consumed offset becomes part of the same transaction as the output produced by the processing. If the transaction commits, the output and the corresponding offset are committed together.

This gives us the core processing model:

Consume the record, process it, produce the result, and include the consumed offset in the same transaction.

That is the important step beyond simply enabling idempotence on a producer.

Our Kafka Test Environment

Now that the application is configured, we can test what happens during a real broker failure. As in the previous At-Most-Once and At-Least-Once experiments, the Kafka cluster runs through Docker Compose in KRaft mode.

The cluster contains three Kafka controllers and three Kafka brokers.

For this experiment, we use two topics:

payments
payment-analytics

Both are recreated before the test to give us a clean environment. Each topic has:

Partitions:                      3
Replication factor:        3
min.insync.replicas:      2

Finding the Broker We Want to Crash

Because the experiment intentionally targets a failure while records are being produced, we first identify the leader of partition 0. For the payments topic Partition 0 -> Leader: Node 4 we then check payment-analytics. It also has Partition 0 -> Leader: Node 4 this gives us a convenient failure target.

Starting the Failure Test

The Java application is started and waits for us to trigger the test. Using Postman, we invoke the endpoint that causes the controller to generate 10 messages.

Immediately after triggering the request, we stop Broker 1. That means Kafka is actively dealing with records while the broker that leads partition 0 is being removed from the cluster.

The sequence is deliberately simple:

  1. Trigger the endpoint.
  2. Start producing 10 messages.
  3. Stop Broker 1 / Node 4.
  4. Allow Kafka to recover.
  5. Inspect both topics.

Kafka Elects a New Leader

Once Node 4 disappears, Kafka detects that the affected partitions have lost their leader and elects replacements. We describe the payments topic again.

Partition 0 now reports Leader: Node 5, we then inspect payment-analytics its partition 0 also reports Leader: Node 5.

This makes the failure being tested much easier to see than simply showing the Docker commands.

Did We Lose or Duplicate Any Messages?

Now comes the important part of the experiment. We sent 10 messages, First, we count the records in the payments topic. The result is payments = 10 All ten messages reached the input topic. Next, we count the records in payment-analytics. The result is also payment-analytics = 10, so after deliberately stopping the broker that originally led partition 0, the experiment ends with:

Sent:                             10

payments:                    10
payment-analytics:     10

Putting the Pieces Together

What makes this different from the previous At-Least-Once example is that several mechanisms are working together.

The producer is configured for acknowledgement and retries, but it also has idempotence enabled. Kafka can therefore use producer identity and sequence numbers to recognize retries instead of blindly treating every retry as a new successful append. The producer also participates in transactions.

On the consumer side, we use:

isolation.level=read_committed

so the consumer reads committed transactional data.

During transactional processing, the application consumes the payment record, processes it, produces the resulting analytics record, and includes the consumed offset in the transaction.

Exactly-Once vs. At-Least-Once

The difference becomes clearer when we compare this experiment with the previous one. With At-Least-Once, we configured the producer to wait for acknowledgements and retry failures. This improved reliability, but retries introduced the possibility that the same logical record could be written or processed again.

Exactly-Once builds on that reliability model by introducing idempotence and transactional processing.

Instead of treating: consume → process → produce → offset commit as unrelated operations, we make the Kafka operations part of the same transactional unit. Either the transaction commits successfully, or it does not become visible to read_committed consumers as successfully committed work.

One Important Boundary

There is an important detail about the experiment worth keeping clear. The Exactly-Once pipeline demonstrated here is primarily:

Kafka → processing → Kafka

Our transactional boundary covers Kafka records and Kafka consumer offsets.

For example:

payments → processor → payment-analytics

That is different from assuming that enabling Kafka transactions automatically makes an unrelated external database update part of the same Kafka transaction.

The video’s practical test verifies the two Kafka topics and the transactional Kafka processing path. It does not test an atomic Kafka-plus-external-database transaction. That distinction matters when applying this pattern to a production system.

What We Learned

This experiment started with the weakness we observed in At-Least-Once delivery: retrying uncertain writes can create duplicates.

We then configured the producer for idempotence and transactions, configured the consumer to read committed transactional data, and made the consumed offset part of transactional processing.

Finally, we deliberately stopped Node 4 while ten messages were being produced.

Kafka elected Node 5 as the new leader for the affected partition 0 in both topics, and after the system recovered we found exactly ten records in payments and exactly ten in payment-analytics. That practical failure test brings together the major concepts demonstrated throughout the video:

idempotent producer + transactional producer + read_committed consumer + transactional offset commit.

The final slide captures the same configuration visually, including acks=all, retries, enable.idempotence=true, transactional.id, enable.auto.commit=false, and isolation.level=read_committed.

For event-driven systems where duplicate processing can have serious consequences—such as payment and financial event pipelines—understanding these mechanisms is essential.

If you would like to see the entire experiment, including the Kafka cluster, Java configuration, broker shutdown, leader election, and final message counts, you can watch the accompanying video.

Recent Posts