Kafka At-Least-Once Delivery: Understanding the Risk of Duplicates

In the previous article, we explored Kafka’s At-Most-Once delivery and saw what happens when a partition leader crashes while messages are being produced. With a producer configured not to wait for acknowledgements or retry failed deliveries, one of our messages was lost.

This time, we’ll change the configuration and ask Kafka for a stronger guarantee: Don’t silently give up on a message. Keep trying when a retryable failure occurs.

This brings us to At-Least-Once delivery.

At-Least-Once changes the trade-off. Instead of accepting message loss, we make delivery more reliable through acknowledgements and retries. But retries introduce another problem: the same message may be delivered or processed more than once.

We’ll test this by running a three-broker Kafka cluster, producing ten messages to a known partition, deliberately crashing its leader while messages are being sent, and then checking what actually survived.

Prefer to watch? The complete experiment is available here.

What Does At-Least-Once Mean?

At-Least-Once means that the system tries to ensure that a message is delivered one or more times. The important difference from At-Most-Once is what happens when delivery becomes uncertain. With At-Most-Once, we accepted the possibility of losing a message. With At-Least-Once, we retry.

The Tax-Refund Letter Analogy

Think of a tax office sending you an important letter after you file your tax return. The delivery person visits your house while you are away and leaves the letter with a neighbour. The tax office expects a response, but none arrives. Because it cannot be sure that you received the first letter, it sends the same letter again. When you return, you may eventually receive both copies.

Kafka faces the same uncertainty. A broker may have received and replicated a record, yet fail before its acknowledgement reaches the producer. The producer cannot safely know whether the write succeeded, so it retries.

Producer Configuration

The transaction producer used for this experiment changes the important settings from the At-Most-Once case:

props.put(ACKS_CONFIG, "all");
props.put(RETRIES_CONFIG, Integer.MAX_VALUE);
props.put(ENABLE_IDEMPOTENCE_CONFIG, false);

acks=all requires the leader to wait for the required in-sync replicas before acknowledging the write. The test topic uses replication factor 3 and min.insync.replicas=2 retries=Integer.MAX_VALUE gives the producer a very large retry count for retryable failures. Other limits, including delivery timeout, still apply. enable.idempotence=false is intentional so the experiment can illustrate the duplicate risk of non-idempotent retries.

How a Producer Retry Can Create a Duplicate

Suppose the producer sends record B. The leader receives B and replicates it. Before the acknowledgement reaches the producer, the leader crashes. Kafka elects a new leader, and that new leader may already contain B. Because the producer never saw the acknowledgement, it retries B. With idempotence disabled, that retry can create a duplicate.

The Kafka Test Environment

The experiment runs Kafka locally through Docker Compose in KRaft mode with three controllers and three brokers. Before the test, I recreate the transactionIds topic so the log is clean. The topic has three partitions, replication factor 3, and min.insync.replicas=2. I also inspect the internal consumer-offset topic before introducing the broker failure.

Targeting Partition 0

For a controlled failure test, the producer deliberately sends every test transaction ID to partition 0. This is not required by At-Least-Once delivery; it simply gives us a predictable leader to stop.

public void send(String transactionId) {
    transactionKafkaTemplate.sendDefault(0, UUID.randomUUID().toString(), transactionId).thenAccept(result - >
            log.info("Successfully sent transactionId: {}", transactionId))
        .exceptionally(ex - > {
            log.error("Error sending transactionId: {}", transactionId, ex);
            return null;
        });
}

Before the test, describing the topic shows node 4 as the leader of partition 0, with nodes 4, 5, and 6 in the ISR. In the Docker Compose setup, node 4 is Broker 1.

Consumer Configuration

The consumer disables automatic commits and uses Spring Kafka manual acknowledgement. It also polls one record at a time so the behaviour is easier to observe.

props.put(ENABLE_AUTO_COMMIT_CONFIG, false);
props.put(MAX_POLL_RECORDS_CONFIG, 1);
factory.setBatchListener(true);
factory.getContainerProperties().setAckMode(ContainerProperties.AckMode.MANUAL);

After processing, the listener acknowledges the record manually. The sleep in the demo exists only to make processing visible; it is not production logic.

Payment payment = paymentService.findByTransactionId(transactionId);
if (payment == null) {
    acknowledgment.acknowledge();
    Thread.sleep(2000);
    return;
}
payment.setProcessedAt(LocalDateTime.now());
paymentService.update(payment);
Thread.sleep(2000);
acknowledgment.acknowledge();

Crashing the Partition Leader

With the application running, I use Postman to trigger ten messages, T0 through T9. As soon as production begins, I stop Broker 1, the current leader of partition 0, to simulate a broker failure during delivery.

Kafka Elects a New Leader

Kafka stabilizes and elects node 5 as the new leader of partition 0. The ISR is now nodes 5 and 6. Those two surviving replicas still satisfy min.insync.replicas=2.

The Actual Result

After the cluster stabilizes, consuming the topic shows T0, T1, T2, T3, T4, T5, T6, T7, T8, and T9. No message was lost. There was also no duplicate in this particular run. That does not contradict At-Least-Once semantics: duplicates are possible, not mandatory. The failure must occur in the uncertainty window where the record has already been stored but the producer has not received the acknowledgement.

The Consumer Side: A Useful Exercise

The video deliberately leaves the consumer-side failure as an exercise. With automatic commits disabled and manual acknowledgement enabled, stop the consumer after the business operation succeeds but before acknowledgment.acknowledge() executes. Restart the consumer and observe that the previously processed record can be delivered again because its offset was not committed.

This is why At-Least-Once applications often need idempotent business operations or a deduplication strategy. Payments, orders, inventory changes, and other side effects must be safe when the same logical message is encountered again.

What This Experiment Shows

The broker crash demonstrates the reliability side of At-Least-Once delivery. The producer waits for acknowledgement and retries retryable failures; Kafka elects a replacement leader; and in this run all ten records survive. The cost of retrying uncertain work is that the application must be prepared for duplicates. At-Most-Once chooses not to retry uncertain delivery and therefore accepts possible loss. At-Least-Once chooses to retry uncertain delivery and therefore accepts possible duplication. The appropriate trade-off depends on the application’s requirements and whether repeated business effects are safe.

Final Thoughts

At-Least-Once is a way of handling uncertainty in a distributed system. A producer can lose an acknowledgement even though Kafka stored the record; a leader can fail and be replaced; and a consumer can complete a database update but crash before committing its offset. In this experiment, Kafka recovered from the partition-leader failure and all ten records remained intact. The next article in the series will explore Kafka Exactly-Once delivery and how its mechanisms change the duplicate problem.

Watch the complete experiment here.

Recent Posts