Kafka At-Most-Once Delivery: Understanding the Risk of Message Loss

Apache Kafka is often described as a reliable distributed event-streaming platform. But Kafka’s reliability does not mean that every message is automatically guaranteed to reach its destination.

The delivery behavior you get depends on how producers, brokers, and consumers are configured and how they behave when failures occur.

Kafka applications are commonly designed around three delivery semantics:

  • At-Most-Once — a message is processed zero or one time. Message loss is possible.
  • At-Least-Once — a message is processed one or more times. Duplicate processing is possible.
  • Exactly-Once — processing is designed so that each message has exactly one effect.

In this article, I will focus exclusively on At-Most-Once delivery.

Rather than discussing it only theoretically, we’ll deliberately create a failure. We will send messages to a Kafka partition and terminate its leader broker while messages are being produced. Then we’ll inspect the topic to see what survived.

This is the same experiment demonstrated in my YouTube video.

What Does At-Most-Once Actually Mean?

At-Most-Once can be summarized as: A message is delivered zero or one time. The important word is zero.

With this delivery semantic, the system accepts the possibility that a message may disappear during a failure rather than trying repeatedly to ensure that it is delivered.

A useful analogy is someone distributing advertising flyers door-to-door. Imagine a salesperson walking through a neighborhood and dropping flyers into mailboxes. Their job is simply to drop each flyer and continue walking.

They don’t wait for someone inside the house to confirm: “Yes, I received your flyer.” If something happens to the flyer after they leave, they don’t return with another copy. That is roughly the mentality we’re creating with our Kafka producer.

The diagram below captures this relationship: the producer sends without confirmation or retry, while a failure between the producer and Kafka can result in a message never reaching the consumer.

The Producer Configuration

For the experiment, the important producer settings are:

props.put(ACKS_CONFIG, "0");
props.put(RETRIES_CONFIG, 0);
props.put(ENABLE_IDEMPOTENCE_CONFIG, false);

acks=0

Kafka’s acks producer configuration determines what acknowledgement the producer requires from the broker before considering a request complete. With acks=0 the producer does not wait for an acknowledgement from the broker. From the producer’s perspective, once the request has been sent, it can continue. It therefore does not know whether the broker successfully persisted the record. This is essentially a fire-and-forget approach.

Conceptually:

Producer
   |
   | send(record)
   |
   +----------------------> Kafka Broker
   |
   | No acknowledgement required
   |
   +---- continue

retries=0

If delivery encounters a problem, the producer will not retry sending the failed record. Put the two configurations acks=0 and retries=0 together and we get the behavior we want to demonstrate. If something goes wrong at the wrong moment, there is no recovery attempt from the producer.

Building the Kafka Test Environment

For this experiment, I used a Kafka cluster running through Docker Compose in KRaft mode. The environment consists of 3 Kafka controllers and 3 Kafka brokers and the test topic is transactionIds I recreated the topic before running the experiment so that we started from a clean environment. Partitions: 3 Replication factor: 3 min.insync.replicas: 2.

The exact leadership assignment can change. What matters for this experiment is determining which broker currently leads partition 0.

Why Target Partition 0

Normally, a producer can allow Kafka’s partitioning logic to determine where a record should go. For this experiment, however, I intentionally send the transaction IDs to partition 0

transactionKafkaTemplate.sendDefault(
        0,
        UUID.randomUUID().toString(),
        transactionId
);

This isn’t something required for At-Most-Once delivery. It is an experimental control.

I need to know exactly which partition receives the records so that I can identify its leader and terminate that broker while records are being produced.

Now we have a predictable failure target.

Producer
    |
    v
Partition 0
    |
    v
Leader Broker

Finding the Partition Leader

Before producing the records, I describe the transactionIds topic. At the beginning of the experiment, Kafka reports that node 4 is the leader for partition 0.

Looking at the Docker Compose environment tells us that node 4 corresponds to broker-1 so our path is effectively: Java Producer - transactionIds -> Partition 0 -> broker-1 (leader).

What About min.insync.replicas=2?

Our topic has replication.factor=3 min.insync.replicas=2 but the producer in this experiment uses acks=0 the producer isn’t waiting for Kafka to confirm that replicas have received the record. The durability settings of the cluster don’t magically turn a fire-and-forget producer into one that waits for durable replication. A Kafka system’s reliability is the result of multiple configurations working together, not one isolated setting.

The Consumer Side

The consumer configuration in the experiment disables automatic offset commits props.put(ENABLE_AUTO_COMMIT_CONFIG, false); and configures Spring Kafka for manual acknowledgement.

factory.getContainerProperties()
       .setAckMode(ContainerProperties.AckMode.MANUAL);

The consumer explicitly acknowledges the record after processing.

payment.setProcessedAt(LocalDateTime.now());
paymentService.update(payment);

acknowledgment.acknowledge();

The experiment also uses MAX_POLL_RECORDS_CONFIG = 1 to make the behavior easier to observe one record at a time. These consumer settings help control the experiment, but there is an important distinction here.

The message loss we’re trying to demonstrate occurs before the consumer can process the missing record. If the record doesn’t survive the producer/broker side of the pipeline, consumer acknowledgement settings cannot recover it.

Now Let’s Break Kafka

With everything running, we send ten transaction IDs T1...T10 while the producer is sending them, we deliberately stop broker-1 — the broker that currently leads partition 0.

This simulates a broker failure occurring during message delivery. Conceptually, we are trying to create this situation:

Producer
   |
   | T1  T2  T3  T4  T5  T6 ...
   |
   v
+----------------+
| Partition 0    |
| Leader         |  X  <-- broker crashes
| broker-1       |
+----------------+
        |
        X

The timing is intentional. We want a failure to happen while records are in flight.

Kafka Elects a New Leader

A broker failure does not mean that the entire Kafka cluster stops. Because partition 0 has replicas on other brokers, Kafka can elect another replica as its leader. After terminating broker-1, we describe the topic again.

Before: Partition 0 → Leader: node 4

After the failure: Partition 0 → Leader: node 5.

The cluster continues operating. This demonstrates one of Kafka’s important distributed-system properties: partition leadership can move when a broker becomes unavailable. But there is another question. What happened to the messages that were being sent during that transition?

Counting What Actually Reached Kafka

We now consume the records from the topic and inspect what actually exists. Remember that we attempted to produce ten transaction IDs. But when we examine the topic, one transaction ID is missing. In the video experiment, that record is T5. We see later records T6,T7,T8,T9... but T5 is not there. That is the result we were trying to reproduce.

The producer sent the record, but because we configured it not to wait for acknowledgement and not to retry, the application had no mechanism to ensure that the record survived the broker failure. From the producer’s perspective, there was nothing more to do.

From the application’s perspective:

T4 ────────────────> Kafka ✓
T5 ───────> X                 LOST
                 broker failure
                       ↓
                 leader election
T6 ────────────────> Kafka ✓
T7 ────────────────> Kafka ✓

And that’s the danger of At-Most-Once delivery.

“Send Successful” Does Not Necessarily Mean “Stored in Kafka”

There is a subtle lesson here for application developers. Our Java producer uses the asynchronous KafkaTemplate API:

transactionKafkaTemplate
    .sendDefault(0, UUID.randomUUID().toString(), transactionId)
    .thenAccept(result -> {
        log.info("Successfully sent transactionId: {}", transactionId);
    })
    .exceptionally(ex -> {
        log.error("Error sending transactionId: {}", transactionId, ex);
        return null;
    });

It’s tempting to interpret successful completion at the application level as Kafka has safely stored my record. But with the At-Most-Once producer configuration used in this experiment, that is not the guarantee we’re asking Kafka to provide. This is why understanding configuration semantics matters more than simply seeing a successful producer call in application logs.

Why Replication Didn’t Save T5

This experiment also illustrates an easy Kafka misconception. We configured replication.factor=3 so three brokers can hold replicas of our partitions. Why, then, can a record still disappear? Because replication and producer acknowledgement are related but different concerns. For a record to survive a leader failure, it needs to make it into the replicated state from which the new leader continues. But our producer says acks=0 it doesn’t wait for confirmation of that durability and retries=0 means it won’t try again when delivery fails.

So although the partition itself remains available, a particular in-flight record can still be lost. This is an important distinction cluster availability does not automatically imply zero message loss. Kafka can successfully elect another leader and continue serving the partition while an individual record sent around the failure window never becomes part of the surviving log.

At-Most-Once in One Diagram

The final slide from the experiment summarizes the configuration and failure path: acks=0, retries=0, and enable.idempotence=false, followed by a broker-side failure that prevents the record from reaching the consumer. We can reduce the idea to:

              NO CONFIRMATION
              NO RETRY

+----------+                    +----------------+
| Producer | -----------------> | Kafka Broker   |
+----------+       record       +----------------+
                                       |
                                       X
                                  Broker failure
                                       |
                                       v
                                  RECORD LOST
                                       |
                                       X
                                   Consumer

The consumer cannot process a record that never survived in Kafka.

What Did We Learn?

The experiment demonstrates several important things about Kafka. First, delivery semantics are not simply a property of Kafka itself. They emerge from the way producers, brokers, consumers, and application logic are configured, with acks=0 retries=0 the producer behaves essentially as fire-and-forget. It sends the record without waiting for broker acknowledgement and doesn’t retry failures.

Third, replication does not eliminate every message-loss scenario. Having three replicas is valuable for availability and durability, but the producer configuration still determines what assurance the producer requires before considering its work complete.

Fourth, leader election can succeed while a record is still lost. In our experiment, Kafka elected a new leader for partition 0 and continued operating, but T5 disappeared.

Finally, the central trade-off of At-Most-Once is simple:

Avoid duplicate delivery

        ↕

Accept possible message loss

That trade-off may be acceptable for some workloads. For others — especially payments, financial transactions, orders, inventory changes, and other business-critical events — silently losing a record can be far more damaging than processing one twice.

Kafka gives us powerful tools for building reliable distributed systems, but we still have to decide what guarantees our application actually needs and configure the entire message-processing path accordingly.

Recent Posts