Most teams we meet can state their replication factor from memory and assume it answers the durability question. It does not. Replication factor sets how many copies could exist; acks, min.insync.replicas, and your unclean-leader-election policy decide how many copies actually existed at the moment the producer was told "written." Those three settings live in three different places — producer config, topic config, broker config — and are usually owned by three different people. That is why the gap survives until a broker dies.
Here is how the pieces fit, and what each one costs.
The ack path, precisely
A produce request goes to the partition leader. The leader appends to its local log, followers fetch, and the leader tracks which replicas are caught up — the in-sync replica set (ISR). A replica drops out of the ISR when it has not fetched up to the leader's log end offset within replica.lag.time.max.ms (default 30s).
acks decides when the producer gets its response:
acks=0— the client does not wait at all. A record can be lost in the socket buffer. Legitimate only for data you would be willing to drop silently.acks=1— the leader responds after its own append. If that broker's disk or process dies before followers fetch, the record is gone and the producer already moved on. This is the setting behind most "we lost a few seconds of events" postmortems.acks=all— the leader responds once every replica currently in the ISR has the record.
Note the phrase: currently in the ISR. If two of three replicas have fallen behind and dropped out, the ISR is just the leader, and acks=all degrades silently to acks=1. Nothing errors. That is the trap min.insync.replicas exists to close.
min.insync.replicas is the other half of acks=all
min.insync.replicas is a topic-level (or broker-default) setting that says: if the ISR is smaller than this, reject acks=all writes with NotEnoughReplicasException rather than accepting them weakly. It has no effect at acks=0 or acks=1 — the two settings only mean something together.
The standard production shape is RF=3, min.insync.replicas=2, acks=all. That tolerates one replica being down or lagging and still accepts writes, while guaranteeing every acknowledged record is on at least two brokers.
Two configurations that look reasonable and are not:
- RF=3, min.insync=3. Now a single broker restart — an upgrade, a node replacement, a routine rolling bounce — stops writes on every partition it hosts. You have bought a marginal durability increase and paid for it with an availability incident every maintenance window.
- RF=2, min.insync=2. Same problem, worse: any one broker down takes the topic offline for producers. If you cannot afford RF=3, you cannot afford min.insync=2.
Also check default.replication.factor and min.insync.replicas at the broker level, because auto-created topics inherit them. We have found plenty of RF=1 topics in otherwise careful clusters, created implicitly by a Connect worker or a test harness and then quietly promoted to carrying real traffic.
Unclean leader election: the explicit data-loss switch
When every in-sync replica for a partition is unavailable, Kafka has two options: wait for one of them to come back (the partition is offline, producers and consumers for it block), or elect an out-of-sync replica as leader and resume (the partition is available, and every record that replica was missing is permanently gone — including records the producer was told were committed).
unclean.leader.election.enable is false by default in modern Kafka, and for anything that resembles a ledger it should stay there. But it is a genuine choice, not an obvious one: for a telemetry topic where a stale-but-live stream beats a stalled one, enabling it per-topic is a defensible decision. Make it explicitly, per topic, and write it down — the bad version is discovering during an incident that someone enabled it cluster-wide in 2021 to clear an outage and never turned it back off.
What the producer does after the ack path
Durability settings only cover records the broker was actually given. Two client-side failure modes sit upstream:
Buffer loss. The producer batches in memory (linger.ms, batch.size). A process killed with records in the accumulator loses them, acks setting irrelevant. If the data matters, call close() with a timeout on shutdown (and handle SIGTERM so it runs), and do not let an async send() without a callback be your error handling — that is where silent drops hide.
Retry exhaustion. delivery.timeout.ms (default 2 minutes) bounds total time including retries. When it expires you get a TimeoutException, and what happens next is application code. If your callback logs a warning and returns, you have built at-most-once delivery by accident. Decide: block, buffer to disk, or fail the upstream request.
Idempotence (enable.idempotence=true, default since 3.0) is worth confirming rather than assuming: it prevents duplicates from broker-side retries and preserves per-partition ordering, and it is silently incompatible with acks=1 configurations some teams still carry forward in old property files.
A 20-minute audit
For each topic that carries data you would have to explain losing:
- Replication factor,
min.insync.replicas, and theacksof every producer writing to it — gathered in one table. The mismatches are visible immediately. unclean.leader.election.enableat broker and topic level, with a stated reason for everytrue.- Replica placement across racks or AZs (
broker.rackset, and actually populated). RF=3 inside one availability zone is one correlated failure from zero copies. - An alert on
UnderMinIsrPartitionCount— not justUnderReplicatedPartitions. The first one means you are writing at reduced durability right now; the second is often routine. enable.idempotenceand shutdown handling in each producer application.- Consumer-side commit order: committing offsets before processing completes reintroduces loss after all this broker-side care.
Most of what we find in health checks is not an exotic bug. It is acks=1 in a service written three years ago by someone who has since left, writing to a topic that the business now treats as a system of record. The settings are one deploy away from correct; noticing is the hard part.
If you want a second set of eyes on how your clusters and producers are configured, a Kafka health check covers exactly this ground and ranks what it finds by production risk.