Most Kafka estates start as one cluster for one team, then quietly become shared infrastructure. Two years later there are forty topics, a dozen producing services, an analytics group that runs a full-topic re-read every Tuesday, and no quotas anywhere. The cluster is fine until the afternoon someone deploys a consumer that reads from the earliest offset across every partition at once, saturates the brokers' network, and adds 400 ms of produce latency to a payments path that has nothing to do with them.
That is the noisy-neighbor problem, and it is an operations problem before it is a capacity problem. Here is the posture we recommend for shared clusters.
1. Quotas are the only hard boundary you have
ACLs decide who may touch a topic. They say nothing about how hard. Kafka's client quotas are the mechanism that actually bounds blast radius, and they come in two kinds:
- Network bandwidth quotas —
producer_byte_rateandconsumer_byte_rate, in bytes per second, applied per client. When a client exceeds its rate, the broker does not error; it delays the response, so the client slows down and sees higher latency rather than failures. - Request rate quotas —
request_percentage, expressed as a percentage of one broker request-handler thread's time. This is the one that catches the client that is cheap in bytes but expensive in requests: tiny batches, aggressivefetch.max.wait.ms=0polling, metadata storms from a misconfigured client library.
Quotas apply by user, client-id, or the pair, with a default tier and per-entity overrides. The pattern that survives contact with reality:
# A default ceiling for everyone, so an unannounced client cannot take the cluster
kafka-configs --bootstrap-server broker:9092 --alter \
--add-config 'producer_byte_rate=20971520,consumer_byte_rate=41943040,request_percentage=25' \
--entity-type users --entity-default
# A named override for a tenant with a measured, justified need
kafka-configs --bootstrap-server broker:9092 --alter \
--add-config 'producer_byte_rate=104857600,consumer_byte_rate=209715200' \
--entity-type users --entity-name checkout-service
Two caveats worth saying out loud. Quotas are per broker, not per cluster: a 20 MB/s producer quota against six brokers is a 120 MB/s ceiling in aggregate if the client's traffic spreads evenly. And they are enforced against the authenticated principal, so a cluster where every service shares one credential cannot be quota'd meaningfully — quotas depend on the authentication work being done first.
2. Start in observation mode, not enforcement
Setting quotas from a guess is how you throttle the payment service on a Monday morning. Measure first. The broker exposes per-client byte rates and throttle times:
kafka.server:type=BrokerTopicMetricsper topic for the aggregate picturekafka.server:type={Produce,Fetch},user=...,client-id=...forbyte-rateandthrottle-timeper client
Take two weeks of per-client peaks, set the default tier a comfortable multiple above the p99 of your ordinary clients, and give explicit overrides to the handful of genuinely heavy tenants. Then alert on non-zero throttle-time per client. A throttled client is not necessarily an incident — it is the system doing its job — but it is always a conversation: either the tenant's traffic grew legitimately and the quota should move, or something changed that they did not intend.
3. Namespacing makes governance mechanical
ACLs support prefixed resource patterns, and prefixed ACLs are only usable if topic names are predictable. Pick a convention before the cluster has two hundred topics, not after:
<team>.<domain>.<entity>.<event-or-state>.<version>
billing.invoices.invoice.issued.v1
search.catalog.product.snapshot.v2
The payoff is that a team's grant becomes one prefixed ACL rather than a growing list of per-topic rules, cleanup after a decommission is a prefix scan, and a chargeback report is a GROUP BY on the first segment. The version suffix gives you a place to land a breaking schema change without a rename negotiation across five consumers — the companion to the compatibility rules in Schema Evolution Without Breaking Consumers.
The governance piece that pays for itself fastest is not the naming scheme though — it is a registry of ownership. For every topic: owning team, on-call contact, expected throughput, retention and why, downstream consumers. Keep it in the same repo as the topic definitions if you manage topics as code. When the cluster is at 80% disk at 2 a.m., the question is always "whose topic is this and can it lose a day of retention," and the cost of not knowing is measured in hours.
4. Per-tenant visibility, or you will argue from anecdote
Cluster-level dashboards are nearly useless during a multi-tenant incident. You need the same signals sliced by tenant:
- Bytes in and out, per topic and per principal
- Request rate and request-handler idle percentage, with the top request-rate clients named
- Consumer lag in seconds, per group, with the owning team attached
- Disk footprint per topic, which is where an unreviewed
retention.ms=-1shows up
The operational goal is that "the cluster is slow" resolves to "group X started a re-read at 14:02" within a couple of minutes. Without per-principal attribution, that conversation takes an hour and usually ends in a broker restart that fixes nothing.
5. Protect the cluster from a single topic, too
Quotas bound clients; a few topic-level settings bound the damage a single topic can do:
retention.bytesalongsideretention.ms, so a traffic spike cannot fill a disk that a time-based policy alone would have allowedmax.message.bytesheld at a sane ceiling — the 50 MB payload someone wants to push through Kafka is an object-store reference waiting to happen- A partition-count review on creation; every partition is replication traffic, open file handles, and controller metadata, and partition inflation across many tenants is a real cluster-level limit (see Sizing Kafka Partitions)
min.insync.replicasset deliberately per tenant rather than inherited by accident, because the durability/availability trade-off is a tenant decision
Auto-topic-creation should be off on any cluster with more than one tenant. It turns a client-side typo into a permanent topic with default settings and no owner.
6. When to stop sharing
Multi-tenancy is worth real effort — a shared cluster is cheaper, better monitored, and better operated than five neglected ones. But there are honest reasons to split:
- Incompatible durability or latency profiles. A tenant that needs single-digit-millisecond p99 and a tenant that runs hour-long batch re-reads will fight over page cache regardless of quotas, because quotas bound throughput, not cache eviction.
- Regulatory or residency isolation that cannot be argued down to ACLs and encryption.
- Divergent upgrade cadence. One tenant that cannot tolerate a rolling restart this quarter should not freeze everyone else's upgrade path.
- Blast radius that the business will not accept. If one tenant's outage is a board-level event, it can have its own cluster, and the cost is worth it.
What is usually not a good reason: "team A wants their own cluster." That path ends with six clusters, one of which has an unpatched broker nobody remembers owning.
The short version
A shared Kafka cluster becomes a shared-fate cluster unless three things are true: every client authenticates as itself, every client has a quota, and every topic has an owner. Those three make throttling a routine conversation instead of an incident, and they are all achievable in an afternoon of configuration plus a few weeks of measurement.
If you are running one cluster for several teams and you are not sure what the current per-tenant picture looks like, that measurement pass is exactly what a Kafka health check produces — current usage per principal, a proposed quota tier, and the topics that have no owner. Get in touch with a description of the cluster and the tenants on it, and we will tell you what we would look at first.