Skip to main content

Health Checks

Install Dekaf.Extensions.HealthChecks to register Kafka checks with ASP.NET Core's health-check services.

dotnet add package Dekaf.Extensions.HealthChecks

Producer flush checkpoints and broker connectivity​

These checks answer different questions:

SignalHow to observe itWhat success establishes
Producer flush checkpointAddDekafProducerHealthCheck<TKey, TValue>()The current FlushAsync completed within the configured timeout.
Individual delivery outcomeAwait ProduceAsync, or inspect the delivery callback's errorThat delivery succeeded or failed under the configured acknowledgement policy.
Recent delivery failuresAggregate delivery outcomes or producer error metrics over an application-defined time windowWhether the workload meets the application's delivery policy during that window.
Broker connectivityAddDekafBrokerHealthCheck()An active admin request reached the cluster and returned at least one broker.

The producer check reports Healthy when its flush checkpoint completes, even when a queued batch failed delivery. Failed batches leave the producer pipeline too. Concurrent production can leave newer messages queued after that checkpoint; Healthy does not assert that the current queue is empty. An idle producer can also report Healthy while every broker is unavailable, because an empty queue needs no broker request. The result explicitly states that delivery outcomes and broker connectivity are not checked.

Register the producer and broker checks separately when both flush completion and connectivity matter. Their producer and admin clients must already be registered in DI; see Dependency Injection.

using Dekaf.Extensions.HealthChecks;

builder.Services.AddHealthChecks()
.AddDekafProducerHealthCheck<string, string>(
name: "kafka-producer-queue",
options: new DekafProducerHealthCheckOptions
{
Timeout = TimeSpan.FromSeconds(5)
})
.AddDekafBrokerHealthCheck(name: "kafka-broker-connectivity");

Broker reachability does not prove that a particular topic accepts writes, that the producer's credentials permit them, or that earlier messages succeeded. Delivery outcomes remain a separate signal. FireAsync has no delivery result; use a result-returning or callback-based produce API when application health depends on individual delivery outcomes.

Failure and recovery​

The producer check returns Unhealthy if its flush throws or exceeds the timeout. Each invocation evaluates a new flush: the next completed flush returns Healthy. The check retains no delivery-history latch and does not turn past produce failures into a permanent unhealthy state.

If recent delivery failures are part of an application's readiness policy, choose a bounded observation window and an explicit recovery condition, such as failures aging out of that window. The built-in producer check neither installs that policy nor resets delivery counters. Keep its status separate from the application's delivery-failure status.

Migration note​

Earlier descriptions claimed that a successful producer health check proved connectivity or successful delivery. Those claims were incorrect; the underlying check waited for a flush checkpoint. The corrected description states that scope explicitly. Existing Healthy/Unhealthy status behavior and registration signatures remain compatible. Applications that used this check as a connectivity probe should also register the broker check; applications that need delivery assurance must observe delivery results.

Consumer checks​

AddDekafConsumerHealthCheck<TKey, TValue>() evaluates consumer liveness and lag using its configured thresholds. A live group member with no assignment can be a healthy standby. See Consumer Groups and Observability for the related signals.

Share consumer checks​

AddDekafShareConsumerHealthCheck<TKey, TValue>() reports Healthy while a share consumer is a stable member of its share group and its heartbeat is no older than three broker-directed heartbeat intervals. It reads a local status snapshot; it never polls, acquires, or acknowledges records. A share group can have more members than partitions, so an idle member is healthy.

To check the consumer owned by a hosted share consumer service, pass its KafkaShareConsumerServiceKey from Dekaf.Extensions.Hosting. The key identifies that worker's own consumer even when other services share its message types and public service key:

using Dekaf.Extensions.HealthChecks;
using Dekaf.Extensions.Hosting;

builder.Services.AddHealthChecks()
.AddDekafShareConsumerHealthCheck<string, string>(
KafkaShareConsumerServiceKey.For<ShareOrderWorker>("worker-a"),
"order-workers");

Keyed clients​

Each check has an overload that takes the service key of a keyed client, followed by a required check name:

using Dekaf.Extensions.HealthChecks;

builder.Services.AddHealthChecks()
.AddDekafProducerHealthCheck<string, string>("orders", "orders-producer")
.AddDekafConsumerHealthCheck<string, string>("orders", "orders-consumer")
.AddDekafBrokerHealthCheck("orders", "orders-broker")
.AddDekafShareConsumerHealthCheck<string, string>("orders", "orders-share-consumer");