Skip to main content

Connection recovery

RespireOptions.ReconnectPolicy configures replacement of failed multiplexed command connections, including connections owned by Redis Cluster nodes, dedicated connection acquisitions, pub/sub recovery, and Cluster/Sentinel discovery fallback. It is null by default. For command connections, replacement starts immediately, and a failed replacement is retried on the next use. Setting a policy preserves that demand-driven scheduling while adding backoff and limits. It does not create a perpetual command retry loop or replay any accepted Redis command.

var options = new RespireOptions
{
Endpoints = { new RespireEndpoint("localhost") },
ReconnectPolicy = new RespireReconnectPolicy
{
InitialDelay = TimeSpan.FromMilliseconds(250),
BackoffMultiplier = 2,
MaxDelay = TimeSpan.FromSeconds(5),
JitterRatio = 0.2,
MaxAttempts = 10,
},
};

The first replacement waits InitialDelay. Later attempts multiply that delay by BackoffMultiplier for each failed attempt, capped at MaxDelay. Symmetric jitter scales the capped delay by a random factor from 1 - JitterRatio through 1 + JitterRatio; the actual delay is capped again. Zero initial delay stays zero. Maximum delay cannot exceed one day, the multiplier must be finite and at least one, and the jitter ratio must be finite and between zero and one. Invalid policies fail before connection I/O. Policies are immutable records and can be shared across clients.

Attempts count separately for each slot in a multiplexer generation. Concurrent callers share an in-progress replacement; they do not each create another attempt. A successful replacement, including its handshake and required connection identity, resets the count. The counter tracks consecutive replacement failures, not connections that fail after a successful handshake. A peer that repeatedly accepts and then drops connections therefore starts a new episode each time, with the initial delay. There is no healthy duration window or first-command condition for reset. MaxAttempts = null permits unlimited attempts. Exhaustion remains in effect for that failed slot; repeatedly submitting commands does not reset it. Recreate the client to begin a new episode after exhaustion. A newly discovered Cluster generation has its own slots and counters. Healthy slots remain usable when another slot is exhausted.

When every slot at an endpoint has exhausted recovery, acquisition throws RespireReconnectLimitException, a RespireConnectionException. Individual failure events retain the original connection error. Identity setup requires every slot, so it throws when any required slot is exhausted even if another slot can still serve ordinary commands. Failed correction fences retain their server-side identity obligations; a terminal recovery error does not mark an unacknowledged fence complete. Fences connect directly to the captured physical peer with their own ConnectTimeout; they do not acquire or revive an exhausted command slot. A fence may therefore succeed after command recovery is exhausted, or report its own connection/permission error while retaining the owed identity. The command reconnect limit does not cap attempts to discharge that separate correctness obligation. A command that notices a failed connection can still fail while replacement proceeds, as it did before configuring this policy. Each connection attempt retains ConnectTimeout; initial multiplexer setup is not retried by this policy. Command deadlines and caller cancellation still bound their own waits. Cancelling one caller does not cancel recovery shared by other callers. Retiring or disposing the multiplexer cancels its scheduled delays and awaits owned recovery work.

Commands accepted by a failed connection are never replayed. A missing reply can leave the server outcome uncertain. This policy only creates replacement connections; application retry decisions still need to account for command idempotency.

Dedicated connection acquisitions​

Dedicated pools serve blocking operations, watched transactions, and other operations that require an exclusive connection. A healthy idle connection is reused immediately. When no healthy idle connection exists, each rent makes one immediate connection attempt. With a null policy, its failure is returned immediately, preserving the existing on-demand behavior. With a policy, a failed attempt starts a retry episode owned by that rent. The first retry waits InitialDelay; subsequent retries use the same backoff and jitter calculation as command replacements. Only connection establishment and its handshake are retried. No accepted application command, transaction, or blocking operation is replayed.

MaxAttempts counts retries after the initial attempt: a limit of two allows at most three connection attempts for one rent. Concurrent renters have independent budgets and can connect concurrently; the policy is not an endpoint-wide rate limiter. Success ends the episode. Exhaustion throws RespireReconnectLimitException with the last connection failure as its inner exception. It does not disable the pool: a later rent starts fresh. An unlimited policy can keep one rent pending until it connects or is cancelled. That rent retains its acquisition reservation through backoff and handshake cleanup. The pool limits retained idle connections, not concurrent acquisitions; callers should bound total waiting work with cancellation and their own concurrency limits. Connection failures, connection/handshake timeouts, and server errors classified by RespireServerException.IsTransient can retry. Permanent server handshake rejections (including WRONGPASS, NOAUTH, NOPERM, and ordinary ERR), TLS authentication failures, and explicit RespireProtocolException/RespireConfigurationException failures return immediately. A permanent failure encountered after a retry starts ends that episode without marking its attempt budget exhausted.

Connect timeouts and handshake command deadlines remain per attempt. Use caller cancellation to bound the entire rent, including all delays and handshakes. Cancelling one rent does not cancel another. Retirement and disposal cancel pending rents and wait for their owned connection cleanup. Retirement still lets borrowed operations finish; disposal aborts them. Corrective fencing acquisitions retain their own deadline and retry rules and bypass this policy, including its attempt cap.

Pub/sub reconnection and resubscription​

In Cluster mode, sharded subscriptions have their own recovery episode, separate from the regular channel/pattern connection. An attempt restores all affected sharded routes; healthy primaries keep delivering. Sharded exhaustion ends all sharded subscriptions and rejects new sharded subscriptions, while regular subscriptions remain usable. Regular pub/sub exhaustion likewise leaves sharded subscriptions usable.

With a null policy, pub/sub retains its existing schedule: an immediate replacement attempt followed, on failure, by 250 ms exponential waits capped at five seconds. With a policy, the first replacement waits InitialDelay, and every failed connection or resubscription consumes one attempt from the same episode. The shared policy applies its multiplier, jitter, and maximum delay to each scheduled attempt. Initial subscription connection setup is still a single attempt governed by its caller's cancellation.

One recovery loop owns the replacement until every currently registered route has been acknowledged. A socket that connects but fails while resubscribing does not reset the budget or start another watcher. A failed resubscription closes its replacement before the next attempt. Successful resubscription resets the count; a later interruption starts a new episode. Existing gap markers remain ordered before messages received after each route's acknowledgement. Redis cannot replay messages missed during an interruption.

While configured recovery is active, new explicit subscriptions fail with RespireConnectionException; they do not open a competing connection or skip the delay. Retry after a lifecycle event with ReconnectSource = PubSub and SourceState = Connected. Removing existing subscriptions is supported during backoff, and recovery snapshots the remaining routes under the control gate. A caller cancellation during an explicit subscription does not cancel shared recovery.

When the configured limit is exhausted, the client publishes Disconnected with ReconnectExhausted = true and the final connection/resubscription error. All live subscriptions complete with RespireSubscriptionEndReason.ReconnectExhausted and become disposed. Existing enumerators may drain already buffered items before ending; a new enumerator cannot be opened on an ended subscription. New subscriptions throw RespireReconnectLimitException. There is no cooldown or implicit reset: recreate the client to subscribe again. Command connections remain usable subject to their own health.

Disposal cancels configured delays, connection handshakes, and resubscription waits, then drains owned recovery and socket cleanup. Ordered lifecycle delivery happens outside the recovery task so a lifecycle handler can synchronously dispose the client. Keep event handlers short; attempt observations can lag the work they describe. Disposal suppresses queued lifecycle callbacks, but queued measurements still drain because they describe already scheduled attempts or completed exhaustion. They may arrive after disposal.

Lifecycle and telemetry​

ConnectionStateChanged retains its existing endpoint health aggregation and adds:

PropertyMeaning
ReconnectAttemptOne-based configured attempt; zero for transitions without policy metadata
ConnectionSlotSource slot within this endpoint's multiplexer generation, including null-policy recovery; null for endpoint-wide transitions
NextReconnectDelayActual scheduled delay before this attempt; null when no attempt is scheduled
ReconnectExhaustedThis slot, dedicated rent, pub/sub episode, or Cluster discovery round reached its configured limit
ReconnectSourceDedicated for a dedicated rent, Command for a multiplexer event, PubSub for subscription recovery, ClusterDiscovery for topology/endpoint fallback; Unspecified for other paths
SourceStateSource connection state before endpoint health aggregation, when supplied
ReconnectEpisodeIdProcess-local identifier grouping one dedicated rent or Cluster discovery round; null for other paths

Policy attempt events are delivered even if aggregate endpoint health has not changed. An endpoint can remain disconnected because another slot has failed while the source slot starts recovery. Use the metadata for the source attempt, not as an endpoint-wide counter. Do not block event handlers. Synchronous disposal remains supported.

Dedicated recovery does not mark a command slot unhealthy. Its event State remains the aggregate endpoint health, while SourceState describes the dedicated attempt. Dedicated events and measurements are dispatched asynchronously in pool order so observers can dispose the client without blocking the acquisition disposal must drain. Events can lag the corresponding attempt; they are observations, not a mechanism for gating retries. After client disposal, pending dedicated lifecycle events are suppressed. Episode IDs are event metadata only and are not metric tags. Caller cancellation and pool retirement/disposal stop a rent without emitting a dedicated Disconnected failure. Scheduled-attempt measurements may still be delivered after pool shutdown. Queued lifecycle callbacks are skipped once the dispatcher observes that the pool is stopping; callbacks already in flight are not joined because they can dispose it. The ordered observer queue has no capacity limit. Slow or blocked callbacks can accumulate pending observations across concurrent renters; keep handlers short and hand off work to an application queue with an explicit capacity policy. A finite retry limit bounds each rent's retries, not total concurrent rents or the observer queue.

The Respire meter records respire.connection.reconnect.attempt (attempt number) and respire.connection.reconnect.delay (seconds), tagged with server.address, server.port, and respire.connection.source (command, dedicated, or pubsub). Pub/sub attempts have a null ConnectionSlot and use one shared counter for the current connection/resubscription episode. They record scheduled replacement attempts, including waits later cancelled by disposal. They are histograms of events, not live countdown gauges. The counter respire.connection.reconnect.exhausted records each episode stopped by its attempt limit, with the same endpoint tags, so operators can alert on terminal recovery failures. Successful command execution does not record these instruments or inspect policy counters.

Sentinel discovery fallback​

Sentinel resolution keeps its first candidate immediate. With a configured policy, each later configured or learned Sentinel candidate consumes one fallback attempt and waits the policy delay before I/O. MaxAttempts = 1 therefore permits the first candidate plus one fallback. A rejected primary ROLE response consumes that same candidate attempt; peer expansion and primary connection do not start nested budgets. For example, with configured candidates A and B and peers C and D learned from A, MaxAttempts = 1 allows A then B. If both fail, C and D remain untried and the policy exhaustion counter increments. With only A and B available, both still run, but their failure ends through candidate depletion rather than policy exhaustion. The existing finite traversal still stops when candidates run out, even with a null MaxAttempts. Endpoints are not cycled indefinitely.

A successful validated primary ends the resolution. Exhaustion preserves the final RespireConnectionException and its underlying discovery or primary-validation error. A new explicit ConnectAsync call or a new reactive resolution starts a fresh budget. Null policy retains immediate fallback across all available candidates. Caller cancellation bounds every wait, while each candidate retains its existing discovery and connection timeout. Policy delays sit outside those per-candidate timeouts. Supply a caller deadline to bound the total resolution, including backoff and every candidate. No application command is replayed during discovery.

The same attempt/delay/exhaustion instruments carry the candidate's server.address and server.port plus respire.reconnect.scope = sentinel-discovery. Exhaustion is attributed to the final failed fallback. Scheduling is recorded before delay, including waits later cancelled by the caller. Exhaustion is recorded only when the failed fallback uses the budget and at least one candidate remains untried. Successful fallback and running out of candidates do not record exhaustion, even when the final candidate coincides with the attempt limit. Observe the returned connection exception to detect every terminal discovery failure. Lazy Sentinel clients can observe initial connection events through ConnectionStateChanged. Disconnects and READONLY replies retire the current generation; the next operation resolves Sentinel again with the same fallback policy. Validated endpoint changes also increment respire.sentinel.failover. Eager ConnectAsync performs initial discovery before returning the client. No accepted command or WATCH state is replayed during a primary handoff.

Sentinel event monitors apply the same policy to their own subscription connections. MaxAttempts counts replacement attempts, so a failed initial subscription still gets a retry. Scheduled retries and exhaustion use the same instruments with respire.reconnect.scope = sentinel-monitor. An exhausted monitor stops observing that Sentinel's events and logs a warning; discovery still runs on demand when commands observe a disconnect or READONLY. A validated primary publication gives each exhausted monitor a fresh budget, so one long Sentinel outage does not remove event monitoring for the life of the client. A publication that lands while the monitor is still spending its retries counts too: the monitor then resumes as soon as it exhausts, rather than waiting for a second publication.

Cluster discovery fallback​

Each logical Cluster discovery round shares one policy budget across cached owners, known masters, topology queries, configured seeds, tracked connection selection, and dedicated pool selection. Pub/sub endpoint discovery and cluster-wide commands use the same contract. The first candidate is immediate. A discovery episode begins only when a fallback is scheduled; a failure with no scheduled fallback remains visible through physical connection health. After a candidate fails, choosing the next candidate consumes one attempt and waits the configured delay. A topology query and connection to its advertised owner share the candidate until one fails. Connecting other required masters after a successful topology query does not consume fallback attempts.

For example, with MaxAttempts = 1, a failed cached owner may try one known master. If that master also fails, seed fallback cannot start a new budget. Exhaustion throws RespireReconnectLimitException; candidate depletion can instead end with the existing connection error. A later explicit discovery round starts fresh. Concurrent initial connection calls retain seed discovery coalescing: waiting callers reuse its successful result. Failed rounds do not permanently disable the router.

MOVED, ASK, and READONLY are explicit server rejections. Their first replacement consumes attempt one. If recovery cannot establish a replacement, the original server exception is preserved. ASK keeps its temporary target; tracked identity setup and retirement reselection share the ongoing round. A command rejected before acceptance by topology retirement starts fallback at attempt one. Repeated retirement during that command, including an individual rejected batch entry or an unwatched transaction, retains the same budget until completion. A batch READONLY replacement and any later retirement use that same round. Blocking redirects retain it through the next dedicated connection rent, including retirement of the selected pool before the rent completes. Resumable Cluster scans keep that budget across retirement restarts of one page; rebuilding from the immutable cursor does not reset it. Accepted batch entries and watched transactions are never replayed. Dedicated physical connection retries retain their separate per-rent budget; node selection does not reset a discovery budget. Neither path replays a command after ambiguous acceptance or a lost reply.

Null policy preserves existing immediate fallback and retirement limits. Unsupported or ACL-restricted CLUSTER SLOTS replies still allow an initial seed connection and learned redirect routing. A cluster-wide command still requires a complete topology. Configured policy limits fallback during that search; it does not turn unsupported topology into successful cluster-wide results.

Caller cancellation and client disposal stop scheduled waits and connection work. Existing command deadlines remain shared. READONLY recovery retains its overall ConnectTimeout and the time reserved for configured seeds; policy waits consume that recovery time. Whichever is reached first ends recovery: the policy attempt limit, the existing recovery deadline, caller cancellation, or disposal. A delay cannot extend the READONLY deadline. Other discovery waits are bounded by caller cancellation, so supply a caller deadline to bound a whole round across multiple candidates and delays.

Attempt/delay/exhaustion instruments use the candidate endpoint and respire.reconnect.scope = cluster-discovery. Exhaustion is attributed to the final failed candidate, only when another fallback would exceed the budget. Successful fallback at the limit and candidate depletion do not count as policy exhaustion. Lifecycle events use ReconnectSource.ClusterDiscovery, a process-local ReconnectEpisodeId, and a null ConnectionSlot. Episode IDs are scoped by ReconnectSource; include it when correlating events from different recovery paths. Cluster discovery IDs share one counter across clients in this process. SourceState describes discovery; State retains aggregate physical endpoint health. Discovery never inserts a synthetic failed command slot. Measurements and lifecycle callbacks run on an ordered asynchronous queue; keep observers short. Discovery ordering is independent of command-slot and dedicated recovery notification queues; events from different sources have no shared ordering guarantee and can arrive after the triggering operation completes. Configured-seed fallback skips multiplexer generations already rejected by owner or known-master discovery in that round. A replacement generation at the same address remains eligible; skipping a duplicate does not consume another policy attempt. A caller-cancelled round that started fallback emits a terminal Disconnected source event with the cancellation error and the same episode ID; cancellation is not policy exhaustion. A READONLY phase timeout is retryable while the overall recovery deadline remains active. If a later configured seed succeeds, that episode ends with Connected, without retaining the handled phase cancellation as a terminal error. Once replacement selection succeeds, later application errors (such as WRONGTYPE) or cancellation of an accepted command do not change that discovery outcome to Disconnected. The command still fails normally; physical connection health is reported independently. Scheduled measurements survive disposal, while pending lifecycle events are suppressed. The queue preserves every scheduled measurement and source transition and has no capacity limit, matching dedicated recovery notifications. Slow or blocked observers can accumulate notifications from concurrent or repeated rounds. MaxAttempts bounds one discovery round, not the notification queue; handlers should hand off expensive work and return promptly.

Periodic Cluster refresh​

The policy covers command multiplexers, dedicated pools, pub/sub, Cluster discovery, Sentinel fallback, and Sentinel event monitors as described above. Background Cluster refresh uses the same discovery budget for candidate selection and backoff; it does not add a separate nested retry counter.

ClusterTopologyRefreshInterval controls the periodic timer and defaults to 60 seconds. Set it to null, TimeSpan.Zero, or Timeout.InfiniteTimeSpan to disable that timer. Redirects, primary disconnects, and failed-refresh retries can still trigger discovery. Each refresh pass has a 60-second deadline and retains the last published slot map if discovery fails. See Cluster topology refresh for scheduling, jitter, coalescing, and candidate ordering.