Skip to main content

Retry

Re-execute the delegate when it fails, waiting between attempts.

See the exceptions reference for failures the default retry clause handles.

Quick forms​

var exponential = Shield.Retry(3); // equal jitter, 250ms base, 30s cap
var constant = Shield.Retry(3, Backoff.Constant(TimeSpan.FromSeconds(1)));
var linear = Shield.Retry(3, Backoff.Linear(TimeSpan.FromMilliseconds(500)));
var forever = Shield.RetryForever(
Backoff.Exponential(TimeSpan.FromSeconds(1), maxDelay: TimeSpan.FromMinutes(1)));

The count is retries, not attempts: Retry(3) makes up to 4 total attempts — the initial call plus 3 retries. The same reading applies to MaxRetries in the options form, and to every Retry overload on Shield, Shield<T>, and the When… builders.

This executable check removes delays and verifies the count:

var attempts = 0;
var retryWithoutDelay = Shield.Retry(3, Backoff.None);

try
{
await retryWithoutDelay.ExecuteAsync(_ =>
{
attempts++;
return ValueTask.FromException(new HttpRequestException("offline"));
});
}
catch (HttpRequestException)
{
}

if (attempts != 4)
{
throw new InvalidOperationException($"Expected 4 attempts, observed {attempts}.");
}

Defaults​

Bare Retry(n) and RetryForever() use Backoff.Default:

SettingDefault
CurveExponential
Base delay250 ms
Factor2
JitterEqual: each delay is scaled by a value in [0.5, 1.5)
Maximum delay30 seconds

The cap also applies to RetryForever(). Equal jitter prevents callers from retrying in lockstep.

Backoff​

Backoff factories describe the delay sequence:

Backoff.Constant(TimeSpan.FromSeconds(1)); // 1s, 1s, 1s, ...
Backoff.Linear(TimeSpan.FromMilliseconds(500)); // 500ms, 1s, 1.5s, ...
Backoff.Exponential(TimeSpan.FromSeconds(1)); // ~1s, ~2s, ~4s, ... (jittered)
Backoff.Exponential(TimeSpan.FromSeconds(1), jitter: Jitter.Full);
Backoff.Exponential(TimeSpan.FromSeconds(1), jitter: Jitter.Decorrelated);
Backoff.Custom(attempt => TimeSpan.FromMilliseconds(100 * attempt)); // attempt is 1-based
_ = Backoff.None; // no delay between attempts
_ = Backoff.Default; // what bare Retry(n) uses
  • Jitter.None uses the exact curve; Equal scales it by [0.5, 1.5); Full selects [0, curve); Decorrelated selects [base delay, previous delay × 3]. Decorrelated state is isolated per retry execution. Direct backoff consumers can carry the effective preceding delay through GetDelay(attempt, previousDelay); pass zero for the first draw.
  • Exponential(baseDelay, factor = 2.0, maxDelay = null, jitter = Jitter.Equal) is the default curve. Backoff.Default = Exponential(250ms, maxDelay: 30s).
  • Constant(delay, jitter = Jitter.None) and Linear(step, maxDelay = null, jitter = Jitter.None) also accept jitter explicitly.
  • Configuration accepts enum names and Boolean aliases (true = Equal, false = None).
  • Without maxDelay, built-in and custom backoffs clamp only at the runtime timer limit (uint.MaxValue - 1 milliseconds, roughly 49.7 days). Use maxDelay for an application-specific cap.

Full options​

API reference: RetryOptions and RetryOptions<T>.

var retry = Shield.Retry(o =>
{
o.MaxRetries = 5;
o.Backoff = Backoff.Custom(attempt => TimeSpan.FromMilliseconds(100 * attempt));
o.MaxDelay = TimeSpan.FromSeconds(10);
o.OnRetry = e =>
{
logger.LogWarning(e.Exception, "Retry {AttemptNumber} after {Delay}", e.AttemptNumber, e.Delay);
return default;
};
o.DelayGenerator = e => new(e.AttemptNumber == 0 ? TimeSpan.Zero : null);
// ^ return a TimeSpan to override the computed delay, or null to keep it
});
OptionDefaultWhat it does
MaxRetries3Retries after the initial attempt — 3 means up to 4 total attempts; int.MaxValue = forever
BackoffBackoff.DefaultThe delay sequence (see above)
MaxDelaybackoff capAbsolute cap applied to every delay, including DelayGenerator output. When unset, the backoff's cap applies (30s for Backoff.Default); a backoff without a cap falls back to the runtime timer limit
OnRetry—Awaited before each retry sleeps — attempt number, final delay, failure. Return default when the work is synchronous
DelayGenerator—Per-retry override returning ValueTask<TimeSpan?>: a TimeSpan replaces the computed delay, null keeps it. This is how Retry-After support works
HandlesException—Local exception predicate; replaces the ambient clause for this retry
HandlesResult (RetryOptions<T>)—Local result predicate; replaces the ambient clause together with HandlesException

Invalid option values throw KevlarConfigurationException and identify the options type, property, and offending value.

Order per retry: backoff computes the delay → effective cap (MaxDelay ?? Backoff.MaxDelay) clamps it → the awaited DelayGenerator may override it → the awaited OnRetry sees the final delay and handled outcome → a superseded disposable result is disposed → sleep → retry metrics are recorded immediately before the next attempt. Concurrent hedge compositions defer disposal until a replacement outcome exists, so execution-wide suppression can still return the original outcome without returning a disposed result. The generator's null and negative results are ignored, and the effective cap clamps its override. Calling SuppressAdditionalAttempts() from either typed or untyped event stops before retry metrics, disposal, sleep, or another attempt, including attempts in nested child shields.

If cancellation arrives during that sleep, the next attempt never starts. Cancellation surfaces as an OperationCanceledException with no InnerException; the failure that triggered the retry is not retained. OnRetry has already run, but kevlar.retries is not recorded because no next attempt started.

When ordinary result handling triggers another attempt, Kevlar disposes the handled result before the next attempt starts. Concurrent hedge compositions dispose it after its replacement completes. Kevlar prefers IAsyncDisposable.DisposeAsync() when a result implements both disposal interfaces. OnRetry runs first so it can inspect the live result. The final result returned to the caller is never disposed by the retry strategy. Disposal failures are reported through KevlarDiagnostics.OnCallbackError as CallbackErrorKind.ResultDisposal and do not replace the pipeline outcome.

Both hooks return ValueTask. A hook that completes synchronously (return default;, new(value)) costs nothing extra and works with synchronous Execute. A hook that yields is awaited by ExecuteAsync; reached through synchronous Execute, it throws NotSupportedException at that call. See synchronous execution compatibility. Notification-hook exceptions follow the shared callback-failure contract: they are reported and never replace the protected outcome.

RetryOptions and RetryOptions<T> are standalone sibling types. Both expose the same MaxRetries, Backoff, and MaxDelay settings, while their callback properties use distinct RetryEvent and RetryEvent<T> delegates. Configure shared scalar defaults in each options lambda; a RetryOptions<T> instance is not assignable to RetryOptions.

Generators receive the caller token through e.Context.CancellationToken. If cancellation arrives while a generator is awaiting, notification hooks still run after it completes, then the next attempt is suppressed and caller cancellation surfaces. A generator exception surfaces with its original identity and skips later hooks. RetryEvent.Context is pooled execution state: use it only before the returned ValueTask completes; never retain it or its property bag.

On an untyped Shield, the RetryEvent callbacks receive: AttemptNumber (zero-based, so 0 is the first retry after the initial execution), Delay, Exception (null when a handled result triggered the retry), Result (the handled result, boxed as object?) and Context. On a typed Shield<T>, the events are RetryEvent<T> instead: same AttemptNumber/Delay/Context, plus the handled failure as a directly stored typed Outcome<T> — e.Outcome.Result is your T, with no boxing, reconstruction, or cast.

var attemptNumbers = new List<int>();
var numberedRetry = Shield.Retry(options =>
{
options.MaxRetries = 3;
options.Backoff = Backoff.None;
options.OnRetry = retry =>
{
attemptNumbers.Add(retry.AttemptNumber);
return default; // completes synchronously, so synchronous Execute below is fine
};
});

try
{
numberedRetry.Execute(static _ => throw new InvalidOperationException());
}
catch (InvalidOperationException)
{
}

Console.WriteLine(string.Join(",", attemptNumbers)); // 0,1,2

What gets retried​

Whatever the current handling clause says is a failure—ordinary exceptions under the default, or the outcomes selected by your When/WhenResult clause:

var retry = Shield
.When<HttpRequestException>()
.Or<TimeoutExceededException>()
.Retry(5);

The options-lambda form can instead set HandlesException and, on Shield<T>, HandlesResult. Setting either creates a per-strategy override; unspecified outcome kinds are not handled.

Placement in the chain​

var scopedRetry = Shield
.Timeout(TimeSpan.FromSeconds(30)) // total budget: retries must fit inside
.Retry(3)
.Timeout(TimeSpan.FromSeconds(5)); // each attempt gets 5s, and Retry sees the TimeoutExceededException

Retry outside a per-attempt timeout retries timeouts; retry inside a circuit breaker hammers a struggling dependency before the breaker sees the pattern. The composition rules cover this in depth.

Respecting an outer deadline​

Set RespectDeadline = true to stop retrying when the next delay is greater than or equal to KevlarContext.Deadline's remaining budget. The effective delay includes DelayGenerator and MaxDelay. Kevlar returns the last handled result or exception immediately, without disposing that returned result. It emits retry.skipped_deadline and increments kevlar.retries.skipped with reason=deadline; deadline refusals do not increment kevlar.retries.

var deadlineAware = Shield.Timeout(TimeSpan.FromSeconds(1)).Retry(options =>
{
options.MaxRetries = 3;
options.Backoff = Backoff.Constant(TimeSpan.FromMilliseconds(600), jitter: Jitter.None);
options.RespectDeadline = true;
});

If attempts fail immediately, this pipeline performs the initial attempt and one retry at 600 ms. The next 600 ms delay cannot fit, so the second failure surfaces instead of waiting for the outer timeout. A viable delay does not guarantee that the following operation will finish before the deadline; cancellation still applies during that operation.

Both typed and untyped options default to false in 1.x, preserving existing timeout and cancellation behavior. With no enclosing timeout, the option has no effect. The budget is checked before OnRetry and again afterward because an asynchronous callback can consume time. A skipped retry normally does not invoke OnRetry; it can already have run when its duration causes the second check to skip. Caller cancellation retains priority. Describe() includes deadline-aware when enabled. DI configuration also accepts Retry.RespectDeadline.

Shared retry budgets​

Use one RetryBudget for a downstream dependency when failures across many callers should suppress additional traffic. The same instance can be assigned to retry and hedge options on independent shields, including shields created for different partition keys:

var budget = new RetryBudget(maxTokens: 100, tokenRatio: 0.1);
var retry = Shield.Retry(options =>
{
options.MaxRetries = 3;
options.Backoff = Backoff.Exponential(TimeSpan.FromMilliseconds(100));
options.Budget = budget;
});
var hedge = Shield.For<int>().Hedge(options =>
{
options.Delay = TimeSpan.FromMilliseconds(200);
options.Budget = budget;
});

Feedback throttle​

The constructor follows the feedback model in gRPC retry throttling. The balance starts at MaxTokens. Each handled failure subtracts one token, including handled result values and the final attempt after the retry count is exhausted. Each acceptable successful result adds TokenRatio. Unhandled exceptions and caller cancellation do not change the balance. Updates are atomic, and the balance stays between zero and MaxTokens.

Additional attempts are allowed only while Tokens > MaxTokens / 2. Equality suppresses them. The initial attempt always runs, so successful initial traffic can replenish an exhausted budget. No timer replenishes tokens. MaxTokens must be positive; TokenRatio must be finite and at least 0.001. Refunds are rounded down to three decimal places and capped at capacity. Integer thousandths avoid floating-point drift at the threshold. Defaults are 100 and 0.1. Settings are immutable, and Tokens is a thread-safe snapshot.

A denied retry returns its last outcome. It rechecks the shared balance after callbacks and delay, and preserves ownership of a returned disposable result. A denied hedge leaves existing contenders running; their completed outcomes still update the balance, including losing attempts. Cancellation caused by selecting a winning hedge is excluded. This feedback throttle does not reserve tokens when attempts start and does not impose a concurrency or requests-per-second bound. Compose a concurrency limiter, rate limiter, or circuit breaker for those separate controls.

Configure a shared budget on one retry or hedge layer per logical dependency call. Nested policies observe their own attempt outcomes; assigning the same budget to multiple nested layers counts those observations separately. Share across independent callers and partitions to aggregate feedback. The handling predicates also classify terminal outcomes when a budget is configured, so they must be safe to call even when no further retry is allowed.

Replenishing additional-attempt allowance​

Use CreateReplenishing when concurrent callers must share a finite allowance for extra attempts:

var budget = RetryBudget.CreateReplenishing(
maxTokens: 100, replenishmentPeriod: TimeSpan.FromSeconds(10));
var retry = Shield.Retry(options =>
{
options.MaxRetries = 3;
options.Backoff = Backoff.None;
options.Budget = budget;
});
var hedge = Shield.For<int>().Hedge(options =>
{
options.Delay = TimeSpan.FromMilliseconds(200);
options.Budget = budget;
});

The allowance starts at MaxTokens. Fixed windows are anchored at budget creation, and each new window restores the full allowance. Unused tokens do not accumulate. Refill happens lazily when the budget is read or an attempt is considered; there is no background timer. A delayed read does not move later boundaries. The optional timeProvider: belongs to the budget and controls its monotonic windows, independently of any shield's clock. ReplenishmentPeriod exposes the duration; it is null for a feedback throttle. TokenRatio is zero in replenishing mode, and successful or failed outcomes do not change its balance.

Initial attempts remain free, including when the allowance is empty. Checks before callbacks and backoff only inspect availability. Each additional attempt atomically consumes one token at final admission, immediately before its continuation or generated hedge action starts. Concurrent retries and hedges cannot acquire the same token. AllowsAdditionalAttempt and Tokens are snapshots, so neither guarantees that a later launch will succeed.

Cancellation or replay suppression before admission consumes nothing. Cancellation, a failed attempt, or rejection by a downstream circuit breaker or admission limiter after admission does not refund a token. A hedge action generator that throws before producing an action consumes no token. On exhaustion, retries preserve their last outcome and ownership of a returned disposable result; hedging preserves already-running contenders. An OnRetry or OnHedge notification may run before another caller consumes the last token, so notifications do not prove an attempt starts.

Nested layers charge only their own additional attempts. For example, an outer retry consumes one token when it starts the inner layer's initial attempt; that inner initial attempt is free. A later inner retry consumes another token. Completion observations do not charge replenishing mode. This also applies to retry/hedge compositions. Separate independent executions always have their own free initial attempt; the allowance is not a bound on all physical requests.

Keep one caller-owned instance per intended dependency or partition group. Reusing it shares the allowance; creating separate instances gives separate allowances. Existing named DI registration accepts either mode and preserves the instance through shield reloads. HTTP and gRPC shields use the same options; assign the shared instance to their retry/hedge layer. Retries performed inside another HTTP handler, gRPC client, or downstream service are outside Kevlar's accounting.

Fixed windows can admit up to two windows' allowances close to a boundary. Use a rate limiter for total request rate and a concurrency limiter for in-flight work. Place those limits and a circuit breaker inside retry/hedge when each attempt must pass admission; any resulting rejection still consumes an already-admitted extra attempt. Placing a limiter outside controls logical executions. Keep per-execution retry/hedge counts and timeout deadlines as independent bounds.

Budget telemetry​

Descriptions include budget. Both modes emit retry.budget_exhausted or hedge.budget_exhausted on refusal. The corresponding kevlar.retries or kevlar.hedges counter includes refusals tagged reason=budget; exclude that tag when counting attempts that actually started. Normal attempt measurements remain untagged by reason. Synchronous Execute works with retry budgets and synchronous callbacks. See named budget registration for configuration binding.