Retry
Re-execute the delegate when it fails, waiting between attempts.
See the exceptions reference for failures the default retry clause handles.
Quick forms
var exponential = Shield.Retry(3); // equal jitter, 250ms base, 30s cap
var constant = Shield.Retry(3, Backoff.Constant(TimeSpan.FromSeconds(1)));
var linear = Shield.Retry(3, Backoff.Linear(TimeSpan.FromMilliseconds(500)));
var forever = Shield.RetryForever(
Backoff.Exponential(TimeSpan.FromSeconds(1), maxDelay: TimeSpan.FromMinutes(1)));
The count is retries, not attempts: Retry(3) makes up to 4 total attempts — the initial call
plus 3 retries. The same reading applies to MaxRetries in the options form, and to every Retry
overload on Shield, Shield<T>, and the When… builders.
This executable check removes delays and verifies the count:
var attempts = 0;
var retryWithoutDelay = Shield.Retry(3, Backoff.None);
try
{
await retryWithoutDelay.ExecuteAsync(_ =>
{
attempts++;
return ValueTask.FromException(new HttpRequestException("offline"));
});
}
catch (HttpRequestException)
{
}
if (attempts != 4)
{
throw new InvalidOperationException($"Expected 4 attempts, observed {attempts}.");
}
Defaults
Bare Retry(n) and RetryForever() use Backoff.Default:
| Setting | Default |
|---|---|
| Curve | Exponential |
| Base delay | 250 ms |
| Factor | 2 |
| Jitter | Equal: each delay is scaled by a value in [0.5, 1.5) |
| Maximum delay | 30 seconds |
The cap also applies to RetryForever(). Equal jitter prevents callers from retrying in lockstep.
Backoff
Backoff factories describe the delay sequence:
Backoff.Constant(TimeSpan.FromSeconds(1)); // 1s, 1s, 1s, ...
Backoff.Linear(TimeSpan.FromMilliseconds(500)); // 500ms, 1s, 1.5s, ...
Backoff.Exponential(TimeSpan.FromSeconds(1)); // ~1s, ~2s, ~4s, ... (jittered)
Backoff.Exponential(TimeSpan.FromSeconds(1), jitter: Jitter.Full);
Backoff.Exponential(TimeSpan.FromSeconds(1), jitter: Jitter.Decorrelated);
Backoff.Custom(attempt => TimeSpan.FromMilliseconds(100 * attempt)); // attempt is 1-based
_ = Backoff.None; // no delay between attempts
_ = Backoff.Default; // what bare Retry(n) uses
Jitter.Noneuses the exact curve;Equalscales it by [0.5, 1.5);Fullselects [0, curve);Decorrelatedselects [base delay, previous delay × 3]. Decorrelated state is isolated per retry execution. Direct backoff consumers can carry the effective preceding delay throughGetDelay(attempt, previousDelay); pass zero for the first draw.Exponential(baseDelay, factor = 2.0, maxDelay = null, jitter = Jitter.Equal)is the default curve.Backoff.Default=Exponential(250ms, maxDelay: 30s).Constant(delay, jitter = Jitter.None)andLinear(step, maxDelay = null, jitter = Jitter.None)also accept jitter explicitly.- Configuration accepts enum names and Boolean aliases (
true=Equal,false=None). - Without
maxDelay, built-in and custom backoffs clamp only at the runtime timer limit (uint.MaxValue - 1milliseconds, roughly 49.7 days). UsemaxDelayfor an application-specific cap.
Full options
API reference: RetryOptions and RetryOptions<T>.
var retry = Shield.Retry(o =>
{
o.MaxRetries = 5;
o.Backoff = Backoff.Custom(attempt => TimeSpan.FromMilliseconds(100 * attempt));
o.MaxDelay = TimeSpan.FromSeconds(10);
o.OnRetry = e =>
{
logger.LogWarning(e.Exception, "Retry {AttemptNumber} after {Delay}", e.AttemptNumber, e.Delay);
return default;
};
o.DelayGenerator = e => new(e.AttemptNumber == 0 ? TimeSpan.Zero : null);
// ^ return a TimeSpan to override the computed delay, or null to keep it
});
| Option | Default | What it does |
|---|---|---|
MaxRetries | 3 | Retries after the initial attempt — 3 means up to 4 total attempts; int.MaxValue = forever |
Backoff | Backoff.Default | The delay sequence (see above) |
MaxDelay | backoff cap | Absolute cap applied to every delay, including DelayGenerator output. When unset, the backoff's cap applies (30s for Backoff.Default); a backoff without a cap falls back to the runtime timer limit |
OnRetry | — | Awaited before each retry sleeps — attempt number, final delay, failure. Return default when the work is synchronous |
DelayGenerator | — | Per-retry override returning ValueTask<TimeSpan?>: a TimeSpan replaces the computed delay, null keeps it. This is how Retry-After support works |
HandlesException | — | Local exception predicate; replaces the ambient clause for this retry |
HandlesResult (RetryOptions<T>) | — | Local result predicate; replaces the ambient clause together with HandlesException |
Invalid option values throw KevlarConfigurationException
and identify the options type, property, and offending value.
Order per retry: backoff computes the delay → effective cap (MaxDelay ?? Backoff.MaxDelay)
clamps it → the awaited DelayGenerator may override it → the awaited OnRetry sees the final
delay and handled outcome → a superseded disposable result is disposed → sleep → retry metrics
are recorded immediately before the next attempt. Concurrent hedge compositions defer disposal
until a replacement outcome exists, so execution-wide suppression can still return the original
outcome without returning a disposed result. The
generator's null and negative results are ignored, and the effective cap clamps its override.
Calling SuppressAdditionalAttempts() from either typed or untyped event stops before retry
metrics, disposal, sleep, or another attempt, including attempts in nested child shields.
If cancellation arrives during that sleep, the next attempt never starts. Cancellation surfaces as
an OperationCanceledException with no InnerException; the failure that triggered the retry is
not retained. OnRetry has already run, but kevlar.retries is not recorded because no next
attempt started.
When ordinary result handling triggers another attempt, Kevlar disposes the handled result before
the next attempt starts. Concurrent hedge compositions dispose it after its replacement completes.
Kevlar prefers IAsyncDisposable.DisposeAsync() when a result implements both disposal
interfaces. OnRetry runs first so it can inspect the live result. The final result returned to the
caller is never disposed by the retry strategy. Disposal failures are reported through
KevlarDiagnostics.OnCallbackError as CallbackErrorKind.ResultDisposal and do not replace the
pipeline outcome.
Both hooks return ValueTask. A hook that completes synchronously (return default;, new(value))
costs nothing extra and works with synchronous Execute. A hook that yields is awaited by
ExecuteAsync; reached through synchronous Execute, it throws NotSupportedException at that
call. See synchronous execution compatibility.
Notification-hook exceptions follow the shared callback-failure contract:
they are reported and never replace the protected outcome.
RetryOptions and RetryOptions<T> are standalone sibling types. Both expose the same
MaxRetries, Backoff, and MaxDelay settings, while their callback properties use distinct
RetryEvent and RetryEvent<T> delegates. Configure shared scalar defaults in each options
lambda; a RetryOptions<T> instance is not assignable to RetryOptions.
Generators receive the caller token through e.Context.CancellationToken. If cancellation
arrives while a generator is awaiting, notification hooks still run after it completes, then the
next attempt is suppressed and caller cancellation surfaces. A generator exception surfaces with
its original identity and skips later hooks. RetryEvent.Context is pooled execution state: use it
only before the returned ValueTask completes; never retain it or its property bag.
On an untyped Shield, the RetryEvent callbacks receive: AttemptNumber (zero-based, so 0 is the first retry after the initial execution), Delay, Exception (null when a handled result triggered the retry), Result (the handled result, boxed as object?) and Context. On a typed Shield<T>, the events are RetryEvent<T> instead: same AttemptNumber/Delay/Context, plus the handled failure as a directly stored typed Outcome<T> — e.Outcome.Result is your T, with no boxing, reconstruction, or cast.
var attemptNumbers = new List<int>();
var numberedRetry = Shield.Retry(options =>
{
options.MaxRetries = 3;
options.Backoff = Backoff.None;
options.OnRetry = retry =>
{
attemptNumbers.Add(retry.AttemptNumber);
return default; // completes synchronously, so synchronous Execute below is fine
};
});
try
{
numberedRetry.Execute(static _ => throw new InvalidOperationException());
}
catch (InvalidOperationException)
{
}
Console.WriteLine(string.Join(",", attemptNumbers)); // 0,1,2
What gets retried
Whatever the current handling clause says is a failure—ordinary
exceptions under the default, or the outcomes selected by your When/WhenResult clause:
var retry = Shield
.When<HttpRequestException>()
.Or<TimeoutExceededException>()
.Retry(5);
The options-lambda form can instead set HandlesException and, on Shield<T>, HandlesResult.
Setting either creates a per-strategy override;
unspecified outcome kinds are not handled.
Placement in the chain
var scopedRetry = Shield
.Timeout(TimeSpan.FromSeconds(30)) // total budget: retries must fit inside
.Retry(3)
.Timeout(TimeSpan.FromSeconds(5)); // each attempt gets 5s, and Retry sees the TimeoutExceededException
Retry outside a per-attempt timeout retries timeouts; retry inside a circuit breaker hammers a struggling dependency before the breaker sees the pattern. The composition rules cover this in depth.
Respecting an outer deadline
Set RespectDeadline = true to stop retrying when the next delay is greater than or equal to
KevlarContext.Deadline's remaining budget. The effective delay includes DelayGenerator and
MaxDelay. Kevlar returns the last handled result or exception immediately, without disposing
that returned result. It emits retry.skipped_deadline and increments kevlar.retries.skipped
with reason=deadline; deadline refusals do not increment kevlar.retries.
var deadlineAware = Shield.Timeout(TimeSpan.FromSeconds(1)).Retry(options =>
{
options.MaxRetries = 3;
options.Backoff = Backoff.Constant(TimeSpan.FromMilliseconds(600), jitter: Jitter.None);
options.RespectDeadline = true;
});
If attempts fail immediately, this pipeline performs the initial attempt and one retry at 600 ms. The next 600 ms delay cannot fit, so the second failure surfaces instead of waiting for the outer timeout. A viable delay does not guarantee that the following operation will finish before the deadline; cancellation still applies during that operation.
Both typed and untyped options default to false in 1.x, preserving existing timeout and
cancellation behavior. With no enclosing timeout, the option has no effect. The budget is
checked before OnRetry and again afterward because an asynchronous callback can consume time.
A skipped retry normally does not invoke OnRetry; it can already have run when its duration
causes the second check to skip. Caller cancellation retains priority. Describe() includes
deadline-aware when enabled. DI configuration also accepts Retry.RespectDeadline.
Shared retry budgets
Use one RetryBudget for a downstream dependency when failures across many callers should
suppress additional traffic. The same instance can be assigned to retry and hedge options on
independent shields, including shields created for different partition keys:
var budget = new RetryBudget(maxTokens: 100, tokenRatio: 0.1);
var retry = Shield.Retry(options =>
{
options.MaxRetries = 3;
options.Backoff = Backoff.Exponential(TimeSpan.FromMilliseconds(100));
options.Budget = budget;
});
var hedge = Shield.For<int>().Hedge(options =>
{
options.Delay = TimeSpan.FromMilliseconds(200);
options.Budget = budget;
});
Feedback throttle
The constructor follows the feedback model in gRPC retry throttling.
The balance starts at MaxTokens. Each handled failure subtracts one token, including handled
result values and the final attempt after the retry count is exhausted. Each acceptable successful
result adds TokenRatio. Unhandled exceptions and caller cancellation do not change the balance.
Updates are atomic, and the balance stays between zero and MaxTokens.
Additional attempts are allowed only while Tokens > MaxTokens / 2. Equality suppresses them.
The initial attempt always runs, so successful initial traffic can replenish an exhausted budget.
No timer replenishes tokens. MaxTokens must be positive; TokenRatio must be finite and at least
0.001. Refunds are rounded down to three decimal places and capped at capacity. Integer thousandths
avoid floating-point drift at the threshold. Defaults are 100 and 0.1. Settings are immutable, and
Tokens is a thread-safe snapshot.
A denied retry returns its last outcome. It rechecks the shared balance after callbacks and delay, and preserves ownership of a returned disposable result. A denied hedge leaves existing contenders running; their completed outcomes still update the balance, including losing attempts. Cancellation caused by selecting a winning hedge is excluded. This feedback throttle does not reserve tokens when attempts start and does not impose a concurrency or requests-per-second bound. Compose a concurrency limiter, rate limiter, or circuit breaker for those separate controls.
Configure a shared budget on one retry or hedge layer per logical dependency call. Nested policies observe their own attempt outcomes; assigning the same budget to multiple nested layers counts those observations separately. Share across independent callers and partitions to aggregate feedback. The handling predicates also classify terminal outcomes when a budget is configured, so they must be safe to call even when no further retry is allowed.
Replenishing additional-attempt allowance
Use CreateReplenishing when concurrent callers must share a finite allowance for extra attempts:
var budget = RetryBudget.CreateReplenishing(
maxTokens: 100, replenishmentPeriod: TimeSpan.FromSeconds(10));
var retry = Shield.Retry(options =>
{
options.MaxRetries = 3;
options.Backoff = Backoff.None;
options.Budget = budget;
});
var hedge = Shield.For<int>().Hedge(options =>
{
options.Delay = TimeSpan.FromMilliseconds(200);
options.Budget = budget;
});
The allowance starts at MaxTokens. Fixed windows are anchored at budget creation, and each new
window restores the full allowance. Unused tokens do not accumulate. Refill happens lazily when
the budget is read or an attempt is considered; there is no background timer. A delayed read does
not move later boundaries. The optional timeProvider: belongs to the budget and controls its
monotonic windows, independently of any shield's clock. ReplenishmentPeriod exposes the duration;
it is null for a feedback throttle. TokenRatio is zero in replenishing mode, and successful or
failed outcomes do not change its balance.
Initial attempts remain free, including when the allowance is empty. Checks before callbacks and
backoff only inspect availability. Each additional attempt atomically consumes one token at final
admission, immediately before its continuation or generated hedge action starts. Concurrent retries
and hedges cannot acquire the same token. AllowsAdditionalAttempt and Tokens are snapshots, so
neither guarantees that a later launch will succeed.
Cancellation or replay suppression before admission consumes nothing. Cancellation, a failed
attempt, or rejection by a downstream circuit breaker or admission limiter after admission does
not refund a token. A hedge action generator that throws before producing an action consumes no
token. On exhaustion, retries preserve their last outcome and ownership of a returned disposable
result; hedging preserves already-running contenders. An OnRetry or OnHedge notification may
run before another caller consumes the last token, so notifications do not prove an attempt starts.
Nested layers charge only their own additional attempts. For example, an outer retry consumes one token when it starts the inner layer's initial attempt; that inner initial attempt is free. A later inner retry consumes another token. Completion observations do not charge replenishing mode. This also applies to retry/hedge compositions. Separate independent executions always have their own free initial attempt; the allowance is not a bound on all physical requests.
Keep one caller-owned instance per intended dependency or partition group. Reusing it shares the allowance; creating separate instances gives separate allowances. Existing named DI registration accepts either mode and preserves the instance through shield reloads. HTTP and gRPC shields use the same options; assign the shared instance to their retry/hedge layer. Retries performed inside another HTTP handler, gRPC client, or downstream service are outside Kevlar's accounting.
Fixed windows can admit up to two windows' allowances close to a boundary. Use a rate limiter for total request rate and a concurrency limiter for in-flight work. Place those limits and a circuit breaker inside retry/hedge when each attempt must pass admission; any resulting rejection still consumes an already-admitted extra attempt. Placing a limiter outside controls logical executions. Keep per-execution retry/hedge counts and timeout deadlines as independent bounds.
Budget telemetry
Descriptions include budget. Both modes emit retry.budget_exhausted or hedge.budget_exhausted on refusal.
The corresponding kevlar.retries or kevlar.hedges counter includes refusals tagged reason=budget;
exclude that tag when counting attempts that actually started. Normal attempt measurements remain
untagged by reason. Synchronous Execute works with retry budgets and synchronous callbacks.
See named budget registration for configuration binding.