Benchmarks
Kevlar vs Polly v8 across every strategy, measured with BenchmarkDotNet on GitHub Actions and republished automatically by the benchmarks workflow.
Last updated 2026-10-05 04:27 UTC (commit a78926f).
Microbenchmarks on shared CI runners are noisy. The vs Polly column is the median ratio over the most recent runs (up to 10); anything within ±20% is reported as on par. Absolute times move with runner hardware — the ratios are the signal.
Pipeline overhead floor
What an empty pipeline costs per execution — the fixed tax every strategy builds on.
| Scenario | Kevlar | Polly | Kevlar allocated | Polly allocated | Kevlar vs Polly |
|---|---|---|---|---|---|
| Empty pipeline — async | 18.6 ns | 59.9 ns | 0 B | 0 B | 3.6× faster |
| Empty pipeline — reference-state baseline | 17.4 ns | — | 0 B | — | — |
| Empty pipeline — caller-seeded context | 66.1 ns | — | 0 B | — | — |
| EmptyOutcomeState | 10.0 ns | — | 0 B | — | — |
| EmptyTaskOutcomeState | 12.5 ns | — | 0 B | — | — |
| Empty pipeline — zero-closure state overload | 14.9 ns | 60.7 ns | 0 B | 0 B | 4.1× faster |
| Empty pipeline — sync | 9.5 ns | 28.6 ns | 0 B | 0 B | 3.1× faster |
| NestedEmptyAsync | 105 ns | — | 0 B | — | — |
| NestedEmptySync | 96.0 ns | — | 0 B | — | — |
| NestedWithPropertiesAsync | 159 ns | — | 0 B | — | — |
| NestedWithPropertiesSync | 160 ns | — | 0 B | — | — |
Retry
Happy path (judge overhead only) and a recovery path where every call fails twice before succeeding, with backoff disabled so only strategy machinery is measured.
| Scenario | Kevlar | Polly | Kevlar allocated | Polly allocated | Kevlar vs Polly |
|---|---|---|---|---|---|
| Retry(3) — success on first attempt | 140 ns | 258 ns | 0 B | 24 B | 1.8× faster |
| Retry(3) — two failures then success | 4.23 μs | 4.44 μs | 192 B | 328 B | on par |
Timeout
The timeout never fires; this is the cost of arming and disarming the cancellation plumbing on every call.
| Scenario | Kevlar | Polly | Kevlar allocated | Polly allocated | Kevlar vs Polly |
|---|---|---|---|---|---|
| Timeout(10 s) — completes instantly | 170 ns | 218 ns | 0 B | 0 B | on par |
| SynchronousGenerator_HappyPath | 217 ns | — | 0 B | — | — |
| AsynchronousGenerator_HappyPath | 1.67 μs | — | 560 B | — | — |
| AsyncHookConfigured_HappyPath | 212 ns | — | 0 B | — | — |
Circuit breaker
Ratio/sampling bookkeeping while closed, and the fast-fail rejection cost while manually isolated — caught as a thrown exception, and read from the no-throw ExecuteOutcomeAsync API.
| Scenario | Kevlar | Polly | Kevlar allocated | Polly allocated | Kevlar vs Polly |
|---|---|---|---|---|---|
| Isolated circuit — fast-fail rejection thrown | 5.46 μs | 5.55 μs | 1.3 KB | 1.3 KB | on par |
| Isolated circuit — fast-fail rejection via ExecuteOutcomeAsync | 115 ns | 180 ns | 144 B | 200 B | 1.6× faster |
| Ratio breaker, closed — success | 191 ns | 269 ns | 0 B | 24 B | 1.4× faster |
| DynamicDurationConfigured | 209 ns | — | 0 B | — | — |
| AsyncCallbackConfigured | 210 ns | — | 0 B | — | — |
| SlowCallDetectionConfigured | 303 ns | — | 0 B | — | — |
Circuit breaker under contention
Eight workers share one open circuit, so every call is a fast-fail rejection observed as an outcome; this is the per-call cost during an outage.
| Scenario | Kevlar | Polly | Kevlar allocated | Polly allocated | Kevlar vs Polly |
|---|---|---|---|---|---|
| Open circuit — 8 workers rejected | 65.1 ns | 201 ns | 144 B | 200 B | 3.2× faster |
Fallback
Pass-through when the execution succeeds, and substitution when it throws.
| Scenario | Kevlar | Polly | Kevlar allocated | Polly allocated | Kevlar vs Polly |
|---|---|---|---|---|---|
| NoNotification | 2.51 μs | — | 200 B | — | — |
| CompletedAsyncNotification | 2.46 μs | — | 200 B | — | — |
| YieldingAsyncNotification | 6.46 μs | — | 798 B | — | — |
| Fallback — not triggered | 133 ns | 135 ns | 0 B | 0 B | on par |
| SynchronousDelegate_Triggered | 2.33 μs | — | 200 B | — | — |
| Fallback — triggered by exception | 2.47 μs | 2.56 μs | 200 B | 256 B | on par |
| EmptyVoid | 37.9 ns | — | 0 B | — | — |
| VoidPassThrough | 173 ns | — | 0 B | — | — |
Rate limit
Uncontended token-bucket permit acquisition — every call is admitted — and the rejection cost once the budget is exhausted, thrown vs ExecuteOutcomeAsync.
| Scenario | Kevlar | Polly | Kevlar allocated | Polly allocated | Kevlar vs Polly |
|---|---|---|---|---|---|
| Budget exhausted — rejection thrown | 5.44 μs | 70.84 μs | 1.3 KB | 41.4 KB | 13× faster |
| Budget exhausted — rejection via ExecuteOutcomeAsync | 151 ns | 46.33 μs | 136 B | 35.5 KB | 305× faster |
| Token bucket — uncontended acquire | 155 ns | 164 ns | 0 B | 0 B | on par |
| WithHooks_Uncontended | 151 ns | — | 0 B | — | — |
| FrameworkAdapter_Uncontended | 142 ns | — | 0 B | — | — |
| PartitionedFrameworkAdapter_Uncontended | 186 ns | — | 32 B | — | — |
Concurrency limit
A single caller against a large permit count; acquire/release cost with no queueing. The rejection rows hold the only permit and compare thrown vs ExecuteOutcomeAsync.
| Scenario | Kevlar | Polly | Kevlar allocated | Polly allocated | Kevlar vs Polly |
|---|---|---|---|---|---|
| Permits exhausted — rejection thrown | 5.36 μs | 65.21 μs | 1.3 KB | 39.5 KB | 12× faster |
| Permits exhausted — rejection via ExecuteOutcomeAsync | 98.8 ns | 42.70 μs | 136 B | 33.6 KB | 432× faster |
| Concurrency limit — uncontended | 146 ns | 221 ns | 0 B | 40 B | 1.3× faster |
| WithHooks_Uncontended | 151 ns | — | 0 B | — | — |
| Adaptive_Uncontended | 216 ns | — | 0 B | — | — |
Typed result handling
Retry configured to treat a sentinel result as a failure; the returned value never matches, so this is the per-call cost of judging results.
| Scenario | Kevlar | Polly | Kevlar allocated | Polly allocated | Kevlar vs Polly |
|---|---|---|---|---|---|
| Typed retry — result judged, no retry | 141 ns | 269 ns | 0 B | 0 B | 1.4× faster |
Composed pipelines
How per-call overhead scales with pipeline depth when nothing goes wrong.
| Scenario | Kevlar | Polly | Kevlar allocated | Polly allocated | Kevlar vs Polly |
|---|---|---|---|---|---|
| Timeout → Retry → ratio breaker | 331 ns | 689 ns | 0 B | 48 B | 1.8× faster |
| Token bucket → Timeout → Retry → ratio breaker → Concurrency limit | 476 ns | 1.02 μs | 0 B | 88 B | 1.9× faster |
| TokenBucketRatioFiveStrategyChainSync | 478 ns | 1.02 μs | 0 B | 88 B | 2.1× faster |
Environment
- AMD EPYC 7763
- .NET 10.0.12 (10.0.12, 10.0.1226.42308), BenchmarkDotNet 0.15.8
- Times are medians; allocations are per operation.
Reproduce
dotnet run -c Release --project benchmarks/Kevlar.Benchmarks -- --filter '*'
As always with microbenchmarks: measure your own workload before optimizing around these numbers. Nanosecond differences matter in tight loops and high-throughput services; they don't matter around a 50 ms network call.