Skip to main content

Controlled outage and recovery

An open-loop load generator exercises Kevlar against an in-process asynchronous dependency. Each composition gets fresh strategy state and the same healthy, slowdown, outage, and recovery schedule. These measurements describe the configured simulation, not production capacity or a comparison with another library.

Measured 2026-09-28T07:29:35.7634385+00:00 at commit 813cbbc0208420514e23de38e369cfe1174cd3d1.

Request outcomes and latency​

Latency starts at the scheduled arrival, including scheduler delay and queue time. p95/p99 include every submitted outcome, including fast pipeline rejections; success p99 is separate. Harness rejections remain in offered/failed counts but have no submitted latency. Requests belong to their scheduled-arrival phase, even when they finish later.

CompositionPhaseOfferedSuccess / failedp95 / p99 msSuccess p99 msScheduler p99 ms
retryhealthy30003000 / 06.28 / 6.836.831.14
retryslowdown30003000 / 0301.75 / 302.16302.161.14
retryoutage30003 / 299762.63 / 63.6145.851.87
retryrecovery30003000 / 06.28 / 6.836.831.15
retry-breakerhealthy30003000 / 06.26 / 7.187.181.87
retry-breakerslowdown30003000 / 0301.79 / 302.15302.151.20
retry-breakeroutage30000 / 30001.21 / 21.18not observed2.00
retry-breakerrecovery30002995 / 56.26 / 7.037.061.54
limit-retry-breakerhealthy30003000 / 06.26 / 6.896.891.15
limit-retry-breakerslowdown30001600 / 1400598.16 / 599.95600.121.14
limit-retry-breakeroutage30000 / 30001.12 / 20.72not observed1.12
limit-retry-breakerrecovery30002996 / 46.26 / 6.886.881.14
limit-hedge-breakerhealthy30003000 / 06.25 / 6.676.671.33
limit-hedge-breakerslowdown30003000 / 051.92 / 52.5952.591.19
limit-hedge-breakeroutage30000 / 30002.07 / 21.36not observed1.36
limit-hedge-breakerrecovery30002984 / 166.25 / 7.047.041.22

Offered load and downstream amplification​

Submitted means accepted by the harness. Admitted means at least one dependency attempt began. A request can be admitted and later rejected by the circuit breaker on a retry. Attempt counts below belong to logical requests in the row; JSON also records attempts by actual start phase.

CompositionPhaseSubmitted / admittedRejected: harness / limiter / circuitAttemptsAttempts / offeredAttempts / admitted
retryhealthy3000 / 30000 / 0 / 030001.001.00
retryslowdown3000 / 30000 / 0 / 030001.001.00
retryoutage3000 / 30000 / 0 / 089993.003.00
retryrecovery3000 / 30000 / 0 / 030001.001.00
retry-breakerhealthy3000 / 30000 / 0 / 030001.001.00
retry-breakerslowdown3000 / 30000 / 0 / 030001.001.00
retry-breakeroutage3000 / 680 / 0 / 2993840.031.24
retry-breakerrecovery3000 / 29950 / 0 / 529951.001.00
limit-retry-breakerhealthy3000 / 30000 / 0 / 030001.001.00
limit-retry-breakerslowdown3000 / 16080 / 1384 / 1316160.541.00
limit-retry-breakeroutage3000 / 560 / 3 / 2997560.021.00
limit-retry-breakerrecovery3000 / 29960 / 0 / 429961.001.00
limit-hedge-breakerhealthy3000 / 30000 / 0 / 030001.001.00
limit-hedge-breakerslowdown3000 / 30000 / 0 / 060002.002.00
limit-hedge-breakeroutage3000 / 1420 / 0 / 29961460.051.03
limit-hedge-breakerrecovery3000 / 29840 / 0 / 1629840.991.00

Recovery and hedge cleanup​

Recovery time ends at the later of the end of the first fully successful 250 ms scheduled-arrival window and completion of every request in that window, measured from the recovery phase boundary. The window must include at least floor(rate × 0.25), minimum one, successful arrivals. Any harness shedding or pipeline rejection invalidates that window. Not observed means no qualifying complete window.

Cancelled attempts include cancelled hedge losers. Cleanup delay is measured after caller completion, clamped to zero. Every attempt must finish cleanup before a scenario publishes results or the next scenario starts.

CompositionRecovery msCancelledLosers pending at caller completionCleanup p99 / max msPeak logical / downstreamActive after drain
retry250.00000.00 / 0.0031 / 310
retry-breaker500.00000.00 / 0.0031 / 310
limit-retry-breaker500.00000.00 / 0.0033 / 160
limit-hedge-breaker500.003088308821.02 / 21.897 / 120

Configuration and limitations​

  • 100 offered requests/s, 30 seconds per phase, four phases per composition, and a harness cap of 512 active logical requests.
  • Arrivals keep their scheduled times when the runner falls behind; delayed arrivals catch up without waiting for prior responses. Scheduler lag exposes bursts caused by a saturated generator. High lag or harness shedding limits interpretation.
  • Healthy/recovery dependency latency: 5 ms. Slowdown: first attempt 300 ms, additional attempts 30 ms. This intentionally models a slow original replica with a faster alternate; hedging will not help uniformly slow replicas.
  • Outage: every attempt fails after 20 ms. An attempt uses the phase at its start, so outstanding outage attempts can fail after recovery begins.
  • Cancelled attempts perform 20 ms of asynchronous cleanup before completing.
  • Retry: two extra attempts with no backoff. Breaker: five consecutive handled failures, 500 ms break. Limiter: 16 permits, 16 queued. Hedge: one extra attempt after 20 ms; handled failures can launch it earlier.
  • Compositions run sequentially in fixed order, with separate warmup and fresh shield state. This is a scenario demonstration, not an A–B performance acceptance experiment. Shared runners, timer resolution, GC, and process scheduling affect latency. No latency threshold gates CI.
  • Percentiles use nearest rank. Short smoke runs have small samples and exist to validate accounting and cleanup. The simulator omits sockets, serialization, server queues, database locks, and real cancellation delays.

Environment and raw results​

  • Ubuntu 24.04.5 LTS; .NET 10.0.12; X64; 4 visible processors.
  • Current published JSON accompanies this page. The outage workflow also uploads JSON and Markdown results for each run. JSON records all configuration, scheduled phase boundaries, actual phase attempt counts, and wall duration for each composition.

Reproduce​

dotnet run -c Release --project benchmarks/Kevlar.OutageTests -- --phase-seconds 30 --rate 100 --max-inflight 512
python .github/scripts/outage_docs.py --input artifacts/outage/outage-results.json --output artifacts/outage/outage-results.md

Use --phase-seconds 1 for smoke validation. Scheduled/manual runs use 30 seconds per phase. For sustained throughput and allocation measurements, see stress tests.