Skip to main content

Stress Test Results

Long-running stress tests comparing sustained performance between Dekaf and Confluent.Kafka under real-world load.

Last Updated: 2026-09-06 03:46 UTC

info

The paired Dekaf vs Confluent comparison runs on demand (full_run=true from main) and updates this page. Manual dispatches stay Dekaf-only unless full_run explicitly requests the same paired publish path. Tests measure sustained performance over 15+ minutes with real Kafka instances.

At a glance​

Each row is a like-for-like comparison: both clients run the same sustained workload sequentially on the same VM, and repeated samples are aggregated with a geometric mean across both run orders.

Paired same-VM stress runSustained throughputBroker-confirmed messages per second for the same workload.

Bars are scaled within each scenario for a direct client-to-client comparison.

Median client CPU timeCPU cost per messageCPU time needed to deliver one message; shorter bars are better.

Bars are scaled within each scenario for a direct client-to-client comparison.

ScenarioDekafConfluentThroughputCPU per message
Produce — fire-and-forget1,446,139 msg/s995,249 msg/s1.5× faster2.4× less
Produce — fire-and-forget (3 brokers)1,165,881 msg/s713,584 msg/s1.6× faster2.0× less
Produce — acks=all1,540,523 msg/s1,228,208 msg/s1.3× faster2.0× less
Produce — acks=all (3 brokers)1,129,651 msg/s751,176 msg/s1.5× faster2.2× less
Produce — fire-and-forget, idempotent1,490,271 msg/s1,263,552 msg/s1.2× faster2.0× less
Produce — fire-and-forget, idempotent (3 brokers)1,116,744 msg/s803,597 msg/s1.4× faster2.2× less
Produce + consume round-trip2,625,535 msg/s1,641,424 msg/s1.6× faster2.1× less
Produce — transactional (exactly-once) (3 brokers)1,193 msg/s169 msg/s7.1× faster1.1× less
Consume — messages1,751,563 msg/s1,329,577 msg/s1.3× faster1.5× less
Consume — batches1,732,327 msg/s———
Consume — raw bytes3,741,807 msg/s———
Consume — raw byte batches4,094,242 msg/s———

"On par" means within ±5% — differences that small are run-to-run noise. "CPU per message" compares the client CPU cost of delivering one message; "less" means Dekaf needs less CPU. Rows showing "—" have no Confluent counterpart in this run (for example, batch and raw consume APIs that librdkafka does not expose). The full per-run data is below.

Full results​

Each section holds the measured per-run data behind the summary: repeated same-VM samples in both client orders, CPU per message and per request, and throughput drift across the run.

Producer (Fire-and-Forget) (15 minutes, 1000B messages)

Order-Balanced Aggregate

ClientSamplesGeomean comparison msg/sSample rangeMedian CPU μs/msgComparison Ratio
Dekaf21,446,1391,395,659–1,498,4450.721.45x
Confluent2995,249917,674–1,079,3821.721.00x

The aggregate uses the geometric mean across balanced same-VM samples run in both dekaf-first and confluent-first order. Raw ordered samples remain below.

ClientCPU μs/msgCPU μs/requestMessages/secMedian msg/sDriftSlope %/minMB/secAccepted msg/sErrorsStanding cores
Dekaf (3conn)0.62629.062,653,2072,746,307+16.7%+1.80%2530.302,653,20701.65
Dekaf (adaptive)0.65645.152,327,9722,309,915+14.2%+1.66%2220.132,327,97201.51
Dekaf (confluent-first)0.69679.681,487,2351,498,445-0.4%-0.06%1418.341,487,23501.03
Dekaf (dekaf-first)0.75736.541,394,4931,395,659-8.8%-0.64%1329.891,394,49301.04
Confluent (dekaf-first)1.64-1,027,0571,079,382-4.1%-0.68%979.481,027,05701.68
Confluent (confluent-first)1.80-954,617917,674+23.4%+2.35%910.39954,61701.72

Median msg/s is the median sampled client-side throughput interval; it shows steady-state throughput without letting a short late-run stall dominate the whole-run average.

Rows and Comparison Ratio use Median msg/s when available; older result files without interval samples fall back to Messages/sec.

Drift compares last-third with first-third average throughput. Slope is the normalized least-squares trend; steady-state below 85% of peak or slope below -1%/min fails the regression gate.

Messages/sec counts broker-confirmed deliveries (end-offset delta). Accepted msg/s is the client-side append rate — a large gap means messages were buffered or dropped without ever reaching the broker.

Producer (Fire-and-Forget), 3 Brokers (15 minutes, 1000B messages)
ClientCPU μs/msgCPU μs/requestMessages/secMedian msg/sDriftSlope %/minMB/secAccepted msg/sErrorsStanding cores
Dekaf (adaptive)1.01946.531,251,5581,211,244+17.1%+1.69%1193.581,251,55801.27
Dekaf (3conn)1.04957.631,229,4141,188,344-19.2%-1.73%1172.461,229,41401.27
Dekaf1.071019.221,151,5701,165,881+5.8%+0.59%1098.221,151,57001.23
Confluent2.11-723,211713,584-9.1%-0.72%689.71723,21101.53

Median msg/s is the median sampled client-side throughput interval; it shows steady-state throughput without letting a short late-run stall dominate the whole-run average.

Rows and Comparison Ratio use Median msg/s when available; older result files without interval samples fall back to Messages/sec.

Drift compares last-third with first-third average throughput. Slope is the normalized least-squares trend; steady-state below 85% of peak or slope below -1%/min fails the regression gate.

Messages/sec counts broker-confirmed deliveries (end-offset delta). Accepted msg/s is the client-side append rate — a large gap means messages were buffered or dropped without ever reaching the broker.

Producer (Acks All) (15 minutes, 1000B messages)

Order-Balanced Aggregate

ClientSamplesGeomean comparison msg/sSample rangeMedian CPU μs/msgComparison Ratio
Dekaf21,540,5231,515,172–1,566,2980.701.25x
Confluent21,228,2081,153,504–1,307,7511.421.00x

The aggregate uses the geometric mean across balanced same-VM samples run in both dekaf-first and confluent-first order. Raw ordered samples remain below.

ClientCPU μs/msgCPU μs/requestMessages/secMedian msg/sDriftSlope %/minMB/secAccepted msg/sErrorsStanding cores
Dekaf (confluent-first)0.70714.821,557,3261,566,298-1.1%-0.11%1485.181,557,32601.09
Dekaf (dekaf-first)0.70709.761,486,9401,515,172-6.5%-0.58%1418.061,486,94001.04
Confluent (confluent-first)1.37-1,288,5181,307,751-3.5%-0.27%1228.831,288,51801.77
Confluent (dekaf-first)1.46-1,161,4011,153,504+9.4%+0.94%1107.601,161,40101.70

Median msg/s is the median sampled client-side throughput interval; it shows steady-state throughput without letting a short late-run stall dominate the whole-run average.

Rows and Comparison Ratio use Median msg/s when available; older result files without interval samples fall back to Messages/sec.

Drift compares last-third with first-third average throughput. Slope is the normalized least-squares trend; steady-state below 85% of peak or slope below -1%/min fails the regression gate.

Messages/sec counts broker-confirmed deliveries (end-offset delta). Accepted msg/s is the client-side append rate — a large gap means messages were buffered or dropped without ever reaching the broker.

Producer (Acks All), 3 Brokers (15 minutes, 1000B messages)
ClientCPU μs/msgCPU μs/requestMessages/secMedian msg/sDriftSlope %/minMB/secAccepted msg/sErrorsStanding cores
Dekaf0.95937.601,117,8451,129,651+7.7%+0.60%1066.061,117,84501.06
Confluent2.05-751,552751,176+25.0%+2.03%716.74751,55201.54

Median msg/s is the median sampled client-side throughput interval; it shows steady-state throughput without letting a short late-run stall dominate the whole-run average.

Rows and Comparison Ratio use Median msg/s when available; older result files without interval samples fall back to Messages/sec.

Drift compares last-third with first-third average throughput. Slope is the normalized least-squares trend; steady-state below 85% of peak or slope below -1%/min fails the regression gate.

Messages/sec counts broker-confirmed deliveries (end-offset delta). Accepted msg/s is the client-side append rate — a large gap means messages were buffered or dropped without ever reaching the broker.

Producer (Fire-and-Forget, Idempotent) (15 minutes, 1000B messages)

Order-Balanced Aggregate

ClientSamplesGeomean comparison msg/sSample rangeMedian CPU μs/msgComparison Ratio
Dekaf21,490,2711,459,918–1,521,2560.691.18x
Confluent21,263,5521,233,956–1,293,8581.411.00x

The aggregate uses the geometric mean across balanced same-VM samples run in both dekaf-first and confluent-first order. Raw ordered samples remain below.

ClientCPU μs/msgCPU μs/requestMessages/secMedian msg/sDriftSlope %/minMB/secAccepted msg/sErrorsStanding cores
Dekaf (adaptive)0.61611.452,425,5322,479,418+11.5%+1.13%2313.172,425,53201.48
Dekaf (3conn)0.60587.962,274,2892,280,285-4.5%-0.39%2168.932,274,28901.36
Dekaf (confluent-first)0.70704.591,508,1191,521,256-1.3%-0.12%1438.251,508,11901.06
Dekaf (dekaf-first)0.68658.451,452,6801,459,918+0.9%+0.08%1385.381,452,68000.99
Confluent (dekaf-first)1.40-1,222,7681,293,858+2.1%-0.08%1166.121,222,76801.71
Confluent (confluent-first)1.42-1,211,8581,233,956+0.2%-0.02%1155.721,211,85801.72

Median msg/s is the median sampled client-side throughput interval; it shows steady-state throughput without letting a short late-run stall dominate the whole-run average.

Rows and Comparison Ratio use Median msg/s when available; older result files without interval samples fall back to Messages/sec.

Drift compares last-third with first-third average throughput. Slope is the normalized least-squares trend; steady-state below 85% of peak or slope below -1%/min fails the regression gate.

Messages/sec counts broker-confirmed deliveries (end-offset delta). Accepted msg/s is the client-side append rate — a large gap means messages were buffered or dropped without ever reaching the broker.

Producer (Fire-and-Forget, Idempotent), 3 Brokers (15 minutes, 1000B messages)
ClientCPU μs/msgCPU μs/requestMessages/secMedian msg/sDriftSlope %/minMB/secAccepted msg/sErrorsStanding cores
Dekaf (adaptive)0.82815.421,244,3911,255,058-1.7%-0.09%1186.741,244,39101.02
Dekaf (3conn)0.83817.601,239,1711,250,776-0.6%-0.06%1181.771,239,17101.02
Dekaf0.88864.891,110,1011,116,744+0.9%+0.09%1058.671,110,10100.97
Confluent1.97-800,917803,597+4.5%+0.41%763.81800,91701.58

Median msg/s is the median sampled client-side throughput interval; it shows steady-state throughput without letting a short late-run stall dominate the whole-run average.

Rows and Comparison Ratio use Median msg/s when available; older result files without interval samples fall back to Messages/sec.

Drift compares last-third with first-third average throughput. Slope is the normalized least-squares trend; steady-state below 85% of peak or slope below -1%/min fails the regression gate.

Messages/sec counts broker-confirmed deliveries (end-offset delta). Accepted msg/s is the client-side append rate — a large gap means messages were buffered or dropped without ever reaching the broker.

Producer → Consumer Round-Trip Steady State (15 minutes, 128B messages)
ClientCPU μs/msgCPU μs/requestMessages/secMedian msg/sDriftSlope %/minMB/secAccepted msg/sErrorsStanding cores
Dekaf0.883066.001,466,3072,625,535+59.2%+571.55%178.991,466,30701.29
Confluent1.85-127,2871,641,424+10.5%+107.74%15.54127,28700.24

Median msg/s is the median sampled client-side throughput interval; it shows steady-state throughput without letting a short late-run stall dominate the whole-run average.

Rows and Comparison Ratio use Median msg/s when available; older result files without interval samples fall back to Messages/sec.

Drift compares last-third with first-third average throughput. Slope is the normalized least-squares trend; steady-state below 85% of peak or slope below -1%/min fails the regression gate.

Messages/sec counts broker-confirmed deliveries (end-offset delta). Accepted msg/s is the client-side append rate — a large gap means messages were buffered or dropped without ever reaching the broker.

Round-Trip Validation​

ClientExpectedConsumedMissingDuplicatesCorruptOut of OrderWrong PartitionUnexpectedTimed OutResult
Confluent19,792,47719,792,477000000noPASS
Dekaf19,792,47719,792,477000000noPASS
Producer (Transactional EOS), 3 Brokers (15 minutes, 1000B messages)
ClientCPU μs/msgCPU μs/requestMessages/secMedian msg/sDriftSlope %/minMB/secAccepted msg/sErrorsStanding cores
Dekaf228.84228.848891,193+0.8%+0.09%0.851,18500.27
Confluent258.82-126169+10.1%+0.94%0.1216800.04

Median msg/s is the median sampled client-side throughput interval; it shows steady-state throughput without letting a short late-run stall dominate the whole-run average.

Rows and Comparison Ratio use Median msg/s when available; older result files without interval samples fall back to Messages/sec.

Drift compares last-third with first-third average throughput. Slope is the normalized least-squares trend; steady-state below 85% of peak or slope below -1%/min fails the regression gate.

Messages/sec counts broker-confirmed deliveries (end-offset delta). Accepted msg/s is the client-side append rate — a large gap means messages were buffered or dropped without ever reaching the broker.

Transaction Verification​

ClientAcceptedCommittedAbortedDeliveredDuplicatesShortfallAborted leaksUnexpectedMissing sentinelsStatus
Confluent151,100113,40037,700113,40000000PASS
Dekaf1,066,700800,100266,600800,10000000PASS
Consumer (15 minutes, 1000B messages, 16,384B seed batches)
ClientCPU μs/msgCPU μs/requestMessages/secMedian msg/sDriftSlope %/minMB/secAccepted msg/sErrorsStanding cores
Dekaf0.76-1,744,4571,751,563-1.1%-0.10%1663.64-01.32
Confluent1.14-1,302,3861,329,577+5.9%+0.59%1242.05-01.48

Median msg/s is the median sampled client-side throughput interval; it shows steady-state throughput without letting a short late-run stall dominate the whole-run average.

Rows and Comparison Ratio use Median msg/s when available; older result files without interval samples fall back to Messages/sec.

Drift compares last-third with first-third average throughput. Slope is the normalized least-squares trend; steady-state below 85% of peak or slope below -1%/min fails the regression gate.

Consumer (Batch) (15 minutes, 1000B messages, 16,384B seed batches)
ClientCPU μs/msgCPU μs/requestMessages/secMedian msg/sDriftSlope %/minMB/secAccepted msg/sErrorsStanding cores
Dekaf0.78-1,740,0151,732,327-6.8%-0.67%1659.41-01.36

Median msg/s is the median sampled client-side throughput interval; it shows steady-state throughput without letting a short late-run stall dominate the whole-run average.

Rows and Comparison Ratio use Median msg/s when available; older result files without interval samples fall back to Messages/sec.

Drift compares last-third with first-third average throughput. Slope is the normalized least-squares trend; steady-state below 85% of peak or slope below -1%/min fails the regression gate.

Consumer (Raw Bytes) (15 minutes, 1000B messages, 16,384B seed batches)
ClientCPU μs/msgCPU μs/requestMessages/secMedian msg/sDriftSlope %/minMB/secAccepted msg/sErrorsStanding cores
Dekaf0.42-3,718,6413,741,807+2.6%+0.25%3546.37-01.57

Median msg/s is the median sampled client-side throughput interval; it shows steady-state throughput without letting a short late-run stall dominate the whole-run average.

Rows and Comparison Ratio use Median msg/s when available; older result files without interval samples fall back to Messages/sec.

Drift compares last-third with first-third average throughput. Slope is the normalized least-squares trend; steady-state below 85% of peak or slope below -1%/min fails the regression gate.

Consumer (Raw Batch) (15 minutes, 1000B messages, 16,384B seed batches)
ClientCPU μs/msgCPU μs/requestMessages/secMedian msg/sDriftSlope %/minMB/secAccepted msg/sErrorsStanding cores
Dekaf0.34-4,168,3814,094,242+0.6%+0.14%3975.28-01.44

Median msg/s is the median sampled client-side throughput interval; it shows steady-state throughput without letting a short late-run stall dominate the whole-run average.

Rows and Comparison Ratio use Median msg/s when available; older result files without interval samples fall back to Messages/sec.

Drift compares last-third with first-third average throughput. Slope is the normalized least-squares trend; steady-state below 85% of peak or slope below -1%/min fails the regression gate.

Memory & GC statistics — latest run
ClientScenarioGen0Gen1Gen2Total AllocatedAlloc/msg
ConfluentConsumer230219312663.62 GB2.38 KB
ConfluentProducer (Fire-and-Forget)2317963511109.33 GB1.26 KB
ConfluentProducer (Fire-and-Forget)192267111031.03 GB1.26 KB
ConfluentProducer (Fire-and-Forget), 3 Brokers14962200781.11 GB1.26 KB
ConfluentProducer (Acks All)2606143611254.39 GB1.26 KB
ConfluentProducer (Acks All)282816111391.65 GB1.26 KB
ConfluentProducer (Acks All), 3 Brokers16556711811.72 GB1.26 KB
ConfluentProducer (Fire-and-Forget, Idempotent)267165111308.82 GB1.26 KB
ConfluentProducer (Fire-and-Forget, Idempotent)2704044411320.72 GB1.26 KB
ConfluentProducer (Fire-and-Forget, Idempotent), 3 Brokers17419400865.04 GB1.26 KB
ConfluentProducer → Consumer Round-Trip Steady State69150017.56 GB953 B
ConfluentProducer (Transactional EOS), 3 Brokers10410309.75 MB2.10 KB
DekafConsumer260797242959.98 GB1.98 KB
DekafConsumer (Batch)25986522952.68 GB1.98 KB
DekafConsumer (Raw Bytes)521492.08 MB0 B
DekafConsumer (Raw Batch)921989.48 MB0 B
DekafProducer (Fire-and-Forget)22542788.32 MB1 B
DekafProducer (Fire-and-Forget)22232192.89 MB0 B
DekafProducer (Fire-and-Forget), 3 Brokers14532150.72 MB0 B
DekafProducer (Acks All)22431791.56 MB1 B
DekafProducer (Acks All)22721164.98 MB0 B
DekafProducer (Acks All), 3 Brokers13232232.53 MB0 B
DekafProducer (Fire-and-Forget, Idempotent)20721152.70 MB0 B
DekafProducer (Fire-and-Forget, Idempotent)20332755.63 MB1 B
DekafProducer (Fire-and-Forget, Idempotent), 3 Brokers15132135.65 MB0 B
DekafProducer → Consumer Round-Trip Steady State1071312.81 GB153 B
DekafProducer (Transactional EOS), 3 Brokers8411175.07 MB172 B
Dekaf (3conn)Producer (Fire-and-Forget)327311.27 GB1 B
Dekaf (3conn)Producer (Fire-and-Forget), 3 Brokers18852737.70 MB1 B
Dekaf (3conn)Producer (Fire-and-Forget, Idempotent)276211.14 GB1 B
Dekaf (3conn)Producer (Fire-and-Forget, Idempotent), 3 Brokers15732670.81 MB1 B
Dekaf (adaptive)Producer (Fire-and-Forget)307321.13 GB1 B
Dekaf (adaptive)Producer (Fire-and-Forget), 3 Brokers17053723.16 MB1 B
Dekaf (adaptive)Producer (Fire-and-Forget, Idempotent)315321.17 GB1 B
Dekaf (adaptive)Producer (Fire-and-Forget, Idempotent), 3 Brokers14832675.10 MB1 B

Confluent.Kafka uses native librdkafka; .NET GC allocation counters exclude unmanaged allocations.


About These Tests​

Stress tests measure sustained performance over extended periods against real Kafka brokers, with both clients paired on the same VM for a fair comparison.

Methodology — how these numbers are produced
  • Real Kafka: Tests run against actual Apache Kafka instances
  • CPU Isolation: Brokers are pinned to dedicated cores and the client under test to its own cores, so the client — not the broker — is the measured bottleneck
  • RAM-backed Broker Logs: Kafka log dirs are mounted on tmpfs so disk I/O never caps broker ingestion
  • Delivered Throughput: producer tables report broker-confirmed throughput, measured as the end-offset delta across all partitions — not the client-side append rate, which can run far ahead of what the broker ever accepts
  • Median Interval Throughput: table order and comparison ratios use median sampled client-side msg/s when available, which is less sensitive to short late-run stalls than the whole-run mean
  • Same-VM Pairing: comparable Dekaf and Confluent scenarios run sequentially inside one job/VM; 1-broker producer acceptance lanes run twice in opposite client orders and publish a geometric-mean aggregate, while other lanes alternate order by workflow run number
  • Backpressure Parity: both producers are bounded to the same 512 MB local buffer (Dekaf BufferMemory, librdkafka queue.buffering.max) and block on a full buffer, so neither client can absorb an unbounded backlog into RAM
  • Consumer Loop Replay: Consumer tests re-read a pre-seeded topic (seek to beginning when drained) instead of racing a live feeder, so the consumer itself is measured; table headings report the 16KB seed batch size because it amplifies per-batch costs relative to well-batched workloads
  • Delivery Latency Sampling: 1 in 1000 produced messages is awaited end-to-end to record true broker round-trip latency
  • Adaptive-Connections Row: the four paired fire-and-forget/idempotent producer lanes also run one Dekaf pass with the library default (adaptive connection scaling enabled, one connection to start), labelled Dekaf (adaptive); like the 3-connection control it is excluded from the headline comparison, which stays pinned to one connection to match Confluent
  • Round-Trip Correctness: Bounded sequenced payloads are consumed back and checked for corruption, wrong partitions, gaps, duplicates, and reordering
  • Round-Trip CPU Scope: CPU time covers both bulk production and consumer validation; it is not a producer-only metric
  • Round-Trip Alloc Scope: the GC/alloc window likewise spans production plus consume-side validation; values are deliberately consumed as byte[] on both clients for parity, so each consumed payload is materialized as a fresh array (~152 B at 128 B messages) — the expected allocation floor for this lane, not a leak
  • CPU Efficiency: CPU time per message differentiates client efficiency even at equal throughput
  • Noise-Aware Trends: each scenario's throughput, CPU per message and Dekaf delivery-latency p50/p95/p99 are compared with its last 10 matching runs using a median ± 2×MAD band; one adverse excursion warns and two consecutive regressions fail the workflow, and paired lanes fail only when the same-run Dekaf/Confluent ratio regressed too
  • Latency Product Bars: p95 within 3× the configured delivery-latency target and p50/p99 within 2× the same-run Confluent control are reported as product goals, not gates; a lane that misses a bar shows a warning until the trend band moves it
  • Parallel Execution: Each scenario runs in its own isolated environment
  • Both Clients: Direct comparison between Dekaf and Confluent.Kafka
  • Memory Monitoring: Tracks GC behavior and memory usage over time
  • Error Rates: Ensures stability under load