vs h2o (quicly, GRO patch, 12 threads, native) 1 KB object
What the four clients are, and what each table measures
- h3x
- this project, 32 worker threads, each multiplexing its share of the connections over one UDP socket (connections share the thread's 4-tuple, distinguished by QUIC connection ID)
- h3x spc
- the same binary with
--socket-per-conn: one UDP socket per connection, so every connection has a unique 4-tuple; cross-connection sendmmsg batching is off in this mode - h2o-httpclient
- h2o's reference client, single-threaded and single-connection; run here
as one process per connection (own socket) with
-C= m streams, stopped at 10 s, completions counted from status lines. At 256-512 connections that is 256-512 processes on 32 CPUs (scheduler oversubscribed, fork storm at cell start), so treat that column there as a floor, not a ceiling - h2load
- nghttp2's load tool on the ngtcp2 stack, one socket per connection
- req/s
- completed responses per second. This is the only table where the four clients are directly comparable, so it is the only one rank-coloured green-to-red across the row
- failed requests
- requests the client gave up on. Against h3x in shared-socket mode these are the 4-tuple churn described under "Why h3x drops requests"; the connection dies, whatever was in flight on it fails, and the pool re-dials
- server UDP drops
- datagrams the kernel discarded on the server's socket because
its receive buffer was full, read per cell from
/proc/net/udp. Every drop we measured was a receive-buffer overflow (the per-socket counter, snmpRcvbufErrorsand snmpInErrorswere exactly equal). These cost throughput, not correctness: QUIC retransmits them. All three servers run the stock 212992-byte buffer (net.core.rmem_max), so this column is a drain-rate difference, not a buffer-size one
throughput
| conns | m | h3x req/s | h3x spc req/s | h2o-httpclient req/s | h2load req/s |
|---|---|---|---|---|---|
| 64 | 1 | 878,718 ●100% | 743,94285% | 652,00074% | 5,1510.6% |
| 64 | 2 | 1,116,933 ●100% | 951,34585% | 881,36379% | 13,9461.2% |
| 64 | 8 | 2,246,583 ●100% | 2,164,76696% | 1,459,97065% | 80,7453.6% |
| 64 | 16 | 3,070,93099% | 3,100,605 ●100% | 1,662,67254% | 213,7856.9% |
| 64 | 32 | 3,185,573 ●100% | 3,055,43396% | 1,894,13259% | 496,83216% |
| 64 | 64 | 3,544,322 ●100% | 3,065,78886% | 1,854,70052% | 3,392,87296% |
| 128 | 1 | 957,304 ●100% | 657,33969% | 638,59167% | 12,9111.3% |
| 128 | 2 | 896,64297% | 730,99279% | 922,887 ●100% | 21,5152.3% |
| 128 | 8 | 2,305,510 ●100% | 1,862,75281% | 1,502,51065% | 190,0748.2% |
| 128 | 16 | 2,957,664 ●100% | 2,598,56788% | 1,638,07955% | 406,79214% |
| 128 | 32 | 3,289,423 ●100% | 3,076,12794% | 1,728,77653% | 846,50726% |
| 128 | 64 | 3,327,313 ●100% | 2,881,63687% | 1,755,58153% | 3,318,701100% |
| 256 | 1 | 793,314 ●100% | 578,60173% | 583,26974% | 20,7452.6% |
| 256 | 2 | 863,101 ●100% | 676,78978% | 854,95299% | 44,0815.1% |
| 256 | 8 | 2,261,353 ●100% | 1,799,04380% | 1,481,89766% | 245,55211% |
| 256 | 16 | 3,070,014 ●100% | 2,152,62470% | 1,514,19049% | 586,29619% |
| 256 | 32 | 3,061,832 ●100% | 2,435,09080% | 1,576,93552% | 1,335,83144% |
| 256 | 64 | 3,031,436 ●100% | 2,564,53785% | 1,660,58255% | 2,956,72298% |
| 512 | 1 | 835,002 ●100% | 607,34173% | 510,14561% | 29,8403.6% |
| 512 | 2 | 892,699 ●100% | 647,11172% | 761,26085% | 57,5376.4% |
| 512 | 8 | 2,144,369 ●100% | 1,629,36376% | 1,337,89062% | 344,42616% |
| 512 | 16 | 2,529,035 ●100% | 2,180,33286% | 1,372,62754% | 820,52432% |
| 512 | 32 | 2,485,235 ●100% | 2,378,07596% | 1,424,76857% | 1,420,61157% |
| 512 | 64 | 2,484,22094% | 2,414,49391% | 1,631,56761% | 2,654,307 ●100% |
failed requests summed over m
| conns | per socket | h3x | h3x spc | h2o-httpclient | h2load |
|---|---|---|---|---|---|
| 64 | 2 | 0 | 0 | 19 | 0 |
| 128 | 4 | 0 | 0 | 31 | 0 |
| 256 | 8 | 0 | 0 | 30 | 0 |
| 512 | 16 | 0 | 0 | 54 | 0 |
per-cell breakdown, all 24 cells
| conns | m | h3x | h3x spc | h2o-httpclient | h2load |
|---|---|---|---|---|---|
| 64 | 1 | 0 | 0 | 1 | 0 |
| 64 | 2 | 0 | 0 | 3 | 0 |
| 64 | 8 | 0 | 0 | 4 | 0 |
| 64 | 16 | 0 | 0 | 6 | 0 |
| 64 | 32 | 0 | 0 | 4 | 0 |
| 64 | 64 | 0 | 0 | 1 | 0 |
| 128 | 1 | 0 | 0 | 2 | 0 |
| 128 | 2 | 0 | 0 | 6 | 0 |
| 128 | 8 | 0 | 0 | 6 | 0 |
| 128 | 16 | 0 | 0 | 5 | 0 |
| 128 | 32 | 0 | 0 | 6 | 0 |
| 128 | 64 | 0 | 0 | 6 | 0 |
| 256 | 1 | 0 | 0 | 7 | 0 |
| 256 | 2 | 0 | 0 | 6 | 0 |
| 256 | 8 | 0 | 0 | 7 | 0 |
| 256 | 16 | 0 | 0 | 6 | 0 |
| 256 | 32 | 0 | 0 | 3 | 0 |
| 256 | 64 | 0 | 0 | 1 | 0 |
| 512 | 1 | 0 | 0 | 9 | 0 |
| 512 | 2 | 0 | 0 | 10 | 0 |
| 512 | 8 | 0 | 0 | 11 | 0 |
| 512 | 16 | 0 | 0 | 8 | 0 |
| 512 | 32 | 0 | 0 | 5 | 0 |
| 512 | 64 | 0 | 0 | 11 | 0 |
server UDP datagrams dropped receive-buffer overflow, summed over m
| conns | per socket | h3x | h3x spc | h2o-httpclient | h2load |
|---|---|---|---|---|---|
| 64 | 2 | 0 | 0 | 0 | 0 |
| 128 | 4 | 20 | 0 | 0 | 0 |
| 256 | 8 | 36,652 | 83,428 | 86 | 0 |
| 512 | 16 | 460,277 | 1,079,736 | 54,310 | 1,459 |
per-cell breakdown, all 24 cells
| conns | m | h3x | h3x spc | h2o-httpclient | h2load |
|---|---|---|---|---|---|
| 64 | 1 | 0 | 0 | 0 | 0 |
| 64 | 2 | 0 | 0 | 0 | 0 |
| 64 | 8 | 0 | 0 | 0 | 0 |
| 64 | 16 | 0 | 0 | 0 | 0 |
| 64 | 32 | 0 | 0 | 0 | 0 |
| 64 | 64 | 0 | 0 | 0 | 0 |
| 128 | 1 | 0 | 0 | 0 | 0 |
| 128 | 2 | 0 | 0 | 0 | 0 |
| 128 | 8 | 0 | 0 | 0 | 0 |
| 128 | 16 | 0 | 0 | 0 | 0 |
| 128 | 32 | 20 | 0 | 0 | 0 |
| 128 | 64 | 0 | 0 | 0 | 0 |
| 256 | 1 | 0 | 0 | 0 | 0 |
| 256 | 2 | 9,651 | 46,155 | 0 | 0 |
| 256 | 8 | 1,460 | 8,181 | 0 | 0 |
| 256 | 16 | 0 | 17,328 | 0 | 0 |
| 256 | 32 | 6,026 | 11,105 | 0 | 0 |
| 256 | 64 | 19,515 | 659 | 86 | 0 |
| 512 | 1 | 13,083 | 76,769 | 0 | 0 |
| 512 | 2 | 155,199 | 320,308 | 0 | 48 |
| 512 | 8 | 85,953 | 348,563 | 0 | 0 |
| 512 | 16 | 26,539 | 208,880 | 0 | 0 |
| 512 | 32 | 78,478 | 80,526 | 9,137 | 0 |
| 512 | 64 | 101,025 | 44,690 | 45,173 | 1,411 |
Methodology & caveats
h2load failed 0 requests in all nine grids, on every server at every payload. Its collapse at
m=1-2 happens only against h2o and only at 1 KB (5k-14k req/s, recovering to 3.39M by m=64): a
pairing-specific interaction between ngtcp2 and the quicly server at low concurrency, not general
client behaviour, and it does not reappear at the larger payloads. h2o-httpclient's failures stay
small throughout (128-319 per grid, process-teardown artifacts) except against nginx at 128 KB
(2,773). --socket-per-conn costs h3x throughput at 1 KB (+16.1% against h2o, +25.0%
against nginx) and buys reliability everywhere; at 128 KB the two modes are within noise of each
other on throughput and spc still fails far less.
All 864 cells are one sitting (2026-07-29, bench/grid-v3.sh), same box, same servers, so
columns and payloads are directly comparable. Each payload is served at its own URL
(/, /20k.html, /128k.html) rather than by swapping one
file's contents, and every cell re-derives the served object size from the client's own byte
counters. A round of earlier sweeps here was invalidated by silently benchmarking a 20 KB object
against 1 KB tables; naming the payload in the request makes that failure structurally impossible
rather than merely checked for.
Four cells at 128 KB drift 2-3% from the nominal object size and are reported by the generator rather than hidden. Three are nginx h3x cells failing 11-23% of requests, where bytes from partial bodies count toward throughput while their requests count as failed, inflating bytes-per-completed. The fourth is haproxy h2load at 512x64, where 32,768 streams were still in flight at the 10 s cutoff, deflating it by the same mechanism in reverse. Both are artifacts of the measurement window. A real payload mix-up here would be 20x off, not 3%.
Single unpinned runs. One 10 s run per cell, no repeats. A control experiment across two grids, comparing cells whose configuration was byte-identical, saw swings from -17.6% to +13.5% with no consistent sign. So differences under roughly 20% are not distinguishable from run-to-run variance, and only the large effects should be read as real — which here means the payload-size reversal (h3x 21-1 to h2load 15-9 against h2o), the 4.6x nginx gap at 128 KB, the drop counts, and h2load's clean failure record. Separating anything smaller needs repeated samples per cell.
Why h3x (shared-socket mode) drops requests
The trigger is multiplexing many QUIC connections over one UDP socket per worker. Differential
tests against haproxy: one connection per socket is flawless (32 conns / 32 sockets: 7.19M requests,
0 drops, 0 re-dials; thread-count controlled separately with 4 conns / 4 sockets, also 0/0), while
8 connections per socket, same 32 connections and same multiplexing, drops 1,428 and re-dials
constantly (355k requests on resumed connections in 5 s, with no reconnect flag). Confirmed at scale by the spc column, which fails 0 requests against nginx and 1,904 against
haproxy versus shared-socket h3x's 3,914 and 53,372. Servers differ in how they tolerate multiple QUIC
connections sharing a 4-tuple: h2o accepts it (this is h2o's own client architecture), nginx churns
moderately, haproxy constantly. Every drop surfaces as "I/O error" at stream attach: a shared-socket
connection dies, whatever was in flight on it fails, the pool re-dials with the cached ticket, and
the run continues, so throughput barely dips while the drop counter grows. The close reason on the
wire is still unread (CONNECTION_CLOSE is encrypted; reading it needs SSLKEYLOGFILE wiring), and
against nginx there is additionally a smaller classic handshake-timeout tail under saturation (2,559
churn errors vs 489 timeouts in the probed cell). Fix: --socket-per-conn (one socket
per connection, the h2load model), the spc column.
vs nginx (OpenSSL-QUIC, docker, 12 workers) 1 KB object
What the four clients are, and what each table measures
- h3x
- this project, 32 worker threads, each multiplexing its share of the connections over one UDP socket (connections share the thread's 4-tuple, distinguished by QUIC connection ID)
- h3x spc
- the same binary with
--socket-per-conn: one UDP socket per connection, so every connection has a unique 4-tuple; cross-connection sendmmsg batching is off in this mode - h2o-httpclient
- h2o's reference client, single-threaded and single-connection; run here
as one process per connection (own socket) with
-C= m streams, stopped at 10 s, completions counted from status lines. At 256-512 connections that is 256-512 processes on 32 CPUs (scheduler oversubscribed, fork storm at cell start), so treat that column there as a floor, not a ceiling - h2load
- nghttp2's load tool on the ngtcp2 stack, one socket per connection
- req/s
- completed responses per second. This is the only table where the four clients are directly comparable, so it is the only one rank-coloured green-to-red across the row
- failed requests
- requests the client gave up on. Against h3x in shared-socket mode these are the 4-tuple churn described under "Why h3x drops requests"; the connection dies, whatever was in flight on it fails, and the pool re-dials
- server UDP drops
- datagrams the kernel discarded on the server's socket because
its receive buffer was full, read per cell from
/proc/net/udp. Every drop we measured was a receive-buffer overflow (the per-socket counter, snmpRcvbufErrorsand snmpInErrorswere exactly equal). These cost throughput, not correctness: QUIC retransmits them. All three servers run the stock 212992-byte buffer (net.core.rmem_max), so this column is a drain-rate difference, not a buffer-size one
throughput
| conns | m | h3x req/s | h3x spc req/s | h2o-httpclient req/s | h2load req/s |
|---|---|---|---|---|---|
| 64 | 1 | 534,046 ●100% | 418,84278% | 433,17481% | 465,10987% |
| 64 | 2 | 531,20066% | 441,93855% | 507,00863% | 803,137 ●100% |
| 64 | 8 | 826,67562% | 739,90955% | 584,26244% | 1,333,435 ●100% |
| 64 | 16 | 1,026,35767% | 881,94257% | 597,10939% | 1,541,618 ●100% |
| 64 | 32 | 1,167,41069% | 946,38256% | 626,60537% | 1,687,661 ●100% |
| 64 | 64 | 1,254,75572% | 968,40156% | 695,29140% | 1,732,569 ●100% |
| 128 | 1 | 555,39597% | 360,43163% | 464,47981% | 573,442 ●100% |
| 128 | 2 | 545,89365% | 412,39649% | 484,39658% | 839,093 ●100% |
| 128 | 8 | 876,45268% | 673,32752% | 528,88541% | 1,290,878 ●100% |
| 128 | 16 | 1,064,02573% | 781,74653% | 557,22538% | 1,464,982 ●100% |
| 128 | 32 | 1,213,36278% | 921,40659% | 654,30042% | 1,563,092 ●100% |
| 128 | 64 | 1,171,14871% | 862,90253% | 716,58144% | 1,642,759 ●100% |
| 256 | 1 | 517,75697% | 375,91970% | 439,51682% | 534,139 ●100% |
| 256 | 2 | 570,11170% | 348,89743% | 470,40958% | 815,479 ●100% |
| 256 | 8 | 868,18469% | 599,55748% | 507,67040% | 1,262,079 ●100% |
| 256 | 16 | 992,49972% | 741,42054% | 616,02345% | 1,382,786 ●100% |
| 256 | 32 | 1,139,00779% | 852,64259% | 706,44049% | 1,450,029 ●100% |
| 256 | 64 | 1,187,25575% | 930,39459% | 770,54849% | 1,574,516 ●100% |
| 512 | 1 | 587,593 ●100% | 358,28161% | 404,55069% | 517,79388% |
| 512 | 2 | 588,36779% | 393,86053% | 440,01059% | 742,717 ●100% |
| 512 | 8 | 890,48976% | 610,91852% | 579,84949% | 1,174,973 ●100% |
| 512 | 16 | 1,013,69877% | 681,48252% | 680,18952% | 1,313,310 ●100% |
| 512 | 32 | 1,050,21578% | 821,25561% | 720,44253% | 1,354,465 ●100% |
| 512 | 64 | 962,50470% | 877,66163% | 767,33155% | 1,382,747 ●100% |
failed requests summed over m
| conns | per socket | h3x | h3x spc | h2o-httpclient | h2load |
|---|---|---|---|---|---|
| 64 | 2 | 3 | 0 | 5 | 0 |
| 128 | 4 | 350 | 0 | 14 | 0 |
| 256 | 8 | 1,869 | 0 | 12 | 0 |
| 512 | 16 | 1,692 | 0 | 20 | 0 |
per-cell breakdown, all 24 cells
| conns | m | h3x | h3x spc | h2o-httpclient | h2load |
|---|---|---|---|---|---|
| 64 | 1 | 0 | 0 | 0 | 0 |
| 64 | 2 | 3 | 0 | 1 | 0 |
| 64 | 8 | 0 | 0 | 3 | 0 |
| 64 | 16 | 0 | 0 | 0 | 0 |
| 64 | 32 | 0 | 0 | 1 | 0 |
| 64 | 64 | 0 | 0 | 0 | 0 |
| 128 | 1 | 6 | 0 | 1 | 0 |
| 128 | 2 | 10 | 0 | 2 | 0 |
| 128 | 8 | 60 | 0 | 4 | 0 |
| 128 | 16 | 82 | 0 | 2 | 0 |
| 128 | 32 | 0 | 0 | 3 | 0 |
| 128 | 64 | 192 | 0 | 2 | 0 |
| 256 | 1 | 0 | 0 | 1 | 0 |
| 256 | 2 | 10 | 0 | 3 | 0 |
| 256 | 8 | 72 | 0 | 2 | 0 |
| 256 | 16 | 0 | 0 | 4 | 0 |
| 256 | 32 | 415 | 0 | 1 | 0 |
| 256 | 64 | 1,372 | 0 | 1 | 0 |
| 512 | 1 | 0 | 0 | 6 | 0 |
| 512 | 2 | 48 | 0 | 1 | 0 |
| 512 | 8 | 185 | 0 | 6 | 0 |
| 512 | 16 | 300 | 0 | 6 | 0 |
| 512 | 32 | 531 | 0 | 1 | 0 |
| 512 | 64 | 628 | 0 | 0 | 0 |
server UDP datagrams dropped receive-buffer overflow, summed over m
| conns | per socket | h3x | h3x spc | h2o-httpclient | h2load |
|---|---|---|---|---|---|
| 64 | 2 | 27,218 | 100,838 | 207,098 | 138 |
| 128 | 4 | 85,062 | 414,658 | 769,979 | 1,416 |
| 256 | 8 | 568,862 | 974,746 | 1,557,677 | 3,770 |
| 512 | 16 | 1,694,869 | 1,733,211 | 2,449,586 | 40,292 |
per-cell breakdown, all 24 cells
| conns | m | h3x | h3x spc | h2o-httpclient | h2load |
|---|---|---|---|---|---|
| 64 | 1 | 0 | 0 | 0 | 0 |
| 64 | 2 | 0 | 0 | 0 | 0 |
| 64 | 8 | 0 | 0 | 0 | 2 |
| 64 | 16 | 0 | 0 | 0 | 6 |
| 64 | 32 | 22,500 | 15,312 | 19,499 | 10 |
| 64 | 64 | 4,718 | 85,526 | 187,599 | 120 |
| 128 | 1 | 0 | 0 | 0 | 0 |
| 128 | 2 | 0 | 158 | 0 | 0 |
| 128 | 8 | 0 | 19,355 | 0 | 224 |
| 128 | 16 | 268 | 58,791 | 31,686 | 51 |
| 128 | 32 | 17,580 | 126,709 | 233,704 | 411 |
| 128 | 64 | 67,214 | 209,645 | 504,589 | 730 |
| 256 | 1 | 0 | 0 | 0 | 0 |
| 256 | 2 | 104 | 79,355 | 0 | 0 |
| 256 | 8 | 43,859 | 198,799 | 15,817 | 399 |
| 256 | 16 | 135,745 | 239,907 | 225,286 | 1,783 |
| 256 | 32 | 176,973 | 226,583 | 515,796 | 1,464 |
| 256 | 64 | 212,181 | 230,102 | 800,778 | 124 |
| 512 | 1 | 36,558 | 57,582 | 0 | 10,114 |
| 512 | 2 | 179,019 | 310,382 | 0 | 1,118 |
| 512 | 8 | 355,667 | 411,926 | 191,572 | 9,643 |
| 512 | 16 | 409,815 | 327,348 | 464,700 | 9,271 |
| 512 | 32 | 366,039 | 314,080 | 793,768 | 5,502 |
| 512 | 64 | 347,771 | 311,893 | 999,546 | 4,644 |
Methodology & caveats
h2load failed 0 requests in all nine grids, on every server at every payload. Its collapse at
m=1-2 happens only against h2o and only at 1 KB (5k-14k req/s, recovering to 3.39M by m=64): a
pairing-specific interaction between ngtcp2 and the quicly server at low concurrency, not general
client behaviour, and it does not reappear at the larger payloads. h2o-httpclient's failures stay
small throughout (128-319 per grid, process-teardown artifacts) except against nginx at 128 KB
(2,773). --socket-per-conn costs h3x throughput at 1 KB (+16.1% against h2o, +25.0%
against nginx) and buys reliability everywhere; at 128 KB the two modes are within noise of each
other on throughput and spc still fails far less.
All 864 cells are one sitting (2026-07-29, bench/grid-v3.sh), same box, same servers, so
columns and payloads are directly comparable. Each payload is served at its own URL
(/, /20k.html, /128k.html) rather than by swapping one
file's contents, and every cell re-derives the served object size from the client's own byte
counters. A round of earlier sweeps here was invalidated by silently benchmarking a 20 KB object
against 1 KB tables; naming the payload in the request makes that failure structurally impossible
rather than merely checked for.
Four cells at 128 KB drift 2-3% from the nominal object size and are reported by the generator rather than hidden. Three are nginx h3x cells failing 11-23% of requests, where bytes from partial bodies count toward throughput while their requests count as failed, inflating bytes-per-completed. The fourth is haproxy h2load at 512x64, where 32,768 streams were still in flight at the 10 s cutoff, deflating it by the same mechanism in reverse. Both are artifacts of the measurement window. A real payload mix-up here would be 20x off, not 3%.
Single unpinned runs. One 10 s run per cell, no repeats. A control experiment across two grids, comparing cells whose configuration was byte-identical, saw swings from -17.6% to +13.5% with no consistent sign. So differences under roughly 20% are not distinguishable from run-to-run variance, and only the large effects should be read as real — which here means the payload-size reversal (h3x 21-1 to h2load 15-9 against h2o), the 4.6x nginx gap at 128 KB, the drop counts, and h2load's clean failure record. Separating anything smaller needs repeated samples per cell.
Why h3x (shared-socket mode) drops requests
The trigger is multiplexing many QUIC connections over one UDP socket per worker. Differential
tests against haproxy: one connection per socket is flawless (32 conns / 32 sockets: 7.19M requests,
0 drops, 0 re-dials; thread-count controlled separately with 4 conns / 4 sockets, also 0/0), while
8 connections per socket, same 32 connections and same multiplexing, drops 1,428 and re-dials
constantly (355k requests on resumed connections in 5 s, with no reconnect flag). Confirmed at scale by the spc column, which fails 0 requests against nginx and 1,904 against
haproxy versus shared-socket h3x's 3,914 and 53,372. Servers differ in how they tolerate multiple QUIC
connections sharing a 4-tuple: h2o accepts it (this is h2o's own client architecture), nginx churns
moderately, haproxy constantly. Every drop surfaces as "I/O error" at stream attach: a shared-socket
connection dies, whatever was in flight on it fails, the pool re-dials with the cached ticket, and
the run continues, so throughput barely dips while the drop counter grows. The close reason on the
wire is still unread (CONNECTION_CLOSE is encrypted; reading it needs SSLKEYLOGFILE wiring), and
against nginx there is additionally a smaller classic handshake-timeout tail under saturation (2,559
churn errors vs 489 timeouts in the probed cell). Fix: --socket-per-conn (one socket
per connection, the h2load model), the spc column.
vs haproxy (HAProxy QUIC, docker, nbthread 12) 1 KB object
What the four clients are, and what each table measures
- h3x
- this project, 32 worker threads, each multiplexing its share of the connections over one UDP socket (connections share the thread's 4-tuple, distinguished by QUIC connection ID)
- h3x spc
- the same binary with
--socket-per-conn: one UDP socket per connection, so every connection has a unique 4-tuple; cross-connection sendmmsg batching is off in this mode - h2o-httpclient
- h2o's reference client, single-threaded and single-connection; run here
as one process per connection (own socket) with
-C= m streams, stopped at 10 s, completions counted from status lines. At 256-512 connections that is 256-512 processes on 32 CPUs (scheduler oversubscribed, fork storm at cell start), so treat that column there as a floor, not a ceiling - h2load
- nghttp2's load tool on the ngtcp2 stack, one socket per connection
- req/s
- completed responses per second. This is the only table where the four clients are directly comparable, so it is the only one rank-coloured green-to-red across the row
- failed requests
- requests the client gave up on. Against h3x in shared-socket mode these are the 4-tuple churn described under "Why h3x drops requests"; the connection dies, whatever was in flight on it fails, and the pool re-dials
- server UDP drops
- datagrams the kernel discarded on the server's socket because
its receive buffer was full, read per cell from
/proc/net/udp. Every drop we measured was a receive-buffer overflow (the per-socket counter, snmpRcvbufErrorsand snmpInErrorswere exactly equal). These cost throughput, not correctness: QUIC retransmits them. All three servers run the stock 212992-byte buffer (net.core.rmem_max), so this column is a drain-rate difference, not a buffer-size one
throughput
| conns | m | h3x req/s | h3x spc req/s | h2o-httpclient req/s | h2load req/s |
|---|---|---|---|---|---|
| 64 | 1 | 483,390 ●100% | 477,95499% | 473,12998% | 402,26383% |
| 64 | 2 | 572,909 ●100% | 569,88299% | 450,59679% | 508,84789% |
| 64 | 8 | 802,167 ●100% | 771,70096% | 571,21971% | 714,85689% |
| 64 | 16 | 1,166,51699% | 1,182,932 ●100% | 766,40265% | 960,78681% |
| 64 | 32 | 1,319,32796% | 1,368,573 ●100% | 922,06567% | 1,178,51286% |
| 64 | 64 | 1,407,43597% | 1,454,157 ●100% | 1,066,48273% | 1,324,69291% |
| 128 | 1 | 507,401 ●100% | 487,01496% | 390,08177% | 453,82589% |
| 128 | 2 | 530,028 ●100% | 522,06298% | 421,01479% | 528,194100% |
| 128 | 8 | 905,64499% | 911,513 ●100% | 515,59157% | 665,32873% |
| 128 | 16 | 1,114,501 ●100% | 1,090,61198% | 734,84466% | 918,96082% |
| 128 | 32 | 1,247,85998% | 1,279,499 ●100% | 870,01268% | 1,085,44685% |
| 128 | 64 | 1,277,492 ●100% | 1,275,612100% | 960,96175% | 1,185,73893% |
| 256 | 1 | 516,309 ●100% | 483,39194% | 360,50470% | 477,82093% |
| 256 | 2 | 557,771 ●100% | 546,12398% | 403,09572% | 537,27696% |
| 256 | 8 | 922,13999% | 927,344 ●100% | 473,09351% | 595,78064% |
| 256 | 16 | 1,064,59598% | 1,086,774 ●100% | 648,54060% | 837,03777% |
| 256 | 32 | 1,108,32897% | 1,138,650 ●100% | 691,71161% | 907,23880% |
| 256 | 64 | 1,104,288 ●100% | 1,096,28899% | 691,17263% | 887,88680% |
| 512 | 1 | 502,341 ●100% | 482,05496% | 312,77662% | 465,31293% |
| 512 | 2 | 496,32599% | 498,923 ●100% | 361,23272% | 495,11999% |
| 512 | 8 | 763,73894% | 814,546 ●100% | 405,94750% | 530,55365% |
| 512 | 16 | 972,437 ●100% | 958,79999% | 506,69752% | 632,38765% |
| 512 | 32 | 869,58788% | 993,699 ●100% | 570,64357% | 696,43670% |
| 512 | 64 | 709,00274% | 962,962 ●100% | 604,42763% | 826,52586% |
failed requests summed over m
| conns | per socket | h3x | h3x spc | h2o-httpclient | h2load |
|---|---|---|---|---|---|
| 64 | 2 | 497 | 0 | 12 | 0 |
| 128 | 4 | 5,410 | 0 | 21 | 0 |
| 256 | 8 | 15,636 | 72 | 54 | 0 |
| 512 | 16 | 31,829 | 1,832 | 44 | 0 |
per-cell breakdown, all 24 cells
| conns | m | h3x | h3x spc | h2o-httpclient | h2load |
|---|---|---|---|---|---|
| 64 | 1 | 8 | 0 | 3 | 0 |
| 64 | 2 | 9 | 0 | 3 | 0 |
| 64 | 8 | 48 | 0 | 4 | 0 |
| 64 | 16 | 192 | 0 | 1 | 0 |
| 64 | 32 | 48 | 0 | 0 | 0 |
| 64 | 64 | 192 | 0 | 1 | 0 |
| 128 | 1 | 15 | 0 | 3 | 0 |
| 128 | 2 | 91 | 0 | 6 | 0 |
| 128 | 8 | 253 | 0 | 6 | 0 |
| 128 | 16 | 740 | 0 | 3 | 0 |
| 128 | 32 | 1,262 | 0 | 2 | 0 |
| 128 | 64 | 3,049 | 0 | 1 | 0 |
| 256 | 1 | 91 | 0 | 2 | 0 |
| 256 | 2 | 197 | 0 | 6 | 0 |
| 256 | 8 | 514 | 72 | 3 | 0 |
| 256 | 16 | 767 | 0 | 3 | 0 |
| 256 | 32 | 3,758 | 0 | 36 | 0 |
| 256 | 64 | 10,309 | 0 | 4 | 0 |
| 512 | 1 | 204 | 0 | 1 | 0 |
| 512 | 2 | 250 | 136 | 4 | 0 |
| 512 | 8 | 1,603 | 136 | 13 | 0 |
| 512 | 16 | 2,977 | 0 | 18 | 0 |
| 512 | 32 | 11,460 | 530 | 2 | 0 |
| 512 | 64 | 15,335 | 1,030 | 6 | 0 |
server UDP datagrams dropped receive-buffer overflow, summed over m
| conns | per socket | h3x | h3x spc | h2o-httpclient | h2load |
|---|---|---|---|---|---|
| 64 | 2 | 0 | 0 | 0 | 0 |
| 128 | 4 | 0 | 0 | 0 | 0 |
| 256 | 8 | 0 | 0 | 0 | 0 |
| 512 | 16 | 0 | 3 | 0 | 43 |
per-cell breakdown, all 24 cells
| conns | m | h3x | h3x spc | h2o-httpclient | h2load |
|---|---|---|---|---|---|
| 64 | 1 | 0 | 0 | 0 | 0 |
| 64 | 2 | 0 | 0 | 0 | 0 |
| 64 | 8 | 0 | 0 | 0 | 0 |
| 64 | 16 | 0 | 0 | 0 | 0 |
| 64 | 32 | 0 | 0 | 0 | 0 |
| 64 | 64 | 0 | 0 | 0 | 0 |
| 128 | 1 | 0 | 0 | 0 | 0 |
| 128 | 2 | 0 | 0 | 0 | 0 |
| 128 | 8 | 0 | 0 | 0 | 0 |
| 128 | 16 | 0 | 0 | 0 | 0 |
| 128 | 32 | 0 | 0 | 0 | 0 |
| 128 | 64 | 0 | 0 | 0 | 0 |
| 256 | 1 | 0 | 0 | 0 | 0 |
| 256 | 2 | 0 | 0 | 0 | 0 |
| 256 | 8 | 0 | 0 | 0 | 0 |
| 256 | 16 | 0 | 0 | 0 | 0 |
| 256 | 32 | 0 | 0 | 0 | 0 |
| 256 | 64 | 0 | 0 | 0 | 0 |
| 512 | 1 | 0 | 0 | 0 | 0 |
| 512 | 2 | 0 | 0 | 0 | 31 |
| 512 | 8 | 0 | 3 | 0 | 0 |
| 512 | 16 | 0 | 0 | 0 | 12 |
| 512 | 32 | 0 | 0 | 0 | 0 |
| 512 | 64 | 0 | 0 | 0 | 0 |
Methodology & caveats
h2load failed 0 requests in all nine grids, on every server at every payload. Its collapse at
m=1-2 happens only against h2o and only at 1 KB (5k-14k req/s, recovering to 3.39M by m=64): a
pairing-specific interaction between ngtcp2 and the quicly server at low concurrency, not general
client behaviour, and it does not reappear at the larger payloads. h2o-httpclient's failures stay
small throughout (128-319 per grid, process-teardown artifacts) except against nginx at 128 KB
(2,773). --socket-per-conn costs h3x throughput at 1 KB (+16.1% against h2o, +25.0%
against nginx) and buys reliability everywhere; at 128 KB the two modes are within noise of each
other on throughput and spc still fails far less.
All 864 cells are one sitting (2026-07-29, bench/grid-v3.sh), same box, same servers, so
columns and payloads are directly comparable. Each payload is served at its own URL
(/, /20k.html, /128k.html) rather than by swapping one
file's contents, and every cell re-derives the served object size from the client's own byte
counters. A round of earlier sweeps here was invalidated by silently benchmarking a 20 KB object
against 1 KB tables; naming the payload in the request makes that failure structurally impossible
rather than merely checked for.
Four cells at 128 KB drift 2-3% from the nominal object size and are reported by the generator rather than hidden. Three are nginx h3x cells failing 11-23% of requests, where bytes from partial bodies count toward throughput while their requests count as failed, inflating bytes-per-completed. The fourth is haproxy h2load at 512x64, where 32,768 streams were still in flight at the 10 s cutoff, deflating it by the same mechanism in reverse. Both are artifacts of the measurement window. A real payload mix-up here would be 20x off, not 3%.
Single unpinned runs. One 10 s run per cell, no repeats. A control experiment across two grids, comparing cells whose configuration was byte-identical, saw swings from -17.6% to +13.5% with no consistent sign. So differences under roughly 20% are not distinguishable from run-to-run variance, and only the large effects should be read as real — which here means the payload-size reversal (h3x 21-1 to h2load 15-9 against h2o), the 4.6x nginx gap at 128 KB, the drop counts, and h2load's clean failure record. Separating anything smaller needs repeated samples per cell.
Why h3x (shared-socket mode) drops requests
The trigger is multiplexing many QUIC connections over one UDP socket per worker. Differential
tests against haproxy: one connection per socket is flawless (32 conns / 32 sockets: 7.19M requests,
0 drops, 0 re-dials; thread-count controlled separately with 4 conns / 4 sockets, also 0/0), while
8 connections per socket, same 32 connections and same multiplexing, drops 1,428 and re-dials
constantly (355k requests on resumed connections in 5 s, with no reconnect flag). Confirmed at scale by the spc column, which fails 0 requests against nginx and 1,904 against
haproxy versus shared-socket h3x's 3,914 and 53,372. Servers differ in how they tolerate multiple QUIC
connections sharing a 4-tuple: h2o accepts it (this is h2o's own client architecture), nginx churns
moderately, haproxy constantly. Every drop surfaces as "I/O error" at stream attach: a shared-socket
connection dies, whatever was in flight on it fails, the pool re-dials with the cached ticket, and
the run continues, so throughput barely dips while the drop counter grows. The close reason on the
wire is still unread (CONNECTION_CLOSE is encrypted; reading it needs SSLKEYLOGFILE wiring), and
against nginx there is additionally a smaller classic handshake-timeout tail under saturation (2,559
churn errors vs 489 timeouts in the probed cell). Fix: --socket-per-conn (one socket
per connection, the h2load model), the spc column.
vs h2o (quicly, GRO patch, 12 threads, native) 20 KB object
What the four clients are, and what each table measures
- h3x
- this project, 32 worker threads, each multiplexing its share of the connections over one UDP socket (connections share the thread's 4-tuple, distinguished by QUIC connection ID)
- h3x spc
- the same binary with
--socket-per-conn: one UDP socket per connection, so every connection has a unique 4-tuple; cross-connection sendmmsg batching is off in this mode - h2o-httpclient
- h2o's reference client, single-threaded and single-connection; run here
as one process per connection (own socket) with
-C= m streams, stopped at 10 s, completions counted from status lines. At 256-512 connections that is 256-512 processes on 32 CPUs (scheduler oversubscribed, fork storm at cell start), so treat that column there as a floor, not a ceiling - h2load
- nghttp2's load tool on the ngtcp2 stack, one socket per connection
- req/s
- completed responses per second. This is the only table where the four clients are directly comparable, so it is the only one rank-coloured green-to-red across the row
- failed requests
- requests the client gave up on. Against h3x in shared-socket mode these are the 4-tuple churn described under "Why h3x drops requests"; the connection dies, whatever was in flight on it fails, and the pool re-dials
- server UDP drops
- datagrams the kernel discarded on the server's socket because
its receive buffer was full, read per cell from
/proc/net/udp. Every drop we measured was a receive-buffer overflow (the per-socket counter, snmpRcvbufErrorsand snmpInErrorswere exactly equal). These cost throughput, not correctness: QUIC retransmits them. All three servers run the stock 212992-byte buffer (net.core.rmem_max), so this column is a drain-rate difference, not a buffer-size one
throughput
| conns | m | h3x req/s | h3x spc req/s | h2o-httpclient req/s | h2load req/s |
|---|---|---|---|---|---|
| 64 | 1 | 336,806 ●100% | 298,80289% | 316,36694% | 10,4613.1% |
| 64 | 2 | 434,960 ●100% | 332,73176% | 342,92679% | 16,5353.8% |
| 64 | 8 | 532,672 ●100% | 464,36887% | 417,94878% | 184,46435% |
| 64 | 16 | 493,875 ●100% | 451,17191% | 393,02980% | 383,31078% |
| 64 | 32 | 450,73494% | 481,861 ●100% | 369,21577% | 470,06698% |
| 64 | 64 | 466,85495% | 425,47187% | 360,12073% | 490,950 ●100% |
| 128 | 1 | 359,267 ●100% | 299,94483% | 288,37880% | 9,4802.6% |
| 128 | 2 | 413,894 ●100% | 329,78280% | 322,38478% | 34,2268.3% |
| 128 | 8 | 449,052 ●100% | 421,42794% | 361,39380% | 209,27647% |
| 128 | 16 | 460,141 ●100% | 424,98792% | 346,62175% | 251,31155% |
| 128 | 32 | 472,13998% | 438,93491% | 344,15272% | 479,848 ●100% |
| 128 | 64 | 460,02395% | 411,05884% | 335,17869% | 486,682 ●100% |
| 256 | 1 | 391,918 ●100% | 298,19176% | 255,52365% | 26,6956.8% |
| 256 | 2 | 394,455 ●100% | 292,51974% | 288,75373% | 64,94616% |
| 256 | 8 | 425,223 ●100% | 343,17981% | 350,01882% | 190,36045% |
| 256 | 16 | 420,895 ●100% | 383,61191% | 331,58679% | 259,10062% |
| 256 | 32 | 461,06694% | 380,55978% | 333,25268% | 489,136 ●100% |
| 256 | 64 | 437,07089% | 379,84977% | 335,95368% | 493,414 ●100% |
| 512 | 1 | 369,770 ●100% | 285,16677% | 242,04165% | 33,1189.0% |
| 512 | 2 | 391,401 ●100% | 299,99877% | 277,04171% | 80,92621% |
| 512 | 8 | 399,783 ●100% | 386,28097% | 344,26386% | 220,15555% |
| 512 | 16 | 407,005 ●100% | 350,84286% | 345,28985% | 380,18993% |
| 512 | 32 | 423,36688% | 381,31279% | 345,59972% | 481,730 ●100% |
| 512 | 64 | 400,45784% | 350,42974% | 331,48970% | 476,193 ●100% |
failed requests summed over m
| conns | per socket | h3x | h3x spc | h2o-httpclient | h2load |
|---|---|---|---|---|---|
| 64 | 2 | 0 | 0 | 8 | 0 |
| 128 | 4 | 0 | 0 | 5 | 0 |
| 256 | 8 | 0 | 0 | 18 | 0 |
| 512 | 16 | 15,552 | 2,816 | 97 | 0 |
per-cell breakdown, all 24 cells
| conns | m | h3x | h3x spc | h2o-httpclient | h2load |
|---|---|---|---|---|---|
| 64 | 1 | 0 | 0 | 2 | 0 |
| 64 | 2 | 0 | 0 | 2 | 0 |
| 64 | 8 | 0 | 0 | 1 | 0 |
| 64 | 16 | 0 | 0 | 0 | 0 |
| 64 | 32 | 0 | 0 | 1 | 0 |
| 64 | 64 | 0 | 0 | 2 | 0 |
| 128 | 1 | 0 | 0 | 0 | 0 |
| 128 | 2 | 0 | 0 | 1 | 0 |
| 128 | 8 | 0 | 0 | 1 | 0 |
| 128 | 16 | 0 | 0 | 0 | 0 |
| 128 | 32 | 0 | 0 | 2 | 0 |
| 128 | 64 | 0 | 0 | 1 | 0 |
| 256 | 1 | 0 | 0 | 2 | 0 |
| 256 | 2 | 0 | 0 | 6 | 0 |
| 256 | 8 | 0 | 0 | 1 | 0 |
| 256 | 16 | 0 | 0 | 5 | 0 |
| 256 | 32 | 0 | 0 | 3 | 0 |
| 256 | 64 | 0 | 0 | 1 | 0 |
| 512 | 1 | 0 | 0 | 7 | 0 |
| 512 | 2 | 0 | 0 | 5 | 0 |
| 512 | 8 | 0 | 0 | 10 | 0 |
| 512 | 16 | 288 | 0 | 4 | 0 |
| 512 | 32 | 5,056 | 0 | 3 | 0 |
| 512 | 64 | 10,208 | 2,816 | 68 | 0 |
server UDP datagrams dropped receive-buffer overflow, summed over m
| conns | per socket | h3x | h3x spc | h2o-httpclient | h2load |
|---|---|---|---|---|---|
| 64 | 2 | 222,538 | 379,163 | 279 | 171,240 |
| 128 | 4 | 534,599 | 1,223,558 | 15,190 | 624,778 |
| 256 | 8 | 909,086 | 1,922,044 | 924,611 | 1,544,397 |
| 512 | 16 | 1,814,035 | 3,068,343 | 2,882,954 | 2,875,339 |
per-cell breakdown, all 24 cells
| conns | m | h3x | h3x spc | h2o-httpclient | h2load |
|---|---|---|---|---|---|
| 64 | 1 | 0 | 0 | 0 | 0 |
| 64 | 2 | 0 | 0 | 0 | 0 |
| 64 | 8 | 0 | 1,599 | 0 | 0 |
| 64 | 16 | 18,629 | 68,368 | 0 | 2,139 |
| 64 | 32 | 191,234 | 127,622 | 0 | 132,406 |
| 64 | 64 | 12,675 | 181,574 | 279 | 36,695 |
| 128 | 1 | 0 | 0 | 0 | 0 |
| 128 | 2 | 0 | 0 | 0 | 0 |
| 128 | 8 | 47,561 | 165,496 | 15 | 373 |
| 128 | 16 | 25,772 | 318,437 | 5,123 | 12,327 |
| 128 | 32 | 183,198 | 367,615 | 2,343 | 313,260 |
| 128 | 64 | 278,068 | 372,010 | 7,709 | 298,818 |
| 256 | 1 | 0 | 16 | 0 | 0 |
| 256 | 2 | 19,858 | 190,339 | 0 | 0 |
| 256 | 8 | 97,134 | 514,651 | 86,037 | 4,746 |
| 256 | 16 | 249,023 | 463,533 | 272,558 | 73,372 |
| 256 | 32 | 257,924 | 374,837 | 275,577 | 700,589 |
| 256 | 64 | 285,147 | 378,668 | 290,439 | 765,690 |
| 512 | 1 | 80,592 | 216,537 | 0 | 0 |
| 512 | 2 | 204,317 | 618,390 | 29 | 22 |
| 512 | 8 | 296,816 | 726,153 | 512,779 | 31,870 |
| 512 | 16 | 348,697 | 484,521 | 770,437 | 538,754 |
| 512 | 32 | 421,085 | 515,175 | 803,620 | 1,065,063 |
| 512 | 64 | 462,528 | 507,567 | 796,089 | 1,239,630 |
Methodology & caveats
h2load failed 0 requests in all nine grids, on every server at every payload. Its collapse at
m=1-2 happens only against h2o and only at 1 KB (5k-14k req/s, recovering to 3.39M by m=64): a
pairing-specific interaction between ngtcp2 and the quicly server at low concurrency, not general
client behaviour, and it does not reappear at the larger payloads. h2o-httpclient's failures stay
small throughout (128-319 per grid, process-teardown artifacts) except against nginx at 128 KB
(2,773). --socket-per-conn costs h3x throughput at 1 KB (+16.1% against h2o, +25.0%
against nginx) and buys reliability everywhere; at 128 KB the two modes are within noise of each
other on throughput and spc still fails far less.
All 864 cells are one sitting (2026-07-29, bench/grid-v3.sh), same box, same servers, so
columns and payloads are directly comparable. Each payload is served at its own URL
(/, /20k.html, /128k.html) rather than by swapping one
file's contents, and every cell re-derives the served object size from the client's own byte
counters. A round of earlier sweeps here was invalidated by silently benchmarking a 20 KB object
against 1 KB tables; naming the payload in the request makes that failure structurally impossible
rather than merely checked for.
Four cells at 128 KB drift 2-3% from the nominal object size and are reported by the generator rather than hidden. Three are nginx h3x cells failing 11-23% of requests, where bytes from partial bodies count toward throughput while their requests count as failed, inflating bytes-per-completed. The fourth is haproxy h2load at 512x64, where 32,768 streams were still in flight at the 10 s cutoff, deflating it by the same mechanism in reverse. Both are artifacts of the measurement window. A real payload mix-up here would be 20x off, not 3%.
Single unpinned runs. One 10 s run per cell, no repeats. A control experiment across two grids, comparing cells whose configuration was byte-identical, saw swings from -17.6% to +13.5% with no consistent sign. So differences under roughly 20% are not distinguishable from run-to-run variance, and only the large effects should be read as real — which here means the payload-size reversal (h3x 21-1 to h2load 15-9 against h2o), the 4.6x nginx gap at 128 KB, the drop counts, and h2load's clean failure record. Separating anything smaller needs repeated samples per cell.
Why h3x (shared-socket mode) drops requests
The trigger is multiplexing many QUIC connections over one UDP socket per worker. Differential
tests against haproxy: one connection per socket is flawless (32 conns / 32 sockets: 7.19M requests,
0 drops, 0 re-dials; thread-count controlled separately with 4 conns / 4 sockets, also 0/0), while
8 connections per socket, same 32 connections and same multiplexing, drops 1,428 and re-dials
constantly (355k requests on resumed connections in 5 s, with no reconnect flag). Confirmed at scale by the spc column, which fails 0 requests against nginx and 1,904 against
haproxy versus shared-socket h3x's 3,914 and 53,372. Servers differ in how they tolerate multiple QUIC
connections sharing a 4-tuple: h2o accepts it (this is h2o's own client architecture), nginx churns
moderately, haproxy constantly. Every drop surfaces as "I/O error" at stream attach: a shared-socket
connection dies, whatever was in flight on it fails, the pool re-dials with the cached ticket, and
the run continues, so throughput barely dips while the drop counter grows. The close reason on the
wire is still unread (CONNECTION_CLOSE is encrypted; reading it needs SSLKEYLOGFILE wiring), and
against nginx there is additionally a smaller classic handshake-timeout tail under saturation (2,559
churn errors vs 489 timeouts in the probed cell). Fix: --socket-per-conn (one socket
per connection, the h2load model), the spc column.
vs nginx (OpenSSL-QUIC, docker, 12 workers) 20 KB object
What the four clients are, and what each table measures
- h3x
- this project, 32 worker threads, each multiplexing its share of the connections over one UDP socket (connections share the thread's 4-tuple, distinguished by QUIC connection ID)
- h3x spc
- the same binary with
--socket-per-conn: one UDP socket per connection, so every connection has a unique 4-tuple; cross-connection sendmmsg batching is off in this mode - h2o-httpclient
- h2o's reference client, single-threaded and single-connection; run here
as one process per connection (own socket) with
-C= m streams, stopped at 10 s, completions counted from status lines. At 256-512 connections that is 256-512 processes on 32 CPUs (scheduler oversubscribed, fork storm at cell start), so treat that column there as a floor, not a ceiling - h2load
- nghttp2's load tool on the ngtcp2 stack, one socket per connection
- req/s
- completed responses per second. This is the only table where the four clients are directly comparable, so it is the only one rank-coloured green-to-red across the row
- failed requests
- requests the client gave up on. Against h3x in shared-socket mode these are the 4-tuple churn described under "Why h3x drops requests"; the connection dies, whatever was in flight on it fails, and the pool re-dials
- server UDP drops
- datagrams the kernel discarded on the server's socket because
its receive buffer was full, read per cell from
/proc/net/udp. Every drop we measured was a receive-buffer overflow (the per-socket counter, snmpRcvbufErrorsand snmpInErrorswere exactly equal). These cost throughput, not correctness: QUIC retransmits them. All three servers run the stock 212992-byte buffer (net.core.rmem_max), so this column is a drain-rate difference, not a buffer-size one
throughput
| conns | m | h3x req/s | h3x spc req/s | h2o-httpclient req/s | h2load req/s |
|---|---|---|---|---|---|
| 64 | 1 | 155,55838% | 131,70732% | 139,17134% | 409,176 ●100% |
| 64 | 2 | 162,46327% | 131,38822% | 150,02925% | 604,494 ●100% |
| 64 | 8 | 181,73425% | 167,18723% | 158,91822% | 716,904 ●100% |
| 64 | 16 | 185,00524% | 172,95723% | 153,46720% | 760,968 ●100% |
| 64 | 32 | 170,39424% | 166,47023% | 150,29421% | 709,980 ●100% |
| 64 | 64 | 170,93725% | 145,37521% | 148,76322% | 676,841 ●100% |
| 128 | 1 | 163,18537% | 114,24526% | 143,01333% | 435,359 ●100% |
| 128 | 2 | 167,24130% | 113,79021% | 144,26126% | 553,072 ●100% |
| 128 | 8 | 181,20826% | 146,72921% | 157,07622% | 702,337 ●100% |
| 128 | 16 | 178,80027% | 149,16522% | 157,40423% | 673,472 ●100% |
| 128 | 32 | 157,80624% | 149,11423% | 155,31924% | 659,355 ●100% |
| 128 | 64 | 158,34625% | 137,74422% | 152,72924% | 632,363 ●100% |
| 256 | 1 | 169,27640% | 113,96027% | 140,20933% | 424,702 ●100% |
| 256 | 2 | 163,58532% | 115,86423% | 141,14328% | 508,250 ●100% |
| 256 | 8 | 156,53128% | 138,91525% | 152,95927% | 564,654 ●100% |
| 256 | 16 | 118,87219% | 149,23324% | 152,53025% | 619,790 ●100% |
| 256 | 32 | 92,70215% | 145,42723% | 156,82325% | 639,054 ●100% |
| 256 | 64 | 157,92626% | 157,92526% | 149,56925% | 601,263 ●100% |
| 512 | 1 | 162,98145% | 121,14134% | 136,77138% | 361,521 ●100% |
| 512 | 2 | 157,80234% | 127,25927% | 144,98731% | 465,419 ●100% |
| 512 | 8 | 125,11221% | 154,67626% | 159,93327% | 594,704 ●100% |
| 512 | 16 | 131,73721% | 151,30724% | 154,54724% | 632,579 ●100% |
| 512 | 32 | 111,32118% | 145,07924% | 156,57526% | 609,370 ●100% |
| 512 | 64 | 101,13818% | 131,52223% | 154,18327% | 573,272 ●100% |
failed requests summed over m
| conns | per socket | h3x | h3x spc | h2o-httpclient | h2load |
|---|---|---|---|---|---|
| 64 | 2 | 131 | 96 | 1 | 0 |
| 128 | 4 | 518 | 20 | 2 | 0 |
| 256 | 8 | 2,558 | 256 | 182 | 0 |
| 512 | 16 | 5,153 | 128 | 1,066 | 0 |
per-cell breakdown, all 24 cells
| conns | m | h3x | h3x spc | h2o-httpclient | h2load |
|---|---|---|---|---|---|
| 64 | 1 | 0 | 0 | 0 | 0 |
| 64 | 2 | 3 | 0 | 0 | 0 |
| 64 | 8 | 0 | 0 | 0 | 0 |
| 64 | 16 | 0 | 0 | 0 | 0 |
| 64 | 32 | 128 | 0 | 0 | 0 |
| 64 | 64 | 0 | 96 | 1 | 0 |
| 128 | 1 | 0 | 0 | 0 | 0 |
| 128 | 2 | 42 | 0 | 0 | 0 |
| 128 | 8 | 0 | 20 | 0 | 0 |
| 128 | 16 | 40 | 0 | 0 | 0 |
| 128 | 32 | 276 | 0 | 2 | 0 |
| 128 | 64 | 160 | 0 | 0 | 0 |
| 256 | 1 | 6 | 0 | 0 | 0 |
| 256 | 2 | 26 | 0 | 0 | 0 |
| 256 | 8 | 36 | 0 | 1 | 0 |
| 256 | 16 | 72 | 0 | 17 | 0 |
| 256 | 32 | 2,258 | 0 | 34 | 0 |
| 256 | 64 | 160 | 256 | 130 | 0 |
| 512 | 1 | 0 | 0 | 2 | 0 |
| 512 | 2 | 34 | 0 | 5 | 0 |
| 512 | 8 | 418 | 0 | 19 | 0 |
| 512 | 16 | 0 | 128 | 48 | 0 |
| 512 | 32 | 1,073 | 0 | 352 | 0 |
| 512 | 64 | 3,628 | 0 | 640 | 0 |
server UDP datagrams dropped receive-buffer overflow, summed over m
| conns | per socket | h3x | h3x spc | h2o-httpclient | h2load |
|---|---|---|---|---|---|
| 64 | 2 | 3,989,373 | 3,430,223 | 4,735,576 | 74,856 |
| 128 | 4 | 5,108,991 | 4,062,377 | 7,672,989 | 135,803 |
| 256 | 8 | 5,885,146 | 4,777,832 | 9,780,754 | 282,494 |
| 512 | 16 | 7,747,500 | 5,670,282 | 12,028,900 | 639,753 |
per-cell breakdown, all 24 cells
| conns | m | h3x | h3x spc | h2o-httpclient | h2load |
|---|---|---|---|---|---|
| 64 | 1 | 0 | 0 | 0 | 0 |
| 64 | 2 | 0 | 3,776 | 0 | 0 |
| 64 | 8 | 316,804 | 447,219 | 467,378 | 6,293 |
| 64 | 16 | 762,690 | 851,748 | 1,260,602 | 14,885 |
| 64 | 32 | 1,256,897 | 929,252 | 1,492,122 | 18,318 |
| 64 | 64 | 1,652,982 | 1,198,228 | 1,515,474 | 35,360 |
| 128 | 1 | 0 | 68,185 | 0 | 0 |
| 128 | 2 | 95,533 | 372,811 | 20,683 | 628 |
| 128 | 8 | 651,904 | 1,008,342 | 1,298,877 | 11,133 |
| 128 | 16 | 1,331,541 | 827,430 | 1,918,358 | 31,796 |
| 128 | 32 | 1,415,110 | 791,291 | 2,152,649 | 31,269 |
| 128 | 64 | 1,614,903 | 994,318 | 2,282,422 | 60,977 |
| 256 | 1 | 71,234 | 370,409 | 0 | 128 |
| 256 | 2 | 402,536 | 743,050 | 205,283 | 3,148 |
| 256 | 8 | 1,097,242 | 954,799 | 1,977,580 | 34,510 |
| 256 | 16 | 1,078,893 | 724,823 | 2,350,188 | 49,931 |
| 256 | 32 | 1,545,356 | 874,161 | 2,590,957 | 83,618 |
| 256 | 64 | 1,689,885 | 1,110,590 | 2,656,746 | 111,159 |
| 512 | 1 | 289,146 | 681,500 | 89,572 | 9,911 |
| 512 | 2 | 660,990 | 1,216,446 | 782,562 | 16,954 |
| 512 | 8 | 1,202,086 | 913,319 | 2,368,241 | 98,699 |
| 512 | 16 | 1,826,283 | 770,509 | 2,702,707 | 151,592 |
| 512 | 32 | 1,839,448 | 1,071,438 | 2,989,925 | 161,829 |
| 512 | 64 | 1,929,547 | 1,017,070 | 3,095,893 | 200,768 |
Methodology & caveats
h2load failed 0 requests in all nine grids, on every server at every payload. Its collapse at
m=1-2 happens only against h2o and only at 1 KB (5k-14k req/s, recovering to 3.39M by m=64): a
pairing-specific interaction between ngtcp2 and the quicly server at low concurrency, not general
client behaviour, and it does not reappear at the larger payloads. h2o-httpclient's failures stay
small throughout (128-319 per grid, process-teardown artifacts) except against nginx at 128 KB
(2,773). --socket-per-conn costs h3x throughput at 1 KB (+16.1% against h2o, +25.0%
against nginx) and buys reliability everywhere; at 128 KB the two modes are within noise of each
other on throughput and spc still fails far less.
All 864 cells are one sitting (2026-07-29, bench/grid-v3.sh), same box, same servers, so
columns and payloads are directly comparable. Each payload is served at its own URL
(/, /20k.html, /128k.html) rather than by swapping one
file's contents, and every cell re-derives the served object size from the client's own byte
counters. A round of earlier sweeps here was invalidated by silently benchmarking a 20 KB object
against 1 KB tables; naming the payload in the request makes that failure structurally impossible
rather than merely checked for.
Four cells at 128 KB drift 2-3% from the nominal object size and are reported by the generator rather than hidden. Three are nginx h3x cells failing 11-23% of requests, where bytes from partial bodies count toward throughput while their requests count as failed, inflating bytes-per-completed. The fourth is haproxy h2load at 512x64, where 32,768 streams were still in flight at the 10 s cutoff, deflating it by the same mechanism in reverse. Both are artifacts of the measurement window. A real payload mix-up here would be 20x off, not 3%.
Single unpinned runs. One 10 s run per cell, no repeats. A control experiment across two grids, comparing cells whose configuration was byte-identical, saw swings from -17.6% to +13.5% with no consistent sign. So differences under roughly 20% are not distinguishable from run-to-run variance, and only the large effects should be read as real — which here means the payload-size reversal (h3x 21-1 to h2load 15-9 against h2o), the 4.6x nginx gap at 128 KB, the drop counts, and h2load's clean failure record. Separating anything smaller needs repeated samples per cell.
Why h3x (shared-socket mode) drops requests
The trigger is multiplexing many QUIC connections over one UDP socket per worker. Differential
tests against haproxy: one connection per socket is flawless (32 conns / 32 sockets: 7.19M requests,
0 drops, 0 re-dials; thread-count controlled separately with 4 conns / 4 sockets, also 0/0), while
8 connections per socket, same 32 connections and same multiplexing, drops 1,428 and re-dials
constantly (355k requests on resumed connections in 5 s, with no reconnect flag). Confirmed at scale by the spc column, which fails 0 requests against nginx and 1,904 against
haproxy versus shared-socket h3x's 3,914 and 53,372. Servers differ in how they tolerate multiple QUIC
connections sharing a 4-tuple: h2o accepts it (this is h2o's own client architecture), nginx churns
moderately, haproxy constantly. Every drop surfaces as "I/O error" at stream attach: a shared-socket
connection dies, whatever was in flight on it fails, the pool re-dials with the cached ticket, and
the run continues, so throughput barely dips while the drop counter grows. The close reason on the
wire is still unread (CONNECTION_CLOSE is encrypted; reading it needs SSLKEYLOGFILE wiring), and
against nginx there is additionally a smaller classic handshake-timeout tail under saturation (2,559
churn errors vs 489 timeouts in the probed cell). Fix: --socket-per-conn (one socket
per connection, the h2load model), the spc column.
vs haproxy (HAProxy QUIC, docker, nbthread 12) 20 KB object
What the four clients are, and what each table measures
- h3x
- this project, 32 worker threads, each multiplexing its share of the connections over one UDP socket (connections share the thread's 4-tuple, distinguished by QUIC connection ID)
- h3x spc
- the same binary with
--socket-per-conn: one UDP socket per connection, so every connection has a unique 4-tuple; cross-connection sendmmsg batching is off in this mode - h2o-httpclient
- h2o's reference client, single-threaded and single-connection; run here
as one process per connection (own socket) with
-C= m streams, stopped at 10 s, completions counted from status lines. At 256-512 connections that is 256-512 processes on 32 CPUs (scheduler oversubscribed, fork storm at cell start), so treat that column there as a floor, not a ceiling - h2load
- nghttp2's load tool on the ngtcp2 stack, one socket per connection
- req/s
- completed responses per second. This is the only table where the four clients are directly comparable, so it is the only one rank-coloured green-to-red across the row
- failed requests
- requests the client gave up on. Against h3x in shared-socket mode these are the 4-tuple churn described under "Why h3x drops requests"; the connection dies, whatever was in flight on it fails, and the pool re-dials
- server UDP drops
- datagrams the kernel discarded on the server's socket because
its receive buffer was full, read per cell from
/proc/net/udp. Every drop we measured was a receive-buffer overflow (the per-socket counter, snmpRcvbufErrorsand snmpInErrorswere exactly equal). These cost throughput, not correctness: QUIC retransmits them. All three servers run the stock 212992-byte buffer (net.core.rmem_max), so this column is a drain-rate difference, not a buffer-size one
throughput
| conns | m | h3x req/s | h3x spc req/s | h2o-httpclient req/s | h2load req/s |
|---|---|---|---|---|---|
| 64 | 1 | 249,684100% | 245,09498% | 222,84189% | 250,896 ●100% |
| 64 | 2 | 271,10296% | 271,01796% | 230,32382% | 281,769 ●100% |
| 64 | 8 | 303,39494% | 318,53498% | 296,15192% | 323,531 ●100% |
| 64 | 16 | 359,337 ●100% | 345,91796% | 284,94579% | 330,85092% |
| 64 | 32 | 336,90396% | 350,871 ●100% | 269,44577% | 303,41586% |
| 64 | 64 | 315,25894% | 333,684 ●100% | 246,47774% | 274,54282% |
| 128 | 1 | 268,71297% | 257,70193% | 222,34080% | 277,845 ●100% |
| 128 | 2 | 264,05891% | 270,21993% | 233,75781% | 290,049 ●100% |
| 128 | 8 | 284,18387% | 291,22489% | 262,63480% | 327,034 ●100% |
| 128 | 16 | 305,03091% | 335,867 ●100% | 244,70273% | 316,09394% |
| 128 | 32 | 285,28690% | 316,355 ●100% | 238,42275% | 285,51890% |
| 128 | 64 | 277,75592% | 300,402 ●100% | 218,88773% | 258,76786% |
| 256 | 1 | 265,39198% | 258,18795% | 208,17277% | 271,854 ●100% |
| 256 | 2 | 256,10187% | 241,02882% | 218,09174% | 293,809 ●100% |
| 256 | 8 | 268,83688% | 273,30689% | 235,32277% | 306,904 ●100% |
| 256 | 16 | 276,45089% | 300,83197% | 228,19874% | 309,398 ●100% |
| 256 | 32 | 266,984 ●100% | 266,019100% | 201,93976% | 261,72998% |
| 256 | 64 | 251,09899% | 252,405 ●100% | 161,39564% | 189,74075% |
| 512 | 1 | 246,61993% | 246,14993% | 184,03170% | 264,445 ●100% |
| 512 | 2 | 238,01986% | 236,96286% | 193,81170% | 275,489 ●100% |
| 512 | 8 | 241,05486% | 261,48793% | 211,80175% | 280,563 ●100% |
| 512 | 16 | 247,43386% | 289,368 ●100% | 195,37568% | 250,67387% |
| 512 | 32 | 230,19595% | 242,358 ●100% | 176,55673% | 209,25386% |
| 512 | 64 | 207,99391% | 227,615 ●100% | 129,62557% | 149,50966% |
failed requests summed over m
| conns | per socket | h3x | h3x spc | h2o-httpclient | h2load |
|---|---|---|---|---|---|
| 64 | 2 | 702 | 0 | 2 | 0 |
| 128 | 4 | 3,492 | 0 | 44 | 0 |
| 256 | 8 | 9,068 | 18 | 72 | 0 |
| 512 | 16 | 10,130 | 884 | 23 | 0 |
per-cell breakdown, all 24 cells
| conns | m | h3x | h3x spc | h2o-httpclient | h2load |
|---|---|---|---|---|---|
| 64 | 1 | 6 | 0 | 0 | 0 |
| 64 | 2 | 12 | 0 | 1 | 0 |
| 64 | 8 | 36 | 0 | 0 | 0 |
| 64 | 16 | 72 | 0 | 0 | 0 |
| 64 | 32 | 192 | 0 | 1 | 0 |
| 64 | 64 | 384 | 0 | 0 | 0 |
| 128 | 1 | 26 | 0 | 0 | 0 |
| 128 | 2 | 28 | 0 | 2 | 0 |
| 128 | 8 | 144 | 0 | 3 | 0 |
| 128 | 16 | 276 | 0 | 2 | 0 |
| 128 | 32 | 738 | 0 | 36 | 0 |
| 128 | 64 | 2,280 | 0 | 1 | 0 |
| 256 | 1 | 52 | 0 | 2 | 0 |
| 256 | 2 | 126 | 18 | 5 | 0 |
| 256 | 8 | 369 | 0 | 1 | 0 |
| 256 | 16 | 1,296 | 0 | 0 | 0 |
| 256 | 32 | 1,035 | 0 | 0 | 0 |
| 256 | 64 | 6,190 | 0 | 64 | 0 |
| 512 | 1 | 81 | 0 | 7 | 0 |
| 512 | 2 | 120 | 204 | 4 | 0 |
| 512 | 8 | 1,113 | 136 | 10 | 0 |
| 512 | 16 | 1,200 | 0 | 0 | 0 |
| 512 | 32 | 3,462 | 544 | 2 | 0 |
| 512 | 64 | 4,154 | 0 | 0 | 0 |
server UDP datagrams dropped receive-buffer overflow, summed over m
| conns | per socket | h3x | h3x spc | h2o-httpclient | h2load |
|---|---|---|---|---|---|
| 64 | 2 | 0 | 0 | 0 | 0 |
| 128 | 4 | 0 | 0 | 0 | 0 |
| 256 | 8 | 0 | 0 | 0 | 0 |
| 512 | 16 | 0 | 1 | 0 | 57 |
per-cell breakdown, all 24 cells
| conns | m | h3x | h3x spc | h2o-httpclient | h2load |
|---|---|---|---|---|---|
| 64 | 1 | 0 | 0 | 0 | 0 |
| 64 | 2 | 0 | 0 | 0 | 0 |
| 64 | 8 | 0 | 0 | 0 | 0 |
| 64 | 16 | 0 | 0 | 0 | 0 |
| 64 | 32 | 0 | 0 | 0 | 0 |
| 64 | 64 | 0 | 0 | 0 | 0 |
| 128 | 1 | 0 | 0 | 0 | 0 |
| 128 | 2 | 0 | 0 | 0 | 0 |
| 128 | 8 | 0 | 0 | 0 | 0 |
| 128 | 16 | 0 | 0 | 0 | 0 |
| 128 | 32 | 0 | 0 | 0 | 0 |
| 128 | 64 | 0 | 0 | 0 | 0 |
| 256 | 1 | 0 | 0 | 0 | 0 |
| 256 | 2 | 0 | 0 | 0 | 0 |
| 256 | 8 | 0 | 0 | 0 | 0 |
| 256 | 16 | 0 | 0 | 0 | 0 |
| 256 | 32 | 0 | 0 | 0 | 0 |
| 256 | 64 | 0 | 0 | 0 | 0 |
| 512 | 1 | 0 | 0 | 0 | 14 |
| 512 | 2 | 0 | 1 | 0 | 0 |
| 512 | 8 | 0 | 0 | 0 | 0 |
| 512 | 16 | 0 | 0 | 0 | 0 |
| 512 | 32 | 0 | 0 | 0 | 0 |
| 512 | 64 | 0 | 0 | 0 | 43 |
Methodology & caveats
h2load failed 0 requests in all nine grids, on every server at every payload. Its collapse at
m=1-2 happens only against h2o and only at 1 KB (5k-14k req/s, recovering to 3.39M by m=64): a
pairing-specific interaction between ngtcp2 and the quicly server at low concurrency, not general
client behaviour, and it does not reappear at the larger payloads. h2o-httpclient's failures stay
small throughout (128-319 per grid, process-teardown artifacts) except against nginx at 128 KB
(2,773). --socket-per-conn costs h3x throughput at 1 KB (+16.1% against h2o, +25.0%
against nginx) and buys reliability everywhere; at 128 KB the two modes are within noise of each
other on throughput and spc still fails far less.
All 864 cells are one sitting (2026-07-29, bench/grid-v3.sh), same box, same servers, so
columns and payloads are directly comparable. Each payload is served at its own URL
(/, /20k.html, /128k.html) rather than by swapping one
file's contents, and every cell re-derives the served object size from the client's own byte
counters. A round of earlier sweeps here was invalidated by silently benchmarking a 20 KB object
against 1 KB tables; naming the payload in the request makes that failure structurally impossible
rather than merely checked for.
Four cells at 128 KB drift 2-3% from the nominal object size and are reported by the generator rather than hidden. Three are nginx h3x cells failing 11-23% of requests, where bytes from partial bodies count toward throughput while their requests count as failed, inflating bytes-per-completed. The fourth is haproxy h2load at 512x64, where 32,768 streams were still in flight at the 10 s cutoff, deflating it by the same mechanism in reverse. Both are artifacts of the measurement window. A real payload mix-up here would be 20x off, not 3%.
Single unpinned runs. One 10 s run per cell, no repeats. A control experiment across two grids, comparing cells whose configuration was byte-identical, saw swings from -17.6% to +13.5% with no consistent sign. So differences under roughly 20% are not distinguishable from run-to-run variance, and only the large effects should be read as real — which here means the payload-size reversal (h3x 21-1 to h2load 15-9 against h2o), the 4.6x nginx gap at 128 KB, the drop counts, and h2load's clean failure record. Separating anything smaller needs repeated samples per cell.
Why h3x (shared-socket mode) drops requests
The trigger is multiplexing many QUIC connections over one UDP socket per worker. Differential
tests against haproxy: one connection per socket is flawless (32 conns / 32 sockets: 7.19M requests,
0 drops, 0 re-dials; thread-count controlled separately with 4 conns / 4 sockets, also 0/0), while
8 connections per socket, same 32 connections and same multiplexing, drops 1,428 and re-dials
constantly (355k requests on resumed connections in 5 s, with no reconnect flag). Confirmed at scale by the spc column, which fails 0 requests against nginx and 1,904 against
haproxy versus shared-socket h3x's 3,914 and 53,372. Servers differ in how they tolerate multiple QUIC
connections sharing a 4-tuple: h2o accepts it (this is h2o's own client architecture), nginx churns
moderately, haproxy constantly. Every drop surfaces as "I/O error" at stream attach: a shared-socket
connection dies, whatever was in flight on it fails, the pool re-dials with the cached ticket, and
the run continues, so throughput barely dips while the drop counter grows. The close reason on the
wire is still unread (CONNECTION_CLOSE is encrypted; reading it needs SSLKEYLOGFILE wiring), and
against nginx there is additionally a smaller classic handshake-timeout tail under saturation (2,559
churn errors vs 489 timeouts in the probed cell). Fix: --socket-per-conn (one socket
per connection, the h2load model), the spc column.
vs h2o (quicly, GRO patch, 12 threads, native) 128 KB object
What the four clients are, and what each table measures
- h3x
- this project, 32 worker threads, each multiplexing its share of the connections over one UDP socket (connections share the thread's 4-tuple, distinguished by QUIC connection ID)
- h3x spc
- the same binary with
--socket-per-conn: one UDP socket per connection, so every connection has a unique 4-tuple; cross-connection sendmmsg batching is off in this mode - h2o-httpclient
- h2o's reference client, single-threaded and single-connection; run here
as one process per connection (own socket) with
-C= m streams, stopped at 10 s, completions counted from status lines. At 256-512 connections that is 256-512 processes on 32 CPUs (scheduler oversubscribed, fork storm at cell start), so treat that column there as a floor, not a ceiling - h2load
- nghttp2's load tool on the ngtcp2 stack, one socket per connection
- req/s
- completed responses per second. This is the only table where the four clients are directly comparable, so it is the only one rank-coloured green-to-red across the row
- failed requests
- requests the client gave up on. Against h3x in shared-socket mode these are the 4-tuple churn described under "Why h3x drops requests"; the connection dies, whatever was in flight on it fails, and the pool re-dials
- server UDP drops
- datagrams the kernel discarded on the server's socket because
its receive buffer was full, read per cell from
/proc/net/udp. Every drop we measured was a receive-buffer overflow (the per-socket counter, snmpRcvbufErrorsand snmpInErrorswere exactly equal). These cost throughput, not correctness: QUIC retransmits them. All three servers run the stock 212992-byte buffer (net.core.rmem_max), so this column is a drain-rate difference, not a buffer-size one
throughput
| conns | m | h3x req/s | h3x spc req/s | h2o-httpclient req/s | h2load req/s |
|---|---|---|---|---|---|
| 64 | 1 | 90,274 ●100% | 74,92983% | 76,94385% | 31,48435% |
| 64 | 2 | 83,750 ●100% | 75,35790% | 71,59185% | 53,38364% |
| 64 | 8 | 81,79590% | 78,60486% | 68,15675% | 91,271 ●100% |
| 64 | 16 | 79,20790% | 75,61086% | 69,72279% | 87,860 ●100% |
| 64 | 32 | 76,81886% | 76,96987% | 68,46377% | 88,873 ●100% |
| 64 | 64 | 77,92286% | 76,27284% | 70,22077% | 90,896 ●100% |
| 128 | 1 | 87,531 ●100% | 73,20384% | 75,35086% | 42,32748% |
| 128 | 2 | 76,713 ●100% | 74,83398% | 70,58392% | 41,40854% |
| 128 | 8 | 77,23386% | 76,56085% | 64,39672% | 89,720 ●100% |
| 128 | 16 | 78,18390% | 71,64582% | 62,31671% | 87,264 ●100% |
| 128 | 32 | 77,15688% | 75,13986% | 65,17474% | 87,503 ●100% |
| 128 | 64 | 72,328 ●100% | 72,066100% | 63,60388% | 62,94687% |
| 256 | 1 | 79,149 ●100% | 72,55692% | 70,26989% | 44,59656% |
| 256 | 2 | 76,944 ●100% | 70,77692% | 67,14287% | 48,64763% |
| 256 | 8 | 76,36787% | 71,50481% | 62,89672% | 87,760 ●100% |
| 256 | 16 | 78,39087% | 67,57575% | 62,79870% | 89,810 ●100% |
| 256 | 32 | 70,77379% | 71,74381% | 62,34170% | 89,045 ●100% |
| 256 | 64 | 77,96787% | 65,24173% | 62,79470% | 89,337 ●100% |
| 512 | 1 | 73,913 ●100% | 66,64290% | 69,62894% | 45,91962% |
| 512 | 2 | 73,548 ●100% | 66,02890% | 66,58791% | 58,87880% |
| 512 | 8 | 73,39083% | 68,54178% | 64,33873% | 87,931 ●100% |
| 512 | 16 | 66,76972% | 66,57172% | 64,68669% | 93,102 ●100% |
| 512 | 32 | 60,55470% | 60,40770% | 62,29072% | 86,060 ●100% |
| 512 | 64 | 65,31174% | 57,77765% | 61,97670% | 88,572 ●100% |
failed requests summed over m
| conns | per socket | h3x | h3x spc | h2o-httpclient | h2load |
|---|---|---|---|---|---|
| 64 | 2 | 0 | 0 | 1 | 0 |
| 128 | 4 | 0 | 372 | 7 | 0 |
| 256 | 8 | 0 | 0 | 5 | 0 |
| 512 | 16 | 17,774 | 2,902 | 306 | 0 |
per-cell breakdown, all 24 cells
| conns | m | h3x | h3x spc | h2o-httpclient | h2load |
|---|---|---|---|---|---|
| 64 | 1 | 0 | 0 | 0 | 0 |
| 64 | 2 | 0 | 0 | 1 | 0 |
| 64 | 8 | 0 | 0 | 0 | 0 |
| 64 | 16 | 0 | 0 | 0 | 0 |
| 64 | 32 | 0 | 0 | 0 | 0 |
| 64 | 64 | 0 | 0 | 0 | 0 |
| 128 | 1 | 0 | 0 | 3 | 0 |
| 128 | 2 | 0 | 0 | 1 | 0 |
| 128 | 8 | 0 | 0 | 1 | 0 |
| 128 | 16 | 0 | 0 | 0 | 0 |
| 128 | 32 | 0 | 116 | 1 | 0 |
| 128 | 64 | 0 | 256 | 1 | 0 |
| 256 | 1 | 0 | 0 | 1 | 0 |
| 256 | 2 | 0 | 0 | 0 | 0 |
| 256 | 8 | 0 | 0 | 2 | 0 |
| 256 | 16 | 0 | 0 | 1 | 0 |
| 256 | 32 | 0 | 0 | 1 | 0 |
| 256 | 64 | 0 | 0 | 0 | 0 |
| 512 | 1 | 0 | 0 | 2 | 0 |
| 512 | 2 | 0 | 0 | 4 | 0 |
| 512 | 8 | 144 | 0 | 8 | 0 |
| 512 | 16 | 4,496 | 136 | 100 | 0 |
| 512 | 32 | 3,616 | 0 | 64 | 0 |
| 512 | 64 | 9,518 | 2,766 | 128 | 0 |
server UDP datagrams dropped receive-buffer overflow, summed over m
| conns | per socket | h3x | h3x spc | h2o-httpclient | h2load |
|---|---|---|---|---|---|
| 64 | 2 | 852,248 | 1,468,144 | 2,449 | 441,036 |
| 128 | 4 | 1,640,036 | 2,354,994 | 108,028 | 984,692 |
| 256 | 8 | 3,002,871 | 3,167,511 | 2,068,645 | 4,206,645 |
| 512 | 16 | 3,933,678 | 4,002,390 | 5,256,972 | 7,064,975 |
per-cell breakdown, all 24 cells
| conns | m | h3x | h3x spc | h2o-httpclient | h2load |
|---|---|---|---|---|---|
| 64 | 1 | 0 | 2,880 | 0 | 0 |
| 64 | 2 | 7,708 | 96,891 | 1 | 562 |
| 64 | 8 | 25,605 | 148,952 | 653 | 70,311 |
| 64 | 16 | 121,666 | 257,242 | 888 | 143,314 |
| 64 | 32 | 541,253 | 503,776 | 85 | 162,392 |
| 64 | 64 | 156,016 | 458,403 | 822 | 64,457 |
| 128 | 1 | 1,274 | 100,242 | 0 | 0 |
| 128 | 2 | 39,828 | 353,699 | 85 | 7,879 |
| 128 | 8 | 440,254 | 419,023 | 10,265 | 171,516 |
| 128 | 16 | 89,538 | 466,685 | 17,318 | 211,909 |
| 128 | 32 | 100,456 | 446,122 | 28,854 | 225,775 |
| 128 | 64 | 968,686 | 569,223 | 51,506 | 367,613 |
| 256 | 1 | 49,658 | 257,226 | 38 | 35 |
| 256 | 2 | 160,422 | 614,489 | 148,146 | 59,484 |
| 256 | 8 | 451,369 | 545,421 | 495,207 | 962,172 |
| 256 | 16 | 418,730 | 506,169 | 462,069 | 1,171,973 |
| 256 | 32 | 1,106,811 | 614,290 | 475,982 | 997,359 |
| 256 | 64 | 815,881 | 629,916 | 487,203 | 1,015,622 |
| 512 | 1 | 143,288 | 513,310 | 131,467 | 8,986 |
| 512 | 2 | 252,588 | 495,703 | 670,371 | 282,276 |
| 512 | 8 | 681,397 | 571,687 | 1,079,018 | 1,393,228 |
| 512 | 16 | 763,548 | 652,826 | 1,090,218 | 1,476,084 |
| 512 | 32 | 857,140 | 808,604 | 1,186,167 | 2,018,750 |
| 512 | 64 | 1,235,717 | 960,260 | 1,099,731 | 1,885,651 |
Methodology & caveats
h2load failed 0 requests in all nine grids, on every server at every payload. Its collapse at
m=1-2 happens only against h2o and only at 1 KB (5k-14k req/s, recovering to 3.39M by m=64): a
pairing-specific interaction between ngtcp2 and the quicly server at low concurrency, not general
client behaviour, and it does not reappear at the larger payloads. h2o-httpclient's failures stay
small throughout (128-319 per grid, process-teardown artifacts) except against nginx at 128 KB
(2,773). --socket-per-conn costs h3x throughput at 1 KB (+16.1% against h2o, +25.0%
against nginx) and buys reliability everywhere; at 128 KB the two modes are within noise of each
other on throughput and spc still fails far less.
All 864 cells are one sitting (2026-07-29, bench/grid-v3.sh), same box, same servers, so
columns and payloads are directly comparable. Each payload is served at its own URL
(/, /20k.html, /128k.html) rather than by swapping one
file's contents, and every cell re-derives the served object size from the client's own byte
counters. A round of earlier sweeps here was invalidated by silently benchmarking a 20 KB object
against 1 KB tables; naming the payload in the request makes that failure structurally impossible
rather than merely checked for.
Four cells at 128 KB drift 2-3% from the nominal object size and are reported by the generator rather than hidden. Three are nginx h3x cells failing 11-23% of requests, where bytes from partial bodies count toward throughput while their requests count as failed, inflating bytes-per-completed. The fourth is haproxy h2load at 512x64, where 32,768 streams were still in flight at the 10 s cutoff, deflating it by the same mechanism in reverse. Both are artifacts of the measurement window. A real payload mix-up here would be 20x off, not 3%.
Single unpinned runs. One 10 s run per cell, no repeats. A control experiment across two grids, comparing cells whose configuration was byte-identical, saw swings from -17.6% to +13.5% with no consistent sign. So differences under roughly 20% are not distinguishable from run-to-run variance, and only the large effects should be read as real — which here means the payload-size reversal (h3x 21-1 to h2load 15-9 against h2o), the 4.6x nginx gap at 128 KB, the drop counts, and h2load's clean failure record. Separating anything smaller needs repeated samples per cell.
Why h3x (shared-socket mode) drops requests
The trigger is multiplexing many QUIC connections over one UDP socket per worker. Differential
tests against haproxy: one connection per socket is flawless (32 conns / 32 sockets: 7.19M requests,
0 drops, 0 re-dials; thread-count controlled separately with 4 conns / 4 sockets, also 0/0), while
8 connections per socket, same 32 connections and same multiplexing, drops 1,428 and re-dials
constantly (355k requests on resumed connections in 5 s, with no reconnect flag). Confirmed at scale by the spc column, which fails 0 requests against nginx and 1,904 against
haproxy versus shared-socket h3x's 3,914 and 53,372. Servers differ in how they tolerate multiple QUIC
connections sharing a 4-tuple: h2o accepts it (this is h2o's own client architecture), nginx churns
moderately, haproxy constantly. Every drop surfaces as "I/O error" at stream attach: a shared-socket
connection dies, whatever was in flight on it fails, the pool re-dials with the cached ticket, and
the run continues, so throughput barely dips while the drop counter grows. The close reason on the
wire is still unread (CONNECTION_CLOSE is encrypted; reading it needs SSLKEYLOGFILE wiring), and
against nginx there is additionally a smaller classic handshake-timeout tail under saturation (2,559
churn errors vs 489 timeouts in the probed cell). Fix: --socket-per-conn (one socket
per connection, the h2load model), the spc column.
vs nginx (OpenSSL-QUIC, docker, 12 workers) 128 KB object
What the four clients are, and what each table measures
- h3x
- this project, 32 worker threads, each multiplexing its share of the connections over one UDP socket (connections share the thread's 4-tuple, distinguished by QUIC connection ID)
- h3x spc
- the same binary with
--socket-per-conn: one UDP socket per connection, so every connection has a unique 4-tuple; cross-connection sendmmsg batching is off in this mode - h2o-httpclient
- h2o's reference client, single-threaded and single-connection; run here
as one process per connection (own socket) with
-C= m streams, stopped at 10 s, completions counted from status lines. At 256-512 connections that is 256-512 processes on 32 CPUs (scheduler oversubscribed, fork storm at cell start), so treat that column there as a floor, not a ceiling - h2load
- nghttp2's load tool on the ngtcp2 stack, one socket per connection
- req/s
- completed responses per second. This is the only table where the four clients are directly comparable, so it is the only one rank-coloured green-to-red across the row
- failed requests
- requests the client gave up on. Against h3x in shared-socket mode these are the 4-tuple churn described under "Why h3x drops requests"; the connection dies, whatever was in flight on it fails, and the pool re-dials
- server UDP drops
- datagrams the kernel discarded on the server's socket because
its receive buffer was full, read per cell from
/proc/net/udp. Every drop we measured was a receive-buffer overflow (the per-socket counter, snmpRcvbufErrorsand snmpInErrorswere exactly equal). These cost throughput, not correctness: QUIC retransmits them. All three servers run the stock 212992-byte buffer (net.core.rmem_max), so this column is a drain-rate difference, not a buffer-size one
throughput
| conns | m | h3x req/s | h3x spc req/s | h2o-httpclient req/s | h2load req/s |
|---|---|---|---|---|---|
| 64 | 1 | 28,80823% | 25,74321% | 29,20624% | 123,006 ●100% |
| 64 | 2 | 30,76222% | 27,38620% | 28,56621% | 136,923 ●100% |
| 64 | 8 | 31,00721% | 27,12919% | 27,56619% | 144,525 ●100% |
| 64 | 16 | 21,66316% | 25,50319% | 26,66620% | 134,782 ●100% |
| 64 | 32 | 8,0256.4% | 27,14222% | 27,71322% | 125,194 ●100% |
| 64 | 64 | 7,6616.3% | 20,24117% | 27,68923% | 121,496 ●100% |
| 128 | 1 | 31,17825% | 25,33620% | 28,42122% | 126,890 ●100% |
| 128 | 2 | 29,76721% | 26,11718% | 27,70419% | 142,262 ●100% |
| 128 | 8 | 27,61023% | 24,83021% | 28,56224% | 118,204 ●100% |
| 128 | 16 | 24,55821% | 23,15619% | 27,61423% | 119,575 ●100% |
| 128 | 32 | 28,42224% | 26,15222% | 27,50823% | 118,001 ●100% |
| 128 | 64 | 7,8777.2% | 7,3636.7% | 28,01326% | 109,579 ●100% |
| 256 | 1 | 29,72528% | 25,29324% | 27,95526% | 106,151 ●100% |
| 256 | 2 | 27,26324% | 25,93523% | 28,42625% | 112,042 ●100% |
| 256 | 8 | 23,60522% | 21,37420% | 28,29626% | 109,609 ●100% |
| 256 | 16 | 18,14716% | 27,00923% | 27,93224% | 115,797 ●100% |
| 256 | 32 | 7,4316.6% | 24,82822% | 28,79226% | 112,378 ●100% |
| 256 | 64 | 6,7256.3% | 7,5207.1% | 29,72628% | 106,404 ●100% |
| 512 | 1 | 29,77134% | 26,96130% | 28,88533% | 88,404 ●100% |
| 512 | 2 | 25,39925% | 26,96326% | 28,77328% | 102,474 ●100% |
| 512 | 8 | 19,28617% | 22,56020% | 27,53724% | 113,900 ●100% |
| 512 | 16 | 19,52818% | 23,74422% | 28,52026% | 109,851 ●100% |
| 512 | 32 | 7,0926.6% | 7,6997.2% | 29,22527% | 106,862 ●100% |
| 512 | 64 | 7,7747.6% | 14,76614% | 29,82729% | 102,614 ●100% |
failed requests summed over m
| conns | per socket | h3x | h3x spc | h2o-httpclient | h2load |
|---|---|---|---|---|---|
| 64 | 2 | 10,231 | 3,845 | 1 | 0 |
| 128 | 4 | 10,706 | 4,520 | 16 | 0 |
| 256 | 8 | 21,966 | 7,654 | 648 | 0 |
| 512 | 16 | 26,849 | 15,573 | 2,108 | 0 |
per-cell breakdown, all 24 cells
| conns | m | h3x | h3x spc | h2o-httpclient | h2load |
|---|---|---|---|---|---|
| 64 | 1 | 4 | 0 | 0 | 0 |
| 64 | 2 | 3 | 0 | 0 | 0 |
| 64 | 8 | 0 | 0 | 1 | 0 |
| 64 | 16 | 0 | 0 | 0 | 0 |
| 64 | 32 | 282 | 48 | 0 | 0 |
| 64 | 64 | 9,942 | 3,797 | 0 | 0 |
| 128 | 1 | 36 | 0 | 0 | 0 |
| 128 | 2 | 0 | 5 | 0 | 0 |
| 128 | 8 | 69 | 0 | 0 | 0 |
| 128 | 16 | 460 | 40 | 16 | 0 |
| 128 | 32 | 436 | 466 | 0 | 0 |
| 128 | 64 | 9,705 | 4,009 | 0 | 0 |
| 256 | 1 | 5 | 0 | 0 | 0 |
| 256 | 2 | 0 | 0 | 0 | 0 |
| 256 | 8 | 269 | 0 | 8 | 0 |
| 256 | 16 | 169 | 0 | 32 | 0 |
| 256 | 32 | 1,704 | 570 | 160 | 0 |
| 256 | 64 | 19,819 | 7,084 | 448 | 0 |
| 512 | 1 | 43 | 0 | 0 | 0 |
| 512 | 2 | 17 | 0 | 12 | 0 |
| 512 | 8 | 421 | 274 | 64 | 0 |
| 512 | 16 | 2,095 | 421 | 208 | 0 |
| 512 | 32 | 4,931 | 2,429 | 608 | 0 |
| 512 | 64 | 19,342 | 12,449 | 1,216 | 0 |
server UDP datagrams dropped receive-buffer overflow, summed over m
| conns | per socket | h3x | h3x spc | h2o-httpclient | h2load |
|---|---|---|---|---|---|
| 64 | 2 | 9,192,852 | 6,131,667 | 9,503,132 | 352,528 |
| 128 | 4 | 10,838,648 | 6,292,446 | 13,345,794 | 577,778 |
| 256 | 8 | 13,011,327 | 6,988,227 | 16,000,040 | 962,648 |
| 512 | 16 | 15,392,160 | 8,134,403 | 19,578,280 | 1,697,301 |
per-cell breakdown, all 24 cells
| conns | m | h3x | h3x spc | h2o-httpclient | h2load |
|---|---|---|---|---|---|
| 64 | 1 | 27,729 | 93,102 | 8 | 13 |
| 64 | 2 | 207,775 | 225,905 | 95,423 | 3,382 |
| 64 | 8 | 1,302,372 | 996,924 | 1,800,376 | 20,766 |
| 64 | 16 | 2,163,197 | 1,218,668 | 2,175,235 | 71,015 |
| 64 | 32 | 2,443,554 | 1,558,594 | 2,645,424 | 117,325 |
| 64 | 64 | 3,048,225 | 2,038,474 | 2,786,666 | 140,027 |
| 128 | 1 | 158,847 | 382,432 | 76,448 | 2,101 |
| 128 | 2 | 745,480 | 813,399 | 776,901 | 7,746 |
| 128 | 8 | 1,743,557 | 836,825 | 2,553,367 | 43,388 |
| 128 | 16 | 1,996,322 | 1,030,202 | 2,925,867 | 90,793 |
| 128 | 32 | 2,856,927 | 1,575,019 | 3,317,590 | 128,551 |
| 128 | 64 | 3,337,515 | 1,654,569 | 3,695,621 | 305,199 |
| 256 | 1 | 539,581 | 673,471 | 542,368 | 5,656 |
| 256 | 2 | 1,215,134 | 919,708 | 1,610,845 | 16,173 |
| 256 | 8 | 2,098,832 | 960,695 | 2,919,725 | 116,027 |
| 256 | 16 | 2,506,391 | 1,183,237 | 3,340,408 | 177,644 |
| 256 | 32 | 3,123,831 | 1,443,844 | 3,745,701 | 284,577 |
| 256 | 64 | 3,527,558 | 1,807,272 | 3,840,993 | 362,571 |
| 512 | 1 | 888,419 | 836,038 | 1,222,679 | 13,404 |
| 512 | 2 | 1,517,862 | 891,269 | 2,272,851 | 26,436 |
| 512 | 8 | 2,467,379 | 997,950 | 3,468,947 | 239,165 |
| 512 | 16 | 3,033,099 | 1,367,543 | 3,896,526 | 497,305 |
| 512 | 32 | 3,394,243 | 1,863,944 | 4,114,758 | 447,726 |
| 512 | 64 | 4,091,158 | 2,177,659 | 4,602,519 | 473,265 |
Methodology & caveats
h2load failed 0 requests in all nine grids, on every server at every payload. Its collapse at
m=1-2 happens only against h2o and only at 1 KB (5k-14k req/s, recovering to 3.39M by m=64): a
pairing-specific interaction between ngtcp2 and the quicly server at low concurrency, not general
client behaviour, and it does not reappear at the larger payloads. h2o-httpclient's failures stay
small throughout (128-319 per grid, process-teardown artifacts) except against nginx at 128 KB
(2,773). --socket-per-conn costs h3x throughput at 1 KB (+16.1% against h2o, +25.0%
against nginx) and buys reliability everywhere; at 128 KB the two modes are within noise of each
other on throughput and spc still fails far less.
All 864 cells are one sitting (2026-07-29, bench/grid-v3.sh), same box, same servers, so
columns and payloads are directly comparable. Each payload is served at its own URL
(/, /20k.html, /128k.html) rather than by swapping one
file's contents, and every cell re-derives the served object size from the client's own byte
counters. A round of earlier sweeps here was invalidated by silently benchmarking a 20 KB object
against 1 KB tables; naming the payload in the request makes that failure structurally impossible
rather than merely checked for.
Four cells at 128 KB drift 2-3% from the nominal object size and are reported by the generator rather than hidden. Three are nginx h3x cells failing 11-23% of requests, where bytes from partial bodies count toward throughput while their requests count as failed, inflating bytes-per-completed. The fourth is haproxy h2load at 512x64, where 32,768 streams were still in flight at the 10 s cutoff, deflating it by the same mechanism in reverse. Both are artifacts of the measurement window. A real payload mix-up here would be 20x off, not 3%.
Single unpinned runs. One 10 s run per cell, no repeats. A control experiment across two grids, comparing cells whose configuration was byte-identical, saw swings from -17.6% to +13.5% with no consistent sign. So differences under roughly 20% are not distinguishable from run-to-run variance, and only the large effects should be read as real — which here means the payload-size reversal (h3x 21-1 to h2load 15-9 against h2o), the 4.6x nginx gap at 128 KB, the drop counts, and h2load's clean failure record. Separating anything smaller needs repeated samples per cell.
Why h3x (shared-socket mode) drops requests
The trigger is multiplexing many QUIC connections over one UDP socket per worker. Differential
tests against haproxy: one connection per socket is flawless (32 conns / 32 sockets: 7.19M requests,
0 drops, 0 re-dials; thread-count controlled separately with 4 conns / 4 sockets, also 0/0), while
8 connections per socket, same 32 connections and same multiplexing, drops 1,428 and re-dials
constantly (355k requests on resumed connections in 5 s, with no reconnect flag). Confirmed at scale by the spc column, which fails 0 requests against nginx and 1,904 against
haproxy versus shared-socket h3x's 3,914 and 53,372. Servers differ in how they tolerate multiple QUIC
connections sharing a 4-tuple: h2o accepts it (this is h2o's own client architecture), nginx churns
moderately, haproxy constantly. Every drop surfaces as "I/O error" at stream attach: a shared-socket
connection dies, whatever was in flight on it fails, the pool re-dials with the cached ticket, and
the run continues, so throughput barely dips while the drop counter grows. The close reason on the
wire is still unread (CONNECTION_CLOSE is encrypted; reading it needs SSLKEYLOGFILE wiring), and
against nginx there is additionally a smaller classic handshake-timeout tail under saturation (2,559
churn errors vs 489 timeouts in the probed cell). Fix: --socket-per-conn (one socket
per connection, the h2load model), the spc column.
vs haproxy (HAProxy QUIC, docker, nbthread 12) 128 KB object
What the four clients are, and what each table measures
- h3x
- this project, 32 worker threads, each multiplexing its share of the connections over one UDP socket (connections share the thread's 4-tuple, distinguished by QUIC connection ID)
- h3x spc
- the same binary with
--socket-per-conn: one UDP socket per connection, so every connection has a unique 4-tuple; cross-connection sendmmsg batching is off in this mode - h2o-httpclient
- h2o's reference client, single-threaded and single-connection; run here
as one process per connection (own socket) with
-C= m streams, stopped at 10 s, completions counted from status lines. At 256-512 connections that is 256-512 processes on 32 CPUs (scheduler oversubscribed, fork storm at cell start), so treat that column there as a floor, not a ceiling - h2load
- nghttp2's load tool on the ngtcp2 stack, one socket per connection
- req/s
- completed responses per second. This is the only table where the four clients are directly comparable, so it is the only one rank-coloured green-to-red across the row
- failed requests
- requests the client gave up on. Against h3x in shared-socket mode these are the 4-tuple churn described under "Why h3x drops requests"; the connection dies, whatever was in flight on it fails, and the pool re-dials
- server UDP drops
- datagrams the kernel discarded on the server's socket because
its receive buffer was full, read per cell from
/proc/net/udp. Every drop we measured was a receive-buffer overflow (the per-socket counter, snmpRcvbufErrorsand snmpInErrorswere exactly equal). These cost throughput, not correctness: QUIC retransmits them. All three servers run the stock 212992-byte buffer (net.core.rmem_max), so this column is a drain-rate difference, not a buffer-size one
throughput
| conns | m | h3x req/s | h3x spc req/s | h2o-httpclient req/s | h2load req/s |
|---|---|---|---|---|---|
| 64 | 1 | 80,39090% | 79,41089% | 69,65178% | 88,840 ●100% |
| 64 | 2 | 74,00087% | 80,93495% | 67,72679% | 85,281 ●100% |
| 64 | 8 | 73,89384% | 73,51583% | 68,59678% | 88,163 ●100% |
| 64 | 16 | 71,09685% | 74,94290% | 66,41680% | 83,169 ●100% |
| 64 | 32 | 74,19295% | 70,65991% | 65,01883% | 77,882 ●100% |
| 64 | 64 | 66,04485% | 69,15589% | 61,74380% | 77,344 ●100% |
| 128 | 1 | 66,65180% | 69,12783% | 63,33676% | 83,288 ●100% |
| 128 | 2 | 64,12377% | 65,28878% | 62,80775% | 83,432 ●100% |
| 128 | 8 | 64,24281% | 66,73985% | 64,78282% | 78,878 ●100% |
| 128 | 16 | 64,98183% | 67,68587% | 61,84079% | 78,209 ●100% |
| 128 | 32 | 62,15480% | 64,20783% | 59,10577% | 77,245 ●100% |
| 128 | 64 | 61,61090% | 64,03093% | 57,27783% | 68,729 ●100% |
| 256 | 1 | 60,70081% | 63,79285% | 58,52278% | 75,082 ●100% |
| 256 | 2 | 56,88973% | 63,33481% | 58,95776% | 77,812 ●100% |
| 256 | 8 | 58,75973% | 64,91481% | 62,38578% | 80,014 ●100% |
| 256 | 16 | 59,24778% | 64,32784% | 60,94280% | 76,173 ●100% |
| 256 | 32 | 60,43583% | 62,52386% | 57,59079% | 72,452 ●100% |
| 256 | 64 | 58,49188% | 60,72491% | 54,90683% | 66,541 ●100% |
| 512 | 1 | 57,31281% | 65,29392% | 55,95179% | 70,696 ●100% |
| 512 | 2 | 54,54075% | 60,13183% | 58,81081% | 72,719 ●100% |
| 512 | 8 | 57,77778% | 61,56383% | 58,32779% | 73,959 ●100% |
| 512 | 16 | 57,10377% | 61,48483% | 58,20878% | 74,433 ●100% |
| 512 | 32 | 56,28282% | 57,63484% | 56,48683% | 68,239 ●100% |
| 512 | 64 | 55,66086% | 59,45192% | 54,87685% | 64,788 ●100% |
failed requests summed over m
| conns | per socket | h3x | h3x spc | h2o-httpclient | h2load |
|---|---|---|---|---|---|
| 64 | 2 | 910 | 0 | 3 | 0 |
| 128 | 4 | 2,732 | 0 | 4 | 0 |
| 256 | 8 | 6,242 | 0 | 69 | 0 |
| 512 | 16 | 6,575 | 915 | 65 | 0 |
per-cell breakdown, all 24 cells
| conns | m | h3x | h3x spc | h2o-httpclient | h2load |
|---|---|---|---|---|---|
| 64 | 1 | 10 | 0 | 1 | 0 |
| 64 | 2 | 12 | 0 | 0 | 0 |
| 64 | 8 | 48 | 0 | 0 | 0 |
| 64 | 16 | 120 | 0 | 1 | 0 |
| 64 | 32 | 336 | 0 | 0 | 0 |
| 64 | 64 | 384 | 0 | 1 | 0 |
| 128 | 1 | 23 | 0 | 0 | 0 |
| 128 | 2 | 50 | 0 | 3 | 0 |
| 128 | 8 | 218 | 0 | 0 | 0 |
| 128 | 16 | 291 | 0 | 1 | 0 |
| 128 | 32 | 516 | 0 | 0 | 0 |
| 128 | 64 | 1,634 | 0 | 0 | 0 |
| 256 | 1 | 76 | 0 | 0 | 0 |
| 256 | 2 | 187 | 0 | 5 | 0 |
| 256 | 8 | 514 | 0 | 0 | 0 |
| 256 | 16 | 442 | 0 | 0 | 0 |
| 256 | 32 | 1,765 | 0 | 0 | 0 |
| 256 | 64 | 3,258 | 0 | 64 | 0 |
| 512 | 1 | 114 | 0 | 1 | 0 |
| 512 | 2 | 112 | 102 | 6 | 0 |
| 512 | 8 | 754 | 813 | 9 | 0 |
| 512 | 16 | 433 | 0 | 16 | 0 |
| 512 | 32 | 1,700 | 0 | 33 | 0 |
| 512 | 64 | 3,462 | 0 | 0 | 0 |
server UDP datagrams dropped receive-buffer overflow, summed over m
| conns | per socket | h3x | h3x spc | h2o-httpclient | h2load |
|---|---|---|---|---|---|
| 64 | 2 | 0 | 0 | 0 | 0 |
| 128 | 4 | 0 | 0 | 0 | 0 |
| 256 | 8 | 0 | 0 | 0 | 0 |
| 512 | 16 | 0 | 0 | 0 | 55 |
per-cell breakdown, all 24 cells
| conns | m | h3x | h3x spc | h2o-httpclient | h2load |
|---|---|---|---|---|---|
| 64 | 1 | 0 | 0 | 0 | 0 |
| 64 | 2 | 0 | 0 | 0 | 0 |
| 64 | 8 | 0 | 0 | 0 | 0 |
| 64 | 16 | 0 | 0 | 0 | 0 |
| 64 | 32 | 0 | 0 | 0 | 0 |
| 64 | 64 | 0 | 0 | 0 | 0 |
| 128 | 1 | 0 | 0 | 0 | 0 |
| 128 | 2 | 0 | 0 | 0 | 0 |
| 128 | 8 | 0 | 0 | 0 | 0 |
| 128 | 16 | 0 | 0 | 0 | 0 |
| 128 | 32 | 0 | 0 | 0 | 0 |
| 128 | 64 | 0 | 0 | 0 | 0 |
| 256 | 1 | 0 | 0 | 0 | 0 |
| 256 | 2 | 0 | 0 | 0 | 0 |
| 256 | 8 | 0 | 0 | 0 | 0 |
| 256 | 16 | 0 | 0 | 0 | 0 |
| 256 | 32 | 0 | 0 | 0 | 0 |
| 256 | 64 | 0 | 0 | 0 | 0 |
| 512 | 1 | 0 | 0 | 0 | 0 |
| 512 | 2 | 0 | 0 | 0 | 0 |
| 512 | 8 | 0 | 0 | 0 | 6 |
| 512 | 16 | 0 | 0 | 0 | 41 |
| 512 | 32 | 0 | 0 | 0 | 8 |
| 512 | 64 | 0 | 0 | 0 | 0 |
Methodology & caveats
h2load failed 0 requests in all nine grids, on every server at every payload. Its collapse at
m=1-2 happens only against h2o and only at 1 KB (5k-14k req/s, recovering to 3.39M by m=64): a
pairing-specific interaction between ngtcp2 and the quicly server at low concurrency, not general
client behaviour, and it does not reappear at the larger payloads. h2o-httpclient's failures stay
small throughout (128-319 per grid, process-teardown artifacts) except against nginx at 128 KB
(2,773). --socket-per-conn costs h3x throughput at 1 KB (+16.1% against h2o, +25.0%
against nginx) and buys reliability everywhere; at 128 KB the two modes are within noise of each
other on throughput and spc still fails far less.
All 864 cells are one sitting (2026-07-29, bench/grid-v3.sh), same box, same servers, so
columns and payloads are directly comparable. Each payload is served at its own URL
(/, /20k.html, /128k.html) rather than by swapping one
file's contents, and every cell re-derives the served object size from the client's own byte
counters. A round of earlier sweeps here was invalidated by silently benchmarking a 20 KB object
against 1 KB tables; naming the payload in the request makes that failure structurally impossible
rather than merely checked for.
Four cells at 128 KB drift 2-3% from the nominal object size and are reported by the generator rather than hidden. Three are nginx h3x cells failing 11-23% of requests, where bytes from partial bodies count toward throughput while their requests count as failed, inflating bytes-per-completed. The fourth is haproxy h2load at 512x64, where 32,768 streams were still in flight at the 10 s cutoff, deflating it by the same mechanism in reverse. Both are artifacts of the measurement window. A real payload mix-up here would be 20x off, not 3%.
Single unpinned runs. One 10 s run per cell, no repeats. A control experiment across two grids, comparing cells whose configuration was byte-identical, saw swings from -17.6% to +13.5% with no consistent sign. So differences under roughly 20% are not distinguishable from run-to-run variance, and only the large effects should be read as real — which here means the payload-size reversal (h3x 21-1 to h2load 15-9 against h2o), the 4.6x nginx gap at 128 KB, the drop counts, and h2load's clean failure record. Separating anything smaller needs repeated samples per cell.
Why h3x (shared-socket mode) drops requests
The trigger is multiplexing many QUIC connections over one UDP socket per worker. Differential
tests against haproxy: one connection per socket is flawless (32 conns / 32 sockets: 7.19M requests,
0 drops, 0 re-dials; thread-count controlled separately with 4 conns / 4 sockets, also 0/0), while
8 connections per socket, same 32 connections and same multiplexing, drops 1,428 and re-dials
constantly (355k requests on resumed connections in 5 s, with no reconnect flag). Confirmed at scale by the spc column, which fails 0 requests against nginx and 1,904 against
haproxy versus shared-socket h3x's 3,914 and 53,372. Servers differ in how they tolerate multiple QUIC
connections sharing a 4-tuple: h2o accepts it (this is h2o's own client architecture), nginx churns
moderately, haproxy constantly. Every drop surfaces as "I/O error" at stream attach: a shared-socket
connection dies, whatever was in flight on it fails, the pool re-dials with the cached ticket, and
the run continues, so throughput barely dips while the drop counter grows. The close reason on the
wire is still unread (CONNECTION_CLOSE is encrypted; reading it needs SSLKEYLOGFILE wiring), and
against nginx there is additionally a smaller classic handshake-timeout tail under saturation (2,559
churn errors vs 489 timeouts in the probed cell). Fix: --socket-per-conn (one socket
per connection, the h2load model), the spc column.
Key findings
Object size decides the winner, more than the server does. Every conclusion below was different when only the 1 KB grid existed. h3x leads 21 of 24 cells against h2o at 1 KB; at 128 KB h2load takes nginx and haproxy 24-0 and h2o 15-9. Read any single payload on its own and you will draw the wrong conclusion.
| cells won | h2o | nginx | haproxy | ||||||
|---|---|---|---|---|---|---|---|---|---|
| 1 KB | 20 KB | 128 KB | 1 KB | 20 KB | 128 KB | 1 KB | 20 KB | 128 KB | |
| h3x | 21 | 16 | 9 | 2 | 0 | 0 | 12 | 2 | 0 |
| h3x spc | 1 | 1 | 0 | 0 | 0 | 0 | 12 | 9 | 0 |
| h2o-httpclient | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| h2load | 1 | 7 | 15 | 22 | 24 | 24 | 0 | 13 | 24 |
- h3x's home ground
- small objects, few connections, high m. Its best result anywhere is 3.54M req/s at 1 KB, 64 connections, m=64 against h2o, 1.9x h2o's own reference client in that cell
- why it does not survive scale-up
- h3x's advantage comes from packing many small requests
into one QUIC packet (see
--send-batch). At 1 KB a batch of requests shares a datagram; at 128 KB a single response spans roughly 110 datagrams and there is nothing left to pack. The machinery that wins the 1 KB grid has no purchase on the 128 KB one - h2load's home ground
- everything else. It wins 22 of 24 nginx cells even at 1 KB, and by 128 KB it wins every cell on nginx and haproxy. Its margin on nginx at 128 KB is the largest in the whole matrix: 144,525 req/s against h3x's 31,178, a factor of 4.6
- the one h2load weakness
- a ngtcp2-x-quicly low-concurrency stall specific to the h2o pairing, and only at 1 KB: 5,151 req/s at m=1 where h3x does 878,718. It recovers to 3.39M by m=64, and the stall does not appear at 20 KB or 128 KB
- reliability
- h2load failed 0 requests in all nine grids — 216 cells, every server,
every payload. Nothing else on this page can say that. h3x's failures grow with payload against h2o
(0 to 15,552 to 17,774) and nginx (3,914 to 8,360 to 69,752), and shrink against haproxy (53,372 to
23,392 to 16,459).
--socket-per-conncuts them hard everywhere but no longer to zero: 31,592 against nginx at 128 KB - UDP drops are a separate axis and rank servers, not clients
- haproxy discards essentially nothing at any payload (0, 0, 0 for h3x across the three sizes) while failing the most requests. h2o and nginx discard millions while failing far fewer. On nginx, h2load causes an order of magnitude fewer server drops than h3x at every size (45,616 / 1.1M / 3.6M against 2.4M / 22.7M / 48.4M); on h2o that advantage inverts above 1 KB
- server capability
- best-client peak: at 1 KB h2o 3.54M > nginx 1.73M > haproxy 1.45M; at 128 KB nginx 144,525 > h2o 93,102 > haproxy 88,840 req/s. Even the server ranking flips with object size
Source layout & what it reuses from h2o
h3x is a thin load-generator shell over h2o's client stack. Everything hard (QUIC, TLS, HTTP/3,
the batched UDP I/O) is h2o library code; h3x adds only the load-generation logic on top. It links
one h2o library target, libh2o-evloop, which bundles quicly, picotls, and the HTTP/3
client.
- src/main.c
- entry point: CLI parsing, config validation, CPU-count detection (honoring Docker cpuset/quota), spawns and joins the worker threads
- src/worker.c
- per-thread setup and the event loop: builds the HTTP/3 context (quicly +
picotls), certificate verification, QUIC transport tuning, UDP socket(s) and connection pools; runs
the closed loop with connection-establishment pacing. Both socket modes (shared and
--socket-per-conn) live here - src/driver.c
- the per-request lifecycle: dispatches requests round-robin across the worker's connections, the on_connect → on_head → on_body callbacks that fill each request and consume its response, run-budget checks (count or duration), and graceful drain
- src/requests.c
- parser for
--requests.http files, turning method / path / headers / body templates into a round-robin request mix - src/tls.c
- session-resumption callbacks: the in-memory ticket/token cache that lets churned connections resume with 0-RTT
- src/stats.c
- per-request latency samples and the merged end-of-run summary (throughput, percentiles)
- src/h3x.h
- shared config / worker / request structs and the cross-file prototypes
Reused from h2o (all via libh2o-evloop): quicly for the QUIC transport;
picotls for TLS 1.3; h2o's lib/http3 for HTTP/3 framing and QPACK; its
httpclient.c / http3client.c client state machine, into which h3x's
callbacks plug; and its lib/common event loop (epoll), socket pool, timers, DNS, and
multithread queue. h3x carries one local patch to that library (see patches/): UDP GRO
on the receive path plus cross-connection sendmmsg batching on send. What h3x itself contributes is
only the shared-nothing worker threads, the closed-loop concurrency driver, connection churn with
in-memory 0-RTT resumption, the request-file parser, and merged latency stats.
Reproduce
Everything runs on one box over loopback. Build h3x and the h2o server, start the servers, then run the sweep. Full source and scripts are in the repo.
# build the client and the h2o server binary
git submodule update --init
git -C deps/h2o apply "$(pwd)/patches/h2o-udp-gro-send-batch.patch"
cmake -S . -B build && cmake --build build
# start the servers (h2o native on :14433; the rest are containers)
build/deps/h2o/h2o -c bench/h2o.conf &
docker start nginx-h3 haproxy-h3 # :14434 / :14435
# one cell, by hand: 512 connections x 8 streams for 10s against h2o, send-batch = half the
# worker target ((512/32) x 8 / 2 = 64)
build/h3x -k -t 32 --connections 512 -m 8 -d 10 --send-batch 64 https://127.0.0.1:14433/
# the whole grid on this page: all four clients x three servers x 24 cells, with per-cell UDP
# drop accounting. Resumable - re-running appends only the cells a log is missing and leaves
# finished grids alone. Writes bench/matrix-v3-.log. Takes about 75 minutes.
bash bench/grid-v3.sh
# a subset, for one cell or one column
CONNS=512 MS=64 bash bench/grid-v3.sh
python3 bench/gen-results.py # rebuilds this page from the logs
Every client here served the same 1 KB bench/doc_root/index.html. Each table cell is
one 10 s run; the raw per-run logs are the bench/matrix-v3-*.log files the page is
built from, one per server with all four client tags interleaved.