1. Allocaotrs:

2. Servers:

ARM server: 16 cores, 60 GB ram.
Architecture:                aarch64
  CPU op-mode(s):            64-bit
  Byte Order:                Little Endian
CPU(s):                      16
  On-line CPU(s) list:       0-15
Vendor ID:                   ARM
  Model name:                Neoverse-V2
    BIOS Model name:         Axion  CPU @ 3.0GHz
    BIOS CPU family:         257
    Model:                   1
    Thread(s) per core:      1
    Core(s) per socket:      16
    Socket(s):               1
    Stepping:                r0p1
    BogoMIPS:                2000.00
    Flags:                   fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma lrcpc dcpop sha3 sm3 sm4 asimddp sha512 sve asimdfhm dit usca
                             t ilrcpc flagm sb paca pacg dcpodp sve2 sveaes svepmull svebitperm svesha3 svesm4 flagm2 frint svei8mm svebf16 i8mm bf16 dgh rng bti
Caches (sum of all):         
  L1d:                       1 MiB (16 instances)
  L1i:                       1 MiB (16 instances)
  L2:                        32 MiB (16 instances)
  L3:                        80 MiB (1 instance)
NUMA:                        
  NUMA node(s):              1
  NUMA node0 CPU(s):         0-15
Vulnerabilities:             
  Gather data sampling:      Not affected
  Indirect target selection: Not affected
  Itlb multihit:             Not affected
  L1tf:                      Not affected
  Mds:                       Not affected
  Meltdown:                  Not affected
  Mmio stale data:           Not affected
  Reg file data sampling:    Not affected
  Retbleed:                  Not affected
  Spec rstack overflow:      Not affected
  Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
  Spectre v1:                Mitigation; __user pointer sanitization
  Spectre v2:                Mitigation; CSV2, BHB
  Srbds:                     Not affected
  Tsa:                       Not affected
  Tsx async abort:           Not affected
      


x86 server: 16 cores, 60 GB ram.
Architecture:                x86_64
  CPU op-mode(s):            32-bit, 64-bit
  Address sizes:             52 bits physical, 57 bits virtual
  Byte Order:                Little Endian
CPU(s):                      16
  On-line CPU(s) list:       0-15
Vendor ID:                   GenuineIntel
  Model name:                INTEL(R) XEON(R) PLATINUM 8581C CPU @ 2.30GHz
    CPU family:              6
    Model:                   207
    Thread(s) per core:      2
    Core(s) per socket:      8
    Socket(s):               1
    Stepping:                2
    BogoMIPS:                4600.00
    Flags:                   fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush mmx fxsr sse sse2 ss ht syscall nx pdpe1gb rdtscp lm constant_tsc rep_g
                             ood nopl xtopology nonstop_tsc cpuid tsc_known_freq pni pclmulqdq monitor ssse3 fma cx16 pcid sse4_1 sse4_2 x2apic movbe popcnt aes xsave avx f16c rdra
                             nd hypervisor lahf_lm abm 3dnowprefetch ssbd ibrs ibpb stibp ibrs_enhanced fsgsbase tsc_adjust bmi1 hle avx2 smep bmi2 erms invpcid rtm avx512f avx512d
                             q rdseed adx smap avx512ifma clflushopt clwb avx512cd sha_ni avx512bw avx512vl xsaveopt xsavec xgetbv1 xsaves avx_vnni avx512_bf16 wbnoinvd arat avx512
                             vbmi umip avx512_vbmi2 gfni vaes vpclmulqdq avx512_vnni avx512_bitalg avx512_vpopcntdq la57 rdpid cldemote movdiri movdir64b fsrm md_clear serialize ts
                             xldtrk amx_bf16 avx512_fp16 amx_tile amx_int8 arch_capabilities
Virtualization features:     
  Hypervisor vendor:         KVM
  Virtualization type:       full
Caches (sum of all):         
  L1d:                       384 KiB (8 instances)
  L1i:                       256 KiB (8 instances)
  L2:                        16 MiB (8 instances)
  L3:                        260 MiB (1 instance)
NUMA:                        
  NUMA node(s):              1
  NUMA node0 CPU(s):         0-15
Vulnerabilities:             
  Gather data sampling:      Not affected
  Indirect target selection: Not affected
  Itlb multihit:             Not affected
  L1tf:                      Not affected
  Mds:                       Not affected
  Meltdown:                  Not affected
  Mmio stale data:           Not affected
  Reg file data sampling:    Not affected
  Retbleed:                  Not affected
  Spec rstack overflow:      Not affected
  Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
  Spectre v1:                Mitigation; usercopy/swapgs barriers and __user pointer sanitization
  Spectre v2:                Mitigation; Enhanced / Automatic IBRS; IBPB conditional; PBRSB-eIBRS SW sequence; BHI BHI_DIS_S
  Srbds:                     Not affected
  Tsa:                       Not affected
  Tsx async abort:           Not affected
  Vmscape:                   Not affected
      


3. Thread types:

  • Event loop thread: AKA FastThreadLocalThread in Netty.
  • Platform thread.

4. Threads count:

  • 32 threads, which equals to 2 × 16 cores, aligned with the default configuration commonly used in production env.

5. Java:

  • OpenJDK-21.0.12.1.

6. JVM args:

  • -XX:InitialRAMPercentage=40.0 -XX:MaxRAMPercentage=40.0
  • -Dio.netty.leakDetection.level=disabled
  • -dsa -da

7. Data:

  • SOCKET_PROXY: Generated by WebSocketProxyPattern.java, the data was copied from netty’s AllocationPatternSimulator, which derived from a web socket proxy service.
  • API_GATEWAY: Generated by ApiGatewayPattern.java, which generates sizes based on a log-normal distribution, which is commonly used to model network traffic patterns, the median size ≈ 2 KiB; the mean size ≈ 3.3 KiB; the P95 ≈ 10KiB; the P99 ≈ 20 KiB, the sizes are constrained to be between 8 bytes to 1 MiB.

8. Benchmark code:

9. Code base:

10. Max live buffers per thread:

  • MAX_LIVE_BUFFERS: [128, 1024, 4096, 8192, 16384, 32768, 65536].

11. Switch on read-write:

  • enableReadWrite: if true, enables read and write on each buffer to simulate real-world usage.