summaryrefslogtreecommitdiff
path: root/.gitignore
diff options
context:
space:
mode:
authorDmitry Ilvokhin <d@ilvokhin.com>2026-10-08 14:59:19 +0000
committerDmitry Ilvokhin <d@ilvokhin.com>2026-10-08 17:14:01 +0000
commite74880552a4f332e3b50967469bad80751e0b0b3 (patch)
tree90bad823a842164f2c8a46d0fb157c695d387803 /.gitignore
parent6ac7b5244639f93de3d1629ace99b549fe7eea52 (diff)
downloadbenchmarks-e74880552a4f332e3b50967469bad80751e0b0b3.tar.gz
benchmarks-e74880552a4f332e3b50967469bad80751e0b0b3.tar.bz2
benchmarks-e74880552a4f332e3b50967469bad80751e0b0b3.zip
Add RingBuffer benchmarkHEADmaster
I stumbled across a curious blog post by Erik Rigtorp and decided to verify results myself. https://rigtorp.se/ringbuffer/ On AMD Ryzen 7 8700G I got around following 50% speedup for cached version with all member fields alignment. $ bin/ring_buffer --benchmark_min_time=2s ------------------------------------------------------ Benchmark Time ------------------------------------------------------ BM_PushPopNoAlign/threads:2 7.66 ns BM_PushPopCachelineAlign/threads:2 9.57 ns BM_PushPopCachedNoAlign/threads:2 6.82 ns BM_PushPopCachedCachelineAlign/threads:2 5.07 ns As a side note, it is interesting to see writer/reader indexes alignment for simple implementation is a regression, not an improvement on this hardware. This fact is mentioned in Low Latency Trading Insights by Henrique Bucher. The explanation from the book is following: in case when L3 is shared between cores there is not much false sharing going on, but there are more L3 fetches. This explanation sounds plausible and I also was able to verify it experimentally. $ perf stat \ -e l1-dcache-loads \ -e l2_cache_req_stat.dc_access_in_l2 \ -e ls_dmnd_fills_from_sys.local_ccx \ bin/ring_buffer \ --benchmark_min_time=100000000x \ --benchmark_filter=BM_PushPopNoAlign ------------------------------------------------------ Benchmark Time ------------------------------------------------------ BM_PushPopNoAlign/threads:2 7.47 ns ------------------------------------------------------ Value Counter ------------------------------------------------------ 2707584731 l1-dcache-loads:u 119412724 l2_cache_req_stat.dc_access_in_l2:u 75608672 ls_dmnd_fills_from_sys.local_ccx:u $ perf stat \ -e l1-dcache-loads \ -e l2_cache_req_stat.dc_access_in_l2 \ -e ls_dmnd_fills_from_sys.local_ccx \ bin/ring_buffer \ --benchmark_min_time=100000000x \ --benchmark_filter=BM_PushPopCachelineAlign ------------------------------------------------------ Benchmark Time ------------------------------------------------------ BM_PushPopCachelineAlign/threads:2 9.50 ns ------------------------------------------------------ Value Counter ------------------------------------------------------ 4066582808 l1-dcache-loads:u 216776537 l2_cache_req_stat.dc_access_in_l2:u 117305061 ls_dmnd_fills_from_sys.local_ccx:u ------------------------------------------------------ Counter Delta ------------------------------------------------------ l1-dcache-loads +33.4% l2_cache_req_stat.dc_access_in_l2 +44.9% ls_dmnd_fills_from_sys.local_ccx +35.5% Same hypothesis is also confirmed, when benchmark is pinned to the cores. When threads are running on diferent physical cores there is not a lot of difference in time. $ taskset --cpu-list 0,1 \ bin/ring_buffer \ --benchmark_min_time=3s ------------------------------------------------------ Benchmark Time ------------------------------------------------------ BM_PushPopNoAlign/threads:2 7.76 ns BM_PushPopCachelineAlign/threads:2 7.64 ns But when threads are running on the same physical core (hyperhthreading) there a noticable slowdown due more cache fetches from all levels. $ taskset --cpu-list 0,8 \ bin/ring_buffer \ --benchmark_min_time=3s ------------------------------------------------------ Benchmark Time ------------------------------------------------------ BM_PushPopNoAlign/threads:2 3.42 ns BM_PushPopCachelineAlign/threads:2 5.13 ns On Apple M4 difference is much more noticeable. ------------------------------------------------------ Benchmark Time ------------------------------------------------------ BM_PushPopNoAlign/threads:2 50.7 ns BM_PushPopCachelineAlign/threads:2 43.7 ns BM_PushPopCachedNoAlign/threads:2 7.76 ns BM_PushPopCachedCachelineAlign/threads:2 2.04 ns Unfortunately, there are much less observability tools available for macOS, so I was unable to dig deeper into results.
Diffstat (limited to '.gitignore')
0 files changed, 0 insertions, 0 deletions