summaryrefslogtreecommitdiff
path: root/src
AgeCommit message (Collapse)Author
8 hoursAdd RingBuffer benchmarkHEADmasterDmitry Ilvokhin
I stumbled across a curious blog post by Erik Rigtorp and decided to verify results myself. https://rigtorp.se/ringbuffer/ On AMD Ryzen 7 8700G I got around following 50% speedup for cached version with all member fields alignment. $ bin/ring_buffer --benchmark_min_time=2s ------------------------------------------------------ Benchmark Time ------------------------------------------------------ BM_PushPopNoAlign/threads:2 7.66 ns BM_PushPopCachelineAlign/threads:2 9.57 ns BM_PushPopCachedNoAlign/threads:2 6.82 ns BM_PushPopCachedCachelineAlign/threads:2 5.07 ns As a side note, it is interesting to see writer/reader indexes alignment for simple implementation is a regression, not an improvement on this hardware. This fact is mentioned in Low Latency Trading Insights by Henrique Bucher. The explanation from the book is following: in case when L3 is shared between cores there is not much false sharing going on, but there are more L3 fetches. This explanation sounds plausible and I also was able to verify it experimentally. $ perf stat \ -e l1-dcache-loads \ -e l2_cache_req_stat.dc_access_in_l2 \ -e ls_dmnd_fills_from_sys.local_ccx \ bin/ring_buffer \ --benchmark_min_time=100000000x \ --benchmark_filter=BM_PushPopNoAlign ------------------------------------------------------ Benchmark Time ------------------------------------------------------ BM_PushPopNoAlign/threads:2 7.47 ns ------------------------------------------------------ Value Counter ------------------------------------------------------ 2707584731 l1-dcache-loads:u 119412724 l2_cache_req_stat.dc_access_in_l2:u 75608672 ls_dmnd_fills_from_sys.local_ccx:u $ perf stat \ -e l1-dcache-loads \ -e l2_cache_req_stat.dc_access_in_l2 \ -e ls_dmnd_fills_from_sys.local_ccx \ bin/ring_buffer \ --benchmark_min_time=100000000x \ --benchmark_filter=BM_PushPopCachelineAlign ------------------------------------------------------ Benchmark Time ------------------------------------------------------ BM_PushPopCachelineAlign/threads:2 9.50 ns ------------------------------------------------------ Value Counter ------------------------------------------------------ 4066582808 l1-dcache-loads:u 216776537 l2_cache_req_stat.dc_access_in_l2:u 117305061 ls_dmnd_fills_from_sys.local_ccx:u ------------------------------------------------------ Counter Delta ------------------------------------------------------ l1-dcache-loads +33.4% l2_cache_req_stat.dc_access_in_l2 +44.9% ls_dmnd_fills_from_sys.local_ccx +35.5% Same hypothesis is also confirmed, when benchmark is pinned to the cores. When threads are running on diferent physical cores there is not a lot of difference in time. $ taskset --cpu-list 0,1 \ bin/ring_buffer \ --benchmark_min_time=3s ------------------------------------------------------ Benchmark Time ------------------------------------------------------ BM_PushPopNoAlign/threads:2 7.76 ns BM_PushPopCachelineAlign/threads:2 7.64 ns But when threads are running on the same physical core (hyperhthreading) there a noticable slowdown due more cache fetches from all levels. $ taskset --cpu-list 0,8 \ bin/ring_buffer \ --benchmark_min_time=3s ------------------------------------------------------ Benchmark Time ------------------------------------------------------ BM_PushPopNoAlign/threads:2 3.42 ns BM_PushPopCachelineAlign/threads:2 5.13 ns On Apple M4 difference is much more noticeable. ------------------------------------------------------ Benchmark Time ------------------------------------------------------ BM_PushPopNoAlign/threads:2 50.7 ns BM_PushPopCachelineAlign/threads:2 43.7 ns BM_PushPopCachedNoAlign/threads:2 7.76 ns BM_PushPopCachedCachelineAlign/threads:2 2.04 ns Unfortunately, there are much less observability tools available for macOS, so I was unable to dig deeper into results.
2025-08-26Initial commitDmitry Ilvokhin
Make stub files for C++ benchmarks.