summaryrefslogtreecommitdiff
AgeCommit message (Collapse)Author
7 hoursAdd RingBuffer benchmarkHEADmasterDmitry Ilvokhin
I stumbled across a curious blog post by Erik Rigtorp and decided to verify results myself. https://rigtorp.se/ringbuffer/ On AMD Ryzen 7 8700G I got around following 50% speedup for cached version with all member fields alignment. $ bin/ring_buffer --benchmark_min_time=2s ------------------------------------------------------ Benchmark Time ------------------------------------------------------ BM_PushPopNoAlign/threads:2 7.66 ns BM_PushPopCachelineAlign/threads:2 9.57 ns BM_PushPopCachedNoAlign/threads:2 6.82 ns BM_PushPopCachedCachelineAlign/threads:2 5.07 ns As a side note, it is interesting to see writer/reader indexes alignment for simple implementation is a regression, not an improvement on this hardware. This fact is mentioned in Low Latency Trading Insights by Henrique Bucher. The explanation from the book is following: in case when L3 is shared between cores there is not much false sharing going on, but there are more L3 fetches. This explanation sounds plausible and I also was able to verify it experimentally. $ perf stat \ -e l1-dcache-loads \ -e l2_cache_req_stat.dc_access_in_l2 \ -e ls_dmnd_fills_from_sys.local_ccx \ bin/ring_buffer \ --benchmark_min_time=100000000x \ --benchmark_filter=BM_PushPopNoAlign ------------------------------------------------------ Benchmark Time ------------------------------------------------------ BM_PushPopNoAlign/threads:2 7.47 ns ------------------------------------------------------ Value Counter ------------------------------------------------------ 2707584731 l1-dcache-loads:u 119412724 l2_cache_req_stat.dc_access_in_l2:u 75608672 ls_dmnd_fills_from_sys.local_ccx:u $ perf stat \ -e l1-dcache-loads \ -e l2_cache_req_stat.dc_access_in_l2 \ -e ls_dmnd_fills_from_sys.local_ccx \ bin/ring_buffer \ --benchmark_min_time=100000000x \ --benchmark_filter=BM_PushPopCachelineAlign ------------------------------------------------------ Benchmark Time ------------------------------------------------------ BM_PushPopCachelineAlign/threads:2 9.50 ns ------------------------------------------------------ Value Counter ------------------------------------------------------ 4066582808 l1-dcache-loads:u 216776537 l2_cache_req_stat.dc_access_in_l2:u 117305061 ls_dmnd_fills_from_sys.local_ccx:u ------------------------------------------------------ Counter Delta ------------------------------------------------------ l1-dcache-loads +33.4% l2_cache_req_stat.dc_access_in_l2 +44.9% ls_dmnd_fills_from_sys.local_ccx +35.5% Same hypothesis is also confirmed, when benchmark is pinned to the cores. When threads are running on diferent physical cores there is not a lot of difference in time. $ taskset --cpu-list 0,1 \ bin/ring_buffer \ --benchmark_min_time=3s ------------------------------------------------------ Benchmark Time ------------------------------------------------------ BM_PushPopNoAlign/threads:2 7.76 ns BM_PushPopCachelineAlign/threads:2 7.64 ns But when threads are running on the same physical core (hyperhthreading) there a noticable slowdown due more cache fetches from all levels. $ taskset --cpu-list 0,8 \ bin/ring_buffer \ --benchmark_min_time=3s ------------------------------------------------------ Benchmark Time ------------------------------------------------------ BM_PushPopNoAlign/threads:2 3.42 ns BM_PushPopCachelineAlign/threads:2 5.13 ns On Apple M4 difference is much more noticeable. ------------------------------------------------------ Benchmark Time ------------------------------------------------------ BM_PushPopNoAlign/threads:2 50.7 ns BM_PushPopCachelineAlign/threads:2 43.7 ns BM_PushPopCachedNoAlign/threads:2 7.76 ns BM_PushPopCachedCachelineAlign/threads:2 2.04 ns Unfortunately, there are much less observability tools available for macOS, so I was unable to dig deeper into results.
11 hoursFix bin directory handlingDmitry Ilvokhin
Every time bin modify time changes targets are rebuit. This is not how it is supposed to work. Specify bin as order-only prerequisite to avoid this problem.
2025-08-26Initial commitDmitry Ilvokhin
Make stub files for C++ benchmarks.