diff options
| author | Dmitry Ilvokhin <d@ilvokhin.com> | 2026-10-08 14:59:19 +0000 |
|---|---|---|
| committer | Dmitry Ilvokhin <d@ilvokhin.com> | 2026-10-08 17:14:01 +0000 |
| commit | e74880552a4f332e3b50967469bad80751e0b0b3 (patch) | |
| tree | 90bad823a842164f2c8a46d0fb157c695d387803 /.gitignore | |
| parent | 6ac7b5244639f93de3d1629ace99b549fe7eea52 (diff) | |
| download | benchmarks-e74880552a4f332e3b50967469bad80751e0b0b3.tar.gz benchmarks-e74880552a4f332e3b50967469bad80751e0b0b3.tar.bz2 benchmarks-e74880552a4f332e3b50967469bad80751e0b0b3.zip | |
I stumbled across a curious blog post by Erik Rigtorp and decided to
verify results myself.
https://rigtorp.se/ringbuffer/
On AMD Ryzen 7 8700G I got around following 50% speedup for cached
version with all member fields alignment.
$ bin/ring_buffer --benchmark_min_time=2s
------------------------------------------------------
Benchmark Time
------------------------------------------------------
BM_PushPopNoAlign/threads:2 7.66 ns
BM_PushPopCachelineAlign/threads:2 9.57 ns
BM_PushPopCachedNoAlign/threads:2 6.82 ns
BM_PushPopCachedCachelineAlign/threads:2 5.07 ns
As a side note, it is interesting to see writer/reader indexes alignment
for simple implementation is a regression, not an improvement on this
hardware. This fact is mentioned in Low Latency Trading Insights by
Henrique Bucher. The explanation from the book is following: in case
when L3 is shared between cores there is not much false sharing going
on, but there are more L3 fetches. This explanation sounds plausible
and I also was able to verify it experimentally.
$ perf stat \
-e l1-dcache-loads \
-e l2_cache_req_stat.dc_access_in_l2 \
-e ls_dmnd_fills_from_sys.local_ccx \
bin/ring_buffer \
--benchmark_min_time=100000000x \
--benchmark_filter=BM_PushPopNoAlign
------------------------------------------------------
Benchmark Time
------------------------------------------------------
BM_PushPopNoAlign/threads:2 7.47 ns
------------------------------------------------------
Value Counter
------------------------------------------------------
2707584731 l1-dcache-loads:u
119412724 l2_cache_req_stat.dc_access_in_l2:u
75608672 ls_dmnd_fills_from_sys.local_ccx:u
$ perf stat \
-e l1-dcache-loads \
-e l2_cache_req_stat.dc_access_in_l2 \
-e ls_dmnd_fills_from_sys.local_ccx \
bin/ring_buffer \
--benchmark_min_time=100000000x \
--benchmark_filter=BM_PushPopCachelineAlign
------------------------------------------------------
Benchmark Time
------------------------------------------------------
BM_PushPopCachelineAlign/threads:2 9.50 ns
------------------------------------------------------
Value Counter
------------------------------------------------------
4066582808 l1-dcache-loads:u
216776537 l2_cache_req_stat.dc_access_in_l2:u
117305061 ls_dmnd_fills_from_sys.local_ccx:u
------------------------------------------------------
Counter Delta
------------------------------------------------------
l1-dcache-loads +33.4%
l2_cache_req_stat.dc_access_in_l2 +44.9%
ls_dmnd_fills_from_sys.local_ccx +35.5%
Same hypothesis is also confirmed, when benchmark is pinned to the
cores. When threads are running on diferent physical cores there is not
a lot of difference in time.
$ taskset --cpu-list 0,1 \
bin/ring_buffer \
--benchmark_min_time=3s
------------------------------------------------------
Benchmark Time
------------------------------------------------------
BM_PushPopNoAlign/threads:2 7.76 ns
BM_PushPopCachelineAlign/threads:2 7.64 ns
But when threads are running on the same physical core (hyperhthreading)
there a noticable slowdown due more cache fetches from all levels.
$ taskset --cpu-list 0,8 \
bin/ring_buffer \
--benchmark_min_time=3s
------------------------------------------------------
Benchmark Time
------------------------------------------------------
BM_PushPopNoAlign/threads:2 3.42 ns
BM_PushPopCachelineAlign/threads:2 5.13 ns
On Apple M4 difference is much more noticeable.
------------------------------------------------------
Benchmark Time
------------------------------------------------------
BM_PushPopNoAlign/threads:2 50.7 ns
BM_PushPopCachelineAlign/threads:2 43.7 ns
BM_PushPopCachedNoAlign/threads:2 7.76 ns
BM_PushPopCachedCachelineAlign/threads:2 2.04 ns
Unfortunately, there are much less observability tools available for
macOS, so I was unable to dig deeper into results.
Diffstat (limited to '.gitignore')
0 files changed, 0 insertions, 0 deletions