<feed xmlns='http://www.w3.org/2005/Atom'>
<title>benchmarks.git, branch master</title>
<subtitle>Code for various micro benchmarks.</subtitle>
<id>https://git.ilvokhin.com/benchmarks.git/atom/?h=master</id>
<link rel='self' href='https://git.ilvokhin.com/benchmarks.git/atom/?h=master'/>
<link rel='alternate' type='text/html' href='https://git.ilvokhin.com/benchmarks.git/'/>
<updated>2026-10-08T17:14:01Z</updated>
<entry>
<title>Add RingBuffer benchmark</title>
<updated>2026-10-08T17:14:01Z</updated>
<author>
<name>Dmitry Ilvokhin</name>
<email>d@ilvokhin.com</email>
</author>
<published>2026-10-08T14:59:19Z</published>
<link rel='alternate' type='text/html' href='https://git.ilvokhin.com/benchmarks.git/commit/?id=e74880552a4f332e3b50967469bad80751e0b0b3'/>
<id>urn:sha1:e74880552a4f332e3b50967469bad80751e0b0b3</id>
<content type='text'>
I stumbled across a curious blog post by Erik Rigtorp and decided to
verify results myself.

    https://rigtorp.se/ringbuffer/

On AMD Ryzen 7 8700G I got around following 50% speedup for cached
version with all member fields alignment.

    $ bin/ring_buffer --benchmark_min_time=2s

------------------------------------------------------
Benchmark                                         Time
------------------------------------------------------
BM_PushPopNoAlign/threads:2                    7.66 ns
BM_PushPopCachelineAlign/threads:2             9.57 ns
BM_PushPopCachedNoAlign/threads:2              6.82 ns
BM_PushPopCachedCachelineAlign/threads:2       5.07 ns

As a side note, it is interesting to see writer/reader indexes alignment
for simple implementation is a regression, not an improvement on this
hardware. This fact is mentioned in Low Latency Trading Insights by
Henrique Bucher. The explanation from the book is following: in case
when L3 is shared between cores there is not much false sharing going
on, but there are more L3 fetches. This explanation sounds plausible
and I also was able to verify it experimentally.

    $ perf stat \
        -e l1-dcache-loads \
        -e l2_cache_req_stat.dc_access_in_l2 \
        -e ls_dmnd_fills_from_sys.local_ccx \
        bin/ring_buffer \
            --benchmark_min_time=100000000x \
            --benchmark_filter=BM_PushPopNoAlign

------------------------------------------------------
Benchmark                                         Time
------------------------------------------------------
BM_PushPopNoAlign/threads:2                    7.47 ns

------------------------------------------------------
Value                  Counter
------------------------------------------------------
2707584731     l1-dcache-loads:u
 119412724     l2_cache_req_stat.dc_access_in_l2:u
  75608672     ls_dmnd_fills_from_sys.local_ccx:u

    $ perf stat \
        -e l1-dcache-loads \
        -e l2_cache_req_stat.dc_access_in_l2 \
        -e ls_dmnd_fills_from_sys.local_ccx \
        bin/ring_buffer \
            --benchmark_min_time=100000000x \
            --benchmark_filter=BM_PushPopCachelineAlign

------------------------------------------------------
Benchmark                                         Time
------------------------------------------------------
BM_PushPopCachelineAlign/threads:2             9.50 ns

------------------------------------------------------
Value                  Counter
------------------------------------------------------
4066582808      l1-dcache-loads:u
 216776537      l2_cache_req_stat.dc_access_in_l2:u
 117305061      ls_dmnd_fills_from_sys.local_ccx:u

------------------------------------------------------
Counter                                          Delta
------------------------------------------------------
l1-dcache-loads                                 +33.4%
l2_cache_req_stat.dc_access_in_l2               +44.9%
ls_dmnd_fills_from_sys.local_ccx                +35.5%

Same hypothesis is also confirmed, when benchmark is pinned to the
cores. When threads are running on diferent physical cores there is not
a lot of difference in time.

    $ taskset --cpu-list 0,1 \
        bin/ring_buffer \
            --benchmark_min_time=3s

------------------------------------------------------
Benchmark                                         Time
------------------------------------------------------
BM_PushPopNoAlign/threads:2                    7.76 ns
BM_PushPopCachelineAlign/threads:2             7.64 ns

But when threads are running on the same physical core (hyperhthreading)
there a noticable slowdown due more cache fetches from all levels.

    $ taskset --cpu-list 0,8 \
        bin/ring_buffer \
            --benchmark_min_time=3s

------------------------------------------------------
Benchmark                                         Time
------------------------------------------------------
BM_PushPopNoAlign/threads:2                    3.42 ns
BM_PushPopCachelineAlign/threads:2             5.13 ns

On Apple M4 difference is much more noticeable.

------------------------------------------------------
Benchmark                                         Time
------------------------------------------------------
BM_PushPopNoAlign/threads:2                    50.7 ns
BM_PushPopCachelineAlign/threads:2             43.7 ns
BM_PushPopCachedNoAlign/threads:2              7.76 ns
BM_PushPopCachedCachelineAlign/threads:2       2.04 ns

Unfortunately, there are much less observability tools available for
macOS, so I was unable to dig deeper into results.
</content>
</entry>
<entry>
<title>Fix bin directory handling</title>
<updated>2026-10-08T12:46:25Z</updated>
<author>
<name>Dmitry Ilvokhin</name>
<email>d@ilvokhin.com</email>
</author>
<published>2026-10-08T12:46:25Z</published>
<link rel='alternate' type='text/html' href='https://git.ilvokhin.com/benchmarks.git/commit/?id=6ac7b5244639f93de3d1629ace99b549fe7eea52'/>
<id>urn:sha1:6ac7b5244639f93de3d1629ace99b549fe7eea52</id>
<content type='text'>
Every time bin modify time changes targets are rebuit. This is not how
it is supposed to work.

Specify bin as order-only prerequisite to avoid this problem.
</content>
</entry>
<entry>
<title>Initial commit</title>
<updated>2025-08-26T21:19:54Z</updated>
<author>
<name>Dmitry Ilvokhin</name>
<email>d@ilvokhin.com</email>
</author>
<published>2025-08-26T21:19:54Z</published>
<link rel='alternate' type='text/html' href='https://git.ilvokhin.com/benchmarks.git/commit/?id=d10b4c21779344acd57a60b059a941e4ebdf2b01'/>
<id>urn:sha1:d10b4c21779344acd57a60b059a941e4ebdf2b01</id>
<content type='text'>
Make stub files for C++ benchmarks.
</content>
</entry>
</feed>
