simd-loop reference: build scalar floor as -O2 -std=c++14, matching the other datasets

#4
by VanishD - opened

The scalar reference is two things at once: the code the agent is shown as its
starting point, and the floor a failed attempt falls back to when scoring
geomean@1. For those roles to mean the same thing in every dataset, the same
source has to be built the same way β€” and it was not.

ncnn       reference-scalar   -O2 -std=c++14
llama.cpp  reference-scalar   -O2 -std=c++14
kleidiai   reference-scalar   -O2 -std=c++14
simd-loop  reference          -O2 -std=c++14 -fno-vectorize -fno-slp-vectorize   <-- odd one out

NEON is mandatory in ARMv8-A, so clang needs no -march to use it: at -O2 the
loop vectorizer emits fmla v.4s from plain scalar C++ (verified on
clang++-18 / Graviton3 β€” 2 NEON instructions with the flags removed, 0 with
them). SVE still requires an explicit -march=...+sve, so nothing here gains
SVE. The effect of the two extra flags was therefore to leave simd-loop's floor
one vectorisation generation behind the other three datasets', which makes
cross-dataset speedups incomparable.

Measured on c7g.xlarge, same protocol, 28 simd-loop definitions: removing the
flags raises the floor on the 8 loops the compiler can actually vectorise
(loop_024 0.046 -> 0.622, loop_114 0.248 -> 1.742, loop_108 0.117 -> 0.421,
loop_035, loop_110, loop_002, loop_113, loop_109 between 2.1x and 3.4x) and
leaves the other 19 within 1% β€” those are sorts, reductions and gathers the
vectorizer cannot touch, so nothing is silently re-benchmarked that should not
be.

Only compile_flags changes; every other field is byte-identical.
scripts/gen_simd_loop_harness.py is updated to match in the code repo, so
regenerating will not reintroduce the flags.

ArtificialRay7579 changed pull request status to merged

Sign up or log in to comment