CPU SIMD backend#

Does aether vectorize on CPU too?

Yes: independent of any GPU, aether has its own CPU SIMD packet layer, aether::simd::Packet<DataT, Width>, which wraps the machine’s real vector registers (AVX-512, AVX2 or SSE2, whichever the build targets) and falls back to an ordinary scalar loop when none apply — the same source compiles either way.

Packet: one register, several lanes#

aether::simd::Packet (aether/backend/cpu/simd/Packet.h) is specialized per data type and register width — 8-wide double/16-wide float on AVX-512, 4-wide/8-wide on AVX2, 2-wide/4-wide on SSE2, and a 1-wide scalar fallback everywhere else — each backed by the matching platform intrinsic type. Every specialization offers the same small surface: zero()/broadcast(v)/load()/maskLoad(), store()/maskStore(), the arithmetic operators plus fmadd, and comparison-to-mask. aether::simd::PreferredWidth<T> names the widest packet the current build target actually supports for T, so code that wants “the best available width” rarely hard-codes a number. aether::simd::PacketMask carries the active-lane mask a partially-full TAIL packet (a sample count not evenly divisible by the packet width) needs, so a loop’s last, partial packet is handled the same way as every full one, not as a special case you write by hand.

Driving a loop: packetFor#

aether/backend/cpu/Tiled.h supplies the loop drivers built on top of Packet: aether::packetFor/aether::packetFlatFor walk a [begin, end) range packet by packet — context-aware, so the same call becomes an omp for nowait inside an existing parallel region, a fresh omp parallel for when asked explicitly, or a plain serial loop otherwise. aether::packetCapture materializes an expression into a register-resident packet item; aether::packetEval/aether:: packetEvalParallel assign an expression into a destination one packet at a time. aether::packetBatchedFor/aether::optimalTileSize add an L2-cache-aware tiled form on top of the same flat loop, sizing each tile from the destination’s L2 cache footprint rather than a fixed constant.

This is the CPU-side counterpart to the device-side, register-resident bundles Host vs device code introduces — different hardware, same idea: process several samples per unit of work, held in fast, on-chip storage rather than read one at a time from memory.

A runnable example#

Broadcasting one scalar across every lane of a packet at the build’s preferred double width:

TEST_F(PacketTest, BroadcastFillsEveryLane)
{
    constexpr std::size_t W = PreferredWidth<double>;
    auto p = Packet<double, W>::broadcast(3.5);
    double out[W];
    Packet<double, W>::store(out, p);
    for (std::size_t k = 0; k < W; ++k)
        EXPECT_EQ(out[k], 3.5);
}

Every one of the W lanes store() writes back out holds the same 3.5 — W itself is whatever this build’s PreferredWidth<double> resolves to, 1 on a build with no wider ISA available.