API reference — host#

OpenMP + SIMD host dispatch: packet-batched launchers with masked and flag-aware variants for the pure-C++ path.

eagle::cpu::Host#

struct Host#

Public Static Functions

template<std::size_t W, typename MaskT, typename KernelFunc>
static inline void dispatchMasked(const aether::PacketIndex<W> &pi, const MaskT &active, KernelFunc &&kernel)#

Dispatch active mask lanes to a scalar kernel.

template<std::size_t W, typename DataT>
static inline aether::simd::PacketMask<DataT, W> applyTail(aether::simd::PacketMask<DataT, W> mask, const aether::PacketIndex<W> &pi)#

Combine a mask with the tail mask for partial packets.

Templated on DataT like packetLoad/packetStore/loadMask &#8212; this was previously the one function in the file hard-typed to Real (double), so it could not accept a mask built over any other DataT (e.g. SoftDouble). DataT is deduced from mask rather than listed before W so every existing call site (applyTail<W>(...) in launchIfNot/launchIf/ packetLaunchIfNot below, and eagle/util/Slice.h’s applyTail<W>(PacketMask<Real,W>::allTrue(), pi)) keeps compiling unchanged &#8212; their single explicit template argument still binds to W, exactly as before. Zero behavior change for the wide-double production path.

template<typename KernelFunc>
static inline void launch(const idx_t &size, KernelFunc &&kernel, idx_t bytesPerSample = 0)#

Launcher wrapper — SIMD packet-based, processes W samples per iteration. Uses context-aware packetBatchedFor, which always passes parallel=false here — this call never opens its own #pragma omp parallel region:

  • If the caller is already inside an existing #pragma omp parallel region, work is shared across that region’s threads via #pragma omp for (aether backend/cpu/Tiled.h’s packetBatchedFor).

  • Otherwise it runs serially (SIMD-vectorised, single-threaded) — a caller expecting multithreading here without an enclosing region gets a silent full serialization: it compiles, produces correct values, and loses all threading.

Region-scoping is the caller’s responsibility. Consecutive launch/launchIf/launchIfNot/packetLaunch* calls made inside ONE such region have no implicit barrier between them — the work-share arm uses nowait (same aether citation above), so a later call’s tiles may start before an earlier call’s tiles have all finished on every thread; insert an explicit #pragma omp barrier between them if that ordering matters.

Production exemplar of correct use: a caller that opens a #pragma omp parallel region and calls Host::launch from inside it — that is what makes the call actually Host::launch call actually multithread. Do NOT copy docs/examples/02_host_dispatch.cpp (a standalone, region-free correctness demo) as a threading example.

Parameters:

bytesPerSample – Caller’s per-sample working-set estimate. 0 (default) falls back to aether’s compile-time DEFAULT_TILE_SIZE. A non-zero value drives runtime L2-derived tile sizing — pass the same value across every host kernel call in a step so each thread’s tile slice stays L2-resident across kernels.

template<typename KernelFunc>
static inline void launchIfNot(const idx_t &size, CRefArrT<bool> terminated, KernelFunc &&kernel, idx_t bytesPerSample = 0)#

Launcher wrapper, excluding items where the flag is true.

Loads W boolean flags per packet into a PacketMask and skips the entire packet when no lanes are active, avoiding per-sample branch mispredictions. bytesPerSample follows the same semantics as launch.

template<typename KernelFunc>
static inline void launchIf(const idx_t &size, const CRefArrT<bool> terminated, KernelFunc &&kernel, idx_t bytesPerSample = 0)#

Launcher wrapper, including only items where the flag is true.

Loads W boolean flags per packet into a PacketMask and skips the entire packet when no lanes are active. bytesPerSample follows the same semantics as launch.

template<typename KernelFunc>
static inline void packetLaunch(const idx_t &size, KernelFunc &&kernel, idx_t bytesPerSample = 0)#

Packet-aware launch — passes PacketIndex directly to kernel. The kernel receives a PacketIndex<W> and operates on W samples at once. bytesPerSample follows the same semantics as launch.

template<typename KernelFunc>
static inline void packetLaunchIfNot(const idx_t &size, CRefArrT<bool> terminated, KernelFunc &&kernel, idx_t bytesPerSample = 0)#

Packet-aware launch excluding terminated samples. Skips entire packets where all lanes are terminated. For mixed packets, the kernel receives the full PacketIndex and must handle masking internally (or use the provided mask). bytesPerSample follows the same semantics as launch.

Note

The low-level SIMD helpers packetLoad / packetStore / loadMask carry C++20 requires-constrained overloads that the Sphinx C++ domain cannot render; they are described in Module: cpu and documented inline in eagle/cpu/Host.h.