API reference — host#
OpenMP + SIMD host dispatch: packet-batched launchers with masked and flag-aware variants for the pure-C++ path.
eagle::cpu::Host#
-
struct Host#
Public Static Functions
-
template<std::size_t W, typename MaskT, typename KernelFunc>
static inline void dispatchMasked(const aether::PacketIndex<W> &pi, const MaskT &active, KernelFunc &&kernel)# Dispatch active mask lanes to a scalar kernel.
-
template<std::size_t W, typename DataT>
static inline aether::simd::PacketMask<DataT, W> applyTail(aether::simd::PacketMask<DataT, W> mask, const aether::PacketIndex<W> &pi)# Combine a mask with the tail mask for partial packets.
Templated on
DataTlikepacketLoad/packetStore/loadMask— this was previously the one function in the file hard-typed toReal(double), so it could not accept a mask built over any otherDataT(e.g. SoftDouble).DataTis deduced frommaskrather than listed beforeWso every existing call site (applyTail<W>(...)inlaunchIfNot/launchIf/packetLaunchIfNotbelow, andeagle/util/Slice.h’sapplyTail<W>(PacketMask<Real,W>::allTrue(), pi)) keeps compiling unchanged — their single explicit template argument still binds toW, exactly as before. Zero behavior change for the wide-doubleproduction path.
-
template<typename KernelFunc>
static inline void launch(const idx_t &size, KernelFunc &&kernel, idx_t bytesPerSample = 0)# Launcher wrapper — SIMD packet-based, processes W samples per iteration. Uses context-aware
packetBatchedFor, which always passesparallel=falsehere — this call never opens its own#pragma omp parallelregion:If the caller is already inside an existing
#pragma omp parallelregion, work is shared across that region’s threads via#pragma omp for(aetherbackend/cpu/Tiled.h’spacketBatchedFor).Otherwise it runs serially (SIMD-vectorised, single-threaded) — a caller expecting multithreading here without an enclosing region gets a silent full serialization: it compiles, produces correct values, and loses all threading.
Region-scoping is the caller’s responsibility. Consecutive
launch/launchIf/launchIfNot/packetLaunch*calls made inside ONE such region have no implicit barrier between them — the work-share arm usesnowait(same aether citation above), so a later call’s tiles may start before an earlier call’s tiles have all finished on every thread; insert an explicit#pragma omp barrierbetween them if that ordering matters.Production exemplar of correct use: a caller that opens a
#pragma omp parallelregion and callsHost::launchfrom inside it — that is what makes the call actuallyHost::launchcall actually multithread. Do NOT copydocs/examples/02_host_dispatch.cpp(a standalone, region-free correctness demo) as a threading example.- Parameters:
bytesPerSample – Caller’s per-sample working-set estimate.
0(default) falls back to aether’s compile-timeDEFAULT_TILE_SIZE. A non-zero value drives runtime L2-derived tile sizing — pass the same value across every host kernel call in a step so each thread’s tile slice stays L2-resident across kernels.
-
template<typename KernelFunc>
static inline void launchIfNot(const idx_t &size, CRefArrT<bool> terminated, KernelFunc &&kernel, idx_t bytesPerSample = 0)# Launcher wrapper, excluding items where the flag is true.
Loads W boolean flags per packet into a PacketMask and skips the entire packet when no lanes are active, avoiding per-sample branch mispredictions.
bytesPerSamplefollows the same semantics aslaunch.
-
template<typename KernelFunc>
static inline void launchIf(const idx_t &size, const CRefArrT<bool> terminated, KernelFunc &&kernel, idx_t bytesPerSample = 0)# Launcher wrapper, including only items where the flag is true.
Loads W boolean flags per packet into a PacketMask and skips the entire packet when no lanes are active.
bytesPerSamplefollows the same semantics aslaunch.
-
template<typename KernelFunc>
static inline void packetLaunch(const idx_t &size, KernelFunc &&kernel, idx_t bytesPerSample = 0)# Packet-aware launch — passes PacketIndex directly to kernel. The kernel receives a PacketIndex<W> and operates on W samples at once.
bytesPerSamplefollows the same semantics aslaunch.
-
template<typename KernelFunc>
static inline void packetLaunchIfNot(const idx_t &size, CRefArrT<bool> terminated, KernelFunc &&kernel, idx_t bytesPerSample = 0)# Packet-aware launch excluding terminated samples. Skips entire packets where all lanes are terminated. For mixed packets, the kernel receives the full PacketIndex and must handle masking internally (or use the provided mask).
bytesPerSamplefollows the same semantics aslaunch.
-
template<std::size_t W, typename MaskT, typename KernelFunc>
Note
The low-level SIMD helpers packetLoad / packetStore / loadMask
carry C++20 requires-constrained overloads that the Sphinx C++ domain
cannot render; they are described in Module: cpu and
documented inline in eagle/cpu/Host.h.