API reference — util

Contents

API reference — util#

Global typedefs, logging, error handling, observer-pattern primitives, array slicing, and OpenMP thread-count management.

Global typedefs#

namespace eagle#

Top-level eagle namespace.

Templated slice - class to act as a slice of a given reference object.

Templated observer for observer pattern - used to track move semantics with custom references.

Typedefs

using Real = double#

Scalar floating-point type used throughout eagle.

using idx_t = aether::idx_t#

Index type (from aether — aether::idx_t).

using mInt_t = int#

General-purpose signed integer type for metadata fields.

using uint_t = unsigned int#

Unsigned integer type.

using dims_t = aether::dims_t#

Dimension type (from aether).

using vecdim_t = aether::dims_t#

Vector-dimension type (alias for dims_t).

template<typename DT, idx_t VD>
using VecNTArray = aether::Array<DT, VD>#

aether SoA array with element type DT and vector extent VD.

aether merges a two-tier scalar::Array/vector::Array design into ONE Array<T, Es...> whose extents pack carries the static (vector) modes and whose trailing dynamic mode is the batch &#8212; so a fixed-width vector::Array<DT,VD> is spelled Array<DT, VD> and its scalar sibling is simply Array<DT>.

template<typename DT, idx_t VD>
using VecNT = aether::Item<DT, VD>#

aether register-resident item with element type DT and dimension VD (the per-item vector width).

using Vec3R = VecNT<Real, 3>#

3-D double-precision vector. Same type as aether::Vec3d (aether/aether.h) under a different name.

using Vec6R = VecNT<Real, 6>#

6-D double-precision vector (position + velocity).

using Vec3RArray = VecNTArray<Real, 3>#

SoA array of 3-D double-precision vectors.

using Vec6RArray = VecNTArray<Real, 6>#

SoA array of 6-D double-precision vectors.

template<typename ComponentT>
using GRefArrT = typename aether::Array<ComponentT>::ViewT#

Shorthand for the device-side view type of an aether scalar array.

aether collapses an earlier GRef/WRef/GRef::HandleT two-tier into the one aether::View (aether/view/View.h): the View IS the lightweight descriptor, and “global vs work reference” is now just which chunk the View was built over. Array<T>::ViewT is that View type.

template<typename ComponentT>
using CRefArrT = typename aether::Array<ComponentT>::ConstViewT#

Read-only counterpart of GRefArrT — Array<T>::ConstViewT.

aether has no implicit View<T> -> View<const T> conversion; call .as_const() on a mutable view to reach this type.

using nativeStream_t = aether::Stream#

Native stream handle threaded through eagle’s copy paths.

aether declares aether::Stream only in CUDA-backed builds (aether/chunk/Copy.h); the pure-C++ arm keeps the neutral integer tag eagle::cpu::StreamRef::native() already returns, so a “stream” argument threads through EAGLE_CPU_ONLY code unchanged.

Functions

inline aether::Device hostDevice()#

The aether Device eagle’s HOST-resident buffers live on — the same one aether::Array gives its own host chunk.

inline aether::Device deviceDevice()#

The aether Device eagle’s DEVICE-resident buffers live on. In an EAGLE_CPU_ONLY build there is only one copy, so this is hostDevice() — mirroring aether::Array’s own single-chunk AETHER_CPP_MODE arm.

template<typename T>
inline GRefArrT<T> spanView(T *data, idx_t n, aether::Device dev)#

Non-owning scalar View over a raw, unit-stride span of n elements &#8212; an earlier scalar::Array<T>::GRef(ptr, size) constructor.

aether’s make_view() is HOST-ONLY (it throws aether::Error on a bad span, and device code cannot throw — aether/view/MakeView.h), so a span view that must also be buildable inside a kernel is assembled from View’s public (data, mapping, device) constructor, exactly as aether::make_work_view does for shared memory.

template<std::size_t C, typename T>
inline auto packedSpanView(T *data, idx_t n, aether::Device dev)#

spanView over a component-packed (SoA) span: C components of n samples each, component c starting at data + c * n.

The pitch is n exactly — NOT aether::Array’s quantised capacity(). This is the shape a caller carving a C*n scratch slot chooses for itself, and aether::View::component<I>() reads the pitch back off the mapping, so the two never disagree.

template<typename T, std::size_t... Es>
inline aether::Array<T, Es...> makeArray(std::size_t n, const T &initVal = T{})#

Allocate an aether::Array of n samples with every element set to initVal.

An earlier Array(size, initVal = 0) ALWAYS filled (zero by default); aether’s Array(n) deliberately leaves the storage UNINITIALISED (it allocates a Chunk and nothing else). Every eagle site that relied on the earlier implicit zero-fill goes through here, so the fill is explicit and greppable rather than an inherited accident. Fills the HOST copy; call upload() to reach the device one, exactly as aether’s explicit-transfer contract requires.

template<typename T, std::size_t... Es>
inline aether::Array<T, Es...> cloneArray(const aether::Array<T, Es...> &src)#

Deep copy of an aether::Array.

aether::Array owns move-only Chunks, so it has no copy constructor and no clone(). Both arrays get the SAME quantised capacity() for the same samples(), so one flat span copy reproduces the pitch exactly. Copies the HOST copy and then upload()s, matching aether’s explicit-transfer contract (a device-only write not yet downloaded is NOT carried over — An earlier clone copied both sides).

Namespace eagle::util#

Includes the nested eagle::util::observable (the observer-pattern array) and eagle::util::slice (array slicing) namespaces — Breathe’s :members: already descends into them, so they are not repeated as separate directives below (doing so re-registered the same Doxygen anchor IDs twice and tripped docutils’ “Duplicate explicit target name”).

namespace util#

Enums

enum Level#

Logging severity levels, ordered from least to most verbose.

Values:

enumerator OFF#

Suppress all log output.

enumerator ERROR#

Error conditions; does not throw — use EAGLE_THROW for that.

enumerator WARN#

Warnings: unexpected but recoverable situations.

enumerator INFO#

Informational messages (default level).

enumerator DEBUG#

Detailed diagnostic output.

enumerator TRACE#

Per-step trace; very verbose.

enumerator ALL#

Alias for TRACE; enables everything.

Functions

inline void setLogLevel(const Level &level)#

Set the active log level.

Parameters:

level – Messages with severity above level are suppressed.

inline void disableLogging()#

Disable all log output (equivalent to setLogLevel(OFF)).

static std::string getLevelTag(const Level &level)#

Return an ANSI-coloured log-level tag string for terminal output.

Parameters:

level – The severity level to render.

Returns:

Bracketed, coloured tag string, e.g. "* [INFO ]".

template<typename ...Args>
void log(const Level &level, const std::string &msg, Args... args)#

Log a formatted message at the given severity level.

The message is suppressed if level > currentLogLevel. Formatting follows printf conventions; the format string is msg.

Template Parameters:

Args – Variadic pack of printf-compatible argument types.

Parameters:
  • level – Severity level of the message.

  • msg – printf-style format string.

  • args – Format arguments.

template<typename ...Args>
void log_fn(const Level &level, const std::string &func, const std::string &msg, Args... args)#

Log a formatted message prefixed with a function name.

Prepends [func] to msg and delegates to log(). Used internally by the EAGLE_INFO, EAGLE_DEBUG, … macros.

Template Parameters:

Args – Variadic pack of printf-compatible argument types.

Parameters:
  • level – Severity level of the message.

  • func – Name of the calling function (typically __func__).

  • msg – printf-style format string.

  • args – Format arguments.

Variables

Level currentLogLevel = INFO#

Active log level; messages above this level are suppressed.

class CaptureGuardState#
#include <CaptureGuard.h>

Per-thread capture-depth counter + deferred-destroy queue (STOP-THE-LINE incident fix).

The fix is never making the illegal call while a capture is in flight, not unswallowing its error. Every capture-opening call site (StreamCapturer, CaptureConditional) brackets itself with a :class:CaptureScope, which increments/decrements ONE process-wide depth counter here. Every CUDA-touching teardown site calls :meth:CaptureGuardState::destroyOrDefer instead of calling the driver directly: if the depth is 0 (no capture anywhere in the process), it runs immediately &#8212; byte-identical to today’s behavior. If the depth is > 0, the destroy is queued and runs the moment the OUTERMOST capture scope of that thread exits (CaptureScope’s destructor / explicit release()), plus a thread-exit backstop so a queue that never drains naturally (e.g. an aborted capture) does not leak forever.

The defect this closes

A CUDA-touching teardown (cudaGraphExecDestroy, cudaStreamDestroy, cudaGraphDestroy, cudaFree, …) that fires WHILE any stream on the process is actively being captured INVALIDATES that capture (cudaErrorStreamCaptureInvalidated) even when the destroyed handle has nothing to do with the capturing stream. Every such call site in this codebase already swallows its OWN error (EAGLE_CHECK_NOTHROW / a bare unchecked call) so the destroy itself never throws &#8212; correct, and by design (a throwing destructor unwinding during teardown would abort the process). But swallowing the ERROR does not undo the DAMAGE: the capture is poisoned, the NEXT operation on it fails, and the native state left behind can crash much later, at an unrelated GC or interpreter exit (the exact shape root-caused in the incident this fixes: a dead Launcher’s cudaGraphExecDestroy, run by Python’s garbage collector DURING a later, unrelated active capture, invalidated it).

Thread-safety

The capture depth and the deferred-destroy queue are PER THREAD (thread_local): captures are opened in cudaStreamCaptureModeThreadLocal, so a capture on thread A is not invalidated by a free, allocation or destroy on thread B, and B’s teardowns therefore run immediately. Only the capturing thread defers its own teardowns, and drains them when ITS outermost capture scope exits. A second top-level capture begun on a thread that already has one open is refused (:meth:requireIdleThread): a capture session is begun, filled and ended by one thread, one at a time per thread. Whole-device synchronisation by any thread invalidates every open capture (a CUDA rule, any mode); stream-level synchronisation is fine.

Note

Pure host-side bookkeeping &#8212; no CUDA type appears in this header, so it compiles identically in EAGLE_CPU_ONLY mode (there is no capture concept there; nothing calls into it).

Public Functions

inline void enter()#

Enter one capture scope. Call ONLY after the underlying cudaStream*Capture* call has SUCCEEDED (a failed begin never opened a capture, so it must not raise the depth).

inline void requireIdleThread() const#

Refuse a top-level capture on a thread that already has one open. Call BEFORE the underlying begin-capture call. A nested weave (CaptureConditional) is part of the open session and does not call this.

Throws:

std::logic_error – If this thread’s capture depth is > 0.

inline void exit()#

Exit one capture scope. Symmetric with :meth:enter &#8212; callers must not call this without a matching prior enter() (:class:CaptureScope enforces that pairing). Drains the deferred queue once the depth returns to 0.

inline void destroyOrDefer(std::function<void()> op)#

Run op now if no capture is active anywhere in the process (depth 0) &#8212; identical timing to today’s inline destroy. Otherwise enqueue it to run once the outermost capture scope exits (or at process exit, as a backstop). op must be the FULLY self-contained destroy call (its own error policy already decided at the call site, e.g. wrapped in EAGLE_CHECK_NOTHROW there) &#8212; this method does not add or remove error handling.

inline int depth() const#

Current capture depth (tests / diagnostics).

inline std::size_t pendingCount() const#

Number of destroys queued, not yet run (tests / diagnostics).

Public Static Functions

static inline CaptureGuardState &instance()#

The single process-wide instance (Meyers singleton &#8212; safe init order across translation units, destroyed at process exit, which is exactly when the backstop drain below needs to run).

class CaptureScope#
#include <CaptureGuard.h>

RAII capture-scope token: enter()s :class:CaptureGuardState on construction, exit()s on destruction or an explicit :meth:release.

Intended as an std::optional<CaptureScope> MEMBER of a capture-opening class (StreamCapturer, CaptureConditional): .emplace() it right after the underlying begin-capture call succeeds, .reset() it (calling :meth:release) right after the underlying end-capture call runs &#8212; regardless of whether that call reports success, since the capture SESSION is over either way. Exception-safe even if a caller never reaches an explicit release: the owning object’s OWN destructor still destroys this member (whenever that object itself is destroyed), so the depth is always eventually decremented.

Public Functions

inline void release()#

Exit the scope early. Idempotent &#8212; calling it again (or on a default-constructed-then-never-entered token, which never happens here since the constructor always enters) is a no-op.

template<typename ChildT>
struct Observable#
#include <Observable.h>

CRTP class for Observability traits.

Public Types

using Self = ChildT#

Declare observer type.

Public Functions

inline void registerObserver(ObserverT *obs)#

Register a new observer.

inline void deregisterObserver(ObserverT *obs)#

Deregister the given observer.

inline void notifyObservers()#

Notify our observers of a move.

inline std::vector<ObserverT*> &observers()#

Expose the observers.

inline void finalizeMove(Self &&other)#

Handle to perform a full move.

template<typename T>
class Observer#
#include <Observer.h>

Move semantics observer.

Public Functions

Observer() = default#

Default construction.

inline Observer(T &obj)#

Construct by directly observing an object.

inline Observer(T *obj)#

Construct with the given pointer to object (if not nullptr).

Observer(Observer &other) = delete#

Copy construction is forbidden.

inline Observer(Observer &&other)#

Move construction is allowed.

Observer &operator=(Observer &other) = delete#

Copy assignment is forbidden.

inline Observer &operator=(Observer &&other)#

Move assignment is allowed.

inline void observe(T &obj)#

Observe.

inline void onNotify(T *newobj)#

Update internal refernece on the move.

inline T *operator->()#

Expose the log.

inline ~Observer()#

Destructor calls de-registration.

inline Observer clone() const#

Create a clone of this observer.

template<typename SliceT>
class RefSlice#
#include <Slice.h>

Reference Slice.

There is no separate work/MaybeVolatile pair — see eagle::filtering::RefScanner for the reasoning. The index array is now a HELD aether::View rather than a base class: aether’s Array has no Ref<work, MaybeVolatile> family to inherit from, and a View is a value descriptor, so composition is the direct translation.

Subclassed by eagle::filtering::RefSlice< Self >

Public Types

using ParentT = GRefArrT<idx_t>#

The index view this slice reads its true indices from.

using T = typename SliceT::T::GRef#

Expose inherent sliced type — the sliced object’s own view.

using ViewT = typename T::reference#

Expose the accessed element reference types.

Public Functions

inline const ParentT &indexes() const#

Return the view that points to the indexes.

inline ParentT rawIndexes() const#

Return the view that points to the indexes (by value — a View copy is a ref-level clone()).

inline ViewT operator[](const SampleIndex &i)#

Access the slice elements.

inline idx_t vectorOffset() const#

Alias the core size as slice offset (for vector elements).

inline idx_t coreSize() const#

Expose the size of the inner object.

inline idx_t size() const#

return the size of the slice

inline RefSlice clone() const#

Clone this reference object.

inline void updateNumElements(const idx_t newn)#

Update the number of elements in this slice.

Public Members

T coredata_#

Data members made public for PODification.

Public Static Functions

static inline RefSlice make(const ParentT &idxs, const T &ref, const ParentT &sz)#

Factory method to construct from the index view and the view being sliced.

template<typename MyClass>
class Slice : public aether::Array<idx_t>#
#include <Slice.h>

Slice.

Subclassed by eagle::filtering::FilteringSlice< MyClass >

Public Types

using T = MyClass#

Expose inherently sliced type.

using GRef = RefSlice<Self>#

Expose the one non-owning reference tier.

Public Functions

Slice() = delete#

Default constructor is forbidden.

inline Slice(T *myobj = nullptr)#

Construct from the given object pointer.

Every index starts at the object size — deliberately out of range, the “not yet scattered” sentinel Array{sz, sz} sets.

inline Slice(T &myobj)#

Construct from the given object.

inline Slice(T &myobj, ParentT &&idxs, const idx_t sz)#

Construct from the given object and indexes and size.

inline Slice(ParentT &&idxs, CoreT &&coredata, ParentT &&sz)#

Move construct from data types.

Slice(Slice &other) = delete#

Copy constructor is forbidden.

inline Slice(Slice &&other)#

Move constructor.

Slice &operator=(Slice &other) = delete#

Copy assignment is forbidden.

inline Slice &operator=(Slice &&other)#

Move assignment operator.

inline idx_t coreSize() const#

Expose the size of the inner object.

inline idx_t size() const#

Expose the size of the slice.

inline idx_t numIndexes() const#

Number of index slots this slice owns.

inline GRef hostRef() const#

Return a host reference.

inline GRef deviceRef() const#

Return a device Reference.

inline void upload(const StreamT &stream = 0, const bool &includeSliceData = true)#

Async memcpy from host to device - indexes.

inline void download(const StreamT &stream = 0, const bool &includeSliceData = true)#

Async memcpy from device to host - Indexes.

inline void downloadSize(const StreamT &stream = 0)#

Focused async D2H of just the slice-size scalar.

Used by callers that need size() cheaply on host without pulling the full index array. After issuing this download and syncing the stream, size() returns the up-to-date active count.

inline void clearDevice()#

Mark the device copy stale.

DEVIATION (see the port report): aether::Array owns its device chunk for its whole lifetime and exposes no release, so this no longer frees device memory — it is a no-op kept for source compatibility.

inline Slice clone() const#

Clone this slice.

inline void updateCoreData(T &obj)#

update the core data

inline const CoreT &core() const#

Expose the core data.

struct Threads#
#include <Threads.h>

Collection of thread/warp utilities.

Public Static Functions

static inline idx_t blockLeader()#

Return the minimum index active thread in this block (i.e. the lowest thread ID that hasn’t returned somewhere before the invokation of this function).

static inline idx_t warpLeader(const bool &condition = true)#

Return the warp leader - i.e. the minimum active thread ID in the current warp that meets the given condition. By default, the lowest Lane ID in the warp of the active thread is returned (i.e. if return has not been invoked before this function).

namespace observable#
template<typename DataT>
class Array : public eagle::util::Observable<Array<DataT>>, public aether::Array<DataT>#
#include <ObservableArray.h>

aether scalar array wrapped with observer-pattern notifications on mutation.

The public base is aether::Array<DataT>. aether merges what used to be a scalar::Array/vector::Array two-tier into ONE Array<T, Es...>, and its non-owning face is aether::View (hostView()/ deviceView()), not a Ref<work, MaybeVolatile> family — so this class exposes hostRef()/deviceRef() returning GRefArrT<DataT> and drops WRef/VolatileRef entirely (neither was ever instantiated in-tree).

Public Types

using GRef = GRefArrT<DataT>#

The one non-owning reference tier (aether’s View).

Public Functions

Array() = delete#

Default construction is forbidden.

inline Array(const idx_t &sz, const DataT &initVal = 0)#

Construct from size and initial value.

aether’s Array(n) leaves the storage UNINITIALISED; eagle::makeArray restores a zero fill explicitly.

Array(Array &other) = delete#

Copy construction is forbidden.

inline Array(ParentT &&other)#

Move construction is allowed.

Array &operator=(Array &other) = delete#

Copy assignment is forbidden.

inline Array &operator=(ParentT &&other)#

Move assignment is allowed.

inline idx_t size() const#

Number of samples (not the total element count). For a scalar array aether::Array::size() agrees, but the name is pinned here so a later vector-valued sibling cannot silently change it.

inline GRef hostRef() const#

Non-owning view over the host copy.

inline GRef deviceRef() const#

Non-owning view over the device copy.

inline Array clone() const#

Create a clone of this observable array.

namespace slice#

Slice edit kernel.

Functions

template<typename SliceT, typename ScannerT>
void CUDAscatter(typename SliceT::GRef slice, const typename ScannerT::GRef scanner)#

Edit the slice from the given scanner type - CUDA kernel.

SASS register footprint: ~11 regs (sm_61).

__launch_bounds__(EAGLE_BLOCKSIZE,

4)

documents the cap for Launcher::setLogicalSize re-tuning.

template<typename SliceT, typename ScannerT>
inline void OMPscatter(typename SliceT::GRef slice, const typename ScannerT::GRef scanner)#

Edit the slice from the given scanner type - OpenMP.

Context-aware via aether::packetFor: inside an existing parallel region the per-sample loop is omp for work-shared, otherwise the primitive opens its own parallel region. The body is dispatched scalar-per-lane today via dispatchMasked; the packet shape is preserved so a future SIMD port can replace the scalar dispatch with packet-shaped writes without touching the outer loop structure.

Error codes — eagle::err#

enum eagle::err::EagleErrorTypes#

Bit-flag error codes reported from device kernels via EAGLE_GPU_THROW. Domain libraries layered on eagle keep their own flag + codes and wrap these macros, exactly as eagle used to wrap this header’s.

Values:

enumerator NO_ERROR#

No error.

enumerator FAILED_WARP_LEADER#

Warp leader failed a required condition.

enumerator ERR_MAX#

Sentinel — maximum error flag value.