API reference — util#
Global typedefs, logging, error handling, observer-pattern primitives, array slicing, and OpenMP thread-count management.
Global typedefs#
-
namespace eagle#
Top-level eagle namespace.
Templated slice - class to act as a slice of a given reference object.
Templated observer for observer pattern - used to track move semantics with custom references.
Typedefs
-
using Real = double#
Scalar floating-point type used throughout eagle.
-
using idx_t = aether::idx_t#
Index type (from aether —
aether::idx_t).
-
using mInt_t = int#
General-purpose signed integer type for metadata fields.
-
using uint_t = unsigned int#
Unsigned integer type.
-
using dims_t = aether::dims_t#
Dimension type (from aether).
-
template<typename DT, idx_t VD>
using VecNTArray = aether::Array<DT, VD># aether SoA array with element type
DTand vector extentVD.aether merges a two-tier
scalar::Array/vector::Arraydesign into ONEArray<T, Es...>whose extents pack carries the static (vector) modes and whose trailing dynamic mode is the batch — so a fixed-widthvector::Array<DT,VD>is spelledArray<DT, VD>and its scalar sibling is simplyArray<DT>.
-
template<typename DT, idx_t VD>
using VecNT = aether::Item<DT, VD># aether register-resident item with element type
DTand dimensionVD(the per-item vector width).
-
using Vec3R = VecNT<Real, 3>#
3-D double-precision vector. Same type as
aether::Vec3d(aether/aether.h) under a different name.
-
using Vec3RArray = VecNTArray<Real, 3>#
SoA array of 3-D double-precision vectors.
-
using Vec6RArray = VecNTArray<Real, 6>#
SoA array of 6-D double-precision vectors.
-
template<typename ComponentT>
using GRefArrT = typename aether::Array<ComponentT>::ViewT# Shorthand for the device-side view type of an aether scalar array.
aether collapses an earlier
GRef/WRef/GRef::HandleTtwo-tier into the oneaether::View(aether/view/View.h): the View IS the lightweight descriptor, and “global vs work reference” is now just which chunk the View was built over.Array<T>::ViewTis that View type.
-
template<typename ComponentT>
using CRefArrT = typename aether::Array<ComponentT>::ConstViewT# Read-only counterpart of
GRefArrT—Array<T>::ConstViewT.aether has no implicit
View<T>->View<const T>conversion; call.as_const()on a mutable view to reach this type.
-
using nativeStream_t = aether::Stream#
Native stream handle threaded through eagle’s copy paths.
aether declares
aether::Streamonly in CUDA-backed builds (aether/chunk/Copy.h); the pure-C++ arm keeps the neutral integer tageagle::cpu::StreamRef::native()already returns, so a “stream” argument threads throughEAGLE_CPU_ONLYcode unchanged.
Functions
-
inline aether::Device hostDevice()#
The aether
Deviceeagle’s HOST-resident buffers live on — the same oneaether::Arraygives its own host chunk.
-
inline aether::Device deviceDevice()#
The aether
Deviceeagle’s DEVICE-resident buffers live on. In anEAGLE_CPU_ONLYbuild there is only one copy, so this ishostDevice()— mirroringaether::Array’s own single-chunkAETHER_CPP_MODEarm.
-
template<typename T>
inline GRefArrT<T> spanView(T *data, idx_t n, aether::Device dev)# Non-owning scalar
Viewover a raw, unit-stride span ofnelements — an earlierscalar::Array<T>::GRef(ptr, size)constructor.aether’s
make_view()is HOST-ONLY (it throwsaether::Erroron a bad span, and device code cannot throw —aether/view/MakeView.h), so a span view that must also be buildable inside a kernel is assembled fromView’s public(data, mapping, device)constructor, exactly asaether::make_work_viewdoes for shared memory.
-
template<std::size_t C, typename T>
inline auto packedSpanView(T *data, idx_t n, aether::Device dev)# spanViewover a component-packed (SoA) span:Ccomponents ofnsamples each, componentcstarting atdata + c * n.The pitch is
nexactly — NOTaether::Array’s quantisedcapacity(). This is the shape a caller carving aC*nscratch slot chooses for itself, andaether::View::component<I>()reads the pitch back off the mapping, so the two never disagree.
-
template<typename T, std::size_t... Es>
inline aether::Array<T, Es...> makeArray(std::size_t n, const T &initVal = T{})# Allocate an
aether::Arrayofnsamples with every element set toinitVal.An earlier
Array(size, initVal = 0)ALWAYS filled (zero by default); aether’sArray(n)deliberately leaves the storage UNINITIALISED (it allocates aChunkand nothing else). Every eagle site that relied on the earlier implicit zero-fill goes through here, so the fill is explicit and greppable rather than an inherited accident. Fills the HOST copy; callupload()to reach the device one, exactly as aether’s explicit-transfer contract requires.
-
template<typename T, std::size_t... Es>
inline aether::Array<T, Es...> cloneArray(const aether::Array<T, Es...> &src)# Deep copy of an
aether::Array.aether::Arrayowns move-onlyChunks, so it has no copy constructor and noclone(). Both arrays get the SAME quantisedcapacity()for the samesamples(), so one flat span copy reproduces the pitch exactly. Copies the HOST copy and thenupload()s, matching aether’s explicit-transfer contract (a device-only write not yet downloaded is NOT carried over — An earlier clone copied both sides).
-
using Real = double#
Namespace eagle::util#
Includes the nested eagle::util::observable (the observer-pattern array)
and eagle::util::slice (array slicing) namespaces — Breathe’s :members:
already descends into them, so they are not repeated as separate directives
below (doing so re-registered the same Doxygen anchor IDs twice and tripped
docutils’ “Duplicate explicit target name”).
-
namespace util#
Enums
-
enum Level#
Logging severity levels, ordered from least to most verbose.
Values:
-
enumerator OFF#
Suppress all log output.
-
enumerator ERROR#
Error conditions; does not throw — use
EAGLE_THROWfor that.
-
enumerator WARN#
Warnings: unexpected but recoverable situations.
-
enumerator INFO#
Informational messages (default level).
-
enumerator DEBUG#
Detailed diagnostic output.
-
enumerator TRACE#
Per-step trace; very verbose.
-
enumerator ALL#
Alias for TRACE; enables everything.
-
enumerator OFF#
Functions
-
inline void setLogLevel(const Level &level)#
Set the active log level.
- Parameters:
level – Messages with severity above
levelare suppressed.
-
inline void disableLogging()#
Disable all log output (equivalent to
setLogLevel(OFF)).
-
static std::string getLevelTag(const Level &level)#
Return an ANSI-coloured log-level tag string for terminal output.
- Parameters:
level – The severity level to render.
- Returns:
Bracketed, coloured tag string, e.g.
"* [INFO ]".
-
template<typename ...Args>
void log(const Level &level, const std::string &msg, Args... args)# Log a formatted message at the given severity level.
The message is suppressed if
level > currentLogLevel. Formatting followsprintfconventions; the format string ismsg.- Template Parameters:
Args – Variadic pack of
printf-compatible argument types.- Parameters:
level – Severity level of the message.
msg –
printf-style format string.args – Format arguments.
-
template<typename ...Args>
void log_fn(const Level &level, const std::string &func, const std::string &msg, Args... args)# Log a formatted message prefixed with a function name.
Prepends
[func]tomsgand delegates tolog(). Used internally by theEAGLE_INFO,EAGLE_DEBUG, … macros.- Template Parameters:
Args – Variadic pack of
printf-compatible argument types.- Parameters:
level – Severity level of the message.
func – Name of the calling function (typically
__func__).msg –
printf-style format string.args – Format arguments.
-
class CaptureGuardState#
- #include <CaptureGuard.h>
Per-thread capture-depth counter + deferred-destroy queue (STOP-THE-LINE incident fix).
The fix is never making the illegal call while a capture is in flight, not unswallowing its error. Every capture-opening call site (
StreamCapturer,CaptureConditional) brackets itself with a :class:CaptureScope, which increments/decrements ONE process-wide depth counter here. Every CUDA-touching teardown site calls :meth:CaptureGuardState::destroyOrDeferinstead of calling the driver directly: if the depth is 0 (no capture anywhere in the process), it runs immediately — byte-identical to today’s behavior. If the depth is > 0, the destroy is queued and runs the moment the OUTERMOST capture scope of that thread exits (CaptureScope’s destructor / explicitrelease()), plus a thread-exit backstop so a queue that never drains naturally (e.g. an aborted capture) does not leak forever.- The defect this closes
A CUDA-touching teardown (
cudaGraphExecDestroy,cudaStreamDestroy,cudaGraphDestroy,cudaFree, …) that fires WHILE any stream on the process is actively being captured INVALIDATES that capture (cudaErrorStreamCaptureInvalidated) even when the destroyed handle has nothing to do with the capturing stream. Every such call site in this codebase already swallows its OWN error (EAGLE_CHECK_NOTHROW/ a bare unchecked call) so the destroy itself never throws — correct, and by design (a throwing destructor unwinding during teardown would abort the process). But swallowing the ERROR does not undo the DAMAGE: the capture is poisoned, the NEXT operation on it fails, and the native state left behind can crash much later, at an unrelated GC or interpreter exit (the exact shape root-caused in the incident this fixes: a deadLauncher’scudaGraphExecDestroy, run by Python’s garbage collector DURING a later, unrelated active capture, invalidated it).
- Thread-safety
The capture depth and the deferred-destroy queue are PER THREAD (
thread_local): captures are opened incudaStreamCaptureModeThreadLocal, so a capture on thread A is not invalidated by a free, allocation or destroy on thread B, and B’s teardowns therefore run immediately. Only the capturing thread defers its own teardowns, and drains them when ITS outermost capture scope exits. A second top-level capture begun on a thread that already has one open is refused (:meth:requireIdleThread): a capture session is begun, filled and ended by one thread, one at a time per thread. Whole-device synchronisation by any thread invalidates every open capture (a CUDA rule, any mode); stream-level synchronisation is fine.
Note
Pure host-side bookkeeping — no CUDA type appears in this header, so it compiles identically in
EAGLE_CPU_ONLYmode (there is no capture concept there; nothing calls into it).Public Functions
-
inline void enter()#
Enter one capture scope. Call ONLY after the underlying
cudaStream*Capture*call has SUCCEEDED (a failed begin never opened a capture, so it must not raise the depth).
-
inline void requireIdleThread() const#
Refuse a top-level capture on a thread that already has one open. Call BEFORE the underlying begin-capture call. A nested weave (
CaptureConditional) is part of the open session and does not call this.- Throws:
std::logic_error – If this thread’s capture depth is > 0.
-
inline void exit()#
Exit one capture scope. Symmetric with :meth:
enter— callers must not call this without a matching priorenter()(:class:CaptureScopeenforces that pairing). Drains the deferred queue once the depth returns to 0.
-
inline void destroyOrDefer(std::function<void()> op)#
Run
opnow if no capture is active anywhere in the process (depth 0) — identical timing to today’s inline destroy. Otherwise enqueue it to run once the outermost capture scope exits (or at process exit, as a backstop).opmust be the FULLY self-contained destroy call (its own error policy already decided at the call site, e.g. wrapped inEAGLE_CHECK_NOTHROWthere) — this method does not add or remove error handling.
-
inline int depth() const#
Current capture depth (tests / diagnostics).
-
inline std::size_t pendingCount() const#
Number of destroys queued, not yet run (tests / diagnostics).
Public Static Functions
-
static inline CaptureGuardState &instance()#
The single process-wide instance (Meyers singleton — safe init order across translation units, destroyed at process exit, which is exactly when the backstop drain below needs to run).
-
class CaptureScope#
- #include <CaptureGuard.h>
RAII capture-scope token:
enter()s :class:CaptureGuardStateon construction,exit()s on destruction or an explicit :meth:release.Intended as an
std::optional<CaptureScope>MEMBER of a capture-opening class (StreamCapturer,CaptureConditional):.emplace()it right after the underlying begin-capture call succeeds,.reset()it (calling :meth:release) right after the underlying end-capture call runs — regardless of whether that call reports success, since the capture SESSION is over either way. Exception-safe even if a caller never reaches an explicit release: the owning object’s OWN destructor still destroys this member (whenever that object itself is destroyed), so the depth is always eventually decremented.Public Functions
-
inline void release()#
Exit the scope early. Idempotent — calling it again (or on a default-constructed-then-never-entered token, which never happens here since the constructor always enters) is a no-op.
-
inline void release()#
-
template<typename ChildT>
struct Observable# - #include <Observable.h>
CRTP class for Observability traits.
-
template<typename T>
class Observer# - #include <Observer.h>
Move semantics observer.
-
template<typename SliceT>
class RefSlice# - #include <Slice.h>
Reference Slice.
There is no separate
work/MaybeVolatilepair — seeeagle::filtering::RefScannerfor the reasoning. The index array is now a HELDaether::Viewrather than a base class: aether’sArrayhas noRef<work, MaybeVolatile>family to inherit from, and aViewis a value descriptor, so composition is the direct translation.Subclassed by eagle::filtering::RefSlice< Self >
Public Types
-
template<typename MyClass>
class Slice : public aether::Array<idx_t># - #include <Slice.h>
Subclassed by eagle::filtering::FilteringSlice< MyClass >
Public Types
Public Functions
-
Slice() = delete#
Default constructor is forbidden.
-
inline Slice(T *myobj = nullptr)#
Construct from the given object pointer.
Every index starts at the object size — deliberately out of range, the “not yet scattered” sentinel
Array{sz, sz}sets.
-
inline Slice(T &myobj, ParentT &&idxs, const idx_t sz)#
Construct from the given object and indexes and size.
-
inline Slice(ParentT &&idxs, CoreT &&coredata, ParentT &&sz)#
Move construct from data types.
-
inline void upload(const StreamT &stream = 0, const bool &includeSliceData = true)#
Async memcpy from host to device - indexes.
-
inline void download(const StreamT &stream = 0, const bool &includeSliceData = true)#
Async memcpy from device to host - Indexes.
-
inline void downloadSize(const StreamT &stream = 0)#
Focused async D2H of just the slice-size scalar.
Used by callers that need
size()cheaply on host without pulling the full index array. After issuing this download and syncing the stream,size()returns the up-to-date active count.
-
inline void clearDevice()#
Mark the device copy stale.
DEVIATION (see the port report):
aether::Arrayowns its device chunk for its whole lifetime and exposes no release, so this no longer frees device memory — it is a no-op kept for source compatibility.
-
inline const CoreT &core() const#
Expose the core data.
-
Slice() = delete#
-
struct Threads#
- #include <Threads.h>
Collection of thread/warp utilities.
Public Static Functions
-
static inline idx_t blockLeader()#
Return the minimum index active thread in this block (i.e. the lowest thread ID that hasn’t returned somewhere before the invokation of this function).
-
static inline idx_t warpLeader(const bool &condition = true)#
Return the warp leader - i.e. the minimum active thread ID in the current warp that meets the given condition. By default, the lowest Lane ID in the warp of the active thread is returned (i.e. if return has not been invoked before this function).
-
static inline idx_t blockLeader()#
-
namespace observable#
-
template<typename DataT>
class Array : public eagle::util::Observable<Array<DataT>>, public aether::Array<DataT># - #include <ObservableArray.h>
aether scalar array wrapped with observer-pattern notifications on mutation.
The public base is
aether::Array<DataT>. aether merges what used to be ascalar::Array/vector::Arraytwo-tier into ONEArray<T, Es...>, and its non-owning face isaether::View(hostView()/deviceView()), not aRef<work, MaybeVolatile>family — so this class exposeshostRef()/deviceRef()returningGRefArrT<DataT>and dropsWRef/VolatileRefentirely (neither was ever instantiated in-tree).Public Functions
-
Array() = delete#
Default construction is forbidden.
-
inline Array(const idx_t &sz, const DataT &initVal = 0)#
Construct from size and initial value.
aether’s
Array(n)leaves the storage UNINITIALISED;eagle::makeArrayrestores a zero fill explicitly.
-
inline Array(ParentT &&other)#
Move construction is allowed.
-
Array() = delete#
-
template<typename DataT>
-
namespace slice#
Slice edit kernel.
Functions
-
template<typename SliceT, typename ScannerT>
void CUDAscatter(typename SliceT::GRef slice, const typename ScannerT::GRef scanner)# Edit the slice from the given scanner type - CUDA kernel.
SASS register footprint: ~11 regs (sm_61).
documents the cap for__launch_bounds__(EAGLE_BLOCKSIZE,
4)
Launcher::setLogicalSizere-tuning.
-
template<typename SliceT, typename ScannerT>
inline void OMPscatter(typename SliceT::GRef slice, const typename ScannerT::GRef scanner)# Edit the slice from the given scanner type - OpenMP.
Context-aware via
aether::packetFor: inside an existing parallel region the per-sample loop isomp forwork-shared, otherwise the primitive opens its own parallel region. The body is dispatched scalar-per-lane today viadispatchMasked; the packet shape is preserved so a future SIMD port can replace the scalar dispatch with packet-shaped writes without touching the outer loop structure.
-
template<typename SliceT, typename ScannerT>
-
enum Level#
Error codes — eagle::err#
-
enum eagle::err::EagleErrorTypes#
Bit-flag error codes reported from device kernels via
EAGLE_GPU_THROW. Domain libraries layered on eagle keep their own flag + codes and wrap these macros, exactly as eagle used to wrap this header’s.Values:
-
enumerator NO_ERROR#
No error.
-
enumerator FAILED_WARP_LEADER#
Warp leader failed a required condition.
-
enumerator ERR_MAX#
Sentinel — maximum error flag value.
-
enumerator NO_ERROR#