API reference — memory#

The owning and non-owning containers: Chunk is raw owned bytes, View/Item/RuntimeView/TableHandle are typed, zero-copy handles over memory, and Array is the owning convenience type built on top of them. See Views and Items and Memory ownership: Chunk, Array, View.

aether::Chunk#

class Chunk#

Move-only owning byte allocation on a single Device.

Public Functions

Chunk() = default#

Default-constructed empty chunk: nullptr, size 0, kDLCPU[0], alignment 0.

inline Chunk(Chunk &&other) noexcept#

Move constructor: leaves other an empty chunk.

inline Chunk &operator=(Chunk &&other) noexcept#

Move assignment: frees this chunk’s own memory first (a no-op when !owns()), leaves other empty.

inline ~Chunk()#

Frees the backing allocation. Never throws.

inline std::byte *data() const#

Raw byte pointer (nullptr for an empty/zero-byte chunk).

inline std::size_t size() const#

Allocation size in bytes.

inline Device device() const#

Device this chunk’s memory lives on.

inline std::size_t alignment() const#

Alignment this chunk was allocated with.

inline bool owns() const#

true for a chunk from allocate() (frees on destruction); false for a chunk from borrow() (never frees).

Public Static Functions

static inline Chunk allocate(Device device, std::size_t bytes, std::size_t alignment = 256)#

Allocate bytes bytes on device, aligned to alignment (default 256 — the CUDA contract alignment).

bytes == 0 ALWAYS returns a valid empty chunk (data() == nullptr), for any device kind whatsoever, WITHOUT validating that kind against this build’s backend — no allocator call is made at all. Validation (and the real allocation) only happens once bytes > 0: requesting a kDLCUDA/kDLCUDAHost chunk in a build with no CUDA backend (AETHER_CPP_MODE, or CUDA simply not compiled in) throws, and so does any device kind this library does not implement at all (kDLOpenCL, …).

Throws:

aether::Error – on an unsupported device kind (for this build, or altogether) with bytes > 0, a kDLCUDA alignment request finer than the 256-byte cudaMalloc contract, or an allocation failure reported by the underlying allocator.

static inline Chunk borrow(void *ptr, std::size_t bytes, Device device, std::size_t alignment = 256)#

Wrap bytes bytes of CALLER-OWNED memory at ptr as a NON-OWNING Chunk

— the deleter is a no-op, so this chunk’s destruction (and a move-assignment’s “free the old

chunk first” step) never touches

ptr. ptr must outlive every use of the returned chunk (and every view built over it) — the caller’s responsibility, exactly as for the raw- pointer make_view overload this pairs with (aether/view/View.h).

Unlike allocate(), bytes == 0 does NOT bypass validation here — there is no allocator call to skip in the first place — but a null ptr is only rejected when bytes > 0 (mirrors make_view’s own “null pointer with a non-zero span” check), so

borrow(nullptr, 0,

dev)

is a valid empty, non-owning chunk.

Throws:

aether::Error – if ptr == nullptr and bytes > 0.

aether::View#

template<class T, class Extents, class Layout = layout_right, bool Volatile = false, bool ReadOnly = false>
class View : public aether::Expression<View<T, Extents, layout_right, false, false>, T>#

Typed, non-owning view of T over Extents, addressed through Layout (default layout_right, the SoA anchor).

Gains the expression leaf protocol unconditionally: isLeaf, element_extents (the static prefix — see detail::ViewLeafExtents above) and eval<Is...>(SampleIndex) (= (*this)(Is..., i.global())), plus operator[](SampleIndex) returning a SampleRef assignment proxy. View rebases onto Expression<View<...>, T> (CRTP).

Volatile is a 4th, defaulted (false) NTTP — every existing View<T,Extents,Layout> spelling is unaffected. When true, pointer/reference become volatile T*/volatile T& so every load and store through this view’s operator()/eval() is a genuine memory access that the compiler may not CSE, hoist, or route through a stack-backed mirror (“volatile” is a view property, not a separate handle tier). Reached only via as_volatile() below; make_view()/Array never produce one directly.

ReadOnly is a 5th, defaulted (false) NTTP. When true, reference becomes T (by value, not T& — see below) and every device-compiled operator()/eval() load routes through __ldg (#if defined(__CUDA_ARCH__); the host path is unchanged, __ldg has no host meaning). Returning by value rather than by reference is deliberate, not incidental: __ldg reads through the read-only data cache into a register — there is no addressable lvalue to hand back — and it also means a ReadOnly view is naturally non-assignable at the element level (SampleRef::operator=’s to.eval<Is...>(i) = ... fails to compile, an lvalue-required error, since eval() no longer returns a reference): the type system enforces “read-only” for free, with no extra guard needed. ReadOnly and Volatile are contradictory (a genuine memory access that also claims cache-only read-through makes no sense) and combining them is a static_assert error. Reached only via as_readonly() below.

Public Types

using pointer = std::conditional_t<Volatile, volatile T*, T*>#

The (possibly volatile-qualified) underlying pointer type.

using reference = std::conditional_t<ReadOnly, T, std::conditional_t<Volatile, volatile T&, T&>>#

The element access type: T by value when ReadOnly (see class docstring), else the (possibly volatile-qualified) element reference type.

using element_extents = detail::ViewLeafExtents<Extents>#

The static prefix of Extents (drops the trailing — batch/ sample — mode). See detail::ViewLeafExtents.

Public Functions

inline constexpr pointer data() const#

The underlying (typed, possibly volatile-qualified) pointer.

inline constexpr Device device() const#

The device this view’s memory lives on.

inline constexpr offset_t extent(std::size_t i) const#

The runtime extent of mode i (offset_t — part of the narrow address chain; see aether/index/Offset.h).

inline constexpr offset_t samples() const#

The batch-mode extent, extent(rank() - 1). Provided unconditionally (any rank >= 1); meaningful for the batched views this library actually builds — every *View alias appends a trailing dynamic sample mode. Calling it on a rank-0 view is a caller error (rank() - 1 underflows std::size_t).

inline constexpr offset_t size() const#

Product of every extent — the logical element count.

Returns offset_t, so the ubiquitous guard if (i.global() >= v.samples()) return; is a single 32-bit ISETP rather than a two-instruction 64-bit compare holding an extra register pair live. Never wraps: no view can exist whose span exceeds offset_max (the make_view guard below), and size() <= required_span_size() for both layouts.

inline constexpr const mapping_type &mapping() const#

The underlying layout mapping.

template<class ...Idxs>
inline constexpr reference operator()(Idxs... idxs) const#

Element access — data()[mapping()(idxs...)] (on the host the same offset at full width, detail::elementOffset). When ReadOnly, the device-compiled load routes through __ldg (the read-only data-cache hint, SASS LDG.E.CI); the host path is unchanged (__ldg has no host meaning).

inline constexpr View<const T, Extents, Layout, Volatile, ReadOnly> as_const() const#

A read-only view of the same data (Volatile- and ReadOnly-preserving).

inline constexpr View<T, Extents, Layout, true, ReadOnly> as_volatile() const#

A volatile-qualified view of the same data — every operator()/eval() access through the result becomes a real load/store the compiler may not CSE, hoist, or stack-mirror. Intended for shared-memory scratch that multiple statements deliberately re-read. Calling this on an already-volatile view is a static-assert error — as_volatile() is not idempotent-by-design, it is the single entry point from the (default) non-volatile side.

inline constexpr View<T, Extents, Layout, Volatile, true> as_readonly() const#

A read-only-qualified view of the same data — every device-compiled operator()/eval() load through the result routes through __ldg (see the class docstring above for why reference becomes T by value here, and why that alone makes the result non-assignable). Calling this on an already-read-only or already-volatile view is a static_assert error — mirrors as_volatile()’s own single-entry-point convention.

template<std::size_t I>
inline constexpr View<T, extents<dyn>, layout_stride> component()#

A writable rank-1 view of static component I of a rank-2 (component, sample) view — e.g. component 0 (“x”) of an aether::Array<double,3>’s hostView().

A Vec3dArr (Array<double,3>) stores its x/y/z components as three separate sample-length runs (the SoA layout — see this class’s own docstring), so “the x component” is itself an ordinary rank-1 array, once you know where it starts and how far apart consecutive samples are. This reads both numbers off this view’s own mapping (start = I * mapping().stride(0), spacing = mapping().stride(1)) rather than assuming samples(): an owning Array pitches components at its (possibly larger, capacity-rounded) capacity(), not samples(), while a view built by hand over a scratch slot pitches however its builder chose — so the two can never silently disagree. Writes through the result land directly in this view’s own backing memory, no copy involved: write via arr.hostView().component<0>, arr.upload(), and the new values are visible through arr.deviceView().

Template Parameters:

I – Which component (0-based). Checked against this view’s static component extent at compile time (a static_assert, not a runtime check) — the component count is always known at compile time for a View, so an out-of-range I is always a programming error, never a data-dependent one.

template<std::size_t I>
inline constexpr View<const T, extents<dyn>, layout_stride> component() const#

Read-only counterpart of component() (see above): identical pointer arithmetic, returns a View<const T, ...> so the result cannot be written through.

template<std::size_t... Is>
inline constexpr reference eval(const SampleIndex &i) const#

Is... matches element_extents (the static prefix); i.global() supplies the trailing (batch/sample) index — eval<Is...>(i) = (*this)(Is..., i.global()).

A member function template, so — unlike element_extents above — this is only instantiated when actually called (by expression machinery: Sum/CWiseScale/detail::RecursiveAssign), which is exactly where this invariant belongs: a View whose Extents does not carry dyn as the trailing mode compiles fine (e.g. reshaped<dyn, Minor>(...), tests/test_Reshape.*) right up until someone actually tries to use it as an expression leaf.

constexpr SampleRef<View> operator[](const SampleIndex &i)#

Assignment-target entry point: v[i] = expr / v[i] += expr / v[i] -= expr / v[i].get() (register materialization). Declared here, defined out-of-line in aether/expr/Assign.h (see SampleRef’s forward declaration above for why).

constexpr ConstSampleRef<const View> operator[](const SampleIndex &i) const#

The const read path: cv[i].get() / cv[i].eval<Is...> — no = += -= (this overload returns ConstSampleRef<const View>, which declares none); no implicit conversion is added. const View is itself a conforming expression leaf — see the static_assert below and aether::aether_expression’s own remove_cv_t note (aether/expr/Expression.h). Declared here, defined out-of-line in aether/expr/Assign.h (see ConstSampleRef’s forward declaration above for why).

The CONST overload’s out-of-line definition (see view/View.h’s forward declaration + this file’s docstring).

template<std::size_t W>
constexpr BundleRef<View, W> operator[](const BundleIndex<W> &bi)#

Assignment-target entry point: v[bi] = expr / v[bi] += expr / v[bi] -= expr / v[bi].get() (register materialization into a DeviceBundle<T,W,...>) for a BundleIndex<W> (W consecutive samples, one thread’s worth). Declared here, defined out-of-line in aether/backend/cuda/bundle/Assign.h (see BundleRef’s forward declaration above for why).

template<std::size_t W>
constexpr ConstBundleRef<const View, W> operator[](const BundleIndex<W> &bi) const#

Const counterpart of the bundle entry point above: cv[bi].get() only, no assignment operators. Declared here, defined out-of-line in aether/backend/cuda/bundle/Assign.h (see ConstBundleRef’s forward declaration above for why).

template<std::size_t W>
constexpr ConstBundleRef<const View<T, Extents, Layout, Volatile, ReadOnly>, W> operator[](const BundleIndex<W> &bi) const#

The CONST bundle overload’s out-of-line definition.

Public Static Functions

static inline constexpr std::size_t rank()#

Number of modes. A compile-time mode count, not an offset — stays std::size_t.

Public Static Attributes

static constexpr bool isLeaf = true#

View is always a leaf.

static constexpr bool isVolatile = Volatile#

Whether this view’s element access is volatile-qualified.

static constexpr bool isReadOnly = ReadOnly#

Whether this view’s device loads route through __ldg.

aether::Item#

template<class T, std::size_t... Es>
class Item : public aether::Expression<Item<T, Es...>, T>#

Register-resident value with all-static extents Es... (no dyn mode is permitted — Array, not Item, carries the batch mode). Row-major element order (matches layout_right): for Item<T,R,C>, (r, c) maps to offset r*C + c.

Rebases Item onto the Expression CRTP: Item is always a conforming expression leaf (isLeaf = true, element_extents = extents<Es...> — trivially all-static, matching Item’s own storage shape) and gains a converting constructor + assignment operator from any matching expression (aether::aether_expression), evaluated at SampleIndex::make(0) since an Item is sample-free. Both are declared here but defined out-of-line in aether/expr/Assign.h — Item itself must not #include that header (it in turn needs Item’s full definition to materialize SampleRef::get(), a genuine two-way need that a physical #include cycle cannot express; the declare/define split breaks the cycle without weakening either side’s compile-time checking).

Public Types

using element_extents = extents<Es...>#

An Item’s own shape is its element shape.

Public Functions

template<class E>
constexpr Item &operator=(const E &expr)#

Assignment from any matching expression, evaluated at SampleIndex::make(0). Defined out-of-line in aether/expr/Assign.h — see the ctor’s note above.

Public Static Attributes

static constexpr bool isLeaf = true#

Item is always a leaf.

static constexpr std::size_t Rank = sizeof...(Es)#

Number of modes.

static constexpr std::size_t Size = (Es * ... * std::size_t{1})#

Total element count — the product of Es....

aether::RuntimeView#

struct RuntimeView#

Fixed-capacity, trivially-copyable runtime tensor descriptor. See the file docstring for the field rationale.

Public Functions

inline constexpr offset_t size() const#

Product of extents[0..rank) — the logical element count. Device-legal: callable from the runtime evaluator kernel. By construction every RuntimeView this library produces (fromView below, aether/interop/DLPack.h’s fromDLPack) has already been guarded, so this narrow product does not itself re-derive that guard — see those call sites for the wide check.

template<class StaticViewT>
StaticViewT as() const#

Checked promotion onto a static fast-path view: throws aether::Error unless dtype, rank, every static extent of StaticViewT, the stride pattern (row-major/layout_right only — the one static layout aether’s static View actually uses), and the device (a known, backend-enabled kind) all match; otherwise returns the static view, zero-copy (same data pointer, no allocation, no data movement).

Host-only (throws); never called from device code, so it carries no AETHER_DEVICEHOST().

Guarded rather than split into a separate header: a split is impossible here without changing RuntimeView’s public shape from a member rt.as<T> to a free function, breaking every existing call site. An entirely unannotated function — even a bodiless declaration, never mind a call — is rejected outright by NVRTC’s JIT mode (“host functions are not allowed in JIT mode”).

Guards on __CUDACC_RTC__ rather than __CUDA_ARCH__: __CUDA_ARCH__ is also defined during nvcc’s own device-compilation pass of an ordinary .cu file (not just NVRTC) — and that pass still fully parses the whole translation unit, including host-only call sites elsewhere in the same .cu file (guarding on __CUDA_ARCH__ broke existing host-side .as<> calls in nvcc’s device pass — a real build regression, not an NVRTC-only one). __CUDACC_RTC__

is the macro that actually distinguishes “compiled

by NVRTC” from “compiled by nvcc” (host or device pass alike) — NVRTC defines it, nvcc never does — so guarding on it removes

as<> only from NVRTC’s view, leaving every nvcc-compiled .cu/.cpp TU (host and device pass) exactly as it always was.

Public Members

void *data = nullptr#

Untyped data pointer (never owning — mirrors View).

DType dtype = {}#

The element dtype (DLPack-native, aether/dtype/DType.h).

Device device = {}#

The device this descriptor’s memory lives on.

std::size_t rank = 0#

Number of modes (<= RuntimeRankCap). A runtime count, not an offset — stays std::size_t (mirrors extents::Rank).

detail::Carray<offset_t, RuntimeRankCap> extents = {}#

Per-mode extents, element units, offset_t. Only indices [0, rank) are meaningful.

detail::Carray<offset_t, RuntimeRankCap> strides = {}#

Per-mode strides, element units (not bytes — matches DLPack’s own convention) — may be non-contiguous. Only indices [0, rank) are meaningful.

aether::TableHandle#

template<class T, class Carrier = Plain>
class TableHandle#

aether::Array#

template<class T, std::size_t... Es>
class Array#

Public Types

using Extents = extents<Es..., dyn>#

extents<Es..., dyn> — the symmetry rule.

using ViewT = View<T, Extents, layout_stride>#

hostView()/deviceView()’s return type: pitch = capacity() (quantised), shape = samples().

using PackedViewT = View<T, Extents, layout_right>#

hostPacked()/devicePacked()’s return type: the original compact accessor — valid only when capacity() == samples().

Public Functions

inline explicit Array(std::size_t n_samples)#

Allocate storage for n_samples items of shape Es.... capacity() is n_samples rounded up to kCapacityAlignment — unconditional, even for a fresh Array: growth later must never re-stride an Array that never grew.

inline std::size_t samples() const#

Number of samples (the batch mode’s current extent — vector- size()-like; unaffected by any reserved headroom).

inline std::size_t size() const#

Total element count — (Es * ...) * samples().

inline std::size_t capacity() const#

How many samples the backing Chunk(s) actually hold room for — always a multiple of kCapacityAlignment.

inline void reserve(std::size_t n)#

Grow capacity() to at least n (quantised up to kCapacityAlignment) — a no-op when n <= capacity(). Otherwise reallocates the backing Chunk(s) and copies the existing samples() real elements, pitch-aware (component c’s samples() elements move from offset c*capacity() to c*newCapacity), preserving samples(). data() pointer identity is not preserved across a realloc (host-throw on allocation failure — from Chunk::allocate).

inline void resize(std::size_t n)#

Set samples() to n — no reallocation (pointer/stride stable) while n <= capacity(); otherwise reserve(n) first (realloc + pitched copy of the old samples() real elements), then adopts n as the new samples(). Shrinking (n < samples()) never reallocates and never releases capacity().

inline void push_back(const Item<T, Es...> &item)#

Append one Item<T,Es...>, host-side, amortised-doubling reserve() when samples() == capacity() (a fresh, zero-capacity Array grows to kCapacityAlignment). Writes through the host chunk only — mirrors upload()/download()’s explicit-transfer contract; upload() afterwards to reach the device copy in CUDA mode.

inline void clear()#

samples() <- 0; capacity() is kept (no deallocation).

inline ViewT hostView()#

A writable view over the host-resident copy — shape samples(), pitch capacity() (shape from size, strides from capacity). Sample extent is 0 for a samples() == 0 Array — any per-sample expression-template assignment over it is a no-op by construction (nothing to iterate).

inline ConstViewT hostView() const#

A read-only view over the host-resident copy (see hostView()).

inline PackedViewT hostPacked()#

The compact layout_right accessor, available only when capacity() == samples() (a full, non-reserving Array).

Throws:

aether::Error – if capacity() != samples().

inline ConstPackedViewT hostPacked() const#

This is an overloaded member function, provided for convenience. It differs from the above function only in what argument(s) it accepts.

inline ViewT hostSpare()#

A writable view over the reserved tail [samples(), capacity()) of the host-resident copy — the target a caller (or, on the device side, AppendCounter) writes new samples into before commit()ing them.

inline ConstViewT hostSpare() const#

This is an overloaded member function, provided for convenience. It differs from the above function only in what argument(s) it accepts.

inline ViewT deviceView()#

AETHER_CPP_MODE: the same single chunk as hostView().

inline void commit(std::size_t count)#

Adopt count newly-written spare samples — samples() <- samples() + count (host-side; bounds-checked against spare capacity, host-throw on overflow). The device-side atomic append counter that produces count is not here.

inline void upload()#

Blocking host -> device copy. No-op in AETHER_CPP_MODE.

inline void download()#

Blocking device -> host copy. No-op in AETHER_CPP_MODE.