API reference — memory#
The owning and non-owning containers: Chunk is raw owned bytes,
View/Item/RuntimeView/TableHandle are typed, zero-copy
handles over memory, and Array is the owning convenience type built on
top of them. See Views and Items and Memory ownership: Chunk, Array, View.
aether::Chunk#
-
class Chunk#
Move-only owning byte allocation on a single
Device.Public Functions
-
Chunk() = default#
Default-constructed empty chunk: nullptr, size 0, kDLCPU[0], alignment 0.
-
inline Chunk &operator=(Chunk &&other) noexcept#
Move assignment: frees this chunk’s own memory first (a no-op when
!owns()), leavesotherempty.
-
inline ~Chunk()#
Frees the backing allocation. Never throws.
-
inline std::byte *data() const#
Raw byte pointer (
nullptrfor an empty/zero-byte chunk).
-
inline std::size_t size() const#
Allocation size in bytes.
-
inline std::size_t alignment() const#
Alignment this chunk was allocated with.
-
inline bool owns() const#
truefor a chunk fromallocate()(frees on destruction);falsefor a chunk fromborrow()(never frees).
Public Static Functions
-
static inline Chunk allocate(Device device, std::size_t bytes, std::size_t alignment = 256)#
Allocate
bytesbytes ondevice, aligned toalignment(default 256 — the CUDA contract alignment).bytes == 0ALWAYS returns a valid empty chunk (data() == nullptr), for any device kind whatsoever, WITHOUT validating that kind against this build’s backend — no allocator call is made at all. Validation (and the real allocation) only happens oncebytes > 0: requesting akDLCUDA/kDLCUDAHostchunk in a build with no CUDA backend (AETHER_CPP_MODE, or CUDA simply not compiled in) throws, and so does any device kind this library does not implement at all (kDLOpenCL, …).- Throws:
aether::Error – on an unsupported device kind (for this build, or altogether) with
bytes > 0, akDLCUDAalignment request finer than the 256-bytecudaMalloccontract, or an allocation failure reported by the underlying allocator.
-
static inline Chunk borrow(void *ptr, std::size_t bytes, Device device, std::size_t alignment = 256)#
Wrap
bytesbytes of CALLER-OWNED memory atptras a NON-OWNINGChunk— the deleter is a no-op, so this chunk’s destruction (and a move-assignment’s “free the old
chunk first” step) never touches
ptr.ptrmust outlive every use of the returned chunk (and every view built over it) — the caller’s responsibility, exactly as for the raw- pointermake_viewoverload this pairs with (aether/view/View.h).Unlike
allocate(),bytes == 0does NOT bypass validation here — there is no allocator call to skip in the first place — but a nullptris only rejected whenbytes > 0(mirrorsmake_view’s own “null pointer with a non-zero span” check), sois a valid empty, non-owning chunk.borrow(nullptr, 0,
dev)
- Throws:
aether::Error – if
ptr == nullptrandbytes > 0.
-
Chunk() = default#
aether::View#
-
template<class T, class Extents, class Layout = layout_right, bool Volatile = false, bool ReadOnly = false>
class View : public aether::Expression<View<T, Extents, layout_right, false, false>, T># Typed, non-owning view of
ToverExtents, addressed throughLayout(defaultlayout_right, the SoA anchor).Gains the expression leaf protocol unconditionally:
isLeaf,element_extents(the static prefix — seedetail::ViewLeafExtentsabove) andeval<Is...>(SampleIndex)(= (*this)(Is..., i.global())), plusoperator[](SampleIndex)returning aSampleRefassignment proxy.Viewrebases ontoExpression<View<...>, T>(CRTP).Volatileis a 4th, defaulted (false) NTTP — every existingView<T,Extents,Layout>spelling is unaffected. Whentrue,pointer/referencebecomevolatile T*/volatile T&so every load and store through this view’soperator()/eval()is a genuine memory access that the compiler may not CSE, hoist, or route through a stack-backed mirror (“volatile” is a view property, not a separate handle tier). Reached only viaas_volatile()below;make_view()/Arraynever produce one directly.ReadOnlyis a 5th, defaulted (false) NTTP. Whentrue,referencebecomesT(by value, notT&— see below) and every device-compiledoperator()/eval()load routes through__ldg(#if defined(__CUDA_ARCH__); the host path is unchanged,__ldghas no host meaning). Returning by value rather than by reference is deliberate, not incidental:__ldgreads through the read-only data cache into a register — there is no addressable lvalue to hand back — and it also means aReadOnlyview is naturally non-assignable at the element level (SampleRef::operator=’sto.eval<Is...>(i) = ...fails to compile, an lvalue-required error, sinceeval()no longer returns a reference): the type system enforces “read-only” for free, with no extra guard needed.ReadOnlyandVolatileare contradictory (a genuine memory access that also claims cache-only read-through makes no sense) and combining them is astatic_asserterror. Reached only viaas_readonly()below.Public Types
-
using pointer = std::conditional_t<Volatile, volatile T*, T*>#
The (possibly volatile-qualified) underlying pointer type.
Public Functions
-
inline constexpr offset_t extent(std::size_t i) const#
The runtime extent of mode
i(offset_t— part of the narrow address chain; seeaether/index/Offset.h).
-
inline constexpr offset_t samples() const#
The batch-mode extent,
extent(rank() - 1). Provided unconditionally (any rank >= 1); meaningful for the batched views this library actually builds — every*Viewalias appends a trailing dynamic sample mode. Calling it on a rank-0 view is a caller error (rank() - 1underflowsstd::size_t).
-
inline constexpr offset_t size() const#
Product of every extent — the logical element count.
Returns
offset_t, so the ubiquitous guardif (i.global() >= v.samples()) return;is a single 32-bitISETPrather than a two-instruction 64-bit compare holding an extra register pair live. Never wraps: no view can exist whose span exceedsoffset_max(themake_viewguard below), andsize() <= required_span_size()for both layouts.
-
inline constexpr const mapping_type &mapping() const#
The underlying layout mapping.
-
template<class ...Idxs>
inline constexpr reference operator()(Idxs... idxs) const# Element access —
data()[mapping()(idxs...)](on the host the same offset at full width,detail::elementOffset). WhenReadOnly, the device-compiled load routes through__ldg(the read-only data-cache hint, SASSLDG.E.CI); the host path is unchanged (__ldghas no host meaning).
-
inline constexpr View<const T, Extents, Layout, Volatile, ReadOnly> as_const() const#
A read-only view of the same data (Volatile- and ReadOnly-preserving).
-
inline constexpr View<T, Extents, Layout, true, ReadOnly> as_volatile() const#
A volatile-qualified view of the same data — every
operator()/eval()access through the result becomes a real load/store the compiler may not CSE, hoist, or stack-mirror. Intended for shared-memory scratch that multiple statements deliberately re-read. Calling this on an already-volatile view is a static-assert error —as_volatile()is not idempotent-by-design, it is the single entry point from the (default) non-volatile side.
-
inline constexpr View<T, Extents, Layout, Volatile, true> as_readonly() const#
A read-only-qualified view of the same data — every device-compiled
operator()/eval()load through the result routes through__ldg(see the class docstring above for whyreferencebecomesTby value here, and why that alone makes the result non-assignable). Calling this on an already-read-only or already-volatile view is astatic_asserterror — mirrorsas_volatile()’s own single-entry-point convention.
-
template<std::size_t I>
inline constexpr View<T, extents<dyn>, layout_stride> component()# A writable rank-1 view of static component
Iof a rank-2 (component, sample) view — e.g. component 0 (“x”) of anaether::Array<double,3>’shostView().A
Vec3dArr(Array<double,3>) stores its x/y/z components as three separate sample-length runs (the SoA layout — see this class’s own docstring), so “the x component” is itself an ordinary rank-1 array, once you know where it starts and how far apart consecutive samples are. This reads both numbers off this view’s own mapping (start =I * mapping().stride(0), spacing =mapping().stride(1)) rather than assumingsamples(): an owningArraypitches components at its (possibly larger, capacity-rounded)capacity(), notsamples(), while a view built by hand over a scratch slot pitches however its builder chose — so the two can never silently disagree. Writes through the result land directly in this view’s own backing memory, no copy involved: write viaarr.hostView().component<0>,arr.upload(), and the new values are visible througharr.deviceView().- Template Parameters:
I – Which component (0-based). Checked against this view’s static component extent at compile time (a
static_assert, not a runtime check) — the component count is always known at compile time for aView, so an out-of-rangeIis always a programming error, never a data-dependent one.
-
template<std::size_t I>
inline constexpr View<const T, extents<dyn>, layout_stride> component() const# Read-only counterpart of
component()(see above): identical pointer arithmetic, returns aView<const T, ...>so the result cannot be written through.
-
template<std::size_t... Is>
inline constexpr reference eval(const SampleIndex &i) const# Is...matcheselement_extents(the static prefix);i.global()supplies the trailing (batch/sample) index —eval<Is...>(i) = (*this)(Is..., i.global()).A member function template, so — unlike
element_extentsabove — this is only instantiated when actually called (by expression machinery:Sum/CWiseScale/detail::RecursiveAssign), which is exactly where this invariant belongs: aViewwhose Extents does not carrydynas the trailing mode compiles fine (e.g.reshaped<dyn, Minor>(...),tests/test_Reshape.*) right up until someone actually tries to use it as an expression leaf.
-
constexpr SampleRef<View> operator[](const SampleIndex &i)#
Assignment-target entry point:
v[i] = expr/v[i] += expr/v[i] -= expr/v[i].get()(register materialization). Declared here, defined out-of-line inaether/expr/Assign.h(seeSampleRef’s forward declaration above for why).
-
constexpr ConstSampleRef<const View> operator[](const SampleIndex &i) const#
The const read path:
cv[i].get()/cv[i].eval<Is...>— no= += -=(this overload returnsConstSampleRef<const View>, which declares none); no implicit conversion is added.const Viewis itself a conforming expression leaf — see thestatic_assertbelow andaether::aether_expression’s ownremove_cv_tnote (aether/expr/Expression.h). Declared here, defined out-of-line inaether/expr/Assign.h(seeConstSampleRef’s forward declaration above for why).The CONST overload’s out-of-line definition (see
view/View.h’s forward declaration + this file’s docstring).
-
template<std::size_t W>
constexpr BundleRef<View, W> operator[](const BundleIndex<W> &bi)# Assignment-target entry point:
v[bi] = expr/v[bi] += expr/v[bi] -= expr/v[bi].get()(register materialization into aDeviceBundle<T,W,...>) for aBundleIndex<W>(W consecutive samples, one thread’s worth). Declared here, defined out-of-line inaether/backend/cuda/bundle/Assign.h(seeBundleRef’s forward declaration above for why).
-
template<std::size_t W>
constexpr ConstBundleRef<const View, W> operator[](const BundleIndex<W> &bi) const# Const counterpart of the bundle entry point above:
cv[bi].get()only, no assignment operators. Declared here, defined out-of-line inaether/backend/cuda/bundle/Assign.h(seeConstBundleRef’s forward declaration above for why).
Public Static Functions
-
static inline constexpr std::size_t rank()#
Number of modes. A compile-time mode count, not an offset — stays
std::size_t.
-
using pointer = std::conditional_t<Volatile, volatile T*, T*>#
aether::Item#
-
template<class T, std::size_t... Es>
class Item : public aether::Expression<Item<T, Es...>, T># Register-resident value with all-static extents
Es...(nodynmode is permitted —Array, notItem, carries the batch mode). Row-major element order (matcheslayout_right): forItem<T,R,C>,(r, c)maps to offsetr*C + c.Rebases
Itemonto theExpressionCRTP:Itemis always a conforming expression leaf (isLeaf = true,element_extents = extents<Es...>— trivially all-static, matchingItem’s own storage shape) and gains a converting constructor + assignment operator from any matching expression (aether::aether_expression), evaluated atSampleIndex::make(0)since anItemis sample-free. Both are declared here but defined out-of-line inaether/expr/Assign.h—Itemitself must not#includethat header (it in turn needsItem’s full definition to materializeSampleRef::get(), a genuine two-way need that a physical#includecycle cannot express; the declare/define split breaks the cycle without weakening either side’s compile-time checking).Public Functions
aether::RuntimeView#
-
struct RuntimeView#
Fixed-capacity, trivially-copyable runtime tensor descriptor. See the file docstring for the field rationale.
Public Functions
-
inline constexpr offset_t size() const#
Product of
extents[0..rank)— the logical element count. Device-legal: callable from the runtime evaluator kernel. By construction everyRuntimeViewthis library produces (fromViewbelow,aether/interop/DLPack.h’sfromDLPack) has already been guarded, so this narrow product does not itself re-derive that guard — see those call sites for the wide check.
-
template<class StaticViewT>
StaticViewT as() const# Checked promotion onto a static fast-path view: throws
aether::Errorunless dtype, rank, every static extent ofStaticViewT, the stride pattern (row-major/layout_rightonly — the one static layout aether’s staticViewactually uses), and the device (a known, backend-enabled kind) all match; otherwise returns the static view, zero-copy (samedatapointer, no allocation, no data movement).Host-only (throws); never called from device code, so it carries no
AETHER_DEVICEHOST().Guarded rather than split into a separate header: a split is impossible here without changing
RuntimeView’s public shape from a memberrt.as<T>to a free function, breaking every existing call site. An entirely unannotated function — even a bodiless declaration, never mind a call — is rejected outright by NVRTC’s JIT mode (“host functions are not allowed in JIT mode”).Guards on
__CUDACC_RTC__rather than__CUDA_ARCH__:__CUDA_ARCH__is also defined during nvcc’s own device-compilation pass of an ordinary.cufile (not just NVRTC) — and that pass still fully parses the whole translation unit, including host-only call sites elsewhere in the same.cufile (guarding on__CUDA_ARCH__broke existing host-side.as<>calls in nvcc’s device pass — a real build regression, not an NVRTC-only one).__CUDACC_RTC__is the macro that actually distinguishes “compiled
by NVRTC” from “compiled by nvcc” (host or device pass alike) — NVRTC defines it, nvcc never does — so guarding on it removes
as<>only from NVRTC’s view, leaving every nvcc-compiled.cu/.cppTU (host and device pass) exactly as it always was.
Public Members
-
std::size_t rank = 0#
Number of modes (
<= RuntimeRankCap). A runtime count, not an offset — staysstd::size_t(mirrorsextents::Rank).
-
detail::Carray<offset_t, RuntimeRankCap> extents = {}#
Per-mode extents, element units,
offset_t. Only indices[0, rank)are meaningful.
-
detail::Carray<offset_t, RuntimeRankCap> strides = {}#
Per-mode strides, element units (not bytes — matches DLPack’s own convention) — may be non-contiguous. Only indices
[0, rank)are meaningful.
-
inline constexpr offset_t size() const#
aether::TableHandle#
-
template<class T, class Carrier = Plain>
class TableHandle#
aether::Array#
-
template<class T, std::size_t... Es>
class Array# Public Types
-
using ViewT = View<T, Extents, layout_stride>#
hostView()/deviceView()’s return type: pitch =capacity()(quantised), shape =samples().
-
using PackedViewT = View<T, Extents, layout_right>#
hostPacked()/devicePacked()’s return type: the original compact accessor — valid only whencapacity() == samples().
Public Functions
-
inline explicit Array(std::size_t n_samples)#
Allocate storage for
n_samplesitems of shapeEs....capacity()isn_samplesrounded up tokCapacityAlignment— unconditional, even for a fresh Array: growth later must never re-stride an Array that never grew.
-
inline std::size_t samples() const#
Number of samples (the batch mode’s current extent — vector-
size()-like; unaffected by any reserved headroom).
-
inline std::size_t size() const#
Total element count —
(Es * ...) * samples().
-
inline std::size_t capacity() const#
How many samples the backing
Chunk(s) actually hold room for — always a multiple ofkCapacityAlignment.
-
inline void reserve(std::size_t n)#
Grow
capacity()to at leastn(quantised up tokCapacityAlignment) — a no-op whenn <= capacity(). Otherwise reallocates the backingChunk(s) and copies the existingsamples()real elements, pitch-aware (component c’ssamples()elements move from offsetc*capacity()toc*newCapacity), preservingsamples().data()pointer identity is not preserved across a realloc (host-throw on allocation failure — fromChunk::allocate).
-
inline void resize(std::size_t n)#
Set
samples()ton— no reallocation (pointer/stride stable) whilen <= capacity(); otherwisereserve(n)first (realloc + pitched copy of the oldsamples()real elements), then adoptsnas the newsamples(). Shrinking (n < samples()) never reallocates and never releasescapacity().
-
inline void push_back(const Item<T, Es...> &item)#
Append one
Item<T,Es...>, host-side, amortised-doublingreserve()whensamples() == capacity()(a fresh, zero-capacity Array grows tokCapacityAlignment). Writes through the host chunk only — mirrorsupload()/download()’s explicit-transfer contract;upload()afterwards to reach the device copy in CUDA mode.
-
inline void clear()#
samples() <- 0;capacity()is kept (no deallocation).
-
inline ViewT hostView()#
A writable view over the host-resident copy — shape
samples(), pitchcapacity()(shape from size, strides from capacity). Sample extent is0for asamples() == 0Array — any per-sample expression-template assignment over it is a no-op by construction (nothing to iterate).
-
inline ConstViewT hostView() const#
A read-only view over the host-resident copy (see
hostView()).
-
inline PackedViewT hostPacked()#
The compact
layout_rightaccessor, available only whencapacity() == samples()(a full, non-reserving Array).- Throws:
aether::Error – if
capacity() != samples().
-
inline ConstPackedViewT hostPacked() const#
This is an overloaded member function, provided for convenience. It differs from the above function only in what argument(s) it accepts.
-
inline ViewT hostSpare()#
A writable view over the reserved tail
[samples(), capacity())of the host-resident copy — the target a caller (or, on the device side,AppendCounter) writes new samples into beforecommit()ing them.
-
inline ConstViewT hostSpare() const#
This is an overloaded member function, provided for convenience. It differs from the above function only in what argument(s) it accepts.
-
inline void commit(std::size_t count)#
Adopt
countnewly-written spare samples —samples() <- samples() + count(host-side; bounds-checked against spare capacity, host-throw on overflow). The device-side atomic append counter that producescountis not here.
-
inline void upload()#
Blocking host -> device copy. No-op in
AETHER_CPP_MODE.
-
inline void download()#
Blocking device -> host copy. No-op in
AETHER_CPP_MODE.
-
using ViewT = View<T, Extents, layout_stride>#