API reference — backend and interop

Contents

API reference — backend and interop#

The device-residency surface: CPU SIMD packets, the CUDA multi-lane carrier, the DLPack interop bridge, and multi-device/multi-stream residency. See CPU SIMD backend, Multi-device residency and Interoperability.

CPU SIMD (aether::simd)#

template<class DataT, std::size_t Width>
struct Packet#
template<class DataT, std::size_t Width = PreferredWidth<DataT>>
struct PacketMask#

SIMD lane mask — primary template for Width > 1.

Template Parameters:
  • DataT – The scalar type whose packet this mask accompanies.

  • Width – Number of SIMD lanes (default: PreferredWidth<DataT>).

template<class DataT>
struct PacketTraits#

Compile-time SIMD traits bundle for DataT.

Public Static Attributes

static constexpr bool HasSIMD = (Width > 1)#

true when a real (non-scalar-fallback) SIMD width is active.

aether::DeviceBundle (CUDA)#

template<class T, std::size_t W, std::size_t... Es>
class DeviceBundle : public aether::Expression<DeviceBundle<T, W, Es...>, T>#

Register-resident bundle of W samples’ worth of extents<Es...> (L3): one Carray<T,W>-shaped lane-group per element of element_extents, Size elements total. Rebases onto the Expression CRTP exactly like Item/PacketItem (isLeaf = true).

Public Types

using element_extents = extents<Es...>#

L3 element protocol: a DeviceBundle’s own shape IS its element shape.

Public Functions

template<std::size_t... Is>
inline constexpr T &at(std::size_t lane)#

Lane k’s scalar for component Is... (compile-time multi-index) — read/write.

template<std::size_t... Is>
inline constexpr const T &at(std::size_t lane) const#

This is an overloaded member function, provided for convenience. It differs from the above function only in what argument(s) it accepts.

template<std::size_t... Is>
inline constexpr T &eval(const SampleIndex &i)#

Scalar element-protocol hook (see file docstring): i.work() selects the LANE, Is... selects the element — every existing expression node composes over DeviceBundle operands through this, UNCHANGED.

template<std::size_t... Is>
inline constexpr const T &eval(const SampleIndex &i) const#

This is an overloaded member function, provided for convenience. It differs from the above function only in what argument(s) it accepts.

inline constexpr T *data()#

Raw pointer to the underlying storage (element-major, lane-minor — see file docstring).

Public Static Functions

static inline constexpr std::size_t size()#

Total scalar count, Size * W.

Public Static Attributes

static constexpr bool isLeaf = true#

L3 element protocol: DeviceBundle is always a leaf.

static constexpr std::size_t Rank = sizeof...(Es)#

Number of modes.

static constexpr std::size_t Size = (Es * ... * std::size_t{1})#

Total element count — the product of Es... (excludes the W lane factor).

static constexpr std::size_t width = W#

Samples bundled per thread/iteration.

Namespace aether::interop (DLPack bridge)#

namespace interop#

Enums

enum class Access#

How a buffer may be accessed, as its producer declared it.

Values:

enumerator ReadWrite#

The producer allows writes through the buffer.

enumerator ReadOnly#

The producer declared the buffer read-only.

enumerator Unknown#

The producer’s protocol carries no access flag (legacy DLPack).

Functions

inline const char *accessName(Access a)#

The shared spelling of an Access value: "read-write", "read-only" or "unknown".

inline void assumeWritable(BufferView &buffer)#

Record the caller’s assertion that buffer may be written. The reported access is unchanged; writable() honours the assertion only for Access::Unknown (a declared read-only buffer stays read-only).

inline BufferView ownedBuffer(std::string owner, const RuntimeView &view, Access access, std::shared_ptr<void> keepAlive = {})#

A view of memory a library of this family allocated (owner is that library’s name, producer stays empty).

inline BufferView importDLPack(DLManagedTensorVersioned *managed, std::string producer = "dlpack")#

Import a versioned DLPack tensor. Access follows its read-only flag. Ownership of managed passes to the returned view’s keep-alive; on refusal the deleter runs once before the throw.

Throws:

aether::Error – null tensor, a newer major DLPack version, or any rejection of fromDLPack other than a non-zero byte_offset.

inline BufferView importDLPack(DLManagedTensor *managed, std::string producer = "dlpack")#

Import a legacy (pre-1.0) DLPack tensor. Its protocol carries no access flag, so access is Access::Unknown. Same ownership and refusal contract as the versioned overload.

inline BufferView fromArrayInterface(const ArrayInterface &ai)#

Build a BufferView from raw array-interface fields. Byte strides are converted to element strides (they must be multiples of the item size); access follows readOnly. The resulting record is validated exactly like a DLPack import.

Throws:

aether::Error – malformed or unsupported typestr, byte strides that are not a multiple of the item size, or any fromDLPack rejection.

inline std::string dtypeName(const DType &dt)#

A dtype’s readable name ("float64", "int32", "bool", …).

inline std::string deviceTypeName(DLDeviceType t)#

A device type’s readable name ("cpu", "cuda", …).

inline std::vector<std::string> refusals(const BufferView &buffer, const Requirements &req)#

Every refusal requirements raises against buffer, one message per failed check. Each message names the check and both values, in the form "<check>: required <want>, got <have>". Empty when the buffer satisfies every requirement.

inline void require(const BufferView &buffer, const Requirements &req)#

Throw one aether::Error carrying every refusal (joined by "; ") when buffer fails requirements; a no-op otherwise.

inline void requireHostAccess(const BufferView &buffer, const char *op)#

Refuse, naming op, an operation that reads or writes buffer’s memory from the host when that memory is not host-addressable (hostAccessible). Importing, inspecting and re-exporting a record never needs this; dereferencing its data does.

inline DLManagedTensorVersioned *exportDLPack(const BufferView &buffer)#

Export buffer as a heap-allocated versioned DLPack tensor, zero-copy. The read-only flag is set unless the buffer is writable (BufferView::writable()), so an Unknown buffer exports as read-only unless writability was asserted. The tensor’s context holds a share of the keep-alive: the producer stays alive until the consumer calls the tensor’s deleter AND every other view is gone. The caller (or the framework it hands the tensor to) owns the result and calls its deleter once.

inline DLManagedTensor *exportDLPackLegacy(const BufferView &buffer)#

Export buffer as a heap-allocated legacy (pre-1.0) DLPack tensor, zero-copy. The legacy struct has no access flag, so a buffer declared read-only is REFUSED rather than exported as if it were writable (the same rule numpy applies); an Unknown buffer exports unchanged (its consumer sees it as unknown again). bool follows toDLPackLegacy’s kDLUInt policy. Same keep-alive contract as exportDLPack.

Throws:

aether::Error – a read-only buffer, or an unsupported rank, dtype or device.

inline bool hostAccessible(DLDeviceType type)#

true when the host may dereference memory of device type type: kDLCPU, and the page-locked host kinds (kDLCUDAHost, kDLROCMHost) whose address is a host address. Device, managed and every other kind are records only on the host.

inline DLPackImport fromDLPack(DLManagedTensor *managed)#

Import a LEGACY DLManagedTensor*. See the file docstring for ownership/stream-semantics and the rejection list.

inline DLPackImport fromDLPack(DLManagedTensorVersioned *managed)#

Import a DLManagedTensorVersioned*

(“the current standard

DLPack exchange data structure”). Same contract as the legacy overload above.

inline DLManagedTensorVersioned *toDLPack(const RuntimeView &view)#

Export a RuntimeView as a heap-allocated DLManagedTensorVersioned* the caller (or whatever framework it hands the pointer to) owns the result and must eventually call its deleter. strides are exported VERBATIM from view.strides (element units — SoA and any other stride pattern this library can represent round-trips exactly, never forced back to a compact/contiguous shape).

Throws:

aether::Error – if view’s dtype or device is not one this library recognizes (defensive — every RuntimeView this library itself produces already satisfies both, via fromView/fromDLPack).

template<class T, class Extents, class Layout, bool Volatile>
DLManagedTensorVersioned *toDLPack(const View<T, Extents, Layout, Volatile> &v)#

toDLPack(view) overload for a static View — converts via aether::fromView first (host-only, always succeeds).

inline DLManagedTensor *toDLPackLegacy(const RuntimeView &view)#

Export a RuntimeView as a heap-allocated pre-1.0 DLManagedTensor*. The caller (or whatever framework it hands the pointer to) owns the result and must eventually call its deleter — identical ownership contract to aether::interop::toDLPack.

Throws:

aether::Error – same rejection set as toDLPack (unsupported rank/ dtype/device for this build) — defensive, since every RuntimeView this library itself produces already satisfies all three.

template<class T, class Extents, class Layout, bool Volatile>
DLManagedTensor *toDLPackLegacy(const View<T, Extents, Layout, Volatile> &v)#

toDLPackLegacy(view) overload for a static View — converts via aether::fromView first (host-only, always succeeds), mirroring aether::interop::toDLPack’s own static-View overload exactly.

Variables

constexpr const char *kExternalOwner = "external"#

The owner name of borrowed memory.

struct ArrayInterface#
#include <Buffer.h>

The raw fields of a __cuda_array_interface__ (device) or __array_interface__ (host) dictionary, so a language binding can build a BufferView without DLPack.

Public Members

void *data = nullptr#

data[0]: the buffer address.

bool readOnly = false#

data[1]: the producer’s read-only flag.

std::vector<std::int64_t> shape#

shape, in elements.

std::vector<std::int64_t> strides#

strides in BYTES, or empty for C-contiguous (the protocol’s None).

std::string typestr#

typestr, e.g. "<f8", "|u1", "|b1".

DLDevice device = {kDLCPU, 0}#

Where the memory lives (kDLCUDA + ordinal for the device protocol, kDLCPU for the host one).

std::string producer#

The producer’s type name.

std::shared_ptr<void> keepAlive#

Whatever keeps the producer alive (e.g. a reference to the object).

struct BufferView#
#include <Buffer.h>

A zero-copy buffer plus its access, ownership and lifetime record.

view.data already includes the producer’s byte_offset. Copies share the keep-alive, so any copy keeps the producer alive.

Public Functions

inline bool owned() const#

true when the memory belongs to a library of this family.

inline bool writable() const#

true when writes are allowed: declared read-write, or unknown with an explicit writability assertion.

Public Members

RuntimeView view = {}#

The zero-copy descriptor (element-unit strides).

Access access = Access::Unknown#

Access as the producer declared it.

bool assumedWritable = false#

true once a caller asserted writability via assumeWritable.

std::string owner = kExternalOwner#

The allocating library’s name, or kExternalOwner.

std::string producer#

The producer’s type name (e.g. "numpy.ndarray"); empty when owned.

std::shared_ptr<void> keepAlive#

Keeps the producer’s memory alive while any copy of this view lives.

struct DLPackImport#
#include <DLPack.h>

Result of a DLPack import: the zero-copy RuntimeView plus the DLPackOwner whose lifetime must dominate the view’s use.

class DLPackOwner#
#include <DLPack.h>

RAII ownership token for an imported DLPack tensor: calls the producer’s deleter exactly once. Move-only (mirrors Chunk’s ownership discipline) — the RuntimeView a DLPackImport carries remains valid only as long as its DLPackOwner is alive (or until release() is called).

Public Functions

inline void release()#

Invoke the producer’s deleter now — idempotent (a no-op once already released, or on a default-constructed token). See the file docstring’s stream-semantics contract: the CALLER must order any outstanding async device work before this runs.

inline bool owns() const#

true if this token still owns a live tensor (has not been released, moved-from, or default-constructed).

struct Requirements#
#include <Buffer.h>

What a consumer needs from a buffer. Every unset field is not checked. writable is the consumer’s policy switch: set it for a buffer the consumer writes through.

Public Members

std::optional<DType> dtype#

Required element dtype.

std::optional<std::int64_t> count#

Required element count (product of the shape).

std::optional<std::vector<std::int64_t>> shape#

Required shape, in elements.

bool contiguous = false#

Require C-contiguous storage with a unit innermost stride.

std::size_t alignment = 0#

Required byte alignment of the data pointer (0 = unchecked).

std::optional<DLDeviceType> deviceType#

Required device type (e.g. kDLCUDA, kDLCPU).

std::optional<std::int32_t> deviceId#

Required device ordinal.

bool writable = false#

Require writes to be allowed (BufferView::writable()).

namespace detail#

Functions

inline RuntimeView bufferViewFromDLTensor(const DLTensor &t)#

The validating core of both DLPack imports, with byte_offset folded into the data pointer.

inline DLDataType dtypeFromTypestr(const std::string &ts)#

Parse an array-interface typestr into a DLPack dtype.

inline std::string tupleText(const std::int64_t *v, std::size_t n)#
inline bool isCContiguous(const RuntimeView &v)#
inline BufferExportCtx *makeExportCtx(const BufferView &b)#
inline void fillTensor(DLTensor &t, const BufferView &b, BufferExportCtx *ctx, DLDataType dtype)#
inline void checkExportable(const BufferView &b, const char *op)#
inline bool dtypeIsSupported(DLDataType dt)#
inline bool deviceIsKnown(DLDeviceType type)#

true for a device type an import/export of THIS build accepts.

A CUDA build accepts host memory and the two CUDA kinds it can operate on (kDLCUDA, kDLCUDAHost). A pure C++ build (AETHER_CPP_MODE, no AETHER_HAS_CUDA) accepts every device type DLPack defines as METADATA: the record is imported and exported unchanged, device memory is never dereferenced, and only an operation that needs host access refuses it (hostAccessible). That lets a pure C++ consumer hold a CUDA buffer that a separately loaded device backend operates on. A code DLPack does not define is refused in both builds.

inline RuntimeView runtimeViewFromDLTensor(const DLTensor &t)#

Convert one DLTensor (the shared payload of both DLPack struct generations) into a RuntimeView — the validating core BOTH fromDLPack overloads share. Throws aether::Error on any rejection; never touches ownership (the caller wraps this in a try/catch that calls the producer’s deleter on the throw path — see the two fromDLPack overloads below).

inline void deleteExported(DLManagedTensorVersioned *self)#
inline DLDataType legacyDType(const DType &dt)#

The legacy bool policy — see the file docstring. Every other dtype passes through unchanged (identical to what toDLPack exports for the same RuntimeView).

inline void deleteLegacyExported(DLManagedTensor *self)#
struct BufferExportCtx#
#include <Buffer.h>

Context of a BufferView export: shape/strides storage plus a share of the view’s keep-alive.

struct ExportCtx#
#include <DLPack.h>

Heap-owned shape/strides storage for an EXPORTED tensor — DLTensor::shape/strides are raw int64_t*, so something must own that storage for as long as the exported tensor lives; the exported DLManagedTensorVersioned::manager_ctx points at one of these, and its deleter frees both this and the tensor itself.

struct LegacyExportCtx#
#include <DLPackLegacy.h>

Heap-owned shape/strides storage for a LEGACY exported tensor — same rationale as DLPack.h’s own ExportCtx (a DLTensor’s shape/strides are raw pointers that must outlive this call’s stack frame). A SEPARATE type from DLPack.h’s ExportCtx only because the two live in the same detail namespace and each needs its own deleteExported overload keyed to its own tensor generation — the field shapes are otherwise identical.

Multi-device residency#

struct PartitionSpec#

Partition parameters: split the SAMPLE mode into parts blocks, each block pitch rounded up to a pad_to-sample granularity (default 32).

template<class T, std::size_t... Es>
class PartitionedArray#

Single-node multi-GPU partitioned array: one Chunk per device in devices (v1: an explicit list), holding rank r’s PartitionSpec-derived, PADDED-pitch block of a logical n-sample array whose canonical (compact) copy lives in a single host pinned Chunk.

Public Types

using ViewT = View<T, Extents, layout_right>#

The host canonical (COMPACT, n samples) view type — hostView()’s return type.

using DeviceViewT = View<T, Extents, layout_stride>#

deviceView(r)’s return type — see its own docstring for why this is layout_stride, not layout_right.

Public Functions

inline PartitionedArray(std::vector<Device> devices, std::size_t n, std::size_t pad_to = 32)#

Allocate the host pinned chunk (n samples, compact) and one device Chunk per entry of devices (padded_block_samples(n, devices.size(), pad_to) samples each, uniform).

Throws:

aether::Error – if devices.empty() (via padded_block_samples’s own parts must be > 0 check) or pad_to == 0, or on any underlying Chunk::allocate failure (unsupported device kind for this build included — e.g. a kDLCUDA device in an AETHER_CPP_MODE build).

inline std::size_t parts() const#

Number of ranks (== devices.size() at construction).

inline std::size_t samples() const#

Logical (unpadded) total sample count.

inline std::size_t padded() const#

The uniform per-rank PADDED pitch (samples), Partition.h’s closed form.

inline std::size_t realCount(std::size_t r) const#

Rank r’s REAL (valid) sample count — SHORT only for a trailing block.

inline Device device(std::size_t r) const#

Rank r’s Device.

inline ViewT hostView()#

A writable view over the host-resident CANONICAL (compact, n samples) copy — fill it before scatter(), read it after gather().

inline DeviceViewT deviceView(std::size_t r)#

A writable view over rank r’s own device-resident block — realCount(r) samples (never the padding tail; SHORT only for a trailing block).

NOT a compact layout_right view at realCount(r): deviceChunks_[r] is physically laid out at the UNIFORM padded_ pitch (every leading mode’s row-strip is padded_ elements apart — what scatter()/ gather() actually write, copyRows_’s own dstPitchElems/ srcPitchElems). For innerSize_ == 1 (a plain rank-1 array) a compact-at-real view and a padded-stride-at-real view address IDENTICAL bytes (there is no leading mode to mis-stride), so the two shapes only diverge for innerSize_ > 1 WHEN a block is genuinely short (real < padded_) — exactly the case test_PartitionedArray.cpp’s rank-2 (innerSize_ == 3) degenerate- path test exists to catch (it did: a first cut here returned a compact make_view<T,Es...>(deviceChunks_[r], realCount(r)), which silently read component 1’s data from the WRONG offset whenever real != padded_).

inline void scatter()#

Blocking H->each device’s block: every rank’s realCount(r) real samples, never the padding tail.

inline void gather()#

Blocking each device’s block -> H.

template<transport Transport>
inline void moveTo(std::size_t r, Device device, Transport &t)#

D2D: relocate rank r’s ENTIRE device-resident block (the full padded_-pitch storage, real data plus any padding tail) onto device, through transport — never touches hostChunk_. Blocking. r’s Device becomes device; the OLD chunk is freed once the move completes.

Transport need only satisfy aether::transport (constrained here via requires, mirroring Replica<T,Transport>’s own constraint) — StreamTransport’s route(a, b) decides SAME/P2P/STAGED/HOST per aether/residency/PeerAccess.h; in an AETHER_CPP_MODE build every device is kDLCPU, so this compiles and moves data through the ordinary CPU<->CPU aether::copy path (peerRoute always HOST).

Throws:

aether::Error – if r >= parts().

template<transport Transport>
inline void exchange(std::size_t ra, std::size_t rb, Transport &t)#

D2D: swap ranks ra/rb’s device-resident block CONTENTS (the full padded_-pitch storage) through transport — each rank’s OWN Device is UNCHANGED (devices_[ra]/devices_[rb] do not move); only the bytes resident there do. Blocking. A no-op when ra == rb.

Throws:

aether::Error – if ra >= parts() or rb >= parts().

class ReplicaSet#

N Chunks of bytes bytes, one per Device in devices. Move-only (implicit — std::vector<Chunk> cannot be copied since Chunk itself is move-only, so ReplicaSet’s copy members are implicitly deleted and its move members implicitly usable).

Public Functions

inline ReplicaSet(const std::vector<Device> &devices, std::size_t bytes, std::size_t alignment = 256)#

Allocate one Chunk of bytes bytes (aligned to alignment) on each device in devices, in order (v1: an explicit Device list).

inline std::size_t size() const#

Number of replicas.

inline Chunk &chunk(std::size_t i)#

The i-th replica’s chunk.

inline const Chunk &chunk(std::size_t i) const#

The i-th replica’s chunk (read-only).

inline void broadcast(const Chunk &src)#

Blocking broadcast: src’s bytes into EVERY replica in this set, via aether::copy — legal device-pair paths only (aether/chunk/Copy.h’s matrix; CPU<->CUDA direct throws).

class StreamTransport#

The peer-access-aware transport implementation. Route::STAGED moves go through this transport’s own pinned BounceBuffer (aether/residency/PeerAccess.h) explicitly; every other route delegates verbatim to aether::copy/aether::copyAsync (aether/chunk/Copy.h) — no behavior of its own beyond that, so it still satisfies the concept trivially.

A BounceBuffer and any recorded cudaEvents are per-instance (not process-global, unlike PeerAccess.h’s peer-access cache) — reusing one StreamTransport for two moves whose STAGED legs may still be in flight on different streams races on the shared bounce chunk; issuing every Route::STAGED move for one transport instance on a single stream (or fully synchronizing between them) is the safe usage this type assumes.

Public Functions

inline explicit StreamTransport(bool forceStaging)#

forceStaging=true routes every CUDA<->CUDA pair (same device included) through Route::STAGED — a test-only override.

inline Route route(const Device &a, const Device &b) const#

This transport’s Route decision for a move from a to b (aether::peerRoute, honoring forceStaging()).

inline bool forceStaging() const#

true when this transport forces every CUDA<->CUDA pair through Route::STAGED.

inline void copy(Chunk &dst, const Chunk &src)#

Blocking copy: Route::STAGED goes device -> pinned bounce -> device explicitly through this transport’s own BounceBuffer; every other route (SAME, P2P, HOST) delegates verbatim to aether::copy.