API reference — backend and interop#
The device-residency surface: CPU SIMD packets, the CUDA multi-lane carrier, the DLPack interop bridge, and multi-device/multi-stream residency. See CPU SIMD backend, Multi-device residency and Interoperability.
CPU SIMD (aether::simd)#
-
template<class DataT, std::size_t Width>
struct Packet#
aether::DeviceBundle (CUDA)#
-
template<class T, std::size_t W, std::size_t... Es>
class DeviceBundle : public aether::Expression<DeviceBundle<T, W, Es...>, T># Register-resident bundle of
Wsamples’ worth ofextents<Es...>(L3): oneCarray<T,W>-shaped lane-group per element ofelement_extents,Sizeelements total. Rebases onto theExpressionCRTP exactly likeItem/PacketItem(isLeaf = true).Public Types
-
using element_extents = extents<Es...>#
L3 element protocol: a
DeviceBundle’s own shape IS its element shape.
Public Functions
-
template<std::size_t... Is>
inline constexpr T &at(std::size_t lane)# Lane
k’s scalar for componentIs...(compile-time multi-index) — read/write.
-
template<std::size_t... Is>
inline constexpr const T &at(std::size_t lane) const# This is an overloaded member function, provided for convenience. It differs from the above function only in what argument(s) it accepts.
-
template<std::size_t... Is>
inline constexpr T &eval(const SampleIndex &i)# Scalar element-protocol hook (see file docstring):
i.work()selects the LANE,Is...selects the element — every existing expression node composes overDeviceBundleoperands through this, UNCHANGED.
-
template<std::size_t... Is>
inline constexpr const T &eval(const SampleIndex &i) const# This is an overloaded member function, provided for convenience. It differs from the above function only in what argument(s) it accepts.
Public Static Functions
-
static inline constexpr std::size_t size()#
Total scalar count,
Size * W.
Public Static Attributes
-
static constexpr bool isLeaf = true#
L3 element protocol:
DeviceBundleis always a leaf.
-
using element_extents = extents<Es...>#
Namespace aether::interop (DLPack bridge)#
-
namespace interop#
Enums
-
enum class Access#
How a buffer may be accessed, as its producer declared it.
Values:
-
enumerator ReadWrite#
The producer allows writes through the buffer.
-
enumerator ReadOnly#
The producer declared the buffer read-only.
-
enumerator Unknown#
The producer’s protocol carries no access flag (legacy DLPack).
-
enumerator ReadWrite#
Functions
-
inline const char *accessName(Access a)#
The shared spelling of an
Accessvalue:"read-write","read-only"or"unknown".
-
inline void assumeWritable(BufferView &buffer)#
Record the caller’s assertion that
buffermay be written. The reportedaccessis unchanged;writable()honours the assertion only forAccess::Unknown(a declared read-only buffer stays read-only).
A view of memory a library of this family allocated (
owneris that library’s name,producerstays empty).
-
inline BufferView importDLPack(DLManagedTensorVersioned *managed, std::string producer = "dlpack")#
Import a versioned DLPack tensor. Access follows its read-only flag. Ownership of
managedpasses to the returned view’s keep-alive; on refusal the deleter runs once before the throw.- Throws:
aether::Error – null tensor, a newer major DLPack version, or any rejection of
fromDLPackother than a non-zerobyte_offset.
-
inline BufferView importDLPack(DLManagedTensor *managed, std::string producer = "dlpack")#
Import a legacy (pre-1.0) DLPack tensor. Its protocol carries no access flag, so
accessisAccess::Unknown. Same ownership and refusal contract as the versioned overload.
-
inline BufferView fromArrayInterface(const ArrayInterface &ai)#
Build a
BufferViewfrom raw array-interface fields. Byte strides are converted to element strides (they must be multiples of the item size); access followsreadOnly. The resulting record is validated exactly like a DLPack import.- Throws:
aether::Error – malformed or unsupported
typestr, byte strides that are not a multiple of the item size, or anyfromDLPackrejection.
-
inline std::string dtypeName(const DType &dt)#
A dtype’s readable name (
"float64","int32","bool", …).
-
inline std::string deviceTypeName(DLDeviceType t)#
A device type’s readable name (
"cpu","cuda", …).
-
inline std::vector<std::string> refusals(const BufferView &buffer, const Requirements &req)#
Every refusal
requirementsraises againstbuffer, one message per failed check. Each message names the check and both values, in the form"<check>: required <want>, got <have>". Empty when the buffer satisfies every requirement.
-
inline void require(const BufferView &buffer, const Requirements &req)#
Throw one
aether::Errorcarrying every refusal (joined by"; ") whenbufferfailsrequirements; a no-op otherwise.
-
inline void requireHostAccess(const BufferView &buffer, const char *op)#
Refuse, naming
op, an operation that reads or writesbuffer’s memory from the host when that memory is not host-addressable (hostAccessible). Importing, inspecting and re-exporting a record never needs this; dereferencing its data does.
-
inline DLManagedTensorVersioned *exportDLPack(const BufferView &buffer)#
Export
bufferas a heap-allocated versioned DLPack tensor, zero-copy. The read-only flag is set unless the buffer is writable (BufferView::writable()), so anUnknownbuffer exports as read-only unless writability was asserted. The tensor’s context holds a share of the keep-alive: the producer stays alive until the consumer calls the tensor’s deleter AND every other view is gone. The caller (or the framework it hands the tensor to) owns the result and calls its deleter once.
-
inline DLManagedTensor *exportDLPackLegacy(const BufferView &buffer)#
Export
bufferas a heap-allocated legacy (pre-1.0) DLPack tensor, zero-copy. The legacy struct has no access flag, so a buffer declared read-only is REFUSED rather than exported as if it were writable (the same rule numpy applies); anUnknownbuffer exports unchanged (its consumer sees it as unknown again).boolfollowstoDLPackLegacy’skDLUIntpolicy. Same keep-alive contract asexportDLPack.- Throws:
aether::Error – a read-only buffer, or an unsupported rank, dtype or device.
-
inline bool hostAccessible(DLDeviceType type)#
truewhen the host may dereference memory of device typetype:kDLCPU, and the page-locked host kinds (kDLCUDAHost,kDLROCMHost) whose address is a host address. Device, managed and every other kind are records only on the host.
-
inline DLPackImport fromDLPack(DLManagedTensor *managed)#
Import a LEGACY
DLManagedTensor*. See the file docstring for ownership/stream-semantics and the rejection list.
-
inline DLPackImport fromDLPack(DLManagedTensorVersioned *managed)#
Import a
DLManagedTensorVersioned*(“the current standard
DLPack exchange data structure”). Same contract as the legacy overload above.
-
inline DLManagedTensorVersioned *toDLPack(const RuntimeView &view)#
Export a
RuntimeViewas a heap-allocatedDLManagedTensorVersioned*the caller (or whatever framework it hands the pointer to) owns the result and must eventually call itsdeleter.stridesare exported VERBATIM fromview.strides(element units — SoA and any other stride pattern this library can represent round-trips exactly, never forced back to a compact/contiguous shape).- Throws:
aether::Error – if
view’s dtype or device is not one this library recognizes (defensive — everyRuntimeViewthis library itself produces already satisfies both, viafromView/fromDLPack).
-
template<class T, class Extents, class Layout, bool Volatile>
DLManagedTensorVersioned *toDLPack(const View<T, Extents, Layout, Volatile> &v)# toDLPack(view)overload for a staticView— converts viaaether::fromViewfirst (host-only, always succeeds).
-
inline DLManagedTensor *toDLPackLegacy(const RuntimeView &view)#
Export a
RuntimeViewas a heap-allocated pre-1.0DLManagedTensor*. The caller (or whatever framework it hands the pointer to) owns the result and must eventually call itsdeleter— identical ownership contract toaether::interop::toDLPack.- Throws:
aether::Error – same rejection set as
toDLPack(unsupported rank/ dtype/device for this build) — defensive, since everyRuntimeViewthis library itself produces already satisfies all three.
-
template<class T, class Extents, class Layout, bool Volatile>
DLManagedTensor *toDLPackLegacy(const View<T, Extents, Layout, Volatile> &v)# toDLPackLegacy(view)overload for a staticView— converts viaaether::fromViewfirst (host-only, always succeeds), mirroringaether::interop::toDLPack’s own static-Viewoverload exactly.
Variables
-
constexpr const char *kExternalOwner = "external"#
The owner name of borrowed memory.
-
struct ArrayInterface#
- #include <Buffer.h>
The raw fields of a
__cuda_array_interface__(device) or__array_interface__(host) dictionary, so a language binding can build aBufferViewwithout DLPack.Public Members
-
void *data = nullptr#
data[0]: the buffer address.
-
bool readOnly = false#
data[1]: the producer’s read-only flag.
-
std::vector<std::int64_t> shape#
shape, in elements.
-
std::vector<std::int64_t> strides#
stridesin BYTES, or empty for C-contiguous (the protocol’sNone).
-
std::string typestr#
typestr, e.g."<f8","|u1","|b1".
-
DLDevice device = {kDLCPU, 0}#
Where the memory lives (
kDLCUDA+ ordinal for the device protocol,kDLCPUfor the host one).
-
std::string producer#
The producer’s type name.
-
std::shared_ptr<void> keepAlive#
Whatever keeps the producer alive (e.g. a reference to the object).
-
void *data = nullptr#
-
struct BufferView#
- #include <Buffer.h>
A zero-copy buffer plus its access, ownership and lifetime record.
view.dataalready includes the producer’sbyte_offset. Copies share the keep-alive, so any copy keeps the producer alive.Public Functions
-
inline bool owned() const#
truewhen the memory belongs to a library of this family.
-
inline bool writable() const#
truewhen writes are allowed: declared read-write, or unknown with an explicit writability assertion.
Public Members
-
RuntimeView view = {}#
The zero-copy descriptor (element-unit strides).
-
bool assumedWritable = false#
trueonce a caller asserted writability viaassumeWritable.
-
std::string owner = kExternalOwner#
The allocating library’s name, or
kExternalOwner.
-
std::string producer#
The producer’s type name (e.g.
"numpy.ndarray"); empty when owned.
-
std::shared_ptr<void> keepAlive#
Keeps the producer’s memory alive while any copy of this view lives.
-
inline bool owned() const#
-
struct DLPackImport#
- #include <DLPack.h>
Result of a DLPack import: the zero-copy
RuntimeViewplus theDLPackOwnerwhose lifetime must dominate the view’s use.
-
class DLPackOwner#
- #include <DLPack.h>
RAII ownership token for an imported DLPack tensor: calls the producer’s
deleterexactly once. Move-only (mirrorsChunk’s ownership discipline) — theRuntimeViewaDLPackImportcarries remains valid only as long as itsDLPackOwneris alive (or untilrelease()is called).Public Functions
-
inline void release()#
Invoke the producer’s deleter now — idempotent (a no-op once already released, or on a default-constructed token). See the file docstring’s stream-semantics contract: the CALLER must order any outstanding async device work before this runs.
-
inline bool owns() const#
trueif this token still owns a live tensor (has not been released, moved-from, or default-constructed).
-
inline void release()#
-
struct Requirements#
- #include <Buffer.h>
What a consumer needs from a buffer. Every unset field is not checked.
writableis the consumer’s policy switch: set it for a buffer the consumer writes through.Public Members
-
std::optional<std::int64_t> count#
Required element count (product of the shape).
-
std::optional<std::vector<std::int64_t>> shape#
Required shape, in elements.
-
bool contiguous = false#
Require C-contiguous storage with a unit innermost stride.
-
std::size_t alignment = 0#
Required byte alignment of the data pointer (0 = unchecked).
-
std::optional<DLDeviceType> deviceType#
Required device type (e.g.
kDLCUDA,kDLCPU).
-
std::optional<std::int32_t> deviceId#
Required device ordinal.
-
bool writable = false#
Require writes to be allowed (
BufferView::writable()).
-
std::optional<std::int64_t> count#
-
namespace detail#
Functions
-
inline RuntimeView bufferViewFromDLTensor(const DLTensor &t)#
The validating core of both DLPack imports, with
byte_offsetfolded into the data pointer.
-
inline DLDataType dtypeFromTypestr(const std::string &ts)#
Parse an array-interface
typestrinto a DLPack dtype.
-
inline std::string tupleText(const std::int64_t *v, std::size_t n)#
-
inline bool isCContiguous(const RuntimeView &v)#
-
inline BufferExportCtx *makeExportCtx(const BufferView &b)#
-
inline void fillTensor(DLTensor &t, const BufferView &b, BufferExportCtx *ctx, DLDataType dtype)#
-
inline void checkExportable(const BufferView &b, const char *op)#
-
inline bool dtypeIsSupported(DLDataType dt)#
-
inline bool deviceIsKnown(DLDeviceType type)#
truefor a device type an import/export of THIS build accepts.A CUDA build accepts host memory and the two CUDA kinds it can operate on (
kDLCUDA,kDLCUDAHost). A pure C++ build (AETHER_CPP_MODE, noAETHER_HAS_CUDA) accepts every device type DLPack defines as METADATA: the record is imported and exported unchanged, device memory is never dereferenced, and only an operation that needs host access refuses it (hostAccessible). That lets a pure C++ consumer hold a CUDA buffer that a separately loaded device backend operates on. A code DLPack does not define is refused in both builds.
-
inline RuntimeView runtimeViewFromDLTensor(const DLTensor &t)#
Convert one
DLTensor(the shared payload of both DLPack struct generations) into aRuntimeView— the validating core BOTHfromDLPackoverloads share. Throwsaether::Erroron any rejection; never touches ownership (the caller wraps this in a try/catch that calls the producer’s deleter on the throw path — see the twofromDLPackoverloads below).
-
inline void deleteExported(DLManagedTensorVersioned *self)#
-
inline DLDataType legacyDType(const DType &dt)#
The legacy
boolpolicy — see the file docstring. Every other dtype passes through unchanged (identical to whattoDLPackexports for the sameRuntimeView).
-
inline void deleteLegacyExported(DLManagedTensor *self)#
-
struct BufferExportCtx#
- #include <Buffer.h>
Context of a
BufferViewexport: shape/strides storage plus a share of the view’s keep-alive.
-
struct ExportCtx#
- #include <DLPack.h>
Heap-owned
shape/stridesstorage for an EXPORTED tensor —DLTensor::shape/stridesare rawint64_t*, so something must own that storage for as long as the exported tensor lives; the exportedDLManagedTensorVersioned::manager_ctxpoints at one of these, and itsdeleterfrees both this and the tensor itself.
-
struct LegacyExportCtx#
- #include <DLPackLegacy.h>
Heap-owned
shape/stridesstorage for a LEGACY exported tensor — same rationale asDLPack.h’s ownExportCtx(aDLTensor’sshape/stridesare raw pointers that must outlive this call’s stack frame). A SEPARATE type fromDLPack.h’sExportCtxonly because the two live in the samedetailnamespace and each needs its owndeleteExportedoverload keyed to its own tensor generation — the field shapes are otherwise identical.
-
inline RuntimeView bufferViewFromDLTensor(const DLTensor &t)#
-
enum class Access#
Multi-device residency#
-
struct PartitionSpec#
Partition parameters: split the SAMPLE mode into
partsblocks, each block pitch rounded up to apad_to-sample granularity (default 32).
-
template<class T, std::size_t... Es>
class PartitionedArray# Single-node multi-GPU partitioned array: one
Chunkper device indevices(v1: an explicit list), holding rankr’sPartitionSpec-derived, PADDED-pitch block of a logicaln-sample array whose canonical (compact) copy lives in a single host pinnedChunk.Public Types
-
using ViewT = View<T, Extents, layout_right>#
The host canonical (COMPACT,
nsamples) view type —hostView()’s return type.
-
using DeviceViewT = View<T, Extents, layout_stride>#
deviceView(r)’s return type — see its own docstring for why this islayout_stride, notlayout_right.
Public Functions
-
inline PartitionedArray(std::vector<Device> devices, std::size_t n, std::size_t pad_to = 32)#
Allocate the host pinned chunk (
nsamples, compact) and one deviceChunkper entry ofdevices(padded_block_samples(n, devices.size(), pad_to)samples each, uniform).- Throws:
aether::Error – if
devices.empty()(viapadded_block_samples’s ownparts must be > 0check) orpad_to == 0, or on any underlyingChunk::allocatefailure (unsupported device kind for this build included — e.g. akDLCUDAdevice in anAETHER_CPP_MODEbuild).
-
inline std::size_t parts() const#
Number of ranks (==
devices.size()at construction).
-
inline std::size_t samples() const#
Logical (unpadded) total sample count.
-
inline std::size_t padded() const#
The uniform per-rank PADDED pitch (samples),
Partition.h’s closed form.
-
inline std::size_t realCount(std::size_t r) const#
Rank
r’s REAL (valid) sample count — SHORT only for a trailing block.
-
inline ViewT hostView()#
A writable view over the host-resident CANONICAL (compact,
nsamples) copy — fill it beforescatter(), read it aftergather().
-
inline DeviceViewT deviceView(std::size_t r)#
A writable view over rank
r’s own device-resident block —realCount(r)samples (never the padding tail; SHORT only for a trailing block).NOT a compact
layout_rightview atrealCount(r):deviceChunks_[r]is physically laid out at the UNIFORMpadded_pitch (every leading mode’s row-strip ispadded_elements apart — whatscatter()/gather()actually write,copyRows_’s owndstPitchElems/srcPitchElems). ForinnerSize_ == 1(a plain rank-1 array) a compact-at-realview and a padded-stride-at-realview address IDENTICAL bytes (there is no leading mode to mis-stride), so the two shapes only diverge forinnerSize_ > 1WHEN a block is genuinely short (real < padded_) — exactly the casetest_PartitionedArray.cpp’s rank-2 (innerSize_ == 3) degenerate- path test exists to catch (it did: a first cut here returned a compactmake_view<T,Es...>(deviceChunks_[r], realCount(r)), which silently read component 1’s data from the WRONG offset wheneverreal != padded_).
-
inline void scatter()#
Blocking H->each device’s block: every rank’s
realCount(r)real samples, never the padding tail.
-
inline void gather()#
Blocking each device’s block -> H.
-
template<transport Transport>
inline void moveTo(std::size_t r, Device device, Transport &t)# D2D: relocate rank
r’s ENTIRE device-resident block (the fullpadded_-pitch storage, real data plus any padding tail) ontodevice, throughtransport— never toucheshostChunk_. Blocking.r’sDevicebecomesdevice; the OLD chunk is freed once the move completes.Transportneed only satisfyaether::transport(constrained here viarequires, mirroringReplica<T,Transport>’s own constraint) —StreamTransport’sroute(a, b)decides SAME/P2P/STAGED/HOST peraether/residency/PeerAccess.h; in anAETHER_CPP_MODEbuild every device iskDLCPU, so this compiles and moves data through the ordinary CPU<->CPUaether::copypath (peerRoutealways HOST).
-
template<transport Transport>
inline void exchange(std::size_t ra, std::size_t rb, Transport &t)# D2D: swap ranks
ra/rb’s device-resident block CONTENTS (the fullpadded_-pitch storage) throughtransport— each rank’s OWNDeviceis UNCHANGED (devices_[ra]/devices_[rb]do not move); only the bytes resident there do. Blocking. A no-op whenra == rb.
-
using ViewT = View<T, Extents, layout_right>#
-
class ReplicaSet#
N
Chunks ofbytesbytes, one perDeviceindevices. Move-only (implicit —std::vector<Chunk>cannot be copied sinceChunkitself is move-only, soReplicaSet’s copy members are implicitly deleted and its move members implicitly usable).
-
class StreamTransport#
The peer-access-aware
transportimplementation.Route::STAGEDmoves go through this transport’s own pinnedBounceBuffer(aether/residency/PeerAccess.h) explicitly; every other route delegates verbatim toaether::copy/aether::copyAsync(aether/chunk/Copy.h) — no behavior of its own beyond that, so it still satisfies the concept trivially.A
BounceBufferand any recordedcudaEvents are per-instance (not process-global, unlikePeerAccess.h’s peer-access cache) — reusing oneStreamTransportfor two moves whose STAGED legs may still be in flight on different streams races on the shared bounce chunk; issuing everyRoute::STAGEDmove for one transport instance on a single stream (or fully synchronizing between them) is the safe usage this type assumes.Public Functions
-
inline explicit StreamTransport(bool forceStaging)#
forceStaging=trueroutes every CUDA<->CUDA pair (same device included) throughRoute::STAGED— a test-only override.
-
inline Route route(const Device &a, const Device &b) const#
This transport’s
Routedecision for a move fromatob(aether::peerRoute, honoringforceStaging()).
-
inline bool forceStaging() const#
truewhen this transport forces every CUDA<->CUDA pair throughRoute::STAGED.
-
inline explicit StreamTransport(bool forceStaging)#