Memory ownership: Chunk, Array, View#
Who owns this buffer, host or device, and how do I copy it?
An aether::Chunk owns (or, if you asked it to borrow foreign memory,
merely wraps) one byte allocation on exactly one Device, and frees it
itself when it goes out of scope; aether::copy/copyAsync move
bytes between two Chunks along a fixed set of legal device pairs, and
refuse — by throwing — anything outside that set.
Chunk: one allocation, one device#
aether::Chunk (aether/chunk/Chunk.h) is a move-only, host-only
management type: allocation and deallocation are always host-side calls,
even when the bytes themselves live on a GPU. Chunk::allocate(device,
bytes) owns what it allocates and frees it on destruction or move-away;
Chunk::borrow() instead wraps memory you already own (a
std::vector’s backing store, say) with a no-op deleter — owns()
reports false and the Chunk never touches that memory at teardown.
Either way, Chunk carries no element type or shape of its own; those
live in the layers built on top (see Views and Items).
Copying between devices#
aether::copy/aether::copyAsync (aether/chunk/Copy.h) move raw
bytes between two Chunks. The legal blocking paths are CPU↔CPU,
CPU↔pinned (kDLCUDAHost), pinned↔CUDA, and CUDA↔CUDA — a direct
CPU↔CUDA copy is deliberately illegal; you stage it through pinned memory
yourself. copyAsync takes an aether::Stream (see
Host vs device code for what a stream is) and is narrower still: only
the pinned/CUDA pairs are async-capable, since ordinary pageable host
memory cannot be safely copied without blocking.
Array: the owning convenience#
Quickstart already introduced aether::Array — this is what it
owns underneath. On a CUDA build an Array holds TWO Chunks, a
pinned host one and a device one, and .upload()/.download() are
what call copy/copyAsync between them; in a CPU-only build there
is only one Chunk, so .hostView() and .deviceView() alias the
SAME memory and upload/download become harmless no-ops — the same call
sequence works, and means the same thing conceptually, either way.
A runnable example#
TEST_F(ChunkTest, CpuAllocateAndFree)
{
aether::Device cpu(kDLCPU);
auto chunk = aether::Chunk::allocate(cpu, 256);
EXPECT_NE(chunk.data(), nullptr);
EXPECT_EQ(chunk.size(), 256u);
EXPECT_EQ(chunk.device(), cpu);
// Freed by the destructor at scope exit — no crash is the pass condition.
}
There is no explicit free() call to make: chunk releases its 256
bytes automatically when it goes out of scope, the same rule every
move-only RAII type in this library follows.