Multi-device residency#
How do I keep a replica per GPU, or partition data across several?
Use aether::ReplicaSet when you want the SAME data copied onto every
device in a list (rs.broadcast(src)); use aether::partition_view
instead when you want to slice ONE view’s sample dimension into several
non-overlapping blocks, one per device or worker, without allocating or
copying anything at all.
Replicating: ReplicaSet#
aether::ReplicaSet (aether/residency/Replica.h) holds one
aether::Chunk per Device in an explicit list you pass at
construction — there is no automatic device discovery or topology probing
in this v1 surface. rs.chunk(i) accesses the i-th replica’s
Chunk directly, and rs.broadcast(source) copies one source
Chunk’s bytes into every replica, bit-exactly, using
aether::copy/copyAsync under the hood (see
Memory ownership: Chunk, Array, View for the legal device-pair rules that copy obeys).
Partitioning: partition_view#
aether::PartitionSpec / aether::padded_block_samples /
aether::partition_view (aether/residency/Partition.h) instead slice
an EXISTING batched View’s trailing sample dimension into parts
blocks, each block’s start rounded up to a pad_to-sample granularity
(32 by default) — “real extent, padded stride”: every block but possibly
the last has exactly the padded number of real samples, while block
BOUNDARIES still fall at regular, padded intervals regardless of how full
the last one actually is. No allocation and no copy happens here either —
partition_view slices the original view’s own memory in place.
The transport seam#
aether::transport (aether/residency/Transport.h) is a small C++20
concept — a type offering a blocking copy and (on a CUDA build) an
async copyAsync between two Chunks — that ReplicaSet-style
data movement is built against. aether::StreamTransport is the one
implementation this library ships (a thin wrapper over
aether::copy/copyAsync); multi-node transports (NCCL, MPI, a
peer-to-peer fabric) are future consumers of the SAME concept, not
something this page’s surface provides today.
A runnable example#
Three CPU-kind devices standing in for three distinct devices, each
getting its own 256-byte Chunk:
TEST_F(ReplicaTest, ConstructsOneChunkPerDeviceWithTheRequestedSize)
{
// Distinct DEVICE IDS stand in for distinct devices — Chunk::allocate
// dispatches on device KIND, not id, so this exercises the "N chunks on
// N devices" contract entirely within an AETHER_CPP_MODE (kDLCPU-only)
// build.
std::vector<aether::Device> devices{
aether::Device(kDLCPU, 0), aether::Device(kDLCPU, 1), aether::Device(kDLCPU, 2)
};
aether::ReplicaSet rs(devices, 256);
EXPECT_EQ(rs.size(), 3u);
for (std::size_t i = 0; i < devices.size(); ++i) {
EXPECT_EQ(rs.chunk(i).size(), 256u);
EXPECT_EQ(rs.chunk(i).device(), devices[i]);
EXPECT_NE(rs.chunk(i).data(), nullptr);
EXPECT_TRUE(rs.chunk(i).owns());
}
}
On an actual multi-GPU build, aether::Device(kDLCUDA, 0),
aether::Device(kDLCUDA, 1), … would name the real devices; the CPU
build shown here exercises the identical code path against distinct
device IDs of the same kind.