Host vs device code#
What does AETHER_DEVICEHOST mean, and when do I need a GPU kernel entry
point (__global__)?
AETHER_DEVICEHOST() marks an ordinary function so it compiles correctly
and runs correctly whether it is called from plain CPU code or from code
running on the GPU; you only reach for a kernel entry point when you need
to launch brand-new parallel work on the GPU from the CPU, one call
spawning thousands of GPU threads at once. Almost everything in aether —
View indexing, expression-template arithmetic, Item construction —
is written with AETHER_DEVICEHOST() precisely so the same source line
works unchanged on both sides; kernel entry points are the exception, not
the rule.
Vocabulary#
This is the one page where aether’s CUDA vocabulary gets defined; every later page assumes these terms.
Host — the CPU and its ordinary system memory: the code that runs when your program starts, allocates buffers and decides what work to do.
Device — the GPU and its own separate memory. Host code cannot read device memory directly (and vice versa) without an explicit copy.
Kernel — a function launched from the host to run on the device. A single kernel launch runs the SAME function body across many threads at once, each thread working on a different piece of data. A kernel is written with a kernel entry-point qualifier (CUDA’s
__global__) and cannot return a value directly to the caller.Thread / block / grid — a kernel launch is shaped as a grid of thread blocks, each block containing some number of threads. Every thread runs the kernel body once, over its own slice of the data; the block/grid shape is just bookkeeping for “how many threads, grouped how”.
Warp — on CUDA GPUs, threads within a block execute in hardware-scheduled groups of (typically) 32, called a warp. You rarely address a warp directly in aether-level code, but some device machinery is built around warp-sized groups.
Register-resident / lane — a value kept in a thread’s own fast, private on-chip storage (a register) rather than in memory. GPUs have no CPU-style SIMD registers holding several samples side by side, so aether’s device “vectorization” instead gives one thread several register-resident scalars, one per lane — see Multi-device residency and the CPU-side analogue in CPU SIMD backend for the two different mechanisms this same idea takes.
Stream — an ordered queue of GPU work (copies, kernel launches) that lets independent streams’ work overlap. aether exposes a stream as
aether::Stream(a thin alias over the CUDA stream handle).Pinned memory — host memory that the CUDA driver has locked in physical RAM (never swapped, never moved), which is what makes a host↔device copy fast and asynchronous-capable; aether calls a chunk of it a
kDLCUDAHostdevice kind.
The macro surface#
Every qualifier aether uses is spelled AETHER_<NAME>(), declared in
aether/macros.h:
AETHER_DEVICEHOST()— context-aware: on the device compile pass it is a device qualifier, on the host pass a host qualifier. This is what nearly every aether function uses, including the plain example below.AETHER_DEVICE()/AETHER_HOST()/AETHER_DEVICEANDHOST()— the narrower, non-context-aware qualifiers for a function that is ONLY ever device code, ONLY ever host code, or explicitly both, respectively.AETHER_KERNEL()— a kernel entry point (__global__). You write one when you need to start new parallel work from the host; a kernel cannot itself be called the way an ordinary function is, and it returnsvoid— results come back through memory aView(see Views and Items) points at.AETHER_GRID_CONSTANT()— marks a kernel parameter as read-only and shared across the whole launch, letting the compiler cache it more cheaply than a per-thread copy.AETHER_SHARED()— declares a block-local scratch variable, visible to every thread in the same block only.AETHER_NOINLINE()/AETHER_FORCEINLINE()— inlining hints, same idea as on the host.
In a build with no GPU toolchain at all (AETHER_CPP_MODE), every one of
these collapses to a plain host function — the SAME source compiles and
runs as an ordinary CPU library, which is how the runnable example below
executes even though it uses a device-aware qualifier.
Two headers, two audiences#
aether/device.h is the device-safe SLICE of the library — expression
templates, View/Item arithmetic — with every host-only piece
(allocation, error throwing, DLPack interop) left out, so it can be parsed
by tools that only understand device code. aether/aether.h is
everything: it includes aether/device.h and then adds the host-only
machinery on top. Day to day you just #include <aether/aether.h> and
get both; the split mostly matters to codegen tooling that consumes only
the device-safe half.
Two more device-side pieces worth knowing the names of: aether::cuda::
launchConfig (aether/backend/cuda/Launch.h) turns a sample count into
a grid/block shape for a launch, and aether::DeviceBundle
(aether/backend/cuda/DeviceBundle.h) is the register-resident,
multi-lane value a kernel thread uses to process several samples at once —
the device-side counterpart to the CPU packet type in CPU SIMD backend.
A runnable example#
The smoke test below defines a function with AETHER_DEVICEHOST() and
calls it directly from host code — no kernel launch needed, because the
function itself is not a kernel, just an ordinary function that also
happens to be legal to call from device code:
AETHER_DEVICEHOST() int answer()
{
return 42;
}