The parallel-execution model#
Read this before Foundations — the words this documentation uses if terms like “kernel,” “thread,” or “device memory” aren’t already familiar. It explains the one idea everything else in these docs assumes: you write what happens to a single item, and the hardware runs that same logic across every item at once. Nothing here is EAGLE-specific — it’s the same idea whether the “hardware” running your logic is a GPU or a multi-core CPU, which is exactly why EAGLE’s API has one shape for both.
One body, run many times at once#
Ordinary sequential code processes items one at a time, in a loop:
for (int i = 0; i < N; i++) {
result[i] = a[i] + b[i]; // one item, then the next, then the next...
}
On a single CPU core, this genuinely happens one iteration after another.
But the body of that loop — result[i] = a[i] + b[i] — doesn’t care
what order the iterations happen in, or whether they happen at the same
time. That’s the property both GPUs and modern CPUs exploit: instead of
running the body N times in sequence, run it for many values of i
simultaneously, on separate pieces of hardware that all execute the exact
same instructions at once. A kernel (the term used throughout these
docs, and throughout GPU programming generally) is exactly that: the body
of a loop, written once, launched to run across every item at once instead
of one at a time.
What actually runs it, in parallel#
On a GPU, each iteration runs on its own thread — a GPU has thousands of small, simple compute units, and a single kernel launch puts one of them on each item (or a small group of items) so they all execute the same instructions in lockstep.
On a CPU, you have far fewer cores, but each core’s vector hardware (SIMD) can run the same instruction over several values at once. EAGLE calls one such group of items processed together a packet (of a fixed size
W, the CPU’s SIMD width), and each item’s slot within that packet a lane — a packet is the CPU’s version of a GPU thread group. See Foundations — the words this documentation uses for the full walkthrough of this vocabulary in EAGLE’scpu::Hostdispatcher.
Both are the same underlying trick — do the same operation to many items at
once instead of one at a time — implemented on different hardware. That’s
why EAGLE presents one calling shape (write the per-item body, call
launch) for both: the parallel-execution idea is identical; only the
hardware executing it differs, and EAGLE’s dual-mode API is what lets your
code not care which one it’s running on.
Host and device memory are separate#
A CPU (the host) and a GPU (the device) each have their own,
physically separate memory. Data your CPU code has doesn’t automatically
exist where the GPU can read it — it has to be explicitly copied over
(uploaded) before a GPU kernel can use it, and copied back
(downloaded) to read the result on the CPU side. This is why EAGLE’s
APIs have explicit upload() / download() steps (see
Module: filtering for a concrete example) — it’s not
EAGLE-specific bookkeeping, it’s a direct consequence of the GPU having its
own memory.
What you won’t find explained here#
Raw CUDA’s own kernel-launch syntax — the <<<grid, block, shmem,
stream>>> triple-angle-bracket notation you’ll see in
Quick start’s addOne<<<1, N, 0, stream.cuda()>>>(...) — along with
what a “block” and “grid” precisely are, and the full CUDA memory model, are
raw CUDA C++ concepts that EAGLE builds on top of but doesn’t reinvent.
EAGLE’s own API (cpu::Host::launch, GraphPipeline, the Reduction
/ Scanner helpers) is designed so you rarely need to write that syntax
yourself. If you do want the full picture — writing a raw CUDA kernel from
scratch, tuning block/grid sizes, the GPU memory hierarchy in depth —
NVIDIA’s own CUDA C++ Programming Guide is the
authoritative reference; nothing in these docs assumes you’ve read it, but
it’s the right place to go deeper.