Tutorials#
These tutorials build up EAGLE’s core patterns as small, complete C++ programs
– the C++ track alongside the Python tutorials that start at
Thirty seconds: eagle.simulate. Every code block below is literalincluded from a
real source file under docs/examples/ that is compiled and run by
docs/examples/run_examples.sh (the exampletest gate) in both CUDA and
pure-C++ / OpenMP mode. The code you read is exactly the code that produces the
output shown – there are no illustrative-but-untested snippets here.
C1 is CUDA-only – graph capture has no pure-C++ analogue. C2 and C3 are dual-mode: the same source file compiles both against CUDA and against plain C++/OpenMP, with no changes (see Foundations — the words this documentation uses for the full picture). They also work with SoA (structure-of-arrays) data: instead of one array of per-sample structs, each field (position, velocity, …) lives in its own contiguous array, which is what lets both the GPU and the OpenMP/SIMD backend read and write it efficiently.
Note
Build any example directly, with eagle and aether installed into your
environment ($CONDA_PREFIX):
# CUDA mode
nvcc -std=c++20 -arch=sm_61 -I$CONDA_PREFIX/include \
-DEAGLE_BLOCKSIZE=256 -DCUDA_API_PER_THREAD_DEFAULT_STREAM=1 \
-Xcompiler -fPIE -Xcompiler -fopenmp \
docs/examples/01_capture_replay.cu -o 01_capture_replay
# pure-C++ mode (dual-mode examples)
g++ -std=c++23 -fopenmp -O2 -I$CONDA_PREFIX/include \
-DEAGLE_CPU_ONLY=1 -DEAGLE_BLOCKSIZE=256 \
docs/examples/02_host_dispatch.cpp -o 02_host_dispatch
C1 — Capture a launch once, replay it many times#
Record a stream of kernel launches once, then replay the recorded graph as many times as you like – with no per-launch API overhead.
Time: ~3 min · Runs on: CUDA GPU only (graph capture has no pure-C++ analogue) · You need: a CUDA toolchain, the build command above
The kernel we want to run:
// A trivial elementwise kernel — the "work" we want to capture and replay.
__global__ void addOne(int* buf, int n)
{
const int tid = threadIdx.x + blockIdx.x * blockDim.x;
if (tid < n)
buf[tid] += 1;
}
Record the launch into a graph. Everything issued on the stream between
begin() and end() is captured rather than executed eagerly, then handed
to the Graph as a node:
// A stream to record on, and a Graph that will own the captured work.
eagle::cuda::Stream stream;
eagle::cuda::Graph graph;
graph.stream(stream.cuda());
// Everything launched on the stream between begin() and end() is recorded
// into a cudaGraph_t instead of executing eagerly.
eagle::cuda::StreamCapturer capturer(stream.cuda());
capturer.begin();
addOne<<<1, N, 0, stream.cuda()>>>(dBuf, N);
graph.addNode(capturer.end());
Instantiate the graph once and replay it – each launch() is a single
cudaGraphLaunch:
// Instantiate the graph once, then replay it REPLAYS times. Each launch()
// is a single cudaGraphLaunch — no kernel-config marshalling per call.
eagle::cuda::Launcher launcher = graph.launcher();
for (int r = 0; r < REPLAYS; ++r)
launcher.launch();
launcher.synchronize();
What it printed#
01_capture_replay: 10 replays -> buf[0]=10 (expected 10) : OK
What just happened#
StreamCapturerrecordedaddOne’s launch instead of running it: thecudaGraph_tit produced betweenbegin()/end()became one node ongraph.graph.launcher()instantiated that graph exactly once into aLauncher; every laterlaunch()replays the same instantiated graph as a single driver call, not a fresh launch ofaddOne.10 replays of a kernel that adds 1 left every element at 10 – the proof that a replay really re-issues the captured work and not a no-op.
Try this#
Raise REPLAYS to 1000 and rebuild: the loop issuing launcher.launch()
still costs one driver call per iteration, but nothing about the capture
above changes – the graph itself does not grow.
Next#
C2, below – the same replay idea, but for work that also has a pure-C++ / OpenMP side with no CUDA at all.
C2 — Run one loop on the CPU or the GPU, unchanged#
The same source file compiles against CUDA and against plain C++/OpenMP – one dispatch call, two backends, no
#ifdefin the body.
Time: ~3 min · Runs on: CUDA GPU or CPU (OpenMP) – pick the build command · You need: C1 (the capture/replay vocabulary), a C++23 compiler
eagle::cpu::Host::launch runs one task per index through EAGLE’s OpenMP +
SIMD host launcher: consecutive indices are batched into aether SIMD packets and
spread across threads. The same source compiles and runs in CUDA mode
(02_host_dispatch.cu) and pure-C++ mode (02_host_dispatch.cpp) – that is
EAGLE’s dual-mode promise.
// One host task per index i; i.global() is this task's flat index. We write
// the sequence of odd numbers 1, 3, 5, ... whose prefix sums are N*N.
eagle::cpu::Host::launch(N, [&](const auto& i) {
const eagle::idx_t k = i.global();
x[k] = 2.0 * static_cast<double>(k) + 1.0;
});
What it printed#
Both builds print the same line (the prefix sums of the first N odd numbers are N2):
02_host_dispatch: N=1024 sum=1048576 (expected 1048576) : OK
What just happened#
The loop body is one line, written once;
Host::launchis what changes between the CUDA build (a device kernel over the same body) and the pure-C++ build (an OpenMP + SIMD host loop).N=1024 lanes each wrote their own prefix sum; the host and device builds agree on every element, not just the final total.
Nothing in
02_host_dispatch.cu/.cppnamesCUDAorOpenMPdirectly – the backend is a build flag (EAGLE_CPU_ONLY), not a code branch.
Try this#
Change N from 1024 to 1 and rebuild both modes: the printed sum becomes 1
either way – a batch of one is not a special case.
Next#
C3, below – the same dual-mode idea applied to a reduction instead of a per-element write.
C3 — Reduce an array, the same way on either backend#
Sum an aether SoA array on the GPU with a multi-level CUDA reduction, or on the CPU with a cache-padded OpenMP one – same call, same answer.
Time: ~2 min · Runs on: CUDA GPU or CPU (OpenMP) · You need: C2
In CUDA mode, eagle::cuda::Reduction runs a multi-level GPU reduction over
the uploaded array:
// Blocking multi-level reduction on the GPU: returns the scalar sum.
const T total = eagle::cuda::Reduction::reduceBlocking<T, Sum>(arr, T(0), stream);
In pure-C++ mode, eagle::cpu::Reduction runs the cache-padded OpenMP
reduction over the host-resident array – same operator, same result:
// Cache-padded OpenMP reduction over the host-resident array.
const T total = eagle::cpu::Reduction<T, Sum>::reduce(arr.hostView().as_const(), T(0));
What it printed#
03_reduction (CUDA): sum(1..1000)=500500 (expected 500500) : OK
What just happened#
Reductiontook the same aether array and the same combine operator on both backends; only the class name (eagle::cuda::Reductionvs.eagle::cpu::Reduction) differs between the two source files.The GPU build reduces in multiple levels (block, then grid); the CPU build pads each OpenMP thread’s partial sum to its own cache line to avoid false sharing – two different implementations reaching the same
sum(1..1000).Neither build re-sums on the host afterwards: the reduction itself is the answer, not a seed for a second pass.
Try this#
Change the summed range from 1..1000 to 1..1 and rebuild: the expected total becomes 1 on both backends, the smallest possible reduction.
Next#
C4 – embedding a compiled plugin in your own C++ program, in Examples – the other side of EAGLE: a binary ABI a C++ host can load with no aether headers, no Python, and no code generator at build or run time. deeper: CUDA graphs for the concepts behind capture/replay, and the Module: reduce / Module: cpu module guides for the reduction and host APIs.
Python: composing several kernels into one step#
The Python tutorials under tutorials/ (starting at Thirty seconds: eagle.simulate)
build the same core patterns from kernels loaded and compiled through hawk
and eagle’s Python face, one kernel at a time.
A multi-kernel step, declaratively is the one to read for a step
built from several kernels – a propagation, a diagnostic and a termination
event – composed explicitly through eagle.plan.plan,
GraphPipeline and ActiveSet, the same
declarative style as the single-kernel run_until_done()
convenience, one level below it.