Tutorials#

These tutorials build up EAGLE’s core patterns as small, complete C++ programs – the C++ track alongside the Python tutorials that start at Thirty seconds: eagle.simulate. Every code block below is literalincluded from a real source file under docs/examples/ that is compiled and run by docs/examples/run_examples.sh (the exampletest gate) in both CUDA and pure-C++ / OpenMP mode. The code you read is exactly the code that produces the output shown – there are no illustrative-but-untested snippets here.

C1 is CUDA-only – graph capture has no pure-C++ analogue. C2 and C3 are dual-mode: the same source file compiles both against CUDA and against plain C++/OpenMP, with no changes (see Foundations — the words this documentation uses for the full picture). They also work with SoA (structure-of-arrays) data: instead of one array of per-sample structs, each field (position, velocity, …) lives in its own contiguous array, which is what lets both the GPU and the OpenMP/SIMD backend read and write it efficiently.

Note

Build any example directly, with eagle and aether installed into your environment ($CONDA_PREFIX):

# CUDA mode
nvcc -std=c++20 -arch=sm_61 -I$CONDA_PREFIX/include \
     -DEAGLE_BLOCKSIZE=256 -DCUDA_API_PER_THREAD_DEFAULT_STREAM=1 \
     -Xcompiler -fPIE -Xcompiler -fopenmp \
     docs/examples/01_capture_replay.cu -o 01_capture_replay

# pure-C++ mode (dual-mode examples)
g++ -std=c++23 -fopenmp -O2 -I$CONDA_PREFIX/include \
    -DEAGLE_CPU_ONLY=1 -DEAGLE_BLOCKSIZE=256 \
    docs/examples/02_host_dispatch.cpp -o 02_host_dispatch

C1 — Capture a launch once, replay it many times#

Record a stream of kernel launches once, then replay the recorded graph as many times as you like – with no per-launch API overhead.

Time: ~3 min · Runs on: CUDA GPU only (graph capture has no pure-C++ analogue) · You need: a CUDA toolchain, the build command above

The kernel we want to run:

// A trivial elementwise kernel — the "work" we want to capture and replay.
__global__ void addOne(int* buf, int n)
{
    const int tid = threadIdx.x + blockIdx.x * blockDim.x;
    if (tid < n)
        buf[tid] += 1;
}

Record the launch into a graph. Everything issued on the stream between begin() and end() is captured rather than executed eagerly, then handed to the Graph as a node:

    // A stream to record on, and a Graph that will own the captured work.
    eagle::cuda::Stream stream;
    eagle::cuda::Graph graph;
    graph.stream(stream.cuda());

    // Everything launched on the stream between begin() and end() is recorded
    // into a cudaGraph_t instead of executing eagerly.
    eagle::cuda::StreamCapturer capturer(stream.cuda());
    capturer.begin();
    addOne<<<1, N, 0, stream.cuda()>>>(dBuf, N);
    graph.addNode(capturer.end());

Instantiate the graph once and replay it – each launch() is a single cudaGraphLaunch:

    // Instantiate the graph once, then replay it REPLAYS times. Each launch()
    // is a single cudaGraphLaunch — no kernel-config marshalling per call.
    eagle::cuda::Launcher launcher = graph.launcher();
    for (int r = 0; r < REPLAYS; ++r)
        launcher.launch();
    launcher.synchronize();

What it printed#

01_capture_replay: 10 replays -> buf[0]=10 (expected 10) : OK

What just happened#

  • StreamCapturer recorded addOne’s launch instead of running it: the cudaGraph_t it produced between begin()/end() became one node on graph.

  • graph.launcher() instantiated that graph exactly once into a Launcher; every later launch() replays the same instantiated graph as a single driver call, not a fresh launch of addOne.

  • 10 replays of a kernel that adds 1 left every element at 10 – the proof that a replay really re-issues the captured work and not a no-op.

Try this#

Raise REPLAYS to 1000 and rebuild: the loop issuing launcher.launch() still costs one driver call per iteration, but nothing about the capture above changes – the graph itself does not grow.

Next#

C2, below – the same replay idea, but for work that also has a pure-C++ / OpenMP side with no CUDA at all.

C2 — Run one loop on the CPU or the GPU, unchanged#

The same source file compiles against CUDA and against plain C++/OpenMP – one dispatch call, two backends, no #ifdef in the body.

Time: ~3 min · Runs on: CUDA GPU or CPU (OpenMP) – pick the build command · You need: C1 (the capture/replay vocabulary), a C++23 compiler

eagle::cpu::Host::launch runs one task per index through EAGLE’s OpenMP + SIMD host launcher: consecutive indices are batched into aether SIMD packets and spread across threads. The same source compiles and runs in CUDA mode (02_host_dispatch.cu) and pure-C++ mode (02_host_dispatch.cpp) – that is EAGLE’s dual-mode promise.

    // One host task per index i; i.global() is this task's flat index. We write
    // the sequence of odd numbers 1, 3, 5, ... whose prefix sums are N*N.
    eagle::cpu::Host::launch(N, [&](const auto& i) {
        const eagle::idx_t k = i.global();
        x[k] = 2.0 * static_cast<double>(k) + 1.0;
    });

What it printed#

Both builds print the same line (the prefix sums of the first N odd numbers are N2):

02_host_dispatch: N=1024 sum=1048576 (expected 1048576) : OK

What just happened#

  • The loop body is one line, written once; Host::launch is what changes between the CUDA build (a device kernel over the same body) and the pure-C++ build (an OpenMP + SIMD host loop).

  • N=1024 lanes each wrote their own prefix sum; the host and device builds agree on every element, not just the final total.

  • Nothing in 02_host_dispatch.cu/.cpp names CUDA or OpenMP directly – the backend is a build flag (EAGLE_CPU_ONLY), not a code branch.

Try this#

Change N from 1024 to 1 and rebuild both modes: the printed sum becomes 1 either way – a batch of one is not a special case.

Next#

C3, below – the same dual-mode idea applied to a reduction instead of a per-element write.

C3 — Reduce an array, the same way on either backend#

Sum an aether SoA array on the GPU with a multi-level CUDA reduction, or on the CPU with a cache-padded OpenMP one – same call, same answer.

Time: ~2 min · Runs on: CUDA GPU or CPU (OpenMP) · You need: C2

In CUDA mode, eagle::cuda::Reduction runs a multi-level GPU reduction over the uploaded array:

    // Blocking multi-level reduction on the GPU: returns the scalar sum.
    const T total = eagle::cuda::Reduction::reduceBlocking<T, Sum>(arr, T(0), stream);

In pure-C++ mode, eagle::cpu::Reduction runs the cache-padded OpenMP reduction over the host-resident array – same operator, same result:

    // Cache-padded OpenMP reduction over the host-resident array.
    const T total = eagle::cpu::Reduction<T, Sum>::reduce(arr.hostView().as_const(), T(0));

What it printed#

03_reduction (CUDA): sum(1..1000)=500500 (expected 500500) : OK

What just happened#

  • Reduction took the same aether array and the same combine operator on both backends; only the class name (eagle::cuda::Reduction vs. eagle::cpu::Reduction) differs between the two source files.

  • The GPU build reduces in multiple levels (block, then grid); the CPU build pads each OpenMP thread’s partial sum to its own cache line to avoid false sharing – two different implementations reaching the same sum(1..1000).

  • Neither build re-sums on the host afterwards: the reduction itself is the answer, not a seed for a second pass.

Try this#

Change the summed range from 1..1000 to 1..1 and rebuild: the expected total becomes 1 on both backends, the smallest possible reduction.

Next#

C4 – embedding a compiled plugin in your own C++ program, in Examples – the other side of EAGLE: a binary ABI a C++ host can load with no aether headers, no Python, and no code generator at build or run time. deeper: CUDA graphs for the concepts behind capture/replay, and the Module: reduce / Module: cpu module guides for the reduction and host APIs.

Python: composing several kernels into one step#

The Python tutorials under tutorials/ (starting at Thirty seconds: eagle.simulate) build the same core patterns from kernels loaded and compiled through hawk and eagle’s Python face, one kernel at a time. A multi-kernel step, declaratively is the one to read for a step built from several kernels – a propagation, a diagnostic and a termination event – composed explicitly through eagle.plan.plan, GraphPipeline and ActiveSet, the same declarative style as the single-kernel run_until_done() convenience, one level below it.