CUDA graphs#
Note
This page is the C++ class/method reference for the graph engine, with
no kernel code. Working from Python? What a graph is
is the same record-once/replay-many idea, worked end to end with
eagle.until_done() and GraphPipeline.
EAGLE is built around CUDA graphs: record a stream of kernel launches once, assemble them into a dependency graph, and replay the instantiated graph to amortise per-launch overhead across the inner loop of a batched algorithm.
Execution model#
There are three cooperating pieces:
StreamCapturerWraps
cudaStreamBeginCapture/cudaStreamEndCapture. Callbegin()before launching kernels on the associated stream andend()afterwards to obtain acudaGraph_t.GraphCollects captured graphs and host callbacks into a single dependency graph (
addNode,addKernelNode,addHostNode) and instantiates it into aLaunchervialauncher().LauncherThe instantiated, replayable graph.
launch()submits it,synchronize()waits, andsetLogicalSize()changes the logical size — how many samples the next replay actually processes — without rebuilding the graph. Concretely: for each kernel node it re-derives the grid/block launch dimensions for the new sample count and patches them directly into the already-instantiated graph (no re-capture, no newcudaGraph_t) — so the nextlaunch()runs exactly as many threads as the new size needs, no more. A typical reason to shrink it: a batched loop where some samples finish (terminate) before others and no longer need processing, so later replays only need to cover the samples still active.
Direct graph usage#
#include <eagle/eagle.h>
eagle::cuda::Stream stream;
eagle::cuda::Graph graph;
graph.stream(stream.cuda());
// Capture a kernel into the graph
eagle::cuda::StreamCapturer cap(stream.cuda());
cap.begin();
myKernel<<<grid, block, 0, stream.cuda()>>>(args);
graph.addNode(cap.end());
// Add a host callback (runs on the CPU as a graph node)
graph.addHostNode([&result]() { /* post-process on CPU */ });
// Instantiate and launch
eagle::cuda::Launcher launcher = graph.launcher();
launcher.launch();
launcher.synchronize();
Stream management#
eagle::cuda::Stream wraps a cudaStream_t with RAII lifetime management; call
.cuda() to obtain the raw handle:
#include <eagle/eagle.h>
eagle::cuda::Stream myStream; // creates a new CUDA stream
graph.stream(myStream.cuda()); // hand the raw handle to the graph
myStream.synchronize();
eagle::cuda::Event similarly wraps cudaEvent_t for cross-stream
synchronisation — make one stream wait for work already queued on a
different stream, without a full host-side synchronize() on either:
eagle::cuda::Stream streamA, streamB;
eagle::cuda::Event doneOnA;
// ... launch work on streamA ...
doneOnA.record(streamA); // mark this point in streamA's queue
// streamB won't start its next queued work until streamA reaches
// the recorded point — no host-side wait, purely device-side ordering
streamB.waitFor(doneOnA);
// ... launch work on streamB that depends on streamA's result ...
Note
eagle defines CUDA_API_PER_THREAD_DEFAULT_STREAM=1 at compile time, so
stream 0 maps to each thread’s private default stream rather than the
legacy NULL stream.
Conditional graph nodes: skippable regions#
Needs CUDA ≥ 12.3. Every node in a plain graph always runs on replay.
A conditional node lets one region decide, on the device and with no
host round trip, whether it runs at all this replay — useful for a region
that is only sometimes needed, like an event’s follow-up work.
eagle::conditional::ConditionalGroup is the native-node wrapper for it:
populate its own body graph, then add the whole group into a parent
Graph exactly like any other native node.
#include <eagle/conditional.h>
// *livePtr is a device word the caller updates between replays.
eagle::CountGuard guard{ livePtr, /*baseline=*/nullptr }; // runs iff *livePtr != 0
eagle::conditional::ConditionalGroup group(guard);
cudaKernelNodeParams kp = {};
void* args[] = { (void*)&someBuffer };
kp.func = (void*)onlyWhenLiveKernel;
kp.gridDim = { 1, 1, 1 };
kp.blockDim = { 1, 1, 1 };
kp.kernelParams = args;
group.body().addKernelNode(kp, {});
eagle::cuda::Graph g;
g.addNative(std::move(group), {});
g.finalizeNatives();
eagle::cuda::Launcher launcher = g.launcher();
launcher.launch(); // the body fires iff *livePtr != 0, read fresh every replay
CountGuard compares a device word (count) against a baseline
(0 when baseline is nullptr); change only that word between
replays and the body fires or is skipped accordingly, with no rebuild and
no re-capture. eagle/conditional.h is the umbrella header: it also pulls
in the loop-building sibling (eagle::cuda::CaptureConditional, for
weaving a conditional directly into stream capture) and documents the
nesting and single-launcher limits inline.
See also
Module: cuda — full cuda module description.
API reference — cuda — Graph, Launcher, StreamCapturer API.