eagle#
One loop, recorded once, replayed until every sample is done. eagle
runs a kernel over any number of samples — one or a million — without a
launch per sample: write the step once, and eagle captures it into a
single replayable launch, a CUDA graph on the GPU or an OpenMP + SIMD host
loop on the CPU, picked automatically from the data you pass in. A PyTorch
bridge lets that same kernel sit inside loss.backward() like any other
layer. (The name stands for Extensible Adaptive Graph Launch Engine, if
you’re curious.) A header-only C++ core; a Python package over the same
engine.
Thirty seconds: eagle.simulate#
Write one sample’s step as a hawk
kernel, hand it to eagle.simulate(), and it runs every sample until
it is done — on the CPU or the GPU, no loop of your own:
@hawk.kernel
def oscillator(omega: Scalar, t_end: Param, dt: Param, terminated: Terminated,
x: Mutable[Scalar], v: Mutable[Scalar], t: Mutable[Scalar]):
x0, v0 = x, v
x = x0 + dt * v0
v = v0 - dt * omega * omega * x0
t += dt
terminated = t >= t_end
result = eagle.simulate(oscillator, omega=omegas, t_end=1.0, dt=1e-3,
x=x0, v=v0, t=t0, max_steps=10_000)
Pass the kernel’s own arguments as you would call it: an array gives each sample its own value, a plain number is shared by every sample.
The full, runnable version is the Python quickstart — start there.
Which call do I use?#
One kernel (or a step’s worth), run until every sample finishes |
|
A plan already paced by its own finish kernel, one call |
|
Build a kernel into a plan that runs where its data lives |
|
Resolve a launchable by name, then call/stream/graph it |
This is eagle’s own module docstring — the same four lines help(eagle)
or eagle.__dir__() leads with. The full reference, every other call
eagle owns, is Python API reference.
eagle is a header-only C++23 library first: capture a CUDA graph,
dispatch dual-mode CPU/GPU work, and more, with no Python runtime at
all. Foundations defines the
vocabulary (Stream, Graph, Launcher); the
C++ quickstart and
tutorials build up from there —
record a launch once, assemble it into a CUDA graph, and replay it,
the same pattern eagle.simulate builds on in Python.
Where eagle sits#
eagle is the execution layer of the RAPTOR family — it launches what aether lays out in memory, and what raptor’s manifest schema describes:
aether (array layout) ---> eagle (THIS: launch, graphs, host dispatch) <--- raptor (manifest / interop contracts)
^
|
hawk (kernel compiler, consumes eagle to run what it compiles)
C++ (header-only, CMake) and Python (pip install "raptor-eagle[cuda12]", or [cuda13]).
eagle.simulate, end to end, in under a minute. Start here.
Build up eagle’s core patterns step by step: graphs, launch cadence, compaction, multi-kernel steps, training through a kernel — in Python, with a C++ track alongside.
Complete C++ (Doxygen/breathe) and Python (autosummary) API reference.
The RAPTOR family#
Repository |
Role |
|---|---|
The shared contracts: manifest schema, engine/launcher/provider contracts, the interop certification matrix. No hard dependencies. |
|
Zero-cost array layout and GPU memory management (SoA containers). |
|
eagle (this repository) |
GPU execution layer: kernel launch, graph capture, plugin registry, host/device dual mode. |
Symbolic/numeric DSL and kernel compiler: hawk compiles kernels eagle runs; neither imports the other. |