eagle

eagle#

One loop, recorded once, replayed until every sample is done. eagle runs a kernel over any number of samples — one or a million — without a launch per sample: write the step once, and eagle captures it into a single replayable launch, a CUDA graph on the GPU or an OpenMP + SIMD host loop on the CPU, picked automatically from the data you pass in. A PyTorch bridge lets that same kernel sit inside loss.backward() like any other layer. (The name stands for Extensible Adaptive Graph Launch Engine, if you’re curious.) A header-only C++ core; a Python package over the same engine.

Thirty seconds: eagle.simulate#

Write one sample’s step as a hawk kernel, hand it to eagle.simulate(), and it runs every sample until it is done — on the CPU or the GPU, no loop of your own:

@hawk.kernel
def oscillator(omega: Scalar, t_end: Param, dt: Param, terminated: Terminated,
               x: Mutable[Scalar], v: Mutable[Scalar], t: Mutable[Scalar]):
    x0, v0 = x, v
    x = x0 + dt * v0
    v = v0 - dt * omega * omega * x0
    t += dt
    terminated = t >= t_end

result = eagle.simulate(oscillator, omega=omegas, t_end=1.0, dt=1e-3,
                        x=x0, v=v0, t=t0, max_steps=10_000)

Pass the kernel’s own arguments as you would call it: an array gives each sample its own value, a plain number is shared by every sample.

The full, runnable version is the Python quickstart — start there.

Which call do I use?#

One kernel (or a step’s worth), run until every sample finishes

eagle.simulate()

A plan already paced by its own finish kernel, one call

eagle.until_done()

Build a kernel into a plan that runs where its data lives

eagle.deploy() / eagle.plan

Resolve a launchable by name, then call/stream/graph it

eagle.KernelRegistry

This is eagle’s own module docstring — the same four lines help(eagle) or eagle.__dir__() leads with. The full reference, every other call eagle owns, is Python API reference.

C++ users start here

eagle is a header-only C++23 library first: capture a CUDA graph, dispatch dual-mode CPU/GPU work, and more, with no Python runtime at all. Foundations defines the vocabulary (Stream, Graph, Launcher); the C++ quickstart and tutorials build up from there — record a launch once, assemble it into a CUDA graph, and replay it, the same pattern eagle.simulate builds on in Python.

Foundations — the words this documentation uses

Where eagle sits#

eagle is the execution layer of the RAPTOR family — it launches what aether lays out in memory, and what raptor’s manifest schema describes:

aether (array layout)  --->  eagle (THIS: launch, graphs, host dispatch)  <---  raptor (manifest / interop contracts)
                                       ^
                                       |
                                     hawk (kernel compiler, consumes eagle to run what it compiles)
Install

C++ (header-only, CMake) and Python (pip install "raptor-eagle[cuda12]", or [cuda13]).

Installation
Python quickstart

eagle.simulate, end to end, in under a minute. Start here.

Thirty seconds: eagle.simulate
Tutorials

Build up eagle’s core patterns step by step: graphs, launch cadence, compaction, multi-kernel steps, training through a kernel — in Python, with a C++ track alongside.

Tutorials
API reference

Complete C++ (Doxygen/breathe) and Python (autosummary) API reference.

API reference

The RAPTOR family#

Repository

Role

raptor

The shared contracts: manifest schema, engine/launcher/provider contracts, the interop certification matrix. No hard dependencies.

aether

Zero-cost array layout and GPU memory management (SoA containers).

eagle (this repository)

GPU execution layer: kernel launch, graph capture, plugin registry, host/device dual mode.

hawk

Symbolic/numeric DSL and kernel compiler: hawk compiles kernels eagle runs; neither imports the other.