Note

Running this on Colab or a fresh environment? Install the family from PyPI, following eagle’s install guide. The NVIDIA packages come only through the [cuda12] / [cuda13] extras:

# !pip install "raptor-hawk[cuda12]" "raptor-eagle[cuda12]"

How to load a kernel someone else compiled#

Needs a CUDA GPU. Every kernel in the tutorials was written as a hawk function and built by eagle.deploy. This how-to works the other way: a kernel producer ships a compiled artifact — a .ptx file plus a small JSON sidecar, optionally several of them behind one manifest.json — and you load and launch it with no nvcc at runtime and no hawk source in sight.

One deployed kernel, loaded directly#

gravity.ptx + gravity.json is one such artifact: a hand-written CUDA “gravity” force, a = -mu r / |r|^3. LoadedVector reads everything it needs from the sidecar at load time: which inputs are per-sample vectors, which are scalar parameters, and the exact argument order the compiled kernel expects.

import pathlib

import cupy as cp
import numpy as np

from eagle.loaded import LoadedVector

FIXTURES = pathlib.Path("../_fixtures").resolve()
force = LoadedVector(FIXTURES / "gravity.ptx")

arg_spec names every declared argument’s role, in the order the compiled kernel expects them: vec_in a per-sample vector input, uniform a broadcast scalar parameter, out the output plane the launch writes, and (where declared) terminated the per-sample done mask. This is raptor’s manifest-role vocabulary — see raptor’s manifest-schema tutorial for the full set.

{
    "kernel": force.kernel_name,
    "vector_inputs": force.vector_inputs,
    "param_names": force.param_names,
    "arg_spec": force.arg_spec,
    "abi_tag": force.abi_tag,
}
{'kernel': 'raptor_kernel',
 'vector_inputs': ('position',),
 'param_names': ('mu',),
 'arg_spec': [('out', 'out'),
  ('vec_in', 'position'),
  ('terminated', 'terminated'),
  ('uniform', 'mu')],
 'abi_tag': 'aether-abi/1'}

numpy in, cupy in — same answer#

The idiomatic __call__ dispatches on the framework of its inputs (eagle.interop.origin_adapter): a numpy input blocks and returns numpy; a cupy input stays on the device’s current stream and returns cupy. Both reach the identical compiled kernel.

MU = 3.986004418e5
rng = np.random.default_rng(1)
position_np = rng.uniform(7.0e3, 4.2e4, size=(3, 64))

accel_np = force(position=position_np, mu=MU)
accel_cp = force(position=cp.asarray(position_np), mu=MU)

type(accel_np), type(accel_cp)
(numpy.ndarray, cupy.ndarray)
np.testing.assert_array_equal(accel_np, cp.asnumpy(accel_cp))
print("numpy and cupy launches agree bit-for-bit")
numpy and cupy launches agree bit-for-bit

The capturable door#

__call__ allocates its own output. The lower-level launch method takes a pre-allocated output buffer and a terminated mask instead — no allocation happens inside it, which is what makes it legal to record inside a CUDA graph capture (next tutorial). The idiomatic __call__ used above never needed a terminated= argument even though arg_spec declares the role: __call__ auto-supplies an all-False mask when the caller doesn’t pass one. launch() makes no such substitution — pass terminated explicitly, as below.

out = cp.zeros((3, 64), dtype=cp.float64)
terminated = cp.zeros(64, dtype=cp.bool_)
force.launch(out=out, position=cp.asarray(position_np), mu=MU, terminated=terminated)
cp.cuda.runtime.deviceSynchronize()
np.testing.assert_array_equal(cp.asnumpy(out), accel_np)
print("launch() into a caller-owned buffer matches __call__'s own answer")
launch() into a caller-owned buffer matches __call__'s own answer

Several kernels behind one manifest#

A producer that ships more than one kernel together names them in a manifest.json: one artifact + sidecar pair per plugin. load_manifest is the Python analogue of the C++ PluginRegistry::from_manifest, and resolves it into a by-name map of launchables. Here the manifest wraps the same gravity artifact, on its own.

import json
import shutil
import tempfile

workdir = pathlib.Path(tempfile.mkdtemp())
shutil.copy(FIXTURES / "gravity.ptx", workdir / "gravity.ptx")
shutil.copy(FIXTURES / "gravity.json", workdir / "gravity.json")
manifest = {
    "schema_version": 1,
    "pattern": "vector",
    "aether_abi": "aether-abi/1",
    "plugins": [
        {"id": "gravity", "order": 0, "enabled": True, "artifact": "gravity.ptx",
         "sidecar": "gravity.json", "format": "ptx"},
    ],
}
(workdir / "manifest.json").write_text(json.dumps(manifest, indent=2))

from eagle.registry import load_manifest

reg = load_manifest(workdir / "manifest.json")
reg.names()
['gravity']

reg is a plain name -> launchable map. Calling a launchable idiomatically (reg["gravity"](...)) allocates, launches, and returns the result in your tensor framework — numpy in, numpy out. RAPTOR planes are component-major: position below is shaped (3, N) — 3 components, then samples — not the (N, 3) a lot of NumPy/PyTorch code defaults to. A sample-major array is also accepted: a contiguous transpose binds without a copy, and any other sample-major array is copied once with an eagle.LayoutWarning naming the fix. See eagle’s Interoperability contract (“Layouts” section) for the full rule.

position = rng.uniform(7.0e3, 4.2e4, size=(3, 16))  # (3, N) state vectors

accel = reg["gravity"](position=position, mu=MU)
accel.shape, accel.dtype
((3, 16), dtype('float64'))
r = np.linalg.norm(position, axis=0)
expected = -MU * position / r**3
float(np.max(np.abs(accel - expected) / np.abs(expected)))
4.836451403850155e-16

Next#

  • What a graph is — the same kind of loaded kernel, captured into a replayable CUDA graph.

  • Interoperability — numpy, cupy and torch inputs/outputs side by side.

  • Thirty seconds: eagle.simulate — writing your own kernel instead of loading one, and running it to a million samples with one eagle.simulate call.