Note
Running this on Colab or a fresh environment? Install the family from PyPI,
following eagle’s
install guide.
The NVIDIA packages come only through the [cuda12] / [cuda13] extras:
# !pip install "raptor-hawk[cuda12]" "raptor-eagle[cuda12]"
How to load a kernel someone else compiled#
Needs a CUDA GPU. Every kernel in the tutorials was written as a
hawk function and built by eagle.deploy. This how-to works the other way:
a kernel producer ships a compiled artifact — a .ptx file plus a small
JSON sidecar, optionally several of them behind one manifest.json — and
you load and launch it with no nvcc at runtime and no hawk source in
sight.
One deployed kernel, loaded directly#
gravity.ptx + gravity.json is one such artifact: a hand-written CUDA
“gravity” force, a = -mu r / |r|^3. LoadedVector reads everything it
needs from the sidecar at load time: which inputs are per-sample vectors,
which are scalar parameters, and the exact argument order the compiled
kernel expects.
import pathlib
import cupy as cp
import numpy as np
from eagle.loaded import LoadedVector
FIXTURES = pathlib.Path("../_fixtures").resolve()
force = LoadedVector(FIXTURES / "gravity.ptx")
arg_spec names every declared argument’s role, in the order the compiled
kernel expects them: vec_in a per-sample vector input, uniform a
broadcast scalar parameter, out the output plane the launch writes, and
(where declared) terminated the per-sample done mask. This is raptor’s
manifest-role vocabulary — see raptor’s
manifest-schema tutorial
for the full set.
{
"kernel": force.kernel_name,
"vector_inputs": force.vector_inputs,
"param_names": force.param_names,
"arg_spec": force.arg_spec,
"abi_tag": force.abi_tag,
}
{'kernel': 'raptor_kernel',
'vector_inputs': ('position',),
'param_names': ('mu',),
'arg_spec': [('out', 'out'),
('vec_in', 'position'),
('terminated', 'terminated'),
('uniform', 'mu')],
'abi_tag': 'aether-abi/1'}
numpy in, cupy in — same answer#
The idiomatic __call__ dispatches on the framework of its inputs
(eagle.interop.origin_adapter): a numpy input blocks and returns numpy; a
cupy input stays on the device’s current stream and returns cupy. Both reach
the identical compiled kernel.
MU = 3.986004418e5
rng = np.random.default_rng(1)
position_np = rng.uniform(7.0e3, 4.2e4, size=(3, 64))
accel_np = force(position=position_np, mu=MU)
accel_cp = force(position=cp.asarray(position_np), mu=MU)
type(accel_np), type(accel_cp)
(numpy.ndarray, cupy.ndarray)
np.testing.assert_array_equal(accel_np, cp.asnumpy(accel_cp))
print("numpy and cupy launches agree bit-for-bit")
numpy and cupy launches agree bit-for-bit
The capturable door#
__call__ allocates its own output. The lower-level launch method takes a
pre-allocated output buffer and a terminated mask instead — no
allocation happens inside it, which is what makes it legal to record inside
a CUDA graph capture (next tutorial). The idiomatic __call__ used above
never needed a terminated= argument even though arg_spec declares the
role: __call__ auto-supplies an all-False mask when the caller doesn’t
pass one. launch() makes no such substitution — pass terminated
explicitly, as below.
out = cp.zeros((3, 64), dtype=cp.float64)
terminated = cp.zeros(64, dtype=cp.bool_)
force.launch(out=out, position=cp.asarray(position_np), mu=MU, terminated=terminated)
cp.cuda.runtime.deviceSynchronize()
np.testing.assert_array_equal(cp.asnumpy(out), accel_np)
print("launch() into a caller-owned buffer matches __call__'s own answer")
launch() into a caller-owned buffer matches __call__'s own answer
Several kernels behind one manifest#
A producer that ships more than one kernel together names them in a
manifest.json: one artifact + sidecar pair per plugin. load_manifest is
the Python analogue of the C++ PluginRegistry::from_manifest, and resolves
it into a by-name map of launchables. Here the manifest wraps the same
gravity artifact, on its own.
import json
import shutil
import tempfile
workdir = pathlib.Path(tempfile.mkdtemp())
shutil.copy(FIXTURES / "gravity.ptx", workdir / "gravity.ptx")
shutil.copy(FIXTURES / "gravity.json", workdir / "gravity.json")
manifest = {
"schema_version": 1,
"pattern": "vector",
"aether_abi": "aether-abi/1",
"plugins": [
{"id": "gravity", "order": 0, "enabled": True, "artifact": "gravity.ptx",
"sidecar": "gravity.json", "format": "ptx"},
],
}
(workdir / "manifest.json").write_text(json.dumps(manifest, indent=2))
from eagle.registry import load_manifest
reg = load_manifest(workdir / "manifest.json")
reg.names()
['gravity']
reg is a plain name -> launchable map. Calling a launchable
idiomatically (reg["gravity"](...)) allocates, launches, and returns
the result in your tensor framework — numpy in, numpy out. RAPTOR planes
are component-major: position below is shaped (3, N) — 3
components, then samples — not the (N, 3) a lot of NumPy/PyTorch code
defaults to. A sample-major array is also accepted: a contiguous transpose
binds without a copy, and any other sample-major array is copied once with
an eagle.LayoutWarning naming the fix. See eagle’s
Interoperability contract (“Layouts” section) for the full rule.
position = rng.uniform(7.0e3, 4.2e4, size=(3, 16)) # (3, N) state vectors
accel = reg["gravity"](position=position, mu=MU)
accel.shape, accel.dtype
((3, 16), dtype('float64'))
r = np.linalg.norm(position, axis=0)
expected = -MU * position / r**3
float(np.max(np.abs(accel - expected) / np.abs(expected)))
4.836451403850155e-16
Next#
What a graph is — the same kind of loaded kernel, captured into a replayable CUDA graph.
Interoperability — numpy, cupy and torch inputs/outputs side by side.
Thirty seconds: eagle.simulate — writing your own kernel instead of loading one, and running it to a million samples with one
eagle.simulatecall.