Interoperability#

Needs a CUDA GPU for the cupy cells; the numpy and torch-CPU cells run anywhere. eagle never asks which array library you use. Every launchable accepts a numpy array, a cupy array, or a PyTorch tensor as input and returns a result in that same framework — the exchange is DLPack: DLPack is the contract. A tensor that already lives on the GPU is imported to cupy as a zero-copy device view; a host tensor is uploaded once.

Hide code cell source

import pathlib
import warnings

import cupy as cp
import numpy as np

# torch's own `import pynvml` below trips a FutureWarning that bakes the
# site-packages path into its message; silence just that one (category +
# message), not warnings generally.
warnings.filterwarnings(
    "ignore", category=FutureWarning, message=r"The pynvml package is deprecated.*"
)
import torch

from eagle.interop import detect
from eagle.loaded import LoadedVector

FIXTURES = pathlib.Path("userguide/_fixtures").resolve()
force = LoadedVector(FIXTURES / "gravity.ptx")

MU = 3.986004418e5
rng = np.random.default_rng(4)
position_np = rng.uniform(7.0e3, 4.2e4, size=(3, 16))

numpy — host arrays, blocking#

accel_np = force(position=position_np, mu=MU)
type(accel_np), detect(position_np).name
(numpy.ndarray, 'numpy')

cupy — device arrays, non-blocking on the current stream#

position_cp = cp.asarray(position_np)
accel_cp = force(position=position_cp, mu=MU)
cp.cuda.runtime.deviceSynchronize()
type(accel_cp), detect(position_cp).name
(cupy.ndarray, 'cupy')

torch — CPU tensor here (this build’s torch has no CUDA); a CUDA torch#

tensor takes the exact same force(position=position_torch, mu=MU) call and is imported zero-copy through __dlpack__, with no code difference — eagle dispatches on type(x).__module__ and the tensor’s own DLPack device tag, not on how the process was started.

position_torch = torch.from_numpy(position_np)
accel_torch = force(position=position_torch, mu=MU)
type(accel_torch), detect(position_torch).name
(torch.Tensor, 'torch')
np.testing.assert_allclose(accel_np, cp.asnumpy(accel_cp))
np.testing.assert_allclose(accel_np, accel_torch.numpy())
print("numpy, cupy and torch all launch the same kernel and agree")
numpy, cupy and torch all launch the same kernel and agree

NVIDIA Warp — the same call, zero-copy both ways#

NVIDIA Warp is a Python framework for writing GPU kernels directly, closer to raw CUDA than an array library; it is used here only as another DLPack producer, not taught in full.

A wp.array needs no dedicated adapter: eagle already recognizes any __dlpack__ + __dlpack_device__ producer, so the exact same force(position=position_warp, mu=MU) call dispatches it. Warp isn’t a “strong” output framework here (it doesn’t own a result the way torch does), so the result comes back as the neutral device type, cupy; getting a wp.array back is one explicit, zero-copy step — eagle.interop.import_buffer + wp.from_dlpack, the layer behind raptor’s WP-IN-CUDA-ALIAS / WP-OUT-CUDA-ALIAS rows.

import warp as wp
from eagle.interop import import_buffer

# suppress the init banner (it prints a local cache path)
wp.config.log_level = wp.LOG_WARNING
wp.init()
position_warp = wp.array(position_np, dtype=wp.float64, device="cuda:0")
accel_from_warp_input = force(position=position_warp, mu=MU)
print(type(accel_from_warp_input), detect(position_warp).name)

# explicit zero-copy round trip back to a wp.array
accel_warp = wp.from_dlpack(import_buffer(accel_from_warp_input))
np.testing.assert_allclose(accel_np, accel_warp.numpy())
same_storage = accel_warp.ptr == accel_from_warp_input.data.ptr
print("numpy and NVIDIA Warp agree; shares storage with the cupy result:", same_storage)
<class 'cupy.ndarray'> dlpack
numpy and NVIDIA Warp agree; shares storage with the cupy result: True

What’s certified today#

raptor’s interop certification matrix is the family’s single source of truth for which (framework, kind) combinations are exercised end to end — see the interoperability protocols reference. The table below is generated from that matrix, not typed by hand, so it cannot drift from it.

from raptor.conformance import interop as _interop

_KINDS = {"ALIAS": "zero-copy", "COPY": "copies once", "STREAM": "stream-ordered"}
rows = sorted(
    (d.framework, _KINDS[d.kind])
    for d in _interop.ROWS.values()
    if d.framework in _interop.CERTIFIED and d.kind in _KINDS
)
print(f"{'framework':<10} {'crossing'}")
for framework, kind in rows:
    print(f"{framework:<10} {kind}")
framework  crossing
cupy       stream-ordered
torch      copies once
torch      copies once
torch      stream-ordered
torch      stream-ordered
torch      stream-ordered
torch      zero-copy
torch      zero-copy
torch      zero-copy
torch      zero-copy
warp       stream-ordered
warp       zero-copy
warp       zero-copy

Roadmap

JAX, TensorFlow/Keras, and any other DLPack-capable framework are foreseen — eagle.interop.detect already recognizes a bare DLPack producer — but are not yet certified by raptor’s interop matrix. Treat only the frameworks in the table above as supported until a row for a new one lands there.