Zero-copy interop#
One hawk kernel (y = k * x, a per-sample vector scale), fed from and
handed back into NumPy, CuPy and PyTorch arrays with no copy where
raptor.conformance.interop’s own declaration says there is none. This
notebook runs the --host arm, which needs no GPU, cupy or torch — the
kernel runs on CPU threads through eagle.exec.HostTeam, reading and
writing the caller’s own numpy arrays directly.
python examples/zero_copy_interop.py --host
Needs hawk and eagle installed from PyPI.
zero_copy_interop from examples/zero_copy_interop.py
The kernel this notebook hands to every framework is a plain per-sample vector scale:
import inspect
print(inspect.getsource(demo.step_kernel))
def step_kernel():
"""Trace ``y = k * x``: a ``Vector[3]`` input, a uniform, a mutable output."""
import hawk
@hawk.kernel
def zc_scale(x: hawk.Vector[3], k: hawk.Param, y: hawk.Mutable[hawk.Vector[3]]):
y = k * x
return zc_scale
By hand: seed a batch of positions, then let hawk publish the compiled kernel into a manifest raptor validates.
import pathlib, tempfile
import numpy as np
import hawk.artifact
from raptor.schema import validate_manifest
import eagle.exec as eexec
from eagle import plan as eplan
positions = demo.seed_positions(64, seed=0)
bundle = hawk.build(
[demo.step_kernel()], pathlib.Path(tempfile.mkdtemp()),
mode="float32", targets=("host",))
validate_manifest(bundle.manifest)
print("exec_targets:", bundle.manifest["exec_targets"])
exec_targets: ['host']
Then hand it to eagle — this is the one call that actually launches the kernel, reading and writing the caller’s own numpy arrays:
plugin = next(iter(hawk.artifact.plugins(bundle).values()))
plan = eplan.plan(plugin, structure=eexec.HostTeam)
x = np.ascontiguousarray(positions, dtype=np.float32)
y = np.zeros_like(x)
ptr_before = x.ctypes.data
plan.bind(x=x, k=2.0, y=y).launch()
print("input array address unchanged:", x.ctypes.data == ptr_before)
input array address unchanged: True
The pieces above, assembled into one call — building, running and checking the crossing, for reference:
results = demo.run_demo(target="host")
target : host (manifest exec_targets ['host'])
batch x k : 64 x 2.0
host arm: plain numpy in/out through eagle.exec.HostTeam -- no row in the matrix governs this crossing (eagle's own host currency)
input array pointer stable across launch = True
result vs numpy reference: max|rel diff| = 0.000e+00 (tolerance 1e-05)
ptr_stable is the thing this crossing actually promises: the host array’s
address never changes across the launch — the kernel wrote into a separate
output array, it did not reallocate the input. No row in the certification
matrix governs this leg (plain numpy through eagle.exec.HostTeam is
eagle’s own host currency, not a framework crossing).
The device rows this notebook does not execute#
--device additionally certifies three rows, run here as read, not
executed — this build’s torch is CPU-only:
T-IN-CUDA-ALIAS/T-OUT-CUDA-ALIAS— a CUDA-resident torch tensor crosses into and out of the kernel as a zero-copy view (eagle.to_cupy), checked by pointer identity and by mutating the input in place and relaunching the same bound plan.T-IN-CPU-COPY— a CPU-resident torch tensor crosses through the identical adapter, which uploads it; the row declares a copy, not an alias, and the example asserts the pointers differ.
A CUDA-capable torch build exchanges device tensors the identical way,
through the same __dlpack__ protocol a cupy array uses — nothing about
the crossing mechanism changes between the two frameworks on the device
side. See Interoperability for what each row
certifies and which repository’s suite proves it.