Example: a drag kernel over a large ensemble#
One task: compile a small force kernel once, then evaluate it over a large ensemble of samples on the GPU and on the host, and check the two agree. This is the shape a batch propagator calls a force model in — many independent samples, one kernel, one launch.
Time: ~5 min · Runs on: GPU only (the comparison needs one) · You need: Your first kernel
from hawk import Mutable, Param, Scalar, Vector
from hawk.ext import Kind, Output
from hawk.math import norm
# two SINKS: a named Mutable `speed`, and the returned slot `acc`
ACCEL = Kind("drag", output=Output.returned(Vector[3], slot="acc"))
@ACCEL
def drag(velocity: Vector[3], k: Param, speed: Mutable[Scalar]):
speed = norm(velocity)
return (-k * norm(velocity)) * velocity
drag
<Kernel drag slots=4>
One plan, both targets#
eagle.deploy(drag) compiles drag for host and device through hawk’s
build cache and returns a plan — an object that runs this kernel on
whatever array you hand it, a cupy array on the GPU or a numpy array on
the CPU, with no manifest, registry or plugin object to assemble by
hand. plan.run returns a dict keyed by plane name when a kernel
declares several sinks: {"acc": ..., "speed": ...} here — acc the
returned slot, speed the named diagnostic.
import numpy as np
import cupy as cp
import eagle
plan = eagle.deploy(drag)
rng = np.random.default_rng(1)
n = 50_000
velocity = rng.normal(size=(3, n)) # (3, n_samples), random velocities
out_gpu = plan.run(velocity=cp.asarray(velocity), k=0.05)
acc_gpu, speed_gpu = out_gpu["acc"], out_gpu["speed"]
print(f"{n} samples, GPU. speed[:3] = {cp.asnumpy(speed_gpu)[:3]}")
50000 samples, GPU. speed[:3] = [1.71802099 1.91535842 1.02095317]
out_host = plan.run(velocity=velocity, k=0.05) # same plan, numpy in: CPU
acc_host, speed_host = out_host["acc"], out_host["speed"]
np.testing.assert_allclose(speed_host, cp.asnumpy(speed_gpu))
np.testing.assert_allclose(acc_host, cp.asnumpy(acc_gpu))
print("host and GPU results agree across all", n, "samples")
host and GPU results agree across all 50000 samples
What just happened#
One
eagle.deploy(drag)plan ran the SAME compiled kernel over 50,000 samples on the GPU and on the host, from the SAME Python call — only the array type (cupyvsnumpy) changed.A kernel with two sinks (a named
Mutableand a returned slot) returns both fromplan.run, as a dict keyed by plane name —{"acc": ..., "speed": ...}.GPU and host results agree exactly, not approximately: the same compiled arithmetic ran both times.