Example: a drag kernel over a large ensemble

Example: a drag kernel over a large ensemble#

One task: compile a small force kernel once, then evaluate it over a large ensemble of samples on the GPU and on the host, and check the two agree. This is the shape a batch propagator calls a force model in — many independent samples, one kernel, one launch.

Time: ~5 min · Runs on: GPU only (the comparison needs one) · You need: Your first kernel

from hawk import Mutable, Param, Scalar, Vector
from hawk.ext import Kind, Output
from hawk.math import norm

# two SINKS: a named Mutable `speed`, and the returned slot `acc`
ACCEL = Kind("drag", output=Output.returned(Vector[3], slot="acc"))


@ACCEL
def drag(velocity: Vector[3], k: Param, speed: Mutable[Scalar]):
    speed = norm(velocity)
    return (-k * norm(velocity)) * velocity


drag
<Kernel drag slots=4>

One plan, both targets#

eagle.deploy(drag) compiles drag for host and device through hawk’s build cache and returns a plan — an object that runs this kernel on whatever array you hand it, a cupy array on the GPU or a numpy array on the CPU, with no manifest, registry or plugin object to assemble by hand. plan.run returns a dict keyed by plane name when a kernel declares several sinks: {"acc": ..., "speed": ...} here — acc the returned slot, speed the named diagnostic.

import numpy as np
import cupy as cp

import eagle

plan = eagle.deploy(drag)
rng = np.random.default_rng(1)
n = 50_000
velocity = rng.normal(size=(3, n))          # (3, n_samples), random velocities

out_gpu = plan.run(velocity=cp.asarray(velocity), k=0.05)
acc_gpu, speed_gpu = out_gpu["acc"], out_gpu["speed"]
print(f"{n} samples, GPU. speed[:3] = {cp.asnumpy(speed_gpu)[:3]}")
50000 samples, GPU. speed[:3] = [1.71802099 1.91535842 1.02095317]
out_host = plan.run(velocity=velocity, k=0.05)   # same plan, numpy in: CPU
acc_host, speed_host = out_host["acc"], out_host["speed"]

np.testing.assert_allclose(speed_host, cp.asnumpy(speed_gpu))
np.testing.assert_allclose(acc_host, cp.asnumpy(acc_gpu))
print("host and GPU results agree across all", n, "samples")
host and GPU results agree across all 50000 samples

What just happened#

  • One eagle.deploy(drag) plan ran the SAME compiled kernel over 50,000 samples on the GPU and on the host, from the SAME Python call — only the array type (cupy vs numpy) changed.

  • A kernel with two sinks (a named Mutable and a returned slot) returns both from plan.run, as a dict keyed by plane name — {"acc": ..., "speed": ...}.

  • GPU and host results agree exactly, not approximately: the same compiled arithmetic ran both times.