Example: hawk next to NVIDIA Warp#
One task, written twice: a batch of projectiles integrated under gravity until each one hits the ground, each stopping at its own step.
Time: ~8 min · Runs on: GPU only (both tools need one) · You need: Stop when you’re done
NVIDIA Warp expresses this with explicit thread-level control: a thread index, a done-flag array, an early return, and a guarded write-back. hawk expresses it by declaring what each argument is, with termination as part of the kernel’s signature. Both run on the GPU here, and the two results agree exactly. See hawk next to Warp for the narrative version of this comparison.
import numpy as np
G = 9.81
DT = 0.02
N = 2048
H_MAX = 20.0
rng = np.random.default_rng(3)
y0 = rng.uniform(2.0, H_MAX, size=N) # initial height
x0 = np.zeros(N) # initial downrange distance
vx0 = rng.uniform(-3.0, 3.0, size=N) # initial horizontal speed
vy0 = np.zeros(N) # initial vertical speed
# enough steps for the highest sample to reach the ground, plus a margin
MAX_STEPS = int(np.ceil(np.sqrt(2 * H_MAX / G) / DT)) + 20
MAX_STEPS
121
Warp: idiomatic#
Warp’s own vocabulary is thread-level: wp.tid() gets this thread’s
index, and the kernel decides for itself what a “done” sample means. A
done flag array, an early return for a finished sample, and a guarded
write-back for one that crosses the ground this step — the arrays are
updated in place.
import warp as wp
wp.config.log_level = wp.LOG_WARNING # suppress the init banner (local cache path)
wp.init()
DEVICE = "cuda:0"
The kernel itself, thread-level and explicit:
@wp.kernel
def warp_step(x: wp.array(dtype=wp.float64), y: wp.array(dtype=wp.float64),
vx: wp.array(dtype=wp.float64), vy: wp.array(dtype=wp.float64),
done: wp.array(dtype=wp.int32), g: wp.float64, dt: wp.float64):
i = wp.tid()
if done[i] == 1:
return # early return: a finished sample costs nothing further
new_vy = vy[i] - g * dt
new_y = y[i] + new_vy * dt
new_x = x[i] + vx[i] * dt
if new_y <= wp.float64(0.0):
new_y = wp.float64(0.0)
done[i] = 1
# guarded write-back: only this still-live-this-step sample's state moves
x[i] = new_x
y[i] = new_y
vy[i] = new_vy
x_w = wp.array(x0, dtype=wp.float64, device=DEVICE)
y_w = wp.array(y0, dtype=wp.float64, device=DEVICE)
vx_w = wp.array(vx0, dtype=wp.float64, device=DEVICE)
vy_w = wp.array(vy0, dtype=wp.float64, device=DEVICE)
done_w = wp.zeros(N, dtype=wp.int32, device=DEVICE)
for _ in range(MAX_STEPS):
wp.launch(warp_step, dim=N, inputs=[x_w, y_w, vx_w, vy_w, done_w, G, DT],
device=DEVICE)
wp.synchronize_device(DEVICE)
print("all samples landed:", bool((done_w.numpy() == 1).all()))
all samples landed: True
hawk: termination in the signature#
terminated: Terminated is a declared argument role, not a value the
body reads and branches on: a sample flagged terminated going into a
launch keeps its output plane’s prior value, enforced by the runtime
itself, regardless of what the body computed for it.
This kernel keeps SEPARATE in/out planes (x/x_out, and so on) and a
host-side copy-back loop, on purpose: stop when you’re done’s
Mutable-plus-eagle.simulate idiom is the one to reach for in your own
code (it needs none of the loop below), but it hides exactly the
in-place-update step this page is comparing against Warp’s own explicit
loop — showing it here keeps the two tools at the same level of
abstraction, one in-place buffer update per launch, same as Warp’s.
import cupy as cp
import hawk
from hawk import Mutable, Param, Scalar, Terminated
from hawk.math import maximum
@hawk.kernel
def hawk_step(x: Scalar, y: Scalar, vx: Scalar, vy: Scalar, g: Param, dt: Param,
terminated: Terminated, x_out: Mutable[Scalar], y_out: Mutable[Scalar],
vx_out: Mutable[Scalar], vy_out: Mutable[Scalar]):
new_vy = vy - g * dt
new_y = maximum(y + new_vy * dt, 0.0) # clamp at the ground, Warp's guard too
x_out = x + vx * dt
y_out = new_y
vx_out = vx
vy_out = new_vy
eagle.deploy compiles it; the deployed plan binds persistent buffers:
import eagle
the_plan = eagle.deploy(hawk_step)
x_s, y_s, vx_s, vy_s = cp.asarray(x0), cp.asarray(y0), cp.asarray(vx0), cp.asarray(vy0)
x_o, y_o, vx_o, vy_o = cp.zeros(N), cp.zeros(N), cp.zeros(N), cp.zeros(N)
terminated_h = cp.zeros(N, dtype=cp.bool_)
bound = the_plan.bind(x=x_s, y=y_s, vx=vx_s, vy=vy_s, g=G, dt=DT,
terminated=terminated_h, x_out=x_o, y_out=y_o,
vx_out=vx_o, vy_out=vy_o)
for _ in range(MAX_STEPS):
bound.launch()
terminated_h[:] = terminated_h | (y_o <= 0.0)
x_s[:], y_s[:], vx_s[:], vy_s[:] = x_o, y_o, vx_o, vy_o
cp.cuda.get_current_stream().synchronize()
print("all samples landed:", bool(cp.asnumpy(terminated_h).all()))
all samples landed: True
y_warp, x_warp = y_w.numpy(), x_w.numpy()
y_hawk, x_hawk = cp.asnumpy(y_o), cp.asnumpy(x_o)
max_dy = float(np.max(np.abs(y_warp - y_hawk)))
max_dx = float(np.max(np.abs(x_warp - x_hawk)))
np.testing.assert_array_equal(y_warp, y_hawk)
np.testing.assert_array_equal(x_warp, x_hawk)
msg = f"Warp and hawk agree exactly over {N} samples:"
print(msg, f"max|dy|={max_dy}, max|dx|={max_dx}")
Warp and hawk agree exactly over 2048 samples: max|dy|=0.0, max|dx|=0.0
They work together#
A wp.array handed to a hawk kernel through eagle’s zero-copy DLPack
exchange (see Interoperability) — no copy, no intermediate
framework. cp.from_dlpack views Warp’s own device memory; the pointer
identity below proves it.
heights = np.array([10.0, 9.0, 8.0, 7.0, 6.0])
height_wp = wp.array(heights, dtype=wp.float64, device=DEVICE)
height_cp = cp.from_dlpack(height_wp) # zero-copy view of Warp's own buffer
print("zero-copy (same pointer):", height_cp.data.ptr == height_wp.ptr)
n5 = height_wp.shape[0]
y_out5 = cp.zeros(n5)
handoff = the_plan.bind(
x=cp.zeros(n5), y=height_cp, vx=cp.zeros(n5), vy=cp.zeros(n5), g=G, dt=DT,
terminated=cp.zeros(n5, dtype=cp.bool_),
x_out=cp.zeros(n5), y_out=y_out5, vx_out=cp.zeros(n5), vy_out=cp.zeros(n5))
handoff.launch()
cp.cuda.get_current_stream().synchronize()
print("one hawk step from a Warp-owned height plane:", y_out5.get())
zero-copy (same pointer): True
one hawk step from a Warp-owned height plane: [9.996076 8.996076 7.996076 6.996076 5.996076]
Where Warp shines#
Warp is the better choice when the job calls for its own strengths: a built-in geometry, mesh and simulation library, and reverse-mode gradients of a whole recorded simulation through its tape, rather than of one kernel at a time. Neither tool replaces the other — they pass arrays zero-copy and are reaching for different parts of the same problem.
What just happened#
The same physics, written against two different vocabularies, agrees exactly: Warp’s explicit thread-level control, and hawk’s declared
terminatedrole that the runtime enforces for you.eagle.deploy(...).bind(...)packs persistent device buffers once;.launch()replays the same compiled kernel every step with no new allocation — the same shape Warp’s own in-place arrays use.A
wp.arrayand acupyarray holding a hawk kernel’s output share memory with no copy, through the DLPack protocol both already speak.
Next#
Compiling and caching — what eagle.deploy’s cache
actually keys on, and when a change forces a real rebuild.