RAPTOR#
Write the code for one sample. Run a million of them.
A million samples, each stopping at its own step: 156 ms and 50 MiB on a modest workstation GPU, 3.9× faster than NVIDIA Warp and 47× faster than PyTorch or CuPy. The same Python source runs on your CPU.
pip install "raptor-hawk[cuda12]" "raptor-eagle[cuda12]"
What you write#
import cupy as cp
import eagle
import hawk
from hawk import Mutable, Param, Scalar, Terminated
@hawk.kernel
def oscillator(
omega: Scalar, t_end: Param, dt: Param,
terminated: Terminated,
x: Mutable[Scalar], v: Mutable[Scalar],
t: Mutable[Scalar],
):
x, v = x + dt * v, v - dt * omega * omega * x
t += dt
terminated = t >= t_end
n = 1_000_000
result = eagle.simulate(
oscillator,
omega=cp.linspace(1.0, 3.0, n), t_end=1.0, dt=1e-3,
x=cp.ones(n), v=cp.zeros(n), t=cp.zeros(n),
max_steps=10_000,
)
print(result.status, result.steps, result.x[:3])
# needs a GPU to run (1,000,000 cupy samples); the output is not shown here
What happened
Pass the kernel’s arguments as you would call it; give an array where each sample has its own value.
eagle captured the step loop as a CUDA graph and replayed it on the GPU, so one call did all the launching.
Samples that finished were dropped from the next launch, so the batch shrank as it ran.
Pass NumPy arrays instead of CuPy and the same call runs on your CPU threads.
How it compares#
Log scale, from 97.3 ms to 11.8 s.
Memory, as a multiple of what the samples need
NVIDIA Warp is an excellent per-thread kernel compiler: when every sample runs all 1000 steps it ties hawk + eagle (693 ms against 699 ms).
On memory, hawk + eagle's lower-level building blocks go lighter than Warp: 1.3× the minimum against Warp's 1.7×. eagle.simulate keeps room to compact the batch and uses 1.3×.
Measured on Quadro P2000, a development GPU. Every number, every arm.
Full numbers, every arm
tool |
wall (N = 1,000,000, spread stop steps, up to 1000) |
|---|---|
eagle.simulate (eagle picks the launch mode) |
156 ms |
Warp per-thread kernel |
609 ms |
JAX jit + vmap(while_loop) |
1.36 s |
CuPy masked |
7.38 s |
PyTorch masked |
7.40 s |
Tensor libraries compute every sample, every step – a masked array keeps computing the samples that already finished.
hawk + eagle skips finished samples and packs the live ones into the next launch, so the batch keeps shrinking as samples finish.
Warp is excellent at per-thread kernels: it ties hawk + eagle here on a fully dense batch (693 ms vs 699 ms) – complementary tools, not a hierarchy.
RAPTOR adds, from the same source: per-sample Python, device-side stopping, forward- and reverse-mode derivatives, and the identical kernel on the CPU.
On the harder RK7(8) family (N = 1,000,000 adaptive Kepler orbits, every one inside its error bound): hawk + eagle persistent 13.0 s, Warp per-thread kernel 26.3 s, JAX jit(vmap(while_loop)) 40.2 s, CuPy masked 182 s, PyTorch masked 184 s. hawk + eagle on the CPU alone, no GPU at all: 7.15 s – the fastest number on this card, because this GPU’s FP64 throughput is modest. Honest, and worth knowing before reaching for a GPU.
Measured on Quadro P2000, a development GPU. Every number, every arm.
On your CPU#
No GPU needed: the same call runs on your CPU’s threads. Here is one desktop CPU running the same batch.
1,000,000 samples, each stopping at its own step (up to 1,000). Log scale, from 90.5 ms to 101 s. Measured on Intel Xeon W-2125 (4 cores, 8 logical CPUs), card of 2026-10-08. Every tool, every N, and where each one fits.
Four libraries#
Install
pip install raptor-core # the protocol spine: Python 3.9+, no dependencies
pip install raptor-hawk # hawk, CPU only (Linux x86_64, CPython 3.9-3.14, host g++ 11+)
pip install "raptor-hawk[cuda12]" "raptor-eagle[cuda12]" # GPU route; use [cuda13] on both for CUDA 13
Important
NVIDIA’s GPU packages come only through the extras. raptor-hawk[cuda12] pulls cuda-bindings 12,
nvidia-cuda-nvrtc-cu12 and nvidia-cuda-cccl-cu12; raptor-hawk[cuda13] pulls cuda-bindings 13,
nvidia-cuda-nvrtc 13 and nvidia-cuda-cccl 13 (CUDA 13’s wheels have no -cu13 suffix);
raptor-eagle[cuda12] / [cuda13] pull CuPy (cupy-cuda12x / cupy-cuda13x, with the CUDA headers CuPy compiles against). Without an extra pip installs no NVIDIA package: you get the CPU route, or the GPU route through a CUDA setup you already have. Pick the extra matching the CUDA version your driver reports (nvidia-smi, top right). No nvcc or CUDA
toolkit is needed.
Platforms: built and tested on Linux x86_64 only so far (CPython 3.9–3.14, including free-threaded 3.13t and 3.14t), on NVIDIA GPUs from Pascal (Quadro P2000) and Turing (Tesla T4). There are no wheels for macOS, Windows or ARM yet, and WSL2 is untested. raptor-core and aether-dsc are pure Python and install anywhere. On free-threaded 3.13t and 3.14t, hawk and eagle run GIL-free.
raptor-hawk pulls aether-dsc (the sealed C++ headers hawk compiles against) automatically. eagle alone:
pip install "raptor-eagle[cuda12]" ([torch] adds PyTorch interop); its wheel ships GPU code for every NVIDIA
architecture from Pascal (sm_60) through Blackwell plus PTX for newer GPUs. aether is a header-only C++ library used through CMake;
aether-dsc on PyPI is not something C++ users install. To build everything from source (contributors, C++ users,
your own toolchain), see Building from source.
Citing
Each repository is cited independently: its own CITATION.cff carries the metadata (GitHub shows it under “Cite
this repository”), and every tagged release is archived on Zenodo: hawk doi:10.5281/zenodo.23250242, eagle doi:10.5281/zenodo.23250240, aether doi:10.5281/zenodo.23250238, raptor doi:10.5281/zenodo.23250234.
License
All four repositories — aether, raptor, eagle and hawk — are licensed under the Apache License, Version 2.0: free
for any use, commercial or not, with attribution. See each repository’s own LICENSE for the full text.