How RAPTOR compares#
The front page gives the headline numbers; this page is organised by the question a reader actually has, each answered from the same committed cards — never a number typed by hand.
Fastest arm, every batch size#
The full 80-row breakdown lives on eagle’s own performance page; this is the compact version: only the arm that won each cell.
Fastest arm, every batch size (GPU). Device: Quadro P2000 (compute capability 6.1, 8 SMs).
configuration |
N = 1,000 |
N = 10,000 |
N = 100,000 |
N = 1,000,000 |
|---|---|---|---|---|
spread, up to 100 steps |
eagle auto (eagle picks the launch mode) (125 µs) |
eagle auto (eagle picks the launch mode) (340 µs) |
eagle.simulate (eagle picks the launch mode) (2.02 ms) |
CPU OpenMP (15.7 ms) |
spread, up to 1000 steps |
CPU OpenMP (437 µs) |
CPU OpenMP (1.78 ms) |
CPU OpenMP (15.0 ms) |
CPU OpenMP (149 ms) |
uniform, 1000 steps each |
CPU OpenMP (419 µs) |
CPU OpenMP (2.20 ms) |
CPU OpenMP (19.5 ms) |
CPU OpenMP (191 ms) |
Fastest arm, every batch size (CPU). Device: Intel(R) Xeon(R) W-2125 CPU @ 4.00GHz (4 cores, 8 logical CPUs).
configuration |
N = 1,000 |
N = 10,000 |
N = 100,000 |
N = 1,000,000 |
|---|---|---|---|---|
spread, up to 100 steps |
eagle host, termination loop, up to 8 threads (55.7 µs) |
eagle host, termination loop, up to 8 threads (270 µs) |
eagle host, termination loop, up to 8 threads (1.68 ms) |
eagle host, termination loop, up to 8 threads (16.5 ms) |
spread, up to 1000 steps |
eagle host, termination loop, up to 8 threads (287 µs) |
eagle host, termination loop, up to 8 threads (2.36 ms) |
eagle host, termination loop, up to 8 threads (14.9 ms) |
eagle host, termination loop, up to 8 threads (145 ms) |
uniform, 1000 steps each |
eagle host, termination loop, up to 8 threads (303 µs) |
eagle host, termination loop, up to 8 threads (2.01 ms) |
eagle host, termination loop, up to 8 threads (22.5 ms) |
eagle host, termination loop, up to 8 threads (194 ms) |
Every arm, every column, every N: eagle’s performance page.
Organized by question instead — “my samples finish at different times”, “my batch is dense”, “I need derivatives”, “I want the CPU too”:
My samples finish at different times#
A batch where each sample stops on its own step (the “spread” distribution here) is exactly what eagle’s compaction targets: once enough samples finish, the live ones are packed into the next launch instead of carrying the whole batch along.
N = 1,000: fastest is
cpu_openmp(CPU OpenMP), 437 µs, of 15 arms run.N = 10,000: fastest is
cpu_openmp(CPU OpenMP), 1.78 ms, of 15 arms run.N = 100,000: fastest is
cpu_openmp(CPU OpenMP), 15.0 ms, of 15 arms run.N = 1,000,000: fastest is
cpu_openmp(CPU OpenMP), 149 ms, of 15 arms run.
My batch is dense (every sample runs the same number of steps)#
At N = 1,000,000 with nothing to compact, warp_kernel (693 ms) and eagle_graph_auto (699 ms) are a tie – a plain per-thread kernel is the better fit once nothing finishes early, and hawk + eagle’s plain graph arm matches it.
My batches are small#
At the smallest size this card runs (N = 1,000), launch and compaction overhead can cost more than they save:
spread, 100 steps: fastest is
eagle_graph_auto(eagle auto (eagle picks the launch mode)), 125 µs.spread, 1000 steps: fastest is
cpu_openmp(CPU OpenMP), 437 µs.uniform, 1000 steps: fastest is
cpu_openmp(CPU OpenMP), 419 µs.
Will it fit on my GPU?#
Same cell as the speed comparison above (N = 1,000,000, spread stop steps, up to 1000): peak GPU memory the driver accounted to the process, against the state the workload actually needs.
tool |
GPU memory |
× minimum |
|---|---|---|
eagle.simulate |
50 MiB |
1.31× |
hawk + eagle graph (no compaction) |
50 MiB |
1.31× |
NVIDIA Warp |
64 MiB |
1.68× |
JAX |
128 MiB |
3.36× |
CuPy |
194 MiB |
5.09× |
PyTorch |
262 MiB |
6.87× |
Array libraries like CuPy, PyTorch and JAX keep a temporary for every operation in a step, so their reservations stack up across the batch; hawk + eagle and NVIDIA Warp keep only the per-sample state the workload itself needs. eagle.simulate (1.3×) is lighter than NVIDIA Warp (1.7×).
My steps are adaptive#
The RK7(8) card integrates an adaptive-step orbit propagator (13-stage Runge-Kutta-Fehlberg, error-controlled) rather than a fixed-step scheme – the same comparison, a heavier per-step workload:
fastest GPU arm: hawk + eagle persistent, 13.0 s
Warp per-thread kernel: 26.3 s
JAX jit(vmap(while_loop)): 40.2 s
CuPy masked: 182 s
PyTorch masked: 184 s
eagle CPU (OpenMP): 7.15 s – beats every GPU arm above on this card
I need derivatives#
Every hawk kernel carries its reverse-mode (hawk.diff.vjp) and forward-mode (hawk.diff.jvp) derivative alongside the primal, generated from the same source and run on the GPU or CPU like the kernel itself. raptor’s own demo (examples/autodiff_vs_torch.py, a hand-run script, not a committed card) measured hawk’s vjp/jvp against torch.autograd on a Quadro P2000 – see raptor’s README for those numbers with their own method and caveats.
I want the CPU too#
Every kernel in this family runs unchanged on CPU threads (OpenMP), no GPU, no CUDA toolchain:
RK4 oscillators, N = 1,000,000: eagle host, termination loop, up to 8 threads, 145 ms.
On the RK7(8) card, the CPU arm (eagle CPU (OpenMP), 7.15 s) beats every GPU arm recorded there – this GPU’s FP64 throughput is modest; a data-centre GPU is expected to flip this back.
Every number on this page traces to a committed card: eagle’s performance page.