How RAPTOR compares#

The front page gives the headline numbers; this page is organised by the question a reader actually has, each answered from the same committed cards — never a number typed by hand.

Fastest arm, every batch size#

The full 80-row breakdown lives on eagle’s own performance page; this is the compact version: only the arm that won each cell.

Fastest arm, every batch size (GPU). Device: Quadro P2000 (compute capability 6.1, 8 SMs).

configuration

N = 1,000

N = 10,000

N = 100,000

N = 1,000,000

spread, up to 100 steps

eagle auto (eagle picks the launch mode) (125 µs)

eagle auto (eagle picks the launch mode) (340 µs)

eagle.simulate (eagle picks the launch mode) (2.02 ms)

CPU OpenMP (15.7 ms)

spread, up to 1000 steps

CPU OpenMP (437 µs)

CPU OpenMP (1.78 ms)

CPU OpenMP (15.0 ms)

CPU OpenMP (149 ms)

uniform, 1000 steps each

CPU OpenMP (419 µs)

CPU OpenMP (2.20 ms)

CPU OpenMP (19.5 ms)

CPU OpenMP (191 ms)

Fastest arm, every batch size (CPU). Device: Intel(R) Xeon(R) W-2125 CPU @ 4.00GHz (4 cores, 8 logical CPUs).

configuration

N = 1,000

N = 10,000

N = 100,000

N = 1,000,000

spread, up to 100 steps

eagle host, termination loop, up to 8 threads (55.7 µs)

eagle host, termination loop, up to 8 threads (270 µs)

eagle host, termination loop, up to 8 threads (1.68 ms)

eagle host, termination loop, up to 8 threads (16.5 ms)

spread, up to 1000 steps

eagle host, termination loop, up to 8 threads (287 µs)

eagle host, termination loop, up to 8 threads (2.36 ms)

eagle host, termination loop, up to 8 threads (14.9 ms)

eagle host, termination loop, up to 8 threads (145 ms)

uniform, 1000 steps each

eagle host, termination loop, up to 8 threads (303 µs)

eagle host, termination loop, up to 8 threads (2.01 ms)

eagle host, termination loop, up to 8 threads (22.5 ms)

eagle host, termination loop, up to 8 threads (194 ms)

Every arm, every column, every N: eagle’s performance page.

Organized by question instead — “my samples finish at different times”, “my batch is dense”, “I need derivatives”, “I want the CPU too”:

My samples finish at different times#

A batch where each sample stops on its own step (the “spread” distribution here) is exactly what eagle’s compaction targets: once enough samples finish, the live ones are packed into the next launch instead of carrying the whole batch along.

  • N = 1,000: fastest is cpu_openmp (CPU OpenMP), 437 µs, of 15 arms run.

  • N = 10,000: fastest is cpu_openmp (CPU OpenMP), 1.78 ms, of 15 arms run.

  • N = 100,000: fastest is cpu_openmp (CPU OpenMP), 15.0 ms, of 15 arms run.

  • N = 1,000,000: fastest is cpu_openmp (CPU OpenMP), 149 ms, of 15 arms run.

Full table, this card.

My batch is dense (every sample runs the same number of steps)#

At N = 1,000,000 with nothing to compact, warp_kernel (693 ms) and eagle_graph_auto (699 ms) are a tie – a plain per-thread kernel is the better fit once nothing finishes early, and hawk + eagle’s plain graph arm matches it.

Full table, this card.

My batches are small#

At the smallest size this card runs (N = 1,000), launch and compaction overhead can cost more than they save:

  • spread, 100 steps: fastest is eagle_graph_auto (eagle auto (eagle picks the launch mode)), 125 µs.

  • spread, 1000 steps: fastest is cpu_openmp (CPU OpenMP), 437 µs.

  • uniform, 1000 steps: fastest is cpu_openmp (CPU OpenMP), 419 µs.

Will it fit on my GPU?#

Same cell as the speed comparison above (N = 1,000,000, spread stop steps, up to 1000): peak GPU memory the driver accounted to the process, against the state the workload actually needs.

tool

GPU memory

× minimum

eagle.simulate

50 MiB

1.31×

hawk + eagle graph (no compaction)

50 MiB

1.31×

NVIDIA Warp

64 MiB

1.68×

JAX

128 MiB

3.36×

CuPy

194 MiB

5.09×

PyTorch

262 MiB

6.87×

Array libraries like CuPy, PyTorch and JAX keep a temporary for every operation in a step, so their reservations stack up across the batch; hawk + eagle and NVIDIA Warp keep only the per-sample state the workload itself needs. eagle.simulate (1.3×) is lighter than NVIDIA Warp (1.7×).

My steps are adaptive#

The RK7(8) card integrates an adaptive-step orbit propagator (13-stage Runge-Kutta-Fehlberg, error-controlled) rather than a fixed-step scheme – the same comparison, a heavier per-step workload:

  • fastest GPU arm: hawk + eagle persistent, 13.0 s

  • Warp per-thread kernel: 26.3 s

  • JAX jit(vmap(while_loop)): 40.2 s

  • CuPy masked: 182 s

  • PyTorch masked: 184 s

  • eagle CPU (OpenMP): 7.15 s – beats every GPU arm above on this card

Full table, this card.

I need derivatives#

Every hawk kernel carries its reverse-mode (hawk.diff.vjp) and forward-mode (hawk.diff.jvp) derivative alongside the primal, generated from the same source and run on the GPU or CPU like the kernel itself. raptor’s own demo (examples/autodiff_vs_torch.py, a hand-run script, not a committed card) measured hawk’s vjp/jvp against torch.autograd on a Quadro P2000 – see raptor’s README for those numbers with their own method and caveats.

I want the CPU too#

Every kernel in this family runs unchanged on CPU threads (OpenMP), no GPU, no CUDA toolchain:

  • RK4 oscillators, N = 1,000,000: eagle host, termination loop, up to 8 threads, 145 ms.

  • On the RK7(8) card, the CPU arm (eagle CPU (OpenMP), 7.15 s) beats every GPU arm recorded there – this GPU’s FP64 throughput is modest; a data-centre GPU is expected to flip this back.

Every number on this page traces to a committed card: eagle’s performance page.