Performance#
RAPTOR lets you focus on the application while still making efficient use of GPUs. This page measures that on two workloads, a fixed-step one and an adaptive one, against other common ways of writing the same computation, on both a GPU and a CPU, and publishes every number the run produced, including the regimes where another approach is faster.
Why this workload#
A large class of scientific code integrates many independent problems at once: one trajectory, one particle or one parameter sample per GPU thread, each with its own state, its own control flow and its own stopping point. Written for the GPU, this class tends to lose efficiency in four places:
launch overhead, when every step is a separate kernel launch from the host;
host round trips, when the host has to read a device value every step to decide whether to continue;
finished samples, when samples that have stopped keep occupying the batch and keep being computed;
idle lanes, when a launch is too small to fill the device.
The benchmark is built to exercise all four. It integrates a batch of damped harmonic oscillators, each with its own frequency, damping ratio and initial state, using a fixed-step classical Runge-Kutta (RK4) scheme. Each sample stops after its own number of steps. In the spread configurations those step counts are drawn log-uniformly over two decades, so most of the batch finishes early; in the uniform configuration every sample runs the full count, a dense, all-active batch. The batch size N runs from 10³ to 10⁶.
The step is written once, as a hawk kernel, one sample at a time:
@hawk.kernel
def oscillator_step(omega: Scalar, zeta: Scalar, nstop: Scalar, dt: Param,
terminated: Terminated, x: Mutable[Scalar],
v: Mutable[Scalar], k: Mutable[Scalar]):
# One classical RK4 step of x'' + 2 zeta omega x' + omega^2 x = 0.
x0 = x
v0 = v
w2 = omega * omega
c = 2.0 * zeta * omega
h2 = 0.5 * dt
h6 = dt * 0.16666666666666666
a1 = -w2 * x0 - c * v0
x2 = x0 + h2 * v0
v2 = v0 + h2 * a1
a2 = -w2 * x2 - c * v2
x3 = x0 + h2 * v2
v3 = v0 + h2 * a2
a3 = -w2 * x3 - c * v3
x4 = x0 + dt * v3
v4 = v0 + dt * a3
a4 = -w2 * x4 - c * v4
x = x0 + h6 * (v0 + 2.0 * v2 + 2.0 * v3 + v4)
v = v0 + h6 * (a1 + 2.0 * a2 + 2.0 * a3 + a4)
k1 = k + 1.0
k = k1
terminated = k1 >= nstop # the sample stops after its own nstop steps
terminated: Terminated is the per-sample stop mask: a finished sample keeps
its state. x, v and k are read as plain variables holding last launch’s
value and reassigned to the next one, which is what updates the state in
place, launch after launch.
The ways of running it#
All arms do the same arithmetic in the same precision (float64) on the same samples, on the same device; the card’s arm table lists each one’s loop structure, and the card records every arm’s code.
hawk + eagle. The kernel is built with eagle.deploy() and run until every
sample has finished with eagle.until_done(): the step, which finishes
its own samples, is captured once as the body of eagle.repeat_while(),
one CUDA graph whose WHILE node repeats it until the finished count reaches
N. One graph launch runs the whole integration; the host does not read
anything until it is done.
# eagle_graph_steps: the same one call, on the K-steps kernel; forced
# off the fast path (it is a fixed-K kernel, not "auto", so it would
# never be fast-routed, but the row's device-loop labels/byte model
# assume the band -- forced explicitly so a future kernel change
# cannot silently drift this arm onto the fast path unnoticed)
s = _GpuArm(n)
s.runner = eagle.until_done(eagle.deploy(oscillator_step), max_steps=max_steps,
dt=DT, _fast_mode=False, **s.planes())
The plain @hawk.kernel decorator builds the automatic kernel. By default
eagle.until_done() auto-routes it past this device-loop/”band” shape
entirely: at or below the device’s own latency-regime capacity it takes one
fused launch of the whole run, above it the persistent (work-stealing)
launch – both no-graph, no-map, one host-device sync total. The band shown
above (each launch running as many steps per sample as the policy picks, 8
to 64, with the state in registers) is still reachable, and is what this
row forces explicitly so its loop-structure description, byte model and
kernel count stay meaningful; it is also what any artifact falls back to
when a reorder is requested, or when the fast entries are unavailable. The
other hawk + eagle arms vary one thing at a time:
compaction: the single step traced under a kind whose guard reads an active-set index map (hawk
Guard(active_set=True)); every 16 steps, when a sample finished since the last time, aneagle.ActiveSetrecomputes the map on the device, inside the same graph, so the launch covers only the samples still running;compaction + reorder: the same loop, plus an occasional physical reorder of every per-sample plane once the live samples spread thin over their warps (
eagle.ActiveSet(mask, reorder=theta)), restored to sample order at the end;K fused steps per launch:
@hawk.kernel(steps=16), with and without compaction (every 64 steps);eagle auto (eagle picks the launch mode):
@hawk.kernel(steps="auto")on the active-set kind – this row is the one left oneagle.until_done’s own default pick, so it takes the fast, no-graph launch described above whenever capacity allows and the band (compacting after every launch in which a sample finished) only above it; the card records which one each run actually took;hawk + eagle persistent: the plain automatic kernel forced through hawk’s persist entry – one launch, no graph, no map, no policy kernel, a grid sized from the SM count in which each lane takes the next unfinished sample off a counter whenever its own finishes;
eagle.simulate: the same automatic kernel run by
eagle.simulate()(its bound form,eagle.simulation, so the repetitions reuse one build), given its state and parameters by name – it auto-routes the same way;eager launch loop: the single step (
.step, one step per launch) launched from a Python loop with a host read of the finished count after every step: the most direct loop, every intermediate state one host read away.
# eagle_simulate: the model, its state and its parameters by name; it
# runs until every sample has finished
m = _GpuArm(n)
m.runner = eagle.simulation(oscillator_step, omega=m.omega, zeta=m.zeta,
nstop=m.nstop, dt=DT, x=m.x, v=m.v, k=m.k,
max_steps=max_steps)
CuPy and PyTorch, masked. The same arithmetic as array expressions over
the whole batch, with where keeping finished samples unchanged and a host
check that some sample is still running after every step; PyTorch also runs
with a block of 16 steps captured as a CUDA graph and replayed. This is how
the computation is commonly written in an array framework: no kernel to
write, and elementwise kernels that stream memory efficiently. Finished
samples are still computed and then discarded.
JAX and Warp. JAX compiles a per-sample lax.while_loop, batched with
vmap, into one executable; Warp runs an explicit per-thread kernel, one
thread per sample, each stopping at its own step.
CPU OpenMP. The same hawk kernel: its planes on the host select the host
side of the same eagle.deploy, run by hawk + eagle’s OpenMP host team on at most
8 threads. It needs no device and no transfers, and it is the correctness
reference the other arms are checked against.
Method#
One block per arm. For each configuration and N, every arm is timed in its own block, one arm after the other, with no interleaving. A block is two warm-up runs (excluded; the first also verifies the result), then untimed back-to-back runs of that arm until 0.5 s has elapsed (at least one run), so the clocks are settled, then 5 timed runs. Every reported time is the mean of the middle 3 of the 5 runs (the single highest and the single lowest dropped), shown with the range of those 3; the median, the interquartile range and every run are kept in the card’s JSON.
Cool-down gate and clock record. Before each arm’s block the card waits until the GPU temperature is within 3 °C of the idle temperature it read at start and the GPU reads idle, or for at most 20 s (the wait, the temperature and whether it timed out are recorded per block). Right after each timed run it reads the SM clock and the clock-throttle reasons through NVML, outside the timed region; the tables show the range of the SM clock per arm, and the FP32 table adds the ratio in SM clock cycles (wall time times the clock read after the run) next to the wall ratio. Wall time stays the headline number. Without NVML the gate is skipped and no clock is recorded.
Agreement before timing. The warm-up run is also the verification run: every arm must reach exactly the expected step count for every sample, and its final state must match the CPU arm within the absolute tolerance recorded in the card. The largest difference each arm showed is recorded per row.
Times. Wall is the integration alone, from inputs resident on the device to the device synchronised. End-to-end adds uploading the inputs from host memory and downloading the results. Kernel-only is the sum of the durations of the GPU kernels one run executes, measured with Nsight Systems in a separate traced pass (for the graph arm it includes hawk + eagle’s loop-control kernels); tracing makes kernels run slightly longer, so at the largest N it can exceed the untraced wall time. For the CPU arm it is the time spent inside the host-team launches.
Counts. FLOPs and bytes are counted analytically from the kernel’s operations, loads and stores; the counting rules are stored in the card. Useful FLOP/s counts only the steps of samples that are still running, the same yardstick for every arm. Issued B/s counts the bytes each arm’s own kernels move, including the passes over finished samples. Both are also given as a fraction of the device’s peak, computed from its reported properties as described in the card. This device does not expose DRAM byte counters to the profiler, so bytes are analytic only.
Device state. The card records the device’s clock, temperature and utilisation, sampled during the repetitions.
Compaction cadence sweep.
eagle_graphchecks its stop guard every step;eagle_graph_compactchecks it once every 16 steps (the cadence its compaction runs on) and reads an active-set map built from it. So the main table’s two graph rows differ in both the cadence and the map, not one variable. The sweep at the end of the results below holds the cadence fixed at a few values (K = 8, 16, 32) and times two arms at each: the same plain step kernel aseagle_graph, looped K steps between guard checks with no map, againsteagle_graph_compactat that K. The only structural difference within a pair is the active-set map.
Results#
The RK4 oscillator workload runs on both a GPU and a CPU; each device gets
its own card, generated by its own script (perf_card.py / cpu_card.py),
so the arm lists differ (the GPU card also includes a CPU OpenMP row as its
correctness reference, at 8 threads only; the CPU card is the fuller
CPU-side comparison, with 1- and 8-thread hawk + eagle rows plus Numba, JAX,
multiprocessing and plain Python).
GPU (Quadro P2000, Tesla T4)#
The GPU card is measured on two devices: the Quadro P2000, the reference card for every ratio quoted on this page, and a Tesla T4 (Kaggle, power-capped at 70 W, CUDA 12.6). On the T4 we observed significant SM clock fluctuations between and within arm blocks, so its cells move more from run to run than the P2000’s; the card records the SM clock of every timed run.
Measured on#
device |
compute capability / CPU model |
date |
section |
|---|---|---|---|
Quadro P2000 |
6.1 |
2026-10-08 |
|
Tesla T4 |
7.5 |
2026-10-08 |
GPU (Quadro P2000)#
Device: Quadro P2000 (compute capability 6.1, 8 SMs). Peak FP64: 9.48e10 FLOP/s. Peak DRAM bandwidth: 1.40e11 B/s. Mean of the middle 3 of 5 runs of each arm (the highest and the lowest dropped), range of those 3 in parentheses.
Each arm’s block starts once the GPU is within 3 °C of its idle 63 °C and reads idle, or after 20 s; the SM clock is read right after each timed run and its range is in the last column.
The FP64 peak above is at the device’s reported maximum clock (1480.5 MHz). During the timed repetitions the SM clock read 1075 MHz; at that clock the FP64 peak is 6.88e10 FLOP/s, and every FP64 percentage below is 1.38× larger against it.
“% of issue peak” (the eagle arms only) counts DP instructions (DFMA/DMUL/DADD/DSETP/MUFU), not useful FLOPs: the step’s SASS has 29 of them, against the device’s DP issue-slot throughput at the observed clock – a kernel that is already at the issue pipe’s limit can still read well under 100% of FP64 peak, because useful FLOPs/step is smaller than issued DP instructions/step.
eagle_graph checks its stop guard every step; eagle_graph_compact checks it once every 16 steps (the cadence its compaction runs on) and reads an active-set map built from it. The two rows below differ in both the cadence and the map; the compaction cadence sweep further down holds the cadence fixed at a few values and isolates the map alone.
PyTorch, CUDA graph of 16 steps captures 16 masked steps once as a torch.cuda.CUDAGraph over static buffers and replays it, with one host check of “any sample running” per replay: it removes the per-kernel launch overhead and keeps the masking (every sample is computed every step, finished ones held by torch.where, so the steps a replay runs past the last stop change nothing). JAX’s while loop is one compiled executable with no host check per step; batched by vmap, it runs until the slowest sample stops. Warp runs one thread per sample, each stopping at its own step.
PyTorch torch.compile (reduce-overhead): needs compute capability 7.0 or newer (Triton, Inductor’s GPU code generator); this GPU is 6.1, so it was not run.
Every arm runs each sample to its own stop and no further than max_steps (S): eagle and Warp enforce the cap per sample, JAX in its while-loop condition, PyTorch and CuPy by the outer step count (a batch-level cap); an arm that runs whole blocks of steps per launch or replay stops at the first block boundary at or after it.
What each arm needed to get there: Warp, a hand-written step counter and cap inside the kernel; JAX, a cap in the while_loop condition; eagle, the kernel source unchanged, with its loop shape, unrolling and device entry chosen by hawk + eagle (per device, per batch).
Cap fixture (checked before any timing counts): N = 4,096, 8 samples with stop step 2S = 256, S = 128; every arm (15) reported k = min(stop step, S) per sample and the same finished count (4,088): passed.
arm |
loop structure |
|---|---|
eagle graph (device loop) |
per launch over the batch, inside one CUDA graph: the plain decorator’s kernel, which finishes its own samples and runs the steps per launch eagle’s policy picks (8 to 64), guard checked on the device every launch |
eagle graph + compaction |
as the graph arm, the step launched over the active-set map only (recomputed every 16 steps) |
eagle graph + compaction + reorder |
as the compaction arm, plus a physical reorder of the per-sample planes when the live samples spread thin |
eagle graph, K fused steps per launch |
per launch over the batch, inside one CUDA graph: the step kernel runs up to K = 16 steps per sample with the state in registers, each sample leaving at its own stop step; guard checked on the device every launch |
eagle graph, K fused steps per launch + compaction |
as the K-steps arm, the launch over the active-set map only (recomputed every 64 steps = 4 launches) |
eagle auto (eagle picks the launch mode) |
the active-set kernel under eagle’s automatic policy, which picks the launch mode from the batch: one launch running every sample to its end for a small batch, the persistent launch above that, and the compaction loop (active-set map, steps per launch picked on the device) only when the caller asks for compaction or reordering; the row records the mode taken |
hawk + eagle persistent |
ONE launch, no graph, no map, no policy kernel: a grid sized from the SM count and the kernel’s occupancy, each lane fetching base + atomicAdd(counter, 1) whenever its sample is done, running the whole budget’s worth of steps per sample it picks up |
eagle.simulate (eagle picks the launch mode) |
as the automatic arm: eagle.simulate takes the kernel, its state and its parameters by name and eagle picks the launch mode the same way |
eager launch loop |
per step over the batch: one kernel launch from Python, a host read of the finished count every step |
CuPy masked |
per step over the batch: whole-batch array expressions, finished samples held by cupy.where, a host check every step; the cap is batch-level (the outer step loop runs at most max_steps) |
PyTorch masked |
per step over the batch: whole-batch tensor expressions, finished samples held by torch.where, a host check every step; the cap is batch-level (the outer step loop runs at most max_steps) |
PyTorch, CUDA graph of 16 steps |
as PyTorch masked, 16 steps captured as one CUDA graph over static buffers, replayed with one host check per replay |
JAX jit + vmap(while_loop) |
per step over the batch, inside one compiled executable: the vmapped while loop runs until the slowest sample stops or the step cap is reached, finished samples held by a select |
Warp per-thread kernel |
per sample: one launch, each thread loops over its own sample’s steps and stops at its own stop step or the step cap, whichever comes first |
CPU OpenMP |
per step over the batch on the host: one host-team launch per step, the kernel marks and counts its finished samples |
Spread stop steps (log-uniform in [1, 100]), up to 100 steps#
N |
arm |
wall (range) |
kernel-only |
end-to-end |
sample·steps/s |
useful FLOP/s (% peak) |
issued B/s (% peak) |
device memory (× minimum) |
host memory |
% issue peak |
lane util |
launches (kernel/memcpy/API) |
SM clock (MHz) |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
1,000 |
eagle graph (device loop) |
517 µs (517 µs–519 µs) |
198 µs |
620 µs |
4.26e7 |
1.88e9 (2 %) |
9.72e8 (0.69 %) |
2 MiB (52.4×) |
2.12 MiB |
3.6 % |
– |
2/8/40 |
1075–1075 |
1,000 |
eagle graph + compaction |
817 µs (813 µs–821 µs) |
620 µs |
966 µs |
2.70e7 |
1.19e9 (1.3 %) |
2.74e9 (2 %) |
2 MiB (52.4×) |
2.51 MiB |
2.3 % |
– |
100/2/18 |
1075–1075 |
1,000 |
eagle graph + compaction + reorder |
921 µs (919 µs–923 µs) |
640 µs |
1.13 ms |
2.40e7 |
1.05e9 (1.1 %) |
2.94e9 (2.1 %) |
2 MiB (52.4×) |
2.25 MiB |
2 % |
– |
100/3/26 |
1075–1075 |
1,000 |
eagle graph, K fused steps per launch |
345 µs (341 µs–349 µs) |
129 µs |
509 µs |
6.40e7 |
2.82e9 (3 %) |
8.98e8 (0.64 %) |
2 MiB (52.4×) |
2.12 MiB |
– |
– |
7/1/14 |
1075–1075 |
1,000 |
eagle graph, K fused steps per launch + compaction |
334 µs (332 µs–337 µs) |
157 µs |
496 µs |
6.61e7 |
2.91e9 (3.1 %) |
8.96e8 (0.64 %) |
2 MiB (52.4×) |
2.25 MiB |
– |
– |
7/2/18 |
1075–1075 |
1,000 |
eagle auto (eagle picks the launch mode) |
125 µs (124 µs–126 µs) |
– |
262 µs |
1.77e8 |
7.77e9 (8.2 %) |
7.34e9 (5.2 %) |
2 MiB (52.4×) |
2.12 MiB |
– |
– |
– |
1075–1075 |
1,000 |
hawk + eagle persistent |
203 µs (198 µs–207 µs) |
– |
304 µs |
1.09e8 |
4.78e9 (5 %) |
4.51e9 (3.2 %) |
2 MiB (52.4×) |
2.12 MiB |
– |
0.15821063701923077 |
– |
1075–1075 |
1,000 |
eagle.simulate (eagle picks the launch mode) |
130 µs (129 µs–130 µs) |
– |
268 µs |
1.70e8 |
7.49e9 (7.9 %) |
7.07e9 (5 %) |
2 MiB (52.4×) |
2.12 MiB |
– |
– |
– |
1075–1075 |
1,000 |
eager launch loop |
1.73 ms (1.72 ms–1.74 ms) |
370 µs |
1.82 ms |
1.28e7 |
5.61e8 (0.59 %) |
2.42e9 (1.7 %) |
2 MiB (52.4×) |
0.125 MiB |
1.1 % |
– |
100/100/502 |
1075–1075 |
1,000 |
CuPy masked |
33 ms (32.7 ms–33.3 ms) |
8.03 ms |
33.1 ms |
6.69e5 |
2.94e7 (0.031 %) |
2.73e9 (2 %) |
2 MiB (52.4×) |
3.62 MiB |
– |
– |
4404/100/4705 |
1075–1075 |
1,000 |
PyTorch masked |
22.2 ms (22 ms–22.4 ms) |
11.7 ms |
22.4 ms |
9.93e5 |
4.37e7 (0.046 %) |
4.06e9 (2.9 %) |
2 MiB (52.4×) |
0.125 MiB |
– |
– |
100/100/4505 |
1075–1075 |
1,000 |
PyTorch, CUDA graph of 16 steps |
12.3 ms (12.3 ms–12.3 ms) |
13.1 ms |
12.5 ms |
1.79e6 |
7.87e7 (0.083 %) |
8.22e9 (5.9 %) |
4 MiB (105×) |
1.38 MiB |
– |
– |
100/28/26 |
1075–1075 |
1,000 |
JAX jit + vmap(while_loop) |
2.13 ms (2 ms–2.26 ms) |
618 µs |
3.25 ms |
1.04e7 |
4.56e8 (0.48 %) |
4.15e9 (3 %) |
0 MiB (0×) |
0 MiB |
– |
– |
100/103/418 |
1075–1075 |
1,000 |
Warp per-thread kernel |
199 µs (196 µs–203 µs) |
143 µs |
450 µs |
1.11e8 |
4.89e9 (5.2 %) |
3.22e8 (0.23 %) |
32 MiB (839×) |
0 MiB |
– |
– |
1/0/4 |
1075–1075 |
1,000 |
CPU OpenMP |
198 µs (198 µs–199 µs) |
– |
209 µs |
1.11e8 |
4.89e9 (–) |
2.11e10 (–) |
– |
0 MiB |
– |
– |
– |
– |
10,000 |
eagle graph (device loop) |
1.21 ms (1.2 ms–1.21 ms) |
798 µs |
1.42 ms |
1.90e8 |
8.34e9 (8.8 %) |
4.20e9 (3 %) |
2 MiB (5.24×) |
2.75 MiB |
16 % |
– |
2/8/40 |
1075–1075 |
10,000 |
eagle graph + compaction |
1.16 ms (1.15 ms–1.16 ms) |
924 µs |
1.38 ms |
1.98e8 |
8.70e9 (9.2 %) |
1.99e10 (14 %) |
2 MiB (5.24×) |
2.88 MiB |
17 % |
– |
100/2/18 |
1075–1075 |
10,000 |
eagle graph + compaction + reorder |
1.43 ms (1.43 ms–1.44 ms) |
1.04 ms |
1.72 ms |
1.59e8 |
7.02e9 (7.4 %) |
1.94e10 (14 %) |
2 MiB (5.24×) |
3.12 MiB |
13 % |
– |
100/4/56 |
1075–1075 |
10,000 |
eagle graph, K fused steps per launch |
871 µs (868 µs–874 µs) |
716 µs |
1.06 ms |
2.62e8 |
1.15e10 (12 %) |
3.58e9 (2.6 %) |
2 MiB (5.24×) |
2.62 MiB |
– |
– |
7/1/14 |
1075–1075 |
10,000 |
eagle graph, K fused steps per launch + compaction |
766 µs (765 µs–769 µs) |
597 µs |
989 µs |
2.98e8 |
1.31e10 (14 %) |
3.95e9 (2.8 %) |
2 MiB (5.24×) |
2.75 MiB |
– |
– |
7/2/18 |
1075–1075 |
10,000 |
eagle auto (eagle picks the launch mode) |
340 µs (340 µs–341 µs) |
– |
548 µs |
6.72e8 |
2.96e10 (31 %) |
2.79e10 (20 %) |
2 MiB (5.24×) |
2.62 MiB |
– |
– |
– |
1075–1075 |
10,000 |
hawk + eagle persistent |
365 µs (362 µs–372 µs) |
– |
538 µs |
6.25e8 |
2.75e10 (29 %) |
2.59e10 (19 %) |
2 MiB (5.24×) |
2.75 MiB |
– |
0.5913234721260388 |
– |
1075–1075 |
10,000 |
eagle.simulate (eagle picks the launch mode) |
381 µs (380 µs–382 µs) |
– |
748 µs |
6.00e8 |
2.64e10 (28 %) |
2.49e10 (18 %) |
2 MiB (5.24×) |
2.75 MiB |
– |
– |
– |
1075–1075 |
10,000 |
eager launch loop |
2.49 ms (2.49 ms–2.5 ms) |
1.09 ms |
2.66 ms |
9.16e7 |
4.03e9 (4.3 %) |
1.69e10 (12 %) |
2 MiB (5.24×) |
0.766 MiB |
7.7 % |
– |
100/100/502 |
1075–1075 |
10,000 |
CuPy masked |
39.8 ms (36.2 ms–44.5 ms) |
10.4 ms |
40.1 ms |
5.74e6 |
2.53e8 (0.27 %) |
2.26e10 (16 %) |
2 MiB (5.24×) |
4 MiB |
– |
– |
4404/100/4705 |
1075–1075 |
10,000 |
PyTorch masked |
24.6 ms (24.5 ms–24.7 ms) |
13.2 ms |
24.9 ms |
9.28e6 |
4.08e8 (0.43 %) |
3.66e10 (26 %) |
2 MiB (5.24×) |
0.75 MiB |
– |
– |
100/100/4505 |
1075–1075 |
10,000 |
PyTorch, CUDA graph of 16 steps |
15.5 ms (15.5 ms–15.5 ms) |
16.4 ms |
15.8 ms |
1.47e7 |
6.47e8 (0.68 %) |
6.52e10 (47 %) |
4 MiB (10.5×) |
2 MiB |
– |
– |
100/28/26 |
1075–1075 |
10,000 |
JAX jit + vmap(while_loop) |
3.01 ms (2.96 ms–3.05 ms) |
– |
4.25 ms |
7.60e7 |
3.34e9 (3.5 %) |
2.93e10 (21 %) |
0 MiB (0×) |
0.25 MiB |
– |
– |
– |
1075–1075 |
10,000 |
Warp per-thread kernel |
719 µs (714 µs–722 µs) |
666 µs |
1.04 ms |
3.18e8 |
1.40e10 (15 %) |
8.90e8 (0.64 %) |
32 MiB (83.9×) |
0.625 MiB |
– |
– |
1/0/4 |
1075–1075 |
10,000 |
CPU OpenMP |
428 µs (426 µs–428 µs) |
– |
488 µs |
5.35e8 |
2.35e10 (–) |
9.86e10 (–) |
– |
1 MiB |
– |
– |
– |
– |
100,000 |
eagle graph (device loop) |
7.01 ms (7.01 ms–7.01 ms) |
6.65 ms |
7.96 ms |
3.26e8 |
1.44e10 (15 %) |
7.22e9 (5.2 %) |
6 MiB (1.57×) |
8.88 MiB |
28 % |
– |
2/8/40 |
1075–1075 |
100,000 |
eagle graph + compaction |
5.22 ms (5.22 ms–5.22 ms) |
5.25 ms |
6.21 ms |
4.38e8 |
1.93e10 (20 %) |
4.41e10 (31 %) |
8 MiB (2.1×) |
8.88 MiB |
37 % |
– |
100/2/18 |
1075–1075 |
100,000 |
eagle graph + compaction + reorder |
5.79 ms (5.78 ms–5.79 ms) |
5.85 ms |
6.84 ms |
3.95e8 |
1.74e10 (18 %) |
4.80e10 (34 %) |
10 MiB (2.62×) |
9.25 MiB |
33 % |
– |
100/4/58 |
1075–1075 |
100,000 |
eagle graph, K fused steps per launch |
6.63 ms (6.6 ms–6.68 ms) |
6.45 ms |
7.59 ms |
3.45e8 |
1.52e10 (16 %) |
4.70e9 (3.4 %) |
6 MiB (1.57×) |
8.12 MiB |
– |
– |
7/1/14 |
1075–1075 |
100,000 |
eagle graph, K fused steps per launch + compaction |
5.21 ms (5.19 ms–5.23 ms) |
5.08 ms |
6.21 ms |
4.40e8 |
1.93e10 (20 %) |
5.82e9 (4.2 %) |
8 MiB (2.1×) |
8.88 MiB |
– |
– |
7/2/18 |
1075–1075 |
100,000 |
eagle auto (eagle picks the launch mode) |
2.02 ms (2.02 ms–2.03 ms) |
– |
3.06 ms |
1.13e9 |
4.97e10 (52 %) |
4.69e10 (33 %) |
6 MiB (1.57×) |
8.25 MiB |
– |
– |
– |
1075–1075 |
100,000 |
hawk + eagle persistent |
2.05 ms (2.05 ms–2.05 ms) |
– |
2.98 ms |
1.12e9 |
4.92e10 (52 %) |
4.64e10 (33 %) |
6 MiB (1.57×) |
8.25 MiB |
– |
0.9063332438772265 |
– |
1075–1075 |
100,000 |
eagle.simulate (eagle picks the launch mode) |
2.02 ms (2.01 ms–2.02 ms) |
– |
3 ms |
1.13e9 |
4.99e10 (53 %) |
4.71e10 (34 %) |
6 MiB (1.57×) |
8.25 MiB |
– |
– |
– |
1075–1075 |
100,000 |
eager launch loop |
9.78 ms (9.78 ms–9.79 ms) |
8.4 ms |
10.7 ms |
2.34e8 |
1.03e10 (11 %) |
4.31e10 (31 %) |
6 MiB (1.57×) |
6.38 MiB |
20 % |
– |
100/100/502 |
1075–1075 |
100,000 |
CuPy masked |
77.1 ms (77.1 ms–77.2 ms) |
76.4 ms |
78.2 ms |
2.97e7 |
1.31e9 (1.4 %) |
1.17e11 (83 %) |
24 MiB (6.29×) |
7.5 MiB |
– |
– |
4404/100/4705 |
1075–1075 |
100,000 |
PyTorch masked |
67.4 ms (67.4 ms–67.4 ms) |
65.5 ms |
68.8 ms |
3.39e7 |
1.49e9 (1.6 %) |
1.34e11 (95 %) |
26 MiB (6.82×) |
6.88 MiB |
– |
– |
100/100/4505 |
1075–1075 |
100,000 |
PyTorch, CUDA graph of 16 steps |
72.2 ms (72.2 ms–72.3 ms) |
72.7 ms |
73.7 ms |
3.17e7 |
1.39e9 (1.5 %) |
1.40e11 (1e+02 %) |
26 MiB (6.82×) |
7.62 MiB |
– |
– |
100/28/26 |
1075–1075 |
100,000 |
JAX jit + vmap(while_loop) |
17.1 ms (17.1 ms–17.2 ms) |
16.1 ms |
19.4 ms |
1.33e8 |
5.87e9 (6.2 %) |
5.15e10 (37 %) |
12 MiB (3.15×) |
10.3 MiB |
– |
– |
100/103/519 |
1075–1075 |
100,000 |
Warp per-thread kernel |
6.41 ms (6.39 ms–6.45 ms) |
6.41 ms |
7.81 ms |
3.57e8 |
1.57e10 (17 %) |
9.98e8 (0.71 %) |
32 MiB (8.39×) |
7.01 MiB |
– |
– |
1/0/4 |
1075–1075 |
100,000 |
CPU OpenMP |
3.01 ms (1.88 ms–5.27 ms) |
– |
3.68 ms |
7.60e8 |
3.34e10 (–) |
1.40e11 (–) |
– |
10.9 MiB |
– |
– |
– |
– |
1,000,000 |
eagle graph (device loop) |
64.4 ms (64.3 ms–64.4 ms) |
64.1 ms |
71.4 ms |
3.57e8 |
1.57e10 (17 %) |
7.88e9 (5.6 %) |
50 MiB (1.31×) |
70.7 MiB |
30 % |
– |
2/8/40 |
1075–1075 |
1,000,000 |
eagle graph + compaction |
45.8 ms (45.7 ms–45.8 ms) |
48.2 ms |
52.9 ms |
5.03e8 |
2.21e10 (23 %) |
5.05e10 (36 %) |
66 MiB (1.73×) |
70.6 MiB |
42 % |
– |
100/2/18 |
1075–1075 |
1,000,000 |
eagle graph + compaction + reorder |
45.5 ms (45.5 ms–45.5 ms) |
49.9 ms |
52.7 ms |
5.05e8 |
2.22e10 (23 %) |
6.13e10 (44 %) |
94 MiB (2.46×) |
70.7 MiB |
43 % |
– |
100/4/58 |
1075–1075 |
1,000,000 |
eagle graph, K fused steps per launch |
63.3 ms (63.3 ms–63.4 ms) |
63.2 ms |
70.4 ms |
3.63e8 |
1.60e10 (17 %) |
4.92e9 (3.5 %) |
50 MiB (1.31×) |
64.2 MiB |
– |
– |
7/1/14 |
1075–1075 |
1,000,000 |
eagle graph, K fused steps per launch + compaction |
48.8 ms (48.8 ms–48.8 ms) |
48.9 ms |
55.9 ms |
4.71e8 |
2.07e10 (22 %) |
6.22e9 (4.4 %) |
66 MiB (1.73×) |
70.7 MiB |
– |
– |
7/2/18 |
1075–1075 |
1,000,000 |
eagle auto (eagle picks the launch mode) |
17.8 ms (17.8 ms–17.9 ms) |
– |
24.8 ms |
1.29e9 |
5.67e10 (60 %) |
5.34e10 (38 %) |
50 MiB (1.31×) |
64.1 MiB |
– |
– |
– |
1075–1075 |
1,000,000 |
hawk + eagle persistent |
17.9 ms (17.9 ms–17.9 ms) |
– |
24.8 ms |
1.29e9 |
5.65e10 (60 %) |
5.33e10 (38 %) |
50 MiB (1.31×) |
64.1 MiB |
– |
0.9837367603842367 |
– |
1075–1075 |
1,000,000 |
eagle.simulate (eagle picks the launch mode) |
17.9 ms (17.9 ms–17.9 ms) |
– |
24.8 ms |
1.29e9 |
5.67e10 (60 %) |
5.34e10 (38 %) |
50 MiB (1.31×) |
64 MiB |
– |
– |
– |
1075–1075 |
1,000,000 |
eager launch loop |
77.8 ms (77.8 ms–77.8 ms) |
77.8 ms |
84.8 ms |
2.96e8 |
1.30e10 (14 %) |
5.43e10 (39 %) |
50 MiB (1.31×) |
62.4 MiB |
25 % |
– |
100/100/502 |
1075–1075 |
1,000,000 |
CuPy masked |
738 ms (738 ms–739 ms) |
736 ms |
746 ms |
3.11e7 |
1.37e9 (1.4 %) |
1.22e11 (87 %) |
194 MiB (5.09×) |
69 MiB |
– |
– |
4404/100/4705 |
1075–1075 |
1,000,000 |
PyTorch masked |
740 ms (740 ms–740 ms) |
738 ms |
748 ms |
3.11e7 |
1.37e9 (1.4 %) |
1.22e11 (87 %) |
262 MiB (6.87×) |
68.7 MiB |
– |
– |
100/100/4605 |
1075–1075 |
1,000,000 |
PyTorch, CUDA graph of 16 steps |
827 ms (827 ms–827 ms) |
825 ms |
843 ms |
2.78e7 |
1.22e9 (1.3 %) |
1.23e11 (87 %) |
264 MiB (6.92×) |
68.8 MiB |
– |
– |
100/28/26 |
1075–1075 |
1,000,000 |
JAX jit + vmap(while_loop) |
138 ms (138 ms–138 ms) |
137 ms |
152 ms |
1.67e8 |
7.34e9 (7.7 %) |
6.40e10 (46 %) |
128 MiB (3.36×) |
101 MiB |
– |
– |
100/103/519 |
1075–1075 |
1,000,000 |
Warp per-thread kernel |
62.1 ms (62 ms–62.1 ms) |
62 ms |
71.7 ms |
3.71e8 |
1.63e10 (17 %) |
1.03e9 (0.74 %) |
64 MiB (1.68×) |
68.6 MiB |
– |
– |
1/0/4 |
1075–1075 |
1,000,000 |
CPU OpenMP |
15.7 ms (15.7 ms–15.7 ms) |
– |
22.9 ms |
1.46e9 |
6.44e10 (–) |
2.69e11 (–) |
– |
109 MiB |
– |
– |
– |
– |
Fastest arm by mean wall time (middle 10 of 12 runs):
N = 1,000: eagle auto (eagle picks the launch mode) (125 µs), then eagle.simulate (eagle picks the launch mode) (130 µs, 1.04× the time)
N = 10,000: eagle auto (eagle picks the launch mode) (340 µs), then hawk + eagle persistent (365 µs, 1.07× the time)
N = 100,000: eagle.simulate (eagle picks the launch mode) (2.02 ms), then eagle auto (eagle picks the launch mode) (2.02 ms, 1× the time)
N = 1,000,000: CPU OpenMP (15.7 ms), then eagle auto (eagle picks the launch mode) (17.8 ms, 1.14× the time)
Spread stop steps (log-uniform in [10, 1000]), up to 1000 steps#
N |
arm |
wall (range) |
kernel-only |
end-to-end |
sample·steps/s |
useful FLOP/s (% peak) |
issued B/s (% peak) |
device memory (× minimum) |
host memory |
% issue peak |
lane util |
launches (kernel/memcpy/API) |
SM clock (MHz) |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
1,000 |
eagle graph (device loop) |
1.62 ms (1.62 ms–1.63 ms) |
1.2 ms |
1.73 ms |
1.29e8 |
5.70e9 (6 %) |
1.15e9 (0.82 %) |
2 MiB (52.4×) |
2.26 MiB |
11 % |
– |
16/8/40 |
1075–1075 |
1,000 |
eagle graph + compaction |
6.53 ms (6.51 ms–6.55 ms) |
5.72 ms |
6.74 ms |
3.22e7 |
1.42e9 (1.5 %) |
2.81e9 (2 %) |
2 MiB (52.4×) |
2.38 MiB |
2.7 % |
– |
1000/2/18 |
1075–1075 |
1,000 |
eagle graph + compaction + reorder |
6.65 ms (6.65 ms–6.65 ms) |
5.86 ms |
6.86 ms |
3.16e7 |
1.39e9 (1.5 %) |
2.87e9 (2 %) |
2 MiB (52.4×) |
2.25 MiB |
2.7 % |
– |
1000/3/26 |
1075–1075 |
1,000 |
eagle graph, K fused steps per launch |
1.51 ms (1.51 ms–1.51 ms) |
1.17 ms |
1.62 ms |
1.39e8 |
6.13e9 (6.5 %) |
1.74e9 (1.2 %) |
2 MiB (52.4×) |
2 MiB |
– |
– |
63/1/14 |
1075–1075 |
1,000 |
eagle graph, K fused steps per launch + compaction |
1.6 ms (1.6 ms–1.6 ms) |
1.26 ms |
1.76 ms |
1.31e8 |
5.77e9 (6.1 %) |
9.90e8 (0.71 %) |
2 MiB (52.4×) |
2.27 MiB |
– |
– |
63/2/18 |
1075–1075 |
1,000 |
eagle auto (eagle picks the launch mode) |
879 µs (878 µs–880 µs) |
– |
1.02 ms |
2.39e8 |
1.05e10 (11 %) |
9.60e9 (6.9 %) |
2 MiB (52.4×) |
2.12 MiB |
– |
– |
– |
1075–1075 |
1,000 |
hawk + eagle persistent |
1.4 ms (1.38 ms–1.41 ms) |
– |
1.49 ms |
1.50e8 |
6.62e9 (7 %) |
6.04e9 (4.3 %) |
2 MiB (52.4×) |
2.12 MiB |
– |
0.22181753019815417 |
– |
1075–1075 |
1,000 |
eagle.simulate (eagle picks the launch mode) |
884 µs (884 µs–884 µs) |
– |
1.02 ms |
2.38e8 |
1.05e10 (11 %) |
9.55e9 (6.8 %) |
2 MiB (52.4×) |
2.25 MiB |
– |
– |
– |
1075–1075 |
1,000 |
eager launch loop |
17.2 ms (17.1 ms–17.3 ms) |
3.67 ms |
17.3 ms |
1.22e7 |
5.38e8 (0.57 %) |
2.41e9 (1.7 %) |
2 MiB (52.4×) |
0.125 MiB |
1 % |
– |
1000/1000/5002 |
1075–1075 |
1,000 |
CuPy masked |
340 ms (337 ms–343 ms) |
80.1 ms |
340 ms |
6.18e5 |
2.72e7 (0.029 %) |
2.65e9 (1.9 %) |
2 MiB (52.4×) |
3.76 MiB |
– |
– |
44004/1000/47005 |
1075–1075 |
1,000 |
PyTorch masked |
222 ms (222 ms–223 ms) |
117 ms |
223 ms |
9.45e5 |
4.16e7 (0.044 %) |
4.05e9 (2.9 %) |
2 MiB (52.4×) |
0.125 MiB |
– |
– |
1000/1000/45005 |
1075–1075 |
1,000 |
PyTorch, CUDA graph of 16 steps |
111 ms (111 ms–111 ms) |
118 ms |
111 ms |
1.89e6 |
8.33e7 (0.088 %) |
8.21e9 (5.9 %) |
4 MiB (105×) |
1.38 MiB |
– |
– |
1000/252/194 |
1075–1075 |
1,000 |
JAX jit + vmap(while_loop) |
17.8 ms (17.8 ms–17.9 ms) |
6.12 ms |
19.1 ms |
1.18e7 |
5.19e8 (0.55 %) |
4.94e9 (3.5 %) |
0 MiB (0×) |
0 MiB |
– |
– |
1000/1003/4018 |
1075–1075 |
1,000 |
Warp per-thread kernel |
1.38 ms (1.37 ms–1.39 ms) |
1.34 ms |
1.62 ms |
1.52e8 |
6.70e9 (7.1 %) |
4.63e7 (0.033 %) |
32 MiB (839×) |
0 MiB |
– |
– |
1/0/4 |
1075–1075 |
1,000 |
CPU OpenMP |
437 µs (433 µs–440 µs) |
– |
448 µs |
4.81e8 |
2.12e10 (–) |
9.47e10 (–) |
– |
0.141 MiB |
– |
– |
– |
– |
10,000 |
eagle graph (device loop) |
7.61 ms (7.59 ms–7.62 ms) |
7.15 ms |
7.8 ms |
2.87e8 |
1.26e10 (13 %) |
2.39e9 (1.7 %) |
2 MiB (5.24×) |
2.62 MiB |
24 % |
– |
16/8/41 |
1075–1075 |
10,000 |
eagle graph + compaction |
9.25 ms (9.25 ms–9.26 ms) |
8.39 ms |
9.49 ms |
2.36e8 |
1.04e10 (11 %) |
2.05e10 (15 %) |
2 MiB (5.24×) |
2.75 MiB |
20 % |
– |
1000/2/18 |
1075–1075 |
10,000 |
eagle graph + compaction + reorder |
9.64 ms (9.58 ms–9.7 ms) |
8.33 ms |
9.95 ms |
2.27e8 |
9.97e9 (11 %) |
2.03e10 (15 %) |
2 MiB (5.24×) |
3 MiB |
19 % |
– |
1000/4/56 |
1075–1075 |
10,000 |
eagle graph, K fused steps per launch |
7.5 ms (7.49 ms–7.5 ms) |
7.16 ms |
7.69 ms |
2.91e8 |
1.28e10 (14 %) |
3.53e9 (2.5 %) |
2 MiB (5.24×) |
2.62 MiB |
– |
– |
63/1/14 |
1075–1075 |
10,000 |
eagle graph, K fused steps per launch + compaction |
3.39 ms (3.38 ms–3.39 ms) |
3.1 ms |
3.61 ms |
6.45e8 |
2.84e10 (30 %) |
4.81e9 (3.4 %) |
2 MiB (5.24×) |
2.75 MiB |
– |
– |
63/2/18 |
1075–1075 |
10,000 |
eagle auto (eagle picks the launch mode) |
2.54 ms (2.53 ms–2.55 ms) |
– |
2.76 ms |
8.59e8 |
3.78e10 (40 %) |
3.45e10 (25 %) |
2 MiB (5.24×) |
2.62 MiB |
– |
– |
– |
1075–1075 |
10,000 |
hawk + eagle persistent |
2.57 ms (2.56 ms–2.58 ms) |
– |
2.75 ms |
8.49e8 |
3.74e10 (39 %) |
3.41e10 (24 %) |
2 MiB (5.24×) |
2.62 MiB |
– |
0.652826923076923 |
– |
1075–1075 |
10,000 |
eagle.simulate (eagle picks the launch mode) |
2.54 ms (2.54 ms–2.54 ms) |
– |
2.76 ms |
8.59e8 |
3.78e10 (40 %) |
3.45e10 (25 %) |
2 MiB (5.24×) |
2.62 MiB |
– |
– |
– |
1075–1075 |
10,000 |
eager launch loop |
24.7 ms (24.6 ms–24.8 ms) |
11 ms |
24.9 ms |
8.84e7 |
3.89e9 (4.1 %) |
1.69e10 (12 %) |
2 MiB (5.24×) |
0.625 MiB |
7.4 % |
– |
1000/1000/5002 |
1075–1075 |
10,000 |
CuPy masked |
329 ms (328 ms–330 ms) |
104 ms |
329 ms |
6.63e6 |
2.92e8 (0.31 %) |
2.74e10 (20 %) |
2 MiB (5.24×) |
4 MiB |
– |
– |
44004/1000/47005 |
1075–1075 |
10,000 |
PyTorch masked |
229 ms (225 ms–232 ms) |
130 ms |
230 ms |
9.51e6 |
4.19e8 (0.44 %) |
3.93e10 (28 %) |
2 MiB (5.24×) |
0.75 MiB |
– |
– |
1000/1000/45005 |
1075–1075 |
10,000 |
PyTorch, CUDA graph of 16 steps |
140 ms (140 ms–140 ms) |
147 ms |
140 ms |
1.56e7 |
6.85e8 (0.72 %) |
6.51e10 (46 %) |
4 MiB (10.5×) |
1.75 MiB |
– |
– |
1000/252/194 |
1075–1075 |
10,000 |
JAX jit + vmap(while_loop) |
29.5 ms (28.8 ms–29.9 ms) |
19.2 ms |
31.1 ms |
7.39e7 |
3.25e9 (3.4 %) |
2.98e10 (21 %) |
0 MiB (0×) |
0 MiB |
– |
– |
1000/1003/5019 |
1075–1075 |
10,000 |
Warp per-thread kernel |
6.64 ms (6.6 ms–6.7 ms) |
6.56 ms |
6.97 ms |
3.29e8 |
1.45e10 (15 %) |
9.64e7 (0.069 %) |
32 MiB (83.9×) |
0.625 MiB |
– |
– |
1/0/4 |
1075–1075 |
10,000 |
CPU OpenMP |
1.78 ms (1.78 ms–1.79 ms) |
– |
1.84 ms |
1.23e9 |
5.39e10 (–) |
2.34e11 (–) |
– |
1.12 MiB |
– |
– |
– |
– |
100,000 |
eagle graph (device loop) |
66.6 ms (66.4 ms–66.7 ms) |
66.3 ms |
67.7 ms |
3.25e8 |
1.43e10 (15 %) |
2.74e9 (2 %) |
6 MiB (1.57×) |
8.88 MiB |
27 % |
– |
16/8/41 |
1075–1075 |
100,000 |
eagle graph + compaction |
45.7 ms (45.6 ms–45.7 ms) |
46.5 ms |
46.8 ms |
4.74e8 |
2.09e10 (22 %) |
4.12e10 (29 %) |
8 MiB (2.1×) |
8.88 MiB |
40 % |
– |
1000/2/18 |
1075–1075 |
100,000 |
eagle graph + compaction + reorder |
34.1 ms (34.1 ms–34.2 ms) |
36.3 ms |
35.3 ms |
6.35e8 |
2.79e10 (29 %) |
5.70e10 (41 %) |
10 MiB (2.62×) |
9.01 MiB |
53 % |
– |
1000/4/58 |
1075–1075 |
100,000 |
eagle graph, K fused steps per launch |
64.8 ms (64.7 ms–64.8 ms) |
64.5 ms |
65.8 ms |
3.34e8 |
1.47e10 (16 %) |
4.08e9 (2.9 %) |
6 MiB (1.57×) |
8.12 MiB |
– |
– |
63/1/14 |
1075–1075 |
100,000 |
eagle graph, K fused steps per launch + compaction |
20.5 ms (20.5 ms–20.5 ms) |
20.4 ms |
21.5 ms |
1.06e9 |
4.65e10 (49 %) |
7.91e9 (5.6 %) |
8 MiB (2.1×) |
9 MiB |
– |
– |
63/2/18 |
1075–1075 |
100,000 |
eagle auto (eagle picks the launch mode) |
16.8 ms (16.8 ms–16.9 ms) |
– |
17.9 ms |
1.29e9 |
5.66e10 (60 %) |
5.16e10 (37 %) |
6 MiB (1.57×) |
8.12 MiB |
– |
– |
– |
1075–1075 |
100,000 |
hawk + eagle persistent |
16.9 ms (16.9 ms–16.9 ms) |
– |
17.8 ms |
1.28e9 |
5.64e10 (60 %) |
5.15e10 (37 %) |
6 MiB (1.57×) |
8.38 MiB |
– |
0.9260717314861655 |
– |
1075–1075 |
100,000 |
eagle.simulate (eagle picks the launch mode) |
16.8 ms (16.8 ms–16.8 ms) |
– |
17.8 ms |
1.29e9 |
5.66e10 (60 %) |
5.17e10 (37 %) |
6 MiB (1.57×) |
8.12 MiB |
– |
– |
– |
1075–1075 |
100,000 |
eager launch loop |
97.8 ms (97.7 ms–97.8 ms) |
83.8 ms |
98.8 ms |
2.22e8 |
9.75e9 (10 %) |
4.26e10 (30 %) |
6 MiB (1.57×) |
6.25 MiB |
19 % |
– |
1000/1000/5002 |
1075–1075 |
100,000 |
CuPy masked |
768 ms (768 ms–768 ms) |
763 ms |
769 ms |
2.82e7 |
1.24e9 (1.3 %) |
1.17e11 (84 %) |
24 MiB (6.29×) |
7.5 MiB |
– |
– |
44004/1000/47005 |
1075–1075 |
100,000 |
PyTorch masked |
674 ms (674 ms–674 ms) |
655 ms |
676 ms |
3.21e7 |
1.41e9 (1.5 %) |
1.34e11 (95 %) |
26 MiB (6.82×) |
6.88 MiB |
– |
– |
1000/1000/45005 |
1075–1075 |
100,000 |
PyTorch, CUDA graph of 16 steps |
650 ms (650 ms–650 ms) |
654 ms |
651 ms |
3.33e7 |
1.47e9 (1.5 %) |
1.40e11 (1e+02 %) |
26 MiB (6.82×) |
7.5 MiB |
– |
– |
1000/252/194 |
1075–1075 |
100,000 |
JAX jit + vmap(while_loop) |
168 ms (168 ms–168 ms) |
160 ms |
171 ms |
1.29e8 |
5.68e9 (6 %) |
5.25e10 (37 %) |
12 MiB (3.15×) |
5.38 MiB |
– |
– |
1000/1003/5019 |
1075–1075 |
100,000 |
Warp per-thread kernel |
64 ms (63.5 ms–64.2 ms) |
63.3 ms |
65.5 ms |
3.39e8 |
1.49e10 (16 %) |
1.00e8 (0.071 %) |
32 MiB (8.39×) |
6.88 MiB |
– |
– |
1/0/4 |
1075–1075 |
100,000 |
CPU OpenMP |
15 ms (14.9 ms–15 ms) |
– |
15.7 ms |
1.44e9 |
6.36e10 (–) |
2.78e11 (–) |
– |
11 MiB |
– |
– |
– |
– |
1,000,000 |
eagle graph (device loop) |
625 ms (625 ms–626 ms) |
625 ms |
633 ms |
3.46e8 |
1.52e10 (16 %) |
2.92e9 (2.1 %) |
50 MiB (1.31×) |
70.7 MiB |
29 % |
– |
16/8/40 |
1075–1075 |
1,000,000 |
eagle graph + compaction |
415 ms (415 ms–415 ms) |
439 ms |
423 ms |
5.21e8 |
2.29e10 (24 %) |
4.53e10 (32 %) |
66 MiB (1.73×) |
70.7 MiB |
44 % |
– |
1000/2/18 |
1075–1075 |
1,000,000 |
eagle graph + compaction + reorder |
279 ms (279 ms–279 ms) |
311 ms |
287 ms |
7.74e8 |
3.41e10 (36 %) |
6.96e10 (50 %) |
94 MiB (2.46×) |
70.7 MiB |
65 % |
– |
1000/4/58 |
1075–1075 |
1,000,000 |
eagle graph, K fused steps per launch |
627 ms (627 ms–627 ms) |
628 ms |
634 ms |
3.45e8 |
1.52e10 (16 %) |
4.21e9 (3 %) |
50 MiB (1.31×) |
64 MiB |
– |
– |
63/1/14 |
1075–1075 |
1,000,000 |
eagle graph, K fused steps per launch + compaction |
185 ms (185 ms–185 ms) |
188 ms |
192 ms |
1.17e9 |
5.14e10 (54 %) |
8.73e9 (6.2 %) |
66 MiB (1.73×) |
70.7 MiB |
– |
– |
63/2/18 |
1075–1075 |
1,000,000 |
eagle auto (eagle picks the launch mode) |
156 ms (156 ms–156 ms) |
– |
163 ms |
1.39e9 |
6.12e10 (65 %) |
5.58e10 (40 %) |
50 MiB (1.31×) |
64.1 MiB |
– |
– |
– |
1075–1075 |
1,000,000 |
hawk + eagle persistent |
156 ms (156 ms–156 ms) |
– |
163 ms |
1.39e9 |
6.11e10 (65 %) |
5.58e10 (40 %) |
50 MiB (1.31×) |
64.1 MiB |
– |
0.9920576938021418 |
– |
1075–1075 |
1,000,000 |
eagle.simulate (eagle picks the launch mode) |
156 ms (156 ms–156 ms) |
– |
163 ms |
1.39e9 |
6.11e10 (65 %) |
5.58e10 (40 %) |
50 MiB (1.31×) |
64 MiB |
– |
– |
– |
1075–1075 |
1,000,000 |
eager launch loop |
774 ms (773 ms–774 ms) |
773 ms |
781 ms |
2.80e8 |
1.23e10 (13 %) |
5.38e10 (38 %) |
50 MiB (1.31×) |
62.1 MiB |
24 % |
– |
1000/1000/5002 |
1075–1075 |
1,000,000 |
CuPy masked |
7.38 s (7.38 s–7.38 s) |
7.35 s |
7.39 s |
2.93e7 |
1.29e9 (1.4 %) |
1.22e11 (87 %) |
194 MiB (5.09×) |
68.9 MiB |
– |
– |
44004/1000/47005 |
1075–1075 |
1,000,000 |
PyTorch masked |
7.4 s (7.4 s–7.4 s) |
7.37 s |
7.41 s |
2.92e7 |
1.29e9 (1.4 %) |
1.22e11 (87 %) |
262 MiB (6.87×) |
68.7 MiB |
– |
– |
1000/1000/46005 |
1075–1075 |
1,000,000 |
PyTorch, CUDA graph of 16 steps |
7.44 s (7.44 s–7.44 s) |
7.42 s |
7.45 s |
2.91e7 |
1.28e9 (1.4 %) |
1.23e11 (87 %) |
264 MiB (6.92×) |
68.8 MiB |
– |
– |
1000/252/194 |
1075–1075 |
1,000,000 |
JAX jit + vmap(while_loop) |
1.36 s (1.36 s–1.36 s) |
1.36 s |
1.37 s |
1.59e8 |
7.00e9 (7.4 %) |
6.48e10 (46 %) |
128 MiB (3.36×) |
101 MiB |
– |
– |
1000/1003/5019 |
1075–1075 |
1,000,000 |
Warp per-thread kernel |
609 ms (608 ms–609 ms) |
615 ms |
624 ms |
3.55e8 |
1.56e10 (17 %) |
1.05e8 (0.075 %) |
64 MiB (1.68×) |
68.7 MiB |
– |
– |
1/0/4 |
1075–1075 |
1,000,000 |
CPU OpenMP |
149 ms (148 ms–150 ms) |
– |
156 ms |
1.45e9 |
6.39e10 (–) |
2.80e11 (–) |
– |
109 MiB |
– |
– |
– |
– |
Fastest arm by mean wall time (middle 10 of 12 runs):
N = 1,000: CPU OpenMP (437 µs), then eagle auto (eagle picks the launch mode) (879 µs, 2.01× the time)
N = 10,000: CPU OpenMP (1.78 ms), then eagle auto (eagle picks the launch mode) (2.54 ms, 1.43× the time)
N = 100,000: CPU OpenMP (15 ms), then eagle.simulate (eagle picks the launch mode) (16.8 ms, 1.12× the time)
N = 1,000,000: CPU OpenMP (149 ms), then eagle auto (eagle picks the launch mode) (156 ms, 1.04× the time)
Uniform stop step: every sample runs 1000 steps#
N |
arm |
wall (range) |
kernel-only |
end-to-end |
sample·steps/s |
useful FLOP/s (% peak) |
issued B/s (% peak) |
device memory (× minimum) |
host memory |
% issue peak |
lane util |
launches (kernel/memcpy/API) |
SM clock (MHz) |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
1,000 |
eagle graph (device loop) |
1.32 ms (1.31 ms–1.32 ms) |
967 µs |
1.43 ms |
7.60e8 |
3.35e10 (35 %) |
9.44e8 (0.67 %) |
2 MiB (52.4×) |
2.12 MiB |
64 % |
– |
16/8/40 |
1075–1075 |
1,000 |
eagle graph + compaction |
5.01 ms (5.01 ms–5.01 ms) |
4.69 ms |
5.18 ms |
2.00e8 |
8.78e9 (9.3 %) |
1.54e10 (11 %) |
2 MiB (52.4×) |
2.25 MiB |
17 % |
– |
1000/2/18 |
1075–1075 |
1,000 |
eagle graph + compaction + reorder |
5.08 ms (5.08 ms–5.09 ms) |
4.69 ms |
5.29 ms |
1.97e8 |
8.66e9 (9.1 %) |
1.52e10 (11 %) |
2 MiB (52.4×) |
2.38 MiB |
17 % |
– |
1000/3/26 |
1075–1075 |
1,000 |
eagle graph, K fused steps per launch |
1.52 ms (1.52 ms–1.52 ms) |
1.18 ms |
1.63 ms |
6.58e8 |
2.90e10 (31 %) |
3.03e9 (2.2 %) |
2 MiB (52.4×) |
2 MiB |
– |
– |
63/1/14 |
1075–1075 |
1,000 |
eagle graph, K fused steps per launch + compaction |
1.38 ms (1.38 ms–1.38 ms) |
1.12 ms |
1.53 ms |
7.25e8 |
3.19e10 (34 %) |
3.57e9 (2.5 %) |
2 MiB (52.4×) |
2.12 MiB |
– |
– |
63/2/18 |
1075–1075 |
1,000 |
eagle auto (eagle picks the launch mode) |
893 µs (892 µs–894 µs) |
– |
1.04 ms |
1.12e9 |
4.93e10 (52 %) |
4.48e10 (32 %) |
2 MiB (52.4×) |
2.12 MiB |
– |
– |
– |
1075–1075 |
1,000 |
hawk + eagle persistent |
1.52 ms (1.52 ms–1.52 ms) |
– |
1.63 ms |
6.59e8 |
2.90e10 (31 %) |
2.64e10 (19 %) |
2 MiB (52.4×) |
2.12 MiB |
– |
0.9247750946969697 |
– |
1075–1075 |
1,000 |
eagle.simulate (eagle picks the launch mode) |
896 µs (894 µs–897 µs) |
– |
1.04 ms |
1.12e9 |
4.91e10 (52 %) |
4.47e10 (32 %) |
2 MiB (52.4×) |
2.12 MiB |
– |
– |
– |
1075–1075 |
1,000 |
eager launch loop |
17.2 ms (17.1 ms–17.3 ms) |
3.71 ms |
17.3 ms |
5.80e7 |
2.55e9 (2.7 %) |
4.24e9 (3 %) |
2 MiB (52.4×) |
0.258 MiB |
4.9 % |
– |
1000/1000/5002 |
1075–1075 |
1,000 |
CuPy masked |
331 ms (328 ms–333 ms) |
80.1 ms |
331 ms |
3.02e6 |
1.33e8 (0.14 %) |
2.72e9 (1.9 %) |
2 MiB (52.4×) |
3.62 MiB |
– |
– |
44004/1000/47005 |
1075–1075 |
1,000 |
PyTorch masked |
223 ms (222 ms–224 ms) |
117 ms |
224 ms |
4.48e6 |
1.97e8 (0.21 %) |
4.04e9 (2.9 %) |
2 MiB (52.4×) |
0.25 MiB |
– |
– |
1000/1000/45005 |
1075–1075 |
1,000 |
PyTorch, CUDA graph of 16 steps |
111 ms (111 ms–111 ms) |
118 ms |
111 ms |
9.02e6 |
3.97e8 (0.42 %) |
8.23e9 (5.9 %) |
4 MiB (105×) |
1.38 MiB |
– |
– |
1000/252/194 |
1075–1075 |
1,000 |
JAX jit + vmap(while_loop) |
18.1 ms (17.9 ms–18.5 ms) |
6.12 ms |
19.4 ms |
5.53e7 |
2.43e9 (2.6 %) |
4.87e9 (3.5 %) |
0 MiB (0×) |
0 MiB |
– |
– |
1000/1003/4018 |
1075–1075 |
1,000 |
Warp per-thread kernel |
1.49 ms (1.49 ms–1.49 ms) |
1.43 ms |
1.73 ms |
6.71e8 |
2.95e10 (31 %) |
4.29e7 (0.031 %) |
32 MiB (839×) |
0.133 MiB |
– |
– |
1/0/4 |
1075–1075 |
1,000 |
CPU OpenMP |
419 µs (418 µs–421 µs) |
– |
431 µs |
2.38e9 |
1.05e11 (–) |
1.74e11 (–) |
– |
0 MiB |
– |
– |
– |
– |
10,000 |
eagle graph (device loop) |
7.74 ms (7.73 ms–7.75 ms) |
7.36 ms |
7.94 ms |
1.29e9 |
5.69e10 (60 %) |
1.60e9 (1.1 %) |
2 MiB (5.24×) |
2.76 MiB |
1.1e+02 % |
– |
16/8/41 |
1075–1075 |
10,000 |
eagle graph + compaction |
12.8 ms (12.8 ms–12.8 ms) |
12.4 ms |
13 ms |
7.81e8 |
3.44e10 (36 %) |
6.04e10 (43 %) |
2 MiB (5.24×) |
2.75 MiB |
66 % |
– |
1000/2/18 |
1075–1075 |
10,000 |
eagle graph + compaction + reorder |
13.3 ms (13.3 ms–13.3 ms) |
12.5 ms |
13.6 ms |
7.52e8 |
3.31e10 (35 %) |
5.82e10 (41 %) |
2 MiB (5.24×) |
3 MiB |
63 % |
– |
1000/3/26 |
1075–1075 |
10,000 |
eagle graph, K fused steps per launch |
7.96 ms (7.96 ms–7.97 ms) |
7.63 ms |
8.15 ms |
1.26e9 |
5.52e10 (58 %) |
5.78e9 (4.1 %) |
2 MiB (5.24×) |
2.62 MiB |
– |
– |
63/1/14 |
1075–1075 |
10,000 |
eagle graph, K fused steps per launch + compaction |
7.89 ms (7.85 ms–7.92 ms) |
7.58 ms |
8.17 ms |
1.27e9 |
5.58e10 (59 %) |
6.24e9 (4.4 %) |
2 MiB (5.24×) |
2.62 MiB |
– |
– |
63/2/18 |
1075–1075 |
10,000 |
eagle auto (eagle picks the launch mode) |
7.22 ms (7.22 ms–7.22 ms) |
– |
7.44 ms |
1.39e9 |
6.10e10 (64 %) |
5.55e10 (40 %) |
2 MiB (5.24×) |
2.75 MiB |
– |
– |
– |
1075–1075 |
10,000 |
hawk + eagle persistent |
7.26 ms (7.25 ms–7.26 ms) |
– |
7.44 ms |
1.38e9 |
6.06e10 (64 %) |
5.52e10 (39 %) |
2 MiB (5.24×) |
2.62 MiB |
– |
0.9940579193811075 |
– |
1075–1075 |
10,000 |
eagle.simulate (eagle picks the launch mode) |
7.22 ms (7.22 ms–7.22 ms) |
– |
7.44 ms |
1.39e9 |
6.09e10 (64 %) |
5.54e10 (40 %) |
2 MiB (5.24×) |
2.75 MiB |
– |
– |
– |
1075–1075 |
10,000 |
eager launch loop |
25.3 ms (25.1 ms–25.6 ms) |
11.7 ms |
25.6 ms |
3.95e8 |
1.74e10 (18 %) |
2.89e10 (21 %) |
2 MiB (5.24×) |
0.625 MiB |
33 % |
– |
1000/1000/5002 |
1075–1075 |
10,000 |
CuPy masked |
339 ms (339 ms–339 ms) |
104 ms |
339 ms |
2.95e7 |
1.30e9 (1.4 %) |
2.66e10 (19 %) |
2 MiB (5.24×) |
4.12 MiB |
– |
– |
44004/1000/47005 |
1075–1075 |
10,000 |
PyTorch masked |
245 ms (242 ms–247 ms) |
130 ms |
246 ms |
4.08e7 |
1.80e9 (1.9 %) |
3.68e10 (26 %) |
2 MiB (5.24×) |
0.625 MiB |
– |
– |
1000/1000/45005 |
1075–1075 |
10,000 |
PyTorch, CUDA graph of 16 steps |
140 ms (140 ms–140 ms) |
147 ms |
140 ms |
7.14e7 |
3.14e9 (3.3 %) |
6.51e10 (46 %) |
4 MiB (10.5×) |
1.88 MiB |
– |
– |
1000/252/194 |
1075–1075 |
10,000 |
JAX jit + vmap(while_loop) |
28.9 ms (28.6 ms–29.3 ms) |
19.1 ms |
30.4 ms |
3.45e8 |
1.52e10 (16 %) |
3.04e10 (22 %) |
0 MiB (0×) |
0 MiB |
– |
– |
1000/1003/5019 |
1075–1075 |
10,000 |
Warp per-thread kernel |
7.22 ms (7.22 ms–7.23 ms) |
7.16 ms |
7.57 ms |
1.38e9 |
6.09e10 (64 %) |
8.86e7 (0.063 %) |
32 MiB (83.9×) |
0.625 MiB |
– |
– |
1/0/4 |
1075–1075 |
10,000 |
CPU OpenMP |
2.2 ms (2.19 ms–2.2 ms) |
– |
2.26 ms |
4.55e9 |
2.00e11 (–) |
3.32e11 (–) |
– |
1.25 MiB |
– |
– |
– |
– |
100,000 |
eagle graph (device loop) |
72.3 ms (72.3 ms–72.4 ms) |
72 ms |
73.4 ms |
1.38e9 |
6.08e10 (64 %) |
1.72e9 (1.2 %) |
6 MiB (1.57×) |
8.25 MiB |
1.2e+02 % |
– |
16/8/41 |
1075–1075 |
100,000 |
eagle graph + compaction |
95.1 ms (95 ms–95.1 ms) |
99.8 ms |
96.2 ms |
1.05e9 |
4.63e10 (49 %) |
8.13e10 (58 %) |
8 MiB (2.1×) |
8.5 MiB |
89 % |
– |
1000/2/18 |
1075–1075 |
100,000 |
eagle graph + compaction + reorder |
95.4 ms (95.3 ms–95.4 ms) |
100 ms |
96.6 ms |
1.05e9 |
4.61e10 (49 %) |
8.11e10 (58 %) |
10 MiB (2.62×) |
8.75 MiB |
88 % |
– |
1000/3/26 |
1075–1075 |
100,000 |
eagle graph, K fused steps per launch |
72.5 ms (72.5 ms–72.5 ms) |
72.4 ms |
73.6 ms |
1.38e9 |
6.07e10 (64 %) |
6.35e9 (4.5 %) |
6 MiB (1.57×) |
8.12 MiB |
– |
– |
63/1/14 |
1075–1075 |
100,000 |
eagle graph, K fused steps per launch + compaction |
72.6 ms (72.6 ms–72.7 ms) |
72.7 ms |
73.7 ms |
1.38e9 |
6.06e10 (64 %) |
6.77e9 (4.8 %) |
8 MiB (2.1×) |
8.25 MiB |
– |
– |
63/2/18 |
1075–1075 |
100,000 |
eagle auto (eagle picks the launch mode) |
70.8 ms (70.8 ms–70.8 ms) |
– |
71.8 ms |
1.41e9 |
6.22e10 (66 %) |
5.66e10 (40 %) |
6 MiB (1.57×) |
8.12 MiB |
– |
– |
– |
1075–1075 |
100,000 |
hawk + eagle persistent |
70.8 ms (70.8 ms–70.9 ms) |
– |
71.9 ms |
1.41e9 |
6.21e10 (66 %) |
5.65e10 (40 %) |
6 MiB (1.57×) |
8.25 MiB |
– |
0.99959312561415 |
– |
1075–1075 |
100,000 |
eagle.simulate (eagle picks the launch mode) |
70.8 ms (70.8 ms–70.8 ms) |
– |
71.9 ms |
1.41e9 |
6.21e10 (66 %) |
5.65e10 (40 %) |
6 MiB (1.57×) |
8.25 MiB |
– |
– |
– |
1075–1075 |
100,000 |
eager launch loop |
108 ms (108 ms–108 ms) |
95.6 ms |
109 ms |
9.24e8 |
4.06e10 (43 %) |
6.74e10 (48 %) |
6 MiB (1.57×) |
6.25 MiB |
78 % |
– |
1000/1000/5002 |
1075–1075 |
100,000 |
CuPy masked |
767 ms (767 ms–767 ms) |
764 ms |
768 ms |
1.30e8 |
5.73e9 (6.1 %) |
1.17e11 (84 %) |
24 MiB (6.29×) |
7.62 MiB |
– |
– |
44004/1000/47005 |
1075–1075 |
100,000 |
PyTorch masked |
673 ms (673 ms–674 ms) |
655 ms |
675 ms |
1.48e8 |
6.53e9 (6.9 %) |
1.34e11 (95 %) |
26 MiB (6.82×) |
6.25 MiB |
– |
– |
1000/1000/45005 |
1075–1075 |
100,000 |
PyTorch, CUDA graph of 16 steps |
649 ms (649 ms–649 ms) |
654 ms |
651 ms |
1.54e8 |
6.78e9 (7.2 %) |
1.40e11 (1e+02 %) |
26 MiB (6.82×) |
7.62 MiB |
– |
– |
1000/252/194 |
1075–1075 |
100,000 |
JAX jit + vmap(while_loop) |
168 ms (167 ms–168 ms) |
160 ms |
170 ms |
5.96e8 |
2.62e10 (28 %) |
5.25e10 (37 %) |
12 MiB (3.15×) |
5.75 MiB |
– |
– |
1000/1003/5019 |
1075–1075 |
100,000 |
Warp per-thread kernel |
71.6 ms (71.6 ms–71.6 ms) |
71.4 ms |
73.2 ms |
1.40e9 |
6.15e10 (65 %) |
8.94e7 (0.064 %) |
32 MiB (8.39×) |
6.88 MiB |
– |
– |
1/0/4 |
1075–1075 |
100,000 |
CPU OpenMP |
19.5 ms (19.4 ms–19.5 ms) |
– |
20.2 ms |
5.14e9 |
2.26e11 (–) |
3.75e11 (–) |
– |
10.9 MiB |
– |
– |
– |
– |
1,000,000 |
eagle graph (device loop) |
703 ms (703 ms–703 ms) |
703 ms |
710 ms |
1.42e9 |
6.26e10 (66 %) |
1.77e9 (1.3 %) |
50 MiB (1.31×) |
64.1 MiB |
1.2e+02 % |
– |
16/8/41 |
1075–1075 |
1,000,000 |
eagle graph + compaction |
862 ms (862 ms–862 ms) |
879 ms |
869 ms |
1.16e9 |
5.11e10 (54 %) |
8.97e10 (64 %) |
66 MiB (1.73×) |
64.2 MiB |
98 % |
– |
1000/2/18 |
1075–1075 |
1,000,000 |
eagle graph + compaction + reorder |
863 ms (863 ms–863 ms) |
879 ms |
871 ms |
1.16e9 |
5.10e10 (54 %) |
8.96e10 (64 %) |
94 MiB (2.46×) |
64.5 MiB |
98 % |
– |
1000/3/26 |
1075–1075 |
1,000,000 |
eagle graph, K fused steps per launch |
710 ms (710 ms–710 ms) |
710 ms |
717 ms |
1.41e9 |
6.20e10 (65 %) |
6.48e9 (4.6 %) |
50 MiB (1.31×) |
64.1 MiB |
– |
– |
63/1/14 |
1075–1075 |
1,000,000 |
eagle graph, K fused steps per launch + compaction |
710 ms (710 ms–710 ms) |
711 ms |
717 ms |
1.41e9 |
6.20e10 (65 %) |
6.93e9 (4.9 %) |
66 MiB (1.73×) |
64.2 MiB |
– |
– |
63/2/18 |
1075–1075 |
1,000,000 |
eagle auto (eagle picks the launch mode) |
699 ms (699 ms–699 ms) |
– |
706 ms |
1.43e9 |
6.30e10 (66 %) |
5.73e10 (41 %) |
50 MiB (1.31×) |
64.1 MiB |
– |
– |
– |
1075–1075 |
1,000,000 |
hawk + eagle persistent |
699 ms (699 ms–699 ms) |
– |
706 ms |
1.43e9 |
6.30e10 (66 %) |
5.73e10 (41 %) |
50 MiB (1.31×) |
64 MiB |
– |
0.999939075711995 |
– |
1075–1075 |
1,000,000 |
eagle.simulate (eagle picks the launch mode) |
699 ms (699 ms–699 ms) |
– |
706 ms |
1.43e9 |
6.30e10 (66 %) |
5.73e10 (41 %) |
50 MiB (1.31×) |
64 MiB |
– |
– |
– |
1075–1075 |
1,000,000 |
eager launch loop |
875 ms (875 ms–875 ms) |
875 ms |
882 ms |
1.14e9 |
5.03e10 (53 %) |
8.34e10 (60 %) |
50 MiB (1.31×) |
62 MiB |
96 % |
– |
1000/1000/5002 |
1075–1075 |
1,000,000 |
CuPy masked |
7.38 s (7.38 s–7.38 s) |
7.35 s |
7.39 s |
1.35e8 |
5.96e9 (6.3 %) |
1.22e11 (87 %) |
194 MiB (5.09×) |
64.9 MiB |
– |
– |
44004/1000/47005 |
1075–1075 |
1,000,000 |
PyTorch masked |
7.4 s (7.4 s–7.4 s) |
7.37 s |
7.41 s |
1.35e8 |
5.95e9 (6.3 %) |
1.22e11 (87 %) |
262 MiB (6.87×) |
62 MiB |
– |
– |
1000/1000/46005 |
1075–1075 |
1,000,000 |
PyTorch, CUDA graph of 16 steps |
7.44 s (7.44 s–7.44 s) |
7.42 s |
7.45 s |
1.34e8 |
5.92e9 (6.2 %) |
1.23e11 (87 %) |
264 MiB (6.92×) |
63.4 MiB |
– |
– |
1000/252/194 |
1075–1075 |
1,000,000 |
JAX jit + vmap(while_loop) |
1.36 s (1.36 s–1.36 s) |
1.36 s |
1.37 s |
7.36e8 |
3.24e10 (34 %) |
6.48e10 (46 %) |
128 MiB (3.36×) |
98 MiB |
– |
– |
1000/1003/5019 |
1075–1075 |
1,000,000 |
Warp per-thread kernel |
693 ms (692 ms–693 ms) |
692 ms |
704 ms |
1.44e9 |
6.35e10 (67 %) |
9.24e7 (0.066 %) |
64 MiB (1.68×) |
68.7 MiB |
– |
– |
1/0/4 |
1075–1075 |
1,000,000 |
CPU OpenMP |
191 ms (190 ms–191 ms) |
– |
198 ms |
5.25e9 |
2.31e11 (–) |
3.83e11 (–) |
– |
109 MiB |
– |
– |
– |
– |
Fastest arm by mean wall time (middle 10 of 12 runs):
N = 1,000: CPU OpenMP (419 µs), then eagle auto (eagle picks the launch mode) (893 µs, 2.13× the time)
N = 10,000: CPU OpenMP (2.2 ms), then eagle auto (eagle picks the launch mode) (7.22 ms, 3.29× the time)
N = 100,000: CPU OpenMP (19.5 ms), then eagle auto (eagle picks the launch mode) (70.8 ms, 3.64× the time)
N = 1,000,000: CPU OpenMP (191 ms), then Warp per-thread kernel (693 ms, 3.63× the time)
Compaction cadence sweep (isolates the active-set map)#
the active-set map alone. For a given K, matched_cadence (the plain per-sample step, no map) and compact (eagle_graph_compact’s active-set step) check the loop’s stop guard at the SAME cadence, once every K steps; the only structural difference between the two rows at a given K is whether the step kernel reads the active-set map. The main table’s eagle_graph row checks its guard every step, so its difference from eagle_graph_compact there mixes the cadence change with the map – this sweep does not.
Config: Spread stop steps (log-uniform in [10, 1000]), up to 1000 steps.
N |
K (steps between guard checks) |
matched cadence, no map (range) |
eagle graph + compaction (range) |
ratio (matched/compact, |
|---|---|---|---|---|
100,000 |
8 |
82 ms (81.9 ms–82 ms) |
49.2 ms (49.2 ms–49.2 ms) |
1.67× |
100,000 |
16 |
81.4 ms (81.4 ms–81.4 ms) |
45.6 ms (45.6 ms–45.6 ms) |
1.78× |
100,000 |
32 |
81.7 ms (81.7 ms–81.7 ms) |
44 ms (44 ms–44 ms) |
1.86× |
1,000,000 |
8 |
757 ms (757 ms–757 ms) |
436 ms (436 ms–436 ms) |
1.74× |
1,000,000 |
16 |
757 ms (757 ms–757 ms) |
415 ms (415 ms–415 ms) |
1.82× |
1,000,000 |
32 |
758 ms (758 ms–758 ms) |
406 ms (406 ms–406 ms) |
1.87× |
FP32: eagle vs Warp (float32, S = 1000)#
Every arm runs the same oscillator in float32: eagle.simulate with scalar_type=”float32”, NVIDIA Warp’s per-thread kernel and JAX’s vmapped loop, all over the same float32 inputs. Each cell runs 2 untimed runs per arm first (eagle’s automatic arm picks its entry over them: the size rule, then a measured run of the other entry, keeping the faster; the entry it settled on is in the last column), then each arm in its own block (no interleaving): an untimed ramp of back-to-back runs for 0.5 s, then 5 timed runs. Before any timing counts, eagle and JAX reach step counts identical to Warp’s and states within absolute 0.0001 of Warp’s (the largest differences are in the JSON). Each time is the mean of the middle 3 of the 5 runs (highest and lowest dropped), with the range of those 3; eagle / Warp is the ratio of these means: below 1 eagle is faster. Before each arm’s block the GPU is cooled down: the block starts once the temperature is within 3 °C of the idle 63 °C and the GPU is idle, or after 20 s. Right after each timed run the SM clock is read (its range is shown per arm); cycles are the wall time times that clock, and eagle / Warp (cycles) is their ratio.
Every arm runs each sample to its own stop and no further than max_steps (S): eagle and Warp enforce the cap per sample, JAX in its while-loop condition, PyTorch and CuPy by the outer step count (a batch-level cap); an arm that runs whole blocks of steps per launch or replay stops at the first block boundary at or after it.
What each arm needed to get there: Warp, a hand-written step counter and cap inside the kernel; JAX, a cap in the while_loop condition; eagle, the kernel source unchanged, with its loop shape, unrolling and device entry chosen by hawk + eagle (per device, per batch).
Cap fixture (checked before any timing counts): N = 4,096, 8 samples with stop step 2S = 256, S = 128; every arm (3) reported k = min(stop step, S) per sample and the same finished count (4,088): passed.
distribution |
N |
eagle.simulate (range) |
Warp (range) |
JAX (range) |
eagle / Warp |
eagle / Warp (cycles) |
eagle entry |
|---|---|---|---|---|---|---|---|
uniform |
10,000 |
311 µs (310 µs–312 µs), 1075–1075 MHz |
353 µs (352 µs–355 µs), 1075–1075 MHz |
19.7 ms (19.7 ms–19.8 ms), 1075–1075 MHz |
0.879× |
0.879× |
fused_one |
uniform |
100,000 |
2.59 ms (2.59 ms–2.59 ms), 1075–1075 MHz |
2.66 ms (2.66 ms–2.66 ms), 1075–1075 MHz |
45.9 ms (45.6 ms–46 ms), 1075–1075 MHz |
0.972× |
0.972× |
fused_one |
uniform |
1,000,000 |
24.9 ms (24.9 ms–24.9 ms), 1075–1075 MHz |
25.5 ms (25.5 ms–25.6 ms), 1075–1075 MHz |
371 ms (371 ms–371 ms), 1075–1075 MHz |
0.976× |
0.976× |
fused_one |
spread |
10,000 |
270 µs (269 µs–272 µs), 1075–1075 MHz |
311 µs (311 µs–311 µs), 1075–1075 MHz |
20.3 ms (20.1 ms–20.7 ms), 1075–1075 MHz |
0.867× |
0.867× |
persist |
spread |
100,000 |
1.14 ms (1.13 ms–1.14 ms), 1075–1075 MHz |
2.38 ms (2.38 ms–2.39 ms), 1075–1075 MHz |
47.1 ms (46 ms–49.2 ms), 1075–1075 MHz |
0.477× |
0.477× |
persist |
spread |
1,000,000 |
9.17 ms (9.16 ms–9.17 ms), 1075–1075 MHz |
22.5 ms (22.5 ms–22.6 ms), 1075–1075 MHz |
371 ms (371 ms–371 ms), 1075–1075 MHz |
0.407× |
0.407× |
persist |
Fixed overhead per run (N = 64, S = 100)#
item 3: wall at a cell small enough that per-run overhead dominates (N=64, S=100), each arm’s own kernel-only vs wall where the main nsys pass covers this cell’s N (it does not by default, so kernel_only_s is usually null here)
arm |
wall (range) |
kernel-only |
|---|---|---|
eagle graph (device loop) |
409 µs (409 µs–410 µs) |
– |
eagle graph + compaction |
687 µs (686 µs–688 µs) |
– |
eagle graph + compaction + reorder |
747 µs (743 µs–750 µs) |
– |
eagle graph, K fused steps per launch |
264 µs (263 µs–265 µs) |
– |
eagle graph, K fused steps per launch + compaction |
286 µs (284 µs–287 µs) |
– |
eagle auto (eagle picks the launch mode) |
117 µs (117 µs–117 µs) |
– |
hawk + eagle persistent |
137 µs (136 µs–139 µs) |
– |
eagle.simulate (eagle picks the launch mode) |
117 µs (116 µs–118 µs) |
– |
eager launch loop |
1.77 ms (1.76 ms–1.77 ms) |
– |
CuPy masked |
32.3 ms (32.3 ms–32.4 ms) |
– |
PyTorch masked |
22 ms (21.9 ms–22 ms) |
– |
PyTorch, CUDA graph of 16 steps |
9.73 ms (9.73 ms–9.74 ms) |
– |
JAX jit + vmap(while_loop) |
1.63 ms (1.62 ms–1.63 ms) |
– |
Warp per-thread kernel |
135 µs (129 µs–146 µs) |
– |
CPU OpenMP |
167 µs (166 µs–168 µs) |
– |
Where each tool fits#
Each tool below is written the way its users write it, and each brings something of its own. Times are walls (mean of the middle 3 of 5 runs) at the largest N; ratios are against the eagle graph arm (below 1 = less time). The fastest arm per N is listed under each table above.
eagle graph (device loop). The step written once as a hawk kernel; the whole loop runs on the device as one CUDA graph with a device-side stop guard, so a run is one launch with no host round trip per step. On this card: N = 1,000,000, spread S=100: 64.4 ms; spread S=1000: 625 ms; uniform S=1000: 703 ms. Fastest in 0 of 12 cells.
eagle graph + compaction. Adds an active-set map: the launch covers only the samples still running, rebuilt every 16 steps; it is made for batches that thin out (the spread configurations). On this card: N = 1,000,000, spread S=100: 45.8 ms (0.711× the eagle graph time); spread S=1000: 415 ms (0.664× the eagle graph time); uniform S=1000: 862 ms (1.23× the eagle graph time). Fastest in 0 of 12 cells.
eagle graph + compaction + reorder. Adds a physical reorder when the live samples spread thin over their warps, so the mapped reads stay contiguous; the final restore to sample order is inside its wall. On this card: N = 1,000,000, spread S=100: 45.5 ms (0.707× the eagle graph time); spread S=1000: 279 ms (0.447× the eagle graph time); uniform S=1000: 863 ms (1.23× the eagle graph time). Fastest in 0 of 12 cells.
eagle graph, K fused steps per launch. The same step with one decorator argument, @hawk.kernel(steps=16): each launch runs that many steps per sample with the state kept in registers, as Warp’s per-thread loop does, and the launches, guard checks and plane reads shrink by the same factor; its results are bit-identical to the eagle graph arm’s in every cell. K is chosen by hand here; the eagle_graph_auto arm picks it at run time. On this card: N = 1,000,000, spread S=100: 63.3 ms (0.984× the eagle graph time); spread S=1000: 627 ms (1× the eagle graph time); uniform S=1000: 710 ms (1.01× the eagle graph time). Fastest in 0 of 12 cells.
eagle graph, K fused steps per launch + compaction. The K-steps kernel over the active-set map: the register-resident steps of the arm above, and the launches cover only the samples still running, so it fits batches that thin out. On this card: N = 1,000,000, spread S=100: 48.8 ms (0.758× the eagle graph time); spread S=1000: 185 ms (0.296× the eagle graph time); uniform S=1000: 710 ms (1.01× the eagle graph time). Fastest in 0 of 12 cells.
eagle auto (eagle picks the launch mode). The same step with @hawk.kernel(steps=”auto”) on the active-set kind, and nothing to tune: it picks the steps per launch itself after every launch, on the device, so a batch whose samples all run to the end gets Warp-like fusion of up to 64 steps per launch, and a batch that thins out gets shorter launches with the active set compacted after each one; no sample takes more steps than the cap. On this card: N = 1,000,000, spread S=100: 17.8 ms (0.277× the eagle graph time); spread S=1000: 156 ms (0.249× the eagle graph time); uniform S=1000: 699 ms (0.994× the eagle graph time). Fastest in 2 of 12 cells.
hawk + eagle persistent. The plain (not active-set) automatic kernel, forced through hawk’s persist entry: one launch, no graph, no map, no policy kernel, no cap at 64 – a grid of SMs blocks of 256 lanes each steal the next unfinished sample off a counter when their own finishes, so idle warps refill without ever rebuilding an index map; made for batches too large for one resident wave. On this card: N = 1,000,000, spread S=100: 17.9 ms (0.278× the eagle graph time); spread S=1000: 156 ms (0.249× the eagle graph time); uniform S=1000: 699 ms (0.994× the eagle graph time). Fastest in 0 of 12 cells.
eagle.simulate (eagle picks the launch mode). The same automatic kernel through eagle.simulate: the model, its state and its parameters by name, with no plan or plane binding to write; eagle deploys the kernel and builds the same device loop. On this card: N = 1,000,000, spread S=100: 17.9 ms (0.277× the eagle graph time); spread S=1000: 156 ms (0.249× the eagle graph time); uniform S=1000: 699 ms (0.994× the eagle graph time). Fastest in 1 of 12 cells.
eager launch loop. The same kernels launched from Python one step at a time; the host sees the state after every step. On this card: N = 1,000,000, spread S=100: 77.8 ms (1.21× the eagle graph time); spread S=1000: 774 ms (1.24× the eagle graph time); uniform S=1000: 875 ms (1.24× the eagle graph time). Fastest in 0 of 12 cells.
CuPy masked. NumPy-like array code on GPUs with no kernel to write; every sample is computed every step, so it fits dense batches where all samples run to the end (the uniform configuration). On this card: N = 1,000,000, spread S=100: 738 ms (11.5× the eagle graph time); spread S=1000: 7.38 s (11.8× the eagle graph time); uniform S=1000: 7.38 s (10.5× the eagle graph time). Fastest in 0 of 12 cells.
PyTorch masked. The same array code in PyTorch tensors; it fits where the computation sits next to a PyTorch model or needs autograd. On this card: N = 1,000,000, spread S=100: 740 ms (11.5× the eagle graph time); spread S=1000: 7.4 s (11.8× the eagle graph time); uniform S=1000: 7.4 s (10.5× the eagle graph time). Fastest in 0 of 12 cells.
PyTorch, CUDA graph of 16 steps. The same PyTorch code with its launch overhead removed by hand: 16 steps captured as one CUDA graph and replayed, one host check per replay, the masking unchanged. On this card: N = 1,000,000, spread S=100: 827 ms (12.8× the eagle graph time); spread S=1000: 7.44 s (11.9× the eagle graph time); uniform S=1000: 7.44 s (10.6× the eagle graph time). Fastest in 0 of 12 cells.
JAX jit + vmap(while_loop). The whole loop compiled once, a per-sample while loop batched with vmap; the same code runs on CPUs and GPUs and differentiates with jax.grad. On this card: N = 1,000,000, spread S=100: 138 ms (2.14× the eagle graph time); spread S=1000: 1.36 s (2.17× the eagle graph time); uniform S=1000: 1.36 s (1.93× the eagle graph time). Fastest in 0 of 12 cells.
Warp per-thread kernel. Explicit per-thread kernels written in Python, maintained by NVIDIA and differentiable through wp.Tape; each thread runs its own sample and stops at its own step, so finished samples cost nothing. On this card: N = 1,000,000, spread S=100: 62.1 ms (0.964× the eagle graph time); spread S=1000: 609 ms (0.973× the eagle graph time); uniform S=1000: 693 ms (0.985× the eagle graph time). Fastest in 0 of 12 cells.
CPU OpenMP. The same hawk kernel on the host through eagle’s OpenMP team, for machines without a GPU. On this card: N = 1,000,000, spread S=100: 15.7 ms (0.244× the eagle graph time); spread S=1000: 149 ms (0.238× the eagle graph time); uniform S=1000: 191 ms (0.271× the eagle graph time). Fastest in 9 of 12 cells.
Compile time#
Its own pass, never part of any wall. Each measurement is a fresh process: the imports and the GPU context start-up run untimed, then the timer spans the first call to a ready kernel (trace + code generation + compile + module load, or the cache lookup). Cold: every cache empty – hawk’s compile cache (HAWK_CACHE_DIR), CuPy’s kernel cache (CUPY_CACHE_DIR), the CUDA driver’s PTX cache (CUDA_CACHE_PATH), PyTorch Inductor’s and Triton’s caches (TORCHINDUCTOR_CACHE_DIR, TRITON_CACHE_DIR) and JAX’s persistent compilation cache (minimum compile time and entry size 0), each a fresh directory. Warm: a new process on the caches the first cold run populated. eagle arms: eagle.deploy of the arm’s hawk kernel, which returns once the device build is done (C++/CUDA emission, nvcc to PTX, driver load; eagle’s own small CuPy kernels compile beside it) while the host compiler builds the host side on another thread, and the arm’s set-up on a 64-sample batch (eagle’s graph built, one step run). CuPy: one masked step on 64 samples (every elementwise and reduction kernel compiled on first use). Warp: the kernel module’s code generation, compile and load on the first launch (64 samples), cached in WARP_CACHE_PATH. PyTorch torch.compile, where it runs: three steps at N (the first call traces and compiles, the second records the CUDA graph, the third replays it), per N (dynamic=False). JAX: lower + compile from abstract shapes, per N (XLA specialises on the shape). CPU OpenMP: the same eagle.deploy as the eagle graph arm and one host run, which waits for the host build. Medians of 3 cold and 3 warm processes.
arm |
compiled |
cold (median of 3) |
warm (median of 3) |
|---|---|---|---|
eagle graph (device loop) |
eagle.deploy of the hawk step kernel (the device build; the host side builds in the background), the device-loop graph |
1.18 s |
127 ms |
eagle graph + compaction |
eagle.deploy of the hawk active-set step kernel (the device build; the host side builds in the background), the compaction-loop graph |
1.05 s |
111 ms |
eagle graph + compaction + reorder |
eagle.deploy of the hawk active-set step kernel (the device build; the host side builds in the background), the compaction + reorder loop graph |
1.13 s |
107 ms |
eagle graph, K fused steps per launch |
eagle.deploy of the hawk K-steps kernel (the device build; the host side builds in the background), the device-loop graph |
1.06 s |
120 ms |
eagle graph, K fused steps per launch + compaction |
eagle.deploy of the hawk active-set K-steps kernel (the device build; the host side builds in the background), the compaction-loop graph |
1.05 s |
116 ms |
eagle auto (eagle picks the launch mode) |
eagle.deploy of the hawk active-set automatic kernel (the device build; the host side builds in the background), the policy kernel, the policy-loop graph |
1.11 s |
119 ms |
hawk + eagle persistent |
eagle.deploy of the hawk automatic kernel (the device build; the host side builds in the background), the persistent launch |
1.11 s |
115 ms |
eagle.simulate (eagle picks the launch mode) |
eagle.simulate’s deploy of the hawk active-set automatic kernel (the device build; the host side builds in the background), the policy kernel, the policy-loop graph |
1.12 s |
109 ms |
eager launch loop |
eagle.deploy of the hawk single step kernel (the device build; the host side builds in the background) |
1.04 s |
100 ms |
CuPy masked |
CuPy’s elementwise and reduction kernels of one masked step |
715 ms |
24.4 ms |
Warp per-thread kernel |
Warp kernel module (code generation, NVRTC compile, load) |
1.49 s |
56.8 ms |
JAX jit + vmap(while_loop) |
XLA executable for N = 1,000 |
246 ms |
45.5 ms |
JAX jit + vmap(while_loop) |
XLA executable for N = 10,000 |
234 ms |
46.1 ms |
JAX jit + vmap(while_loop) |
XLA executable for N = 100,000 |
273 ms |
45.7 ms |
JAX jit + vmap(while_loop) |
XLA executable for N = 1,000,000 |
257 ms |
44.9 ms |
CPU OpenMP |
eagle.deploy of the hawk step kernel and its first host run, which waits for the host build |
1.37 s |
125 ms |
No compile step: PyTorch masked, PyTorch, CUDA graph of 16 steps (PyTorch’s eager kernels ship prebuilt; the CUDA-graph capture is per N and recorded per configuration as torch_graph_capture_s).
Memory method#
A separate pass after the timing pass. Each (arm, N, configuration) runs in a fresh process: imports, the arm’s kernels built and a warm-up run on a 64-sample batch (every compile happens here; JAX also compiles for shape N from abstract shapes), every caching allocator trimmed (CuPy’s pools freed, torch.cuda.empty_cache(), Warp’s pool release threshold set to 0), then the baseline; then the batch’s inputs generated, the arm set up for N (eagle builds its graph, PyTorch captures its CUDA graph, whose private memory pool counts) and one full run with upload and download. Device memory = the memory the CUDA driver accounts to the process (NVML per-process used memory) after the run minus at the baseline. The caching allocators of CuPy, PyTorch and JAX are left at their defaults and are not trimmed after the baseline; each keeps the memory it reserved until trimmed, so the reading after the run is the allocator’s high-water reservation, the same measure for every arm. A disabled pool returns memory between operations, so its peak could only be caught by sampling, which misses short peaks; the held reservation needs no sampling. Caveats: the reading is what the process holds, so each allocator’s rounding and growth policy count (JAX’s allocator grows in regions that can exceed the request; PyTorch’s rounds blocks up; a captured CUDA graph keeps a private pool; Warp’s stream-ordered pool reserves in chunks of tens of MiB); the driver accounts in pages of about 2 MiB, so small-N rows read 0 or one page. The per-library counters (CuPy pool, torch.cuda.max_memory_reserved/allocated, JAX memory_stats) are recorded in the JSON as a cross-check. Host memory = peak resident set size above the baseline (/proc/self/clear_refs reset at the baseline, then VmHWM). Minimum = the state the workload needs, N × (2 state + 3 per-sample scalars: omega, zeta, stop step) × 8 bytes; the factor is device memory / minimum.
Device memory the process already held at the baseline (CUDA context, loaded kernels and libraries, the warm-up’s leftovers; at N = 1,000,000, first configuration): eagle graph (device loop) 56 MiB; eagle graph + compaction 56 MiB; eagle graph + compaction + reorder 56 MiB; eagle graph, K fused steps per launch 56 MiB; eagle graph, K fused steps per launch + compaction 56 MiB; eagle auto (eagle picks the launch mode) 56 MiB; hawk + eagle persistent 56 MiB; eagle.simulate (eagle picks the launch mode) 56 MiB; eager launch loop 56 MiB; CuPy masked 56 MiB; PyTorch masked 60 MiB; PyTorch, CUDA graph of 16 steps 86 MiB; JAX jit + vmap(while_loop) 64 MiB; Warp per-thread kernel 56 MiB.
GPU (Tesla T4)#
Device: Tesla T4 (compute capability 7.5, 40 SMs). Peak FP64: 2.54e11 FLOP/s. Peak DRAM bandwidth: 3.20e11 B/s. Mean of the middle 3 of 5 runs of each arm (the highest and the lowest dropped), range of those 3 in parentheses.
Each arm’s block starts once the GPU is within 3 °C of its idle 67 °C and reads idle, or after 20 s; the SM clock is read right after each timed run and its range is in the last column.
The FP64 peak above is at the device’s reported maximum clock (1590 MHz). During the timed repetitions the SM clock read 585 MHz; at that clock the FP64 peak is 9.36e10 FLOP/s, and every FP64 percentage below is 2.72× larger against it.
“% of issue peak” (the eagle arms only) counts DP instructions (DFMA/DMUL/DADD/DSETP/MUFU), not useful FLOPs: the step’s SASS has 29 of them, against the device’s DP issue-slot throughput at the observed clock – a kernel that is already at the issue pipe’s limit can still read well under 100% of FP64 peak, because useful FLOPs/step is smaller than issued DP instructions/step.
eagle_graph checks its stop guard every step; eagle_graph_compact checks it once every 16 steps (the cadence its compaction runs on) and reads an active-set map built from it. The two rows below differ in both the cadence and the map; the compaction cadence sweep further down holds the cadence fixed at a few values and isolates the map alone.
PyTorch, CUDA graph of 16 steps captures 16 masked steps once as a torch.cuda.CUDAGraph over static buffers and replays it, with one host check of “any sample running” per replay: it removes the per-kernel launch overhead and keeps the masking (every sample is computed every step, finished ones held by torch.where, so the steps a replay runs past the last stop change nothing). JAX’s while loop is one compiled executable with no host check per step; batched by vmap, it runs until the slowest sample stops. Warp runs one thread per sample, each stopping at its own step.
Every arm runs each sample to its own stop and no further than max_steps (S): eagle and Warp enforce the cap per sample, JAX in its while-loop condition, PyTorch and CuPy by the outer step count (a batch-level cap); an arm that runs whole blocks of steps per launch or replay stops at the first block boundary at or after it.
What each arm needed to get there: Warp, a hand-written step counter and cap inside the kernel; JAX, a cap in the while_loop condition; eagle, the kernel source unchanged, with its loop shape, unrolling and device entry chosen by hawk + eagle (per device, per batch).
Cap fixture (checked before any timing counts): N = 4,096, 8 samples with stop step 2S = 256, S = 128; every arm (16) reported k = min(stop step, S) per sample and the same finished count (4,088): passed.
arm |
loop structure |
|---|---|
eagle graph (device loop) |
per launch over the batch, inside one CUDA graph: the plain decorator’s kernel, which finishes its own samples and runs the steps per launch eagle’s policy picks (8 to 64), guard checked on the device every launch |
eagle graph + compaction |
as the graph arm, the step launched over the active-set map only (recomputed every 16 steps) |
eagle graph + compaction + reorder |
as the compaction arm, plus a physical reorder of the per-sample planes when the live samples spread thin |
eagle graph, K fused steps per launch |
per launch over the batch, inside one CUDA graph: the step kernel runs up to K = 16 steps per sample with the state in registers, each sample leaving at its own stop step; guard checked on the device every launch |
eagle graph, K fused steps per launch + compaction |
as the K-steps arm, the launch over the active-set map only (recomputed every 64 steps = 4 launches) |
eagle auto (eagle picks the launch mode) |
the active-set kernel under eagle’s automatic policy, which picks the launch mode from the batch: one launch running every sample to its end for a small batch, the persistent launch above that, and the compaction loop (active-set map, steps per launch picked on the device) only when the caller asks for compaction or reordering; the row records the mode taken |
hawk + eagle persistent |
ONE launch, no graph, no map, no policy kernel: a grid sized from the SM count and the kernel’s occupancy, each lane fetching base + atomicAdd(counter, 1) whenever its sample is done, running the whole budget’s worth of steps per sample it picks up |
eagle.simulate (eagle picks the launch mode) |
as the automatic arm: eagle.simulate takes the kernel, its state and its parameters by name and eagle picks the launch mode the same way |
eager launch loop |
per step over the batch: one kernel launch from Python, a host read of the finished count every step |
CuPy masked |
per step over the batch: whole-batch array expressions, finished samples held by cupy.where, a host check every step; the cap is batch-level (the outer step loop runs at most max_steps) |
PyTorch masked |
per step over the batch: whole-batch tensor expressions, finished samples held by torch.where, a host check every step; the cap is batch-level (the outer step loop runs at most max_steps) |
PyTorch, CUDA graph of 16 steps |
as PyTorch masked, 16 steps captured as one CUDA graph over static buffers, replayed with one host check per replay |
PyTorch torch.compile (reduce-overhead) |
as PyTorch masked, the step compiled by Inductor and replayed as a CUDA graph, a host check every step |
JAX jit + vmap(while_loop) |
per step over the batch, inside one compiled executable: the vmapped while loop runs until the slowest sample stops or the step cap is reached, finished samples held by a select |
Warp per-thread kernel |
per sample: one launch, each thread loops over its own sample’s steps and stops at its own stop step or the step cap, whichever comes first |
CPU OpenMP |
per step over the batch on the host: one host-team launch per step, the kernel marks and counts its finished samples |
Spread stop steps (log-uniform in [1, 100]), up to 100 steps#
N |
arm |
wall (range) |
kernel-only |
end-to-end |
sample·steps/s |
useful FLOP/s (% peak) |
issued B/s (% peak) |
device memory (× minimum) |
host memory |
% issue peak |
lane util |
launches (kernel/memcpy/API) |
SM clock (MHz) |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
1,000 |
eagle graph (device loop) |
836 µs (815 µs–849 µs) |
– |
1.05 ms |
2.64e7 |
1.16e9 (0.46 %) |
6.01e8 (0.19 %) |
2 MiB (52.4×) |
2.12 MiB |
1.6 % |
– |
– |
735–735 |
1,000 |
eagle graph + compaction |
750 µs (729 µs–771 µs) |
– |
1.03 ms |
2.94e7 |
1.29e9 (0.51 %) |
2.99e9 (0.93 %) |
2 MiB (52.4×) |
2.26 MiB |
1.8 % |
– |
– |
1590–1590 |
1,000 |
eagle graph + compaction + reorder |
858 µs (850 µs–866 µs) |
– |
1.23 ms |
2.57e7 |
1.13e9 (0.44 %) |
3.16e9 (0.99 %) |
2 MiB (52.4×) |
2.26 MiB |
1.6 % |
– |
– |
1590–1590 |
1,000 |
eagle graph, K fused steps per launch |
292 µs (290 µs–294 µs) |
– |
495 µs |
7.57e7 |
3.33e9 (1.3 %) |
1.06e9 (0.33 %) |
2 MiB (52.4×) |
2.09 MiB |
– |
– |
– |
1590–1590 |
1,000 |
eagle graph, K fused steps per launch + compaction |
349 µs (348 µs–351 µs) |
– |
616 µs |
6.32e7 |
2.78e9 (1.1 %) |
8.57e8 (0.27 %) |
2 MiB (52.4×) |
2.2 MiB |
– |
– |
– |
1590–1590 |
1,000 |
eagle auto (eagle picks the launch mode) |
127 µs (124 µs–128 µs) |
– |
382 µs |
1.74e8 |
7.65e9 (3 %) |
7.23e9 (2.3 %) |
2 MiB (52.4×) |
2.12 MiB |
– |
– |
– |
1005–1005 |
1,000 |
hawk + eagle persistent |
214 µs (212 µs–218 µs) |
– |
427 µs |
1.03e8 |
4.53e9 (1.8 %) |
4.28e9 (1.3 %) |
2 MiB (52.4×) |
2.12 MiB |
– |
0.0532833751619171 |
– |
1275–1275 |
1,000 |
eagle.simulate (eagle picks the launch mode) |
123 µs (122 µs–125 µs) |
– |
376 µs |
1.79e8 |
7.88e9 (3.1 %) |
7.44e9 (2.3 %) |
2 MiB (52.4×) |
2.12 MiB |
– |
– |
– |
1275–1275 |
1,000 |
eager launch loop |
2.9 ms (2.85 ms–2.96 ms) |
– |
3.08 ms |
7.61e6 |
3.35e8 (0.13 %) |
1.44e9 (0.45 %) |
2 MiB (52.4×) |
0.164 MiB |
0.47 % |
– |
– |
825–1275 |
1,000 |
CuPy masked |
58.3 ms (56.1 ms–61.1 ms) |
– |
58.5 ms |
3.79e5 |
1.67e7 (0.0065 %) |
1.55e9 (0.48 %) |
2 MiB (52.4×) |
3.98 MiB |
– |
– |
– |
645–660 |
1,000 |
PyTorch masked |
42.2 ms (41.4 ms–43.4 ms) |
– |
42.5 ms |
5.22e5 |
2.30e7 (0.009 %) |
2.13e9 (0.67 %) |
2 MiB (52.4×) |
0.125 MiB |
– |
– |
– |
1005–1005 |
1,000 |
PyTorch, CUDA graph of 16 steps |
10 ms (10 ms–10 ms) |
– |
10.3 ms |
2.20e6 |
9.69e7 (0.038 %) |
1.01e10 (3.2 %) |
4 MiB (105×) |
0.469 MiB |
– |
– |
– |
1590–1590 |
1,000 |
PyTorch torch.compile (reduce-overhead) |
13.9 ms (13.7 ms–14.1 ms) |
– |
14.1 ms |
1.59e6 |
7.00e7 (0.028 %) |
6.34e8 (0.2 %) |
2 MiB (52.4×) |
32.6 MiB |
– |
– |
– |
1590–1590 |
1,000 |
JAX jit + vmap(while_loop) |
3.04 ms (3.02 ms–3.09 ms) |
– |
5.21 ms |
7.24e6 |
3.19e8 (0.13 %) |
2.90e9 (0.91 %) |
0 MiB (0×) |
0.0781 MiB |
– |
– |
– |
855–885 |
1,000 |
Warp per-thread kernel |
387 µs (386 µs–388 µs) |
– |
765 µs |
5.70e7 |
2.51e9 (0.99 %) |
1.65e8 (0.052 %) |
32 MiB (839×) |
0.0977 MiB |
– |
– |
– |
1065–1065 |
1,000 |
CPU OpenMP |
359 µs (353 µs–366 µs) |
– |
374 µs |
6.14e7 |
2.70e9 (–) |
1.17e10 (–) |
– |
0.0898 MiB |
– |
– |
– |
– |
10,000 |
eagle graph (device loop) |
958 µs (931 µs–993 µs) |
– |
1.33 ms |
2.39e8 |
1.05e10 (4.1 %) |
5.28e9 (1.7 %) |
2 MiB (5.24×) |
2.67 MiB |
15 % |
– |
– |
1140–1140 |
10,000 |
eagle graph + compaction |
856 µs (852 µs–864 µs) |
– |
1.27 ms |
2.67e8 |
1.17e10 (4.6 %) |
2.69e10 (8.4 %) |
2 MiB (5.24×) |
2.81 MiB |
17 % |
– |
– |
1590–1590 |
10,000 |
eagle graph + compaction + reorder |
1.25 ms (1.21 ms–1.26 ms) |
– |
1.79 ms |
1.83e8 |
8.07e9 (3.2 %) |
2.23e10 (7 %) |
2 MiB (5.24×) |
3.1 MiB |
11 % |
– |
– |
1590–1590 |
10,000 |
eagle graph, K fused steps per launch |
474 µs (469 µs–479 µs) |
– |
814 µs |
4.82e8 |
2.12e10 (8.3 %) |
6.57e9 (2.1 %) |
2 MiB (5.24×) |
2.65 MiB |
– |
– |
– |
1590–1590 |
10,000 |
eagle graph, K fused steps per launch + compaction |
522 µs (507 µs–545 µs) |
– |
951 µs |
4.38e8 |
1.93e10 (7.6 %) |
5.80e9 (1.8 %) |
2 MiB (5.24×) |
2.75 MiB |
– |
– |
– |
1590–1590 |
10,000 |
eagle auto (eagle picks the launch mode) |
270 µs (267 µs–272 µs) |
– |
682 µs |
8.46e8 |
3.72e10 (15 %) |
3.51e10 (11 %) |
2 MiB (5.24×) |
2.66 MiB |
– |
– |
– |
1590–1590 |
10,000 |
hawk + eagle persistent |
326 µs (312 µs–345 µs) |
– |
705 µs |
7.01e8 |
3.08e10 (12 %) |
2.91e10 (9.1 %) |
2 MiB (5.24×) |
2.67 MiB |
– |
0.2101060762180118 |
– |
1590–1590 |
10,000 |
eagle.simulate (eagle picks the launch mode) |
280 µs (275 µs–283 µs) |
– |
699 µs |
8.18e8 |
3.60e10 (14 %) |
3.39e10 (11 %) |
2 MiB (5.24×) |
2.76 MiB |
– |
– |
– |
1590–1590 |
10,000 |
eager launch loop |
2.86 ms (2.8 ms–2.89 ms) |
– |
3.22 ms |
8.00e7 |
3.52e9 (1.4 %) |
1.47e10 (4.6 %) |
2 MiB (5.24×) |
0.68 MiB |
5 % |
– |
– |
1590–1590 |
10,000 |
CuPy masked |
57.6 ms (57.2 ms–57.9 ms) |
– |
57.9 ms |
3.97e6 |
1.75e8 (0.069 %) |
1.57e10 (4.9 %) |
2 MiB (5.24×) |
3.98 MiB |
– |
– |
– |
825–1590 |
10,000 |
PyTorch masked |
41.9 ms (41.7 ms–42.1 ms) |
– |
42.4 ms |
5.46e6 |
2.40e8 (0.094 %) |
2.15e10 (6.7 %) |
2 MiB (5.24×) |
0.734 MiB |
– |
– |
– |
1020–1020 |
10,000 |
PyTorch, CUDA graph of 16 steps |
9.85 ms (9.83 ms–9.87 ms) |
– |
10.3 ms |
2.32e7 |
1.02e9 (0.4 %) |
1.03e11 (32 %) |
4 MiB (10.5×) |
0.852 MiB |
– |
– |
– |
1590–1590 |
10,000 |
PyTorch torch.compile (reduce-overhead) |
17.6 ms (17.4 ms–17.8 ms) |
– |
18.1 ms |
1.30e7 |
5.70e8 (0.22 %) |
4.99e9 (1.6 %) |
2 MiB (5.24×) |
33.3 MiB |
– |
– |
– |
1590–1590 |
10,000 |
JAX jit + vmap(while_loop) |
3.25 ms (3.23 ms–3.27 ms) |
– |
5.32 ms |
7.04e7 |
3.10e9 (1.2 %) |
2.72e10 (8.5 %) |
0 MiB (0×) |
0.801 MiB |
– |
– |
– |
1095–1125 |
10,000 |
Warp per-thread kernel |
399 µs (395 µs–403 µs) |
– |
967 µs |
5.73e8 |
2.52e10 (9.9 %) |
1.60e9 (0.5 %) |
32 MiB (83.9×) |
0.746 MiB |
– |
– |
– |
1095–1095 |
10,000 |
CPU OpenMP |
1.22 ms (1.19 ms–1.26 ms) |
– |
1.3 ms |
1.87e8 |
8.23e9 (–) |
3.45e10 (–) |
– |
1.22 MiB |
– |
– |
– |
– |
100,000 |
eagle graph (device loop) |
2.69 ms (2.66 ms–2.71 ms) |
– |
4.43 ms |
8.52e8 |
3.75e10 (15 %) |
1.89e10 (5.9 %) |
6 MiB (1.57×) |
8.26 MiB |
53 % |
– |
– |
1590–1590 |
100,000 |
eagle graph + compaction |
2.36 ms (2.34 ms–2.38 ms) |
– |
4.21 ms |
9.71e8 |
4.27e10 (17 %) |
9.77e10 (31 %) |
8 MiB (2.1×) |
8.96 MiB |
60 % |
– |
– |
1590–1590 |
100,000 |
eagle graph + compaction + reorder |
2.74 ms (2.72 ms–2.75 ms) |
– |
4.74 ms |
8.34e8 |
3.67e10 (14 %) |
1.01e11 (32 %) |
10 MiB (2.62×) |
9.07 MiB |
52 % |
– |
– |
1590–1590 |
100,000 |
eagle graph, K fused steps per launch |
2.26 ms (2.25 ms–2.26 ms) |
– |
4.07 ms |
1.01e9 |
4.46e10 (18 %) |
1.38e10 (4.3 %) |
6 MiB (1.57×) |
8.25 MiB |
– |
– |
– |
1590–1590 |
100,000 |
eagle graph, K fused steps per launch + compaction |
1.98 ms (1.96 ms–1.99 ms) |
– |
3.85 ms |
1.16e9 |
5.09e10 (20 %) |
1.53e10 (4.8 %) |
8 MiB (2.1×) |
8.93 MiB |
– |
– |
– |
1590–1590 |
100,000 |
eagle auto (eagle picks the launch mode) |
911 µs (907 µs–914 µs) |
– |
2.76 ms |
2.51e9 |
1.11e11 (43 %) |
1.04e11 (33 %) |
6 MiB (1.57×) |
8.38 MiB |
– |
– |
– |
1380–1380 |
100,000 |
hawk + eagle persistent |
856 µs (851 µs–863 µs) |
– |
2.7 ms |
2.67e9 |
1.18e11 (46 %) |
1.11e11 (35 %) |
6 MiB (1.57×) |
8.27 MiB |
– |
0.7549041995981985 |
– |
1590–1590 |
100,000 |
eagle.simulate (eagle picks the launch mode) |
983 µs (908 µs–1.12 ms) |
– |
2.8 ms |
2.33e9 |
1.02e11 (40 %) |
9.65e10 (30 %) |
6 MiB (1.57×) |
8.32 MiB |
– |
– |
– |
1095–1425 |
100,000 |
eager launch loop |
5.33 ms (5.24 ms–5.49 ms) |
– |
7.11 ms |
4.29e8 |
1.89e10 (7.4 %) |
7.91e10 (25 %) |
6 MiB (1.57×) |
6.31 MiB |
27 % |
– |
– |
1590–1590 |
100,000 |
CuPy masked |
58 ms (57.2 ms–59.2 ms) |
– |
59.8 ms |
3.95e7 |
1.74e9 (0.68 %) |
1.55e11 (49 %) |
24 MiB (6.29×) |
7.9 MiB |
– |
– |
– |
1335–1425 |
100,000 |
PyTorch masked |
45 ms (44.2 ms–45.9 ms) |
– |
47.3 ms |
5.08e7 |
2.24e9 (0.88 %) |
2.00e11 (63 %) |
26 MiB (6.82×) |
6.91 MiB |
– |
– |
– |
1590–1590 |
100,000 |
PyTorch, CUDA graph of 16 steps |
29.3 ms (29.2 ms–29.3 ms) |
– |
31.6 ms |
7.82e7 |
3.44e9 (1.4 %) |
3.46e11 (1.1e+02 %) |
26 MiB (6.82×) |
6.93 MiB |
– |
– |
– |
1425–1530 |
100,000 |
PyTorch torch.compile (reduce-overhead) |
19.7 ms (19.6 ms–19.8 ms) |
– |
21.9 ms |
1.16e8 |
5.12e9 (2 %) |
4.47e10 (14 %) |
8 MiB (2.1×) |
37 MiB |
– |
– |
– |
600–600 |
100,000 |
JAX jit + vmap(while_loop) |
7.5 ms (7.37 ms–7.69 ms) |
– |
11.1 ms |
3.05e8 |
1.34e10 (5.3 %) |
1.18e11 (37 %) |
12 MiB (3.15×) |
6.91 MiB |
– |
– |
– |
1485–1590 |
100,000 |
Warp per-thread kernel |
2.13 ms (2.13 ms–2.14 ms) |
– |
4.38 ms |
1.07e9 |
4.72e10 (19 %) |
3.00e9 (0.94 %) |
32 MiB (8.39×) |
6.92 MiB |
– |
– |
– |
1590–1590 |
100,000 |
CPU OpenMP |
6.92 ms (6.68 ms–7.05 ms) |
– |
7.75 ms |
3.30e8 |
1.45e10 (–) |
6.09e10 (–) |
– |
11 MiB |
– |
– |
– |
– |
1,000,000 |
eagle graph (device loop) |
20 ms (19.9 ms–20.2 ms) |
– |
34.6 ms |
1.15e9 |
5.06e10 (20 %) |
2.53e10 (7.9 %) |
50 MiB (1.31×) |
64.1 MiB |
71 % |
– |
– |
1590–1590 |
1,000,000 |
eagle graph + compaction |
25.6 ms (25.6 ms–25.6 ms) |
– |
40.4 ms |
9.00e8 |
3.96e10 (16 %) |
9.04e10 (28 %) |
66 MiB (1.73×) |
70.8 MiB |
56 % |
– |
– |
1440–1545 |
1,000,000 |
eagle graph + compaction + reorder |
18.2 ms (18.2 ms–18.2 ms) |
– |
32.9 ms |
1.26e9 |
5.56e10 (22 %) |
1.53e11 (48 %) |
94 MiB (2.46×) |
70.8 MiB |
78 % |
– |
– |
1530–1575 |
1,000,000 |
eagle graph, K fused steps per launch |
19.3 ms (19.3 ms–19.3 ms) |
– |
33.8 ms |
1.19e9 |
5.25e10 (21 %) |
1.62e10 (5 %) |
50 MiB (1.31×) |
64.1 MiB |
– |
– |
– |
1590–1590 |
1,000,000 |
eagle graph, K fused steps per launch + compaction |
15.3 ms (15.2 ms–15.3 ms) |
– |
30.1 ms |
1.51e9 |
6.63e10 (26 %) |
1.99e10 (6.2 %) |
66 MiB (1.73×) |
70.8 MiB |
– |
– |
– |
1590–1590 |
1,000,000 |
eagle auto (eagle picks the launch mode) |
6.53 ms (6.52 ms–6.55 ms) |
– |
20.9 ms |
3.52e9 |
1.55e11 (61 %) |
1.46e11 (46 %) |
50 MiB (1.31×) |
64.1 MiB |
– |
– |
– |
1590–1590 |
1,000,000 |
hawk + eagle persistent |
6.61 ms (6.61 ms–6.61 ms) |
– |
21 ms |
3.48e9 |
1.53e11 (60 %) |
1.44e11 (45 %) |
50 MiB (1.31×) |
64.1 MiB |
– |
0.9389432931442274 |
– |
1410–1590 |
1,000,000 |
eagle.simulate (eagle picks the launch mode) |
6.51 ms (6.5 ms–6.52 ms) |
– |
21.3 ms |
3.53e9 |
1.55e11 (61 %) |
1.46e11 (46 %) |
50 MiB (1.31×) |
64.1 MiB |
– |
– |
– |
1590–1590 |
1,000,000 |
eager launch loop |
27.2 ms (27 ms–27.5 ms) |
– |
41.7 ms |
8.44e8 |
3.71e10 (15 %) |
1.55e11 (48 %) |
50 MiB (1.31×) |
62.1 MiB |
52 % |
– |
– |
1485–1560 |
1,000,000 |
CuPy masked |
369 ms (369 ms–370 ms) |
– |
384 ms |
6.23e7 |
2.74e9 (1.1 %) |
2.44e11 (76 %) |
194 MiB (5.09×) |
68.8 MiB |
– |
– |
– |
1245–1260 |
1,000,000 |
PyTorch masked |
386 ms (386 ms–386 ms) |
– |
411 ms |
5.96e7 |
2.62e9 (1 %) |
2.33e11 (73 %) |
262 MiB (6.87×) |
68.7 MiB |
– |
– |
– |
1200–1395 |
1,000,000 |
PyTorch, CUDA graph of 16 steps |
428 ms (428 ms–428 ms) |
– |
454 ms |
5.37e7 |
2.36e9 (0.93 %) |
2.37e11 (74 %) |
262 MiB (6.87×) |
68.8 MiB |
– |
– |
– |
1185–1380 |
1,000,000 |
PyTorch torch.compile (reduce-overhead) |
50.5 ms (50.4 ms–50.7 ms) |
– |
75.2 ms |
4.55e8 |
2.00e10 (7.9 %) |
1.74e11 (54 %) |
80 MiB (2.1×) |
77.6 MiB |
– |
– |
– |
1440–1530 |
1,000,000 |
JAX jit + vmap(while_loop) |
57.9 ms (57.7 ms–58.1 ms) |
– |
74.8 ms |
3.97e8 |
1.75e10 (6.9 %) |
1.52e11 (48 %) |
128 MiB (3.36×) |
101 MiB |
– |
– |
– |
1230–1365 |
1,000,000 |
Warp per-thread kernel |
19.7 ms (19.7 ms–19.7 ms) |
– |
50.9 ms |
1.17e9 |
5.14e10 (20 %) |
3.25e9 (1 %) |
64 MiB (1.68×) |
68.7 MiB |
– |
– |
– |
1590–1590 |
1,000,000 |
CPU OpenMP |
60 ms (59.1 ms–60.8 ms) |
– |
77.1 ms |
3.84e8 |
1.69e10 (–) |
7.04e10 (–) |
– |
109 MiB |
– |
– |
– |
– |
Fastest arm by mean wall time (middle 10 of 12 runs):
N = 1,000: eagle.simulate (eagle picks the launch mode) (123 µs), then eagle auto (eagle picks the launch mode) (127 µs, 1.03× the time)
N = 10,000: eagle auto (eagle picks the launch mode) (270 µs), then eagle.simulate (eagle picks the launch mode) (280 µs, 1.03× the time)
N = 100,000: hawk + eagle persistent (856 µs), then eagle auto (eagle picks the launch mode) (911 µs, 1.06× the time)
N = 1,000,000: eagle.simulate (eagle picks the launch mode) (6.51 ms), then eagle auto (eagle picks the launch mode) (6.53 ms, 1× the time)
Spread stop steps (log-uniform in [10, 1000]), up to 1000 steps#
N |
arm |
wall (range) |
kernel-only |
end-to-end |
sample·steps/s |
useful FLOP/s (% peak) |
issued B/s (% peak) |
device memory (× minimum) |
host memory |
% issue peak |
lane util |
launches (kernel/memcpy/API) |
SM clock (MHz) |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
1,000 |
eagle graph (device loop) |
1.28 ms (1.28 ms–1.28 ms) |
– |
1.48 ms |
1.64e8 |
7.22e9 (2.8 %) |
1.45e9 (0.45 %) |
2 MiB (52.4×) |
2.12 MiB |
10 % |
– |
– |
1590–1590 |
1,000 |
eagle graph + compaction |
4.64 ms (4.62 ms–4.65 ms) |
– |
4.94 ms |
4.53e7 |
1.99e9 (0.78 %) |
3.95e9 (1.2 %) |
2 MiB (52.4×) |
2.19 MiB |
2.8 % |
– |
– |
1590–1590 |
1,000 |
eagle graph + compaction + reorder |
4.88 ms (4.88 ms–4.88 ms) |
– |
5.25 ms |
4.31e7 |
1.90e9 (0.75 %) |
3.91e9 (1.2 %) |
2 MiB (52.4×) |
2.26 MiB |
2.7 % |
– |
– |
1590–1590 |
1,000 |
eagle graph, K fused steps per launch |
1.04 ms (1.03 ms–1.04 ms) |
– |
1.25 ms |
2.03e8 |
8.93e9 (3.5 %) |
2.53e9 (0.79 %) |
2 MiB (52.4×) |
2.1 MiB |
– |
– |
– |
1590–1590 |
1,000 |
eagle graph, K fused steps per launch + compaction |
1.23 ms (1.21 ms–1.26 ms) |
– |
1.52 ms |
1.71e8 |
7.52e9 (3 %) |
1.29e9 (0.4 %) |
2 MiB (52.4×) |
2.26 MiB |
– |
– |
– |
1590–1590 |
1,000 |
eagle auto (eagle picks the launch mode) |
502 µs (498 µs–509 µs) |
– |
757 µs |
4.19e8 |
1.84e10 (7.2 %) |
1.68e10 (5.3 %) |
2 MiB (52.4×) |
2.12 MiB |
– |
– |
– |
1590–1590 |
1,000 |
hawk + eagle persistent |
1.05 ms (1.01 ms–1.11 ms) |
– |
1.24 ms |
2.01e8 |
8.84e9 (3.5 %) |
8.07e9 (2.5 %) |
2 MiB (52.4×) |
2.18 MiB |
– |
0.17357174622982158 |
– |
1590–1590 |
1,000 |
eagle.simulate (eagle picks the launch mode) |
507 µs (503 µs–509 µs) |
– |
757 µs |
4.15e8 |
1.83e10 (7.2 %) |
1.67e10 (5.2 %) |
2 MiB (52.4×) |
2.19 MiB |
– |
– |
– |
1590–1590 |
1,000 |
eager launch loop |
30.4 ms (29.7 ms–31.5 ms) |
– |
30.7 ms |
6.90e6 |
3.04e8 (0.12 %) |
1.36e9 (0.42 %) |
2 MiB (52.4×) |
0.102 MiB |
0.43 % |
– |
– |
585–585 |
1,000 |
CuPy masked |
575 ms (573 ms–578 ms) |
– |
576 ms |
3.65e5 |
1.61e7 (0.0063 %) |
1.57e9 (0.49 %) |
2 MiB (52.4×) |
4.03 MiB |
– |
– |
– |
585–585 |
1,000 |
PyTorch masked |
407 ms (407 ms–408 ms) |
– |
408 ms |
5.16e5 |
2.27e7 (0.0089 %) |
2.21e9 (0.69 %) |
2 MiB (52.4×) |
0.121 MiB |
– |
– |
– |
1065–1065 |
1,000 |
PyTorch, CUDA graph of 16 steps |
89.1 ms (89.1 ms–89.2 ms) |
– |
89.5 ms |
2.36e6 |
1.04e8 (0.041 %) |
1.02e10 (3.2 %) |
4 MiB (105×) |
0.469 MiB |
– |
– |
– |
1590–1590 |
1,000 |
PyTorch torch.compile (reduce-overhead) |
142 ms (142 ms–142 ms) |
– |
142 ms |
1.48e6 |
6.52e7 (0.026 %) |
6.20e8 (0.19 %) |
2 MiB (52.4×) |
32.8 MiB |
– |
– |
– |
585–585 |
1,000 |
JAX jit + vmap(while_loop) |
25.6 ms (25.4 ms–25.7 ms) |
– |
27.9 ms |
8.21e6 |
3.61e8 (0.14 %) |
3.44e9 (1.1 %) |
0 MiB (0×) |
0.0742 MiB |
– |
– |
– |
1005–1005 |
1,000 |
Warp per-thread kernel |
2.08 ms (2.08 ms–2.08 ms) |
– |
2.55 ms |
1.01e8 |
4.44e9 (1.7 %) |
3.07e7 (0.0096 %) |
32 MiB (839×) |
0.0352 MiB |
– |
– |
– |
1590–1590 |
1,000 |
CPU OpenMP |
1.17 ms (1.15 ms–1.19 ms) |
– |
1.19 ms |
1.80e8 |
7.92e9 (–) |
3.55e10 (–) |
– |
0.219 MiB |
– |
– |
– |
– |
10,000 |
eagle graph (device loop) |
2.95 ms (2.95 ms–2.95 ms) |
– |
3.31 ms |
7.40e8 |
3.25e10 (13 %) |
6.17e9 (1.9 %) |
2 MiB (5.24×) |
2.73 MiB |
46 % |
– |
– |
1590–1590 |
10,000 |
eagle graph + compaction |
5.87 ms (5.86 ms–5.88 ms) |
– |
6.3 ms |
3.72e8 |
1.64e10 (6.4 %) |
3.23e10 (10 %) |
2 MiB (5.24×) |
2.83 MiB |
23 % |
– |
– |
1590–1590 |
10,000 |
eagle graph + compaction + reorder |
6.47 ms (6.46 ms–6.48 ms) |
– |
6.99 ms |
3.37e8 |
1.48e10 (5.8 %) |
3.03e10 (9.5 %) |
2 MiB (5.24×) |
3.04 MiB |
21 % |
– |
– |
1590–1590 |
10,000 |
eagle graph, K fused steps per launch |
2.72 ms (2.71 ms–2.73 ms) |
– |
3.07 ms |
8.02e8 |
3.53e10 (14 %) |
9.72e9 (3 %) |
2 MiB (5.24×) |
2.66 MiB |
– |
– |
– |
1590–1590 |
10,000 |
eagle graph, K fused steps per launch + compaction |
2.05 ms (2.04 ms–2.05 ms) |
– |
2.47 ms |
1.07e9 |
4.69e10 (18 %) |
7.95e9 (2.5 %) |
2 MiB (5.24×) |
2.73 MiB |
– |
– |
– |
1590–1590 |
10,000 |
eagle auto (eagle picks the launch mode) |
2.15 ms (2.14 ms–2.16 ms) |
– |
2.57 ms |
1.02e9 |
4.47e10 (18 %) |
4.08e10 (13 %) |
2 MiB (5.24×) |
2.68 MiB |
– |
– |
– |
1590–1590 |
10,000 |
hawk + eagle persistent |
2.16 ms (2.15 ms–2.19 ms) |
– |
2.52 ms |
1.01e9 |
4.44e10 (17 %) |
4.05e10 (13 %) |
2 MiB (5.24×) |
2.68 MiB |
– |
0.23828475965858043 |
– |
1590–1590 |
10,000 |
eagle.simulate (eagle picks the launch mode) |
2.15 ms (2.13 ms–2.16 ms) |
– |
2.57 ms |
1.02e9 |
4.47e10 (18 %) |
4.08e10 (13 %) |
2 MiB (5.24×) |
2.75 MiB |
– |
– |
– |
1590–1590 |
10,000 |
eager launch loop |
37.7 ms (37.5 ms–37.9 ms) |
– |
38 ms |
5.80e7 |
2.55e9 (1 %) |
1.11e10 (3.5 %) |
2 MiB (5.24×) |
0.68 MiB |
3.6 % |
– |
– |
585–585 |
10,000 |
CuPy masked |
570 ms (568 ms–571 ms) |
– |
570 ms |
3.83e6 |
1.68e8 (0.066 %) |
1.58e10 (4.9 %) |
2 MiB (5.24×) |
4.16 MiB |
– |
– |
– |
585–585 |
10,000 |
PyTorch masked |
402 ms (401 ms–403 ms) |
– |
403 ms |
5.43e6 |
2.39e8 (0.094 %) |
2.24e10 (7 %) |
2 MiB (5.24×) |
0.738 MiB |
– |
– |
– |
1065–1080 |
10,000 |
PyTorch, CUDA graph of 16 steps |
88.2 ms (88.1 ms–88.2 ms) |
– |
88.7 ms |
2.48e7 |
1.09e9 (0.43 %) |
1.03e11 (32 %) |
4 MiB (10.5×) |
0.855 MiB |
– |
– |
– |
1590–1590 |
10,000 |
PyTorch torch.compile (reduce-overhead) |
197 ms (196 ms–198 ms) |
– |
197 ms |
1.11e7 |
4.88e8 (0.19 %) |
4.47e9 (1.4 %) |
2 MiB (5.24×) |
33.4 MiB |
– |
– |
– |
1260–1260 |
10,000 |
JAX jit + vmap(while_loop) |
28.1 ms (28 ms–28.1 ms) |
– |
30.4 ms |
7.78e7 |
3.42e9 (1.3 %) |
3.14e10 (9.8 %) |
0 MiB (0×) |
0.773 MiB |
– |
– |
– |
1380–1380 |
10,000 |
Warp per-thread kernel |
2.24 ms (2.22 ms–2.28 ms) |
– |
2.85 ms |
9.74e8 |
4.29e10 (17 %) |
2.86e8 (0.089 %) |
32 MiB (83.9×) |
0.688 MiB |
– |
– |
– |
1590–1590 |
10,000 |
CPU OpenMP |
7.92 ms (7.69 ms–8.06 ms) |
– |
8 ms |
2.76e8 |
1.21e10 (–) |
5.27e10 (–) |
– |
1.2 MiB |
– |
– |
– |
– |
100,000 |
eagle graph (device loop) |
20.8 ms (20.7 ms–20.9 ms) |
– |
22.5 ms |
1.04e9 |
4.58e10 (18 %) |
8.78e9 (2.7 %) |
6 MiB (1.57×) |
8.26 MiB |
65 % |
– |
– |
1590–1590 |
100,000 |
eagle graph + compaction |
19.2 ms (19.1 ms–19.4 ms) |
– |
21.1 ms |
1.13e9 |
4.96e10 (19 %) |
9.80e10 (31 %) |
8 MiB (2.1×) |
8.96 MiB |
70 % |
– |
– |
1500–1575 |
100,000 |
eagle graph + compaction + reorder |
15.9 ms (15.9 ms–16.1 ms) |
– |
17.9 ms |
1.36e9 |
5.98e10 (23 %) |
1.22e11 (38 %) |
10 MiB (2.62×) |
9.01 MiB |
84 % |
– |
– |
1590–1590 |
100,000 |
eagle graph, K fused steps per launch |
20.7 ms (20.6 ms–20.7 ms) |
– |
22.5 ms |
1.05e9 |
4.61e10 (18 %) |
1.28e10 (4 %) |
6 MiB (1.57×) |
8.27 MiB |
– |
– |
– |
1590–1590 |
100,000 |
eagle graph, K fused steps per launch + compaction |
8.89 ms (8.84 ms–8.93 ms) |
– |
10.7 ms |
2.44e9 |
1.07e11 (42 %) |
1.82e10 (5.7 %) |
8 MiB (2.1×) |
9.09 MiB |
– |
– |
– |
1590–1590 |
100,000 |
eagle auto (eagle picks the launch mode) |
6.22 ms (6.19 ms–6.24 ms) |
– |
8.04 ms |
3.48e9 |
1.53e11 (60 %) |
1.40e11 (44 %) |
6 MiB (1.57×) |
8.28 MiB |
– |
– |
– |
1590–1590 |
100,000 |
hawk + eagle persistent |
6.31 ms (6.29 ms–6.33 ms) |
– |
8.09 ms |
3.43e9 |
1.51e11 (59 %) |
1.38e11 (43 %) |
6 MiB (1.57×) |
8.33 MiB |
– |
0.7890996166226245 |
– |
1590–1590 |
100,000 |
eagle.simulate (eagle picks the launch mode) |
6.17 ms (6.15 ms–6.19 ms) |
– |
7.99 ms |
3.51e9 |
1.54e11 (61 %) |
1.41e11 (44 %) |
6 MiB (1.57×) |
8.28 MiB |
– |
– |
– |
1590–1590 |
100,000 |
eager launch loop |
52.5 ms (52.4 ms–52.6 ms) |
– |
54.4 ms |
4.12e8 |
1.81e10 (7.1 %) |
7.93e10 (25 %) |
6 MiB (1.57×) |
6.33 MiB |
26 % |
– |
– |
1590–1590 |
100,000 |
CuPy masked |
570 ms (568 ms–571 ms) |
– |
572 ms |
3.80e7 |
1.67e9 (0.66 %) |
1.58e11 (49 %) |
24 MiB (6.29×) |
7.76 MiB |
– |
– |
– |
1590–1590 |
100,000 |
PyTorch masked |
444 ms (441 ms–447 ms) |
– |
446 ms |
4.88e7 |
2.15e9 (0.84 %) |
2.03e11 (63 %) |
26 MiB (6.82×) |
6.91 MiB |
– |
– |
– |
1575–1590 |
100,000 |
PyTorch, CUDA graph of 16 steps |
268 ms (268 ms–268 ms) |
– |
270 ms |
8.08e7 |
3.56e9 (1.4 %) |
3.40e11 (1.1e+02 %) |
26 MiB (6.82×) |
6.93 MiB |
– |
– |
– |
1380–1380 |
100,000 |
PyTorch torch.compile (reduce-overhead) |
188 ms (185 ms–189 ms) |
– |
190 ms |
1.15e8 |
5.08e9 (2 %) |
4.69e10 (15 %) |
8 MiB (2.1×) |
37.1 MiB |
– |
– |
– |
660–660 |
100,000 |
JAX jit + vmap(while_loop) |
69.5 ms (69 ms–69.9 ms) |
– |
72.9 ms |
3.12e8 |
1.37e10 (5.4 %) |
1.27e11 (40 %) |
12 MiB (3.15×) |
6.88 MiB |
– |
– |
– |
1560–1590 |
100,000 |
Warp per-thread kernel |
20.2 ms (20.2 ms–20.2 ms) |
– |
22.4 ms |
1.07e9 |
4.72e10 (19 %) |
3.17e8 (0.099 %) |
32 MiB (8.39×) |
6.93 MiB |
– |
– |
– |
1590–1590 |
100,000 |
CPU OpenMP |
55.6 ms (53.9 ms–56.9 ms) |
– |
56.5 ms |
3.89e8 |
1.71e10 (–) |
7.49e10 (–) |
– |
11 MiB |
– |
– |
– |
– |
1,000,000 |
eagle graph (device loop) |
189 ms (189 ms–189 ms) |
– |
203 ms |
1.15e9 |
5.04e10 (20 %) |
9.67e9 (3 %) |
50 MiB (1.31×) |
64.1 MiB |
71 % |
– |
– |
1590–1590 |
1,000,000 |
eagle graph + compaction |
253 ms (253 ms–253 ms) |
– |
268 ms |
8.55e8 |
3.76e10 (15 %) |
7.43e10 (23 %) |
66 MiB (1.73×) |
70.8 MiB |
53 % |
– |
– |
1290–1380 |
1,000,000 |
eagle graph + compaction + reorder |
124 ms (124 ms–124 ms) |
– |
139 ms |
1.74e9 |
7.66e10 (30 %) |
1.57e11 (49 %) |
94 MiB (2.46×) |
70.8 MiB |
1.1e+02 % |
– |
– |
1320–1350 |
1,000,000 |
eagle graph, K fused steps per launch |
189 ms (189 ms–189 ms) |
– |
204 ms |
1.14e9 |
5.03e10 (20 %) |
1.39e10 (4.4 %) |
50 MiB (1.31×) |
64.1 MiB |
– |
– |
– |
1590–1590 |
1,000,000 |
eagle graph, K fused steps per launch + compaction |
60.6 ms (60.3 ms–60.8 ms) |
– |
75.4 ms |
3.57e9 |
1.57e11 (62 %) |
2.67e10 (8.3 %) |
66 MiB (1.73×) |
70.8 MiB |
– |
– |
– |
1470–1545 |
1,000,000 |
eagle auto (eagle picks the launch mode) |
50.1 ms (50.1 ms–50.1 ms) |
– |
64.5 ms |
4.32e9 |
1.90e11 (75 %) |
1.73e11 (54 %) |
50 MiB (1.31×) |
64.1 MiB |
– |
– |
– |
1590–1590 |
1,000,000 |
hawk + eagle persistent |
50.2 ms (50.2 ms–50.2 ms) |
– |
64.5 ms |
4.31e9 |
1.90e11 (75 %) |
1.73e11 (54 %) |
50 MiB (1.31×) |
64.1 MiB |
– |
0.9611841620724161 |
– |
1590–1590 |
1,000,000 |
eagle.simulate (eagle picks the launch mode) |
50.1 ms (50.1 ms–50.1 ms) |
– |
64.3 ms |
4.32e9 |
1.90e11 (75 %) |
1.73e11 (54 %) |
50 MiB (1.31×) |
64.1 MiB |
– |
– |
– |
1590–1590 |
1,000,000 |
eager launch loop |
299 ms (299 ms–300 ms) |
– |
314 ms |
7.22e8 |
3.18e10 (12 %) |
1.39e11 (43 %) |
50 MiB (1.31×) |
62.1 MiB |
45 % |
– |
– |
1485–1500 |
1,000,000 |
CuPy masked |
3.68 s (3.68 s–3.68 s) |
– |
3.7 s |
5.87e7 |
2.58e9 (1 %) |
2.45e11 (76 %) |
194 MiB (5.09×) |
68.7 MiB |
– |
– |
– |
1095–1185 |
1,000,000 |
PyTorch masked |
3.85 s (3.85 s–3.85 s) |
– |
3.88 s |
5.61e7 |
2.47e9 (0.97 %) |
2.34e11 (73 %) |
262 MiB (6.87×) |
68.7 MiB |
– |
– |
– |
1155–1335 |
1,000,000 |
PyTorch, CUDA graph of 16 steps |
3.84 s (3.84 s–3.85 s) |
– |
3.87 s |
5.63e7 |
2.48e9 (0.97 %) |
2.37e11 (74 %) |
262 MiB (6.87×) |
68.8 MiB |
– |
– |
– |
1050–1350 |
1,000,000 |
PyTorch torch.compile (reduce-overhead) |
502 ms (501 ms–503 ms) |
– |
527 ms |
4.31e8 |
1.90e10 (7.5 %) |
1.75e11 (55 %) |
80 MiB (2.1×) |
77.5 MiB |
– |
– |
– |
1425–1500 |
1,000,000 |
JAX jit + vmap(while_loop) |
594 ms (590 ms–596 ms) |
– |
611 ms |
3.64e8 |
1.60e10 (6.3 %) |
1.48e11 (46 %) |
128 MiB (3.36×) |
101 MiB |
– |
– |
– |
1260–1380 |
1,000,000 |
Warp per-thread kernel |
194 ms (194 ms–194 ms) |
– |
226 ms |
1.12e9 |
4.91e10 (19 %) |
3.30e8 (0.1 %) |
64 MiB (1.68×) |
68.7 MiB |
– |
– |
– |
1590–1590 |
1,000,000 |
CPU OpenMP |
548 ms (545 ms–551 ms) |
– |
565 ms |
3.95e8 |
1.74e10 (–) |
7.60e10 (–) |
– |
109 MiB |
– |
– |
– |
– |
Fastest arm by mean wall time (middle 10 of 12 runs):
N = 1,000: eagle auto (eagle picks the launch mode) (502 µs), then eagle.simulate (eagle picks the launch mode) (507 µs, 1.01× the time)
N = 10,000: eagle graph, K fused steps per launch + compaction (2.05 ms), then eagle.simulate (eagle picks the launch mode) (2.15 ms, 1.05× the time)
N = 100,000: eagle.simulate (eagle picks the launch mode) (6.17 ms), then eagle auto (eagle picks the launch mode) (6.22 ms, 1.01× the time)
N = 1,000,000: eagle auto (eagle picks the launch mode) (50.1 ms), then eagle.simulate (eagle picks the launch mode) (50.1 ms, 1× the time)
Uniform stop step: every sample runs 1000 steps#
N |
arm |
wall (range) |
kernel-only |
end-to-end |
sample·steps/s |
useful FLOP/s (% peak) |
issued B/s (% peak) |
device memory (× minimum) |
host memory |
% issue peak |
lane util |
launches (kernel/memcpy/API) |
SM clock (MHz) |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
1,000 |
eagle graph (device loop) |
1.08 ms (1.07 ms–1.09 ms) |
– |
1.29 ms |
9.25e8 |
4.07e10 (16 %) |
1.15e9 (0.36 %) |
2 MiB (52.4×) |
2.12 MiB |
57 % |
– |
– |
1545–1545 |
1,000 |
eagle graph + compaction |
3.53 ms (3.51 ms–3.55 ms) |
– |
3.84 ms |
2.83e8 |
1.25e10 (4.9 %) |
2.19e10 (6.8 %) |
2 MiB (52.4×) |
2.18 MiB |
18 % |
– |
– |
1590–1590 |
1,000 |
eagle graph + compaction + reorder |
3.61 ms (3.61 ms–3.61 ms) |
– |
4 ms |
2.77e8 |
1.22e10 (4.8 %) |
2.14e10 (6.7 %) |
2 MiB (52.4×) |
2.26 MiB |
17 % |
– |
– |
1590–1590 |
1,000 |
eagle graph, K fused steps per launch |
1.04 ms (1.02 ms–1.06 ms) |
– |
1.26 ms |
9.64e8 |
4.24e10 (17 %) |
4.43e9 (1.4 %) |
2 MiB (52.4×) |
2.16 MiB |
– |
– |
– |
1590–1590 |
1,000 |
eagle graph, K fused steps per launch + compaction |
988 µs (984 µs–995 µs) |
– |
1.28 ms |
1.01e9 |
4.45e10 (17 %) |
4.98e9 (1.6 %) |
2 MiB (52.4×) |
2.2 MiB |
– |
– |
– |
1590–1590 |
1,000 |
eagle auto (eagle picks the launch mode) |
505 µs (497 µs–520 µs) |
– |
771 µs |
1.98e9 |
8.72e10 (34 %) |
7.93e10 (25 %) |
2 MiB (52.4×) |
2.19 MiB |
– |
– |
– |
1590–1590 |
1,000 |
hawk + eagle persistent |
1.13 ms (959 µs–1.22 ms) |
– |
1.33 ms |
8.86e8 |
3.90e10 (15 %) |
3.55e10 (11 %) |
2 MiB (52.4×) |
2.11 MiB |
– |
0.7443311737804879 |
– |
1590–1590 |
1,000 |
eagle.simulate (eagle picks the launch mode) |
512 µs (510 µs–513 µs) |
– |
771 µs |
1.95e9 |
8.60e10 (34 %) |
7.82e10 (24 %) |
2 MiB (52.4×) |
2.16 MiB |
– |
– |
– |
1575–1590 |
1,000 |
eager launch loop |
29.1 ms (29.1 ms–29.2 ms) |
– |
29.3 ms |
3.43e7 |
1.51e9 (0.59 %) |
2.50e9 (0.78 %) |
2 MiB (52.4×) |
0.16 MiB |
2.1 % |
– |
– |
585–615 |
1,000 |
CuPy masked |
570 ms (569 ms–571 ms) |
– |
570 ms |
1.76e6 |
7.72e7 (0.03 %) |
1.58e9 (0.49 %) |
2 MiB (52.4×) |
4.06 MiB |
– |
– |
– |
585–585 |
1,000 |
PyTorch masked |
408 ms (405 ms–413 ms) |
– |
408 ms |
2.45e6 |
1.08e8 (0.042 %) |
2.21e9 (0.69 %) |
2 MiB (52.4×) |
0.121 MiB |
– |
– |
– |
1050–1080 |
1,000 |
PyTorch, CUDA graph of 16 steps |
89 ms (89 ms–89.1 ms) |
– |
89.3 ms |
1.12e7 |
4.94e8 (0.19 %) |
1.02e10 (3.2 %) |
4 MiB (105×) |
0.469 MiB |
– |
– |
– |
1590–1590 |
1,000 |
PyTorch torch.compile (reduce-overhead) |
139 ms (138 ms–139 ms) |
– |
139 ms |
7.22e6 |
3.18e8 (0.12 %) |
6.35e8 (0.2 %) |
2 MiB (52.4×) |
32.5 MiB |
– |
– |
– |
585–585 |
1,000 |
JAX jit + vmap(while_loop) |
26.1 ms (25.8 ms–26.7 ms) |
– |
28.5 ms |
3.82e7 |
1.68e9 (0.66 %) |
3.37e9 (1.1 %) |
0 MiB (0×) |
0.125 MiB |
– |
– |
– |
975–975 |
1,000 |
Warp per-thread kernel |
2.4 ms (2.37 ms–2.42 ms) |
– |
3.01 ms |
4.17e8 |
1.83e10 (7.2 %) |
2.67e7 (0.0083 %) |
32 MiB (839×) |
0.105 MiB |
– |
– |
– |
1590–1590 |
1,000 |
CPU OpenMP |
1.33 ms (1.28 ms–1.42 ms) |
– |
1.36 ms |
7.49e8 |
3.30e10 (–) |
5.47e10 (–) |
– |
0.0781 MiB |
– |
– |
– |
– |
10,000 |
eagle graph (device loop) |
2.85 ms (2.84 ms–2.85 ms) |
– |
3.23 ms |
3.51e9 |
1.55e11 (61 %) |
4.36e9 (1.4 %) |
2 MiB (5.24×) |
2.68 MiB |
2.2e+02 % |
– |
– |
1590–1590 |
10,000 |
eagle graph + compaction |
5.7 ms (5.67 ms–5.72 ms) |
– |
6.14 ms |
1.75e9 |
7.72e10 (30 %) |
1.36e11 (42 %) |
2 MiB (5.24×) |
2.74 MiB |
1.1e+02 % |
– |
– |
1590–1590 |
10,000 |
eagle graph + compaction + reorder |
6.04 ms (6.03 ms–6.07 ms) |
– |
6.57 ms |
1.66e9 |
7.28e10 (29 %) |
1.28e11 (40 %) |
2 MiB (5.24×) |
2.98 MiB |
1e+02 % |
– |
– |
1590–1590 |
10,000 |
eagle graph, K fused steps per launch |
2.89 ms (2.83 ms–2.95 ms) |
– |
3.31 ms |
3.46e9 |
1.52e11 (60 %) |
1.59e10 (5 %) |
2 MiB (5.24×) |
2.67 MiB |
– |
– |
– |
1590–1590 |
10,000 |
eagle graph, K fused steps per launch + compaction |
2.78 ms (2.77 ms–2.79 ms) |
– |
3.21 ms |
3.60e9 |
1.58e11 (62 %) |
1.77e10 (5.5 %) |
2 MiB (5.24×) |
2.68 MiB |
– |
– |
– |
1590–1590 |
10,000 |
eagle auto (eagle picks the launch mode) |
2.23 ms (2.23 ms–2.23 ms) |
– |
2.66 ms |
4.49e9 |
1.98e11 (78 %) |
1.80e11 (56 %) |
2 MiB (5.24×) |
2.75 MiB |
– |
– |
– |
1590–1590 |
10,000 |
hawk + eagle persistent |
2.29 ms (2.28 ms–2.3 ms) |
– |
2.65 ms |
4.36e9 |
1.92e11 (75 %) |
1.75e11 (55 %) |
2 MiB (5.24×) |
2.74 MiB |
– |
0.9743215604110546 |
– |
1590–1590 |
10,000 |
eagle.simulate (eagle picks the launch mode) |
2.24 ms (2.23 ms–2.24 ms) |
– |
2.68 ms |
4.46e9 |
1.96e11 (77 %) |
1.79e11 (56 %) |
2 MiB (5.24×) |
2.7 MiB |
– |
– |
– |
1590–1590 |
10,000 |
eager launch loop |
39.8 ms (38.9 ms–41.1 ms) |
– |
40.2 ms |
2.51e8 |
1.10e10 (4.3 %) |
1.83e10 (5.7 %) |
2 MiB (5.24×) |
0.754 MiB |
16 % |
– |
– |
585–585 |
10,000 |
CuPy masked |
568 ms (565 ms–574 ms) |
– |
569 ms |
1.76e7 |
7.74e8 (0.3 %) |
1.59e10 (5 %) |
2 MiB (5.24×) |
4.23 MiB |
– |
– |
– |
585–585 |
10,000 |
PyTorch masked |
408 ms (407 ms–410 ms) |
– |
409 ms |
2.45e7 |
1.08e9 (0.42 %) |
2.21e10 (6.9 %) |
2 MiB (5.24×) |
0.676 MiB |
– |
– |
– |
1050–1080 |
10,000 |
PyTorch, CUDA graph of 16 steps |
87.7 ms (87.7 ms–87.7 ms) |
– |
88.3 ms |
1.14e8 |
5.02e9 (2 %) |
1.04e11 (32 %) |
4 MiB (10.5×) |
0.852 MiB |
– |
– |
– |
1590–1590 |
10,000 |
PyTorch torch.compile (reduce-overhead) |
190 ms (188 ms–191 ms) |
– |
191 ms |
5.26e7 |
2.31e9 (0.91 %) |
4.63e9 (1.4 %) |
2 MiB (5.24×) |
33.3 MiB |
– |
– |
– |
1290–1290 |
10,000 |
JAX jit + vmap(while_loop) |
27.9 ms (27.8 ms–27.9 ms) |
– |
30.1 ms |
3.59e8 |
1.58e10 (6.2 %) |
3.16e10 (9.9 %) |
0 MiB (0×) |
0.715 MiB |
– |
– |
– |
1380–1380 |
10,000 |
Warp per-thread kernel |
2.36 ms (2.34 ms–2.37 ms) |
– |
2.95 ms |
4.24e9 |
1.87e11 (73 %) |
2.72e8 (0.085 %) |
32 MiB (83.9×) |
0.75 MiB |
– |
– |
– |
1590–1590 |
10,000 |
CPU OpenMP |
7.73 ms (7.51 ms–8.06 ms) |
– |
7.82 ms |
1.29e9 |
5.69e10 (–) |
9.45e10 (–) |
– |
1.19 MiB |
– |
– |
– |
– |
100,000 |
eagle graph (device loop) |
22.5 ms (22.4 ms–22.5 ms) |
– |
24.2 ms |
4.45e9 |
1.96e11 (77 %) |
5.53e9 (1.7 %) |
6 MiB (1.57×) |
8.26 MiB |
2.8e+02 % |
– |
– |
1590–1590 |
100,000 |
eagle graph + compaction |
38.7 ms (38.4 ms–39.2 ms) |
– |
40.6 ms |
2.58e9 |
1.14e11 (45 %) |
2.00e11 (62 %) |
8 MiB (2.1×) |
8.41 MiB |
1.6e+02 % |
– |
– |
1395–1440 |
100,000 |
eagle graph + compaction + reorder |
38.8 ms (38.5 ms–39 ms) |
– |
40.8 ms |
2.58e9 |
1.13e11 (45 %) |
1.99e11 (62 %) |
10 MiB (2.62×) |
8.65 MiB |
1.6e+02 % |
– |
– |
1395–1440 |
100,000 |
eagle graph, K fused steps per launch |
22.8 ms (22.8 ms–22.8 ms) |
– |
24.6 ms |
4.39e9 |
1.93e11 (76 %) |
2.02e10 (6.3 %) |
6 MiB (1.57×) |
8.25 MiB |
– |
– |
– |
1590–1590 |
100,000 |
eagle graph, K fused steps per launch + compaction |
23 ms (22.9 ms–23.1 ms) |
– |
25 ms |
4.36e9 |
1.92e11 (75 %) |
2.14e10 (6.7 %) |
8 MiB (2.1×) |
8.29 MiB |
– |
– |
– |
1590–1590 |
100,000 |
eagle auto (eagle picks the launch mode) |
21.8 ms (21.7 ms–21.9 ms) |
– |
23.6 ms |
4.59e9 |
2.02e11 (79 %) |
1.84e11 (57 %) |
6 MiB (1.57×) |
8.34 MiB |
– |
– |
– |
1590–1590 |
100,000 |
hawk + eagle persistent |
21.8 ms (21.7 ms–21.8 ms) |
– |
23.5 ms |
4.59e9 |
2.02e11 (79 %) |
1.84e11 (57 %) |
6 MiB (1.57×) |
8.27 MiB |
– |
0.9981729442428579 |
– |
1590–1590 |
100,000 |
eagle.simulate (eagle picks the launch mode) |
21.7 ms (21.7 ms–21.7 ms) |
– |
23.6 ms |
4.62e9 |
2.03e11 (80 %) |
1.85e11 (58 %) |
6 MiB (1.57×) |
8.35 MiB |
– |
– |
– |
1590–1590 |
100,000 |
eager launch loop |
57.1 ms (57 ms–57.1 ms) |
– |
58.8 ms |
1.75e9 |
7.71e10 (30 %) |
1.28e11 (40 %) |
6 MiB (1.57×) |
6.33 MiB |
1.1e+02 % |
– |
– |
1590–1590 |
100,000 |
CuPy masked |
563 ms (563 ms–564 ms) |
– |
565 ms |
1.77e8 |
7.81e9 (3.1 %) |
1.60e11 (50 %) |
24 MiB (6.29×) |
7.74 MiB |
– |
– |
– |
1590–1590 |
100,000 |
PyTorch masked |
438 ms (433 ms–445 ms) |
– |
440 ms |
2.29e8 |
1.01e10 (4 %) |
2.06e11 (64 %) |
26 MiB (6.82×) |
6.29 MiB |
– |
– |
– |
1575–1590 |
100,000 |
PyTorch, CUDA graph of 16 steps |
267 ms (267 ms–267 ms) |
– |
269 ms |
3.75e8 |
1.65e10 (6.5 %) |
3.42e11 (1.1e+02 %) |
26 MiB (6.82×) |
6.57 MiB |
– |
– |
– |
1380–1410 |
100,000 |
PyTorch torch.compile (reduce-overhead) |
184 ms (184 ms–185 ms) |
– |
187 ms |
5.42e8 |
2.39e10 (9.4 %) |
4.77e10 (15 %) |
8 MiB (2.1×) |
36.8 MiB |
– |
– |
– |
675–675 |
100,000 |
JAX jit + vmap(while_loop) |
69.2 ms (68.4 ms–70.6 ms) |
– |
72.6 ms |
1.44e9 |
6.36e10 (25 %) |
1.27e11 (40 %) |
12 MiB (3.15×) |
10.8 MiB |
– |
– |
– |
1575–1590 |
100,000 |
Warp per-thread kernel |
22.5 ms (22.5 ms–22.5 ms) |
– |
24.8 ms |
4.44e9 |
1.95e11 (77 %) |
2.84e8 (0.089 %) |
32 MiB (8.39×) |
6.86 MiB |
– |
– |
– |
1590–1590 |
100,000 |
CPU OpenMP |
74.8 ms (73.6 ms–75.5 ms) |
– |
75.7 ms |
1.34e9 |
5.88e10 (–) |
9.76e10 (–) |
– |
11 MiB |
– |
– |
– |
– |
1,000,000 |
eagle graph (device loop) |
213 ms (213 ms–213 ms) |
– |
228 ms |
4.69e9 |
2.06e11 (81 %) |
5.82e9 (1.8 %) |
50 MiB (1.31×) |
64.1 MiB |
2.9e+02 % |
– |
– |
1590–1590 |
1,000,000 |
eagle graph + compaction |
348 ms (347 ms–348 ms) |
– |
363 ms |
2.87e9 |
1.26e11 (50 %) |
2.22e11 (69 %) |
66 MiB (1.73×) |
64.2 MiB |
1.8e+02 % |
– |
– |
1215–1305 |
1,000,000 |
eagle graph + compaction + reorder |
350 ms (349 ms–350 ms) |
– |
365 ms |
2.86e9 |
1.26e11 (49 %) |
2.21e11 (69 %) |
94 MiB (2.46×) |
64.4 MiB |
1.8e+02 % |
– |
– |
1200–1275 |
1,000,000 |
eagle graph, K fused steps per launch |
215 ms (215 ms–215 ms) |
– |
230 ms |
4.65e9 |
2.05e11 (80 %) |
2.14e10 (6.7 %) |
50 MiB (1.31×) |
64.1 MiB |
– |
– |
– |
1590–1590 |
1,000,000 |
eagle graph, K fused steps per launch + compaction |
215 ms (215 ms–215 ms) |
– |
230 ms |
4.65e9 |
2.05e11 (80 %) |
2.29e10 (7.1 %) |
66 MiB (1.73×) |
64.1 MiB |
– |
– |
– |
1590–1590 |
1,000,000 |
eagle auto (eagle picks the launch mode) |
211 ms (211 ms–211 ms) |
– |
226 ms |
4.73e9 |
2.08e11 (82 %) |
1.90e11 (59 %) |
50 MiB (1.31×) |
64.1 MiB |
– |
– |
– |
1590–1590 |
1,000,000 |
hawk + eagle persistent |
211 ms (211 ms–211 ms) |
– |
226 ms |
4.73e9 |
2.08e11 (82 %) |
1.89e11 (59 %) |
50 MiB (1.31×) |
64.1 MiB |
– |
0.9996442865770644 |
– |
1590–1590 |
1,000,000 |
eagle.simulate (eagle picks the launch mode) |
211 ms (211 ms–211 ms) |
– |
226 ms |
4.73e9 |
2.08e11 (82 %) |
1.90e11 (59 %) |
50 MiB (1.31×) |
64.2 MiB |
– |
– |
– |
1590–1590 |
1,000,000 |
eager launch loop |
355 ms (354 ms–355 ms) |
– |
369 ms |
2.82e9 |
1.24e11 (49 %) |
2.06e11 (64 %) |
50 MiB (1.31×) |
62.1 MiB |
1.7e+02 % |
– |
– |
1275–1290 |
1,000,000 |
CuPy masked |
3.68 s (3.68 s–3.68 s) |
– |
3.7 s |
2.72e8 |
1.20e10 (4.7 %) |
2.45e11 (76 %) |
194 MiB (5.09×) |
65 MiB |
– |
– |
– |
1080–1200 |
1,000,000 |
PyTorch masked |
3.85 s (3.85 s–3.85 s) |
– |
3.87 s |
2.60e8 |
1.14e10 (4.5 %) |
2.34e11 (73 %) |
262 MiB (6.87×) |
62.1 MiB |
– |
– |
– |
1050–1170 |
1,000,000 |
PyTorch, CUDA graph of 16 steps |
3.84 s (3.84 s–3.85 s) |
– |
3.86 s |
2.60e8 |
1.14e10 (4.5 %) |
2.37e11 (74 %) |
262 MiB (6.87×) |
62.5 MiB |
– |
– |
– |
1020–1125 |
1,000,000 |
PyTorch torch.compile (reduce-overhead) |
496 ms (495 ms–497 ms) |
– |
511 ms |
2.02e9 |
8.87e10 (35 %) |
1.77e11 (55 %) |
80 MiB (2.1×) |
71.4 MiB |
– |
– |
– |
1410–1425 |
1,000,000 |
JAX jit + vmap(while_loop) |
588 ms (586 ms–590 ms) |
– |
605 ms |
1.70e9 |
7.48e10 (29 %) |
1.50e11 (47 %) |
128 MiB (3.36×) |
99.1 MiB |
– |
– |
– |
1275–1395 |
1,000,000 |
Warp per-thread kernel |
220 ms (220 ms–220 ms) |
– |
252 ms |
4.55e9 |
2.00e11 (79 %) |
2.91e8 (0.091 %) |
64 MiB (1.68×) |
68.8 MiB |
– |
– |
– |
1590–1590 |
1,000,000 |
CPU OpenMP |
708 ms (705 ms–712 ms) |
– |
726 ms |
1.41e9 |
6.21e10 (–) |
1.03e11 (–) |
– |
109 MiB |
– |
– |
– |
– |
Fastest arm by mean wall time (middle 10 of 12 runs):
N = 1,000: eagle auto (eagle picks the launch mode) (505 µs), then eagle.simulate (eagle picks the launch mode) (512 µs, 1.01× the time)
N = 10,000: eagle auto (eagle picks the launch mode) (2.23 ms), then eagle.simulate (eagle picks the launch mode) (2.24 ms, 1.01× the time)
N = 100,000: eagle.simulate (eagle picks the launch mode) (21.7 ms), then eagle auto (eagle picks the launch mode) (21.8 ms, 1.01× the time)
N = 1,000,000: eagle.simulate (eagle picks the launch mode) (211 ms), then eagle auto (eagle picks the launch mode) (211 ms, 1× the time)
Compaction cadence sweep (isolates the active-set map)#
the active-set map alone. For a given K, matched_cadence (the plain per-sample step, no map) and compact (eagle_graph_compact’s active-set step) check the loop’s stop guard at the SAME cadence, once every K steps; the only structural difference between the two rows at a given K is whether the step kernel reads the active-set map. The main table’s eagle_graph row checks its guard every step, so its difference from eagle_graph_compact there mixes the cadence change with the map – this sweep does not.
Config: Spread stop steps (log-uniform in [10, 1000]), up to 1000 steps.
N |
K (steps between guard checks) |
matched cadence, no map (range) |
eagle graph + compaction (range) |
ratio (matched/compact, |
|---|---|---|---|---|
100,000 |
8 |
30 ms (29.6 ms–30.4 ms) |
22.3 ms (22.2 ms–22.5 ms) |
1.35× |
100,000 |
16 |
29.7 ms (29.7 ms–29.7 ms) |
19.8 ms (19.6 ms–20 ms) |
1.5× |
100,000 |
32 |
29.4 ms (29.2 ms–29.7 ms) |
19 ms (18.7 ms–19.3 ms) |
1.55× |
1,000,000 |
8 |
293 ms (293 ms–293 ms) |
265 ms (264 ms–265 ms) |
1.11× |
1,000,000 |
16 |
293 ms (293 ms–294 ms) |
254 ms (254 ms–254 ms) |
1.16× |
1,000,000 |
32 |
294 ms (293 ms–294 ms) |
248 ms (248 ms–248 ms) |
1.18× |
FP32: eagle vs Warp (float32, S = 1000)#
Every arm runs the same oscillator in float32: eagle.simulate with scalar_type=”float32”, NVIDIA Warp’s per-thread kernel and JAX’s vmapped loop, all over the same float32 inputs. Each cell runs 2 untimed runs per arm first (eagle’s automatic arm picks its entry over them: the size rule, then a measured run of the other entry, keeping the faster; the entry it settled on is in the last column), then each arm in its own block (no interleaving): an untimed ramp of back-to-back runs for 0.5 s, then 5 timed runs. Before any timing counts, eagle and JAX reach step counts identical to Warp’s and states within absolute 0.0001 of Warp’s (the largest differences are in the JSON). Each time is the mean of the middle 3 of the 5 runs (highest and lowest dropped), with the range of those 3; eagle / Warp is the ratio of these means: below 1 eagle is faster. Before each arm’s block the GPU is cooled down: the block starts once the temperature is within 3 °C of the idle 67 °C and the GPU is idle, or after 20 s. Right after each timed run the SM clock is read (its range is shown per arm); cycles are the wall time times that clock, and eagle / Warp (cycles) is their ratio.
Every arm runs each sample to its own stop and no further than max_steps (S): eagle and Warp enforce the cap per sample, JAX in its while-loop condition, PyTorch and CuPy by the outer step count (a batch-level cap); an arm that runs whole blocks of steps per launch or replay stops at the first block boundary at or after it.
What each arm needed to get there: Warp, a hand-written step counter and cap inside the kernel; JAX, a cap in the while_loop condition; eagle, the kernel source unchanged, with its loop shape, unrolling and device entry chosen by hawk + eagle (per device, per batch).
Cap fixture (checked before any timing counts): N = 4,096, 8 samples with stop step 2S = 256, S = 128; every arm (3) reported k = min(stop step, S) per sample and the same finished count (4,088): passed.
distribution |
N |
eagle.simulate (range) |
Warp (range) |
JAX (range) |
eagle / Warp |
eagle / Warp (cycles) |
eagle entry |
|---|---|---|---|---|---|---|---|
uniform |
10,000 |
246 µs (241 µs–254 µs), 645–645 MHz |
259 µs (257 µs–260 µs), 645–645 MHz |
29.7 ms (29.4 ms–29.9 ms), 765–765 MHz |
0.948× |
0.948× |
fused_one |
uniform |
100,000 |
845 µs (836 µs–856 µs), 1290–1560 MHz |
894 µs (893 µs–894 µs), 1170–1170 MHz |
28.8 ms (28.7 ms–28.8 ms), 1065–1065 MHz |
0.946× |
1.23× |
persist |
uniform |
1,000,000 |
6.55 ms (6.31 ms–6.83 ms), 1365–1545 MHz |
6.45 ms (6.26 ms–6.62 ms), 1365–1500 MHz |
204 ms (204 ms–205 ms), 1260–1320 MHz |
1.02× |
1.03× |
fused_one |
spread |
10,000 |
235 µs (227 µs–245 µs), 690–690 MHz |
277 µs (270 µs–283 µs), 585–585 MHz |
29.4 ms (29.3 ms–29.5 ms), 750–750 MHz |
0.847× |
0.999× |
fused_one |
spread |
100,000 |
558 µs (549 µs–567 µs), 870–1125 MHz |
904 µs (900 µs–911 µs), 1065–1065 MHz |
28.7 ms (28.3 ms–29.2 ms), 1080–1080 MHz |
0.617× |
0.642× |
persist |
spread |
1,000,000 |
2.72 ms (2.7 ms–2.73 ms), 1245–1275 MHz |
5.42 ms (5.24 ms–5.51 ms), 1485–1590 MHz |
206 ms (205 ms–207 ms), 1185–1230 MHz |
0.502× |
0.412× |
persist |
Fixed overhead per run (N = 64, S = 100)#
item 3: wall at a cell small enough that per-run overhead dominates (N=64, S=100), each arm’s own kernel-only vs wall where the main nsys pass covers this cell’s N (it does not by default, so kernel_only_s is usually null here)
arm |
wall (range) |
kernel-only |
|---|---|---|
eagle graph (device loop) |
761 µs (747 µs–779 µs) |
– |
eagle graph + compaction |
670 µs (665 µs–674 µs) |
– |
eagle graph + compaction + reorder |
894 µs (885 µs–907 µs) |
– |
eagle graph, K fused steps per launch |
433 µs (419 µs–442 µs) |
– |
eagle graph, K fused steps per launch + compaction |
485 µs (477 µs–492 µs) |
– |
eagle auto (eagle picks the launch mode) |
204 µs (195 µs–212 µs) |
– |
hawk + eagle persistent |
259 µs (251 µs–266 µs) |
– |
eagle.simulate (eagle picks the launch mode) |
201 µs (195 µs–207 µs) |
– |
eager launch loop |
3.06 ms (3.02 ms–3.11 ms) |
– |
CuPy masked |
54.3 ms (54 ms–54.8 ms) |
– |
PyTorch masked |
40.2 ms (40.1 ms–40.3 ms) |
– |
PyTorch, CUDA graph of 16 steps |
7.88 ms (7.87 ms–7.9 ms) |
– |
PyTorch torch.compile (reduce-overhead) |
13.9 ms (13.6 ms–14.4 ms) |
– |
JAX jit + vmap(while_loop) |
3.01 ms (3 ms–3.01 ms) |
– |
Warp per-thread kernel |
266 µs (260 µs–271 µs) |
– |
CPU OpenMP |
208 µs (206 µs–209 µs) |
– |
Where each tool fits#
Each tool below is written the way its users write it, and each brings something of its own. Times are walls (mean of the middle 3 of 5 runs) at the largest N; ratios are against the eagle graph arm (below 1 = less time). The fastest arm per N is listed under each table above.
eagle graph (device loop). The step written once as a hawk kernel; the whole loop runs on the device as one CUDA graph with a device-side stop guard, so a run is one launch with no host round trip per step. On this card: N = 1,000,000, spread S=100: 20 ms; spread S=1000: 189 ms; uniform S=1000: 213 ms. Fastest in 0 of 12 cells.
eagle graph + compaction. Adds an active-set map: the launch covers only the samples still running, rebuilt every 16 steps; it is made for batches that thin out (the spread configurations). On this card: N = 1,000,000, spread S=100: 25.6 ms (1.28× the eagle graph time); spread S=1000: 253 ms (1.34× the eagle graph time); uniform S=1000: 348 ms (1.63× the eagle graph time). Fastest in 0 of 12 cells.
eagle graph + compaction + reorder. Adds a physical reorder when the live samples spread thin over their warps, so the mapped reads stay contiguous; the final restore to sample order is inside its wall. On this card: N = 1,000,000, spread S=100: 18.2 ms (0.909× the eagle graph time); spread S=1000: 124 ms (0.658× the eagle graph time); uniform S=1000: 350 ms (1.64× the eagle graph time). Fastest in 0 of 12 cells.
eagle graph, K fused steps per launch. The same step with one decorator argument, @hawk.kernel(steps=16): each launch runs that many steps per sample with the state kept in registers, as Warp’s per-thread loop does, and the launches, guard checks and plane reads shrink by the same factor; its results are bit-identical to the eagle graph arm’s in every cell. K is chosen by hand here; the eagle_graph_auto arm picks it at run time. On this card: N = 1,000,000, spread S=100: 19.3 ms (0.963× the eagle graph time); spread S=1000: 189 ms (1× the eagle graph time); uniform S=1000: 215 ms (1.01× the eagle graph time). Fastest in 0 of 12 cells.
eagle graph, K fused steps per launch + compaction. The K-steps kernel over the active-set map: the register-resident steps of the arm above, and the launches cover only the samples still running, so it fits batches that thin out. On this card: N = 1,000,000, spread S=100: 15.3 ms (0.762× the eagle graph time); spread S=1000: 60.6 ms (0.321× the eagle graph time); uniform S=1000: 215 ms (1.01× the eagle graph time). Fastest in 1 of 12 cells.
eagle auto (eagle picks the launch mode). The same step with @hawk.kernel(steps=”auto”) on the active-set kind, and nothing to tune: it picks the steps per launch itself after every launch, on the device, so a batch whose samples all run to the end gets Warp-like fusion of up to 64 steps per launch, and a batch that thins out gets shorter launches with the active set compacted after each one; no sample takes more steps than the cap. On this card: N = 1,000,000, spread S=100: 6.53 ms (0.326× the eagle graph time); spread S=1000: 50.1 ms (0.265× the eagle graph time); uniform S=1000: 211 ms (0.991× the eagle graph time). Fastest in 5 of 12 cells.
hawk + eagle persistent. The plain (not active-set) automatic kernel, forced through hawk’s persist entry: one launch, no graph, no map, no policy kernel, no cap at 64 – a grid of SMs blocks of 256 lanes each steal the next unfinished sample off a counter when their own finishes, so idle warps refill without ever rebuilding an index map; made for batches too large for one resident wave. On this card: N = 1,000,000, spread S=100: 6.61 ms (0.33× the eagle graph time); spread S=1000: 50.2 ms (0.266× the eagle graph time); uniform S=1000: 211 ms (0.991× the eagle graph time). Fastest in 1 of 12 cells.
eagle.simulate (eagle picks the launch mode). The same automatic kernel through eagle.simulate: the model, its state and its parameters by name, with no plan or plane binding to write; eagle deploys the kernel and builds the same device loop. On this card: N = 1,000,000, spread S=100: 6.51 ms (0.325× the eagle graph time); spread S=1000: 50.1 ms (0.266× the eagle graph time); uniform S=1000: 211 ms (0.99× the eagle graph time). Fastest in 5 of 12 cells.
eager launch loop. The same kernels launched from Python one step at a time; the host sees the state after every step. On this card: N = 1,000,000, spread S=100: 27.2 ms (1.36× the eagle graph time); spread S=1000: 299 ms (1.59× the eagle graph time); uniform S=1000: 355 ms (1.66× the eagle graph time). Fastest in 0 of 12 cells.
CuPy masked. NumPy-like array code on GPUs with no kernel to write; every sample is computed every step, so it fits dense batches where all samples run to the end (the uniform configuration). On this card: N = 1,000,000, spread S=100: 369 ms (18.5× the eagle graph time); spread S=1000: 3.68 s (19.5× the eagle graph time); uniform S=1000: 3.68 s (17.3× the eagle graph time). Fastest in 0 of 12 cells.
PyTorch masked. The same array code in PyTorch tensors; it fits where the computation sits next to a PyTorch model or needs autograd. On this card: N = 1,000,000, spread S=100: 386 ms (19.3× the eagle graph time); spread S=1000: 3.85 s (20.4× the eagle graph time); uniform S=1000: 3.85 s (18.1× the eagle graph time). Fastest in 0 of 12 cells.
PyTorch, CUDA graph of 16 steps. The same PyTorch code with its launch overhead removed by hand: 16 steps captured as one CUDA graph and replayed, one host check per replay, the masking unchanged. On this card: N = 1,000,000, spread S=100: 428 ms (21.4× the eagle graph time); spread S=1000: 3.84 s (20.4× the eagle graph time); uniform S=1000: 3.84 s (18× the eagle graph time). Fastest in 0 of 12 cells.
PyTorch torch.compile (reduce-overhead). The PyTorch step under torch.compile(mode=”reduce-overhead”): Inductor-fused kernels replayed as a CUDA graph. On this card: N = 1,000,000, spread S=100: 50.5 ms (2.52× the eagle graph time); spread S=1000: 502 ms (2.66× the eagle graph time); uniform S=1000: 496 ms (2.33× the eagle graph time). Fastest in 0 of 12 cells.
JAX jit + vmap(while_loop). The whole loop compiled once, a per-sample while loop batched with vmap; the same code runs on CPUs and GPUs and differentiates with jax.grad. On this card: N = 1,000,000, spread S=100: 57.9 ms (2.89× the eagle graph time); spread S=1000: 594 ms (3.15× the eagle graph time); uniform S=1000: 588 ms (2.76× the eagle graph time). Fastest in 0 of 12 cells.
Warp per-thread kernel. Explicit per-thread kernels written in Python, maintained by NVIDIA and differentiable through wp.Tape; each thread runs its own sample and stops at its own step, so finished samples cost nothing. On this card: N = 1,000,000, spread S=100: 19.7 ms (0.984× the eagle graph time); spread S=1000: 194 ms (1.03× the eagle graph time); uniform S=1000: 220 ms (1.03× the eagle graph time). Fastest in 0 of 12 cells.
CPU OpenMP. The same hawk kernel on the host through eagle’s OpenMP team, for machines without a GPU. On this card: N = 1,000,000, spread S=100: 60 ms (3× the eagle graph time); spread S=1000: 548 ms (2.9× the eagle graph time); uniform S=1000: 708 ms (3.32× the eagle graph time). Fastest in 0 of 12 cells.
Compile time#
Its own pass, never part of any wall. Each measurement is a fresh process: the imports and the GPU context start-up run untimed, then the timer spans the first call to a ready kernel (trace + code generation + compile + module load, or the cache lookup). Cold: every cache empty – hawk’s compile cache (HAWK_CACHE_DIR), CuPy’s kernel cache (CUPY_CACHE_DIR), the CUDA driver’s PTX cache (CUDA_CACHE_PATH), PyTorch Inductor’s and Triton’s caches (TORCHINDUCTOR_CACHE_DIR, TRITON_CACHE_DIR) and JAX’s persistent compilation cache (minimum compile time and entry size 0), each a fresh directory. Warm: a new process on the caches the first cold run populated. eagle arms: eagle.deploy of the arm’s hawk kernel, which returns once the device build is done (C++/CUDA emission, nvcc to PTX, driver load; eagle’s own small CuPy kernels compile beside it) while the host compiler builds the host side on another thread, and the arm’s set-up on a 64-sample batch (eagle’s graph built, one step run). CuPy: one masked step on 64 samples (every elementwise and reduction kernel compiled on first use). Warp: the kernel module’s code generation, compile and load on the first launch (64 samples), cached in WARP_CACHE_PATH. PyTorch torch.compile, where it runs: three steps at N (the first call traces and compiles, the second records the CUDA graph, the third replays it), per N (dynamic=False). JAX: lower + compile from abstract shapes, per N (XLA specialises on the shape). CPU OpenMP: the same eagle.deploy as the eagle graph arm and one host run, which waits for the host build. Medians of 3 cold and 3 warm processes.
arm |
compiled |
cold (median of 3) |
warm (median of 3) |
|---|---|---|---|
eagle graph (device loop) |
eagle.deploy of the hawk step kernel (the device build; the host side builds in the background), the device-loop graph |
1.41 s |
107 ms |
eagle graph + compaction |
eagle.deploy of the hawk active-set step kernel (the device build; the host side builds in the background), the compaction-loop graph |
1.05 s |
102 ms |
eagle graph + compaction + reorder |
eagle.deploy of the hawk active-set step kernel (the device build; the host side builds in the background), the compaction + reorder loop graph |
1.17 s |
107 ms |
eagle graph, K fused steps per launch |
eagle.deploy of the hawk K-steps kernel (the device build; the host side builds in the background), the device-loop graph |
1.25 s |
110 ms |
eagle graph, K fused steps per launch + compaction |
eagle.deploy of the hawk active-set K-steps kernel (the device build; the host side builds in the background), the compaction-loop graph |
1.31 s |
112 ms |
eagle auto (eagle picks the launch mode) |
eagle.deploy of the hawk active-set automatic kernel (the device build; the host side builds in the background), the policy kernel, the policy-loop graph |
1.35 s |
99.6 ms |
hawk + eagle persistent |
eagle.deploy of the hawk automatic kernel (the device build; the host side builds in the background), the persistent launch |
1.36 s |
110 ms |
eagle.simulate (eagle picks the launch mode) |
eagle.simulate’s deploy of the hawk active-set automatic kernel (the device build; the host side builds in the background), the policy kernel, the policy-loop graph |
1.32 s |
113 ms |
eager launch loop |
eagle.deploy of the hawk single step kernel (the device build; the host side builds in the background) |
1.09 s |
88.7 ms |
CuPy masked |
CuPy’s elementwise and reduction kernels of one masked step |
1.01 s |
38.4 ms |
PyTorch torch.compile (reduce-overhead) |
torch.compile (Inductor + Triton) and CUDA-graph record, N = 1,000 |
6.45 s |
4.33 s |
PyTorch torch.compile (reduce-overhead) |
torch.compile (Inductor + Triton) and CUDA-graph record, N = 10,000 |
6.82 s |
4.35 s |
PyTorch torch.compile (reduce-overhead) |
torch.compile (Inductor + Triton) and CUDA-graph record, N = 100,000 |
6.91 s |
4.4 s |
PyTorch torch.compile (reduce-overhead) |
torch.compile (Inductor + Triton) and CUDA-graph record, N = 1,000,000 |
6.92 s |
4.34 s |
Warp per-thread kernel |
Warp kernel module (code generation, NVRTC compile, load) |
855 ms |
150 ms |
JAX jit + vmap(while_loop) |
XLA executable for N = 1,000 |
246 ms |
53.6 ms |
JAX jit + vmap(while_loop) |
XLA executable for N = 10,000 |
224 ms |
53.8 ms |
JAX jit + vmap(while_loop) |
XLA executable for N = 100,000 |
269 ms |
53.4 ms |
JAX jit + vmap(while_loop) |
XLA executable for N = 1,000,000 |
272 ms |
55.9 ms |
CPU OpenMP |
eagle.deploy of the hawk step kernel and its first host run, which waits for the host build |
1.76 s |
106 ms |
No compile step: PyTorch masked, PyTorch, CUDA graph of 16 steps (PyTorch’s eager kernels ship prebuilt; the CUDA-graph capture is per N and recorded per configuration as torch_graph_capture_s).
Memory method#
A separate pass after the timing pass. Each (arm, N, configuration) runs in a fresh process: imports, the arm’s kernels built and a warm-up run on a 64-sample batch (every compile happens here; JAX also compiles for shape N from abstract shapes), every caching allocator trimmed (CuPy’s pools freed, torch.cuda.empty_cache(), Warp’s pool release threshold set to 0), then the baseline; then the batch’s inputs generated, the arm set up for N (eagle builds its graph, PyTorch captures its CUDA graph, whose private memory pool counts) and one full run with upload and download. Device memory = the memory the CUDA driver accounts to the process (NVML per-process used memory) after the run minus at the baseline. The caching allocators of CuPy, PyTorch and JAX are left at their defaults and are not trimmed after the baseline; each keeps the memory it reserved until trimmed, so the reading after the run is the allocator’s high-water reservation, the same measure for every arm. A disabled pool returns memory between operations, so its peak could only be caught by sampling, which misses short peaks; the held reservation needs no sampling. Caveats: the reading is what the process holds, so each allocator’s rounding and growth policy count (JAX’s allocator grows in regions that can exceed the request; PyTorch’s rounds blocks up; a captured CUDA graph keeps a private pool; Warp’s stream-ordered pool reserves in chunks of tens of MiB); the driver accounts in pages of about 2 MiB, so small-N rows read 0 or one page. The per-library counters (CuPy pool, torch.cuda.max_memory_reserved/allocated, JAX memory_stats) are recorded in the JSON as a cross-check. Host memory = peak resident set size above the baseline (/proc/self/clear_refs reset at the baseline, then VmHWM). Minimum = the state the workload needs, N × (2 state + 3 per-sample scalars: omega, zeta, stop step) × 8 bytes; the factor is device memory / minimum.
Device memory the process already held at the baseline (CUDA context, loaded kernels and libraries, the warm-up’s leftovers; at N = 1,000,000, first configuration): eagle graph (device loop) 104 MiB; eagle graph + compaction 104 MiB; eagle graph + compaction + reorder 104 MiB; eagle graph, K fused steps per launch 104 MiB; eagle graph, K fused steps per launch + compaction 104 MiB; eagle auto (eagle picks the launch mode) 104 MiB; hawk + eagle persistent 104 MiB; eagle.simulate (eagle picks the launch mode) 104 MiB; eager launch loop 104 MiB; CuPy masked 104 MiB; PyTorch masked 118 MiB; PyTorch, CUDA graph of 16 steps 144 MiB; PyTorch torch.compile (reduce-overhead) 140 MiB; JAX jit + vmap(while_loop) 112 MiB; Warp per-thread kernel 104 MiB.
spread (up to 100 steps)
Fastest arm by mean wall time (middle 10 of 12 runs):
N = 1,000: eagle.simulate (eagle picks the launch mode) (123 µs), then eagle auto (eagle picks the launch mode) (127 µs, 1.03× the time)
N = 10,000: eagle auto (eagle picks the launch mode) (270 µs), then eagle.simulate (eagle picks the launch mode) (280 µs, 1.03× the time)
N = 100,000: hawk + eagle persistent (856 µs), then eagle auto (eagle picks the launch mode) (911 µs, 1.06× the time)
N = 1,000,000: eagle.simulate (eagle picks the launch mode) (6.51 ms), then eagle auto (eagle picks the launch mode) (6.53 ms, 1× the time)
spread (up to 1000 steps)
Fastest arm by mean wall time (middle 10 of 12 runs):
N = 1,000: eagle auto (eagle picks the launch mode) (502 µs), then eagle.simulate (eagle picks the launch mode) (507 µs, 1.01× the time)
N = 10,000: eagle graph, K fused steps per launch + compaction (2.05 ms), then eagle.simulate (eagle picks the launch mode) (2.15 ms, 1.05× the time)
N = 100,000: eagle.simulate (eagle picks the launch mode) (6.17 ms), then eagle auto (eagle picks the launch mode) (6.22 ms, 1.01× the time)
N = 1,000,000: eagle auto (eagle picks the launch mode) (50.1 ms), then eagle.simulate (eagle picks the launch mode) (50.1 ms, 1× the time)
uniform (up to 1000 steps)
Fastest arm by mean wall time (middle 10 of 12 runs):
N = 1,000: eagle auto (eagle picks the launch mode) (505 µs), then eagle.simulate (eagle picks the launch mode) (512 µs, 1.01× the time)
N = 10,000: eagle auto (eagle picks the launch mode) (2.23 ms), then eagle.simulate (eagle picks the launch mode) (2.24 ms, 1.01× the time)
N = 100,000: eagle.simulate (eagle picks the launch mode) (21.7 ms), then eagle auto (eagle picks the launch mode) (21.8 ms, 1.01× the time)
N = 1,000,000: eagle.simulate (eagle picks the launch mode) (211 ms), then eagle auto (eagle picks the launch mode) (211 ms, 1× the time)
eagle_simulate against its hand-built twin eagle_graph_auto:
spread (up to 100 steps), N = 1,000: wall 0.971× the twin’s time
spread (up to 100 steps), N = 10,000: wall 1.03× the twin’s time
spread (up to 100 steps), N = 100,000: wall 1.08× the twin’s time
spread (up to 100 steps), N = 1,000,000: wall 0.998× the twin’s time
spread (up to 1000 steps), N = 1,000: wall 1.01× the twin’s time
spread (up to 1000 steps), N = 10,000: wall 1× the twin’s time
spread (up to 1000 steps), N = 100,000: wall 0.993× the twin’s time
spread (up to 1000 steps), N = 1,000,000: wall 1× the twin’s time
uniform (up to 1000 steps), N = 1,000: wall 1.01× the twin’s time
uniform (up to 1000 steps), N = 10,000: wall 1.01× the twin’s time
uniform (up to 1000 steps), N = 100,000: wall 0.995× the twin’s time
uniform (up to 1000 steps), N = 1,000,000: wall 1× the twin’s time
CPU (Intel Xeon W-2125)#
Measured on#
device |
compute capability / CPU model |
date |
section |
|---|---|---|---|
Intel Xeon W-2125 CPU @ 4.00GHz |
Intel(R) Xeon(R) W-2125 CPU @ 4.00GHz |
2026-10-08 |
CPU (Intel Xeon W-2125 CPU @ 4.00GHz)#
CPU: Intel(R) Xeon(R) W-2125 CPU @ 4.00GHz (4 cores, 8 logical CPUs, 8 float64 lanes per vector instruction). Peak FP64: 5.12e11 FLOP/s (4 cores × 2 FMA units × 8 float64 lanes × 2 FLOP per FMA × 4 GHz base clock). Without SIMD the same cores peak at 6.40e10 FLOP/s. Medians of 5 interleaved repetitions; IQR in parentheses.
Same workload as the GPU card (damped oscillators, classical RK4, float64, every sample stops at its own step). “8 threads” uses all 8 logical CPUs (4 cores × 2 hardware threads). CPU busy = process CPU seconds / wall (all threads and worker processes; spin-waiting counts). Compile times are not in any wall; they have their own table below. 0.2 s pause before every timed repetition, so the previous arm’s thread pools stop spin-waiting first, then 0.05 s of untimed runs of the arm itself, so the timed block starts on a ramped clock (the machine’s busy cores during the pause and the clock at the timer start are recorded per row); a repetition whose run is shorter than 0.1 s repeats it back to back until the repetition spans at least 0.1 s.
The arms run in 3 groups, each interleaved on its own: (eagle_term8, eagle_term1, eagle_compact8, eagle_compact1, eagle_reorder8, eagle_reorder1, python_loops); (eagle_term8, torch_masked, mp_numpy); (eagle_term8, numpy_masked, numba_prange, jax_shard8, jax_vmap). eagle at 8 threads runs in every group as the common reference; its row below is from the first group. Its median wall differed between groups by at most 49 % in any cell (spread S=1000, N=10,000).
SIMD. eagle’s host build of the step (hawk host profile native-vector-math: -O3 -ffp-contract=fast -march=native -mprefer-vector-width=512 -fno-trapping-math -DAETHER_HOST_VECTOR_MATH -std=c++23 -shared -fPIC -DAETHER_CPP_MODE) is vectorised: its entry function holds scalar float64 instructions (vaddsd ×43, vfmadd132sd ×33, vfmadd231sd ×3, vfnmsub132sd ×7, vfnmsub231sd ×5, vmulsd ×24) and packed ones (vaddpd ×14, vfmadd132pd ×43, vfmadd231pd ×5, vfnmadd231pd ×8, vfnmsub132pd ×6, vfnmsub231pd ×2, vmulpd ×20). The active-set build of the step (compaction and reorder arms) is vectorised (packed float64 instructions: vaddpd ×8, vfmadd132pd ×21, vfmadd231pd ×3, vfnmsub132pd ×6, vfnmsub231pd ×2, vmulpd ×12). hawk at commit 8dd8034. Numba’s compiled loop uses scalar float64 instructions only (each sample’s steps are a dependent chain). NumPy and PyTorch use their libraries’ own SIMD kernels for each array operation; JAX’s XLA code was not inspected.
Every arm runs each sample to its own stop, and the stop steps here never exceed max_steps (S); the arms that carry a cap enforce it as well: eagle per sample, JAX in its while-loop condition, the masked arms by their outer step count (a batch-level cap).
What each arm needed to get there: JAX, a cap in the while_loop condition; eagle, the kernel source unchanged, with its loop shape, unrolling and host entry chosen by hawk + eagle (per device, per batch).
arm |
loop structure |
threading |
|---|---|---|
eagle host, termination loop, up to 8 threads |
per sample: eagle.until_done, one host-team pass over dynamic tiles of samples, each sample stepped to its own stop |
OMP_NUM_THREADS=8 (eagle’s OpenMP host team: up to 8 threads; eagle uses one per physical core for a run under eagle._host_loop.SMT_MIN_WORK sample-steps) |
eagle host, termination loop, 1 thread |
per sample: eagle.until_done, one host-team pass over dynamic tiles of samples, each sample stepped to its own stop |
OMP_NUM_THREADS=1 (eagle’s OpenMP host team) |
eagle host, termination + compaction, up to 8 threads |
as the termination loop, the step launched over the active-set map only (recomputed every 16 steps) |
OMP_NUM_THREADS=8 (eagle’s OpenMP host team: up to 8 threads; eagle uses one per physical core for a run under eagle._host_loop.SMT_MIN_WORK sample-steps) |
eagle host, termination + compaction, 1 thread |
as the termination loop, the step launched over the active-set map only (recomputed every 16 steps) |
OMP_NUM_THREADS=1 (eagle’s OpenMP host team) |
eagle host, termination + compaction + reorder, up to 8 threads |
as the compaction loop, plus a reorder of the planes when the live samples spread thin, restored at the end |
OMP_NUM_THREADS=8 (eagle’s OpenMP host team: up to 8 threads; eagle uses one per physical core for a run under eagle._host_loop.SMT_MIN_WORK sample-steps) |
eagle host, termination + compaction + reorder, 1 thread |
as the compaction loop, plus a reorder of the planes when the live samples spread thin, restored at the end |
OMP_NUM_THREADS=1 (eagle’s OpenMP host team) |
Numba prange, 8 threads |
per sample: prange over samples, each runs its own steps to its stop step |
NUMBA_NUM_THREADS=8 and numba.set_num_threads(8) |
JAX sharded over 8 CPU devices |
per step over a shard: 8 shards, each a vmapped while loop running until the shard’s last sample stops or the step cap is reached |
XLA_FLAGS=–xla_force_host_platform_device_count=8, one shard per device |
JAX jit + vmap(while_loop), 1 device |
per step over the batch: the vmapped while loop runs until the last sample stops or the step cap is reached, finished samples held by a select |
XLA’s CPU client defaults, one device (CPU busy column shows the use) |
PyTorch masked, 8 threads |
per step over the batch: whole-batch tensor expressions, finished samples held by torch.where |
torch.set_num_threads(8), OMP_NUM_THREADS=8 |
NumPy masked |
per step over the batch: whole-batch array expressions, finished samples held by np.where |
single-threaded ufuncs; BLAS/OpenMP pools pinned to 1 |
multiprocessing, 8 x NumPy |
per step over a chunk: 8 processes, each the NumPy masked loop over its chunk, stopping when its chunk is done |
multiprocessing.Pool(8), fork; every process pinned to 1 thread |
plain Python loops |
per sample: a Python loop over samples, each runs its own steps to its stop step |
one interpreter thread |
Spread stop steps (log-uniform in [1, 100]), up to 100 steps#
N |
arm |
threads |
wall (IQR) |
sample·steps/s |
useful FLOP/s (% peak) |
CPU busy |
peak memory (× minimum) |
|---|---|---|---|---|---|---|---|
1,000 |
eagle host, termination loop, up to 8 threads |
8 |
55.7 µs (6.95 µs) |
3.96e8 |
1.74e10 (3.4 %) |
4.1 |
20.1 MiB (528×) |
1,000 |
eagle host, termination loop, 1 thread |
1 |
109 µs (636 ns) |
2.02e8 |
8.87e9 (1.7 %) |
1 |
20.1 MiB (526×) |
1,000 |
eagle host, termination + compaction, up to 8 threads |
8 |
597 µs (38.6 µs) |
3.70e7 |
1.63e9 (0.32 %) |
4 |
19.9 MiB (522×) |
1,000 |
eagle host, termination + compaction, 1 thread |
1 |
410 µs (4.68 µs) |
5.38e7 |
2.37e9 (0.46 %) |
1 |
20.3 MiB (532×) |
1,000 |
eagle host, termination + compaction + reorder, up to 8 threads |
8 |
636 µs (37.3 µs) |
3.47e7 |
1.53e9 (0.3 %) |
4 |
20 MiB (525×) |
1,000 |
eagle host, termination + compaction + reorder, 1 thread |
1 |
474 µs (3.93 µs) |
4.66e7 |
2.05e9 (0.4 %) |
1 |
19.6 MiB (515×) |
1,000 |
Numba prange, 8 threads |
8 |
74.8 µs (18.5 µs) |
2.95e8 |
1.30e10 (2.5 %) |
8 |
93.5 MiB (2.45e+03×) |
1,000 |
JAX sharded over 8 CPU devices |
8 |
265 µs (12.2 µs) |
8.32e7 |
3.66e9 (0.72 %) |
4.7 |
35.2 MiB (923×) |
1,000 |
JAX jit + vmap(while_loop), 1 device |
8 |
372 µs (24.5 µs) |
5.93e7 |
2.61e9 (0.51 %) |
1 |
36.2 MiB (949×) |
1,000 |
PyTorch masked, 8 threads |
8 |
8.34 ms (48 µs) |
2.65e6 |
1.16e8 (0.023 %) |
1 |
16.7 MiB (437×) |
1,000 |
NumPy masked |
1 |
4.16 ms (11.1 µs) |
5.30e6 |
2.33e8 (0.046 %) |
1 |
191 MiB (5e+03×) |
1,000 |
multiprocessing, 8 x NumPy |
8 |
7.03 ms (1.11 ms) |
3.14e6 |
1.38e8 (0.027 %) |
5.7 |
190 MiB (4.98e+03×) |
1,000 |
plain Python loops |
1 |
4.9 ms (99.1 µs) |
4.50e6 |
1.98e8 (0.039 %) |
1 |
191 MiB (5e+03×) |
10,000 |
eagle host, termination loop, up to 8 threads |
8 |
270 µs (17.7 µs) |
8.47e8 |
3.73e10 (7.3 %) |
4.1 |
20 MiB (52.5×) |
10,000 |
eagle host, termination loop, 1 thread |
1 |
896 µs (2.62 µs) |
2.55e8 |
1.12e10 (2.2 %) |
1 |
20.2 MiB (52.8×) |
10,000 |
eagle host, termination + compaction, up to 8 threads |
8 |
1.43 ms (22.6 µs) |
1.59e8 |
7.01e9 (1.4 %) |
4 |
19.9 MiB (52.3×) |
10,000 |
eagle host, termination + compaction, 1 thread |
1 |
1.71 ms (44.6 µs) |
1.34e8 |
5.88e9 (1.1 %) |
1 |
19.9 MiB (52.2×) |
10,000 |
eagle host, termination + compaction + reorder, up to 8 threads |
8 |
2.03 ms (57.7 µs) |
1.13e8 |
4.96e9 (0.97 %) |
4 |
19.9 MiB (52.2×) |
10,000 |
eagle host, termination + compaction + reorder, 1 thread |
1 |
2.18 ms (61.1 µs) |
1.05e8 |
4.62e9 (0.9 %) |
1 |
19.9 MiB (52×) |
10,000 |
Numba prange, 8 threads |
8 |
567 µs (179 µs) |
4.03e8 |
1.77e10 (3.5 %) |
7.9 |
93.2 MiB (244×) |
10,000 |
JAX sharded over 8 CPU devices |
8 |
952 µs (170 µs) |
2.40e8 |
1.06e10 (2.1 %) |
5.7 |
31.3 MiB (82.1×) |
10,000 |
JAX jit + vmap(while_loop), 1 device |
8 |
7.03 ms (147 µs) |
3.25e7 |
1.43e9 (0.28 %) |
1.7 |
35.2 MiB (92.4×) |
10,000 |
PyTorch masked, 8 threads |
8 |
26.4 ms (1.07 ms) |
8.66e6 |
3.81e8 (0.074 %) |
1 |
16.9 MiB (44.4×) |
10,000 |
NumPy masked |
1 |
21.8 ms (208 µs) |
1.05e7 |
4.62e8 (0.09 %) |
1 |
191 MiB (500×) |
10,000 |
multiprocessing, 8 x NumPy |
8 |
11 ms (618 µs) |
2.07e7 |
9.11e8 (0.18 %) |
5.9 |
191 MiB (501×) |
10,000 |
plain Python loops |
1 |
51.4 ms (1.88 ms) |
4.45e6 |
1.96e8 (0.038 %) |
1 |
191 MiB (501×) |
100,000 |
eagle host, termination loop, up to 8 threads |
8 |
1.68 ms (507 µs) |
1.36e9 |
6.00e10 (12 %) |
7.9 |
20.2 MiB (5.31×) |
100,000 |
eagle host, termination loop, 1 thread |
1 |
8.66 ms (176 µs) |
2.64e8 |
1.16e10 (2.3 %) |
1 |
20.1 MiB (5.28×) |
100,000 |
eagle host, termination + compaction, up to 8 threads |
8 |
7.72 ms (1.54 ms) |
2.96e8 |
1.30e10 (2.5 %) |
7.8 |
19.9 MiB (5.22×) |
100,000 |
eagle host, termination + compaction, 1 thread |
1 |
19.5 ms (340 µs) |
1.18e8 |
5.17e9 (1 %) |
1 |
19.9 MiB (5.2×) |
100,000 |
eagle host, termination + compaction + reorder, up to 8 threads |
8 |
18.1 ms (521 µs) |
1.26e8 |
5.56e9 (1.1 %) |
7.9 |
19.7 MiB (5.16×) |
100,000 |
eagle host, termination + compaction + reorder, 1 thread |
1 |
21.6 ms (985 µs) |
1.06e8 |
4.65e9 (0.91 %) |
1 |
19.7 MiB (5.18×) |
100,000 |
Numba prange, 8 threads |
8 |
4.86 ms (36.1 µs) |
4.71e8 |
2.07e10 (4 %) |
7.9 |
93.4 MiB (24.5×) |
100,000 |
JAX sharded over 8 CPU devices |
8 |
16.6 ms (512 µs) |
1.38e8 |
6.05e9 (1.2 %) |
7 |
33.9 MiB (8.9×) |
100,000 |
JAX jit + vmap(while_loop), 1 device |
8 |
79.5 ms (10.8 ms) |
2.88e7 |
1.27e9 (0.25 %) |
1.9 |
34.5 MiB (9.03×) |
100,000 |
PyTorch masked, 8 threads |
8 |
320 ms (113 ms) |
7.15e6 |
3.15e8 (0.061 %) |
7.7 |
19.3 MiB (5.07×) |
100,000 |
NumPy masked |
1 |
432 ms (7.29 ms) |
5.30e6 |
2.33e8 (0.046 %) |
0.99 |
191 MiB (50×) |
100,000 |
multiprocessing, 8 x NumPy |
8 |
87.1 ms (2.9 ms) |
2.63e7 |
1.16e9 (0.23 %) |
5.7 |
203 MiB (53.2×) |
1,000,000 |
eagle host, termination loop, up to 8 threads |
8 |
16.5 ms (1.12 ms) |
1.40e9 |
6.15e10 (12 %) |
7.6 |
107 MiB (2.82×) |
1,000,000 |
eagle host, termination loop, 1 thread |
1 |
85.5 ms (1.16 ms) |
2.69e8 |
1.18e10 (2.3 %) |
1 |
107 MiB (2.81×) |
1,000,000 |
eagle host, termination + compaction, up to 8 threads |
8 |
194 ms (58.4 ms) |
1.18e8 |
5.21e9 (1 %) |
7.6 |
111 MiB (2.91×) |
1,000,000 |
eagle host, termination + compaction, 1 thread |
1 |
313 ms (4.89 ms) |
7.35e7 |
3.23e9 (0.63 %) |
1 |
111 MiB (2.91×) |
1,000,000 |
eagle host, termination + compaction + reorder, up to 8 threads |
8 |
215 ms (13.4 ms) |
1.07e8 |
4.71e9 (0.92 %) |
4.9 |
135 MiB (3.54×) |
1,000,000 |
eagle host, termination + compaction + reorder, 1 thread |
1 |
260 ms (32.6 ms) |
8.84e7 |
3.89e9 (0.76 %) |
1 |
135 MiB (3.55×) |
1,000,000 |
Numba prange, 8 threads |
8 |
49.9 ms (13.3 ms) |
4.61e8 |
2.03e10 (4 %) |
7.9 |
93.5 MiB (2.45×) |
1,000,000 |
JAX sharded over 8 CPU devices |
8 |
495 ms (1.14 ms) |
4.65e7 |
2.04e9 (0.4 %) |
7.6 |
197 MiB (5.16×) |
1,000,000 |
JAX jit + vmap(while_loop), 1 device |
8 |
1.01 s (33.6 ms) |
2.28e7 |
1.00e9 (0.2 %) |
2.8 |
183 MiB (4.8×) |
1,000,000 |
PyTorch masked, 8 threads |
8 |
3.4 s (209 ms) |
6.76e6 |
2.98e8 (0.058 %) |
7.7 |
191 MiB (5.01×) |
1,000,000 |
NumPy masked |
1 |
6.65 s (41.4 ms) |
3.46e6 |
1.52e8 (0.03 %) |
1 |
191 MiB (5×) |
1,000,000 |
multiprocessing, 8 x NumPy |
8 |
1.99 s (59.5 ms) |
1.15e7 |
5.07e8 (0.099 %) |
6.9 |
324 MiB (8.48×) |
Fastest arm by median wall time:
N = 1,000: eagle host, termination loop, up to 8 threads (55.7 µs), then Numba prange, 8 threads (74.8 µs, 1.34× the time)
N = 10,000: eagle host, termination loop, up to 8 threads (270 µs), then Numba prange, 8 threads (567 µs, 2.1× the time)
N = 100,000: eagle host, termination loop, up to 8 threads (1.68 ms), then Numba prange, 8 threads (4.86 ms, 2.9× the time)
N = 1,000,000: eagle host, termination loop, up to 8 threads (16.5 ms), then Numba prange, 8 threads (49.9 ms, 3.03× the time)
Spread stop steps (log-uniform in [10, 1000]), up to 1000 steps#
N |
arm |
threads |
wall (IQR) |
sample·steps/s |
useful FLOP/s (% peak) |
CPU busy |
peak memory (× minimum) |
|---|---|---|---|---|---|---|---|
1,000 |
eagle host, termination loop, up to 8 threads |
8 |
287 µs (14.2 µs) |
7.31e8 |
3.22e10 (6.3 %) |
4 |
19.2 MiB (503×) |
1,000 |
eagle host, termination loop, 1 thread |
1 |
877 µs (7.39 µs) |
2.40e8 |
1.05e10 (2.1 %) |
1 |
19.4 MiB (508×) |
1,000 |
eagle host, termination + compaction, up to 8 threads |
8 |
5.09 ms (58.7 µs) |
4.13e7 |
1.82e9 (0.35 %) |
4 |
19.3 MiB (507×) |
1,000 |
eagle host, termination + compaction, 1 thread |
1 |
3.65 ms (61.1 µs) |
5.76e7 |
2.53e9 (0.49 %) |
1 |
20.1 MiB (527×) |
1,000 |
eagle host, termination + compaction + reorder, up to 8 threads |
8 |
5.72 ms (741 µs) |
3.67e7 |
1.62e9 (0.32 %) |
4 |
19.9 MiB (522×) |
1,000 |
eagle host, termination + compaction + reorder, 1 thread |
1 |
4.36 ms (84.3 µs) |
4.82e7 |
2.12e9 (0.41 %) |
1 |
20 MiB (524×) |
1,000 |
Numba prange, 8 threads |
8 |
617 µs (196 µs) |
3.41e8 |
1.50e10 (2.9 %) |
7.9 |
93.5 MiB (2.45e+03×) |
1,000 |
JAX sharded over 8 CPU devices |
8 |
1.19 ms (87.7 µs) |
1.77e8 |
7.79e9 (1.5 %) |
5.8 |
33.8 MiB (885×) |
1,000 |
JAX jit + vmap(while_loop), 1 device |
8 |
3.28 ms (226 µs) |
6.40e7 |
2.82e9 (0.55 %) |
1 |
35 MiB (918×) |
1,000 |
PyTorch masked, 8 threads |
8 |
82 ms (1.35 ms) |
2.56e6 |
1.13e8 (0.022 %) |
1 |
16.8 MiB (441×) |
1,000 |
NumPy masked |
1 |
41.4 ms (697 µs) |
5.08e6 |
2.24e8 (0.044 %) |
1 |
191 MiB (5e+03×) |
1,000 |
multiprocessing, 8 x NumPy |
8 |
49.4 ms (8.27 ms) |
4.25e6 |
1.87e8 (0.037 %) |
7.4 |
190 MiB (4.98e+03×) |
1,000 |
plain Python loops |
1 |
46.7 ms (71.4 µs) |
4.50e6 |
1.98e8 (0.039 %) |
1 |
191 MiB (5.01e+03×) |
10,000 |
eagle host, termination loop, up to 8 threads |
8 |
2.36 ms (1.06 ms) |
9.24e8 |
4.07e10 (7.9 %) |
7.6 |
19.9 MiB (52.3×) |
10,000 |
eagle host, termination loop, 1 thread |
1 |
8.45 ms (46.6 µs) |
2.58e8 |
1.14e10 (2.2 %) |
1 |
20.1 MiB (52.8×) |
10,000 |
eagle host, termination + compaction, up to 8 threads |
8 |
11.9 ms (755 µs) |
1.83e8 |
8.07e9 (1.6 %) |
7.9 |
19.9 MiB (52.1×) |
10,000 |
eagle host, termination + compaction, 1 thread |
1 |
16.1 ms (503 µs) |
1.36e8 |
5.97e9 (1.2 %) |
1 |
20 MiB (52.5×) |
10,000 |
eagle host, termination + compaction + reorder, up to 8 threads |
8 |
12.5 ms (4.56 ms) |
1.75e8 |
7.68e9 (1.5 %) |
7.9 |
20 MiB (52.3×) |
10,000 |
eagle host, termination + compaction + reorder, 1 thread |
1 |
16.9 ms (131 µs) |
1.29e8 |
5.70e9 (1.1 %) |
1 |
19.9 MiB (52.3×) |
10,000 |
Numba prange, 8 threads |
8 |
4.61 ms (71.4 µs) |
4.74e8 |
2.08e10 (4.1 %) |
8 |
93.5 MiB (245×) |
10,000 |
JAX sharded over 8 CPU devices |
8 |
8.29 ms (762 µs) |
2.63e8 |
1.16e10 (2.3 %) |
6.1 |
31.9 MiB (83.5×) |
10,000 |
JAX jit + vmap(while_loop), 1 device |
8 |
60.1 ms (2.11 ms) |
3.63e7 |
1.60e9 (0.31 %) |
1.7 |
35.4 MiB (92.7×) |
10,000 |
PyTorch masked, 8 threads |
8 |
252 ms (653 µs) |
8.65e6 |
3.80e8 (0.074 %) |
1 |
16.9 MiB (44.3×) |
10,000 |
NumPy masked |
1 |
214 ms (4.8 ms) |
1.02e7 |
4.49e8 (0.088 %) |
1 |
191 MiB (500×) |
10,000 |
multiprocessing, 8 x NumPy |
8 |
99.6 ms (5.44 ms) |
2.19e7 |
9.64e8 (0.19 %) |
6.4 |
191 MiB (500×) |
10,000 |
plain Python loops |
1 |
492 ms (11.6 ms) |
4.43e6 |
1.95e8 (0.038 %) |
1 |
191 MiB (501×) |
100,000 |
eagle host, termination loop, up to 8 threads |
8 |
14.9 ms (2.25 ms) |
1.46e9 |
6.41e10 (13 %) |
7.9 |
20 MiB (5.25×) |
100,000 |
eagle host, termination loop, 1 thread |
1 |
83 ms (1.21 ms) |
2.61e8 |
1.15e10 (2.2 %) |
1 |
20.2 MiB (5.29×) |
100,000 |
eagle host, termination + compaction, up to 8 threads |
8 |
91.5 ms (19.1 ms) |
2.37e8 |
1.04e10 (2 %) |
7.6 |
19.8 MiB (5.2×) |
100,000 |
eagle host, termination + compaction, 1 thread |
1 |
185 ms (1.5 ms) |
1.17e8 |
5.16e9 (1 %) |
1 |
20.1 MiB (5.28×) |
100,000 |
eagle host, termination + compaction + reorder, up to 8 threads |
8 |
69.1 ms (15 ms) |
3.14e8 |
1.38e10 (2.7 %) |
7.8 |
19.9 MiB (5.21×) |
100,000 |
eagle host, termination + compaction + reorder, 1 thread |
1 |
139 ms (463 µs) |
1.55e8 |
6.83e9 (1.3 %) |
1 |
20 MiB (5.25×) |
100,000 |
Numba prange, 8 threads |
8 |
48.4 ms (6.58 ms) |
4.47e8 |
1.97e10 (3.8 %) |
7.8 |
93.5 MiB (24.5×) |
100,000 |
JAX sharded over 8 CPU devices |
8 |
153 ms (4.58 ms) |
1.41e8 |
6.22e9 (1.2 %) |
6.9 |
35 MiB (9.17×) |
100,000 |
JAX jit + vmap(while_loop), 1 device |
8 |
674 ms (83.1 ms) |
3.21e7 |
1.41e9 (0.28 %) |
1.9 |
31.8 MiB (8.32×) |
100,000 |
PyTorch masked, 8 threads |
8 |
2.57 s (546 ms) |
8.43e6 |
3.71e8 (0.072 %) |
7.7 |
19.3 MiB (5.07×) |
100,000 |
NumPy masked |
1 |
3.61 s (8.73 ms) |
6.00e6 |
2.64e8 (0.052 %) |
0.99 |
191 MiB (50×) |
100,000 |
multiprocessing, 8 x NumPy |
8 |
747 ms (2.53 ms) |
2.90e7 |
1.28e9 (0.25 %) |
7.2 |
203 MiB (53.1×) |
1,000,000 |
eagle host, termination loop, up to 8 threads |
8 |
145 ms (4.09 ms) |
1.49e9 |
6.57e10 (13 %) |
7.9 |
107 MiB (2.82×) |
1,000,000 |
eagle host, termination loop, 1 thread |
1 |
831 ms (10.5 ms) |
2.60e8 |
1.15e10 (2.2 %) |
1 |
107 MiB (2.82×) |
1,000,000 |
eagle host, termination + compaction, up to 8 threads |
8 |
1.6 s (114 ms) |
1.35e8 |
5.93e9 (1.2 %) |
7.7 |
111 MiB (2.91×) |
1,000,000 |
eagle host, termination + compaction, 1 thread |
1 |
3.04 s (2.14 ms) |
7.10e7 |
3.13e9 (0.61 %) |
1 |
111 MiB (2.91×) |
1,000,000 |
eagle host, termination + compaction + reorder, up to 8 threads |
8 |
940 ms (54.9 ms) |
2.30e8 |
1.01e10 (2 %) |
6.9 |
135 MiB (3.54×) |
1,000,000 |
eagle host, termination + compaction + reorder, 1 thread |
1 |
1.5 s (61.4 ms) |
1.44e8 |
6.33e9 (1.2 %) |
1 |
135 MiB (3.54×) |
1,000,000 |
Numba prange, 8 threads |
8 |
501 ms (14.6 ms) |
4.32e8 |
1.90e10 (3.7 %) |
7.1 |
93.5 MiB (2.45×) |
1,000,000 |
JAX sharded over 8 CPU devices |
8 |
4.97 s (19.3 ms) |
4.35e7 |
1.91e9 (0.37 %) |
7.3 |
213 MiB (5.6×) |
1,000,000 |
JAX jit + vmap(while_loop), 1 device |
8 |
9.84 s (78.2 ms) |
2.20e7 |
9.67e8 (0.19 %) |
2.8 |
153 MiB (4×) |
1,000,000 |
PyTorch masked, 8 threads |
8 |
33.4 s (362 ms) |
6.47e6 |
2.85e8 (0.056 %) |
7.7 |
191 MiB (5.02×) |
1,000,000 |
NumPy masked |
1 |
62.9 s (124 ms) |
3.44e6 |
1.51e8 (0.03 %) |
1 |
191 MiB (5.01×) |
1,000,000 |
multiprocessing, 8 x NumPy |
8 |
19.3 s (43.1 ms) |
1.12e7 |
4.93e8 (0.096 %) |
7.3 |
324 MiB (8.48×) |
Fastest arm by median wall time:
N = 1,000: eagle host, termination loop, up to 8 threads (287 µs), then Numba prange, 8 threads (617 µs, 2.15× the time)
N = 10,000: eagle host, termination loop, up to 8 threads (2.36 ms), then Numba prange, 8 threads (4.61 ms, 1.95× the time)
N = 100,000: eagle host, termination loop, up to 8 threads (14.9 ms), then Numba prange, 8 threads (48.4 ms, 3.26× the time)
N = 1,000,000: eagle host, termination loop, up to 8 threads (145 ms), then Numba prange, 8 threads (501 ms, 3.46× the time)
Uniform stop step: every sample runs 1000 steps#
N |
arm |
threads |
wall (IQR) |
sample·steps/s |
useful FLOP/s (% peak) |
CPU busy |
peak memory (× minimum) |
|---|---|---|---|---|---|---|---|
1,000 |
eagle host, termination loop, up to 8 threads |
8 |
303 µs (9 µs) |
3.31e9 |
1.45e11 (28 %) |
4 |
20.1 MiB (527×) |
1,000 |
eagle host, termination loop, 1 thread |
1 |
910 µs (619 ns) |
1.10e9 |
4.83e10 (9.4 %) |
1 |
20.2 MiB (529×) |
1,000 |
eagle host, termination + compaction, up to 8 threads |
8 |
5.43 ms (569 µs) |
1.84e8 |
8.11e9 (1.6 %) |
4 |
20 MiB (524×) |
1,000 |
eagle host, termination + compaction, 1 thread |
1 |
3.99 ms (81.3 µs) |
2.51e8 |
1.10e10 (2.2 %) |
1 |
19.8 MiB (519×) |
1,000 |
eagle host, termination + compaction + reorder, up to 8 threads |
8 |
4.97 ms (131 µs) |
2.01e8 |
8.85e9 (1.7 %) |
4 |
19.9 MiB (523×) |
1,000 |
eagle host, termination + compaction + reorder, 1 thread |
1 |
4.05 ms (13.8 µs) |
2.47e8 |
1.09e10 (2.1 %) |
1 |
19.9 MiB (522×) |
1,000 |
Numba prange, 8 threads |
8 |
2.1 ms (28.2 µs) |
4.76e8 |
2.09e10 (4.1 %) |
8 |
93.4 MiB (2.45e+03×) |
1,000 |
JAX sharded over 8 CPU devices |
8 |
1.15 ms (56.6 µs) |
8.71e8 |
3.83e10 (7.5 %) |
5.8 |
32.9 MiB (861×) |
1,000 |
JAX jit + vmap(while_loop), 1 device |
8 |
2.99 ms (80 µs) |
3.34e8 |
1.47e10 (2.9 %) |
1 |
36.9 MiB (968×) |
1,000 |
PyTorch masked, 8 threads |
8 |
84.1 ms (548 µs) |
1.19e7 |
5.23e8 (0.1 %) |
1 |
16.7 MiB (439×) |
1,000 |
NumPy masked |
1 |
41.1 ms (2.42 ms) |
2.43e7 |
1.07e9 (0.21 %) |
1 |
191 MiB (5e+03×) |
1,000 |
multiprocessing, 8 x NumPy |
8 |
49.9 ms (9.4 ms) |
2.00e7 |
8.82e8 (0.17 %) |
7.7 |
190 MiB (4.98e+03×) |
1,000 |
plain Python loops |
1 |
222 ms (2.48 ms) |
4.51e6 |
1.98e8 (0.039 %) |
1 |
191 MiB (5.01e+03×) |
10,000 |
eagle host, termination loop, up to 8 threads |
8 |
2.01 ms (6.09 µs) |
4.99e9 |
2.19e11 (43 %) |
8 |
20 MiB (52.4×) |
10,000 |
eagle host, termination loop, 1 thread |
1 |
8.81 ms (31.7 µs) |
1.14e9 |
5.00e10 (9.8 %) |
1 |
20.1 MiB (52.6×) |
10,000 |
eagle host, termination + compaction, up to 8 threads |
8 |
15.6 ms (1.09 ms) |
6.41e8 |
2.82e10 (5.5 %) |
8 |
19.9 MiB (52.2×) |
10,000 |
eagle host, termination + compaction, 1 thread |
1 |
20.3 ms (247 µs) |
4.92e8 |
2.16e10 (4.2 %) |
1 |
20.1 MiB (52.7×) |
10,000 |
eagle host, termination + compaction + reorder, up to 8 threads |
8 |
15.1 ms (509 µs) |
6.63e8 |
2.92e10 (5.7 %) |
7.9 |
19.9 MiB (52.2×) |
10,000 |
eagle host, termination + compaction + reorder, 1 thread |
1 |
19.6 ms (363 µs) |
5.09e8 |
2.24e10 (4.4 %) |
1 |
20 MiB (52.4×) |
10,000 |
Numba prange, 8 threads |
8 |
20.2 ms (1.02 ms) |
4.95e8 |
2.18e10 (4.3 %) |
8 |
93.4 MiB (245×) |
10,000 |
JAX sharded over 8 CPU devices |
8 |
7.74 ms (963 µs) |
1.29e9 |
5.68e10 (11 %) |
6.6 |
32.4 MiB (85×) |
10,000 |
JAX jit + vmap(while_loop), 1 device |
8 |
61.5 ms (2.18 ms) |
1.63e8 |
7.15e9 (1.4 %) |
1.7 |
34.3 MiB (89.9×) |
10,000 |
PyTorch masked, 8 threads |
8 |
254 ms (5.51 ms) |
3.94e7 |
1.74e9 (0.34 %) |
1 |
16.8 MiB (43.9×) |
10,000 |
NumPy masked |
1 |
188 ms (896 µs) |
5.33e7 |
2.34e9 (0.46 %) |
1 |
191 MiB (500×) |
10,000 |
multiprocessing, 8 x NumPy |
8 |
93.7 ms (1.25 ms) |
1.07e8 |
4.70e9 (0.92 %) |
6.6 |
191 MiB (501×) |
100,000 |
eagle host, termination loop, up to 8 threads |
8 |
22.5 ms (896 µs) |
4.44e9 |
1.95e11 (38 %) |
7.5 |
20 MiB (5.25×) |
100,000 |
eagle host, termination loop, 1 thread |
1 |
87 ms (946 µs) |
1.15e9 |
5.05e10 (9.9 %) |
1 |
20.1 MiB (5.28×) |
100,000 |
eagle host, termination + compaction, up to 8 threads |
8 |
96.6 ms (26 ms) |
1.04e9 |
4.56e10 (8.9 %) |
7.9 |
20.2 MiB (5.3×) |
100,000 |
eagle host, termination + compaction, 1 thread |
1 |
216 ms (1.4 ms) |
4.63e8 |
2.04e10 (4 %) |
1 |
20 MiB (5.25×) |
100,000 |
eagle host, termination + compaction + reorder, up to 8 threads |
8 |
103 ms (7.29 ms) |
9.75e8 |
4.29e10 (8.4 %) |
7.9 |
20 MiB (5.24×) |
100,000 |
eagle host, termination + compaction + reorder, 1 thread |
1 |
212 ms (1.31 ms) |
4.72e8 |
2.08e10 (4.1 %) |
1 |
19.8 MiB (5.19×) |
100,000 |
Numba prange, 8 threads |
8 |
216 ms (17 ms) |
4.63e8 |
2.04e10 (4 %) |
7.9 |
93.2 MiB (24.4×) |
100,000 |
JAX sharded over 8 CPU devices |
8 |
160 ms (4.43 ms) |
6.27e8 |
2.76e10 (5.4 %) |
7 |
44.3 MiB (11.6×) |
100,000 |
JAX jit + vmap(while_loop), 1 device |
8 |
723 ms (42.7 ms) |
1.38e8 |
6.08e9 (1.2 %) |
1.9 |
33.1 MiB (8.69×) |
100,000 |
PyTorch masked, 8 threads |
8 |
2.2 s (492 ms) |
4.54e7 |
2.00e9 (0.39 %) |
7.7 |
18.6 MiB (4.89×) |
100,000 |
NumPy masked |
1 |
3.33 s (46.4 ms) |
3.01e7 |
1.32e9 (0.26 %) |
0.99 |
191 MiB (50×) |
100,000 |
multiprocessing, 8 x NumPy |
8 |
733 ms (10.2 ms) |
1.36e8 |
6.00e9 (1.2 %) |
7.2 |
203 MiB (53.2×) |
1,000,000 |
eagle host, termination loop, up to 8 threads |
8 |
194 ms (7.16 ms) |
5.16e9 |
2.27e11 (44 %) |
7.8 |
108 MiB (2.83×) |
1,000,000 |
eagle host, termination loop, 1 thread |
1 |
868 ms (6.43 ms) |
1.15e9 |
5.07e10 (9.9 %) |
1 |
108 MiB (2.82×) |
1,000,000 |
eagle host, termination + compaction, up to 8 threads |
8 |
2.5 s (32.1 ms) |
4.01e8 |
1.76e10 (3.4 %) |
7.7 |
115 MiB (3.01×) |
1,000,000 |
eagle host, termination + compaction, 1 thread |
1 |
3.81 s (30.3 ms) |
2.63e8 |
1.16e10 (2.3 %) |
1 |
115 MiB (3.02×) |
1,000,000 |
eagle host, termination + compaction + reorder, up to 8 threads |
8 |
2.56 s (19 ms) |
3.90e8 |
1.72e10 (3.4 %) |
7.7 |
123 MiB (3.21×) |
1,000,000 |
eagle host, termination + compaction + reorder, 1 thread |
1 |
3.76 s (21.4 ms) |
2.66e8 |
1.17e10 (2.3 %) |
1 |
123 MiB (3.21×) |
1,000,000 |
Numba prange, 8 threads |
8 |
2.17 s (50.4 ms) |
4.60e8 |
2.03e10 (4 %) |
7.4 |
93.4 MiB (2.45×) |
1,000,000 |
JAX sharded over 8 CPU devices |
8 |
4.95 s (2.17 ms) |
2.02e8 |
8.89e9 (1.7 %) |
7.4 |
213 MiB (5.57×) |
1,000,000 |
JAX jit + vmap(while_loop), 1 device |
8 |
9.82 s (111 ms) |
1.02e8 |
4.48e9 (0.87 %) |
2.8 |
145 MiB (3.8×) |
1,000,000 |
PyTorch masked, 8 threads |
8 |
33.8 s (709 ms) |
2.96e7 |
1.30e9 (0.25 %) |
7.7 |
343 MiB (9×) |
1,000,000 |
NumPy masked |
1 |
60.7 s (311 ms) |
1.65e7 |
7.25e8 (0.14 %) |
1 |
191 MiB (5×) |
1,000,000 |
multiprocessing, 8 x NumPy |
8 |
19 s (44.1 ms) |
5.26e7 |
2.32e9 (0.45 %) |
7.3 |
324 MiB (8.5×) |
Fastest arm by median wall time:
N = 1,000: eagle host, termination loop, up to 8 threads (303 µs), then eagle host, termination loop, 1 thread (910 µs, 3.01× the time)
N = 10,000: eagle host, termination loop, up to 8 threads (2.01 ms), then JAX sharded over 8 CPU devices (7.74 ms, 3.86× the time)
N = 100,000: eagle host, termination loop, up to 8 threads (22.5 ms), then eagle host, termination loop, 1 thread (87 ms, 3.86× the time)
N = 1,000,000: eagle host, termination loop, up to 8 threads (194 ms), then eagle host, termination loop, 1 thread (868 ms, 4.48× the time)
Cells not run#
plain Python loops: plain Python runs only at N <= 10,000 with at most 2,500,000 sample-steps per run (spread S=100 N=100,000; spread S=100 N=1,000,000; spread S=1000 N=100,000; spread S=1000 N=1,000,000; uniform S=1000 N=10,000; uniform S=1000 N=100,000; uniform S=1000 N=1,000,000).
Where each tool fits#
Each tool below is written the way its users write it. Two loop structures appear: per step over the batch (eagle’s termination loops, JAX, PyTorch, NumPy, multiprocessing) and per sample over its steps (Numba, plain Python). The per-sample structure touches each sample’s state once for all its steps; the per-step structure streams the batch through memory every step, which is the structure the GPU card’s kernels use. Ratios below are wall-time ratios at the largest N each arm ran (below 1 = less time than eagle at 8 threads).
eagle host, termination loop, up to 8 threads. The GPU card’s eagle graph arm, unchanged, on the host: the same hawk step kernel, eagle.deploy and eagle.until_done call, run by eagle’s OpenMP host team; one code base for GPUs and CPUs. On this card: spread S=100, N=1,000,000: 16.5 ms; spread S=1000, N=1,000,000: 145 ms; uniform S=1000, N=1,000,000: 194 ms. Fastest in 12 of 12 cells.
eagle host, termination loop, 1 thread. The same loop on one thread: the difference to the 8-thread row is the threading alone. On this card: spread S=100, N=1,000,000: 85.5 ms (5.19× eagle 8-thread time); spread S=1000, N=1,000,000: 831 ms (5.74× eagle 8-thread time); uniform S=1000, N=1,000,000: 868 ms (4.48× eagle 8-thread time). Fastest in 0 of 12 cells.
eagle host, termination + compaction, up to 8 threads. The GPU card’s compaction arm on the host: the launch covers only the samples still running, the map rebuilt every 16 steps when a sample finished. On this card: spread S=100, N=1,000,000: 194 ms (11.8× eagle 8-thread time); spread S=1000, N=1,000,000: 1.6 s (11.1× eagle 8-thread time); uniform S=1000, N=1,000,000: 2.5 s (12.9× eagle 8-thread time). Fastest in 0 of 12 cells.
eagle host, termination + compaction, 1 thread. The compaction loop on one thread: the difference to the 8-thread row is the threading alone. On this card: spread S=100, N=1,000,000: 313 ms (19× eagle 8-thread time); spread S=1000, N=1,000,000: 3.04 s (21× eagle 8-thread time); uniform S=1000, N=1,000,000: 3.81 s (19.6× eagle 8-thread time). Fastest in 0 of 12 cells.
eagle host, termination + compaction + reorder, up to 8 threads. The GPU card’s reorder arm on the host: the live samples moved to the front of every plane when they spread thin, restored to sample order at the end (inside the wall). On this card: spread S=100, N=1,000,000: 215 ms (13× eagle 8-thread time); spread S=1000, N=1,000,000: 940 ms (6.49× eagle 8-thread time); uniform S=1000, N=1,000,000: 2.56 s (13.2× eagle 8-thread time). Fastest in 0 of 12 cells.
eagle host, termination + compaction + reorder, 1 thread. The reorder loop on one thread: the difference to the 8-thread row is the threading alone. On this card: spread S=100, N=1,000,000: 260 ms (15.8× eagle 8-thread time); spread S=1000, N=1,000,000: 1.5 s (10.4× eagle 8-thread time); uniform S=1000, N=1,000,000: 3.76 s (19.4× eagle 8-thread time). Fastest in 0 of 12 cells.
Numba prange, 8 threads. Compiles a plain Python loop to machine code with no separate language; each sample keeps its state in registers across all of its steps and stops on its own, so finished samples cost nothing. On this card: spread S=100, N=1,000,000: 49.9 ms (3.03× eagle 8-thread time); spread S=1000, N=1,000,000: 501 ms (3.46× eagle 8-thread time); uniform S=1000, N=1,000,000: 2.17 s (11.2× eagle 8-thread time). Fastest in 0 of 12 cells.
JAX sharded over 8 CPU devices. The JAX arm spread over 8 CPU devices: the batch is sharded and each shard runs its own vmapped while loop, with no cross-shard step synchronisation. On this card: spread S=100, N=1,000,000: 495 ms (30.1× eagle 8-thread time); spread S=1000, N=1,000,000: 4.97 s (34.3× eagle 8-thread time); uniform S=1000, N=1,000,000: 4.95 s (25.5× eagle 8-thread time). Fastest in 0 of 12 cells.
JAX jit + vmap(while_loop), 1 device. Compiles the whole loop once and batches a per-sample while loop with vmap; the same code runs on GPUs and differentiates with jax.grad. On this card: spread S=100, N=1,000,000: 1.01 s (61.3× eagle 8-thread time); spread S=1000, N=1,000,000: 9.84 s (68× eagle 8-thread time); uniform S=1000, N=1,000,000: 9.82 s (50.7× eagle 8-thread time). Fastest in 0 of 12 cells.
PyTorch masked, 8 threads. Runs the masked array version on a multi-threaded tensor library; it fits where the computation already lives next to a PyTorch model or needs autograd. On this card: spread S=100, N=1,000,000: 3.4 s (207× eagle 8-thread time); spread S=1000, N=1,000,000: 33.4 s (231× eagle 8-thread time); uniform S=1000, N=1,000,000: 33.8 s (174× eagle 8-thread time). Fastest in 0 of 12 cells.
NumPy masked. No compile step and no dependency beyond NumPy; masked array code computes every sample every step, so it fits dense batches where all samples run to the end (the uniform configuration). On this card: spread S=100, N=1,000,000: 6.65 s (404× eagle 8-thread time); spread S=1000, N=1,000,000: 62.9 s (434× eagle 8-thread time); uniform S=1000, N=1,000,000: 60.7 s (313× eagle 8-thread time). Fastest in 0 of 12 cells.
multiprocessing, 8 x NumPy. Spreads NumPy over processes with the standard library alone; each chunk stops once its own samples are done. On this card: spread S=100, N=1,000,000: 1.99 s (121× eagle 8-thread time); spread S=1000, N=1,000,000: 19.3 s (133× eagle 8-thread time); uniform S=1000, N=1,000,000: 19 s (98× eagle 8-thread time). Fastest in 0 of 12 cells.
plain Python loops. Nothing to install and each line can be stepped in a debugger; run only at small N here. On this card: spread S=100, N=10,000: 51.4 ms (190× eagle 8-thread time); spread S=1000, N=10,000: 492 ms (208× eagle 8-thread time); uniform S=1000, N=1,000: 222 ms (734× eagle 8-thread time). Fastest in 0 of 12 cells.
Compile time#
Its own pass, never part of any wall. Each measurement is a fresh process: the imports run untimed, then the timer spans the first call to a ready kernel (trace + code generation + compile, or the cache lookup). Cold: an empty cache (hawk’s compile cache pointed at a fresh directory through HAWK_CACHE_DIR; Numba with cache=True into a fresh NUMBA_CACHE_DIR; JAX with its persistent compilation cache in a fresh directory). Warm: a new process on the cache the first cold run populated (hawk cache hit; Numba’s on-disk cache; JAX’s persistent cache with the minimum compile time and entry size set to 0). hawk: eagle.deploy of the kernel (tracing, C++ emission, the g++ host build, eagle’s plan); the active-set variant is a second build of the same step. Numba: the first call on a 64-sample batch (compile plus a negligible run). JAX: lower + compile from abstract shapes, once per N (XLA specialises on the shape). Medians of 3 cold and 3 warm processes.
arm |
compiled |
cold (median of 3) |
warm (median of 3) |
|---|---|---|---|
eagle host, termination loop, up to 8 threads |
hawk step kernel (host build) |
1.42 s |
154 ms |
eagle host, termination + compaction, up to 8 threads |
hawk active-set step kernel (host build) |
1.24 s |
151 ms |
Numba prange, 8 threads |
Numba parallel loop (first call, 64 samples) |
716 ms |
270 ms |
JAX sharded over 8 CPU devices |
XLA executable for N = 1,000 |
140 ms |
40.4 ms |
JAX sharded over 8 CPU devices |
XLA executable for N = 10,000 |
137 ms |
43.8 ms |
JAX sharded over 8 CPU devices |
XLA executable for N = 100,000 |
136 ms |
43 ms |
JAX sharded over 8 CPU devices |
XLA executable for N = 1,000,000 |
155 ms |
42.5 ms |
JAX jit + vmap(while_loop), 1 device |
XLA executable for N = 1,000 |
137 ms |
50.8 ms |
JAX jit + vmap(while_loop), 1 device |
XLA executable for N = 10,000 |
138 ms |
50.7 ms |
JAX jit + vmap(while_loop), 1 device |
XLA executable for N = 100,000 |
155 ms |
52.9 ms |
JAX jit + vmap(while_loop), 1 device |
XLA executable for N = 1,000,000 |
156 ms |
52.1 ms |
No compile step: PyTorch masked, 8 threads, NumPy masked, multiprocessing, 8 x NumPy, plain Python loops.
Same build as another row: eagle host, termination loop, 1 thread (as eagle host, termination loop, up to 8 threads); eagle host, termination + compaction, 1 thread (as eagle host, termination + compaction, up to 8 threads); eagle host, termination + compaction + reorder, up to 8 threads (as eagle host, termination + compaction, up to 8 threads); eagle host, termination + compaction + reorder, 1 thread (as eagle host, termination + compaction, up to 8 threads).
Memory method#
A separate pass after the timing pass. Each (arm, N, configuration) runs in a fresh process: imports, then a warm-up run on a 64-sample batch (every JIT compiles here; JAX also compiles for shape N from abstract shapes), then the baseline (resident set size, VmRSS) and a reset of the process’s peak mark (/proc/self/clear_refs, so ru_maxrss restarts from the baseline); then the batch’s inputs are generated and one full run done. Peak memory = ru_maxrss at the end minus the baseline; for multiprocessing, plus each pool process’s own peak above its baseline (shared pages counted once per process). Minimum = the state the workload needs, N × (2 state + 3 per-sample scalars: omega, zeta, stop step) × 8 bytes; the factor is peak / minimum. Compile overhead = baseline minus the resident size before the arm’s imports-and-warm-up. Below about 1 MiB the allocator reuses memory already resident at the baseline, so small-N peaks can read below the minimum.
Compile/import overhead (resident memory added by the warm-up compile, measured once per arm at the largest N): eagle host, termination loop, up to 8 threads 182 MiB; eagle host, termination loop, 1 thread 182 MiB; eagle host, termination + compaction, up to 8 threads 182 MiB; eagle host, termination + compaction, 1 thread 182 MiB; eagle host, termination + compaction + reorder, up to 8 threads 182 MiB; eagle host, termination + compaction + reorder, 1 thread 182 MiB; Numba prange, 8 threads 109 MiB; JAX sharded over 8 CPU devices 172 MiB; JAX jit + vmap(while_loop), 1 device 169 MiB; PyTorch masked, 8 threads 185 MiB; NumPy masked 11.2 MiB; multiprocessing, 8 x NumPy 12 MiB; plain Python loops 11 MiB.
Reading the results#
Every statement below can be checked against the tables and the “fastest arm” lines above (GPU card unless marked CPU card).
Small batches (N = 10³). Among the GPU arms the automatic-policy arm is fastest in all three configurations (125 µs spread-100, 879 µs spread-1000, 893 µs uniform), with
eagle.simulatenext (130 µs, 884 µs and 896 µs).warp_kernelfollows on spread-100 (199 µs, withhawk + eagle persistentlevel with it at 203 µs) and on spread-1000 (1.38 ms); on uniform the plaineagle_graph(1.32 ms) and the K-steps + compaction arm (1.38 ms) come next, thenwarp_kernelat 1.49 ms. The CPU OpenMP row, which needs no GPU, is faster than every GPU arm on spread-1000 (437 µs againsteagle.simulate’s 884 µs) and on uniform (419 µs against 893 µs), as expected on an FP64-weak card; on spread-100 the automatic-policy arm leads it (125 µs against 198 µs). At the fixed-overhead cell (N = 64, S = 100)eagle.simulatetakes 117 µs andwarp_kernel135 µs.Mid-size batches (N = 10⁴-10⁵). At N = 10⁴ the automatic-policy arm leads spread-100 (340 µs, with
eagle.simulate’s build of the same arm 12 % behind at 381 µs, against Warp’s 719 µs); the automatic-policy arm also leads the GPU arms on spread-1000 (2.54 ms,eagle.simulate2.54 ms, against Warp’s 6.64 ms), with the CPU OpenMP row ahead of it at 1.78 ms. On uniform,eagle.simulate(7.22 ms) andwarp_kernel(7.22 ms) are level; the CPU OpenMP row (2.2 ms) is faster than every GPU arm there. At N = 10⁵ the automatic-policy arm leads spread-100 (2.02 ms,eagle.simulate2.02 ms);eagle.simulateleads the GPU arms on spread-1000 (16.8 ms, against Warp’s 64 ms; the CPU OpenMP row is faster at 15 ms) and uniform (70.8 ms, against Warp’s 71.6 ms, 1.1 % behind); the CPU OpenMP row (19.5 ms) is again faster on uniform.Large batches (N = 10⁶). The automatic-policy arm is the fastest GPU arm on spread-100 (17.8 ms,
eagle.simulate17.9 ms against Warp’s 62.1 ms; the CPU OpenMP row is faster at 15.7 ms) andeagle.simulateties its hand-built twin on spread-1000 (156 ms each, against Warp’s 609 ms; the CPU OpenMP row is faster at 149 ms); on uniformeagle.simulateandwarp_kerneltie (699 ms and 693 ms) and the CPU OpenMP row is faster than both (191 ms). Across every N and configuration in the card,eagle.simulate’s policy arm tracks its hand-built twin – state, params and the stop rule given by name instead of bound into a plan by hand – within 1 % in wall time in every cell but N = 10³ and 10⁴ spread-100 (125 µs vs 130 µs; 340 µs vs 381 µs,eagle.simulate’s per-call host work): the problem-level door costs next to nothing extra.What compaction adds on a thinning batch. On spread-1000, compaction only pays from about N = 10⁵: at N = 10³ and 10⁴
eagle_graph_compactcosts more than plaineagle_graph(6.53 ms vs 1.62 ms at N = 10³; 9.25 ms vs 7.61 ms at N = 10⁴ – the active-set scan has little to recover from at this size), crosses over at N = 10⁵ (45.7 ms vs 66.6 ms, compaction now 1.46× faster) and keeps winning at N = 10⁶ (415 ms vs 625 ms, 1.51×).What compaction costs on a dense batch. On the uniform configuration, where no sample stops early, compaction is overhead with nothing to recover at every N measured: 5.01 ms vs 1.32 ms at N = 10³ (3.8×), 12.8 ms vs 7.74 ms at N = 10⁴ (1.7×), 95.1 ms vs 72.3 ms at N = 10⁵ (1.32×), and 862 ms vs 703 ms at N = 10⁶ (1.23×) – plain
eagle_graphis the better fit for a batch that never thins.What the reorder adds. On spread-1000, the reorder on top of compaction starts paying at N = 10⁵, cutting 45.7 ms to 34.1 ms (1.34×); at N = 10⁶ it cuts 415 ms to 279 ms (1.49×). Below that it costs slightly more than compaction alone (6.65 ms vs 6.53 ms at N = 10³; 9.64 ms vs 9.25 ms at N = 10⁴). On the uniform configuration, with nothing to reorder, it lands within 4 % of compaction alone at every N (863 ms vs 862 ms at N = 10⁶).
The compaction cadence sweep. At N = 10⁶ and K = 16, the cadence-matched map-free arm runs 757 ms against
eagle_graph_compact’s 415 ms (1.82×); the main table’seagle_graphvseagle_graph_compactdifference at the same N is 625 ms vs 415 ms (1.51×). Across K = 8, 16 and 32 the sweep’s ratio climbs from 1.74× to 1.87× at N = 10⁶ and 1.67× to 1.86× at N = 10⁵: the active-set map, not the guard-check cadence, accounts for most of the difference, and compacting less often gives up little of the benefit over this K range.Memory traffic. PyTorch’s CUDA-graph arm reaches the card’s full issued bandwidth at N = 10⁵ uniform (1.40e11 B/s, “1e+02 %” in the table); CuPy masked reaches 87 % of peak at N = 10⁶ uniform. Both arms’ wall time comes from the number of passes they make over the whole batch every step, finished samples included, rather than from slow passes.
Device memory. At N = 10⁶ spread,
eagle.simulate, the plain graph and the persistent arms each hold 50 MiB (1.31× the minimum the planes need) andwarp_kernel64 MiB (1.68×); CuPy masked holds 194 MiB (5.09×) and PyTorch masked 262 MiB (6.87×).Finished samples. In the hawk step a sample that is already finished when a launch starts skips its whole body (no loads, arithmetic or stores), but its warp keeps running while any other lane in it is live; the active-set map in
eagle_graph_compactgoes further and packs the live samples into the first warps, so a finished sample no longer occupies a lane once the map drops it. The bullets above on compaction’s cost and benefit, at the same N on the spread and uniform configurations, show what that is worth and what it costs when there is nothing to pack.CPU card.
eagle host, termination loop, 8 threads(the same hawk step through hawk + eagle’s OpenMP team) is the fastest arm at N = 10⁴ and 10⁵ in all three configurations (270 µs and 1.68 ms spread-100; 2.36 ms and 14.9 ms spread-1000; 2.01 ms and 22.5 ms uniform) and at N = 10⁶ (16.5 ms, 145 ms and 194 ms); at N = 10³ on spread-100 it is also ahead of Numba (55.7 µs against 74.8 µs); Numba’sprangeloop, which keeps each sample’s state in registers across its own steps, is the next-fastest arm at N = 10⁶ on both spread configurations (49.9 ms spread-100; 501 ms spread-1000). CPU compaction does not track the GPU card’s pattern on the thinning spread-1000 workload: at N = 10⁶ it costs 11.0× the plain termination loop (1.6 s vs 145 ms), where the GPU’s compaction instead wins. On the dense uniform workload it costs more than the plain termination loop as well (2.5 s vs 194 ms, 13×) – matching the GPU’s own loss on a batch that never thins, and larger on the CPU.
A second workload: adaptive RK7(8)#
The RK4 oscillator above is a fixed-step problem: every accepted step costs
the same arithmetic, and the four friction points (launch overhead, host
round trips, finished samples, idle lanes) show up against a constant
per-step cost. The second card runs the same class of problem through an
adaptive solver instead: a two-body Kepler orbit (float64, eccentricity
uniform in [0, 0.9]) integrated with Fehlberg’s embedded RKF7(8) pair to
rtol = atol = 1e-10, so each sample also decides, step by step, whether to
accept or shrink – a second, data-dependent source of per-sample
divergence on top of each sample’s own stopping time. eagle.simulate
appears here too, running the same attempt kernel the hand-built hawk + eagle graph arm uses, under its automatic launch policy.
GPU (Quadro P2000)#
Measured on#
device |
compute capability / CPU model |
date |
section |
|---|---|---|---|
Quadro P2000 |
6.1 |
2026-10-08 |
GPU (Quadro P2000)#
Device: Quadro P2000 (compute capability 6.1, 8 SMs). Peak FP64: 94.8 G FLOP/s. Medians of 5 interleaved repetitions (3 for an arm whose warm-up says 5 would take over 60 s); IQR in parentheses.
Workload: two-body Kepler orbits (mu = 1, float64) from periapsis, a = 1, e uniform in [0, 0.9], random orientation; adaptive RKF7(8) (Fehlberg’s tableau, a mixed absolute/relative max-norm error control and a standard step-size controller). Tolerance rtol = atol = 1e-10; final time 64 (10.2 periods). 1106 FLOP per attempt.
arm |
loop structure |
lines of code |
|---|---|---|
eagle graph |
one CUDA graph: device WHILE loop of the attempt kernel, which marks and counts its finished samples and runs the attempts per launch eagle’s policy picks; finished samples skipped by the Terminated guard |
22 |
eagle graph + compaction |
as eagle graph, the attempt launched over the live samples only (index map rebuilt every 16 attempts) |
23 |
eagle graph + compaction + reorder |
as + compaction, live samples moved to the front of every plane when spread thin; restored at the end |
23 |
hawk + eagle persistent |
ONE launch, no graph, no map, no policy kernel: a grid sized from the SM count and the kernel’s occupancy, each lane fetching base + atomicAdd(counter, 1) whenever its sample is done |
22 |
eagle.simulate (eagle picks the launch mode) |
eagle.simulate takes the kernel, its state and its parameters by name; eagle’s automatic policy picks the launch mode from the batch (one launch for a small batch, the persistent launch above it); the row records the mode taken |
21 |
CuPy masked |
host loop: one masked whole-batch attempt (array expressions), host check per attempt |
25 |
PyTorch masked |
host loop: one masked whole-batch attempt (tensor expressions), host check per attempt |
27 |
PyTorch + CUDA graph |
host loop: a CUDA graph of 16 masked attempts replayed, host check per block |
53 |
JAX jit(vmap(while_loop)) |
one compiled executable: per-sample while loop, batched (runs until the slowest sample finishes) |
25 |
Warp per-thread kernel |
one launch: each thread loops over its own sample’s attempts to the final time |
74 |
eagle CPU (OpenMP) |
host loop: one attempt over the batch (OpenMP), the kernel marks and counts its finished samples |
22 |
Lines of code: non-blank, non-comment lines between the # >>> code:NAME and # <<< code:NAME markers of the script; every arm but Warp also uses the shared scheme block (the RKF7(8) stages, the right-hand side, the error norm), counted once in shared_scheme_lines (25 lines).
N |
arm |
wall (IQR) |
kernel-only |
end-to-end |
attempts·samples/s |
useful FLOP/s (% peak) |
device memory (× minimum) |
host memory |
|---|---|---|---|---|---|---|---|---|
1,000 |
eagle graph |
30.7 ms (56.3 µs) |
29.7 ms |
31 ms |
16.4 M |
18.2 G (19%) |
2 MiB (26.2× min.) |
3.25 MiB |
1,000 |
eagle graph + compaction |
34.8 ms (947 µs) |
32.9 ms |
35 ms |
14.5 M |
16.1 G (17%) |
2 MiB (26.2× min.) |
3.25 MiB |
1,000 |
eagle graph + compaction + reorder |
35.4 ms (747 µs) |
34.4 ms |
35.7 ms |
14.2 M |
15.8 G (17%) |
2 MiB (26.2× min.) |
3.25 MiB |
1,000 |
hawk + eagle persistent |
49.1 ms (599 µs) |
49.2 ms |
49.3 ms |
10.3 M |
11.4 G (12%) |
2 MiB (26.2× min.) |
3.25 MiB |
1,000 |
eagle.simulate (eagle picks the launch mode) |
29.7 ms (83.3 µs) |
29.3 ms |
29.9 ms |
17 M |
18.8 G (20%) |
2 MiB (26.2× min.) |
3.25 MiB |
1,000 |
CuPy masked |
8.25 s (486 ms) |
1.88 s |
8.25 s |
61.2 k |
67.7 M (0.071%) |
2 MiB (26.2× min.) |
6.76 MiB |
1,000 |
PyTorch masked |
5.45 s (195 ms) |
2.74 s |
5.45 s |
92.7 k |
103 M (0.11%) |
2 MiB (26.2× min.) |
0.125 MiB |
1,000 |
PyTorch + CUDA graph |
2.96 s (123 µs) |
3.23 s |
2.96 s |
171 k |
189 M (0.2%) |
4 MiB (52.4× min.) |
15.5 MiB |
1,000 |
JAX jit(vmap(while_loop)) |
107 ms (1.21 ms) |
91.6 ms |
108 ms |
4.74 M |
5.24 G (5.5%) |
0 MiB (0× min.) |
0 MiB |
1,000 |
Warp per-thread kernel |
55.3 ms (346 µs) |
55.4 ms |
55.9 ms |
9.14 M |
10.1 G (11%) |
32 MiB (419× min.) |
0.125 MiB |
1,000 |
eagle CPU (OpenMP) |
19.5 ms (180 µs) |
– |
19.6 ms |
25.9 M |
28.6 G (–) |
– |
0.25 MiB |
10,000 |
eagle graph |
246 ms (145 µs) |
245 ms |
247 ms |
20.1 M |
22.2 G (23%) |
2 MiB (2.62× min.) |
5.25 MiB |
10,000 |
eagle graph + compaction |
162 ms (18.3 µs) |
162 ms |
163 ms |
30.5 M |
33.8 G (36%) |
2 MiB (2.62× min.) |
5.38 MiB |
10,000 |
eagle graph + compaction + reorder |
163 ms (177 µs) |
162 ms |
164 ms |
30.4 M |
33.6 G (35%) |
2 MiB (2.62× min.) |
5.62 MiB |
10,000 |
hawk + eagle persistent |
160 ms (522 µs) |
159 ms |
161 ms |
31 M |
34.3 G (36%) |
2 MiB (2.62× min.) |
5 MiB |
10,000 |
eagle.simulate (eagle picks the launch mode) |
160 ms (759 µs) |
160 ms |
161 ms |
31 M |
34.2 G (36%) |
2 MiB (2.62× min.) |
5.12 MiB |
10,000 |
CuPy masked |
8.11 s (219 ms) |
2.85 s |
8.11 s |
611 k |
676 M (0.71%) |
2 MiB (2.62× min.) |
7.5 MiB |
10,000 |
PyTorch masked |
5.63 s (311 ms) |
3.33 s |
5.65 s |
881 k |
974 M (1%) |
2 MiB (2.62× min.) |
2.25 MiB |
10,000 |
PyTorch + CUDA graph |
3.59 s (607 µs) |
3.8 s |
3.59 s |
1.38 M |
1.53 G (1.6%) |
12 MiB (15.7× min.) |
16.4 MiB |
10,000 |
JAX jit(vmap(while_loop)) |
427 ms (1.01 ms) |
417 ms |
429 ms |
11.6 M |
12.8 G (14%) |
0 MiB (0× min.) |
0.125 MiB |
10,000 |
Warp per-thread kernel |
274 ms (1.12 ms) |
275 ms |
275 ms |
18.1 M |
20 G (21%) |
32 MiB (41.9× min.) |
2 MiB |
10,000 |
eagle CPU (OpenMP) |
82.9 ms (3.34 ms) |
– |
83.1 ms |
59.8 M |
66.2 G (–) |
– |
2.63 MiB |
100,000 |
eagle graph |
2.36 s (2.18 ms) |
2.39 s |
2.37 s |
21.1 M |
23.3 G (25%) |
10 MiB (1.31× min.) |
25.8 MiB |
100,000 |
eagle graph + compaction |
1.43 s (15.1 ms) |
1.46 s |
1.44 s |
34.7 M |
38.4 G (41%) |
12 MiB (1.57× min.) |
26 MiB |
100,000 |
eagle graph + compaction + reorder |
1.43 s (14.8 ms) |
1.45 s |
1.43 s |
34.8 M |
38.5 G (41%) |
14 MiB (1.84× min.) |
25.4 MiB |
100,000 |
hawk + eagle persistent |
1.32 s (15.2 ms) |
1.34 s |
1.33 s |
37.6 M |
41.6 G (44%) |
10 MiB (1.31× min.) |
25.9 MiB |
100,000 |
eagle.simulate (eagle picks the launch mode) |
1.32 s (15.8 ms) |
1.34 s |
1.33 s |
37.7 M |
41.6 G (44%) |
10 MiB (1.31× min.) |
26 MiB |
100,000 |
CuPy masked |
19.1 s (7.98 ms) |
19.3 s |
19.1 s |
2.61 M |
2.88 G (3%) |
10 MiB (1.31× min.) |
22.4 MiB |
100,000 |
PyTorch masked |
17.3 s (13.1 ms) |
17.1 s |
17.4 s |
2.87 M |
3.18 G (3.4%) |
10 MiB (1.31× min.) |
20.7 MiB |
100,000 |
PyTorch + CUDA graph |
17.2 s (8.63 ms) |
17.4 s |
17.2 s |
2.9 M |
3.21 G (3.4%) |
116 MiB (15.2× min.) |
33.4 MiB |
100,000 |
JAX jit(vmap(while_loop)) |
4.2 s (809 µs) |
4.24 s |
4.21 s |
11.9 M |
13.1 G (14%) |
32 MiB (4.19× min.) |
26.7 MiB |
100,000 |
Warp per-thread kernel |
2.64 s (2.61 ms) |
2.67 s |
2.64 s |
18.9 M |
20.9 G (22%) |
32 MiB (4.19× min.) |
19.3 MiB |
100,000 |
eagle CPU (OpenMP) |
713 ms (18.5 ms) |
– |
715 ms |
69.9 M |
77.3 G (–) |
– |
25.6 MiB |
1,000,000 |
eagle graph |
23.5 s (3.31 ms) |
23.8 s |
23.7 s |
21.2 M |
23.4 G (25%) |
80 MiB (1.05× min.) |
232 MiB |
1,000,000 |
eagle graph + compaction |
14.2 s (2.86 ms) |
14.4 s |
14.4 s |
35.2 M |
38.9 G (41%) |
96 MiB (1.26× min.) |
232 MiB |
1,000,000 |
eagle graph + compaction + reorder |
14.1 s (3.24 ms) |
14.4 s |
14.2 s |
35.3 M |
39 G (41%) |
124 MiB (1.63× min.) |
232 MiB |
1,000,000 |
hawk + eagle persistent |
13 s (1.5 ms) |
13.1 s |
13 s |
38.4 M |
42.5 G (45%) |
80 MiB (1.05× min.) |
232 MiB |
1,000,000 |
eagle.simulate (eagle picks the launch mode) |
13 s (1.35 ms) |
13.1 s |
13 s |
38.4 M |
42.5 G (45%) |
80 MiB (1.05× min.) |
232 MiB |
1,000,000 |
CuPy masked |
182 s (3.22 ms) |
183 s |
182 s |
2.74 M |
3.02 G (3.2%) |
80 MiB (1.05× min.) |
182 MiB |
1,000,000 |
PyTorch masked |
184 s (16.9 ms) |
183 s |
184 s |
2.72 M |
3 G (3.2%) |
100 MiB (1.31× min.) |
191 MiB |
1,000,000 |
PyTorch + CUDA graph |
184 s (2.02 ms) |
185 s |
184 s |
2.7 M |
2.99 G (3.2%) |
1.17e+03 MiB (15.3× min.) |
192 MiB |
1,000,000 |
JAX jit(vmap(while_loop)) |
40.2 s (9.7 ms) |
40.3 s |
40.3 s |
12.4 M |
13.7 G (14%) |
256 MiB (3.36× min.) |
307 MiB |
1,000,000 |
Warp per-thread kernel |
26.3 s (10.2 ms) |
26.3 s |
26.4 s |
18.9 M |
20.9 G (22%) |
96 MiB (1.26× min.) |
191 MiB |
1,000,000 |
eagle CPU (OpenMP) |
7.15 s (49.4 ms) |
– |
7.17 s |
69.7 M |
77.1 G (–) |
– |
254 MiB |
Agreement with the analytic orbit#
Every arm, every sample: final time reached exactly; position error against the analytic Kepler state (Newton on Kepler’s equation to machine precision) below the stated bound n_acc tol (1 + v_p)(1 + 6 pi n_orbits v_p), derived from the controller’s per-step budget and the Kepler along-track amplification (see truth_bound; not fitted). Step counts are compared with the reference arm (eagle graph): any difference comes from a floating-point difference flipping one accept/reject decision (contraction into fused multiply-adds, pow and sqrt implementations differ between compilers and libraries); the table reports how many samples differ and by how much.
N |
arm |
max position error |
max error / bound |
samples with different step counts |
max Δ accepted / rejected |
|---|---|---|---|---|---|
1,000 |
eagle graph |
1.85e-07 |
6.35e-03 |
0 |
0 / 0 |
1,000 |
eagle graph + compaction |
1.85e-07 |
6.35e-03 |
0 |
0 / 0 |
1,000 |
eagle graph + compaction + reorder |
1.85e-07 |
6.35e-03 |
0 |
0 / 0 |
1,000 |
hawk + eagle persistent |
1.85e-07 |
6.35e-03 |
0 |
0 / 0 |
1,000 |
eagle.simulate (eagle picks the launch mode) |
1.85e-07 |
6.35e-03 |
0 |
0 / 0 |
1,000 |
CuPy masked |
1.85e-07 |
6.35e-03 |
0 |
0 / 0 |
1,000 |
PyTorch masked |
1.85e-07 |
6.35e-03 |
0 |
0 / 0 |
1,000 |
PyTorch + CUDA graph |
1.85e-07 |
6.35e-03 |
0 |
0 / 0 |
1,000 |
JAX jit(vmap(while_loop)) |
1.85e-07 |
6.35e-03 |
0 |
0 / 0 |
1,000 |
Warp per-thread kernel |
1.85e-07 |
6.35e-03 |
0 |
0 / 0 |
1,000 |
eagle CPU (OpenMP) |
1.85e-07 |
6.35e-03 |
0 |
0 / 0 |
10,000 |
eagle graph |
2.02e-07 |
6.57e-03 |
0 |
0 / 0 |
10,000 |
eagle graph + compaction |
2.02e-07 |
6.57e-03 |
0 |
0 / 0 |
10,000 |
eagle graph + compaction + reorder |
2.02e-07 |
6.57e-03 |
0 |
0 / 0 |
10,000 |
hawk + eagle persistent |
2.02e-07 |
6.57e-03 |
0 |
0 / 0 |
10,000 |
eagle.simulate (eagle picks the launch mode) |
2.02e-07 |
6.57e-03 |
0 |
0 / 0 |
10,000 |
CuPy masked |
2.02e-07 |
6.57e-03 |
0 |
0 / 0 |
10,000 |
PyTorch masked |
2.02e-07 |
6.57e-03 |
0 |
0 / 0 |
10,000 |
PyTorch + CUDA graph |
2.02e-07 |
6.57e-03 |
0 |
0 / 0 |
10,000 |
JAX jit(vmap(while_loop)) |
2.02e-07 |
6.57e-03 |
0 |
0 / 0 |
10,000 |
Warp per-thread kernel |
2.02e-07 |
6.57e-03 |
0 |
0 / 0 |
10,000 |
eagle CPU (OpenMP) |
2.02e-07 |
6.57e-03 |
0 |
0 / 0 |
100,000 |
eagle graph |
2.25e-07 |
6.62e-03 |
0 |
0 / 0 |
100,000 |
eagle graph + compaction |
2.25e-07 |
6.62e-03 |
0 |
0 / 0 |
100,000 |
eagle graph + compaction + reorder |
2.25e-07 |
6.62e-03 |
0 |
0 / 0 |
100,000 |
hawk + eagle persistent |
2.25e-07 |
6.62e-03 |
0 |
0 / 0 |
100,000 |
eagle.simulate (eagle picks the launch mode) |
2.25e-07 |
6.62e-03 |
0 |
0 / 0 |
100,000 |
CuPy masked |
2.25e-07 |
6.62e-03 |
0 |
0 / 0 |
100,000 |
PyTorch masked |
2.25e-07 |
6.62e-03 |
0 |
0 / 0 |
100,000 |
PyTorch + CUDA graph |
2.25e-07 |
6.62e-03 |
0 |
0 / 0 |
100,000 |
JAX jit(vmap(while_loop)) |
2.25e-07 |
6.62e-03 |
0 |
0 / 0 |
100,000 |
Warp per-thread kernel |
2.25e-07 |
6.62e-03 |
0 |
0 / 0 |
100,000 |
eagle CPU (OpenMP) |
2.25e-07 |
6.62e-03 |
0 |
0 / 0 |
1,000,000 |
eagle graph |
2.26e-07 |
6.64e-03 |
0 |
0 / 0 |
1,000,000 |
eagle graph + compaction |
2.26e-07 |
6.64e-03 |
0 |
0 / 0 |
1,000,000 |
eagle graph + compaction + reorder |
2.26e-07 |
6.64e-03 |
0 |
0 / 0 |
1,000,000 |
hawk + eagle persistent |
2.26e-07 |
6.64e-03 |
0 |
0 / 0 |
1,000,000 |
eagle.simulate (eagle picks the launch mode) |
2.26e-07 |
6.64e-03 |
0 |
0 / 0 |
1,000,000 |
CuPy masked |
2.26e-07 |
6.64e-03 |
4 |
1 / 2 |
1,000,000 |
PyTorch masked |
2.26e-07 |
6.64e-03 |
2 |
1 / 2 |
1,000,000 |
PyTorch + CUDA graph |
2.26e-07 |
6.64e-03 |
2 |
1 / 2 |
1,000,000 |
JAX jit(vmap(while_loop)) |
2.26e-07 |
6.64e-03 |
4 |
1 / 2 |
1,000,000 |
Warp per-thread kernel |
2.26e-07 |
6.64e-03 |
2 |
1 / 2 |
1,000,000 |
eagle CPU (OpenMP) |
2.26e-07 |
6.64e-03 |
2 |
1 / 1 |
Steps per sample at N = 1,000,000 (reference arm): accepted min 322 / median 371 / p90 571 / max 693; rejected min 1 / median 59 / max 264.
Where each tool fits#
Each tool here solves the same 13-stage adaptive problem to the same tolerance and agrees with the analytic orbit; they differ in how the per-sample control flow is written and run. Numbers below are at N = 1,000,000; the fastest arm there is eagle.simulate (eagle picks the launch mode) (13 s).
eagle graph. the attempt is per-sample scalar code (22 lines with its launch; the kernel finishes its own samples), one kernel per attempt looped on the device: 23.5 s.
eagle graph + compaction. the same kernel launched over the live samples only: 14.2 s; the plain graph takes 1.66× as long.
eagle graph + compaction + reorder. adds the occasional reorder of the planes: 14.1 s; the compacted graph takes 1× as long.
hawk + eagle persistent. the plain (not active-set) attempt kernel, forced through hawk’s persist entry: one launch, no graph, no map, no policy kernel – a grid of SMs blocks of 256 lanes steals the next unfinished sample off a counter when their own finishes: 13 s; the eagle graph takes 1.81× as long.
eagle.simulate (eagle picks the launch mode). the same kernel through eagle.simulate, its state and parameters by name (21 lines with the kernel): 13 s; the eagle graph takes 1.81× as long.
CuPy masked. array code in the NumPy style (25 lines), the control flow as masks over the batch: 182 s.
PyTorch masked. the same array formulation in PyTorch (27 lines), running where a PyTorch model already runs: 184 s.
PyTorch + CUDA graph. records 16 attempts as a CUDA graph and replays them, removing per-kernel launch cost: 184 s; eager PyTorch takes 0.995× as long.
JAX jit(vmap(while_loop)). writes the per-sample loop directly (25 lines) and compiles the whole batch into one executable: 40.2 s.
Warp per-thread kernel. an explicit per-thread loop in a Python-embedded kernel language (74 lines, its own stages): 26.3 s.
eagle CPU (OpenMP). the same hawk kernel on the host’s cores, no GPU needed: 7.15 s.
Compile time#
Its own pass, never part of any wall. Each measurement is a fresh process: imports and the GPU context run untimed, then the timer spans the first call to a ready kernel (trace + code generation + compile + load, or the cache lookup). Cold: every cache empty (HAWK_CACHE_DIR, CUPY_CACHE_DIR, CUDA_CACHE_PATH, WARP_CACHE_PATH and JAX’s persistent compilation cache, each a fresh directory); warm: a new process on the caches the first cold run populated. eagle GPU: eagle.deploy of the hawk attempt kernel and of its active-set variant (host and device builds each), and one eagle graph of each built and run on 64 samples (eagle.simulate deploys the first again, a cache hit). eagle CPU: the same eagle.deploy of the attempt kernel (host and device builds) and one host run. CuPy: one masked attempt on 64 samples (every elementwise kernel compiled on first use). JAX: lower + compile from abstract shapes at N. Warp: the kernel module (code generation, compile, load) and one launch. Medians of 3 cold and 3 warm processes.
arm |
compiled |
cold (median of 3) |
warm (median of 3) |
|---|---|---|---|
eagle graph, eagle graph + compaction, eagle graph + compaction + reorder, hawk + eagle persistent, eagle.simulate (eagle picks the launch mode) |
eagle.deploy of the attempt kernel and its active-set variant |
3.36 s |
527 ms |
eagle CPU (OpenMP) |
eagle.deploy of the attempt kernel, host run |
2.62 s |
302 ms |
CuPy masked |
CuPy elementwise kernels |
2.71 s |
1.14 s |
JAX jit(vmap(while_loop)) |
XLA executable, N = 1,000 |
3.06 s |
281 ms |
JAX jit(vmap(while_loop)) |
XLA executable, N = 10,000 |
2.65 s |
279 ms |
JAX jit(vmap(while_loop)) |
XLA executable, N = 100,000 |
3.47 s |
290 ms |
JAX jit(vmap(while_loop)) |
XLA executable, N = 1,000,000 |
3.32 s |
282 ms |
Warp per-thread kernel |
Warp kernel module |
9.82 s |
62.8 ms |
No compile step: PyTorch masked, PyTorch + CUDA graph (PyTorch’s eager kernels ship prebuilt; the CUDA graph is recorded, not compiled).
Memory method#
A separate pass after the timing pass. Each (arm, N) runs in a fresh process: imports, the arm’s kernels built and a warm-up run on a 64-sample batch (every compile happens here; JAX also compiles for shape N), every caching allocator trimmed (CuPy’s pools, PyTorch’s cache), then the baseline read; then the arm built for N (eagle’s graph, PyTorch’s captured graph) and one full run with upload and download; then the reading again, the allocators still holding their high-water reservation (the same measure for every arm, no sampling needed). Device: the memory NVML accounts to the process on the timed GPU, after minus baseline. Caveats: each allocator’s rounding and growth policy count (JAX’s allocator grows in regions that can exceed the request; PyTorch rounds blocks up and a captured CUDA graph keeps a private pool; Warp’s stream-ordered pool reserves in chunks of tens of MiB, its release threshold set to 0 before the baseline); the driver accounts in pages of about 2 MiB, so small-N rows read 0 or one page. Host: the peak resident set (VmHWM, reset to the current RSS at the baseline) minus the baseline RSS. Minimum: the planes the workload needs, N x 10 (state, time, step, two counters) x 8 bytes.
torch.compile is not run: its GPU code generator (Inductor, through Triton) needs compute capability 7.0 or newer, and this GPU is 6.1. The PyTorch + CUDA graph arm removes the same per-kernel launch cost without generating code.
CPU (Intel Xeon W-2125)#
Measured on#
device |
compute capability / CPU model |
date |
section |
|---|---|---|---|
Intel Xeon W-2125 CPU @ 4.00GHz |
Intel(R) Xeon(R) W-2125 CPU @ 4.00GHz |
2026-10-08 |
CPU (Intel Xeon W-2125 CPU @ 4.00GHz)#
CPU: Intel(R) Xeon(R) W-2125 CPU @ 4.00GHz, 8 logical CPUs. Medians of 5 interleaved repetitions; IQR in parentheses. Workload as the GPU card (1106 FLOP per attempt); masked arms up to N = 100,000.
N |
arm |
threads |
wall (IQR) |
attempts·samples/s |
useful FLOP/s |
host memory |
compile (cold) |
lines of code |
|---|---|---|---|---|---|---|---|---|
1,000 |
eagle CPU, 8 threads |
8 |
19.7 ms (284 µs) |
25.6 M |
28.4 G |
0 MiB |
2.72 s |
22 |
1,000 |
eagle CPU, 1 thread |
1 |
41.6 ms (2.62 ms) |
12.1 M |
13.4 G |
0 MiB |
2.72 s |
22 |
1,000 |
Numba prange |
8 |
71.9 ms (7.88 ms) |
7.02 M |
7.76 G |
0 MiB |
3.55 s |
58 |
1,000 |
JAX jit(vmap(while_loop)) |
8 |
204 ms (1.22 ms) |
2.48 M |
2.74 G |
0.5 MiB |
2.15 s |
25 |
1,000 |
PyTorch masked |
8 |
3.88 s (42.8 ms) |
130 k |
144 M |
1.12 MiB |
– |
27 |
1,000 |
NumPy masked |
1 |
1.1 s (9.97 ms) |
460 k |
509 M |
0.875 MiB |
– |
23 |
10,000 |
eagle CPU, 8 threads |
8 |
85.2 ms (1.03 ms) |
58.2 M |
64.4 G |
1.38 MiB |
2.66 s |
22 |
10,000 |
eagle CPU, 1 thread |
1 |
346 ms (1.33 ms) |
14.3 M |
15.9 G |
1 MiB |
2.78 s |
22 |
10,000 |
Numba prange |
8 |
594 ms (32.1 ms) |
8.35 M |
9.23 G |
0.75 MiB |
3.53 s |
58 |
10,000 |
JAX jit(vmap(while_loop)) |
8 |
1.16 s (18.4 ms) |
4.28 M |
4.74 G |
3.38 MiB |
2.13 s |
25 |
10,000 |
PyTorch masked |
8 |
10.6 s (1.81 s) |
466 k |
515 M |
15.7 MiB |
– |
27 |
10,000 |
NumPy masked |
1 |
9.96 s (3.74 s) |
498 k |
550 M |
8.4 MiB |
– |
23 |
100,000 |
eagle CPU, 8 threads |
8 |
698 ms (9.15 ms) |
71.4 M |
79 G |
15.1 MiB |
2.77 s |
22 |
100,000 |
eagle CPU, 1 thread |
1 |
3.44 s (3.23 ms) |
14.5 M |
16 G |
14.9 MiB |
2.78 s |
22 |
100,000 |
Numba prange |
8 |
5.71 s (23.2 ms) |
8.72 M |
9.65 G |
11.7 MiB |
3.56 s |
58 |
100,000 |
JAX jit(vmap(while_loop)) |
8 |
11.3 s (63.2 ms) |
4.41 M |
4.87 G |
31.8 MiB |
2.16 s |
25 |
100,000 |
PyTorch masked |
8 |
60.8 s (8.59 s) |
819 k |
906 M |
152 MiB |
– |
27 |
100,000 |
NumPy masked |
1 |
103 s (1.02 s) |
485 k |
537 M |
85.4 MiB |
– |
23 |
1,000,000 |
eagle CPU, 8 threads |
8 |
6.82 s (17.8 ms) |
73 M |
80.8 G |
153 MiB |
2.74 s |
22 |
1,000,000 |
eagle CPU, 1 thread |
1 |
34.4 s (48.5 ms) |
14.5 M |
16 G |
153 MiB |
2.78 s |
22 |
1,000,000 |
Numba prange |
8 |
56.5 s (131 ms) |
8.81 M |
9.75 G |
122 MiB |
3.58 s |
58 |
1,000,000 |
JAX jit(vmap(while_loop)) |
8 |
122 s (317 ms) |
4.07 M |
4.5 G |
314 MiB |
2.1 s |
25 |
N |
arm |
max position error |
max error / bound |
samples with different step counts |
|---|---|---|---|---|
1,000 |
eagle CPU, 8 threads |
1.85e-07 |
6.35e-03 |
0 |
1,000 |
eagle CPU, 1 thread |
1.85e-07 |
6.35e-03 |
0 |
1,000 |
Numba prange |
1.85e-07 |
6.35e-03 |
0 |
1,000 |
JAX jit(vmap(while_loop)) |
1.85e-07 |
6.35e-03 |
0 |
1,000 |
PyTorch masked |
1.85e-07 |
6.35e-03 |
0 |
1,000 |
NumPy masked |
1.85e-07 |
6.35e-03 |
0 |
10,000 |
eagle CPU, 8 threads |
2.02e-07 |
6.57e-03 |
0 |
10,000 |
eagle CPU, 1 thread |
2.02e-07 |
6.57e-03 |
0 |
10,000 |
Numba prange |
2.02e-07 |
6.57e-03 |
0 |
10,000 |
JAX jit(vmap(while_loop)) |
2.02e-07 |
6.57e-03 |
0 |
10,000 |
PyTorch masked |
2.02e-07 |
6.57e-03 |
0 |
10,000 |
NumPy masked |
2.02e-07 |
6.57e-03 |
0 |
100,000 |
eagle CPU, 8 threads |
2.25e-07 |
6.62e-03 |
0 |
100,000 |
eagle CPU, 1 thread |
2.25e-07 |
6.62e-03 |
0 |
100,000 |
Numba prange |
2.25e-07 |
6.62e-03 |
0 |
100,000 |
JAX jit(vmap(while_loop)) |
2.25e-07 |
6.62e-03 |
0 |
100,000 |
PyTorch masked |
2.25e-07 |
6.62e-03 |
0 |
100,000 |
NumPy masked |
2.25e-07 |
6.62e-03 |
0 |
1,000,000 |
eagle CPU, 8 threads |
2.26e-07 |
6.64e-03 |
0 |
1,000,000 |
eagle CPU, 1 thread |
2.26e-07 |
6.64e-03 |
0 |
1,000,000 |
Numba prange |
2.26e-07 |
6.64e-03 |
2 |
1,000,000 |
JAX jit(vmap(while_loop)) |
2.26e-07 |
6.64e-03 |
4 |
Reading the RK7(8) results#
Fastest arm. At N = 1,000,000
eagle.simulateandhawk + eagle persistentare level on the GPU card (13.0 s each); plainhawk + eagle graphtakes 1.81× as long (23.5 s). Compaction takes the graph to 14.2 s (1.66× faster), and the reorder adds nothing beyond it (14.1 s, 1× the compaction arm’s time).eagle.simulatehere.eagle.simulateruns the same attempt kernel under eagle’s automatic policy, which picks the launch mode from the batch (one launch for a small batch, the persistent launch above it; the card records the mode taken). From N = 10,000 up it tieshawk + eagle persistent(160 ms vs 160 ms at N = 10,000; 13.0 s each at N = 1,000,000); at N = 1,000 it takes the single launch (29.7 ms against 49.1 ms for the persistent arm and 30.7 ms for the plain graph). Compile time is its own pass: a cold run of the eagle arms (they all deploy the same hawk attempt kernel and its active-set variant) costs 3.36 s, falling to 527 ms warm.GPU vs. CPU. This workload is where the CPU wins outright at every N:
hawk + eagle CPU (OpenMP)reaches 7.15 s at N = 1,000,000 on the GPU card’s own CPU reference row – about twice as fast as every GPU arm, including the 13.0 seagle.simulate– and the dedicated CPU card confirms it independently at 6.82 s (hawk + eagle CPU, 8 threads). (The RK4 dense batch above shows a milder version of the same effect.) An adaptive, branch-heavy per-sample attempt with only 1106 FLOP leaves little for the GPU’s wider lanes to amortize against its own launch and active-set overhead, on a card whose FP64 throughput is modest; Numba’sprangeloop is the fastest other CPU arm at 8.81 M attempts-samples/s (N = 1,000,000), and JAX’s compiledvmap(while_loop)is the fastest arm that also differentiates and runs on a GPU unmodified (40.2 s on this GPU card, 122 s on the CPU card). NVIDIA Warp’s per-thread kernel takes 26.3 s on the GPU.Agreement. Every arm reaches the same final time and matches the analytic Kepler state within the card’s derived bound at every N; only at N = 1,000,000 do a handful of samples (at most 4 of a million) take a different accept/reject path, from a floating-point difference in how each framework contracts or rounds
pow/sqrt, not from a tolerance violation.
Reproduce#
From the eagle repository root, in an environment with eagle, hawk, CuPy, a CUDA toolchain and (optionally) Nsight Systems:
python benchmarks/perf_card/perf_card.py # RK4 oscillators, GPU
python benchmarks/perf_card/cpu_card.py # RK4 oscillators, CPU
python benchmarks/rk78_card/rk78_card.py # RK7(8) Kepler orbits, GPU
python benchmarks/rk78_card/cpu_rk78_card.py # RK7(8) Kepler orbits, CPU
Each run builds the hawk kernel(s), verifies the arms against each other,
times them, and writes its own card_<device>.json together with the
card_<device>.md table shown above. The JSON carries the md5 of the script
that produced it; a test re-renders the table from the JSON so the two
cannot drift. --quick runs a reduced matrix as a smoke test.
To reproduce the cards on another machine (Kaggle, a rented GPU or your own box) with the same pinned software, see benchmarks/reproduce/.
Adding a device. Drop the new card_<slug>.json/.md (GPU) or
cpu_card_<slug>.json/.md (CPU) next to the existing ones under
benchmarks/perf_card/ or benchmarks/rk78_card/ – benchmarks/reproduce/run_cards.sh --publish writes them there directly. Rebuilding the docs (make html or make strict) then regenerates this page’s “Measured on” tables and per-device
sections on its own; re-running a card for a device already listed replaces
that device’s rows the same way. No other file on this page needs editing.
Hardware caveat#
The GPU cards above come from a Quadro P2000, a development, consumer-class card whose FP64 throughput (1:32 of its FP32 rate) and memory bandwidth are recorded in each card, and from a Kaggle Tesla T4; the CPU cards come from an Intel Xeon W-2125 workstation chip (4 cores, 8 logical CPUs). The method transfers to other hardware unchanged; the absolute numbers, and possibly the crossover points between arms, will differ on data-centre GPUs or larger CPUs, whose FP64 throughput, bandwidth and core counts differ from these cards’. A card from such a device would be added beside these, produced by the same scripts.