eagle

Contents

eagle#

eagle — the Python face of the EAGLE launch engine.

Which call do I use?

one kernel (or a step’s worth), run until every sample finishes – simulate() a plan already paced by its own finish kernel, one call – until_done() build a kernel into a plan that runs where its data lives – deploy() / plan resolve a launchable BY NAME, then call/stream/graph it – KernelRegistry

EAGLE owns everything about executing compiled kernels: the framework-polymorphic launch skeleton (LaunchMixin), the capturable kind-dispatched launch() primitive, input-role marshalling (eagle.marshal), the by-value launch ABI (eagle.abi), DLPack framework adapters (eagle.interop), CUDA-graph capture/replay (GraphPipeline), composing many already-built launchables into one launch list (GraphComposer), and the plugin protocol consumers — KernelRegistry / load_manifest() on the Python side, mirroring the C++ host/ registries.

Producers of launchable artifacts (compiled CUDA sources, PTX bundles with sidecar/manifest metadata) live upstream — the code generator is the first — and depend on this package; the eagle core never imports a producer. The optional framework bridges (eagle.frameworks) are the one exception: they import hawk lazily, on first use, and so does eagle.deploy() (the same function as eagle.plan.auto()), which builds a hawk kernel into a plan that runs where its data lives. eagle.simulate() runs a model (one hawk kernel, or the kernels of one step) on every sample until each one finishes, called exactly like the kernel itself.

class eagle.ActiveSet(mask, *, reorder=None, _lazy_scratch=False, _lazy_map=False)[source]#

Bases: object

The live samples of mask (a terminated plane: set = finished), as an ascending index map plus a count – and, with reorder=theta, the occasional physical reorder of the planes it owns. mask is a one-dimensional, C-contiguous bool/uint8 array (cupy or numpy), held by reference and read where it is every time. The map, count and device scratch are allocated once, before any capture, and live as long as this object. reorder is an explicit opt-in: the locality threshold theta (in THETA_RANGE, or True for DEFAULT_THETA) that pays on large, irregularly thinning batches (from about 1e5 samples; below that the map alone is as fast). The set then owns the mask and every plane given to own(). The set starts as the identity (every sample live, count == n).

mask#
n#
map#
count#
theta#
perm#
inv#
span#
fire#
live32#
reorders#
__call__()#

Recompute the map and the count from the mask (and, for a reordering set, live32/fire). Device: enqueued on the current cupy stream (capturable). Host: computed now.

Return type:

None

compact()[source]#

Recompute the map and the count from the mask (and, for a reordering set, live32/fire). Device: enqueued on the current cupy stream (capturable). Host: computed now.

Return type:

None

in_sample_order(plane)[source]#

A copy of plane in sample order (plane[..., inv]); legal while permuted, capture-legal on the device.

indirect(*planes)[source]#

Declare bound per-sample planes the caller keeps in sample order (its kernel reads sample perm[slot] from them). Returns the set.

Return type:

ActiveSet

property live: int#

The live count as a Python int (a device sync for a device set).

materialize_map()[source]#

Allocate the deferred map (see _lazy_map) and set it to the identity. True when it was deferred, so the caller must re-point every bound plan at map; False when it already is real.

Return type:

bool

own(*planes)[source]#

Hand the set per-sample planes to move with the mask at every reorder (a (D, n) array is D planes). Refused once fixed, for a plane another set owns, or one overlapping an owned plane.

Return type:

ActiveSet

property permuted: bool#

Whether a reorder ran since the last restore() / reset() (a device sync for a device set).

planes()[source]#

The two reserved planes as eagle.plan.Plan.bind() keywords.

Return type:

dict

prepare_compaction()[source]#

Allocate the device compaction scratch (once). The first compact() does it when needed; anything that captures a compaction into a graph calls it first. A runner on a no-compaction fast path never pays for it.

Return type:

None

reorder()[source]#

Run the reorder now if the last compaction fired it: owned planes and perm move so live samples come first, inv/the map reset, span = count, reorders += 1. Device: enqueued on the current stream; host: now.

Return type:

None

reorder_if_degraded()[source]#

The reorder as a eagle.Skippable guarded by fire (one IF node under a graph, an eager check elsewhere).

reset()[source]#

Back to the identity: every sample live, count == n (and, for a reordering set, perm/inv identity, span == n, reorder counter zero). Not capture-legal: a plain array write, run before a launch, never recorded.

Return type:

None

restore()[source]#

Put every owned plane back in sample order; perm/inv reset to identity, span = n, the map recomputed. Eager only (refused under stream capture); a no-op for a set that does not reorder.

Return type:

None

when_finished(finished)[source]#

A eagle.Skippable compaction that runs only when finished moved since it last ran (the guard is evaluated on the device under a graph).

exception eagle.BackendUnavailable[source]#

Bases: RuntimeError

The device backend cannot serve this call (see the message for why).

A RuntimeError, so code that already guards a device call with except RuntimeError keeps working.

class eagle.DeviceProps(name, cc_major, cc_minor, sm_count, clock_rate_khz, memory_clock_rate_khz, memory_bus_width_bits, shared_mem_per_block, shared_mem_per_sm, regs_per_block, regs_per_sm, warp_size, peak_bytes_per_s, peak_flops_sp, peak_flops_dp, fp64_ratio, ridge_flops_per_byte_sp, ridge_flops_per_byte_dp)[source]#

Bases: object

Frozen view of one CUDA device’s raw + derived properties.

Field set mirrors eagle::DeviceProps (eagle/DeviceProps.h) exactly — see that header for every formula’s one-line derivation.

Parameters:
  • name (str)

  • cc_major (int)

  • cc_minor (int)

  • sm_count (int)

  • clock_rate_khz (int)

  • memory_clock_rate_khz (int)

  • memory_bus_width_bits (int)

  • shared_mem_per_block (int)

  • shared_mem_per_sm (int)

  • regs_per_block (int)

  • regs_per_sm (int)

  • warp_size (int)

  • peak_bytes_per_s (float)

  • peak_flops_sp (float)

  • peak_flops_dp (float)

  • fp64_ratio (float)

  • ridge_flops_per_byte_sp (float)

  • ridge_flops_per_byte_dp (float)

ridge_flops_per_byte(dtype)[source]#

flops/byte this device balances at for dtype (“float32” or “float64”) — a SELECT between the two C++-precomputed fields above, the same dtype convention eagle::DeviceProps::ridgeFlopsPerByte uses. Zero arithmetic on the Python side.

Return type:

float

Parameters:

dtype (str)

name: str#
cc_major: int#
cc_minor: int#
sm_count: int#
clock_rate_khz: int#
memory_clock_rate_khz: int#
memory_bus_width_bits: int#
shared_mem_per_block: int#
shared_mem_per_sm: int#
regs_per_block: int#
regs_per_sm: int#
warp_size: int#
peak_bytes_per_s: float#
peak_flops_sp: float#
peak_flops_dp: float#
fp64_ratio: float#
ridge_flops_per_byte_sp: float#
ridge_flops_per_byte_dp: float#
class eagle.GraphComposer(member_flags, *, mode='sequenced')[source]#

Bases: object

Compose N members into one launch list, four ways (mode=) – see the module docstring for the pitch and each mode’s guidance.

member_flags (device routing flags / host routing pattern, retained by reference, never rebound):

  • mode="conditional" – required, a uint32 cupy array, one lane per member.

  • the other three modes – any host-side sequence of truthy/falsy values, or None for “every member starts active”.

A caller may mutate member_flags’s contents (via set_routing(), or – mode="conditional" only – direct index assignment) between replays, without recapture, but must never rebind the array itself.

Parameters:

mode (str)

build()[source]#

Perform whatever one-time setup mode needs. Returns self (rejected, every mode, if no members are registered).

mode="conditional" wraps every step in a skippable guard and captures one GraphPipeline once; mode="enabled" registers every member as a plain member of one GraphPipeline, captures once, then applies the initial pattern via set_member_enabled (no recapture from here on); mode="rebuild" captures a super-graph of only the active members, re-capturing from scratch on every later set_routing(); mode="sequenced" is bookkeeping only, since every member already owns its own launchable.

fired_history()[source]#

Per-launch()-call fired-member index sets (frozenset), oldest first, since the last reset_fired_history() (or construction).

Return type:

list

launch(n=1)[source]#

Replay n times. Returns self.

The three flat modes delegate to the one super-pipeline’s GraphPipeline.launch() (a mode="rebuild" composer with an all-inactive pattern owns no pipe, so n replays of nothing is a no-op). mode="sequenced" fires every currently-active registered member (_launch_sequenced()), with the cross-member overlap treatment for stream-owning launchables.

property member_flags#
property mode: str#
node_labels()[source]#

Kernel names of the captured nodes, in declaration order. Same availability caveats as num_nodes().

Return type:

list

property num_members: int#
num_nodes()[source]#

Kernel-node count of the one captured super-graph (after build()). Not available for mode="sequenced" (no single super-graph – introspect each member on its own), nor for a mode="rebuild" composer whose current pattern is all-inactive.

Return type:

int

register(obj, *, name=None, pre_launch=None)[source]#

Register one member; returns its assigned index (registration order).

obj may be:

  • a built launchable – any object with a working .launch(n) (duck-typed). mode="sequenced" only.

  • a plain zero-arg step callable – valid in every mode: one member of this composer’s own super-GraphPipeline under the three flat modes, replayed directly under mode="sequenced".

  • a built mode="sequenced" GraphComposer – the recursion lock. mode="sequenced" outer composers only.

pre_launch (mode="sequenced" only) is an optional zero-arg callable run immediately before this member’s replay(s) – on that member’s own stream if it joins the cross-member overlap group, otherwise host-side. Common rejections, loud: registering after build(); the same object registered twice; a colliding name; nesting a composer under a non-"sequenced" outer composer, or nesting an unbuilt/wrong-mode composer under a "sequenced" one; a pre_launch hook outside mode="sequenced"; an obj that is none of the three accepted shapes.

Return type:

int

reset_fired_history()[source]#

Clear the fired-set history (does not affect current routing).

Return type:

None

set_routing(pattern)[source]#

Update which members are active.

Mechanism + cost depend on mode: "conditional" writes the device flags array in place (no recapture); "enabled" calls GraphPipeline.set_member_enabled per member (no recapture, µs-scale toggle); "rebuild" re-captures a fresh super-graph (paid here, not per replay); "sequenced" updates the host-known launch list.

pattern is any sequence of truthy/falsy values, length == registered member count.

Return type:

GraphComposer

class eagle.GraphPipeline[source]#

Bases: object

A CUDA-graph wrapper: capture a sequence of launches once, replay cheaply. Backed by the eagle._core binding of the eagle::cuda C++ core, so capture/instantiate/replay is the same code a native C++ consumer runs.

add(step, *, name=None)[source]#

Register a zero-arg callable that issues launches on the current stream. The callable may issue one launch or several; what matters is not the count but the stream – build() captures whatever it launches on the stream that is current at the time it runs. name (optional, keyword-only) aliases this step as an individually addressable member: after build(), pipe.set_member_enabled(name, False) (or its registration-order index, always valid – see member()) disables its launches on replay with no recapture. Every step is a member, named or not.

add_concurrent(members, *, names=None)[source]#

Register a fork/join stage: every member is a sibling of every other, and both join before whatever is registered next. members is a sequence of zero-arg callables, len(members) >= 2 – each has the same contract as add()’s step. Co-execution is permitted, never required: replaying members serialized, in registration order, is always a conforming schedule. Within one member its own launches still serialize; fork/join only forks between members. Actual concurrency, if any, is bounded by the device’s own occupancy, never a fixed constant. v1 is series-parallel only: a plain chain of add()/add_concurrent() registrations, each stage joining fully before the next begins. This does not check that members are actually independent – call _concurrent_write_conflict_check() yourself before build() if you want that asserted. names (optional) aliases each member the same way add’s name does – a sequence the same length as members, entries may be None. Every member is individually addressable by its registration-order index regardless of whether it is named.

build()[source]#

Stream-capture the registered launches into an instantiated graph. Immediately before and after each registered member’s launches, this snapshots the in-construction graph’s node set (_core.capture_snapshot_nodes) and records the set difference as that member’s contributed nodes – the only place member boundaries are known at capture time. Toggleability classification (_core.is_node_toggleable) runs separately, right after capture ends: a member is toggleable later only if every one of its nodes is a kernel, memcpy or memset node. A Skippable or RepeatWhile member is marked non-toggleable structurally, from this layer’s own knowledge that it wove a conditional region, without ever driver-querying its nodes (cudaGraphNodeGetType fails on this driver for a conditional/IF node). Either way the member is recorded non-toggleable, never silently accepted.

enqueue(n=1, *, pre_launch=None, seed_events=None)[source]#

Enqueue n replays on this pipeline’s own stream and return immediately – the non-blocking half of launch(). launch(n) is enqueue(n) + synchronize(), verbatim. Use this pair directly only when several independent pipelines should have their device work in flight simultaneously: enqueue them all first, then synchronize them in a second pass, so the total wall time approximates max over the pipelines rather than their sum.

The ordering (see launch()) is applied here, not in synchronize(): an event is recorded on the caller’s current stream (and the legacy Stream.null when different) and waited on this pipeline’s own stream before the replays are issued. pre_launch (optional) is a zero-arg callable run on this pipeline’s own stream after the seed waits and before the replays – the seam for copying per-launch varying data into this pipeline’s buffers. seed_events (optional) are pre-recorded events to wait on instead of recording this pipeline’s own – the shared-event mode eagle.compose.GraphComposer._async_launch_group() needs, one event pair recorded once and waited on every member’s stream. An empty sequence means “I have already ordered this replay myself, issue no seed waits”. HAZARD 1 – one in-flight enqueue per pipeline. A second enqueue before the first has been synchronized re-records the same cached events; overlapping enqueues simply serialize on this pipeline’s single stream, interleaving in a way the caller cannot control. Enqueue once, synchronize, then enqueue again.

HAZARD 2 – the reverse write fence is gone. enqueue returns with the replay still in flight; the ordering events only say replay-after-seeds, nothing about writes issued after:

pipe.enqueue()
buf[...] = new_values  # RACE: the in-flight replay may read these
pipe.synchronize()

is a data race, silently. The discipline is: enqueue -> (only work that does not touch this pipeline’s buffers) -> synchronize -> write; stage new inputs via pre_launch instead. Raises RuntimeError if called before build().

property graph#
introspection_available()[source]#

Whether node introspection works here: the bound core dots the captured graph through cudaGraphDebugDotPrint (always available), so this is ready as soon as build() has run.

Return type:

bool

is_member_toggleable(handle_or_index)[source]#

Whether set_member_enabled() will accept this member (its captured node set is entirely kernel/memcpy/memset). Requires build() to have run.

Return type:

bool

launch(n=1)[source]#

Replay the captured graph n times. launch(n) guarantees all work previously enqueued on the caller’s current cupy stream and the legacy default stream is visible to replay 1.

A stream-ordering fix. The pipeline’s own stream is non_blocking=True precisely so capture never entangles the legacy default stream ordinary cupy/torch launches land on – but that means replay is not otherwise sequenced after host-side seeding (buffer[...] = ...) done on the caller’s current or legacy stream. Under ambient GPU load the seed can lose the race and replay 1 reads pre-seed buffer contents. The fix: record a cupy Event on the caller’s current stream and on the legacy Stream.null (when different) and wait on both, on the pipeline’s own stream, before the replay loop. This orders every consumer of launch() for free, at the facility seam, with no host block (~1-2us/call). Exactly enqueue(n) followed by synchronize() – same events, same order, same host block, same return value. The two halves are also available separately (see enqueue()) for a caller that wants several pipelines’ replays in flight at once.

member(name_or_index)[source]#

Resolve a name or registration-order index to this pipeline’s stable MemberHandle. Available any time after registration; does not require build() (only set_member_enabled() does, since toggleability is a build-time fact).

Return type:

MemberHandle

member_nodes(handle_or_index)[source]#

Raw node handles (as ints) attributed to one registered member by build() – the per-member analogue of node_labels(). A copy; mutating it has no effect. Requires build() to have run.

Return type:

list

node_labels()[source]#

Kernel function names of the captured nodes, in declaration order.

Return type:

list

num_nodes()[source]#

Number of kernel nodes in the captured graph.

Return type:

int

set_member_enabled(handle_or_index, enabled)[source]#

Enable or disable one registered member’s launches on this instantiated graph, effective on the next launch() – no recapture, no change to graph structure (composition mode “enabled”: capture once at full sibling concurrency, then flip driver-level enabled bits between replays). The host-paced sibling of the device-paced conditional (IF) node: where a conditional’s predicate is re-read by an in-graph kernel on every replay, this toggles a driver bit from Python between replays, for zero per-replay gate cost – the right tool when routing decisions are made in Python, not on-device. A member’s captured node set never changes after build(); this only flips whether that fixed node set runs. A disabled member’s nodes become driver-level no-ops (any buffer it would have written is left untouched); an enabled member’s nodes run exactly as captured. Requires CUDA >= 12.3 (the same floor the conditional facility already imposes on the compiled eagle._core extension). Raises NonToggleableMemberError if build() recorded this member as non-toggleable (its node set contains something other than kernel/memcpy/memset – most commonly a Skippable’s conditional node). UnknownMemberError for a bad name or index, and RuntimeError if called before build().

Return type:

GraphPipeline

Parameters:

enabled (bool)

set_node_enabled(node, enabled)[source]#

Enable/disable ONE raw captured node directly, bypassing this pipeline’s own member registry. For a caller that tracks its own node attribution outside add()/add_concurrent()’s member bookkeeping. The same underlying call set_member_enabled() makes per node (Launcher.set_node_enabled); this is only a narrower, public door to it for a caller already holding valid node handles (from capture_snapshot_nodes()).

node must be a real handle from this pipeline’s own capture – passing a foreign or stale one is undefined at the driver level. Requires build() to have run.

Return type:

GraphPipeline

Parameters:
  • node (int)

  • enabled (bool)

property stream: int#

This pipeline’s own cudaStream_t, as a raw integer handle (the family’s shared stream name: the same value a DLPack consumer passes as __dlpack__(stream=...)). For capture composition: a caller that needs to bind a cuBLAS/cuSOLVER-style handle to this pipeline’s stream, or run warm-up passes before build() captures, needs this exact stream – launching on any other stream would race the graph. Read-only, valid for this pipeline’s entire lifetime.

synchronize()[source]#

Block the host until this pipeline’s enqueued replays have completed – the blocking half of launch(). Safe to call with nothing in flight (a cheap no-op query on an already-idle stream). Raises RuntimeError if called before build().

class eagle.KernelRegistry(kernels=None)[source]#

Bases: object

A name -> launchable map for resolving a launchable by name.

Values are any launchable sharing the __call__ + launch protocol. Dict-like: reg[name], name in reg, len(reg), iteration over names, .get / .names.

get(name, default=None)[source]#

Return the launchable registered under name, or default.

Parameters:

name (str)

names()[source]#

The registered names, in insertion order.

Return type:

list

register(kernel, name=None)[source]#

Register kernel under name (default __name__); return the key.

Raises on a duplicate name – pass an explicit name= to disambiguate.

Return type:

str

Parameters:

name (str | None)

exception eagle.LayoutWarning[source]#

Bases: UserWarning

A sample-major plane was copied into eagle’s component-major layout.

Filterable like any warning; make it an error with warnings.filterwarnings("error", category=eagle.LayoutWarning).

eagle.samples_first(x)[source]#

Mark x as sample-major: its FIRST axis holds the samples, (N, w) (a matrix plane (N, R, C)). Zero-copy, honoured for any shape; passes through any eagle or hawk door that binds per-sample planes, and beats the door’s layout=. A shape that contradicts it is refused.

eagle.samples_last(x)[source]#

Mark x as component-major (native): its LAST axis holds the samples, (w, N). Zero-copy, honoured for any shape; passes through any eagle or hawk door and beats the door’s layout=. A shape that contradicts it is refused.

class eagle.LoadedKernel(ptx_path, *, abi_kind)[source]#

Bases: LaunchMixin

Common scaffolding for a PTX-loaded launcher (no nvcc at runtime): reads the shared sidecar fields, rejects a binary-ABI mismatch, and driver-loads the module. Subclasses read their own fields, build self._allowed, and define launch / __call__ over the inherited _dispatch(). The parsed sidecar is kept on self._meta.

abi_tag#

The artifact’s own ABI generation, read by _refuse_v2_door() and by eagle.plan.plan’s legacy-partition bridge.

param_names#

The binding names of params.

class eagle.LoadedPure(ptx_path)[source]#

Bases: LoadedKernel

A precompiled pure kernel loaded from a PTX (+ sidecar) artifact – no nvcc. Drop-in equivalent to the in-process compiled pure launcher via eagle.launch. The Mutable buffers are the handoff (read-modify-written in place), so there is no separate out=.

__call__(**kw)[source]#

Launch the pure kernel and return the updated Mutable buffers as a dict, in the caller’s framework.

launch(*, grid=None, block=None, **kw)[source]#

Capturable single pure launch on the current stream (cupy arrays only). The writable Mutable buffers, any inputs, lookup tables, and the terminated mask must be pre-allocated; read-modify-written in place. block=None defers to eagle’s launch-policy resolver.

class eagle.LoadedVector(ptx_path)[source]#

Bases: LoadedKernel

A precompiled vector kernel loaded from a PTX (+ sidecar) artifact – no nvcc.

Drop-in equivalent to the in-process compiled vector-kernel launcher.

__call__(*, out=None, **kw)[source]#

Allocate, launch, and return the contribution, in the caller’s tensor framework (numpy in -> numpy out, cupy/torch -> same via DLPack): blocking for numpy, non-blocking on the framework’s stream otherwise. Pass out= a pre-allocated device buffer to fill it in place – see as_out_buffer().

launch(*, out, grid=None, block=None, **kw)[source]#

Capturable single launch on the current stream (cupy arrays only). Inputs, out, and any lookup tables must be pre-allocated and passed by name. block=None defers to eagle’s launch-policy resolver.

class eagle.RepeatWhile(step, guard, max_iters)[source]#

Bases: object

A zero-arg step repeated on the device while a SkipGuard holds, at most max_iters times per replay – the loop sibling of Skippable (not a subclass: 0..N runs and 0..1 runs are different contracts).

Under GraphPipeline.build() the step becomes the body of one CUDA WHILE conditional node: a head kernel resets the [remaining, ran] cell and decides entry, the body runs, and a tail kernel decides continuation. launch(1) is one cudaGraphLaunch regardless of iteration count; launch(n) replays the whole loop n times.

max_iters is a plain Python int >= 1 frozen at build as a kernel argument – a new cap needs a new build(). Called directly it is the host arm: while ran < max_iters and guard.evaluate(): step(), writing the same cell so iterations() reads the same in every mode. The cell lives in the guard’s array module (cupy or numpy), allocated at construction; iteration_index exposes its ran word to body kernels as the device-visible iteration index (0 on the first iteration).

Composition: the body step may be a Skippable, or a tuple of such parts run in order (e.g. (k_steps, skippable(compact, guard))); a RepeatWhile may be an add_concurrent member, but nesting one directly inside another RepeatWhile or a Skippable is refused at GraphPipeline.build(). One hidden inside an opaque body callable is not detected structurally: called mid-capture it runs its host arm, whose guard read syncs and fails the capture loudly.

Parameters:
step#
guard#
max_iters#
property iteration_index#

One-element uint32 view of the ran word – the 0-based index of the iteration in flight. Pass it to a body kernel that needs it; same array module as the guard.

iterations()[source]#

Iterations the last run performed (host read of ran; a device sync under a cupy guard, eager/debug paths only). 0 before any run.

Return type:

int

property parts: tuple#

The body as a tuple of parts, run in order (a single-step body is a one-part tuple).

class eagle.RunReport(n, finished, done, launches, steps, compactions, reorders, build_s, launches_by_k=<factory>, exact_steps=False, lane_utilisation=None, mode='band', probe=False)[source]#

Bases: object

What one Runner.run() did.

finished counts samples whose mask is set when the run ends, done is finished == n, launches the loop’s iteration count, and steps the steps the loop ran (launches * K, or the summed k for an automatic artifact). exact_steps says steps is exact (true for K=1, or an automatic artifact that launched nothing, hit the cap early, or ended on a one-step launch). compactions/reorders count what ran, build_s the one-time capture wall (0 on host), launches_by_k each k’s launch count. lane_utilisation (active lanes / warp-iterations, in (0, 1]) is set only for a persist-entry run that counted it; None otherwise (every other mode, or a persist run that skipped the counter). mode is the launch shape this run actually took: "fused_one"/ "persist" for a no-graph fast-path run, "band" for the WHILE-graph/policy path (the only mode before the fast path existed). probe is true for a fast-path run taken to time the entry the batch is not currently using (see Runner._pick_entry()).

Parameters:
  • n (int)

  • finished (int)

  • done (bool)

  • launches (int)

  • steps (int)

  • compactions (int)

  • reorders (int)

  • build_s (float)

  • launches_by_k (dict)

  • exact_steps (bool)

  • lane_utilisation (float | None)

  • mode (str)

  • probe (bool)

exact_steps: bool = False#
lane_utilisation: float | None = None#
mode: str = 'band'#
probe: bool = False#
n: int#
finished: int#
done: bool#
launches: int#
steps: int#
compactions: int#
reorders: int#
build_s: float#
launches_by_k: dict#
class eagle.Runner(plan, *, max_steps, every=None, reorder=None, **planes)[source]#

Bases: object

A bound, built run-until-done loop over one plan, or a step of several plans launched in order (see until_done()).

Attributes: step (the eagle.plan.BoundPlan, or a tuple), loop, pipeline (the device eagle.GraphPipeline, built on first run()), active (the eagle.ActiveSet, or None), finished/terminated, steps_per_launch (K, or "auto"), every/launches_per_compaction. For an automatic artifact fused_steps is the device steps-per-launch word and policy reads the policy’s cells (a sync). host_policy picks how a host run paces automatic launches: "auto" (default: "tiled" when the entry’s fused side is tiled, else "measured"), "tiled" (steps_max steps per launch), "measured" (sweeps one step per launch until a timed fused probe wins) or "band" (the device’s band). host_threads is the host team size: None (default) lets eagle pick it (one thread per physical core for a short run, every logical CPU for a long one — see eagle._host_loop.host_team_size()), an int forces it. Planes are held by reference: a run writes the caller’s arrays in place; a new binding is a new runner.

step#
loop#
pipeline#
active#
finished#
terminated#
n#
steps_per_launch#
every#
launches_per_compaction#
fused_steps#
auto#
host_policy#
host_threads#
property policy: dict | None#

The automatic loop’s policy cells as host ints (a device sync): k, finished_at_launch, steps_done, max_steps, go; None for a fixed artifact.

reset()[source]#

Clear the mask/counter/active set: the next run() starts the batch over.

Return type:

None

run()[source]#

Run until every sample is done (or the cap is reached) and report; the counter is re-synced to the mask first, so a run continues from the current state (a reordering set’s planes land back in sample order).

The counters this needs before/after the launch (not counting the automatic policy’s own steps_done/histogram/ fused_steps reads, still separate – see _build_auto()) are read as ONE packed transfer per snapshot (_packed_read()) instead of a separate full round-trip sync each: a plain, non-compacting run (self.active is None) needs no PRE snapshot at all (_compactions can only ever be 0 without an active set, so it is never even read), and the POST snapshot folds iterations/finished/compactions/reorders into one.

A fast mode (_build_fast()) skips all of the above: one launch IS the run, so _run_fast() handles it entirely.

Return type:

RunReport

class eagle.SimResult(*, state, finished, report, wall_s, allocated)[source]#

Bases: object

What one Simulation.run() produced. state maps each state plane to its array (written in place, plus allocated ones, listed in allocated); each is also an attribute (result.x) and an item (result["x"]). finished is the per-sample mask, done whether every sample did, status "finished"/"max_steps", steps/report/wall_s the loop’s steps, full eagle.RunReport and wall time.

state#
finished#
done#
steps#
status#
report#
n#
wall_s#
allocated#
class eagle.Simulation(model, args, kwargs, *, until=None, max_steps, every=None, reorder=None, scalar_type=None, layout=None)[source]#

Bases: object

A bound, built simulation (see simulation()): run() runs it, reset() starts the batch over. runner is the eagle.Runner driving the model, loop/pipeline its loop and device eagle.GraphPipeline (None on host), plans the plans per kernel, n the batch size, names the state/params/allocated plane names.

runner#
loop#
plans#
n#
names#
active#
finished#
terminated#
every#
property pipeline#

The device eagle.GraphPipeline (built on the first run()); None on the host and before.

reset()[source]#

Clear the finished mask (and the active set): the next run() starts the batch over; the caller re-seeds its state.

Return type:

None

run()[source]#

Run until every sample finishes or the step cap is reached; the state is written in place and returned in the SimResult. A second call continues from the current state.

Return type:

SimResult

class eagle.SkipGuard(count_arr, count_idx=0, baseline_arr=None, baseline_idx=0)[source]#

Bases: object

Region runs iff count_arr[count_idx] != (baseline_arr[baseline_idx] or 0). Arrays are uint32 device (cupy) or host (numpy) arrays; references are retained so pointer lifetime stays pinned to the pipeline.

count_arr#
count_idx#
baseline_arr#
baseline_idx#
evaluate()[source]#

Host read; eager/debug paths ONLY (syncs on cupy).

Return type:

bool

classmethod intent(flags, index)[source]#

Router/user-owned intent flag: flags[index] != 0.

flags is a uint32 array the caller owns and writes between graph replays, one lane per guarded region (e.g. SkipGuard.intent(expert_flags, e) for expert e). No baseline: a plain on/off decision.

Under GraphPipeline.build() the predicate is read fresh from the device every replay, so flipping flags[index] takes effect with no recapture; in eager/numpy mode the same flag is read host-side by evaluate().

index has no default (unlike nonzero()’s 0): it is load-bearing at every call site.

classmethod nonempty_bucket(type_offsets, t)[source]#

Fencepost liveness: type_offsets[t + 1] != type_offsets[t].

classmethod nonzero(n_active)[source]#

Plain nonzero count: n_active[0] != 0.

class eagle.Skippable(step, guard)[source]#

Bases: object

A zero-arg step wrapped with a SkipGuard. Works as a GraphPipeline.add() step or an add_concurrent member – GraphPipeline.build() recognizes it by type and weaves an IF node around it; called directly, it evaluates its guard eagerly: if guard.evaluate(): step().

Parameters:

guard (SkipGuard)

step#
guard#
eagle.assemble_args(arg_spec, *, out, vec, per_sample, terminated, uniforms, n, tables=None, mutables=None, mutable_vec=None, mutable_mat=None, mats=None, wide_in=None, wide_out=None, plan=None)[source]#

Build the kernel argument tuple in the order the signature expects.

The mutable output, input vectors, and matrix inputs are by-value GRef structs (a matrix GRef composes the flat vector one); per-sample scalars, the termination mask, lookup tables, and a pure kernel’s scalar/int Mutable slots are by-value HandleT views; broadcast constants are float64. A vector kernel’s GRef carries its own samples_; a pure kernel takes an explicit nsamples instead. wide_in/wide_out are also HandleT views, each riding its own caller-populated dict so they never collide with an unrelated per-sample-scalar or lookup-table name.

The ABI-shape decision per role is eagle.roles.classify_arg() (single-sourced with the ctypes host path); this function keeps only the construction – which dict a role’s value comes from, and which make_* helper builds it. An unrecognised role raises before any role-specific branch runs.

plan (optional) is the LaunchPlan for THIS signature. When given, the per-argument classification and POD boxes come from it instead of being re-derived here; the resulting argument list is byte-identical either way (read LaunchPlan’s box-reuse note first). Omit plan for fresh boxes per call, as always.

eagle.check_aether_abi(meta, *, kind, name)[source]#

Reject an artifact whose meta["aether_abi"] is not one of ACCEPTED_ABI_TAGS (absent/empty included). kind names the artifact in the error (e.g. ‘vector plugin’, ‘plugin manifest’).

Return type:

None

Parameters:
  • meta (dict)

  • kind (str)

  • name (str)

eagle.compaction_body(step, active, *, every=16, finished=None, reorder=None, steps_per_call=1)[source]#

The eagle.repeat_while() body that runs step every times, then compacts active – only when finished moved, if given, else unconditionally – and, for a set built with reorder=theta over at least MIN_REORDER_SPAN samples, reorders when the compaction fired (reorder=False/True overrides this). The loop guard is checked once per every steps, so it may run up to every - 1 steps past the last sample finishing; size max_iters as ceil(max_steps / every). steps_per_call is the steps one step() call advances (a fused kernel’s K): the body then runs every * K steps between compactions, and a cadence below MIN_EVERY steps is refused.

Return type:

tuple

Parameters:
  • active (ActiveSet)

  • every (int)

  • steps_per_call (int)

eagle.deploy(plugin, *, targets=None, cache_dir=None, scalar_type=None, **plan_kw)#

An AutoPlan for plugin: run it where its data lives. eagle.deploy is this same function.

plugin is a built plugin, a hawk kernel, or a list built into one bundle (returned as a tuple of plans in order). The build uses hawk’s cache under cache_dir (None: hawk’s default); a repeat call compiles nothing. targets (None: both sides when a GPU is usable – the call returns with the device built, the host still building beside it – else ("host",)), cache_dir and scalar_type are refused for an already-built plugin. scalar_type is the precision hawk builds in: "float64" (None, the default) or "float32"; the planes passed at run time must match it.

plan_kw are plan()’s keywords but structure (inner/gather do not apply); a side the plugin lacks is refused when first selected.

Return type:

AutoPlan | tuple[AutoPlan, ...]

eagle.detect(x)[source]#

Identify x’s framework adapter without importing the framework. None for a value we don’t recognize as a tensor (to_cupy() then treats it as a host upload).

Return type:

Adapter | None

eagle.device_props(device=0)[source]#

Query device device’s properties via eagle._core.device_props() and wrap the resulting dict in a frozen DeviceProps — no Python-side computation, only a dict-to-dataclass reshape.

eagle._core (the compiled extension) is imported LOCALLY, here, never at this module’s top — importing eagle on a cuda-free box must not force-load it (the same discipline eagle.pipeline. GraphPipeline.__init__() follows for the same reason).

Return type:

DeviceProps

Parameters:

device (int)

eagle.launch(kind, fn, arg_spec, vector_inputs, params, per_sample, *, kw, grid, block, dt=<class 'numpy.float64'>, **extra)[source]#

Issue ONE capturable launch of any kernel kind on the current stream. kind ("vector" | "pure") selects the ABI binding shape; kind-specific inputs ride extra (a vector kernel’s out=, or a pure kernel’s mutables_decl=, each with optional matrix_inputs=/wide_inputs=/wide_outputs=/ wide_out_exempt=, named identically to pure_prepare()’s). Every array argument must be pre-allocated cupy. Returns the vector out, or the pure Mutable dict merged with wide outputs.

eagle.load_manifest(path)[source]#

Load a deployed manifest.json by its top-level pattern.

The Python analogue of the C++ PluginRegistry::from_manifest – but where that injects the whole ordered set, this returns a by-name result:

  • "vector" / "pure" – each enabled entry’s artifact is path-loaded into LoadedVector / LoadedPure and keyed by its manifest id into a KernelRegistry.

  • "neural_block" – a neural deployment unit: the inline blocks[] compose one ordered layer whose exec references resolve against the plugins[] kernels, returned as a frozen neural-layer wrapper. Recognized but not launch-certified, so it never enters the by-name map.

An unknown or absent pattern is a hard error. Returns KernelRegistry | LoadedNeuralLayer.

eagle.make_gref(arr, n)[source]#

Build a by-value GRef struct over a contiguous (3, N) cupy array: compStride_ == samples_ == n, sampleStride_ == 1.

eagle.make_handle(arr, n)[source]#

Build a by-value scalar-array HandleT over a contiguous (N,) array; n is required since the struct carries its own extent.

eagle.np_dtype(scalar_type)[source]#

Return the numpy dtype of Real-typed arrays for a resolved scalar_type.

SoftDouble is bit-identical IEEE float64 — the array interface is float64 (mirrors eagle.dtypes.np_dtype’s table literal).

Parameters:

scalar_type (str)

eagle.origin_adapter(candidates)[source]#

Pick the adapter the kernel’s output should be returned as, over the state-vector inputs in signature order: the first strong framework (torch/jax/tensorflow) wins (two distinct ones raise TypeError), else the first device adapter, else numpy.

Return type:

Adapter

eagle.pure_prepare(fn, arg_spec, *, vector_inputs, per_sample, params, mutables_decl, mutable_defaults, lookup_counts, kw, mat_shapes=None, vec_widths=None, wide_inputs=None, wide_outputs=None, wide_out_exempt=(), dt=<class 'numpy.float64'>, readonly_mask=False)[source]#

Coerce inputs, bind the writable Mutables, assemble args, launch (no sync). The pure allocate-and-launch body, shared by the in-process compiled pure launcher and the deployed LoadedPure so the two can never drift.

State vectors -> (3, N) f64; matrix inputs -> flat (R*C, N) f64; provided Mutable arrays -> device buffers (also sizing N); per-sample scalars -> (N,); lookup tables -> flat handles; wide inputs -> 2-D (rows, N); wide gradient outputs -> 2-D (rows, N) written in place, never auto-allocated. wide_out_exempt names the wide_outputs that are row-indexed destinations (e.g. an atomic Accum output), skipping N-reconciliation. A Mutable omitted from kw is a broadcast scalar fill or its declared default. Returns the Mutable buffers merged with any wide gradient outputs (RMW in place). Meant to run inside the caller’s launch context (see LaunchMixin._dispatch()).

readonly_mask (default False) is the sidecar-declared opt-in letting an omitted terminated mask be served from the cached all-false mask (see coerce_terminated()); never inferred, since marshal cannot read a kernel’s source.

eagle.repeat_while(step, guard, max_iters)[source]#

Repeat step (a zero-arg callable, a Skippable, or a tuple of such parts) while guard holds, at most max_iters times: one device-side WHILE node under GraphPipeline.build(), a host loop everywhere else.

Return type:

RepeatWhile

Parameters:
eagle.run_until_done(plan, *, max_steps, every=None, reorder=None, **planes)[source]#

until_done() then Runner.run(): the one-call form.

Return type:

RunReport

Parameters:
  • max_steps (int)

  • every (int | None)

eagle.simulate(model, *args, until=None, max_steps, every=None, reorder=None, scalar_type=None, layout=None, **kwargs)[source]#

Run model on every sample until each one finishes (or max_steps is reached) and return the SimResult.

model is a hawk kernel, a list of them (one step, in list order), or plans from eagle.deploy(); until is the kernel that finishes the samples, placed last (at least one kernel must finish). The rest of the call is the kernel’s own: *args/**kwargs bind to its parameters exactly as a call to the kernel function would (positionally, by name, or both), by its own annotations – except that a model built from a PREBUILT plan (no raw kernel to read a call signature from) refuses *args outright, naming the keyword call to use instead; keyword binding always works. A Mutable plane is the state it updates (its initial value; one never read before it is written may be left out, zero-filled instead), a Param a number shared by every sample, a Scalar/Vector/Table plane one value per sample (a number where the kernel allows it, an array otherwise); Terminated is never passed. numpy/numbers run on the CPU, cupy on the GPU. max_steps is required; every/reorder are eagle.until_done()’s compaction cadence and reorder threshold, scalar_type the precision the kernels are built in ("float64", the default, or "float32"; see eagle.deploy()), layout ("samples_first" or "samples_last") which axis holds the samples of a per-sample plane whose shape reads both ways ((w, w)), which is otherwise refused; eagle.samples_first(x) / eagle.samples_last(x) say it for one array and win. All keyword-only like every one of eagle’s own options – after the kernel’s own arguments.

Return type:

SimResult

Parameters:
  • max_steps (int)

  • every (int | None)

eagle.simulation(model, *args, until=None, max_steps, every=None, reorder=None, scalar_type=None, layout=None, **kwargs)[source]#

Bind and build model once; returns the Simulation, whose run() runs it (again, after reset(), without rebuilding). *args/**kwargs are the kernel’s own, bound exactly like a call to it; the rest are simulate()’s own options.

Return type:

Simulation

Parameters:
  • max_steps (int)

  • every (int | None)

eagle.skippable(step, guard)[source]#

Wrap step (a zero-arg callable) so it only runs while guard holds: one IF node under GraphPipeline.build(), an eager guard check everywhere else.

Return type:

Skippable

Parameters:

guard (SkipGuard)

eagle.to_cupy(x)[source]#

Import a framework / DLPack tensor to a cupy array: zero-copy for a CUDA-resident producer, a host upload otherwise. Also the pre-import helper for the capturable launch path, since a DLPack import can synchronize and is unsafe mid-capture.

eagle.until_done(plan, *, max_steps, every=None, reorder=None, **planes)[source]#

Bind plan and build the loop that runs it until every sample is done; returns the Runner (.run() runs it).

**planes binds every name of the plan’s arg_spec by eagle.plan.Plan.bind()’s rules, except the reserved FINISHED_PLANE, active_map and active_count (the runner allocates these); the terminated mask is all-false when omitted. plan may be an eagle.plan.AutoPlan (residency picks host/device), or a list of plans run as one step, in order, over one namespace of planes; each takes one step per launch, at least one finishes the samples, and they share one mask and guard.

max_steps caps the steps a sample may take (a compacting loop may overrun by up to every - 1 steps; RunReport.steps counts them). Guard(active_set=True) compacts every every steps (a multiple of K, at least 4, default the smallest at least 16), and reorder=theta also reorders the bound planes; both are refused on a plain artifact.

An automatic artifact (steps="auto") runs the policy loop instead (see the module docstring): the cap is exact, every= is refused, reorder only with reorder=theta, and the runner owns the FUSED_STEPS_PLANE word.

Return type:

Runner

Parameters:
  • max_steps (int)

  • every (int | None)