eagle#
eagle — the Python face of the EAGLE launch engine.
- Which call do I use?
one kernel (or a step’s worth), run until every sample finishes –
simulate()a plan already paced by its own finish kernel, one call –until_done()build a kernel into a plan that runs where its data lives –deploy()/planresolve a launchable BY NAME, then call/stream/graph it –KernelRegistry
EAGLE owns everything about executing compiled kernels: the
framework-polymorphic launch skeleton (LaunchMixin), the
capturable kind-dispatched launch() primitive,
input-role marshalling (eagle.marshal), the by-value launch ABI
(eagle.abi), DLPack
framework adapters (eagle.interop), CUDA-graph capture/replay
(GraphPipeline), composing many already-built
launchables into one launch list (GraphComposer),
and the plugin protocol consumers —
KernelRegistry / load_manifest()
on the Python side, mirroring the C++ host/ registries.
Producers of launchable artifacts (compiled CUDA sources, PTX bundles with
sidecar/manifest metadata) live upstream — the code generator is the first — and depend
on this package; the eagle core never imports a producer. The optional framework bridges
(eagle.frameworks) are the one exception: they import hawk lazily, on first use,
and so does eagle.deploy() (the same function as eagle.plan.auto()), which
builds a hawk kernel into a plan that runs where its data lives.
eagle.simulate() runs a model (one hawk kernel, or the kernels of one step)
on every sample until each one finishes, called exactly like the kernel itself.
- class eagle.ActiveSet(mask, *, reorder=None, _lazy_scratch=False, _lazy_map=False)[source]#
Bases:
objectThe live samples of
mask(aterminatedplane: set = finished), as an ascending index map plus a count – and, withreorder=theta, the occasional physical reorder of the planes it owns.maskis a one-dimensional, C-contiguousbool/uint8array (cupy or numpy), held by reference and read where it is every time. The map, count and device scratch are allocated once, before any capture, and live as long as this object.reorderis an explicit opt-in: the locality thresholdtheta(inTHETA_RANGE, orTrueforDEFAULT_THETA) that pays on large, irregularly thinning batches (from about 1e5 samples; below that the map alone is as fast). The set then owns the mask and every plane given toown(). The set starts as the identity (every sample live,count == n).- mask#
- n#
- map#
- count#
- theta#
- perm#
- inv#
- span#
- fire#
- live32#
- reorders#
- __call__()#
Recompute the map and the count from the mask (and, for a reordering set,
live32/fire). Device: enqueued on the current cupy stream (capturable). Host: computed now.- Return type:
None
- compact()[source]#
Recompute the map and the count from the mask (and, for a reordering set,
live32/fire). Device: enqueued on the current cupy stream (capturable). Host: computed now.- Return type:
None
- in_sample_order(plane)[source]#
A copy of
planein sample order (plane[..., inv]); legal while permuted, capture-legal on the device.
- indirect(*planes)[source]#
Declare bound per-sample planes the caller keeps in sample order (its kernel reads sample
perm[slot]from them). Returns the set.- Return type:
- property live: int#
The live count as a Python int (a device sync for a device set).
- materialize_map()[source]#
Allocate the deferred map (see
_lazy_map) and set it to the identity.Truewhen it was deferred, so the caller must re-point every bound plan atmap;Falsewhen it already is real.- Return type:
bool
- own(*planes)[source]#
Hand the set per-sample planes to move with the mask at every reorder (a
(D, n)array is D planes). Refused once fixed, for a plane another set owns, or one overlapping an owned plane.- Return type:
- property permuted: bool#
Whether a reorder ran since the last
restore()/reset()(a device sync for a device set).
- planes()[source]#
The two reserved planes as
eagle.plan.Plan.bind()keywords.- Return type:
dict
- prepare_compaction()[source]#
Allocate the device compaction scratch (once). The first
compact()does it when needed; anything that captures a compaction into a graph calls it first. A runner on a no-compaction fast path never pays for it.- Return type:
None
- reorder()[source]#
Run the reorder now if the last compaction fired it: owned planes and
permmove so live samples come first,inv/the map reset,span = count,reorders += 1. Device: enqueued on the current stream; host: now.- Return type:
None
- reorder_if_degraded()[source]#
The reorder as a
eagle.Skippableguarded byfire(one IF node under a graph, an eager check elsewhere).
- reset()[source]#
Back to the identity: every sample live,
count == n(and, for a reordering set,perm/invidentity,span == n, reorder counter zero). Not capture-legal: a plain array write, run before a launch, never recorded.- Return type:
None
- restore()[source]#
Put every owned plane back in sample order;
perm/invreset to identity,span = n, the map recomputed. Eager only (refused under stream capture); a no-op for a set that does not reorder.- Return type:
None
- when_finished(finished)[source]#
A
eagle.Skippablecompaction that runs only whenfinishedmoved since it last ran (the guard is evaluated on the device under a graph).
Bases:
RuntimeErrorThe device backend cannot serve this call (see the message for why).
A
RuntimeError, so code that already guards a device call withexcept RuntimeErrorkeeps working.
- class eagle.DeviceProps(name, cc_major, cc_minor, sm_count, clock_rate_khz, memory_clock_rate_khz, memory_bus_width_bits, shared_mem_per_block, shared_mem_per_sm, regs_per_block, regs_per_sm, warp_size, peak_bytes_per_s, peak_flops_sp, peak_flops_dp, fp64_ratio, ridge_flops_per_byte_sp, ridge_flops_per_byte_dp)[source]#
Bases:
objectFrozen view of one CUDA device’s raw + derived properties.
Field set mirrors
eagle::DeviceProps(eagle/DeviceProps.h) exactly — see that header for every formula’s one-line derivation.- Parameters:
name (str)
cc_major (int)
cc_minor (int)
sm_count (int)
clock_rate_khz (int)
memory_clock_rate_khz (int)
memory_bus_width_bits (int)
shared_mem_per_block (int)
shared_mem_per_sm (int)
regs_per_block (int)
regs_per_sm (int)
warp_size (int)
peak_bytes_per_s (float)
peak_flops_sp (float)
peak_flops_dp (float)
fp64_ratio (float)
ridge_flops_per_byte_sp (float)
ridge_flops_per_byte_dp (float)
- ridge_flops_per_byte(dtype)[source]#
flops/byte this device balances at for
dtype(“float32” or “float64”) — a SELECT between the two C++-precomputed fields above, the same dtype conventioneagle::DeviceProps::ridgeFlopsPerByteuses. Zero arithmetic on the Python side.- Return type:
float- Parameters:
dtype (str)
-
name:
str#
-
cc_major:
int#
-
cc_minor:
int#
-
sm_count:
int#
-
clock_rate_khz:
int#
-
memory_clock_rate_khz:
int#
-
memory_bus_width_bits:
int#
-
regs_per_block:
int#
-
regs_per_sm:
int#
-
warp_size:
int#
-
peak_bytes_per_s:
float#
-
peak_flops_sp:
float#
-
peak_flops_dp:
float#
-
fp64_ratio:
float#
-
ridge_flops_per_byte_sp:
float#
-
ridge_flops_per_byte_dp:
float#
- class eagle.GraphComposer(member_flags, *, mode='sequenced')[source]#
Bases:
objectCompose N members into one launch list, four ways (
mode=) – see the module docstring for the pitch and each mode’s guidance.member_flags(device routing flags / host routing pattern, retained by reference, never rebound):mode="conditional"– required, auint32cupy array, one lane per member.the other three modes – any host-side sequence of truthy/falsy values, or
Nonefor “every member starts active”.
A caller may mutate
member_flags’s contents (viaset_routing(), or –mode="conditional"only – direct index assignment) between replays, without recapture, but must never rebind the array itself.- Parameters:
mode (str)
- build()[source]#
Perform whatever one-time setup
modeneeds. Returnsself(rejected, every mode, if no members are registered).mode="conditional"wraps every step in askippableguard and captures oneGraphPipelineonce;mode="enabled"registers every member as a plain member of oneGraphPipeline, captures once, then applies the initial pattern viaset_member_enabled(no recapture from here on);mode="rebuild"captures a super-graph of only the active members, re-capturing from scratch on every laterset_routing();mode="sequenced"is bookkeeping only, since every member already owns its own launchable.
- fired_history()[source]#
Per-
launch()-call fired-member index sets (frozenset), oldest first, since the lastreset_fired_history()(or construction).- Return type:
list
- launch(n=1)[source]#
Replay
ntimes. Returnsself.The three flat modes delegate to the one super-pipeline’s
GraphPipeline.launch()(amode="rebuild"composer with an all-inactive pattern owns no pipe, sonreplays of nothing is a no-op).mode="sequenced"fires every currently-active registered member (_launch_sequenced()), with the cross-member overlap treatment for stream-owning launchables.
- property member_flags#
- property mode: str#
- node_labels()[source]#
Kernel names of the captured nodes, in declaration order. Same availability caveats as
num_nodes().- Return type:
list
- property num_members: int#
- num_nodes()[source]#
Kernel-node count of the one captured super-graph (after
build()). Not available formode="sequenced"(no single super-graph – introspect each member on its own), nor for amode="rebuild"composer whose current pattern is all-inactive.- Return type:
int
- register(obj, *, name=None, pre_launch=None)[source]#
Register one member; returns its assigned index (registration order).
objmay be:a built launchable – any object with a working
.launch(n)(duck-typed).mode="sequenced"only.a plain zero-arg step callable – valid in every mode: one member of this composer’s own super-
GraphPipelineunder the three flat modes, replayed directly undermode="sequenced".a built
mode="sequenced"GraphComposer– the recursion lock.mode="sequenced"outer composers only.
pre_launch(mode="sequenced"only) is an optional zero-arg callable run immediately before this member’s replay(s) – on that member’s own stream if it joins the cross-member overlap group, otherwise host-side. Common rejections, loud: registering afterbuild(); the same object registered twice; a collidingname; nesting a composer under a non-"sequenced"outer composer, or nesting an unbuilt/wrong-mode composer under a"sequenced"one; apre_launchhook outsidemode="sequenced"; anobjthat is none of the three accepted shapes.- Return type:
int
- reset_fired_history()[source]#
Clear the fired-set history (does not affect current routing).
- Return type:
None
- set_routing(pattern)[source]#
Update which members are active.
Mechanism + cost depend on
mode:"conditional"writes the device flags array in place (no recapture);"enabled"callsGraphPipeline.set_member_enabledper member (no recapture, µs-scale toggle);"rebuild"re-captures a fresh super-graph (paid here, not per replay);"sequenced"updates the host-known launch list.patternis any sequence of truthy/falsy values, length == registered member count.- Return type:
- class eagle.GraphPipeline[source]#
Bases:
objectA CUDA-graph wrapper: capture a sequence of launches once, replay cheaply. Backed by the
eagle._corebinding of theeagle::cudaC++ core, so capture/instantiate/replay is the same code a native C++ consumer runs.- add(step, *, name=None)[source]#
Register a zero-arg callable that issues launches on the current stream. The callable may issue one launch or several; what matters is not the count but the stream –
build()captures whatever it launches on the stream that is current at the time it runs.name(optional, keyword-only) aliases this step as an individually addressable member: afterbuild(),pipe.set_member_enabled(name, False)(or its registration-order index, always valid – seemember()) disables its launches on replay with no recapture. Every step is a member, named or not.
- add_concurrent(members, *, names=None)[source]#
Register a fork/join stage: every member is a sibling of every other, and both join before whatever is registered next.
membersis a sequence of zero-arg callables,len(members) >= 2– each has the same contract asadd()’s step. Co-execution is permitted, never required: replaying members serialized, in registration order, is always a conforming schedule. Within one member its own launches still serialize; fork/join only forks between members. Actual concurrency, if any, is bounded by the device’s own occupancy, never a fixed constant. v1 is series-parallel only: a plain chain ofadd()/add_concurrent()registrations, each stage joining fully before the next begins. This does not check that members are actually independent – call_concurrent_write_conflict_check()yourself beforebuild()if you want that asserted.names(optional) aliases each member the same wayadd’snamedoes – a sequence the same length asmembers, entries may beNone. Every member is individually addressable by its registration-order index regardless of whether it is named.
- build()[source]#
Stream-capture the registered launches into an instantiated graph. Immediately before and after each registered member’s launches, this snapshots the in-construction graph’s node set (
_core.capture_snapshot_nodes) and records the set difference as that member’s contributed nodes – the only place member boundaries are known at capture time. Toggleability classification (_core.is_node_toggleable) runs separately, right after capture ends: a member is toggleable later only if every one of its nodes is a kernel, memcpy or memset node. ASkippableorRepeatWhilemember is marked non-toggleable structurally, from this layer’s own knowledge that it wove a conditional region, without ever driver-querying its nodes (cudaGraphNodeGetTypefails on this driver for a conditional/IF node). Either way the member is recorded non-toggleable, never silently accepted.
- enqueue(n=1, *, pre_launch=None, seed_events=None)[source]#
Enqueue
nreplays on this pipeline’s own stream and return immediately – the non-blocking half oflaunch().launch(n)isenqueue(n)+synchronize(), verbatim. Use this pair directly only when several independent pipelines should have their device work in flight simultaneously: enqueue them all first, then synchronize them in a second pass, so the total wall time approximatesmaxover the pipelines rather than theirsum.The ordering (see
launch()) is applied here, not insynchronize(): an event is recorded on the caller’s current stream (and the legacyStream.nullwhen different) and waited on this pipeline’s own stream before the replays are issued.pre_launch(optional) is a zero-arg callable run on this pipeline’s own stream after the seed waits and before the replays – the seam for copying per-launch varying data into this pipeline’s buffers.seed_events(optional) are pre-recorded events to wait on instead of recording this pipeline’s own – the shared-event modeeagle.compose.GraphComposer._async_launch_group()needs, one event pair recorded once and waited on every member’s stream. An empty sequence means “I have already ordered this replay myself, issue no seed waits”. HAZARD 1 – one in-flight enqueue per pipeline. A secondenqueuebefore the first has been synchronized re-records the same cached events; overlapping enqueues simply serialize on this pipeline’s single stream, interleaving in a way the caller cannot control. Enqueue once, synchronize, then enqueue again.HAZARD 2 – the reverse write fence is gone.
enqueuereturns with the replay still in flight; the ordering events only say replay-after-seeds, nothing about writes issued after:pipe.enqueue() buf[...] = new_values # RACE: the in-flight replay may read these pipe.synchronize()
is a data race, silently. The discipline is:
enqueue-> (only work that does not touch this pipeline’s buffers) ->synchronize-> write; stage new inputs viapre_launchinstead. RaisesRuntimeErrorif called beforebuild().
- property graph#
- introspection_available()[source]#
Whether node introspection works here: the bound core dots the captured graph through
cudaGraphDebugDotPrint(always available), so this is ready as soon asbuild()has run.- Return type:
bool
- is_member_toggleable(handle_or_index)[source]#
Whether
set_member_enabled()will accept this member (its captured node set is entirely kernel/memcpy/memset). Requiresbuild()to have run.- Return type:
bool
- launch(n=1)[source]#
Replay the captured graph
ntimes.launch(n)guarantees all work previously enqueued on the caller’s current cupy stream and the legacy default stream is visible to replay 1.A stream-ordering fix. The pipeline’s own stream is
non_blocking=Trueprecisely so capture never entangles the legacy default stream ordinary cupy/torch launches land on – but that means replay is not otherwise sequenced after host-side seeding (buffer[...] = ...) done on the caller’s current or legacy stream. Under ambient GPU load the seed can lose the race and replay 1 reads pre-seed buffer contents. The fix: record a cupy Event on the caller’s current stream and on the legacyStream.null(when different) and wait on both, on the pipeline’s own stream, before the replay loop. This orders every consumer oflaunch()for free, at the facility seam, with no host block (~1-2us/call). Exactlyenqueue(n)followed bysynchronize()– same events, same order, same host block, same return value. The two halves are also available separately (seeenqueue()) for a caller that wants several pipelines’ replays in flight at once.
- member(name_or_index)[source]#
Resolve a name or registration-order index to this pipeline’s stable
MemberHandle. Available any time after registration; does not requirebuild()(onlyset_member_enabled()does, since toggleability is a build-time fact).- Return type:
MemberHandle
- member_nodes(handle_or_index)[source]#
Raw node handles (as ints) attributed to one registered member by
build()– the per-member analogue ofnode_labels(). A copy; mutating it has no effect. Requiresbuild()to have run.- Return type:
list
- node_labels()[source]#
Kernel function names of the captured nodes, in declaration order.
- Return type:
list
- set_member_enabled(handle_or_index, enabled)[source]#
Enable or disable one registered member’s launches on this instantiated graph, effective on the next
launch()– no recapture, no change to graph structure (composition mode “enabled”: capture once at full sibling concurrency, then flip driver-level enabled bits between replays). The host-paced sibling of the device-paced conditional (IF) node: where a conditional’s predicate is re-read by an in-graph kernel on every replay, this toggles a driver bit from Python between replays, for zero per-replay gate cost – the right tool when routing decisions are made in Python, not on-device. A member’s captured node set never changes afterbuild(); this only flips whether that fixed node set runs. A disabled member’s nodes become driver-level no-ops (any buffer it would have written is left untouched); an enabled member’s nodes run exactly as captured. Requires CUDA >= 12.3 (the same floor the conditional facility already imposes on the compiledeagle._coreextension). RaisesNonToggleableMemberErrorifbuild()recorded this member as non-toggleable (its node set contains something other than kernel/memcpy/memset – most commonly aSkippable’s conditional node).UnknownMemberErrorfor a bad name or index, andRuntimeErrorif called beforebuild().- Return type:
- Parameters:
enabled (bool)
- set_node_enabled(node, enabled)[source]#
Enable/disable ONE raw captured node directly, bypassing this pipeline’s own member registry. For a caller that tracks its own node attribution outside
add()/add_concurrent()’s member bookkeeping. The same underlying callset_member_enabled()makes per node (Launcher.set_node_enabled); this is only a narrower, public door to it for a caller already holding valid node handles (fromcapture_snapshot_nodes()).nodemust be a real handle from this pipeline’s own capture – passing a foreign or stale one is undefined at the driver level. Requiresbuild()to have run.- Return type:
- Parameters:
node (int)
enabled (bool)
- property stream: int#
This pipeline’s own
cudaStream_t, as a raw integer handle (the family’s sharedstreamname: the same value a DLPack consumer passes as__dlpack__(stream=...)). For capture composition: a caller that needs to bind a cuBLAS/cuSOLVER-style handle to this pipeline’s stream, or run warm-up passes beforebuild()captures, needs this exact stream – launching on any other stream would race the graph. Read-only, valid for this pipeline’s entire lifetime.
- class eagle.KernelRegistry(kernels=None)[source]#
Bases:
objectA
name -> launchablemap for resolving a launchable by name.Values are any launchable sharing the
__call__+launchprotocol. Dict-like:reg[name],name in reg,len(reg), iteration over names,.get/.names.
- exception eagle.LayoutWarning[source]#
Bases:
UserWarningA sample-major plane was copied into eagle’s component-major layout.
Filterable like any warning; make it an error with
warnings.filterwarnings("error", category=eagle.LayoutWarning).
- eagle.samples_first(x)[source]#
Mark
xas sample-major: its FIRST axis holds the samples,(N, w)(a matrix plane(N, R, C)). Zero-copy, honoured for any shape; passes through any eagle or hawk door that binds per-sample planes, and beats the door’slayout=. A shape that contradicts it is refused.
- eagle.samples_last(x)[source]#
Mark
xas component-major (native): its LAST axis holds the samples,(w, N). Zero-copy, honoured for any shape; passes through any eagle or hawk door and beats the door’slayout=. A shape that contradicts it is refused.
- class eagle.LoadedKernel(ptx_path, *, abi_kind)[source]#
Bases:
LaunchMixinCommon scaffolding for a PTX-loaded launcher (no nvcc at runtime): reads the shared sidecar fields, rejects a binary-ABI mismatch, and driver-loads the module. Subclasses read their own fields, build
self._allowed, and definelaunch/__call__over the inherited_dispatch(). The parsed sidecar is kept onself._meta.- abi_tag#
The artifact’s own ABI generation, read by
_refuse_v2_door()and byeagle.plan.plan’s legacy-partition bridge.
- param_names#
The binding names of
params.
- class eagle.LoadedPure(ptx_path)[source]#
Bases:
LoadedKernelA precompiled pure kernel loaded from a PTX (+ sidecar) artifact – no nvcc. Drop-in equivalent to the in-process compiled pure launcher via
eagle.launch. TheMutablebuffers are the handoff (read-modify-written in place), so there is no separateout=.- __call__(**kw)[source]#
Launch the pure kernel and return the updated
Mutablebuffers as a dict, in the caller’s framework.
- launch(*, grid=None, block=None, **kw)[source]#
Capturable single pure launch on the current stream (cupy arrays only). The writable
Mutablebuffers, any inputs, lookup tables, and theterminatedmask must be pre-allocated; read-modify-written in place.block=Nonedefers to eagle’s launch-policy resolver.
- class eagle.LoadedVector(ptx_path)[source]#
Bases:
LoadedKernelA precompiled vector kernel loaded from a PTX (+ sidecar) artifact – no nvcc.
Drop-in equivalent to the in-process compiled vector-kernel launcher.
- __call__(*, out=None, **kw)[source]#
Allocate, launch, and return the contribution, in the caller’s tensor framework (numpy in -> numpy out, cupy/torch -> same via DLPack): blocking for numpy, non-blocking on the framework’s stream otherwise. Pass
out=a pre-allocated device buffer to fill it in place – seeas_out_buffer().
- class eagle.RepeatWhile(step, guard, max_iters)[source]#
Bases:
objectA zero-arg step repeated on the device while a
SkipGuardholds, at mostmax_iterstimes per replay – the loop sibling ofSkippable(not a subclass: 0..N runs and 0..1 runs are different contracts).Under
GraphPipeline.build()the step becomes the body of one CUDA WHILE conditional node: a head kernel resets the[remaining, ran]cell and decides entry, the body runs, and a tail kernel decides continuation.launch(1)is onecudaGraphLaunchregardless of iteration count;launch(n)replays the whole loopntimes.max_itersis a plain Pythonint >= 1frozen at build as a kernel argument – a new cap needs a newbuild(). Called directly it is the host arm:while ran < max_iters and guard.evaluate(): step(), writing the same cell soiterations()reads the same in every mode. The cell lives in the guard’s array module (cupy or numpy), allocated at construction;iteration_indexexposes itsranword to body kernels as the device-visible iteration index (0 on the first iteration).Composition: the body
stepmay be aSkippable, or a tuple of such parts run in order (e.g.(k_steps, skippable(compact, guard))); aRepeatWhilemay be anadd_concurrentmember, but nesting one directly inside anotherRepeatWhileor aSkippableis refused atGraphPipeline.build(). One hidden inside an opaque body callable is not detected structurally: called mid-capture it runs its host arm, whose guard read syncs and fails the capture loudly.- Parameters:
guard (SkipGuard)
max_iters (int)
- step#
- guard#
- max_iters#
- property iteration_index#
One-element
uint32view of theranword – the 0-based index of the iteration in flight. Pass it to a body kernel that needs it; same array module as the guard.
- iterations()[source]#
Iterations the last run performed (host read of
ran; a device sync under a cupy guard, eager/debug paths only). 0 before any run.- Return type:
int
- property parts: tuple#
The body as a tuple of parts, run in order (a single-step body is a one-part tuple).
- class eagle.RunReport(n, finished, done, launches, steps, compactions, reorders, build_s, launches_by_k=<factory>, exact_steps=False, lane_utilisation=None, mode='band', probe=False)[source]#
Bases:
objectWhat one
Runner.run()did.finishedcounts samples whose mask is set when the run ends,doneisfinished == n,launchesthe loop’s iteration count, andstepsthe steps the loop ran (launches * K, or the summedkfor an automatic artifact).exact_stepssaysstepsis exact (true forK=1, or an automatic artifact that launched nothing, hit the cap early, or ended on a one-step launch).compactions/reorderscount what ran,build_sthe one-time capture wall (0 on host),launches_by_keachk’s launch count.lane_utilisation(active lanes / warp-iterations, in(0, 1]) is set only for a persist-entry run that counted it;Noneotherwise (every other mode, or a persist run that skipped the counter).modeis the launch shape this run actually took:"fused_one"/"persist"for a no-graph fast-path run,"band"for the WHILE-graph/policy path (the only mode before the fast path existed).probeis true for a fast-path run taken to time the entry the batch is not currently using (seeRunner._pick_entry()).- Parameters:
n (int)
finished (int)
done (bool)
launches (int)
steps (int)
compactions (int)
reorders (int)
build_s (float)
launches_by_k (dict)
exact_steps (bool)
lane_utilisation (float | None)
mode (str)
probe (bool)
-
exact_steps:
bool= False#
-
lane_utilisation:
float|None= None#
-
mode:
str= 'band'#
-
probe:
bool= False#
-
n:
int#
-
finished:
int#
-
done:
bool#
-
launches:
int#
-
steps:
int#
-
compactions:
int#
-
reorders:
int#
-
build_s:
float#
-
launches_by_k:
dict#
- class eagle.Runner(plan, *, max_steps, every=None, reorder=None, **planes)[source]#
Bases:
objectA bound, built run-until-done loop over one plan, or a step of several plans launched in order (see
until_done()).Attributes:
step(theeagle.plan.BoundPlan, or a tuple),loop,pipeline(the deviceeagle.GraphPipeline, built on firstrun()),active(theeagle.ActiveSet, orNone),finished/terminated,steps_per_launch(K, or"auto"),every/launches_per_compaction. For an automatic artifactfused_stepsis the device steps-per-launch word andpolicyreads the policy’s cells (a sync).host_policypicks how a host run paces automatic launches:"auto"(default:"tiled"when the entry’s fused side is tiled, else"measured"),"tiled"(steps_maxsteps per launch),"measured"(sweeps one step per launch until a timed fused probe wins) or"band"(the device’s band).host_threadsis the host team size:None(default) lets eagle pick it (one thread per physical core for a short run, every logical CPU for a long one — seeeagle._host_loop.host_team_size()), an int forces it. Planes are held by reference: a run writes the caller’s arrays in place; a new binding is a new runner.- step#
- loop#
- pipeline#
- active#
- finished#
- terminated#
- n#
- steps_per_launch#
- every#
- launches_per_compaction#
- fused_steps#
- auto#
- host_policy#
- host_threads#
- property policy: dict | None#
The automatic loop’s policy cells as host ints (a device sync):
k,finished_at_launch,steps_done,max_steps,go;Nonefor a fixed artifact.
- reset()[source]#
Clear the mask/counter/active set: the next
run()starts the batch over.- Return type:
None
- run()[source]#
Run until every sample is done (or the cap is reached) and report; the counter is re-synced to the mask first, so a run continues from the current state (a reordering set’s planes land back in sample order).
The counters this needs before/after the launch (not counting the automatic policy’s own
steps_done/histogram/fused_stepsreads, still separate – see_build_auto()) are read as ONE packed transfer per snapshot (_packed_read()) instead of a separate full round-trip sync each: a plain, non-compacting run (self.active is None) needs no PRE snapshot at all (_compactionscan only ever be 0 without an active set, so it is never even read), and the POST snapshot foldsiterations/finished/compactions/reordersinto one.A fast mode (
_build_fast()) skips all of the above: one launch IS the run, so_run_fast()handles it entirely.- Return type:
- class eagle.SimResult(*, state, finished, report, wall_s, allocated)[source]#
Bases:
objectWhat one
Simulation.run()produced.statemaps each state plane to its array (written in place, plus allocated ones, listed inallocated); each is also an attribute (result.x) and an item (result["x"]).finishedis the per-sample mask,donewhether every sample did,status"finished"/"max_steps",steps/report/wall_sthe loop’s steps, fulleagle.RunReportand wall time.- state#
- finished#
- done#
- steps#
- status#
- report#
- n#
- wall_s#
- allocated#
- class eagle.Simulation(model, args, kwargs, *, until=None, max_steps, every=None, reorder=None, scalar_type=None, layout=None)[source]#
Bases:
objectA bound, built simulation (see
simulation()):run()runs it,reset()starts the batch over.runneris theeagle.Runnerdriving the model,loop/pipelineits loop and deviceeagle.GraphPipeline(Noneon host),plansthe plans per kernel,nthe batch size,namesthestate/params/allocatedplane names.- runner#
- loop#
- plans#
- n#
- names#
- active#
- finished#
- terminated#
- every#
- property pipeline#
The device
eagle.GraphPipeline(built on the firstrun());Noneon the host and before.
- class eagle.SkipGuard(count_arr, count_idx=0, baseline_arr=None, baseline_idx=0)[source]#
Bases:
objectRegion runs iff
count_arr[count_idx] != (baseline_arr[baseline_idx] or 0). Arrays are uint32 device (cupy) or host (numpy) arrays; references are retained so pointer lifetime stays pinned to the pipeline.- count_arr#
- count_idx#
- baseline_arr#
- baseline_idx#
- classmethod intent(flags, index)[source]#
Router/user-owned intent flag:
flags[index] != 0.flagsis a uint32 array the caller owns and writes between graph replays, one lane per guarded region (e.g.SkipGuard.intent(expert_flags, e)for experte). No baseline: a plain on/off decision.Under
GraphPipeline.build()the predicate is read fresh from the device every replay, so flippingflags[index]takes effect with no recapture; in eager/numpy mode the same flag is read host-side byevaluate().indexhas no default (unlikenonzero()’s0): it is load-bearing at every call site.
- class eagle.Skippable(step, guard)[source]#
Bases:
objectA zero-arg step wrapped with a
SkipGuard. Works as aGraphPipeline.add()step or anadd_concurrentmember –GraphPipeline.build()recognizes it by type and weaves an IF node around it; called directly, it evaluates its guard eagerly:if guard.evaluate(): step().- Parameters:
guard (SkipGuard)
- step#
- guard#
- eagle.assemble_args(arg_spec, *, out, vec, per_sample, terminated, uniforms, n, tables=None, mutables=None, mutable_vec=None, mutable_mat=None, mats=None, wide_in=None, wide_out=None, plan=None)[source]#
Build the kernel argument tuple in the order the signature expects.
The mutable output, input vectors, and matrix inputs are by-value
GRefstructs (a matrix GRef composes the flat vector one); per-sample scalars, the termination mask, lookup tables, and a pure kernel’s scalar/intMutableslots are by-valueHandleTviews; broadcast constants arefloat64. A vector kernel’sGRefcarries its ownsamples_; a pure kernel takes an explicitnsamplesinstead.wide_in/wide_outare alsoHandleTviews, each riding its own caller-populated dict so they never collide with an unrelated per-sample-scalar or lookup-table name.The ABI-shape decision per role is
eagle.roles.classify_arg()(single-sourced with the ctypes host path); this function keeps only the construction – which dict a role’s value comes from, and whichmake_*helper builds it. An unrecognised role raises before any role-specific branch runs.plan(optional) is theLaunchPlanfor THIS signature. When given, the per-argument classification and POD boxes come from it instead of being re-derived here; the resulting argument list is byte-identical either way (readLaunchPlan’s box-reuse note first). Omitplanfor fresh boxes per call, as always.
- eagle.check_aether_abi(meta, *, kind, name)[source]#
Reject an artifact whose
meta["aether_abi"]is not one ofACCEPTED_ABI_TAGS(absent/empty included).kindnames the artifact in the error (e.g. ‘vector plugin’, ‘plugin manifest’).- Return type:
None- Parameters:
meta (dict)
kind (str)
name (str)
- eagle.compaction_body(step, active, *, every=16, finished=None, reorder=None, steps_per_call=1)[source]#
The
eagle.repeat_while()body that runsstepeverytimes, then compactsactive– only whenfinishedmoved, if given, else unconditionally – and, for a set built withreorder=thetaover at leastMIN_REORDER_SPANsamples, reorders when the compaction fired (reorder=False/Trueoverrides this). The loop guard is checked once pereverysteps, so it may run up toevery - 1steps past the last sample finishing; sizemax_itersasceil(max_steps / every).steps_per_callis the steps onestep()call advances (a fused kernel’sK): the body then runsevery * Ksteps between compactions, and a cadence belowMIN_EVERYsteps is refused.- Return type:
tuple- Parameters:
active (ActiveSet)
every (int)
steps_per_call (int)
- eagle.deploy(plugin, *, targets=None, cache_dir=None, scalar_type=None, **plan_kw)#
An
AutoPlanforplugin: run it where its data lives.eagle.deployis this same function.pluginis a built plugin, a hawk kernel, or a list built into one bundle (returned as a tuple of plans in order). The build uses hawk’s cache undercache_dir(None: hawk’s default); a repeat call compiles nothing.targets(None: both sides when a GPU is usable – the call returns with the device built, the host still building beside it – else("host",)),cache_dirandscalar_typeare refused for an already-built plugin.scalar_typeis the precision hawk builds in:"float64"(None, the default) or"float32"; the planes passed at run time must match it.plan_kwareplan()’s keywords butstructure(inner/gatherdo not apply); a side the plugin lacks is refused when first selected.
- eagle.detect(x)[source]#
Identify
x’s framework adapter without importing the framework.Nonefor a value we don’t recognize as a tensor (to_cupy()then treats it as a host upload).- Return type:
Adapter|None
- eagle.device_props(device=0)[source]#
Query device
device’s properties viaeagle._core.device_props()and wrap the resulting dict in a frozenDeviceProps— no Python-side computation, only a dict-to-dataclass reshape.eagle._core(the compiled extension) is imported LOCALLY, here, never at this module’s top — importingeagleon a cuda-free box must not force-load it (the same disciplineeagle.pipeline. GraphPipeline.__init__()follows for the same reason).- Return type:
- Parameters:
device (int)
- eagle.launch(kind, fn, arg_spec, vector_inputs, params, per_sample, *, kw, grid, block, dt=<class 'numpy.float64'>, **extra)[source]#
Issue ONE capturable launch of any kernel kind on the current stream.
kind("vector"|"pure") selects the ABI binding shape; kind-specific inputs rideextra(a vector kernel’sout=, or a pure kernel’smutables_decl=, each with optionalmatrix_inputs=/wide_inputs=/wide_outputs=/wide_out_exempt=, named identically topure_prepare()’s). Every array argument must be pre-allocated cupy. Returns the vectorout, or the pureMutabledict merged with wide outputs.
- eagle.load_manifest(path)[source]#
Load a deployed
manifest.jsonby its top-levelpattern.The Python analogue of the C++
PluginRegistry::from_manifest– but where that injects the whole ordered set, this returns a by-name result:"vector"/"pure"– each enabled entry’s artifact is path-loaded intoLoadedVector/LoadedPureand keyed by its manifestidinto aKernelRegistry."neural_block"– a neural deployment unit: the inlineblocks[]compose one ordered layer whose exec references resolve against theplugins[]kernels, returned as a frozen neural-layer wrapper. Recognized but not launch-certified, so it never enters the by-name map.
An unknown or absent
patternis a hard error. ReturnsKernelRegistry | LoadedNeuralLayer.
- eagle.make_gref(arr, n)[source]#
Build a by-value
GRefstruct over a contiguous(3, N)cupy array:compStride_ == samples_ == n,sampleStride_ == 1.
- eagle.make_handle(arr, n)[source]#
Build a by-value scalar-array
HandleTover a contiguous(N,)array;nis required since the struct carries its own extent.
- eagle.np_dtype(scalar_type)[source]#
Return the numpy dtype of Real-typed arrays for a resolved
scalar_type.SoftDouble is bit-identical IEEE float64 — the array interface is float64 (mirrors
eagle.dtypes.np_dtype’stableliteral).- Parameters:
scalar_type (str)
- eagle.origin_adapter(candidates)[source]#
Pick the adapter the kernel’s output should be returned as, over the state-vector inputs in signature order: the first strong framework (torch/jax/tensorflow) wins (two distinct ones raise
TypeError), else the first device adapter, else numpy.- Return type:
Adapter
- eagle.pure_prepare(fn, arg_spec, *, vector_inputs, per_sample, params, mutables_decl, mutable_defaults, lookup_counts, kw, mat_shapes=None, vec_widths=None, wide_inputs=None, wide_outputs=None, wide_out_exempt=(), dt=<class 'numpy.float64'>, readonly_mask=False)[source]#
Coerce inputs, bind the writable
Mutables, assemble args, launch (no sync). The pure allocate-and-launch body, shared by the in-process compiled pure launcher and the deployedLoadedPureso the two can never drift.State vectors ->
(3, N)f64; matrix inputs -> flat(R*C, N)f64; providedMutablearrays -> device buffers (also sizing N); per-sample scalars ->(N,); lookup tables -> flat handles; wide inputs -> 2-D(rows, N); wide gradient outputs -> 2-D(rows, N)written in place, never auto-allocated.wide_out_exemptnames thewide_outputsthat are row-indexed destinations (e.g. an atomicAccumoutput), skipping N-reconciliation. AMutableomitted fromkwis a broadcast scalar fill or its declared default. Returns theMutablebuffers merged with any wide gradient outputs (RMW in place). Meant to run inside the caller’s launch context (seeLaunchMixin._dispatch()).readonly_mask(defaultFalse) is the sidecar-declared opt-in letting an omittedterminatedmask be served from the cached all-false mask (seecoerce_terminated()); never inferred, since marshal cannot read a kernel’s source.
- eagle.repeat_while(step, guard, max_iters)[source]#
Repeat
step(a zero-arg callable, aSkippable, or a tuple of such parts) whileguardholds, at mostmax_iterstimes: one device-side WHILE node underGraphPipeline.build(), a host loop everywhere else.- Return type:
- Parameters:
guard (SkipGuard)
max_iters (int)
- eagle.run_until_done(plan, *, max_steps, every=None, reorder=None, **planes)[source]#
until_done()thenRunner.run(): the one-call form.- Return type:
- Parameters:
max_steps (int)
every (int | None)
- eagle.simulate(model, *args, until=None, max_steps, every=None, reorder=None, scalar_type=None, layout=None, **kwargs)[source]#
Run
modelon every sample until each one finishes (ormax_stepsis reached) and return theSimResult.modelis a hawk kernel, a list of them (one step, in list order), or plans fromeagle.deploy();untilis the kernel that finishes the samples, placed last (at least one kernel must finish). The rest of the call is the kernel’s own:*args/**kwargsbind to its parameters exactly as a call to the kernel function would (positionally, by name, or both), by its own annotations – except that a model built from a PREBUILT plan (no raw kernel to read a call signature from) refuses*argsoutright, naming the keyword call to use instead; keyword binding always works. AMutableplane is the state it updates (its initial value; one never read before it is written may be left out, zero-filled instead), aParama number shared by every sample, aScalar/Vector/Tableplane one value per sample (a number where the kernel allows it, an array otherwise);Terminatedis never passed. numpy/numbers run on the CPU, cupy on the GPU.max_stepsis required;every/reorderareeagle.until_done()’s compaction cadence and reorder threshold,scalar_typethe precision the kernels are built in ("float64", the default, or"float32"; seeeagle.deploy()),layout("samples_first"or"samples_last") which axis holds the samples of a per-sample plane whose shape reads both ways ((w, w)), which is otherwise refused;eagle.samples_first(x)/eagle.samples_last(x)say it for one array and win. All keyword-only like every one of eagle’s own options – after the kernel’s own arguments.- Return type:
- Parameters:
max_steps (int)
every (int | None)
- eagle.simulation(model, *args, until=None, max_steps, every=None, reorder=None, scalar_type=None, layout=None, **kwargs)[source]#
Bind and build
modelonce; returns theSimulation, whoserun()runs it (again, afterreset(), without rebuilding).*args/**kwargsare the kernel’s own, bound exactly like a call to it; the rest aresimulate()’s own options.- Return type:
- Parameters:
max_steps (int)
every (int | None)
- eagle.skippable(step, guard)[source]#
Wrap
step(a zero-arg callable) so it only runs whileguardholds: one IF node underGraphPipeline.build(), an eager guard check everywhere else.
- eagle.to_cupy(x)[source]#
Import a framework / DLPack tensor to a cupy array: zero-copy for a CUDA-resident producer, a host upload otherwise. Also the pre-import helper for the capturable
launchpath, since a DLPack import can synchronize and is unsafe mid-capture.
- eagle.until_done(plan, *, max_steps, every=None, reorder=None, **planes)[source]#
Bind
planand build the loop that runs it until every sample is done; returns theRunner(.run()runs it).**planesbinds every name of the plan’sarg_specbyeagle.plan.Plan.bind()’s rules, except the reservedFINISHED_PLANE,active_mapandactive_count(the runner allocates these); theterminatedmask is all-false when omitted.planmay be aneagle.plan.AutoPlan(residency picks host/device), or a list of plans run as one step, in order, over one namespace of planes; each takes one step per launch, at least one finishes the samples, and they share one mask and guard.max_stepscaps the steps a sample may take (a compacting loop may overrun by up toevery - 1steps;RunReport.stepscounts them).Guard(active_set=True)compacts everyeverysteps (a multiple ofK, at least 4, default the smallest at least 16), andreorder=thetaalso reorders the bound planes; both are refused on a plain artifact.An automatic artifact (
steps="auto") runs the policy loop instead (see the module docstring): the cap is exact,every=is refused, reorder only withreorder=theta, and the runner owns theFUSED_STEPS_PLANEword.- Return type:
- Parameters:
max_steps (int)
every (int | None)