Python API reference#
Complete reference for the eagle Python package, produced by Sphinx
autodoc/autosummary from the live package — the signatures shown here
are the ones actually exported by the build you have installed.
eagle owns everything about executing compiled kernels: the
framework-polymorphic launch skeleton, the capturable kind-dispatched
launch primitive, input-role marshalling, the by-value launch ABI, DLPack
framework adapters, CUDA-graph capture/replay, composing several already-built
launchables into one launch list, and the plugin protocol consumers
(KernelRegistry, load_manifest()).
The top-level namespace#
eagle/__init__.py re-exports its whole intended public surface — every
launcher, every DLPack adapter, the by-value ABI helpers, the plugin
registry — directly as eagle.<name>, so this one page is a complete
single-page reference for everything a typical caller needs (see
Thirty seconds: eagle.simulate, Interoperability).
eagle — the Python face of the EAGLE launch engine. |
eagle.deploy is bound on first use (so import eagle imports no hawk)
and is the same function as eagle.plan.auto():
- eagle.deploy(plugin, *, targets=None, cache_dir=None, scalar_type=None, **plan_kw)
An
AutoPlanforplugin: run it where its data lives.eagle.deployis this same function.pluginis a built plugin, a hawk kernel, or a list built into one bundle (returned as a tuple of plans in order). The build uses hawk’s cache undercache_dir(None: hawk’s default); a repeat call compiles nothing.targets(None: both sides when a GPU is usable – the call returns with the device built, the host still building beside it – else("host",)),cache_dirandscalar_typeare refused for an already-built plugin.scalar_typeis the precision hawk builds in:"float64"(None, the default) or"float32"; the planes passed at run time must match it.plan_kwareplan()’s keywords butstructure(inner/gatherdo not apply); a side the plugin lacks is refused when first selected.
Submodules not re-exported at the top level#
A handful of names are reached by their submodule path rather than directly
off eagle — mostly the newer aether-abi/2 bind-by-name door
(eagle.plan, eagle.exec), the shared input-role coercion layer,
and the GPU/CPU device-dispatch internals.
eagle.plan — PLAN in Python, EXECUTION in C++.
plan() resolves a v2 plugin, an execution structure (always named
explicitly, never residency-derived) and a partitioning into a Plan.
Placement legality (eagle.exec.check_placement()) is checked here, at
plan time, against every partition count the plan will launch.
Plan.run() drives the plugin over every partition and returns the
per-sample result, marshalling each role to the ABI shape
eagle.roles.classify_arg() resolves it to — the same classifier the v1
launch paths dispatch on. Output planes are allocated in the artifact’s
declared scalar_type (eagle.dtypes); a single-output plugin
returns the plane itself, never a 1-tuple, and several return a dict
keyed by plane NAME — never a positional tuple, since arg_spec’s own
order is role-grouped then name-sorted, not the kernel’s authored order.
Plan.bind() is the capture-legal door: it packs the caller’s own named
planes into the argument block ONCE and returns a BoundPlan whose
launch() issues the structure’s launches and nothing
else — no upload, allocation, sync or copy home — so a consumer can bind
once and replay the launches as ordinary nodes inside a captured graph.
Under eagle.exec.RankPartition the plan’s partitions stay WHOLE and
the rank cut happens inside the structure, so the same plan describes the
run at any world size; every rank returns the whole output plane(s).
Both doors share one packer (_pack_args() over _pack_one()), so
they cannot disagree about how a role packs.
- class eagle.plan.AutoPlan(plugin, plan_kw, host=None)[source]#
Bases:
objectA plugin with both targets, run where its data lives — the result of
auto().host/deviceare thePlanfor each target, built lazily;select()picks one from the values a call binds (host- or device-resident; the sample count never enters).run()/bind()delegate to the selected plan. A default deploy with a usable GPU returns with the device side built and the host side still building.- Parameters:
plan_kw (dict)
- plugin#
- plan_kw#
- bind(**planes)[source]#
Plan.bind()on the planselect()picks forplanes.- Return type:
- property device: Plan#
The
DeviceKernelplan (built on first use).
- run(**kw)[source]#
Plan.run()on the planselect()picks forkw.
- select(planes)[source]#
The plan
planesselects (residency()): the device when any is device-resident, else the host; both sides refused.- Return type:
- Parameters:
planes (dict)
- class eagle.plan.BoundPlan(plan, arg_spec, planes, n, parts, boxes, addrs, entry, *, device_type, framework, dtype, widths, vec_mutables, mat_mutables, int_uniforms=frozenset({}), active_count=None, copies=())[source]#
Bases:
objectA
Planwith its argument block already packed over the caller’s own planes — the thing a captured consumer holds.It owns the packed host-side mirror block and a reference to every plane those mirrors point at (a captured graph holds raw device pointers, so a plane freed too early reads as a correct answer once and somebody else’s memory afterwards), but no plane itself: nothing here is allocated, uploaded, cast or copied home, at bind or at launch.
The bound block is a snapshot, frozen at bind;
rebind()changes only the slots it is given.- plan#
- device_capacity()[source]#
The safe “no active-set rebuild” bound at this bound plan’s kernel/device (see
eagle._launch_policy.latency_regime_capacity()— the smaller of hardware residency and the issue-pipe saturation point, NOT residency alone).0for a host-team plan, or when device/kernel properties are unavailable; the caller (eagle._until_done’s auto policy) treats0as “unknown”, which never satisfiesn <= capacity. By module name (not a top-level import), matchinglaunch()’s own lazyeagle.launchlookup – this stays import-cycle-free.- Return type:
int
- launch(stream=None, *, grid=None, block=None)[source]#
Issue this plan’s launches — one per partition — and nothing else. Returns
None: the answers are already in the caller’s own planes.streamis the CUDA stream to issue on (an int, or anything carrying.ptr);Noneresolves to the framework’s current stream at launch time, not at bind, since the stream a graph captures on is not the one the plan was bound on. A host_team run has no stream and is refused one.gridis refused (eagle derives it from each partition’scount);block=Nonedefers to eagle’s launch-policy resolver (eagle.launch._resolved_block), the same one the v1 door consults.- Return type:
None
- launch_persist(stream=None, *, grid, counter, util=None, block=256, steps=None, stepsum=None, prepared=None)[source]#
Launch
<kernel>_persistonce – the persistent launch’s whole shape: lanes fetchbase + atomicAdd(counter, 1)until the partition’s owncount, an already-Terminated sample skipped, each running up to the boundfused_stepsword’s budget before writing back and fetching the next; no WHILE graph, no active-set map, no policy kernel – this one call is the whole run.gridis the block count andblockthe block size – the caller’s own choice, both geometry knobs this method does not resolve itself (eagle._launch_policy.persistent_geometry()is the documented rule a caller derives them from, from the entry’s own register-limited occupancy and the SM count;block=256here is only a bare fallback for a caller that does not).counteris a one-element uint32 cupy array the caller has already memset to 0 (not done here, so a caller replaying this inside a captured graph can fold the reset into a graph node instead).utilis an optional 2-element uint64 cupy array (active lanes, warp-iterations*32); omitted, the entry gets a null pointer and skips counting.stepsis an optional one-element uint32 cupy array the caller has zeroed: the entry maxes each sample’s steps taken in this launch into it (the run’s exact step count); omitted, a null pointer and no report. The entry’s trailing ABI isbase, count, nSamples, hawk_next, hawk_util, hawk_steps, hawk_stepsum;stepsumis an optional one-element uint64 cupy array the caller has zeroed, into which the entry adds every step it executes (the run’s total; omitted, a null pointer and no report).Refused for anything but a single, contiguous partition – the only shape the persist entry’s
base/countpair means.prepared(fromprepare_persist()) skips the checks and the packing and only launches: the other arguments are then unread.- Return type:
None- Parameters:
block (int)
- launch_range(stream=None, *, block=None, steps=None, stepsum=None, prepared=None)[source]#
Launch
<kernel>_rangeonce over this bound plan’s single contiguous partition – an active-set automatic kernel’s fast-path equivalent of the plain kernel’s ownlaunch(): the SAME packed argument block the map entry uses (active_map/active_countincluded, unread), plusbase, count, nSamples, hawk_steps, hawk_stepsum; no counter, no util (that is the persist entry’s own shape,launch_persist()).stepsis an optional one-element uint32 cupy array the caller has zeroed, into which the entry maxes each sample’s steps taken (the launch’s exact step count); omitted, a null pointer and no report.stepsumis the persist entry’s one-element uint64 total of executed steps (launch_persist()). The entry itself picks single-step vs fused from the boundfused_stepsword (1 vs anything else); this call does not choose.block(None: resolved) takes the SAME policylaunch()already resolves its own block with (eagle.launch._resolved_block()->eagle._launch_policy.resolve_block, keyed off this plan’s own entry – the range entry shares its cubin/registers, so the same block size applies); the grid followseagle.exec.DeviceKernel’s own grid-from-count rule, overcountsamples densely (one lane per sample, no work-stealing). A count of 0 launches nothing.Refused for anything but a single, contiguous partition – the only shape the range entry’s
base/countpair means.prepared(fromprepare_range()) skips the checks, the block/grid resolution and the packing and only launches.- Return type:
None
- property n: int#
The sample count this block was packed for (frozen at bind).
- property names: tuple#
The bound names, in
arg_specorder.
- property partitions: tuple#
The
eagle.exec.Partitiontripleslaunch()issues, in order (resolved at bind).
- persist_entry()[source]#
This bound plan’s
<kernel>_persistcupyFunction– same compiled unit as the fused entry (hawk’shawk.emit.cuda.persist_entry).Nonewhen unavailable: a host-team plan, a non-automatic kernel, or an artifact built before the persist entry existed. Resolved once, from the loaded device module the plugin’s_keepalivealready holds – no new compile, no new module load.
- property planes#
the arrays the packed block points at (or, for a sample-major plane, the component-major view that was packed).
- Type:
The bound values by name (read-only)
- prepare_persist(*, grid, counter, util=None, block=256, steps=None, stepsum=None)[source]#
launch_persist()’s checks and argument packing done ONCE: returnscall(stream=None)that only issues the driver launch (a resident runner’s per-run cost). The packed block is re-made ifrebind()ran since.- Parameters:
block (int)
- prepare_range(*, block=None, steps=None, stepsum=None)[source]#
launch_range()’s checks, block/grid resolution and argument packing done ONCE: returnscall(stream=None)that only issues the driver launch (a count of 0 launches nothing). The packed block is re-made ifrebind()ran since.
- range_entry()[source]#
This bound plan’s
<kernel>_rangecupyFunction– an automatic kernel’s contiguous-range sibling entry (same cubin as the kernel’s own entry; for an active-set kernelactive_map/active_countride along in the SAME packed argument block, bound but never read), which also reports the launch’s exact step count.Nonewhen unavailable: a host-team plan, a non-automatic kernel, or a unit built without hawk’sHAWK_FAST_ENTRIES=1define. Resolved once, same lazy pattern aspersist_entry().
- rebind(**changed)[source]#
Re-pack only the named slots — one box per name, in one call: for a resident pipeline whose planes move (a double-buffered state swaps) while the plan, partitions and sample count stay the same.
selfis positional-only, asPlan.run’s andPlan.bind’s are.A rebound plane is checked exactly as the original was, at the same frozen
n— a different length is a different run and must be bound afresh. Returns this sameBoundPlan, updated in place; it cannot change a launch already recorded into a graph, only what the NEXT capture or launch sees.
- class eagle.plan.Plan(plugin, structure, access, op, partitions, npartitions, n_samples, inner=None, gather=True)[source]#
Bases:
objectWHERE and HOW
pluginruns — the result ofplan().partitionsis the resolved, explicit tuple ofeagle.exec.Partitiontriples the plan will launch (contiguous, covering[0, n_samples)), orNonewhen deferred: nopartitions=/n_samples=was given toplan(), so the shape is resolved from the sample countrun()observes. Placement legality is already checked byplan();run()never re-checks it.- Parameters:
-
plugin:
object#
-
access:
str#
-
op:
str|None#
-
partitions:
tuple|None#
-
npartitions:
int#
-
n_samples:
int|None#
-
inner:
Structure|None= None# The structure each rank’s share runs through under
eagle.exec.RankPartition(defaultHostTeam);Nonefor every other structure, which has no inner.
-
gather:
bool= True# Whether a
eagle.exec.RankPartitionrun gathers its output planes before returning (defaultTrue, every rank holding the whole plane);Falseis the pre-gather arm, for asserting a rank’s own sub-partition before the gather overwrites it.
- partitions_for(n)[source]#
The concrete
eagle.exec.Partitiontuple this plan runs over a sample count ofn: the resolvedpartitionswhen explicit, else a fresh even split intonpartitionscontiguous partitions.- Return type:
tuple- Parameters:
n (int)
- run(**kw)[source]#
Drive
pluginover every partition throughstructureand return the assembled per-sample result: the declared output plane as a numpy array for a single-output plugin, or adictkeyed by plane NAME when it declares several (never a positional tuple — the same shapeSimResult/eagle.loaded.LoadedPure()already return).Device inputs stay on the device: on a
DeviceKernelplan, an input already in device memory (cupy, or a CUDA torch tensor) is bound where it lies, and the outputs then come back as cupy arrays, never brought home — a caller-supplied contiguous output of the declared dtype is the caller’s own buffer, written in place. With host inputs only, the outputs come home as numpy arrays. A rank-partitioned run always gathers on the host.Every output/mutable plane the caller supplies is written into in place, in the caller’s own layout (
eagle._layout), and still returned. A read-only supplied array (flags.writeable is False, or a torch tensor withrequires_grad=True) raises, naming the argument. An output not supplied is allocated fresh and returned.A value shaped as a plane’s per-sample head (a number or 0-d array for a scalar,
(w,)for a vector) is one sample (eagle._layout): the call runs as a batch of one and every output comes back in its head shape. Mixing one sample with a batch is refused.An automatic kernel (it reads the reserved
fused_stepsword) launched here without that word takes exactly one step (one_step_word()); onlyeagle.until_done()takes several steps per launch.selfis positional-only, so a plugin may declare a plane literally namedselfwithout colliding with this method’s own bound-method parameter.
- bind(**planes)[source]#
Bind the caller’s own planes by name and pack the argument block once — the capture-legal door (see the module docstring).
Every name this plugin’s
arg_specdeclares must be given, as the thing the plan’s structure can address: a cupy array for aeagle.exec.DeviceKernelplan, numpy forHostTeam, a plain scalar for auniform(frozen into its by-value box here).nsamplesis derived, never bound. The reservedfused_stepsword of an automatic kernel may be left out: a one-step word is then bound for you (one_step_word());eagle.until_done()binds its own.Nothing is allocated, coerced or copied, with one exception: a sample-major
(N, w)plane (eagle._layout) binds zero-copy when its transpose is C-contiguous, else is copied here into a component-major buffer thisBoundPlanowns (with aneagle.LayoutWarning) and refreshes before each launch, copying back after for a written plane — capture-legal, no device-wide sync. The output planes are always the caller’s own: this door cannot allocate one, since it would have to outlive the graph that writes it with no handle for the caller to hold.What’s checked here (every refusal names the field): every declared name is bound and no stray one is; each plane is an array of the structure’s own framework (a host array to a device plan is refused, not uploaded); it is C-contiguous; its dtype is the artifact’s declared
scalar_type(an integer/bool plane passes at its own dtype, asrun()does); and its shape agrees with what the artifact declared. An input plane’s component width is NOT checked: a v2 sidecar declares widths for its output planes only.selfis positional-only for the same reasonrun()’s is.- Return type:
- eagle.plan.auto(plugin, *, targets=None, cache_dir=None, scalar_type=None, **plan_kw)[source]#
An
AutoPlanforplugin: run it where its data lives.eagle.deployis this same function.pluginis a built plugin, a hawk kernel, or a list built into one bundle (returned as a tuple of plans in order). The build uses hawk’s cache undercache_dir(None: hawk’s default); a repeat call compiles nothing.targets(None: both sides when a GPU is usable – the call returns with the device built, the host still building beside it – else("host",)),cache_dirandscalar_typeare refused for an already-built plugin.scalar_typeis the precision hawk builds in:"float64"(None, the default) or"float32"; the planes passed at run time must match it.plan_kwareplan()’s keywords butstructure(inner/gatherdo not apply); a side the plugin lacks is refused when first selected.
- eagle.plan.device_resident(value)[source]#
Whether
valueis already in device memory (a cupy array, or anything exporting__cuda_array_interface__); a CPU or grad-tracked tensor raises there, so both answerFalse.- Return type:
bool
- eagle.plan.one_step_word(arg_spec, planes, framework='numpy')[source]#
planesplus a one-cellint64word of1underfused_stepswhen the artifact reads that word and the caller binds none, so a launch outsideeagle.until_done()takes exactly one step.frameworkis where the word lives ("cupy"for a device bind,"numpy"otherwise). Any other artifact getsplanesback unchanged.- Return type:
dict- Parameters:
planes (dict)
framework (str)
- eagle.plan.plan(plugin, *, structure, inner=None, npartitions=1, partitions=None, n_samples=None, _exec_access=None, _exec_op=None, _gather=True)[source]#
Build a
Planfor runningpluginunderstructure(always named explicitly, never residency-derived).structureis one ofeagle.exec.DeviceKernel/HostTeam/RankPartition/DeviceGroup.innernames the structure each rank’s share runs through underRankPartition(HostTeamby default) and is refused for any other structure. The plan’s partitions stay WHOLE either way — the rank cut happens inside the structure, so the same plan describes the run at any world size.partitions(a list of(base, count, n_samples)triples) takes precedence overnpartitionsand is validated for contiguity/coverage; otherwise the plan is deferred to a whole view (npartitions=1, the default) or an even split, resolved oncePlan.run()observes the sample count. The execution axis is the plugin’s own.exec_access/.exec_opdeclaration.Placement legality is checked here, at plan time, against the partition count this plan will launch. A refused plan names the rule.
A plugin declaring
.abi_tag == eagle.exec.ABI_TAG_V1(a legacy v1LoadedVector/LoadedPure) carries no partition triple, so a partition count above 1 is refused naming “legacy”.- Return type:
- Parameters:
npartitions (int)
n_samples (int | None)
_exec_access (str | None)
_exec_op (str | None)
_gather (bool)
- eagle.plan.residency(planes, *, door='eagle.plan.auto', what='kernel')[source]#
Where
planessays the work runs:"device"when any value is device-resident, else"host"; both sides refused, naming one of each (door/whatname the caller).- Return type:
str- Parameters:
door (str)
what (str)
eagle’s execution structures: the thin Python face over the compiled
eagle._core aether-abi/2 primitives (Partition, run_device /
run_host / run_host_serial, check_placement, fold, and the
layout self-check).
This is the EXECUTION half; eagle.plan marshals a plugin’s
arg_spec into by-value ABI structs and imports this module rather than
duplicating it.
RankPartition alone lives in a second, optional compiled extension,
eagle._mpi (built only with -DEAGLE_PYTHON_MPI), imported lazily so
an MPI-less install never pays for it and its absence is reported as a
RuntimeError naming the build switch, never a bare ImportError.
- class eagle.exec.Partition(*args, **kwargs)#
Bases:
objectThe launch partition triple
{base, count, nSamples}, re-exported from the compiled binding.- property base#
(self) -> int
- property count#
(self) -> int
- is_whole#
Is this the whole view – the only shape a legacy aether-abi/1 plugin may be handed?
- property n_samples#
(self) -> int
- whole = <nanobind.nb_func object>#
- eagle.exec.ABI_TAG_V1 = 'aether-abi/1'#
The two ABI generations’ wire tags (byte-identical to
eagle.abi.ABI_TAG_V1/eagle.abi.ABI_TAG_V2).
- eagle.exec.LAYOUT_SYMBOL = 'eagle_layout_sizes'#
The exported symbol name a v2 plugin’s layout self-check carries.
- class eagle.exec.Structure(name)[source]#
Bases:
objectOne of eagle’s execution structures:
device_kernel/host_team/rank_partition/device_group..nameis the wire spelling. Instances are the singletons below – never construct one directly.- Parameters:
name (str)
- name#
- eagle.exec.DeviceKernel = eagle.exec.device_kernel#
The four execution structures.
DeviceGroupis a named placeholder whose.runrefuses “NCCL not implemented”.
- eagle.exec.check_placement(access, structure, npartitions=1)[source]#
Raise
ValueErrornaming the rule if running a body of declaredaccessunderstructureovernpartitionsis illegal.structureis aStructuresingleton or its wire name.- Return type:
None- Parameters:
access (str)
npartitions (int)
- eagle.exec.fold(op, values)[source]#
Combine
valuesunderop(sum/times/max/land) in eagle’s fixed ascending order – never atomics, so the result is reproducible regardless of how the run was partitioned.- Return type:
float- Parameters:
op (str)
- eagle.exec.layout_sizes()[source]#
This build’s layout self-check sizes, in
LAYOUT_SYMBOLorder.- Return type:
list[int]
- eagle.exec.check_layout_sizes(sizes)[source]#
Refuse a v2 plugin’s layout self-check if it disagrees with this build’s own
layout_sizes(), naming the field.A tag alone cannot catch a layout mismatch (a wrong-arch or stale-binding artifact can carry a correct
aether_abitag); this is what catches it instead.sizesis either a partial{field_name: byte_size}dict (only the given keys are checked) or the raw positional sequence a real plugin’s exportedLAYOUT_SYMBOLcarries (every position checked).Raises
ValueErrornaming the disagreeing field, or a length mismatch for the positional form.- Return type:
None
Input-role marshalling for every kernel launcher.
The single home for coercing a launch call’s inputs into the contiguous
device arrays / by-value scalars the aether GRef/HandleT ABI expects: one
implementation per input role (state vector, matrix, per-sample scalar,
lookup table, broadcast uniform, writable Mutable, the terminated
mask, wide per-sample buffer/gradient), shared verbatim by every launch
path so the input contract can never drift between kinds.
Each foreign tensor is imported to cupy through eagle.interop
(zero-copy DLPack for a CUDA-resident tensor, a host upload otherwise); the
cast/contiguate after that is a copy only when not already the target dtype
/ C-contiguous. The launchable’s producer resolves scalar_type to the
Real dtype dt threaded through here (eagle.dtypes.np_dtype()).
- eagle.marshal.require_n(n, *, hint='state-vector input')[source]#
The canonical ‘cannot size the batch’ error, with a kind-appropriate hint.
- eagle.marshal.coerce_vec_inputs(kw, names, n=None, widths=None, dt=<class 'numpy.float64'>)[source]#
State vectors -> contiguous
(W, N)dtdevice arrays; returns(dict, n)with N reconciled.widthsmaps a name to its component count (default 3; a synthesized derivative seed may differ).
- eagle.marshal.coerce_wide_inputs(kw, names, n=None, dt=<class 'numpy.float64'>)[source]#
Wide per-sample flat inputs -> contiguous 2-D
(rows, N)dtdevice arrays; returns(dict, n). Unlike a fixed-width state vector, a wide buffer’s true row count (stride * n_in) is a runtime quantity the sidecar never carries, so only the 2-D shape is checked – the caller must match its ownn_in.
- eagle.marshal.coerce_wide_outputs(kw, names, n=None, dt=<class 'numpy.float64'>, exempt=())[source]#
Wide gradient-scatter outputs (a VJP-derived kernel’s disjoint scatter targets) -> contiguous 2-D
(rows, N)dtdevice arrays, written in place; returns(dict, n). Never auto-allocated (the row count is a runtime quantity, as incoerce_wide_inputs()) – the caller must pass a buffer.exemptnames a declared row-indexed destination (e.g. an atomicAccumoutput’s(rows, 1)buffer) that carries no batch axis, so it binds as-is and skips_reconcile_n.
- eagle.marshal.coerce_mat_inputs(kw, shapes, n=None, dt=<class 'numpy.float64'>)[source]#
Per-sample matrix inputs -> contiguous flat
(R*C, N)dtdevice arrays; returns(dict, n).shapesmaps each name to its declared(R, C); the caller passes either(R, C, N)or flat(R*C, N), both reshaping to the same row-major SoA buffer.
- eagle.marshal.coerce_per_sample(kw, names, n=None, dt=<class 'numpy.float64'>)[source]#
Per-sample scalars -> contiguous
(N,)dtdevice arrays (strict: a 2-D input is rejected, never silently flattened); returns(dict, n).
- eagle.marshal.coerce_tables(kw, lookup_counts, dt=<class 'numpy.float64'>)[source]#
Lookup tables / shared constants -> flat
(count,)dthandles. The row-major flatten is the binding contract (the kernel readshandle[flat_index]); the count is checked against the declaration.
- eagle.marshal.coerce_int_uniform(name, value)[source]#
One
intbroadcast param -> an exactnp.int64, never throughdt(value): a float/double route is bit-exact only below 2^53. Not an exact integer (2.5) is refused rather than truncated;2.0is accepted, matchingfill_mutable().
- eagle.marshal.coerce_uniforms(kw, params, dt=<class 'numpy.float64'>)[source]#
Broadcast
Paramconstants -> their declared by-value type.paramsmay be bare names (v1, all float) or decl-carrying entries (eagle.roles.param_decls()normalizes both); anintparam bindsnp.int64exactly (coerce_int_uniform()), neverdt.
- eagle.marshal.coerce_mutable(name, value, decl, dt=<class 'numpy.float64'>, adapted=None)[source]#
Coerce a provided per-sample
Mutablearray to its device buffer; returns(device_array, n). Afloat/intslot is contiguous(N,)(intalwaysint64); avectorslot a(W, N)SoA array; amatrixslot bound flat(R*C, N)(a contiguous(R, C, N)input reshapes zero-copy, so the kernel’s update lands in the caller’s buffer).A sample-major
vector/matrixslot is adapted first (eagle._layout); whenadaptedis a dict its record is stored there by name, so the caller can write the result back.
- eagle.marshal.fill_mutable(name, value, decl, n, dt=<class 'numpy.float64'>)[source]#
A broadcast / default
Mutablebuffer: a scalarvaluefilled to(N,)asdt/int64. A vector/matrix slot has no scalar fill – it must be passed as an array.
- eagle.marshal.ZERO_MASK_CACHE_CAP = 64#
How many distinct
(device, n)masks may be retained (each pinsnbytes of device memory for the process’s life). Past the cap the cache bypasses – a fresh private mask, same as an undeclared plugin gets.
- eagle.marshal.coerce_terminated(kw, n, *, readonly_mask=False)[source]#
The
terminatedmask -> a(N,)bool device array (all-false when omitted).readonly_mask(defaultFalse) is the sidecar-declared opt-in: whenTrueand the mask is omitted, it is served from the per-(device, n)cache. A passed mask is never cached.
The ONE shared schema-v1 sidecar validator (the Python half of
plugin/sidecar.h).
Every Python loader that reads a sidecar – eagle.loaded.LoadedKernel
and eagle.host_launch.HostPluginLibrary – funnels through
validate_sidecar() before it loads anything (before cupy.RawModule
or ctypes.CDLL), so a rejected artifact never reaches a driver load and
every reject case is exercisable with stub bytes and no GPU.
Absence is lenient here for pattern and scalar_type (a pre-freeze
sidecar never stamped either); per-loader absence rules (e.g.
LoadedVector defaulting an absent pattern to "vector") stay with
the loader. The neural_block clause is imported, not defined here
(raptor.schema.blocks.validate_neural_block_descriptor()): the spine
declares that wire contract, this module owns only the kernel-family checks
(schema_version, arg_spec roles, scalar_type, pattern,
derivative, buffer kind).
The optional ``terminated_readonly`` declaration. A producer may stamp
"terminated_readonly": true on a kernel sidecar to declare that its
kernel only ever reads the terminated mask, letting
eagle.marshal.coerce_terminated() serve an omitted mask from a shared
cache instead of allocating a fresh one. Opt-in only the producer can
assert; absent means “not declared”. Read by
read_terminated_readonly(), not by validate_sidecar() – it is a
Python-side launch hint the C++ registry has no use for.
- eagle.sidecar.TERMINATED_READONLY_KEY = 'terminated_readonly'#
The optional sidecar key a producer stamps to opt its kernel into the shared all-false
terminatedmask. See this module’s docstring.
- eagle.sidecar.read_terminated_readonly(meta, *, name)[source]#
Whether this sidecar declares its
terminatedmask read-only;Falsewhen the key is absent. Present-but-not-a-bool raisesValueError– a truthiness coercion would let a typo like"false"silently enable the shared buffer.- Return type:
bool- Parameters:
meta (dict)
name (str)
- eagle.sidecar.FINISH_KEY = 'finish'#
The optional sidecar key a producer stamps on a kernel that finishes its own samples. Read by
read_finish().
- eagle.sidecar.FINISH_COUNTER = 'finished_count'#
The one counter plane a finishing kernel counts into.
- eagle.sidecar.FINISH_AUTO = 'auto'#
The
finish.stepsvalue of a kernel whose steps per launch are a run-time word (hawk.steps(kernel, "auto")); such a kernel also stampsfinish.steps_max.
- eagle.sidecar.read_finish(meta, *, name)[source]#
The kernel’s
finishdeclaration, orNonewhen it never finishes:{"mask": "<plane>", "counter": "finished_count", "steps": K}(Kan int>= 1, or"auto"paired with an integersteps_max). A malformed declaration raisesValueErrornaming the artifact.- Return type:
dict|None- Parameters:
meta (dict)
name (str)
- eagle.sidecar.PARAMS_SCHEMA_KEY = 'params_schema'#
The sidecar key carrying the
paramsblock’s own wire version. Absent => 1.
- eagle.sidecar.PARAM_DTYPES = ('float', 'int')#
float(aRealargument) orint(the 8-byte signedInt p_<name>slot).- Type:
The uniform element types a v2
paramsentry may declare
- class eagle.sidecar.ParamSpec(name, dtype='float')[source]#
Bases:
objectOne declared broadcast (uniform) parameter: its
nameanddtype("float"|"int"). A v1 sidecar’s bare name yieldsParamSpec(name, "float").- Parameters:
name (str)
dtype (str)
-
name:
str#
-
dtype:
str= 'float'#
- eagle.sidecar.read_params(meta, *, name, required=True)[source]#
The sidecar’s declared broadcast params as a tuple of
ParamSpec. Accepts both wire shapes, keyed onPARAMS_SCHEMA: v1 (absent or1) is bare names (["mu", "k"], every uniform aReal); v2 is decl-carrying ([{"name": "mu", "dtype": "float"}, ...]).Raises loudly on a
params_schemanewer than this build, a shape disagreeing with the declared version, or an unknowndtype– an integer uniform silently bound through adoubleis wrong from 2^53 up, so every ambiguity here fails instead.requiredmirrors the two calling conventions in the tree: the device path readsmeta["params"], the host pathmeta.get("params", []).- Parameters:
meta (dict)
name (str)
required (bool)
- eagle.sidecar.validate_sidecar(meta, *, name)[source]#
Validate a parsed sidecar against the schema-v1 contract; raise on any breach. Mirrors the C++
eagle::plugin::validate_sidecarcheck for check, so an artifact accepted by one language is accepted by the other (pinned by the shared conformance corpus).namenames the artifact in every error. Returnsmetaunchanged, for chaining.Checks, in the C++ order:
schema_version(forward-strict, absent => v1);arg_specroles;scalar_type(absent/empty stays lenient);pattern(value-strict againstRECOGNIZED_PATTERNS, absent stays lenient); theneural_blockclause (pattern-conditional, called fromraptor.schema.blocks.validate_neural_block_descriptor(), before derivative/buffer since it forbids both);derivative’s shape; bufferkind(value-strict againstBUFFER_KINDS).Not checked here: the presence of
kernel/arg_spec– those are required keys a loader reads directly (KeyErrornaming the key), as the C++parse_sidecarenforces at parse rather than validate. Aneural_blockdescriptor is the exception, since no Python loader reads it: its clause requires both keys itself.- Return type:
dict- Parameters:
meta (dict)
name (str)
Canonical plugin arg-spec role vocabulary + schema version (the Python half).
The sidecar arg_spec is an ordered list of [role, name] pairs; role
is one of the fixed strings below. The C++ half is plugin/roles.h, and a
cross-check test (tests/test_roles_vocab.py) asserts the two lists are
identical. Every loader validates each arg_spec role against ROLES
at load time (forward-strict), so the Python Loaded* and the C++
PluginRegistry accept exactly the same vocabulary.
mat_in is a valid schema-v1 role (a pure kernel’s matrix input). A matrix
binds through the same 40-byte GRef mirror as a vector (a matrix GRef is a
width-R*C vector GRef; the R x C shape is in-kernel flat indexing only),
so every loader packs it as that same layout, validating the bound shape
against the sidecar’s mat_shapes / Mutable shape. Only the cupy
device path calls eagle.abi.make_gref; the C++ host PluginRegistry
and the ctypes HostPluginLibrary each build their own byte-compatible
GRefMirror-shaped struct directly. All three constructions produce the
identical 40-byte ABI layout; only the source language and POD type differ.
classify_arg() is the single role -> ABI-shape classifier both launch
paths (eagle.launch.assemble_args and eagle.host_launch.
HostPluginLibrary.run) dispatch on, so an unrecognised role always raises
rather than silently falling through in only one path. The two paths still
build different artifacts (cupy objects vs ctypes structs); only the
classification decision is shared.
- eagle.roles.ROLES = frozenset({'accum_out', 'lookup', 'mat_in', 'mutable', 'nsamples', 'out', 'per_sample', 'terminated', 'uniform', 'vec_in', 'wide_in', 'wide_out'})#
The 12 canonical arg-spec roles (schema v1). Keep in sync with
plugin/roles.h::kPluginArgRoles(set-equality, order-independent).accum_outnames the cross-sample accumulate plane; it resolves to the same ABI aswide_out(seeclassify_arg()/ARG_TAGS).
- eagle.roles.OUTPUT_ROLES = frozenset({'accum_out', 'mutable', 'out', 'wide_out'})#
allocated when the caller supplies none, and returned by
eagle.plan.Plan.run().- Type:
The roles of an output plane
- eagle.roles.INPUT_ROLES = frozenset({'lookup', 'mat_in', 'per_sample', 'terminated', 'vec_in', 'wide_in'})#
The roles of an input plane the caller supplies by name.
- eagle.roles.PER_SAMPLE_ROLES = frozenset({'mat_in', 'mutable', 'out', 'per_sample', 'terminated', 'vec_in'})#
The roles whose plane holds one element (or one column) per sample.
- eagle.roles.STATE_ROLES = frozenset({'accum_out', 'mutable', 'wide_out'})#
The roles of a plane a stepping model updates (
eagle.simulate’sstate).
- eagle.roles.SCHEMA_VERSION = 1#
The current + maximum plugin-schema version this loader understands. A sidecar/manifest tagged with a higher
schema_versionis rejected (forward-strict); an untagged artifact is treated as v1 (backward-lenient). Orthogonal to the"aether-abi/1"ABI tag (seeeagle.abi). Single-sourced withraptor.schema.manifest.SCHEMA_VERSION; the C++ half (plugin/roles.h::kPluginSchemaVersion) stays a source-level literal, cross-checked bytests/test_roles_vocab.py.
- eagle.roles.MAX_SCHEMA_VERSION = 2#
The highest
schema_versioncheck_schema_version()accepts (v1 and v2 both load today, v3+ does not) — a separate constant fromSCHEMA_VERSION, not a bump of it. Attribute-derived, so a “future schema version” anywhere in tests isMAX_SCHEMA_VERSION + 1, never a hardcoded literal.
- eagle.roles.RECOGNIZED_PATTERNS = frozenset({'neural_block', 'pure', 'vector'})#
The plugin families the shared sidecar validator (
eagle.sidecar.validate_sidecar()) can structurally validate.neural_blockis recognized but not launched by any loader. Keep in sync withplugin/roles.h::kRecognizedPatterns; sourced fromraptor.schema.blocks.ALL_PATTERNS.
- eagle.roles.LAUNCH_CERTIFIED_PATTERNS = frozenset({'pure', 'vector'})#
The plugin families a launching entry point will actually bind and run — a subset of
RECOGNIZED_PATTERNS.neural_blockis recognized but never launched; every launching door uses the shared message incheck_launch_certified_pattern(). The Python manifest door (eagle.registry.load_manifest()) dispatches on this set plusneural_blockitself (its sole descriptor-consuming branch). Keep in sync withplugin/roles.h::kLaunchCertifiedPatterns.
- eagle.roles.EXEC_REF_KINDS = frozenset({'kernel'})#
The
kinddiscriminant of an exec reference ({"kind": ..., "kernel": ...}) on aneural_blockdescriptor. One value in v1: the referenced artifact is a plain plugin kernel. A future aggregate/plan-bundle kind is a meaning change, so aschema_versionbump. Keep in sync withplugin/roles.h::kExecRefKinds; sourced fromraptor.schema.manifest.EXEC_REF_KINDS.
- eagle.roles.SCATTER_POLICIES = frozenset({'accumulate', 'unique_write'})#
The declared terminal-write contract of a block’s scatter.
scatter_policydeclares what the committed result MEANS, never the mechanism eagle uses to commit it.unique_write(v1’s sole value): every(target, slot)is written by exactly one source per step, so the commit is a plain store, deterministic with zero atomics (the bit-exact gate mode applies).accumulate: more than one source may write the same(target, slot)per step, and the result is the carried base plus an order-unspecified sum (band-gated, never bit-exact). Keep in sync withplugin/roles.h::kScatterPolicies; sourced fromraptor.schema.blocks.SCATTER_POLICIES.
- eagle.roles.NEURAL_REQUIRED_FIELDS = frozenset({'forward_exec', 'in_degree', 'input_width', 'out_degree', 'output_width', 'param_width', 'scatter_policy', 'state_width'})#
The fields a
neural_blockdescriptor must carry, beyond the general required set (schema_version,kernel,pattern,scalar_type, an emptyarg_spec). Keep in sync withplugin/roles.h::kNeuralRequiredFields; sourced fromraptor.schema.blocks.NEURAL_REQUIRED_FIELDS.
- eagle.roles.NEURAL_EXEC_REF_FIELDS = frozenset({'forward_exec', 'jvp_exec', 'vjp_exec'})#
The keys whose values are exec references.
forward_execis required; the two derivative refs are optional-additive. Keep in sync withplugin/roles.h::kNeuralExecRefFields; sourced fromraptor.schema.blocks.NEURAL_EXEC_REF_FIELDS.
- eagle.roles.NEURAL_FORBIDDEN_FIELDS = frozenset({'aether_abi', 'buffers', 'derivative', 'host_entry', 'mat_shapes', 'mutables'})#
Kernel-machinery fields a descriptor must not carry (each would otherwise be silently ignored rather than fail loudly).
aether_abi,derivative,buffers,mutables,mat_shapes,host_entry— a descriptor binds nothing. Keep in sync withplugin/roles.h::kNeuralForbiddenFields; sourced fromraptor.schema.blocks.NEURAL_FORBIDDEN_FIELDS.
- eagle.roles.BUFFER_KINDS = frozenset({'lookup'})#
The declared-buffer
kindvocabulary (schema v1) — the only kinds a loader can bind; an unrecognized kind is rejected up front rather than silently dropped. Sourced fromraptor.schema.blocks.BUFFER_KINDS.
- eagle.roles.MANIFEST_FORMATS = frozenset({'cubin', 'fatbin', 'ptx'})#
The manifest-entry
formatvocabulary (schema v1) — the artifact container a manifest entry names; an unrecognized value is refused rather than handed to a loader. Keep in sync withplugin/roles.h::kManifestFormats; sourced fromraptor.schema.manifest.MANIFEST_FORMATS.
- eagle.roles.validate_roles(arg_spec, *, name)[source]#
Reject an
arg_speccarrying a role outsideROLES(forward-strict).arg_specis the sidecar’s list of(role, name)pairs;namenames the artifact in the error. Mirrorsis_valid_roleon the C++ side.- Return type:
None- Parameters:
name (str)
- eagle.roles.ARG_TAGS = frozenset({'ACCUM_OUT', 'GREF_MAT', 'GREF_VEC', 'HANDLE', 'NSAMPLES', 'UNIFORM', 'WIDE_IN', 'WIDE_OUT'})#
The ABI-shape tags a role resolves to — the classification both launch paths dispatch on, single-sourced so
eagle.launch.assemble_argsandeagle.host_launch.HostPluginLibrary.runcan never diverge.GREF_VEC/GREF_MATstay distinct (rather than one sharedGREF) because the ctypes host path reads a “mutable” role’s bound value from one of two different caller-populated dicts (self._vecvsself._mat) depending on shape; the cupy device path’s ownmake_grefcall is identical either way.WIDE_IN/WIDE_OUTandACCUM_OUTare each kept distinct fromHANDLE/WIDE_OUTfor the same reason: each reads from its own caller-populated dict (wide_in/wide_out/accum_out), even though construction is byte-identical to a plain scalar-handle pointer.
- eagle.roles.classify_arg(role, name, *, vec_mutables=frozenset({}), mat_mutables=frozenset({}))[source]#
The ABI-shape tag
(role, name)resolves to — one ofARG_TAGS.Extracts the classification only: both launch paths still build their own artifacts from whichever tag comes back (a cupy
GRefnumpy-structured scalar vs a ctypesGRefMirror/ScalarHandlePOD) — that construction stays per-path."out"/"vec_in"->GREF_VEC;"mat_in"->GREF_MAT."mutable"resolves by context:GREF_MATifnameis inmat_mutables,GREF_VECif invec_mutables, elseHANDLE(a scalar/int Mutable)."per_sample"/"lookup"/"terminated"->HANDLE."wide_in"->WIDE_IN;"wide_out"->WIDE_OUT."accum_out"->ACCUM_OUT(the cross-sample accumulate plane; same construction asWIDE_OUT)."nsamples"->NSAMPLES;"uniform"->UNIFORM.Any other
role->ValueError(fail-loud; never a silent skip).
- Return type:
str- Parameters:
role (str)
name (str)
vec_mutables (frozenset)
mat_mutables (frozenset)
- eagle.roles.check_launch_certified_pattern(pattern, *, subject)[source]#
Refuse
patternat a launching door, with the two-branch message this module locks (mirrors the C++check_launch_certified_patterninplugin/roles.hmessage for message).patternoutsideRECOGNIZED_PATTERNS— this build has never heard of the family: the message names the supported set and saysupgrade eagle.patternrecognized but not certified (e.g. aneural_blockdescriptor, not runnable) — the message saysrecognized but not launchable by this loader, with no upgrade suffix.
subjectnames the artifact, followed directly by" pattern '<value>'". An absent/emptypatternis lenient (a pre-freeze artifact never stamped one).- Return type:
None- Parameters:
subject (str)
- eagle.roles.check_schema_version(meta, *, name='<document>', allow_legacy_version_key=False)[source]#
Return the artifact’s plugin-schema version, rejecting one we cannot load.
schema_versionabsent => v1 (backward-lenient). A version greater thanMAX_SCHEMA_VERSIONis rejected (forward-strict); v1 and v2 both load today. Mirrors the C++kPluginSchemaVersiongate (still v1-only pending the twin-site bump).allow_legacy_version_keyscopes the legacyversion-key fallback to manifests only (passTruefromeagle.registry.load_manifest()); the default (False) matches C++ sidecar behaviour, where an unrecognizedversionkey is ignored.Re-exported from raptor, which owns the version-compare wording.
- Return type:
int- Parameters:
meta (dict)
name (str)
allow_legacy_version_key (bool)
- eagle.roles.check_execution_axis(meta, version, *, name='<document>')[source]#
Validate the schema-v2 execution axis, given the already-resolved
version(check_schema_version()’s return). v1 documents must carry none ofexec_targets/exec_access/exec_op; v2 documents must carryexec_targets+exec_access(absence = load refused) and validate their vocabulary +exec_op’s mapreduce-conditional requiredness.Re-exported from raptor, which owns the execution-axis shape. Wired into
eagle.registry.load_manifest().- Return type:
None- Parameters:
meta (dict)
version (int)
name (str)
- eagle.roles.PARAMS_SCHEMA = 2#
The sidecar
paramsblock’s own wire version — the current + maximum shape this build can read.1 (what an absent
params_schemakey means):"params": ["mu", "k"]— a bare name list, every uniform implicitly aReal.2:
"params": [{"name": "mu", "dtype": "float"}, ...]— decl-carrying, the same{name, dtype}shapemutableshas always had, so anintuniform binds through the int binder instead of arriving widened through a double.
Field-local rather than a
schema_versionbump: the C++ side never readsparamsat all (it resolves uniforms byarg_specrole), so no C++ reader can misread this block’s new shape. Python-only by construction, except the code generator’s conformance-gated verbatim copy, which must be bumped together with this one.
- eagle.roles.param_decls(params)[source]#
Normalize a broadcast-param list to
((name, dtype), …).One spelling for “what type is this uniform?”, shared by every eagle consumer of a params list. Accepts:
a bare
str— a v1 sidecar’s name; defaults tofloat.a
(name, dtype)pair.anything with
.name/.dtype(eagle.sidecar.ParamSpecor a code generator’sParamDecl) — so a generatedTraceResult.paramscan be handed straight to a launch.a
{"name", "dtype"}dict — a v2 sidecar entry read raw.
An unrecognised entry raises: a uniform whose type cannot be established is the exact silent-wrong-answer this normalization exists to prevent.
- eagle.roles.param_names(params)[source]#
Just the NAMES of a broadcast-param list (any of the shapes
param_decls()accepts) — the binding-set / kwarg-key view.- Return type:
tuple[str,...]
- eagle.roles.parse_derivative(meta, *, name)[source]#
Return the sidecar’s optional
derivativeblock (orNone).A VJP/JVP derivative artifact carries an additive
derivativeblock describing its role; an ordinary primal/kernel has none. Backward-lenient: an absent block returnsNone. Validates the block’s shape at load (mirroring the C++validate_sidecar):kindmust bevjp/jvp, and a populatedresidualslist is rejected (recompute-only for now). The returned dict is the loaded-kernel metadata surface (LoadedKernel.derivative).- Parameters:
meta (dict)
name (str)
eagle.cuda — the CUDA backend surface (torch.cuda-familiar).
Mirrors the C++ eagle::cuda namespace 1:1 in Python: the nanobind-bound
CUDA-graph machinery from the compiled eagle._core extension — the
capturable graph (Graph), the stream-capture recorder
(StreamCapturer) and its owned result (CapturedGraph), the
CUDA stream wrapper (Stream, the cudaStream_t interop backbone), and
the instantiated/replayable executable handle (Launcher).
Importing this submodule loads the compiled _core extension, so it is kept
out of the top-level eagle import (which must stay usable in a
pure-Python / no-GPU install). Reach these as eagle.cuda.Graph etc.
It also carries the capture-introspection trio a graph recorder needs: the
forked-stream scope (CaptureFork), the node snapshot of a capture
(capture_snapshot_nodes()) and whether a captured node can be toggled
(is_node_toggleable()).
The CPU backend (eagle::cpu — the host graph executor + OMP Scan/Reduction)
is currently C++-only.
- class eagle.cuda.CaptureFork(*args, **kwargs)#
Bases:
object- branch#
Raw cudaStream_t of branch index as a Python int (wrap with cupy.cuda.ExternalStream). Only meaningful between fork() and join().
- fork#
every branch becomes a sibling of every other, and everything captured so far becomes a predecessor of all branches. Calling twice is a no-op.
- Type:
Open the fork
- forked#
True once fork() has run and join() has not.
- join#
the origin waits for every branch, so work issued after the join depends on all of them. Idempotent. An unjoined branch makes StreamCapturer.end() fail and discard the graph.
- Type:
Close the fork
- origin#
Raw cudaStream_t of the origin stream this fork branches from.
- size#
Number of branches.
- class eagle.cuda.CapturedGraph#
Bases:
object- debug_dot#
Write this captured graph to Graphviz dot at path and return the text (matches cupy Graph.debug_dot_str).
- class eagle.cuda.Graph(*args, **kwargs)#
Bases:
object- add_node#
Fold a CapturedGraph in as a child-graph node (consumes it) and harvest its kernel records.
- from_captured = <nanobind.nb_func object>#
- last_node#
- launcher#
Instantiate an exec graph and return a Launcher. Raises RuntimeError if cudaGetLastError() is nonzero after instantiate.
- stream#
- class eagle.cuda.Launcher#
Bases:
object- kernel_node_count#
Number of harvested kernel-node records (cross-check for num_nodes()).
- launch#
Replay the instantiated exec graph once. Raises RuntimeError if cudaGetLastError() is nonzero after the replay; see eagle/python/eagle/pipeline.py’s launch() for the companion stream-ordering fix).
- set_logical_size#
Patch every kernel node’s grid/block for a new logical size.
- set_node_enabled#
Enable/disable one node (by raw handle, as returned by capture_snapshot_nodes()) in this Launcher’s instantiated exec graph. Legal between replays; no recapture, no structure change (mode=”enabled”). Raises RuntimeError (name+code) if the node’s type is not one CUDA supports toggling – check is_node_toggleable() ahead of time.
- stream#
- synchronize#
The low-level launch stack#
eagle.launch(), eagle.LoadedKernel and friends
(eagle.launch, eagle.loaded, eagle.marshal) and the ctypes
host launcher eagle.host_launch are the LOW-LEVEL launch stack: one
compiled aether-abi/1 kernel launched by hand, with its arguments marshalled
per call. New code reaches kernels through eagle.plan.plan(),
eagle.deploy(), eagle.until_done() and eagle.simulate(); the
low-level stack stays for neural-block manifests and for callers that need
exactly one launch and nothing else.
- class eagle.launch.LaunchMixin[source]#
Bases:
objectThe framework-polymorphic launch skeleton, shared by every launcher: owns the
origin -> with origin.launch_context(): ... -> if origin.blocking: deviceSynchronize()wrapper and the unknown-keyword guard – the single source of “one call mirrors the input framework”.
- class eagle.launch.LaunchPlan(arg_spec, vec_mutables=(), mat_mutables=())[source]#
Bases:
objectThe kernel-STATIC half of a launch, resolved ONCE per signature: the ABI shape each
(role, name)binds as, whichMutables are vector- vs matrix-shaped, and a reusable by-value POD box per argument – all thingsassemble_args()used to re-derive and re-allocate on every call. Block-size derivation is deliberately not here: it is a per-call function of the batch size, not the kernel.The boxes are reused, and that is the point.
_boxes[i]is refilled in place on every planned launch, which is safe since a launch packs its by-value arguments synchronously – no launch reads a box afterfn(...)returns. Do NOT hold anassemble_argsresult across a second planned launch of the same signature: that list aliases the plan’s boxes. Omitplanto get fresh boxes per call.Single-threaded launch. Reused per-signature boxes mean two threads launching the same signature concurrently would interleave their refills – not a regression, since the launch path is host-side sequencing behind the GIL and no launch pool in this codebase is threaded.
- arg_spec#
- vec_mutables#
- mat_mutables#
- tags#
- boxes#
- eagle.launch.launch_plan(arg_spec, vec_mutables=(), mat_mutables=())[source]#
The cached
LaunchPlanfor one kernel signature (identity, not equality); a changed signature gets its own plan. Nothing per-call enters the key.- Return type:
- eagle.launch.pure_origin(kw, *, vector_inputs, mutable_names, per_sample)[source]#
The caller’s framework, chosen from every per-sample source in signature order (vector inputs, provided
Mutablearrays, then per-sample scalars) else numpy. Detected up front so the coercion rides that framework’s stream too.
- eagle.launch.DEFAULT_BLOCK = 256#
default launch block size
CPU-plugin launcher – the host twin of eagle.registry, and the
Python peer of the C++ eagle::cpu::PluginRegistry (plugin/host_registry.h).
Where eagle.registry launches CUDA kernels over cupy device arrays,
this module ctypes-loads a host plugin .so and calls its
<kernel>_host entry over host buffers – numpy arrays or torch CPU
tensors, zero-copy via their raw data pointer (eagle.interop.host_ptr()).
It shares the SAME binary ABI as the device path: the entry is
void <kernel>_host(void* const* params, int32_t n), each params[i]
pointing to the same GRefMirror / ScalarHandle POD as
plugin/gref_abi.h, packed in arg_spec order – the array the C++
registry would hand cuLaunchKernel, here handed to a host function
running its own OpenMP loop. A CPU plugin is the same artifact contract as
a GPU plugin, minus the PTX.
- class eagle.host_launch.GRefMirror[source]#
Bases:
Structurectypes mirror of
plugin/gref_abi.hGRefMirror(a 40-byte width-independent POD view).- compStride_#
Structure/Union member
- data_#
Structure/Union member
- deviceId_#
Structure/Union member
- deviceType_#
Structure/Union member
- sampleStride_#
Structure/Union member
- samples_#
Structure/Union member
- class eagle.host_launch.ScalarHandle[source]#
Bases:
Structurectypes mirror of
plugin/gref_abi.hScalarHandle(a 32-byte POD, GRefMirror’s rank-1 sibling; carries both an extent and a device tag, unlike the earlier bare-pointerHandleT).- data#
Structure/Union member
- deviceId#
Structure/Union member
- deviceType#
Structure/Union member
- samples#
Structure/Union member
- stride#
Structure/Union member
- class eagle.host_launch.HostPluginLibrary(so_path, sidecar)[source]#
Bases:
objectA dlopen’d CPU plugin, driven from Python over host buffers.
Mirrors
eagle::cpu::PluginRegistry: bind the kernel’s by-name buffers from host pointers the caller already owns (seeeagle.interop.host_ptr()), thenrun()packs the args inarg_specorder and calls the host entry.so_pathis the plugin shared object;sidecaris the same dict the C++ registry and code generator speak:kernel,aether_abi(required), optionalhost_entry/scalar_type(must be absent/"float64"),arg_spec([role, name]pairs), and optionalmutables/mat_shapes.- Parameters:
sidecar (dict)
- bind_vector(name, ptr, n)[source]#
Bind an
out/vec_in/ vector-mutableSoA buffer by name.- Return type:
- Parameters:
name (str)
ptr (int)
n (int)
- bind_matrix(name, ptr, n, rows, cols)[source]#
Bind a
mat_in/ matrix-mutableflat(R*C, N)buffer.rows/colsare checked against the sidecar’s declared shape;ptraddressesdim = r*C + c, sample-fastest.- Return type:
- Parameters:
name (str)
ptr (int)
n (int)
rows (int)
cols (int)
- bind_handle(name, ptr)[source]#
Bind a
per_sample/terminated/ scalar-mutable/ wide flat buffer by name (all ride the same scalar-handle pointer). Alookuptable usesconsolidate()instead.- Return type:
- Parameters:
name (str)
ptr (int)
- consolidate(name, ptr, count)[source]#
Consolidate a read-only
lookuptable – the host twin of the device registry’sconsolidate(nothing to upload; just records pointer + count).countis checked against the sidecar’s declared size; a declared table must be consolidated beforerun().- Return type:
- Parameters:
name (str)
ptr (int)
count (int)
- bind_uniform(name, value)[source]#
Bind a
uniformscalar by name (the float64 spelling). The change test is bit-exact where==is not:-0.0 == 0.0but the two are not interchangeable in a kernel, so a rebind between them must invalidate the cache;NaN != NaNalready does.- Return type:
- Parameters:
name (str)
value (float)
- bind_uniform_int(name, value)[source]#
Bind a
uniformscalar by name as an exact 64-bit signed integer (ctypes.c_longlong, neverc_double: that widening is bit-exact only below 2^53). A non-integral value is refused, not truncated (2.0is accepted as an integer spelling).- Return type:
- Parameters:
name (str)
value (int)
- run(n)[source]#
Pack the args in
arg_specorder and call the host entry over @p n samples, the same role -> params[] packing the C++ registry does. Returns the sample count run.The pack is CACHED: a hit replays the exact array a prior miss built from the same bindings, skipping the validation the miss path performs (it cannot newly fail, since nothing changed since it passed).
- Return type:
int- Parameters:
n (int)
Framework bridges#
eagle.frameworks.torch turns a hawk per-sample kernel into a
torch.autograd.Function (see Train through a physics kernel with torch).
It is the one eagle module that imports torch, and only when it is imported
itself.
A hawk per-sample kernel as a torch.autograd.Function: function().
Forward runs the kernel; backward runs the reverse-mode kernel hawk
derives from it (hawk.diff.vjp()); forward-mode AD runs the derived
tangent kernel (hawk.diff.jvp()) – nothing is taped, each pass is
one kernel launch. torch.func transforms are not supported: they hand
derivative rules wrapped tensors exposing no storage, and a kernel can
only read a buffer.
CPU tensors run through hawk’s host runtime; CUDA tensors through eagle’s
launch path (eagle.plan.plan()) on torch’s current stream, no sync.
Every plane crosses through eagle.interop.import_buffer()
zero-copy; a non-contiguous input is made contiguous, a wrong-dtype one
refused.
Conventions, off the kernel’s own declaration: inputs are its planes and
Param uniforms in order; outputs are its Mutable planes (bare if
one), PLUS the Terminated mask itself when the kernel’s own body
finishes it (terminated = cond) – the updated mask comes back as an
extra return value, the last one, so a stepping loop reuses the SAME stop
decision the kernel made instead of recomputing it with a second masking
rule of its own. A per-sample plane is (n,)/(width, n); a
sample-major (n, width) input is accepted too (zero-copy transpose where
contiguous), and outputs come back in that same layout – mixing layouts
in one call is refused. A Param’s gradient is the batch sum of its
per-sample partials (held in CUDA, read back to host once per launch). A
Terminated mask is never differentiated: a marked sample’s outputs and
gradient/tangent stay zero, and the mask itself carries no gradient/tangent
of its own.
torch is imported by this submodule only: import eagle stays torch-free.
|
Wrap |
|
A hawk kernel, its derived reverse- and forward-mode kernels, and the |
GEMM helpers#
A capture-legal single-precision GEMM step: raw cuBLAS, bound by ctypes.
Pure ctypes plus a strides-only shape mapping, with zero
code-generator imports – this is eagle runtime, not codegen.
Cross-module references point at eagle.gemm.plan.
eagle.gemm.plan’s matmul_step calls the array module’s
matmul, which cupy refuses to record into a CUDA graph during
stream capture. This module is the direct cuBLAS call that bypasses
that restriction: it never selects a mechanism, never allocates, and is
not wired into matmul_step, which remains the uncaptured
implementation.
Three pieces: load_cublas() resolves the shared library from the
process’s existing mappings; CublasCaptureHandle is one handle
per captured-graph build; sgemm_params() is the row-major to
column-major mapping, a pure function over shapes and strides.
Refusals are typed and loud (GemmCaptureUnavailable and its
two subclasses), never fallbacks.
- eagle.gemm.capture.CUBLAS_OP = {'N': 0, 'T': 1}#
cuBLAS’ own transpose enum, by the letters this module’s mapping speaks.
- class eagle.gemm.capture.CublasCaptureHandle(stream_ptr)[source]#
Bases:
objectOne cuBLAS handle, owned by ONE captured-graph build. The lifetime rules are why this is an object, not a function:
cublasSetStreamis called once, in the constructor, outside the capture region (cupy itself sets the stream per call, which is why it blanket-refuses cuBLAS during capture); the handle outlives its graph (cuBLAS keeps a per-handle workspace, so destroying it early is a use-after-free at replay);warm()runs on the capture stream, aftersetStreamand beforebegin_capture, once per GEMM shape (its workspace allocates lazily, so warming elsewhere leaves acudaMallocinside the capture region); onlysgemm()runs inside the capture region; teardown order is graph, then handle, then pool.alpha/betaare host-side scalars baked into the recorded graph node at enqueue time: changing them means recording again.- warm(params, operands)[source]#
Run this GEMM shape once, outside the capture region. Idempotent per shape, so a caller may warm defensively before every recorded step without cost. Writes the destination exactly as the captured call will, so contents that matter must be re-filled after.
- Return type:
None- Parameters:
params (SgemmParams)
operands (dict)
- sgemm(params, operands)[source]#
Enqueue ONE Sgemm on the handle’s stream – the only call legal inside a capture region.
operandsmaps"a"/"b"/"out"to the caller’s device buffers; which ofa/bis the first cuBLAS operand isparams.operands, since the row-major to column-major mapping swaps them for every case excepttranspose_out.- Return type:
None- Parameters:
params (SgemmParams)
operands (dict)
- exception eagle.gemm.capture.CublasResolutionError[source]#
Bases:
GemmCaptureUnavailableThe cuBLAS library could not be resolved unambiguously. Zero matches means cupy’s own cuBLAS never initialized; more than one means two different
libcublasfiles are mapped, and picking either would risk a wrong answer or a crash far from here.
Bases:
ExceptionBase for “this GEMM cannot be run as a captured raw-cuBLAS step”.
Bases:
GemmCaptureUnavailableThis contraction cannot be expressed as one Sgemm call: a shape, dtype, stride pattern or operand class the mapping does not cover. Permanent for the operands as bound – leave the step uncaptured, not retry.
- class eagle.gemm.capture.SgemmParams(op_a, op_b, m, n, k, lda, ldb, ldc, operands)[source]#
Bases:
NamedTupleEverything one
cublasSgemm_v2call needs, bar the pointers.op_a/ldaandop_b/ldbare named for the cuBLAS argument position, not the caller’s arrays:operandssays which caller buffer fills each position, swapped for every case excepttranspose_out.m/n/kare the column-major extents cuBLAS is told, not necessarily the row-major result’s own shape.- Parameters:
op_a (str)
op_b (str)
m (int)
n (int)
k (int)
lda (int)
ldb (int)
ldc (int)
operands (tuple)
-
op_a:
str# Alias for field number 0
-
op_b:
str# Alias for field number 1
-
m:
int# Alias for field number 2
-
n:
int# Alias for field number 3
-
k:
int# Alias for field number 4
-
lda:
int# Alias for field number 5
-
ldb:
int# Alias for field number 6
-
ldc:
int# Alias for field number 7
-
operands:
tuple# Alias for field number 8
- eagle.gemm.capture.check_plan_capturable(plan, *, operand_class)[source]#
Refuse a lowered plan the captured raw-cuBLAS route cannot run. Two permanent fences: the
"gemv"operand class has no mapping here (its lowering contracts against a ones vector), and every declared buffer must be float32.- Return type:
None- Parameters:
operand_class (str)
- eagle.gemm.capture.load_cublas()[source]#
The
ctypes.CDLLforresolve_cublas_path()’s library, cached per process (the handle-per-graph rule is about the cuBLAS handle, not this library object).
- eagle.gemm.capture.mapped_library_paths(maps_text, soname=('libcublas.so.12', 'libcublas.so.13'))[source]#
Every distinct file path in
maps_textwhose basename issoname(one name, or any of a tuple of names) itself or extended by a real-file suffix – never a bare prefix match (conda’s real file is further-versioned, e.g.libcublas.so.12.9.1.4, andlibcublasLt.so.12must not match). Pure function over a/proc/<pid>/mapsdump: the refusals are provable without a GPU. Sorted and deduplicated.- Return type:
tuple- Parameters:
maps_text (str)
- eagle.gemm.capture.require_float32(specs)[source]#
Refuse unless every
WorkspaceSpecinspecsdeclares float32. Stated as a fence:lower_dense_gemm’s own dtype default is still"float64", so a plan built without an explicit dtype would otherwise reach here declaring a type this step cannot run.- Return type:
None
- eagle.gemm.capture.sgemm_params(a_shape, a_strides, b_shape, b_strides, out_shape, out_strides, itemsize, *, transpose_a, transpose_b, transpose_out)[source]#
Map
matmul_step’s row-major contraction onto one column-major Sgemm. The contraction isout_view = op(a) @ op(b); a row-major(r, c)buffer with row strideld, read column-major with leading dimensionld, is that buffer’s transpose. So an untransposed destination asks cuBLAS forout.T = op(b).T @ op(a).T(operands swapped, flags inverted);transpose_outinverts that inversion back to natural order with opposite flags. Leading dimensions come from the strides (_leading_dimension()); refuses anything one Sgemm cannot address, or any itemsize that is not float32’s.- Return type:
- eagle.gemm.capture.sgemm_params_for_arrays(a, b, out, *, transpose_a=False, transpose_b=False, transpose_out=False)[source]#
sgemm_params()read off three live arrays (numpy or cupy).Also enforces the dtype fence on all three, since here it can be seen.
.stridesis in BYTES for both array modules, which is what the pure function takes.- Return type:
The record shape a recognized contraction hands to eagle.gemm.plan
— not the recognizer itself.
recognize lives in the code generator (a structural walk over its trace
IR); eagle never imports a producer, so this module owns only the shape of
the result: ContractionOperand and RecognizedContraction,
pure frozen dataclasses with no imports beyond dataclasses. A caller
that already has a recognized-contraction result can pass it straight into
eagle.gemm.plan.lower_dense_gemm(); a future producer with its own
recognizer can construct a RecognizedContraction here directly —
neither path depends on the code generator.
- class eagle.gemm.contraction.ContractionOperand(name, role, dims, bound, reads, reduce_stride, reduce_dim, stride_labels, dense, refusals)[source]#
Bases:
objectOne buffer a recognized contraction reads, described from its declaration.
dimsis the declared axis layout((label, size, stride), …)in declared order,()for a positional table with no declared axes.reduce_strideis how far one step of the reduce axis moves this operand’s flat index:0if invariant along it,Noneif the index is not affine in the loop variable at all.reduce_dimnames the declared axis that stride belongs to, when it is exactly one declaration’s named constant.denseis this operand’s half of the contraction’s density verdict;refusalsis why not, one sentence per reason.- Parameters:
name (str)
role (str)
dims (tuple[tuple[str, int, int], ...])
bound (int)
reads (int)
reduce_stride (int | None)
reduce_dim (str | None)
stride_labels (tuple[str, ...])
dense (bool)
refusals (tuple[str, ...])
-
name:
str#
-
role:
str#
-
dims:
tuple[tuple[str,int,int],...]#
-
bound:
int#
-
reads:
int#
-
reduce_stride:
int|None#
-
reduce_dim:
str|None#
-
stride_labels:
tuple[str,...]#
-
dense:
bool#
-
refusals:
tuple[str,...]#
- class eagle.gemm.contraction.RecognizedContraction(carry, reduce_var, reduce_extent, reduce_start, reduce_step, operands, operand_class, term_form, outputs, dense, refusals, staged_fanin, fanin_source, adjoint_accumulates)[source]#
Bases:
objectA reduce loop described as a contraction — the whole seam, frozen.
operand_classis"gemv"for a single-operand row reduction or"gemm"for two operands multiplied along the reduce axis.denseis the AND of every operand’s verdict;refusalscollects their reasons, each prefixed with the operand it came from.term_formis the value property besidedense’s index property:"bare_read","product_of_reads"or"transformed". Both are needed to lower: density says the operands can be walked as matrices, term form says a matmul over those matrices computes what the loop computes.staged_faninis how many lanes read one staged slot — the number of cotangent contributions the reverse pass accumulates into it — carried as declared, withNonewhere the trace derives none.fanin_sourcerecords where it was read from.adjoint_accumulatesis((name, atomic?), …)for the shared accumulate outputs the derived kernel contributes to.- Parameters:
carry (str)
reduce_var (str)
reduce_extent (int | None)
reduce_start (int | None)
reduce_step (int)
operands (tuple[ContractionOperand, ...])
operand_class (str)
term_form (str)
outputs (tuple[tuple[str, str], ...])
dense (bool)
refusals (tuple[str, ...])
staged_fanin (tuple[tuple[str, int | None], ...])
fanin_source (str)
adjoint_accumulates (tuple[tuple[str, bool], ...])
-
carry:
str#
-
reduce_var:
str#
-
reduce_extent:
int|None#
-
reduce_start:
int|None#
-
reduce_step:
int#
-
operands:
tuple[ContractionOperand,...]#
-
operand_class:
str#
-
term_form:
str#
-
outputs:
tuple[tuple[str,str],...]#
-
dense:
bool#
-
refusals:
tuple[str,...]#
-
staged_fanin:
tuple[tuple[str,int|None],...]#
-
fanin_source:
str#
-
adjoint_accumulates:
tuple[tuple[str,bool],...]#
- property operand_count: int#
How many distinct buffers the addend reads (1 or 2 — see
operand_class).
- property operand_names: tuple[str, ...]#
The read operands’ declared names, in first-read order.
The lowered plan — an ordered, backend-parametric sequence of GEMM steps.
eagle.gemm.contraction says what a reduce loop’s record looks like
once recognized; this module is what a lowering actually is, given one. A
LoweredPlan is a frozen tuple of named steps plus the metadata a
caller needs to run them: the mechanism row it came from, its determinism
typing, and the specs of every buffer it touches.
Three properties, each a deliberate exclusion:
backend-parametric. Every step carries both a device form and a host twin, so a plan is checked against a numpy reference without a GPU;
it allocates nothing. Every buffer is the caller’s, bound by name; the plan publishes
WorkspaceSpecs and validates what it is handed against them. In v1 the allocator is the runner;it does not schedule. Steps run in list order, enqueued on the current stream and left there; the caller keeps control of ordering.
What a plan is NOT: a place where a contraction is decided. The steps
arrive already chosen. lower_dense_gemm() derives one from the
record alone, consuming it entirely by duck-typed attribute access, so
a record built by the code generator’s recognize(), or hand-built
to the same shape, drives this module identically.
- class eagle.gemm.plan.CapturedGemmCall(name, params, operands)[source]#
Bases:
objectOne plan step, described as the raw
Sgemma captured graph records.operandsmaps"a"/"b"/"out"to the caller’s own buffers – never copies, since a captured graph records device pointers. Which of"a"/"b"is the first cuBLAS operand isparams.operands, not this mapping.- Parameters:
name (str)
params (object)
operands (dict)
-
name:
str#
-
params:
object#
-
operands:
dict#
- class eagle.gemm.plan.LoweredPlan(steps, mechanism, determinism, inputs=(), outputs=(), workspaces=())[source]#
Bases:
objectAn ordered sequence of steps plus everything needed to run them.
inputs/outputs/workspacesare allWorkspaceSpecs, uniform since the caller allocates and the plan validates all three the same way; the split is about role: inputs arrive filled, outputs leave filled, workspaces are intermediates a later step reads. Construction validates the read/write graph: a step may only read a name that is an input or was written by an earlier step, may only write a declared name, every output must be written, and every workspace must be both written and later read (one nothing reads would be over-specified).- Parameters:
steps (tuple[PlanStep, ...])
mechanism (str)
determinism (str)
inputs (tuple[WorkspaceSpec, ...])
outputs (tuple[WorkspaceSpec, ...])
workspaces (tuple[WorkspaceSpec, ...])
-
mechanism:
str#
-
determinism:
str#
-
inputs:
tuple[WorkspaceSpec,...] = ()#
-
outputs:
tuple[WorkspaceSpec,...] = ()#
-
workspaces:
tuple[WorkspaceSpec,...] = ()#
- property required: tuple[WorkspaceSpec, ...]#
Every buffer the caller must allocate and bind, in declaration order.
- property kinds: tuple[str, ...]#
- __call__(backend, buffers)[source]#
Run every step in order and return
{output name: buffer}.backendis"host"(numpy, sequential) or"device"(cupy, enqueued on the current stream and left there – the caller’s stream ordering is the only ordering).buffersmust already hold every name inrequired, matching its spec: nothing is allocated here, and a missing or mis-shaped binding is refused before any step runs.- Return type:
dict- Parameters:
backend (str)
buffers (dict)
Bases:
ExceptionBase for “this lowering cannot be used” — raised at BUILD time (
PlanUnavailable) or at CALL time (MatmulUnavailable).
- class eagle.gemm.plan.MatmulSpec(a, b, out, transpose_a=False, transpose_b=False, transpose_out=False)[source]#
Bases:
objectWhat a matmul step multiplies, by name, and which way round. A step’s two closures already know this, which was enough while the only consumer was
LoweredPlan.__call__(). Recording a plan into a CUDA graph needs more: cupy refuses every cuBLAS call during stream capture, so the captured form is a rawSgemmissued througheagle.gemm.capture, which needs the operand names, the three transpose flags and the destination – the same descriptionmatmul_step()was built from. Published rather than re-derived from the record, since a plan is what actually runs: a second derivation would be a second thing to keep in step with the first.- Parameters:
a (str)
b (str)
out (str)
transpose_a (bool)
transpose_b (bool)
transpose_out (bool)
-
a:
str#
-
b:
str#
-
out:
str#
-
transpose_a:
bool= False#
-
transpose_b:
bool= False#
-
transpose_out:
bool= False#
Bases:
LoweringUnavailableA plan’s matmul FAILED on the device at call time: the failure a mechanism-selection layer’s
cublas_available()check cannot predict (a cupy that imports cleanly but whose cuBLAS library is missing, mismatched, or cannot create a handle). The contract is that the caller maps it to an uncaptured fallback plus a loud note (matmul_fallback_note()), never a silent retry or a swallowed genuine shape/dtype bug.
- class eagle.gemm.plan.PlanStep(name, kind, reads, writes, device_fn, host_fn, matmul=None)[source]#
Bases:
objectOne step: what it is called, what kind it is, which buffers it reads and writes, and its two forms.
device_fn/host_fntake the bound buffer mapping and return nothing. Two callables rather than one taking an array module, because a generated kernel’s two forms are not the same function with a differentxp(a producer’s own dispatch picks the provider), while a matmul’s are.- Parameters:
name (str)
kind (str)
reads (tuple[str, ...])
writes (tuple[str, ...])
device_fn (object)
host_fn (object)
matmul (object)
-
name:
str#
-
kind:
str#
-
reads:
tuple[str,...]#
-
writes:
tuple[str,...]#
-
device_fn:
object#
-
host_fn:
object#
-
matmul:
object= None# The step’s own description when it is a matmul,
Noneotherwise. Excluded from comparison/hashing, like the two callables beside it: plan equality is used to check a hand-built plan against a derived one.
Bases:
LoweringUnavailableA plan cannot be built for this contraction: a permanent property of the record, not a retry target.
- class eagle.gemm.plan.WorkspaceSpec(name, shape, dtype)[source]#
Bases:
objectOne buffer a plan touches: its name, shape and dtype. The plan never allocates it – this is the declaration the caller allocates against.
dtypeis a normalized name string, not a numpy dtype object, so the spec stays hashable and comparable.- Parameters:
name (str)
shape (tuple[int, ...])
dtype (str)
-
name:
str#
-
shape:
tuple[int,...]#
-
dtype:
str#
- check(array)[source]#
Raise unless
arraymatches this spec (shape and dtype): a plan that wrote throughout=into a wrongly-typed buffer would either raise deep inside the array module or silently downcast. The dtype-name lookup goes through_dtype_name()’s per-descriptor memo, since this runs once per bound buffer per plan call.- Return type:
None
- eagle.gemm.plan.assembly_step(name, *, reads, writes, fn)[source]#
A rearrangement written ONCE against an array module:
fn(buffers, xp). Where a plan puts the reshape/scale/scatter-free rearrangement a matmul cannot express; writing it twice is how the two backends drift apart.- Return type:
- eagle.gemm.plan.captured_gemm_calls(plan, buffers, *, operand_class)[source]#
Describe every step of
planas a capture-legal rawSgemm. The seam a captured build reads a lowered plan through. Runs nothing and allocates nothing: it maps each step’s publishedMatmulSpecand the caller’s bound buffers ontoSgemmParams, a pure function over shapes and strides, and hands the results back for the caller to issue inside its own capture region.Three refusals (
GemmMappingUnavailablethroughout):operand_class == "gemv"or any non-float32 buffer; a step that is not a matmul (skipping it would replay a different computation than the plan describes); a name the caller did not bind. Leading dimensions come from each buffer’s strides, so a plan bound to slice views of larger allocations maps correctly.- Return type:
tuple- Parameters:
operand_class (str)
- eagle.gemm.plan.generated_step(name, *, kernel, reads, writes, bind)[source]#
A launch of one of a producer’s own compiled kernels.
bind(buffers) -> kwargsmaps the plan’s names onto the kernel’s declared parameters. One form here, not two: a producer’s kernels already dispatch host-or-device off their inputs, so the twin a plan needs is the one the kernel already has.- Return type:
- eagle.gemm.plan.lower_dense_gemm(record, *, direction='forward', dtype='float64')[source]#
Build the dense-gemm plan for a recognized contraction, from the record alone.
directionpicks the half:"forward"is the contraction itself,"adjoint"its gradients – two plans rather than one, since the adjoint’s input (the output’s cotangent) does not exist when the forward runs.Everything is declaration-derived: each operand’s declared dims give its matrix,
reduce_dimgives the contracted axis, the free axis gives the free extent, and the GEMM adjoints follow from those shapes.recordis consumed by duck-typed attribute access – seeeagle.gemm.contraction’s docstring for why this module never needs to know how the record was produced. NOT synthesized: an output assembly step. The plan’s output is the contracted result with its axes in operand order; how that maps onto the flat per-sample buffer the generated kernel writes depends on the launch domain’s lane decomposition, which the record does not carry, so v1 hands back the labelled result and leaves the flattening to the caller.- Return type:
- eagle.gemm.plan.matmul_fallback_note(error, *, mechanism='dense_gemm')[source]#
The note emitted when a plan’s matmul fails and falls back, naming the abandoned mechanism and the underlying error.
- Return type:
str- Parameters:
error (BaseException)
mechanism (str)
- eagle.gemm.plan.matmul_step(name, *, a, b, out, transpose_a=False, transpose_b=False, transpose_out=False)[source]#
out = op(a) @ op(b), whereopis a transpose when asked.A transpose is a view in both array modules, so an operand orientation costs nothing; the product is written through
out=the same way.transpose_outwrites through a transposed view of the destination, which is how a gradient lands in its operand’s own declared layout when that layout puts the contracted axis first. The device form re-raisesMatmulUnavailableon every call, since a broken cuBLAS shows up on the first one.- Return type:
- eagle.gemm.plan.verify_staged_fanin(record, *, contracted_extent, operand=None)[source]#
Check a record’s staged fan-in against the extent the adjoint GEMM contracts – a check, never a scale factor. The banded adjoint accumulates
fanincontributions into each staged slot, one per lane that read it. The GEMM adjointĀ = Ȳ @ Bperforms that same sum in one contraction: the axis it sums over is Ȳ’s free axis, whose extent is the other operand’s free extent – exactly the number of lanes sharing a slot. So both routes sum the same contributions and no multiplicative correction applies anywhere.If the declared fan-in disagrees with the contracted extent, the GEMM would sum a different set than the banded route accumulates, surfacing as an unlocalizable twin mismatch.
None(no fan-in derived) is not a disagreement and passes.operandnarrows the check to one staged input, needed when two operands are staged: each one’s adjoint contracts the OTHER’s free extent, so a single shared number would be wrong for one of them.- Return type:
None- Parameters:
contracted_extent (int)
See also
API reference — the C++ API (breathe/Doxygen). Interoperability — numpy, cupy and torch interoperability, worked.