Python API reference

Contents

Python API reference#

Complete reference for the eagle Python package, produced by Sphinx autodoc/autosummary from the live package — the signatures shown here are the ones actually exported by the build you have installed.

eagle owns everything about executing compiled kernels: the framework-polymorphic launch skeleton, the capturable kind-dispatched launch primitive, input-role marshalling, the by-value launch ABI, DLPack framework adapters, CUDA-graph capture/replay, composing several already-built launchables into one launch list, and the plugin protocol consumers (KernelRegistry, load_manifest()).

The top-level namespace#

eagle/__init__.py re-exports its whole intended public surface — every launcher, every DLPack adapter, the by-value ABI helpers, the plugin registry — directly as eagle.<name>, so this one page is a complete single-page reference for everything a typical caller needs (see Thirty seconds: eagle.simulate, Interoperability).

eagle

eagle — the Python face of the EAGLE launch engine.

eagle.deploy is bound on first use (so import eagle imports no hawk) and is the same function as eagle.plan.auto():

eagle.deploy(plugin, *, targets=None, cache_dir=None, scalar_type=None, **plan_kw)

An AutoPlan for plugin: run it where its data lives. eagle.deploy is this same function.

plugin is a built plugin, a hawk kernel, or a list built into one bundle (returned as a tuple of plans in order). The build uses hawk’s cache under cache_dir (None: hawk’s default); a repeat call compiles nothing. targets (None: both sides when a GPU is usable – the call returns with the device built, the host still building beside it – else ("host",)), cache_dir and scalar_type are refused for an already-built plugin. scalar_type is the precision hawk builds in: "float64" (None, the default) or "float32"; the planes passed at run time must match it.

plan_kw are plan()’s keywords but structure (inner/gather do not apply); a side the plugin lacks is refused when first selected.

Return type:

AutoPlan | tuple[AutoPlan, ...]

Submodules not re-exported at the top level#

A handful of names are reached by their submodule path rather than directly off eagle — mostly the newer aether-abi/2 bind-by-name door (eagle.plan, eagle.exec), the shared input-role coercion layer, and the GPU/CPU device-dispatch internals.

eagle.plan — PLAN in Python, EXECUTION in C++.

plan() resolves a v2 plugin, an execution structure (always named explicitly, never residency-derived) and a partitioning into a Plan. Placement legality (eagle.exec.check_placement()) is checked here, at plan time, against every partition count the plan will launch.

Plan.run() drives the plugin over every partition and returns the per-sample result, marshalling each role to the ABI shape eagle.roles.classify_arg() resolves it to — the same classifier the v1 launch paths dispatch on. Output planes are allocated in the artifact’s declared scalar_type (eagle.dtypes); a single-output plugin returns the plane itself, never a 1-tuple, and several return a dict keyed by plane NAME — never a positional tuple, since arg_spec’s own order is role-grouped then name-sorted, not the kernel’s authored order.

Plan.bind() is the capture-legal door: it packs the caller’s own named planes into the argument block ONCE and returns a BoundPlan whose launch() issues the structure’s launches and nothing else — no upload, allocation, sync or copy home — so a consumer can bind once and replay the launches as ordinary nodes inside a captured graph.

Under eagle.exec.RankPartition the plan’s partitions stay WHOLE and the rank cut happens inside the structure, so the same plan describes the run at any world size; every rank returns the whole output plane(s).

Both doors share one packer (_pack_args() over _pack_one()), so they cannot disagree about how a role packs.

class eagle.plan.AutoPlan(plugin, plan_kw, host=None)[source]#

Bases: object

A plugin with both targets, run where its data lives — the result of auto().

host/device are the Plan for each target, built lazily; select() picks one from the values a call binds (host- or device-resident; the sample count never enters). run()/bind() delegate to the selected plan. A default deploy with a usable GPU returns with the device side built and the host side still building.

Parameters:

plan_kw (dict)

plugin#
plan_kw#
bind(**planes)[source]#

Plan.bind() on the plan select() picks for planes.

Return type:

BoundPlan

property device: Plan#

The DeviceKernel plan (built on first use).

property host: Plan#

The HostTeam plan (built on first use).

run(**kw)[source]#

Plan.run() on the plan select() picks for kw.

select(planes)[source]#

The plan planes selects (residency()): the device when any is device-resident, else the host; both sides refused.

Return type:

Plan

Parameters:

planes (dict)

class eagle.plan.BoundPlan(plan, arg_spec, planes, n, parts, boxes, addrs, entry, *, device_type, framework, dtype, widths, vec_mutables, mat_mutables, int_uniforms=frozenset({}), active_count=None, copies=())[source]#

Bases: object

A Plan with its argument block already packed over the caller’s own planes — the thing a captured consumer holds.

It owns the packed host-side mirror block and a reference to every plane those mirrors point at (a captured graph holds raw device pointers, so a plane freed too early reads as a correct answer once and somebody else’s memory afterwards), but no plane itself: nothing here is allocated, uploaded, cast or copied home, at bind or at launch.

The bound block is a snapshot, frozen at bind; rebind() changes only the slots it is given.

plan#
device_capacity()[source]#

The safe “no active-set rebuild” bound at this bound plan’s kernel/device (see eagle._launch_policy.latency_regime_capacity() — the smaller of hardware residency and the issue-pipe saturation point, NOT residency alone). 0 for a host-team plan, or when device/kernel properties are unavailable; the caller (eagle._until_done’s auto policy) treats 0 as “unknown”, which never satisfies n <= capacity. By module name (not a top-level import), matching launch()’s own lazy eagle.launch lookup – this stays import-cycle-free.

Return type:

int

launch(stream=None, *, grid=None, block=None)[source]#

Issue this plan’s launches — one per partition — and nothing else. Returns None: the answers are already in the caller’s own planes.

stream is the CUDA stream to issue on (an int, or anything carrying .ptr); None resolves to the framework’s current stream at launch time, not at bind, since the stream a graph captures on is not the one the plan was bound on. A host_team run has no stream and is refused one.

grid is refused (eagle derives it from each partition’s count); block=None defers to eagle’s launch-policy resolver (eagle.launch._resolved_block), the same one the v1 door consults.

Return type:

None

launch_persist(stream=None, *, grid, counter, util=None, block=256, steps=None, stepsum=None, prepared=None)[source]#

Launch <kernel>_persist once – the persistent launch’s whole shape: lanes fetch base + atomicAdd(counter, 1) until the partition’s own count, an already-Terminated sample skipped, each running up to the bound fused_steps word’s budget before writing back and fetching the next; no WHILE graph, no active-set map, no policy kernel – this one call is the whole run.

grid is the block count and block the block size – the caller’s own choice, both geometry knobs this method does not resolve itself (eagle._launch_policy.persistent_geometry() is the documented rule a caller derives them from, from the entry’s own register-limited occupancy and the SM count; block=256 here is only a bare fallback for a caller that does not). counter is a one-element uint32 cupy array the caller has already memset to 0 (not done here, so a caller replaying this inside a captured graph can fold the reset into a graph node instead). util is an optional 2-element uint64 cupy array (active lanes, warp-iterations*32); omitted, the entry gets a null pointer and skips counting. steps is an optional one-element uint32 cupy array the caller has zeroed: the entry maxes each sample’s steps taken in this launch into it (the run’s exact step count); omitted, a null pointer and no report. The entry’s trailing ABI is base, count, nSamples, hawk_next, hawk_util, hawk_steps, hawk_stepsum; stepsum is an optional one-element uint64 cupy array the caller has zeroed, into which the entry adds every step it executes (the run’s total; omitted, a null pointer and no report).

Refused for anything but a single, contiguous partition – the only shape the persist entry’s base/count pair means. prepared (from prepare_persist()) skips the checks and the packing and only launches: the other arguments are then unread.

Return type:

None

Parameters:

block (int)

launch_range(stream=None, *, block=None, steps=None, stepsum=None, prepared=None)[source]#

Launch <kernel>_range once over this bound plan’s single contiguous partition – an active-set automatic kernel’s fast-path equivalent of the plain kernel’s own launch(): the SAME packed argument block the map entry uses (active_map/active_count included, unread), plus base, count, nSamples, hawk_steps, hawk_stepsum; no counter, no util (that is the persist entry’s own shape, launch_persist()). steps is an optional one-element uint32 cupy array the caller has zeroed, into which the entry maxes each sample’s steps taken (the launch’s exact step count); omitted, a null pointer and no report. stepsum is the persist entry’s one-element uint64 total of executed steps (launch_persist()). The entry itself picks single-step vs fused from the bound fused_steps word (1 vs anything else); this call does not choose.

block (None: resolved) takes the SAME policy launch() already resolves its own block with (eagle.launch._resolved_block() -> eagle._launch_policy.resolve_block, keyed off this plan’s own entry – the range entry shares its cubin/registers, so the same block size applies); the grid follows eagle.exec.DeviceKernel’s own grid-from-count rule, over count samples densely (one lane per sample, no work-stealing). A count of 0 launches nothing.

Refused for anything but a single, contiguous partition – the only shape the range entry’s base/count pair means. prepared (from prepare_range()) skips the checks, the block/grid resolution and the packing and only launches.

Return type:

None

property n: int#

The sample count this block was packed for (frozen at bind).

property names: tuple#

The bound names, in arg_spec order.

property partitions: tuple#

The eagle.exec.Partition triples launch() issues, in order (resolved at bind).

persist_entry()[source]#

This bound plan’s <kernel>_persist cupy Function – same compiled unit as the fused entry (hawk’s hawk.emit.cuda.persist_entry). None when unavailable: a host-team plan, a non-automatic kernel, or an artifact built before the persist entry existed. Resolved once, from the loaded device module the plugin’s _keepalive already holds – no new compile, no new module load.

property planes#

the arrays the packed block points at (or, for a sample-major plane, the component-major view that was packed).

Type:

The bound values by name (read-only)

prepare_persist(*, grid, counter, util=None, block=256, steps=None, stepsum=None)[source]#

launch_persist()’s checks and argument packing done ONCE: returns call(stream=None) that only issues the driver launch (a resident runner’s per-run cost). The packed block is re-made if rebind() ran since.

Parameters:

block (int)

prepare_range(*, block=None, steps=None, stepsum=None)[source]#

launch_range()’s checks, block/grid resolution and argument packing done ONCE: returns call(stream=None) that only issues the driver launch (a count of 0 launches nothing). The packed block is re-made if rebind() ran since.

range_entry()[source]#

This bound plan’s <kernel>_range cupy Function – an automatic kernel’s contiguous-range sibling entry (same cubin as the kernel’s own entry; for an active-set kernel active_map/active_count ride along in the SAME packed argument block, bound but never read), which also reports the launch’s exact step count. None when unavailable: a host-team plan, a non-automatic kernel, or a unit built without hawk’s HAWK_FAST_ENTRIES=1 define. Resolved once, same lazy pattern as persist_entry().

rebind(**changed)[source]#

Re-pack only the named slots — one box per name, in one call: for a resident pipeline whose planes move (a double-buffered state swaps) while the plan, partitions and sample count stay the same. self is positional-only, as Plan.run’s and Plan.bind’s are.

A rebound plane is checked exactly as the original was, at the same frozen n — a different length is a different run and must be bound afresh. Returns this same BoundPlan, updated in place; it cannot change a launch already recorded into a graph, only what the NEXT capture or launch sees.

class eagle.plan.Plan(plugin, structure, access, op, partitions, npartitions, n_samples, inner=None, gather=True)[source]#

Bases: object

WHERE and HOW plugin runs — the result of plan().

partitions is the resolved, explicit tuple of eagle.exec.Partition triples the plan will launch (contiguous, covering [0, n_samples)), or None when deferred: no partitions=/n_samples= was given to plan(), so the shape is resolved from the sample count run() observes. Placement legality is already checked by plan(); run() never re-checks it.

Parameters:
  • plugin (object)

  • structure (Structure)

  • access (str)

  • op (str | None)

  • partitions (tuple | None)

  • npartitions (int)

  • n_samples (int | None)

  • inner (Structure | None)

  • gather (bool)

plugin: object#
structure: Structure#
access: str#
op: str | None#
partitions: tuple | None#
npartitions: int#
n_samples: int | None#
inner: Structure | None = None#

The structure each rank’s share runs through under eagle.exec.RankPartition (default HostTeam); None for every other structure, which has no inner.

gather: bool = True#

Whether a eagle.exec.RankPartition run gathers its output planes before returning (default True, every rank holding the whole plane); False is the pre-gather arm, for asserting a rank’s own sub-partition before the gather overwrites it.

partitions_for(n)[source]#

The concrete eagle.exec.Partition tuple this plan runs over a sample count of n: the resolved partitions when explicit, else a fresh even split into npartitions contiguous partitions.

Return type:

tuple

Parameters:

n (int)

run(**kw)[source]#

Drive plugin over every partition through structure and return the assembled per-sample result: the declared output plane as a numpy array for a single-output plugin, or a dict keyed by plane NAME when it declares several (never a positional tuple — the same shape SimResult/ eagle.loaded.LoadedPure() already return).

Device inputs stay on the device: on a DeviceKernel plan, an input already in device memory (cupy, or a CUDA torch tensor) is bound where it lies, and the outputs then come back as cupy arrays, never brought home — a caller-supplied contiguous output of the declared dtype is the caller’s own buffer, written in place. With host inputs only, the outputs come home as numpy arrays. A rank-partitioned run always gathers on the host.

Every output/mutable plane the caller supplies is written into in place, in the caller’s own layout (eagle._layout), and still returned. A read-only supplied array (flags.writeable is False, or a torch tensor with requires_grad=True) raises, naming the argument. An output not supplied is allocated fresh and returned.

A value shaped as a plane’s per-sample head (a number or 0-d array for a scalar, (w,) for a vector) is one sample (eagle._layout): the call runs as a batch of one and every output comes back in its head shape. Mixing one sample with a batch is refused.

An automatic kernel (it reads the reserved fused_steps word) launched here without that word takes exactly one step (one_step_word()); only eagle.until_done() takes several steps per launch.

self is positional-only, so a plugin may declare a plane literally named self without colliding with this method’s own bound-method parameter.

bind(**planes)[source]#

Bind the caller’s own planes by name and pack the argument block once — the capture-legal door (see the module docstring).

Every name this plugin’s arg_spec declares must be given, as the thing the plan’s structure can address: a cupy array for a eagle.exec.DeviceKernel plan, numpy for HostTeam, a plain scalar for a uniform (frozen into its by-value box here). nsamples is derived, never bound. The reserved fused_steps word of an automatic kernel may be left out: a one-step word is then bound for you (one_step_word()); eagle.until_done() binds its own.

Nothing is allocated, coerced or copied, with one exception: a sample-major (N, w) plane (eagle._layout) binds zero-copy when its transpose is C-contiguous, else is copied here into a component-major buffer this BoundPlan owns (with an eagle.LayoutWarning) and refreshes before each launch, copying back after for a written plane — capture-legal, no device-wide sync. The output planes are always the caller’s own: this door cannot allocate one, since it would have to outlive the graph that writes it with no handle for the caller to hold.

What’s checked here (every refusal names the field): every declared name is bound and no stray one is; each plane is an array of the structure’s own framework (a host array to a device plan is refused, not uploaded); it is C-contiguous; its dtype is the artifact’s declared scalar_type (an integer/bool plane passes at its own dtype, as run() does); and its shape agrees with what the artifact declared. An input plane’s component width is NOT checked: a v2 sidecar declares widths for its output planes only.

self is positional-only for the same reason run()’s is.

Return type:

BoundPlan

eagle.plan.auto(plugin, *, targets=None, cache_dir=None, scalar_type=None, **plan_kw)[source]#

An AutoPlan for plugin: run it where its data lives. eagle.deploy is this same function.

plugin is a built plugin, a hawk kernel, or a list built into one bundle (returned as a tuple of plans in order). The build uses hawk’s cache under cache_dir (None: hawk’s default); a repeat call compiles nothing. targets (None: both sides when a GPU is usable – the call returns with the device built, the host still building beside it – else ("host",)), cache_dir and scalar_type are refused for an already-built plugin. scalar_type is the precision hawk builds in: "float64" (None, the default) or "float32"; the planes passed at run time must match it.

plan_kw are plan()’s keywords but structure (inner/gather do not apply); a side the plugin lacks is refused when first selected.

Return type:

AutoPlan | tuple[AutoPlan, ...]

eagle.plan.device_resident(value)[source]#

Whether value is already in device memory (a cupy array, or anything exporting __cuda_array_interface__); a CPU or grad-tracked tensor raises there, so both answer False.

Return type:

bool

eagle.plan.one_step_word(arg_spec, planes, framework='numpy')[source]#

planes plus a one-cell int64 word of 1 under fused_steps when the artifact reads that word and the caller binds none, so a launch outside eagle.until_done() takes exactly one step. framework is where the word lives ("cupy" for a device bind, "numpy" otherwise). Any other artifact gets planes back unchanged.

Return type:

dict

Parameters:
  • planes (dict)

  • framework (str)

eagle.plan.plan(plugin, *, structure, inner=None, npartitions=1, partitions=None, n_samples=None, _exec_access=None, _exec_op=None, _gather=True)[source]#

Build a Plan for running plugin under structure (always named explicitly, never residency-derived).

structure is one of eagle.exec.DeviceKernel / HostTeam / RankPartition / DeviceGroup. inner names the structure each rank’s share runs through under RankPartition (HostTeam by default) and is refused for any other structure. The plan’s partitions stay WHOLE either way — the rank cut happens inside the structure, so the same plan describes the run at any world size. partitions (a list of (base, count, n_samples) triples) takes precedence over npartitions and is validated for contiguity/coverage; otherwise the plan is deferred to a whole view (npartitions=1, the default) or an even split, resolved once Plan.run() observes the sample count. The execution axis is the plugin’s own .exec_access/.exec_op declaration.

Placement legality is checked here, at plan time, against the partition count this plan will launch. A refused plan names the rule.

A plugin declaring .abi_tag == eagle.exec.ABI_TAG_V1 (a legacy v1 LoadedVector/LoadedPure) carries no partition triple, so a partition count above 1 is refused naming “legacy”.

Return type:

Plan

Parameters:
  • npartitions (int)

  • n_samples (int | None)

  • _exec_access (str | None)

  • _exec_op (str | None)

  • _gather (bool)

eagle.plan.residency(planes, *, door='eagle.plan.auto', what='kernel')[source]#

Where planes says the work runs: "device" when any value is device-resident, else "host"; both sides refused, naming one of each (door/what name the caller).

Return type:

str

Parameters:
  • door (str)

  • what (str)

eagle’s execution structures: the thin Python face over the compiled eagle._core aether-abi/2 primitives (Partition, run_device / run_host / run_host_serial, check_placement, fold, and the layout self-check).

This is the EXECUTION half; eagle.plan marshals a plugin’s arg_spec into by-value ABI structs and imports this module rather than duplicating it.

RankPartition alone lives in a second, optional compiled extension, eagle._mpi (built only with -DEAGLE_PYTHON_MPI), imported lazily so an MPI-less install never pays for it and its absence is reported as a RuntimeError naming the build switch, never a bare ImportError.

class eagle.exec.Partition(*args, **kwargs)#

Bases: object

The launch partition triple {base, count, nSamples}, re-exported from the compiled binding.

property base#

(self) -> int

property count#

(self) -> int

is_whole#

Is this the whole view – the only shape a legacy aether-abi/1 plugin may be handed?

property n_samples#

(self) -> int

whole = <nanobind.nb_func object>#
eagle.exec.ABI_TAG_V1 = 'aether-abi/1'#

The two ABI generations’ wire tags (byte-identical to eagle.abi.ABI_TAG_V1 / eagle.abi.ABI_TAG_V2).

eagle.exec.LAYOUT_SYMBOL = 'eagle_layout_sizes'#

The exported symbol name a v2 plugin’s layout self-check carries.

class eagle.exec.Structure(name)[source]#

Bases: object

One of eagle’s execution structures: device_kernel / host_team / rank_partition / device_group. .name is the wire spelling. Instances are the singletons below – never construct one directly.

Parameters:

name (str)

name#
eagle.exec.DeviceKernel = eagle.exec.device_kernel#

The four execution structures. DeviceGroup is a named placeholder whose .run refuses “NCCL not implemented”.

eagle.exec.check_placement(access, structure, npartitions=1)[source]#

Raise ValueError naming the rule if running a body of declared access under structure over npartitions is illegal. structure is a Structure singleton or its wire name.

Return type:

None

Parameters:
  • access (str)

  • npartitions (int)

eagle.exec.fold(op, values)[source]#

Combine values under op (sum/times/max/land) in eagle’s fixed ascending order – never atomics, so the result is reproducible regardless of how the run was partitioned.

Return type:

float

Parameters:

op (str)

eagle.exec.layout_sizes()[source]#

This build’s layout self-check sizes, in LAYOUT_SYMBOL order.

Return type:

list[int]

eagle.exec.check_layout_sizes(sizes)[source]#

Refuse a v2 plugin’s layout self-check if it disagrees with this build’s own layout_sizes(), naming the field.

A tag alone cannot catch a layout mismatch (a wrong-arch or stale-binding artifact can carry a correct aether_abi tag); this is what catches it instead.

sizes is either a partial {field_name: byte_size} dict (only the given keys are checked) or the raw positional sequence a real plugin’s exported LAYOUT_SYMBOL carries (every position checked).

Raises ValueError naming the disagreeing field, or a length mismatch for the positional form.

Return type:

None

Input-role marshalling for every kernel launcher.

The single home for coercing a launch call’s inputs into the contiguous device arrays / by-value scalars the aether GRef/HandleT ABI expects: one implementation per input role (state vector, matrix, per-sample scalar, lookup table, broadcast uniform, writable Mutable, the terminated mask, wide per-sample buffer/gradient), shared verbatim by every launch path so the input contract can never drift between kinds.

Each foreign tensor is imported to cupy through eagle.interop (zero-copy DLPack for a CUDA-resident tensor, a host upload otherwise); the cast/contiguate after that is a copy only when not already the target dtype / C-contiguous. The launchable’s producer resolves scalar_type to the Real dtype dt threaded through here (eagle.dtypes.np_dtype()).

eagle.marshal.require_n(n, *, hint='state-vector input')[source]#

The canonical ‘cannot size the batch’ error, with a kind-appropriate hint.

eagle.marshal.coerce_vec_inputs(kw, names, n=None, widths=None, dt=<class 'numpy.float64'>)[source]#

State vectors -> contiguous (W, N) dt device arrays; returns (dict, n) with N reconciled. widths maps a name to its component count (default 3; a synthesized derivative seed may differ).

eagle.marshal.coerce_wide_inputs(kw, names, n=None, dt=<class 'numpy.float64'>)[source]#

Wide per-sample flat inputs -> contiguous 2-D (rows, N) dt device arrays; returns (dict, n). Unlike a fixed-width state vector, a wide buffer’s true row count (stride * n_in) is a runtime quantity the sidecar never carries, so only the 2-D shape is checked – the caller must match its own n_in.

eagle.marshal.coerce_wide_outputs(kw, names, n=None, dt=<class 'numpy.float64'>, exempt=())[source]#

Wide gradient-scatter outputs (a VJP-derived kernel’s disjoint scatter targets) -> contiguous 2-D (rows, N) dt device arrays, written in place; returns (dict, n). Never auto-allocated (the row count is a runtime quantity, as in coerce_wide_inputs()) – the caller must pass a buffer. exempt names a declared row-indexed destination (e.g. an atomic Accum output’s (rows, 1) buffer) that carries no batch axis, so it binds as-is and skips _reconcile_n.

eagle.marshal.coerce_mat_inputs(kw, shapes, n=None, dt=<class 'numpy.float64'>)[source]#

Per-sample matrix inputs -> contiguous flat (R*C, N) dt device arrays; returns (dict, n). shapes maps each name to its declared (R, C); the caller passes either (R, C, N) or flat (R*C, N), both reshaping to the same row-major SoA buffer.

eagle.marshal.coerce_per_sample(kw, names, n=None, dt=<class 'numpy.float64'>)[source]#

Per-sample scalars -> contiguous (N,) dt device arrays (strict: a 2-D input is rejected, never silently flattened); returns (dict, n).

eagle.marshal.coerce_tables(kw, lookup_counts, dt=<class 'numpy.float64'>)[source]#

Lookup tables / shared constants -> flat (count,) dt handles. The row-major flatten is the binding contract (the kernel reads handle[flat_index]); the count is checked against the declaration.

eagle.marshal.coerce_int_uniform(name, value)[source]#

One int broadcast param -> an exact np.int64, never through dt(value): a float/double route is bit-exact only below 2^53. Not an exact integer (2.5) is refused rather than truncated; 2.0 is accepted, matching fill_mutable().

eagle.marshal.coerce_uniforms(kw, params, dt=<class 'numpy.float64'>)[source]#

Broadcast Param constants -> their declared by-value type. params may be bare names (v1, all float) or decl-carrying entries (eagle.roles.param_decls() normalizes both); an int param binds np.int64 exactly (coerce_int_uniform()), never dt.

eagle.marshal.coerce_mutable(name, value, decl, dt=<class 'numpy.float64'>, adapted=None)[source]#

Coerce a provided per-sample Mutable array to its device buffer; returns (device_array, n). A float/int slot is contiguous (N,) (int always int64); a vector slot a (W, N) SoA array; a matrix slot bound flat (R*C, N) (a contiguous (R, C, N) input reshapes zero-copy, so the kernel’s update lands in the caller’s buffer).

A sample-major vector/matrix slot is adapted first (eagle._layout); when adapted is a dict its record is stored there by name, so the caller can write the result back.

eagle.marshal.fill_mutable(name, value, decl, n, dt=<class 'numpy.float64'>)[source]#

A broadcast / default Mutable buffer: a scalar value filled to (N,) as dt/int64. A vector/matrix slot has no scalar fill – it must be passed as an array.

eagle.marshal.ZERO_MASK_CACHE_CAP = 64#

How many distinct (device, n) masks may be retained (each pins n bytes of device memory for the process’s life). Past the cap the cache bypasses – a fresh private mask, same as an undeclared plugin gets.

eagle.marshal.coerce_terminated(kw, n, *, readonly_mask=False)[source]#

The terminated mask -> a (N,) bool device array (all-false when omitted). readonly_mask (default False) is the sidecar-declared opt-in: when True and the mask is omitted, it is served from the per-(device, n) cache. A passed mask is never cached.

The ONE shared schema-v1 sidecar validator (the Python half of plugin/sidecar.h).

Every Python loader that reads a sidecar – eagle.loaded.LoadedKernel and eagle.host_launch.HostPluginLibrary – funnels through validate_sidecar() before it loads anything (before cupy.RawModule or ctypes.CDLL), so a rejected artifact never reaches a driver load and every reject case is exercisable with stub bytes and no GPU.

Absence is lenient here for pattern and scalar_type (a pre-freeze sidecar never stamped either); per-loader absence rules (e.g. LoadedVector defaulting an absent pattern to "vector") stay with the loader. The neural_block clause is imported, not defined here (raptor.schema.blocks.validate_neural_block_descriptor()): the spine declares that wire contract, this module owns only the kernel-family checks (schema_version, arg_spec roles, scalar_type, pattern, derivative, buffer kind).

The optional ``terminated_readonly`` declaration. A producer may stamp "terminated_readonly": true on a kernel sidecar to declare that its kernel only ever reads the terminated mask, letting eagle.marshal.coerce_terminated() serve an omitted mask from a shared cache instead of allocating a fresh one. Opt-in only the producer can assert; absent means “not declared”. Read by read_terminated_readonly(), not by validate_sidecar() – it is a Python-side launch hint the C++ registry has no use for.

eagle.sidecar.TERMINATED_READONLY_KEY = 'terminated_readonly'#

The optional sidecar key a producer stamps to opt its kernel into the shared all-false terminated mask. See this module’s docstring.

eagle.sidecar.read_terminated_readonly(meta, *, name)[source]#

Whether this sidecar declares its terminated mask read-only; False when the key is absent. Present-but-not-a-bool raises ValueError – a truthiness coercion would let a typo like "false" silently enable the shared buffer.

Return type:

bool

Parameters:
  • meta (dict)

  • name (str)

eagle.sidecar.FINISH_KEY = 'finish'#

The optional sidecar key a producer stamps on a kernel that finishes its own samples. Read by read_finish().

eagle.sidecar.FINISH_COUNTER = 'finished_count'#

The one counter plane a finishing kernel counts into.

eagle.sidecar.FINISH_AUTO = 'auto'#

The finish.steps value of a kernel whose steps per launch are a run-time word (hawk.steps(kernel, "auto")); such a kernel also stamps finish.steps_max.

eagle.sidecar.read_finish(meta, *, name)[source]#

The kernel’s finish declaration, or None when it never finishes: {"mask": "<plane>", "counter": "finished_count", "steps": K} (K an int >= 1, or "auto" paired with an integer steps_max). A malformed declaration raises ValueError naming the artifact.

Return type:

dict | None

Parameters:
  • meta (dict)

  • name (str)

eagle.sidecar.PARAMS_SCHEMA_KEY = 'params_schema'#

The sidecar key carrying the params block’s own wire version. Absent => 1.

eagle.sidecar.PARAM_DTYPES = ('float', 'int')#

float (a Real argument) or int (the 8-byte signed Int p_<name> slot).

Type:

The uniform element types a v2 params entry may declare

class eagle.sidecar.ParamSpec(name, dtype='float')[source]#

Bases: object

One declared broadcast (uniform) parameter: its name and dtype ("float" | "int"). A v1 sidecar’s bare name yields ParamSpec(name, "float").

Parameters:
  • name (str)

  • dtype (str)

name: str#
dtype: str = 'float'#
eagle.sidecar.read_params(meta, *, name, required=True)[source]#

The sidecar’s declared broadcast params as a tuple of ParamSpec. Accepts both wire shapes, keyed on PARAMS_SCHEMA: v1 (absent or 1) is bare names (["mu", "k"], every uniform a Real); v2 is decl-carrying ([{"name": "mu", "dtype": "float"}, ...]).

Raises loudly on a params_schema newer than this build, a shape disagreeing with the declared version, or an unknown dtype – an integer uniform silently bound through a double is wrong from 2^53 up, so every ambiguity here fails instead. required mirrors the two calling conventions in the tree: the device path reads meta["params"], the host path meta.get("params", []).

Parameters:
  • meta (dict)

  • name (str)

  • required (bool)

eagle.sidecar.validate_sidecar(meta, *, name)[source]#

Validate a parsed sidecar against the schema-v1 contract; raise on any breach. Mirrors the C++ eagle::plugin::validate_sidecar check for check, so an artifact accepted by one language is accepted by the other (pinned by the shared conformance corpus). name names the artifact in every error. Returns meta unchanged, for chaining.

Checks, in the C++ order: schema_version (forward-strict, absent => v1); arg_spec roles; scalar_type (absent/empty stays lenient); pattern (value-strict against RECOGNIZED_PATTERNS, absent stays lenient); the neural_block clause (pattern-conditional, called from raptor.schema.blocks.validate_neural_block_descriptor(), before derivative/buffer since it forbids both); derivative’s shape; buffer kind (value-strict against BUFFER_KINDS).

Not checked here: the presence of kernel / arg_spec – those are required keys a loader reads directly (KeyError naming the key), as the C++ parse_sidecar enforces at parse rather than validate. A neural_block descriptor is the exception, since no Python loader reads it: its clause requires both keys itself.

Return type:

dict

Parameters:
  • meta (dict)

  • name (str)

Canonical plugin arg-spec role vocabulary + schema version (the Python half).

The sidecar arg_spec is an ordered list of [role, name] pairs; role is one of the fixed strings below. The C++ half is plugin/roles.h, and a cross-check test (tests/test_roles_vocab.py) asserts the two lists are identical. Every loader validates each arg_spec role against ROLES at load time (forward-strict), so the Python Loaded* and the C++ PluginRegistry accept exactly the same vocabulary.

mat_in is a valid schema-v1 role (a pure kernel’s matrix input). A matrix binds through the same 40-byte GRef mirror as a vector (a matrix GRef is a width-R*C vector GRef; the R x C shape is in-kernel flat indexing only), so every loader packs it as that same layout, validating the bound shape against the sidecar’s mat_shapes / Mutable shape. Only the cupy device path calls eagle.abi.make_gref; the C++ host PluginRegistry and the ctypes HostPluginLibrary each build their own byte-compatible GRefMirror-shaped struct directly. All three constructions produce the identical 40-byte ABI layout; only the source language and POD type differ.

classify_arg() is the single role -> ABI-shape classifier both launch paths (eagle.launch.assemble_args and eagle.host_launch. HostPluginLibrary.run) dispatch on, so an unrecognised role always raises rather than silently falling through in only one path. The two paths still build different artifacts (cupy objects vs ctypes structs); only the classification decision is shared.

eagle.roles.ROLES = frozenset({'accum_out', 'lookup', 'mat_in', 'mutable', 'nsamples', 'out', 'per_sample', 'terminated', 'uniform', 'vec_in', 'wide_in', 'wide_out'})#

The 12 canonical arg-spec roles (schema v1). Keep in sync with plugin/roles.h::kPluginArgRoles (set-equality, order-independent). accum_out names the cross-sample accumulate plane; it resolves to the same ABI as wide_out (see classify_arg() / ARG_TAGS).

eagle.roles.OUTPUT_ROLES = frozenset({'accum_out', 'mutable', 'out', 'wide_out'})#

allocated when the caller supplies none, and returned by eagle.plan.Plan.run().

Type:

The roles of an output plane

eagle.roles.INPUT_ROLES = frozenset({'lookup', 'mat_in', 'per_sample', 'terminated', 'vec_in', 'wide_in'})#

The roles of an input plane the caller supplies by name.

eagle.roles.PER_SAMPLE_ROLES = frozenset({'mat_in', 'mutable', 'out', 'per_sample', 'terminated', 'vec_in'})#

The roles whose plane holds one element (or one column) per sample.

eagle.roles.STATE_ROLES = frozenset({'accum_out', 'mutable', 'wide_out'})#

The roles of a plane a stepping model updates (eagle.simulate’s state).

eagle.roles.SCHEMA_VERSION = 1#

The current + maximum plugin-schema version this loader understands. A sidecar/manifest tagged with a higher schema_version is rejected (forward-strict); an untagged artifact is treated as v1 (backward-lenient). Orthogonal to the "aether-abi/1" ABI tag (see eagle.abi). Single-sourced with raptor.schema.manifest.SCHEMA_VERSION; the C++ half (plugin/roles.h::kPluginSchemaVersion) stays a source-level literal, cross-checked by tests/test_roles_vocab.py.

eagle.roles.MAX_SCHEMA_VERSION = 2#

The highest schema_version check_schema_version() accepts (v1 and v2 both load today, v3+ does not) — a separate constant from SCHEMA_VERSION, not a bump of it. Attribute-derived, so a “future schema version” anywhere in tests is MAX_SCHEMA_VERSION + 1, never a hardcoded literal.

eagle.roles.RECOGNIZED_PATTERNS = frozenset({'neural_block', 'pure', 'vector'})#

The plugin families the shared sidecar validator (eagle.sidecar.validate_sidecar()) can structurally validate. neural_block is recognized but not launched by any loader. Keep in sync with plugin/roles.h::kRecognizedPatterns; sourced from raptor.schema.blocks.ALL_PATTERNS.

eagle.roles.LAUNCH_CERTIFIED_PATTERNS = frozenset({'pure', 'vector'})#

The plugin families a launching entry point will actually bind and run — a subset of RECOGNIZED_PATTERNS. neural_block is recognized but never launched; every launching door uses the shared message in check_launch_certified_pattern(). The Python manifest door (eagle.registry.load_manifest()) dispatches on this set plus neural_block itself (its sole descriptor-consuming branch). Keep in sync with plugin/roles.h::kLaunchCertifiedPatterns.

eagle.roles.EXEC_REF_KINDS = frozenset({'kernel'})#

The kind discriminant of an exec reference ({"kind": ..., "kernel": ...}) on a neural_block descriptor. One value in v1: the referenced artifact is a plain plugin kernel. A future aggregate/plan-bundle kind is a meaning change, so a schema_version bump. Keep in sync with plugin/roles.h::kExecRefKinds; sourced from raptor.schema.manifest.EXEC_REF_KINDS.

eagle.roles.SCATTER_POLICIES = frozenset({'accumulate', 'unique_write'})#

The declared terminal-write contract of a block’s scatter.

scatter_policy declares what the committed result MEANS, never the mechanism eagle uses to commit it. unique_write (v1’s sole value): every (target, slot) is written by exactly one source per step, so the commit is a plain store, deterministic with zero atomics (the bit-exact gate mode applies). accumulate: more than one source may write the same (target, slot) per step, and the result is the carried base plus an order-unspecified sum (band-gated, never bit-exact). Keep in sync with plugin/roles.h::kScatterPolicies; sourced from raptor.schema.blocks.SCATTER_POLICIES.

eagle.roles.NEURAL_REQUIRED_FIELDS = frozenset({'forward_exec', 'in_degree', 'input_width', 'out_degree', 'output_width', 'param_width', 'scatter_policy', 'state_width'})#

The fields a neural_block descriptor must carry, beyond the general required set (schema_version, kernel, pattern, scalar_type, an empty arg_spec). Keep in sync with plugin/roles.h::kNeuralRequiredFields; sourced from raptor.schema.blocks.NEURAL_REQUIRED_FIELDS.

eagle.roles.NEURAL_EXEC_REF_FIELDS = frozenset({'forward_exec', 'jvp_exec', 'vjp_exec'})#

The keys whose values are exec references. forward_exec is required; the two derivative refs are optional-additive. Keep in sync with plugin/roles.h::kNeuralExecRefFields; sourced from raptor.schema.blocks.NEURAL_EXEC_REF_FIELDS.

eagle.roles.NEURAL_FORBIDDEN_FIELDS = frozenset({'aether_abi', 'buffers', 'derivative', 'host_entry', 'mat_shapes', 'mutables'})#

Kernel-machinery fields a descriptor must not carry (each would otherwise be silently ignored rather than fail loudly). aether_abi, derivative, buffers, mutables, mat_shapes, host_entry — a descriptor binds nothing. Keep in sync with plugin/roles.h::kNeuralForbiddenFields; sourced from raptor.schema.blocks.NEURAL_FORBIDDEN_FIELDS.

eagle.roles.BUFFER_KINDS = frozenset({'lookup'})#

The declared-buffer kind vocabulary (schema v1) — the only kinds a loader can bind; an unrecognized kind is rejected up front rather than silently dropped. Sourced from raptor.schema.blocks.BUFFER_KINDS.

eagle.roles.MANIFEST_FORMATS = frozenset({'cubin', 'fatbin', 'ptx'})#

The manifest-entry format vocabulary (schema v1) — the artifact container a manifest entry names; an unrecognized value is refused rather than handed to a loader. Keep in sync with plugin/roles.h::kManifestFormats; sourced from raptor.schema.manifest.MANIFEST_FORMATS.

eagle.roles.validate_roles(arg_spec, *, name)[source]#

Reject an arg_spec carrying a role outside ROLES (forward-strict).

arg_spec is the sidecar’s list of (role, name) pairs; name names the artifact in the error. Mirrors is_valid_role on the C++ side.

Return type:

None

Parameters:

name (str)

eagle.roles.ARG_TAGS = frozenset({'ACCUM_OUT', 'GREF_MAT', 'GREF_VEC', 'HANDLE', 'NSAMPLES', 'UNIFORM', 'WIDE_IN', 'WIDE_OUT'})#

The ABI-shape tags a role resolves to — the classification both launch paths dispatch on, single-sourced so eagle.launch.assemble_args and eagle.host_launch.HostPluginLibrary.run can never diverge.

GREF_VEC/GREF_MAT stay distinct (rather than one shared GREF) because the ctypes host path reads a “mutable” role’s bound value from one of two different caller-populated dicts (self._vec vs self._mat) depending on shape; the cupy device path’s own make_gref call is identical either way.

WIDE_IN/WIDE_OUT and ACCUM_OUT are each kept distinct from HANDLE/WIDE_OUT for the same reason: each reads from its own caller-populated dict (wide_in/wide_out/accum_out), even though construction is byte-identical to a plain scalar-handle pointer.

eagle.roles.classify_arg(role, name, *, vec_mutables=frozenset({}), mat_mutables=frozenset({}))[source]#

The ABI-shape tag (role, name) resolves to — one of ARG_TAGS.

Extracts the classification only: both launch paths still build their own artifacts from whichever tag comes back (a cupy GRef numpy-structured scalar vs a ctypes GRefMirror/ScalarHandle POD) — that construction stays per-path.

  • "out" / "vec_in" -> GREF_VEC; "mat_in" -> GREF_MAT.

  • "mutable" resolves by context: GREF_MAT if name is in mat_mutables, GREF_VEC if in vec_mutables, else HANDLE (a scalar/int Mutable).

  • "per_sample" / "lookup" / "terminated" -> HANDLE.

  • "wide_in" -> WIDE_IN; "wide_out" -> WIDE_OUT.

  • "accum_out" -> ACCUM_OUT (the cross-sample accumulate plane; same construction as WIDE_OUT).

  • "nsamples" -> NSAMPLES; "uniform" -> UNIFORM.

  • Any other role -> ValueError (fail-loud; never a silent skip).

Return type:

str

Parameters:
  • role (str)

  • name (str)

  • vec_mutables (frozenset)

  • mat_mutables (frozenset)

eagle.roles.check_launch_certified_pattern(pattern, *, subject)[source]#

Refuse pattern at a launching door, with the two-branch message this module locks (mirrors the C++ check_launch_certified_pattern in plugin/roles.h message for message).

  • pattern outside RECOGNIZED_PATTERNS — this build has never heard of the family: the message names the supported set and says upgrade eagle.

  • pattern recognized but not certified (e.g. a neural_block descriptor, not runnable) — the message says recognized but not launchable by this loader, with no upgrade suffix.

subject names the artifact, followed directly by " pattern '<value>'". An absent/empty pattern is lenient (a pre-freeze artifact never stamped one).

Return type:

None

Parameters:

subject (str)

eagle.roles.check_schema_version(meta, *, name='<document>', allow_legacy_version_key=False)[source]#

Return the artifact’s plugin-schema version, rejecting one we cannot load.

schema_version absent => v1 (backward-lenient). A version greater than MAX_SCHEMA_VERSION is rejected (forward-strict); v1 and v2 both load today. Mirrors the C++ kPluginSchemaVersion gate (still v1-only pending the twin-site bump).

allow_legacy_version_key scopes the legacy version-key fallback to manifests only (pass True from eagle.registry.load_manifest()); the default (False) matches C++ sidecar behaviour, where an unrecognized version key is ignored.

Re-exported from raptor, which owns the version-compare wording.

Return type:

int

Parameters:
  • meta (dict)

  • name (str)

  • allow_legacy_version_key (bool)

eagle.roles.check_execution_axis(meta, version, *, name='<document>')[source]#

Validate the schema-v2 execution axis, given the already-resolved version (check_schema_version()’s return). v1 documents must carry none of exec_targets/exec_access/exec_op; v2 documents must carry exec_targets + exec_access (absence = load refused) and validate their vocabulary + exec_op’s mapreduce-conditional requiredness.

Re-exported from raptor, which owns the execution-axis shape. Wired into eagle.registry.load_manifest().

Return type:

None

Parameters:
  • meta (dict)

  • version (int)

  • name (str)

eagle.roles.PARAMS_SCHEMA = 2#

The sidecar params block’s own wire version — the current + maximum shape this build can read.

  • 1 (what an absent params_schema key means): "params": ["mu", "k"] — a bare name list, every uniform implicitly a Real.

  • 2: "params": [{"name": "mu", "dtype": "float"}, ...] — decl-carrying, the same {name, dtype} shape mutables has always had, so an int uniform binds through the int binder instead of arriving widened through a double.

Field-local rather than a schema_version bump: the C++ side never reads params at all (it resolves uniforms by arg_spec role), so no C++ reader can misread this block’s new shape. Python-only by construction, except the code generator’s conformance-gated verbatim copy, which must be bumped together with this one.

eagle.roles.param_decls(params)[source]#

Normalize a broadcast-param list to ((name, dtype), …).

One spelling for “what type is this uniform?”, shared by every eagle consumer of a params list. Accepts:

  • a bare str — a v1 sidecar’s name; defaults to float.

  • a (name, dtype) pair.

  • anything with .name / .dtype (eagle.sidecar.ParamSpec or a code generator’s ParamDecl) — so a generated TraceResult.params can be handed straight to a launch.

  • a {"name", "dtype"} dict — a v2 sidecar entry read raw.

An unrecognised entry raises: a uniform whose type cannot be established is the exact silent-wrong-answer this normalization exists to prevent.

eagle.roles.param_names(params)[source]#

Just the NAMES of a broadcast-param list (any of the shapes param_decls() accepts) — the binding-set / kwarg-key view.

Return type:

tuple[str, ...]

eagle.roles.parse_derivative(meta, *, name)[source]#

Return the sidecar’s optional derivative block (or None).

A VJP/JVP derivative artifact carries an additive derivative block describing its role; an ordinary primal/kernel has none. Backward-lenient: an absent block returns None. Validates the block’s shape at load (mirroring the C++ validate_sidecar): kind must be vjp/jvp, and a populated residuals list is rejected (recompute-only for now). The returned dict is the loaded-kernel metadata surface (LoadedKernel.derivative).

Parameters:
  • meta (dict)

  • name (str)

eagle.cuda — the CUDA backend surface (torch.cuda-familiar).

Mirrors the C++ eagle::cuda namespace 1:1 in Python: the nanobind-bound CUDA-graph machinery from the compiled eagle._core extension — the capturable graph (Graph), the stream-capture recorder (StreamCapturer) and its owned result (CapturedGraph), the CUDA stream wrapper (Stream, the cudaStream_t interop backbone), and the instantiated/replayable executable handle (Launcher).

Importing this submodule loads the compiled _core extension, so it is kept out of the top-level eagle import (which must stay usable in a pure-Python / no-GPU install). Reach these as eagle.cuda.Graph etc.

It also carries the capture-introspection trio a graph recorder needs: the forked-stream scope (CaptureFork), the node snapshot of a capture (capture_snapshot_nodes()) and whether a captured node can be toggled (is_node_toggleable()).

The CPU backend (eagle::cpu — the host graph executor + OMP Scan/Reduction) is currently C++-only.

class eagle.cuda.CaptureFork(*args, **kwargs)#

Bases: object

branch#

Raw cudaStream_t of branch index as a Python int (wrap with cupy.cuda.ExternalStream). Only meaningful between fork() and join().

fork#

every branch becomes a sibling of every other, and everything captured so far becomes a predecessor of all branches. Calling twice is a no-op.

Type:

Open the fork

forked#

True once fork() has run and join() has not.

join#

the origin waits for every branch, so work issued after the join depends on all of them. Idempotent. An unjoined branch makes StreamCapturer.end() fail and discard the graph.

Type:

Close the fork

origin#

Raw cudaStream_t of the origin stream this fork branches from.

size#

Number of branches.

class eagle.cuda.CapturedGraph#

Bases: object

debug_dot#

Write this captured graph to Graphviz dot at path and return the text (matches cupy Graph.debug_dot_str).

class eagle.cuda.Graph(*args, **kwargs)#

Bases: object

add_node#

Fold a CapturedGraph in as a child-graph node (consumes it) and harvest its kernel records.

from_captured = <nanobind.nb_func object>#
last_node#
launcher#

Instantiate an exec graph and return a Launcher. Raises RuntimeError if cudaGetLastError() is nonzero after instantiate.

stream#
class eagle.cuda.Launcher#

Bases: object

kernel_node_count#

Number of harvested kernel-node records (cross-check for num_nodes()).

launch#

Replay the instantiated exec graph once. Raises RuntimeError if cudaGetLastError() is nonzero after the replay; see eagle/python/eagle/pipeline.py’s launch() for the companion stream-ordering fix).

set_logical_size#

Patch every kernel node’s grid/block for a new logical size.

set_node_enabled#

Enable/disable one node (by raw handle, as returned by capture_snapshot_nodes()) in this Launcher’s instantiated exec graph. Legal between replays; no recapture, no structure change (mode=”enabled”). Raises RuntimeError (name+code) if the node’s type is not one CUDA supports toggling – check is_node_toggleable() ahead of time.

stream#
synchronize#
class eagle.cuda.Stream(*args, **kwargs)#

Bases: object

ptr#

Raw cudaStream_t as a Python int (wrap with cupy.cuda.ExternalStream).

synchronize#
class eagle.cuda.StreamCapturer(*args, **kwargs)#

Bases: object

begin#
end#

End capture and return a CapturedGraph owning the cudaGraph_t.

The low-level launch stack#

eagle.launch(), eagle.LoadedKernel and friends (eagle.launch, eagle.loaded, eagle.marshal) and the ctypes host launcher eagle.host_launch are the LOW-LEVEL launch stack: one compiled aether-abi/1 kernel launched by hand, with its arguments marshalled per call. New code reaches kernels through eagle.plan.plan(), eagle.deploy(), eagle.until_done() and eagle.simulate(); the low-level stack stays for neural-block manifests and for callers that need exactly one launch and nothing else.

class eagle.launch.LaunchMixin[source]#

Bases: object

The framework-polymorphic launch skeleton, shared by every launcher: owns the origin -> with origin.launch_context(): ... -> if origin.blocking: deviceSynchronize() wrapper and the unknown-keyword guard – the single source of “one call mirrors the input framework”.

class eagle.launch.LaunchPlan(arg_spec, vec_mutables=(), mat_mutables=())[source]#

Bases: object

The kernel-STATIC half of a launch, resolved ONCE per signature: the ABI shape each (role, name) binds as, which Mutables are vector- vs matrix-shaped, and a reusable by-value POD box per argument – all things assemble_args() used to re-derive and re-allocate on every call. Block-size derivation is deliberately not here: it is a per-call function of the batch size, not the kernel.

The boxes are reused, and that is the point. _boxes[i] is refilled in place on every planned launch, which is safe since a launch packs its by-value arguments synchronously – no launch reads a box after fn(...) returns. Do NOT hold an assemble_args result across a second planned launch of the same signature: that list aliases the plan’s boxes. Omit plan to get fresh boxes per call.

Single-threaded launch. Reused per-signature boxes mean two threads launching the same signature concurrently would interleave their refills – not a regression, since the launch path is host-side sequencing behind the GIL and no launch pool in this codebase is threaded.

arg_spec#
vec_mutables#
mat_mutables#
tags#
boxes#
eagle.launch.launch_plan(arg_spec, vec_mutables=(), mat_mutables=())[source]#

The cached LaunchPlan for one kernel signature (identity, not equality); a changed signature gets its own plan. Nothing per-call enters the key.

Return type:

LaunchPlan

eagle.launch.pure_origin(kw, *, vector_inputs, mutable_names, per_sample)[source]#

The caller’s framework, chosen from every per-sample source in signature order (vector inputs, provided Mutable arrays, then per-sample scalars) else numpy. Detected up front so the coercion rides that framework’s stream too.

eagle.launch.DEFAULT_BLOCK = 256#

default launch block size

CPU-plugin launcher – the host twin of eagle.registry, and the Python peer of the C++ eagle::cpu::PluginRegistry (plugin/host_registry.h).

Where eagle.registry launches CUDA kernels over cupy device arrays, this module ctypes-loads a host plugin .so and calls its <kernel>_host entry over host buffers – numpy arrays or torch CPU tensors, zero-copy via their raw data pointer (eagle.interop.host_ptr()).

It shares the SAME binary ABI as the device path: the entry is void <kernel>_host(void* const* params, int32_t n), each params[i] pointing to the same GRefMirror / ScalarHandle POD as plugin/gref_abi.h, packed in arg_spec order – the array the C++ registry would hand cuLaunchKernel, here handed to a host function running its own OpenMP loop. A CPU plugin is the same artifact contract as a GPU plugin, minus the PTX.

class eagle.host_launch.GRefMirror[source]#

Bases: Structure

ctypes mirror of plugin/gref_abi.h GRefMirror (a 40-byte width-independent POD view).

compStride_#

Structure/Union member

data_#

Structure/Union member

deviceId_#

Structure/Union member

deviceType_#

Structure/Union member

sampleStride_#

Structure/Union member

samples_#

Structure/Union member

class eagle.host_launch.ScalarHandle[source]#

Bases: Structure

ctypes mirror of plugin/gref_abi.h ScalarHandle (a 32-byte POD, GRefMirror’s rank-1 sibling; carries both an extent and a device tag, unlike the earlier bare-pointer HandleT).

data#

Structure/Union member

deviceId#

Structure/Union member

deviceType#

Structure/Union member

samples#

Structure/Union member

stride#

Structure/Union member

class eagle.host_launch.HostPluginLibrary(so_path, sidecar)[source]#

Bases: object

A dlopen’d CPU plugin, driven from Python over host buffers.

Mirrors eagle::cpu::PluginRegistry: bind the kernel’s by-name buffers from host pointers the caller already owns (see eagle.interop.host_ptr()), then run() packs the args in arg_spec order and calls the host entry.

so_path is the plugin shared object; sidecar is the same dict the C++ registry and code generator speak: kernel, aether_abi (required), optional host_entry / scalar_type (must be absent/"float64"), arg_spec ([role, name] pairs), and optional mutables / mat_shapes.

Parameters:

sidecar (dict)

bind_vector(name, ptr, n)[source]#

Bind an out / vec_in / vector-mutable SoA buffer by name.

Return type:

HostPluginLibrary

Parameters:
  • name (str)

  • ptr (int)

  • n (int)

bind_matrix(name, ptr, n, rows, cols)[source]#

Bind a mat_in / matrix-mutable flat (R*C, N) buffer. rows/cols are checked against the sidecar’s declared shape; ptr addresses dim = r*C + c, sample-fastest.

Return type:

HostPluginLibrary

Parameters:
  • name (str)

  • ptr (int)

  • n (int)

  • rows (int)

  • cols (int)

bind_handle(name, ptr)[source]#

Bind a per_sample / terminated / scalar-mutable / wide flat buffer by name (all ride the same scalar-handle pointer). A lookup table uses consolidate() instead.

Return type:

HostPluginLibrary

Parameters:
  • name (str)

  • ptr (int)

consolidate(name, ptr, count)[source]#

Consolidate a read-only lookup table – the host twin of the device registry’s consolidate (nothing to upload; just records pointer + count). count is checked against the sidecar’s declared size; a declared table must be consolidated before run().

Return type:

HostPluginLibrary

Parameters:
  • name (str)

  • ptr (int)

  • count (int)

bind_uniform(name, value)[source]#

Bind a uniform scalar by name (the float64 spelling). The change test is bit-exact where == is not: -0.0 == 0.0 but the two are not interchangeable in a kernel, so a rebind between them must invalidate the cache; NaN != NaN already does.

Return type:

HostPluginLibrary

Parameters:
  • name (str)

  • value (float)

bind_uniform_int(name, value)[source]#

Bind a uniform scalar by name as an exact 64-bit signed integer (ctypes.c_longlong, never c_double: that widening is bit-exact only below 2^53). A non-integral value is refused, not truncated (2.0 is accepted as an integer spelling).

Return type:

HostPluginLibrary

Parameters:
  • name (str)

  • value (int)

run(n)[source]#

Pack the args in arg_spec order and call the host entry over @p n samples, the same role -> params[] packing the C++ registry does. Returns the sample count run.

The pack is CACHED: a hit replays the exact array a prior miss built from the same bindings, skipping the validation the miss path performs (it cannot newly fail, since nothing changed since it passed).

Return type:

int

Parameters:

n (int)

Framework bridges#

eagle.frameworks.torch turns a hawk per-sample kernel into a torch.autograd.Function (see Train through a physics kernel with torch). It is the one eagle module that imports torch, and only when it is imported itself.

A hawk per-sample kernel as a torch.autograd.Function: function().

Forward runs the kernel; backward runs the reverse-mode kernel hawk derives from it (hawk.diff.vjp()); forward-mode AD runs the derived tangent kernel (hawk.diff.jvp()) – nothing is taped, each pass is one kernel launch. torch.func transforms are not supported: they hand derivative rules wrapped tensors exposing no storage, and a kernel can only read a buffer.

CPU tensors run through hawk’s host runtime; CUDA tensors through eagle’s launch path (eagle.plan.plan()) on torch’s current stream, no sync. Every plane crosses through eagle.interop.import_buffer() zero-copy; a non-contiguous input is made contiguous, a wrong-dtype one refused.

Conventions, off the kernel’s own declaration: inputs are its planes and Param uniforms in order; outputs are its Mutable planes (bare if one), PLUS the Terminated mask itself when the kernel’s own body finishes it (terminated = cond) – the updated mask comes back as an extra return value, the last one, so a stepping loop reuses the SAME stop decision the kernel made instead of recomputing it with a second masking rule of its own. A per-sample plane is (n,)/(width, n); a sample-major (n, width) input is accepted too (zero-copy transpose where contiguous), and outputs come back in that same layout – mixing layouts in one call is refused. A Param’s gradient is the batch sum of its per-sample partials (held in CUDA, read back to host once per launch). A Terminated mask is never differentiated: a marked sample’s outputs and gradient/tangent stay zero, and the mask itself carries no gradient/tangent of its own.

torch is imported by this submodule only: import eagle stays torch-free.

function(kernel, *[, wrt, cache_dir, ...])

Wrap kernel as a torch-differentiable callable.

KernelFunction(kernel, *[, wrt, cache_dir, ...])

A hawk kernel, its derived reverse- and forward-mode kernels, and the torch.autograd.Function that applies them (see function()).

GEMM helpers#

A capture-legal single-precision GEMM step: raw cuBLAS, bound by ctypes.

Pure ctypes plus a strides-only shape mapping, with zero code-generator imports – this is eagle runtime, not codegen. Cross-module references point at eagle.gemm.plan.

eagle.gemm.plan’s matmul_step calls the array module’s matmul, which cupy refuses to record into a CUDA graph during stream capture. This module is the direct cuBLAS call that bypasses that restriction: it never selects a mechanism, never allocates, and is not wired into matmul_step, which remains the uncaptured implementation.

Three pieces: load_cublas() resolves the shared library from the process’s existing mappings; CublasCaptureHandle is one handle per captured-graph build; sgemm_params() is the row-major to column-major mapping, a pure function over shapes and strides.

Refusals are typed and loud (GemmCaptureUnavailable and its two subclasses), never fallbacks.

eagle.gemm.capture.CUBLAS_OP = {'N': 0, 'T': 1}#

cuBLAS’ own transpose enum, by the letters this module’s mapping speaks.

class eagle.gemm.capture.CublasCaptureHandle(stream_ptr)[source]#

Bases: object

One cuBLAS handle, owned by ONE captured-graph build. The lifetime rules are why this is an object, not a function: cublasSetStream is called once, in the constructor, outside the capture region (cupy itself sets the stream per call, which is why it blanket-refuses cuBLAS during capture); the handle outlives its graph (cuBLAS keeps a per-handle workspace, so destroying it early is a use-after-free at replay); warm() runs on the capture stream, after setStream and before begin_capture, once per GEMM shape (its workspace allocates lazily, so warming elsewhere leaves a cudaMalloc inside the capture region); only sgemm() runs inside the capture region; teardown order is graph, then handle, then pool. alpha/beta are host-side scalars baked into the recorded graph node at enqueue time: changing them means recording again.

property alive: bool#

Whether the handle is still created (False after destroy()).

warm(params, operands)[source]#

Run this GEMM shape once, outside the capture region. Idempotent per shape, so a caller may warm defensively before every recorded step without cost. Writes the destination exactly as the captured call will, so contents that matter must be re-filled after.

Return type:

None

Parameters:
sgemm(params, operands)[source]#

Enqueue ONE Sgemm on the handle’s stream – the only call legal inside a capture region. operands maps "a"/"b"/"out" to the caller’s device buffers; which of a/b is the first cuBLAS operand is params.operands, since the row-major to column-major mapping swaps them for every case except transpose_out.

Return type:

None

Parameters:
destroy()[source]#

Destroy the handle. Call after destroying the graph that recorded work against it, and before freeing the memory pool. Idempotent.

Return type:

None

exception eagle.gemm.capture.CublasResolutionError[source]#

Bases: GemmCaptureUnavailable

The cuBLAS library could not be resolved unambiguously. Zero matches means cupy’s own cuBLAS never initialized; more than one means two different libcublas files are mapped, and picking either would risk a wrong answer or a crash far from here.

exception eagle.gemm.capture.GemmCaptureUnavailable[source]#

Bases: Exception

Base for “this GEMM cannot be run as a captured raw-cuBLAS step”.

exception eagle.gemm.capture.GemmMappingUnavailable[source]#

Bases: GemmCaptureUnavailable

This contraction cannot be expressed as one Sgemm call: a shape, dtype, stride pattern or operand class the mapping does not cover. Permanent for the operands as bound – leave the step uncaptured, not retry.

class eagle.gemm.capture.SgemmParams(op_a, op_b, m, n, k, lda, ldb, ldc, operands)[source]#

Bases: NamedTuple

Everything one cublasSgemm_v2 call needs, bar the pointers. op_a/lda and op_b/ldb are named for the cuBLAS argument position, not the caller’s arrays: operands says which caller buffer fills each position, swapped for every case except transpose_out. m/n/k are the column-major extents cuBLAS is told, not necessarily the row-major result’s own shape.

Parameters:
  • op_a (str)

  • op_b (str)

  • m (int)

  • n (int)

  • k (int)

  • lda (int)

  • ldb (int)

  • ldc (int)

  • operands (tuple)

op_a: str#

Alias for field number 0

op_b: str#

Alias for field number 1

m: int#

Alias for field number 2

n: int#

Alias for field number 3

k: int#

Alias for field number 4

lda: int#

Alias for field number 5

ldb: int#

Alias for field number 6

ldc: int#

Alias for field number 7

operands: tuple#

Alias for field number 8

eagle.gemm.capture.check_plan_capturable(plan, *, operand_class)[source]#

Refuse a lowered plan the captured raw-cuBLAS route cannot run. Two permanent fences: the "gemv" operand class has no mapping here (its lowering contracts against a ones vector), and every declared buffer must be float32.

Return type:

None

Parameters:

operand_class (str)

eagle.gemm.capture.load_cublas()[source]#

The ctypes.CDLL for resolve_cublas_path()’s library, cached per process (the handle-per-graph rule is about the cuBLAS handle, not this library object).

eagle.gemm.capture.mapped_library_paths(maps_text, soname=('libcublas.so.12', 'libcublas.so.13'))[source]#

Every distinct file path in maps_text whose basename is soname (one name, or any of a tuple of names) itself or extended by a real-file suffix – never a bare prefix match (conda’s real file is further-versioned, e.g. libcublas.so.12.9.1.4, and libcublasLt.so.12 must not match). Pure function over a /proc/<pid>/maps dump: the refusals are provable without a GPU. Sorted and deduplicated.

Return type:

tuple

Parameters:

maps_text (str)

eagle.gemm.capture.require_float32(specs)[source]#

Refuse unless every WorkspaceSpec in specs declares float32. Stated as a fence: lower_dense_gemm’s own dtype default is still "float64", so a plan built without an explicit dtype would otherwise reach here declaring a type this step cannot run.

Return type:

None

eagle.gemm.capture.sgemm_params(a_shape, a_strides, b_shape, b_strides, out_shape, out_strides, itemsize, *, transpose_a, transpose_b, transpose_out)[source]#

Map matmul_step’s row-major contraction onto one column-major Sgemm. The contraction is out_view = op(a) @ op(b); a row-major (r, c) buffer with row stride ld, read column-major with leading dimension ld, is that buffer’s transpose. So an untransposed destination asks cuBLAS for out.T = op(b).T @ op(a).T (operands swapped, flags inverted); transpose_out inverts that inversion back to natural order with opposite flags. Leading dimensions come from the strides (_leading_dimension()); refuses anything one Sgemm cannot address, or any itemsize that is not float32’s.

Return type:

SgemmParams

eagle.gemm.capture.sgemm_params_for_arrays(a, b, out, *, transpose_a=False, transpose_b=False, transpose_out=False)[source]#

sgemm_params() read off three live arrays (numpy or cupy).

Also enforces the dtype fence on all three, since here it can be seen. .strides is in BYTES for both array modules, which is what the pure function takes.

Return type:

SgemmParams

The record shape a recognized contraction hands to eagle.gemm.plan — not the recognizer itself.

recognize lives in the code generator (a structural walk over its trace IR); eagle never imports a producer, so this module owns only the shape of the result: ContractionOperand and RecognizedContraction, pure frozen dataclasses with no imports beyond dataclasses. A caller that already has a recognized-contraction result can pass it straight into eagle.gemm.plan.lower_dense_gemm(); a future producer with its own recognizer can construct a RecognizedContraction here directly — neither path depends on the code generator.

class eagle.gemm.contraction.ContractionOperand(name, role, dims, bound, reads, reduce_stride, reduce_dim, stride_labels, dense, refusals)[source]#

Bases: object

One buffer a recognized contraction reads, described from its declaration.

dims is the declared axis layout ((label, size, stride), …) in declared order, () for a positional table with no declared axes. reduce_stride is how far one step of the reduce axis moves this operand’s flat index: 0 if invariant along it, None if the index is not affine in the loop variable at all. reduce_dim names the declared axis that stride belongs to, when it is exactly one declaration’s named constant.

dense is this operand’s half of the contraction’s density verdict; refusals is why not, one sentence per reason.

Parameters:
  • name (str)

  • role (str)

  • dims (tuple[tuple[str, int, int], ...])

  • bound (int)

  • reads (int)

  • reduce_stride (int | None)

  • reduce_dim (str | None)

  • stride_labels (tuple[str, ...])

  • dense (bool)

  • refusals (tuple[str, ...])

name: str#
role: str#
dims: tuple[tuple[str, int, int], ...]#
bound: int#
reads: int#
reduce_stride: int | None#
reduce_dim: str | None#
stride_labels: tuple[str, ...]#
dense: bool#
refusals: tuple[str, ...]#
class eagle.gemm.contraction.RecognizedContraction(carry, reduce_var, reduce_extent, reduce_start, reduce_step, operands, operand_class, term_form, outputs, dense, refusals, staged_fanin, fanin_source, adjoint_accumulates)[source]#

Bases: object

A reduce loop described as a contraction — the whole seam, frozen.

operand_class is "gemv" for a single-operand row reduction or "gemm" for two operands multiplied along the reduce axis. dense is the AND of every operand’s verdict; refusals collects their reasons, each prefixed with the operand it came from.

term_form is the value property beside dense’s index property: "bare_read", "product_of_reads" or "transformed". Both are needed to lower: density says the operands can be walked as matrices, term form says a matmul over those matrices computes what the loop computes.

staged_fanin is how many lanes read one staged slot — the number of cotangent contributions the reverse pass accumulates into it — carried as declared, with None where the trace derives none. fanin_source records where it was read from. adjoint_accumulates is ((name, atomic?), …) for the shared accumulate outputs the derived kernel contributes to.

Parameters:
  • carry (str)

  • reduce_var (str)

  • reduce_extent (int | None)

  • reduce_start (int | None)

  • reduce_step (int)

  • operands (tuple[ContractionOperand, ...])

  • operand_class (str)

  • term_form (str)

  • outputs (tuple[tuple[str, str], ...])

  • dense (bool)

  • refusals (tuple[str, ...])

  • staged_fanin (tuple[tuple[str, int | None], ...])

  • fanin_source (str)

  • adjoint_accumulates (tuple[tuple[str, bool], ...])

carry: str#
reduce_var: str#
reduce_extent: int | None#
reduce_start: int | None#
reduce_step: int#
operands: tuple[ContractionOperand, ...]#
operand_class: str#
term_form: str#
outputs: tuple[tuple[str, str], ...]#
dense: bool#
refusals: tuple[str, ...]#
staged_fanin: tuple[tuple[str, int | None], ...]#
fanin_source: str#
adjoint_accumulates: tuple[tuple[str, bool], ...]#
property operand_count: int#

How many distinct buffers the addend reads (1 or 2 — see operand_class).

property operand_names: tuple[str, ...]#

The read operands’ declared names, in first-read order.

The lowered plan — an ordered, backend-parametric sequence of GEMM steps.

eagle.gemm.contraction says what a reduce loop’s record looks like once recognized; this module is what a lowering actually is, given one. A LoweredPlan is a frozen tuple of named steps plus the metadata a caller needs to run them: the mechanism row it came from, its determinism typing, and the specs of every buffer it touches.

Three properties, each a deliberate exclusion:

  • backend-parametric. Every step carries both a device form and a host twin, so a plan is checked against a numpy reference without a GPU;

  • it allocates nothing. Every buffer is the caller’s, bound by name; the plan publishes WorkspaceSpec s and validates what it is handed against them. In v1 the allocator is the runner;

  • it does not schedule. Steps run in list order, enqueued on the current stream and left there; the caller keeps control of ordering.

What a plan is NOT: a place where a contraction is decided. The steps arrive already chosen. lower_dense_gemm() derives one from the record alone, consuming it entirely by duck-typed attribute access, so a record built by the code generator’s recognize(), or hand-built to the same shape, drives this module identically.

class eagle.gemm.plan.CapturedGemmCall(name, params, operands)[source]#

Bases: object

One plan step, described as the raw Sgemm a captured graph records. operands maps "a"/"b"/"out" to the caller’s own buffers – never copies, since a captured graph records device pointers. Which of "a"/"b" is the first cuBLAS operand is params.operands, not this mapping.

Parameters:
  • name (str)

  • params (object)

  • operands (dict)

name: str#
params: object#
operands: dict#
class eagle.gemm.plan.LoweredPlan(steps, mechanism, determinism, inputs=(), outputs=(), workspaces=())[source]#

Bases: object

An ordered sequence of steps plus everything needed to run them. inputs/outputs/workspaces are all WorkspaceSpec s, uniform since the caller allocates and the plan validates all three the same way; the split is about role: inputs arrive filled, outputs leave filled, workspaces are intermediates a later step reads. Construction validates the read/write graph: a step may only read a name that is an input or was written by an earlier step, may only write a declared name, every output must be written, and every workspace must be both written and later read (one nothing reads would be over-specified).

Parameters:
steps: tuple[PlanStep, ...]#
mechanism: str#
determinism: str#
inputs: tuple[WorkspaceSpec, ...] = ()#
outputs: tuple[WorkspaceSpec, ...] = ()#
workspaces: tuple[WorkspaceSpec, ...] = ()#
property required: tuple[WorkspaceSpec, ...]#

Every buffer the caller must allocate and bind, in declaration order.

property kinds: tuple[str, ...]#
__call__(backend, buffers)[source]#

Run every step in order and return {output name: buffer}. backend is "host" (numpy, sequential) or "device" (cupy, enqueued on the current stream and left there – the caller’s stream ordering is the only ordering). buffers must already hold every name in required, matching its spec: nothing is allocated here, and a missing or mis-shaped binding is refused before any step runs.

Return type:

dict

Parameters:
  • backend (str)

  • buffers (dict)

exception eagle.gemm.plan.LoweringUnavailable[source]#

Bases: Exception

Base for “this lowering cannot be used” — raised at BUILD time (PlanUnavailable) or at CALL time (MatmulUnavailable).

class eagle.gemm.plan.MatmulSpec(a, b, out, transpose_a=False, transpose_b=False, transpose_out=False)[source]#

Bases: object

What a matmul step multiplies, by name, and which way round. A step’s two closures already know this, which was enough while the only consumer was LoweredPlan.__call__(). Recording a plan into a CUDA graph needs more: cupy refuses every cuBLAS call during stream capture, so the captured form is a raw Sgemm issued through eagle.gemm.capture, which needs the operand names, the three transpose flags and the destination – the same description matmul_step() was built from. Published rather than re-derived from the record, since a plan is what actually runs: a second derivation would be a second thing to keep in step with the first.

Parameters:
  • a (str)

  • b (str)

  • out (str)

  • transpose_a (bool)

  • transpose_b (bool)

  • transpose_out (bool)

a: str#
b: str#
out: str#
transpose_a: bool = False#
transpose_b: bool = False#
transpose_out: bool = False#
exception eagle.gemm.plan.MatmulUnavailable[source]#

Bases: LoweringUnavailable

A plan’s matmul FAILED on the device at call time: the failure a mechanism-selection layer’s cublas_available() check cannot predict (a cupy that imports cleanly but whose cuBLAS library is missing, mismatched, or cannot create a handle). The contract is that the caller maps it to an uncaptured fallback plus a loud note (matmul_fallback_note()), never a silent retry or a swallowed genuine shape/dtype bug.

class eagle.gemm.plan.PlanStep(name, kind, reads, writes, device_fn, host_fn, matmul=None)[source]#

Bases: object

One step: what it is called, what kind it is, which buffers it reads and writes, and its two forms. device_fn/host_fn take the bound buffer mapping and return nothing. Two callables rather than one taking an array module, because a generated kernel’s two forms are not the same function with a different xp (a producer’s own dispatch picks the provider), while a matmul’s are.

Parameters:
  • name (str)

  • kind (str)

  • reads (tuple[str, ...])

  • writes (tuple[str, ...])

  • device_fn (object)

  • host_fn (object)

  • matmul (object)

name: str#
kind: str#
reads: tuple[str, ...]#
writes: tuple[str, ...]#
device_fn: object#
host_fn: object#
matmul: object = None#

The step’s own description when it is a matmul, None otherwise. Excluded from comparison/hashing, like the two callables beside it: plan equality is used to check a hand-built plan against a derived one.

run(buffers, backend)[source]#
Return type:

None

Parameters:

backend (str)

exception eagle.gemm.plan.PlanUnavailable[source]#

Bases: LoweringUnavailable

A plan cannot be built for this contraction: a permanent property of the record, not a retry target.

class eagle.gemm.plan.WorkspaceSpec(name, shape, dtype)[source]#

Bases: object

One buffer a plan touches: its name, shape and dtype. The plan never allocates it – this is the declaration the caller allocates against. dtype is a normalized name string, not a numpy dtype object, so the spec stays hashable and comparable.

Parameters:
  • name (str)

  • shape (tuple[int, ...])

  • dtype (str)

name: str#
shape: tuple[int, ...]#
dtype: str#
check(array)[source]#

Raise unless array matches this spec (shape and dtype): a plan that wrote through out= into a wrongly-typed buffer would either raise deep inside the array module or silently downcast. The dtype-name lookup goes through _dtype_name()’s per-descriptor memo, since this runs once per bound buffer per plan call.

Return type:

None

eagle.gemm.plan.assembly_step(name, *, reads, writes, fn)[source]#

A rearrangement written ONCE against an array module: fn(buffers, xp). Where a plan puts the reshape/scale/scatter-free rearrangement a matmul cannot express; writing it twice is how the two backends drift apart.

Return type:

PlanStep

eagle.gemm.plan.captured_gemm_calls(plan, buffers, *, operand_class)[source]#

Describe every step of plan as a capture-legal raw Sgemm. The seam a captured build reads a lowered plan through. Runs nothing and allocates nothing: it maps each step’s published MatmulSpec and the caller’s bound buffers onto SgemmParams, a pure function over shapes and strides, and hands the results back for the caller to issue inside its own capture region.

Three refusals (GemmMappingUnavailable throughout): operand_class == "gemv" or any non-float32 buffer; a step that is not a matmul (skipping it would replay a different computation than the plan describes); a name the caller did not bind. Leading dimensions come from each buffer’s strides, so a plan bound to slice views of larger allocations maps correctly.

Return type:

tuple

Parameters:

operand_class (str)

eagle.gemm.plan.generated_step(name, *, kernel, reads, writes, bind)[source]#

A launch of one of a producer’s own compiled kernels.

bind(buffers) -> kwargs maps the plan’s names onto the kernel’s declared parameters. One form here, not two: a producer’s kernels already dispatch host-or-device off their inputs, so the twin a plan needs is the one the kernel already has.

Return type:

PlanStep

eagle.gemm.plan.lower_dense_gemm(record, *, direction='forward', dtype='float64')[source]#

Build the dense-gemm plan for a recognized contraction, from the record alone. direction picks the half: "forward" is the contraction itself, "adjoint" its gradients – two plans rather than one, since the adjoint’s input (the output’s cotangent) does not exist when the forward runs.

Everything is declaration-derived: each operand’s declared dims give its matrix, reduce_dim gives the contracted axis, the free axis gives the free extent, and the GEMM adjoints follow from those shapes. record is consumed by duck-typed attribute access – see eagle.gemm.contraction’s docstring for why this module never needs to know how the record was produced. NOT synthesized: an output assembly step. The plan’s output is the contracted result with its axes in operand order; how that maps onto the flat per-sample buffer the generated kernel writes depends on the launch domain’s lane decomposition, which the record does not carry, so v1 hands back the labelled result and leaves the flattening to the caller.

Return type:

LoweredPlan

eagle.gemm.plan.matmul_fallback_note(error, *, mechanism='dense_gemm')[source]#

The note emitted when a plan’s matmul fails and falls back, naming the abandoned mechanism and the underlying error.

Return type:

str

Parameters:
  • error (BaseException)

  • mechanism (str)

eagle.gemm.plan.matmul_step(name, *, a, b, out, transpose_a=False, transpose_b=False, transpose_out=False)[source]#

out = op(a) @ op(b), where op is a transpose when asked.

A transpose is a view in both array modules, so an operand orientation costs nothing; the product is written through out= the same way. transpose_out writes through a transposed view of the destination, which is how a gradient lands in its operand’s own declared layout when that layout puts the contracted axis first. The device form re-raises MatmulUnavailable on every call, since a broken cuBLAS shows up on the first one.

Return type:

PlanStep

eagle.gemm.plan.verify_staged_fanin(record, *, contracted_extent, operand=None)[source]#

Check a record’s staged fan-in against the extent the adjoint GEMM contracts – a check, never a scale factor. The banded adjoint accumulates fanin contributions into each staged slot, one per lane that read it. The GEMM adjoint Ā = Ȳ @ B performs that same sum in one contraction: the axis it sums over is Ȳ’s free axis, whose extent is the other operand’s free extent – exactly the number of lanes sharing a slot. So both routes sum the same contributions and no multiplicative correction applies anywhere.

If the declared fan-in disagrees with the contracted extent, the GEMM would sum a different set than the banded route accumulates, surfacing as an unlocalizable twin mismatch. None (no fan-in derived) is not a disagreement and passes. operand narrows the check to one staged input, needed when two operands are staged: each one’s adjoint contracts the OTHER’s free extent, so a single shared number would be wrong for one of them.

Return type:

None

Parameters:

contracted_extent (int)

See also

API reference — the C++ API (breathe/Doxygen). Interoperability — numpy, cupy and torch interoperability, worked.