Plugin schema (v1 and v2)#
A plugin, in one picture#
The code generator — a separate DSL/JIT-compiler library that EAGLE’s launch engine serves — compiles a kernel and writes it to disk as a small set of files, a producer step. Later, at deploy time, EAGLE reads those files back and turns them into something you can call — a loader step. Neither side needs the other’s source code; they agree only on the files’ shape. Concretely, for one compiled kernel:
build_dir/
escape.ptx <- the compiled GPU code (producer output)
escape.json <- the "sidecar": how to call escape.ptx (producer output)
PTX is an intermediate, architecture-portable GPU assembly format — the
GPU driver JIT-compiles it to real machine code at load time, so one
.ptx file runs on any supported GPU generation without being rebuilt for
each one.
and the sidecar itself, annotated:
{
"format": "ptx",
"schema_version": 1,
"pattern": "pure",
"aether_abi": "aether-abi/1",
"kernel": "raptor_kernel",
"per_sample": ["r_max"],
"arg_spec": [["nsamples", "N"], ["per_sample", "r_max"], ["mutable", "fired"]]
}
Reading it top to bottom: this artifact is PTX-format (The arg_spec role vocabulary
below covers format’s siblings), speaks plugin schema version 1
(this page’s own versioning — see The aether_abi tag is separate for how that differs
from aether_abi), is a pure-shaped kernel (as opposed to vector — the
same two kernel shapes the code generator’s docs describe), was compiled against binary
layout aether-abi/1, exposes the entry point raptor_kernel (the one
name every artifact exports, see kernel below), and
arg_spec is the exact, ordered list of arguments a loader must bind to
call it — each a [role, name] pair. A role (nsamples,
per_sample, mutable, and 6 others) says what kind of argument a
name is; The arg_spec role vocabulary below is the complete list.
When you have several kernels to deploy together, one manifest
(manifest.json) lists them all, each with its own artifact + sidecar
pair — see the deployment manifest below for its
exact shape.
A short glossary before the tables#
Producer — the code that compiles a kernel and writes the artifact + sidecar (the code generator’s deploy pipeline:
bundle.py/compile.py).Loader — the code that reads a manifest/sidecar back and hands you a callable object (EAGLE’s C++
PluginRegistry, or the PythonLoaded*classes viaload_manifest).Sink — where a kernel’s result goes. A pure kernel’s sink is its
mutableslots, written in place; a vector kernel’s sink is a single accumulated output value. (The tables below name the exact JSON field and role for each — this is the plain-language version.)Consolidation — the one-time upload of a read-only lookup table before a run starts, as opposed to per-launch arguments, which travel with every call. (The sidecar’s
buffersarray, in the tables below, is what gets consolidated.)``aether_abi`` — a separate version tag from
schema_version, covering binary struct layout rather than JSON shape; see The aether_abi tag is separate.
Now the exhaustive shapes — for daily reference once the picture above makes sense.
Overview#
EAGLE owns the plugin protocol: the on-disk contract between the producer and every loader, as sketched above. A deployment unit is:
one manifest (
manifest.json) — an ordered list of plugin artifacts;one artifact per plugin (
<id>.ptx/.cubin/.fatbin) — the compiled kernel;one sidecar (
<id>.json) per artifact — its launch signature.
This page freezes both JSON shapes as schema v1. The normative machine-readable definitions are the two JSON Schema files (draft 2020-12) shipped alongside this page:
The schema version is single-sourced from raptor.schema.manifest.SCHEMA_VERSION
(Python; re-exported as eagle.roles.SCHEMA_VERSION) ≡
plugin/roles.h::kPluginSchemaVersion (C++); its current frozen value is
1.
The deployment manifest#
Top-level keys of manifest.json:
Key |
Type |
Meaning |
|---|---|---|
|
int ( |
Plugin-schema version. Absent ⇒ treated as v1 (backward-lenient). Newer than the loader ⇒ rejected (forward-strict). |
|
int ( |
Legacy manifest-version alias (always |
|
string |
Plugin family of the set. The code generator’s |
|
const |
Binary-ABI tag (see The aether_abi tag is separate). Orthogonal to |
|
array |
The set’s plugin entries in declared (injection) order. |
Each plugins entry:
Key |
Type |
Meaning |
|---|---|---|
|
string |
Stable, unique plugin id (from the kernel function name). |
|
int ≥ 0 |
Injection order == list index. |
|
bool |
Whether the registry loads this plugin. |
|
string |
Artifact filename, relative to the manifest directory. |
|
string |
Sidecar filename, relative to the manifest directory. |
|
|
Artifact format: arch-portable PTX, single-arch cubin, or multi-arch fatbin. |
The sidecar#
A sidecar is a shared base plus one pattern-keyed variant. Two variants are
KERNEL variants — vector and pure, the shapes a Bundle emits and
the only shapes any loader launches. A third, neural_block, is a descriptor:
it names kernels rather than being one, and nothing launches it (see
The neural_block descriptor).
The base fields (compile.py::_base_sidecar) are present on every KERNEL variant;
a descriptor carries a deliberately smaller base, spelled out in its own section
below:
Key |
Type |
Meaning |
|---|---|---|
|
|
Artifact format the sidecar accompanies. |
|
int ( |
As for the manifest (backward-lenient / forward-strict). |
|
|
Selects the variant. Python’s |
|
const |
Binary-ABI tag (see The aether_abi tag is separate). |
|
|
Real scalar type the kernel was compiled for. |
|
string |
The |
|
array[string] |
Named vector inputs, in binding order. |
|
array[string] |
Named uniform scalar parameters. |
|
array[string] |
Named per-sample scalar inputs. |
|
array[[role, name]] |
The ordered |
Variant additions#
vector adds:
Key |
Type |
Meaning |
|---|---|---|
|
bool |
Whether the |
|
string |
The accumulate-sink contract name (e.g. |
|
object{string→int} |
Per-name vector width (component count) for host |
|
array |
Declared read-only lookup buffers (may be empty; see below). |
pure adds:
Key |
Type |
Meaning |
|---|---|---|
|
array[object] |
The writable per-sample state that is the pure output. Each element is
|
|
object{string→int} |
Per-name vector width (includes vector Mutables). |
|
array |
Declared read-only lookup buffers (may be empty). |
|
array[string] |
Only when the kernel has matrix inputs: the |
|
object{string→[R, C]} |
Only when the kernel has matrix inputs: the |
Each buffers element (compile.py::_buffers_field) is
{name: string, kind: "lookup", dtype: "float", shape: array[int], count: int},
where count == prod(shape) is the flat element count the registry uploads once
at consolidation and binds by value as a flat handle.
The neural_block descriptor#
A neural_block sidecar describes a block, not a kernel. It carries no
artifact of its own and no bindings; it declares the block’s shape and names the
kernels that implement it. Those kernels are ordinary pure plugins, deployed and
loaded through the normal doors — so a neural deployment is a set of familiar
artifacts plus a descriptor that says how they fit together.
That distinction has a hard consequence, and it is the reason this variant exists as
a separate family rather than as extra keys on pure:
Important
neural_block is recognized but never launch-certified. The shared
validator understands and checks it in full; no loader in either language will
bind or run it. A descriptor arriving at any launching entry point is refused
with “is recognized but not launchable by this loader” — deliberately without
an “upgrade eagle” hint, because a newer eagle is not the remedy. Only a genuinely
unrecognized pattern earns that hint.
Descriptor fields#
Required, from the shared base: schema_version, pattern, kernel (the
descriptor’s own wire identity), scalar_type, and arg_spec — which must be
present and empty ([]). scalar_type is stricter here than on a kernel
variant: it is required (no pre-freeze leniency — a family minted after the freeze
has no backward to be compatible with) and must be "float32" or "float64".
Required, descriptor-specific — single-sourced as
eagle.roles.NEURAL_REQUIRED_FIELDS / plugin/roles.h::kNeuralRequiredFields,
which both validators iterate:
Key |
Type |
Meaning |
|---|---|---|
|
int > 0 |
Number of source slots each block gathers from. Renamed from
|
|
int > 0 |
Number of target slots each block scatters to. Renamed from
|
|
int > 0 |
Width of the gathered input vector the block consumes. |
|
int > 0 |
Width of the result the block commits. Must not exceed |
|
int > 0 |
Width of the block’s per-block state row. |
|
int >= 0 |
Width of the block’s parameter slice. Zero is legal: a block with no learned parameters is meaningful. |
|
|
The declared terminal-write contract (see below). |
|
exec reference |
The kernel implementing the block’s forward evaluation. |
Optional: vjp_exec and jvp_exec, two more exec references naming the
derivative artifacts. Absent is always fine; present is validated exactly as
strictly as forward_exec.
An exec reference is {"kind": "kernel", "kernel": "<id>"}. kind is a
value-strict discriminant with one v1 value: the reference names a plain plugin
kernel. It exists as a discriminant rather than a bare string because a future
aggregate or plan-bundle reference would reshape the object (carrying a digest
alongside kind), which is a meaning change and therefore a schema_version
bump — not something a v1 reader should try to interpret.
What a descriptor must not carry#
aether_abi, derivative, host_entry, and any populated buffers,
mutables or mat_shapes are forbidden on a descriptor, and a populated
arg_spec is forbidden too. Every one of them is kernel machinery that a loader
would simply ignore on a descriptor, and a field that is silently ignored is how a
confused producer ships something that does not do what it says. aether_abi in
particular belongs on the artifacts the descriptor references — they carry their own
tags, and a copy on the descriptor is a second spelling that can drift. A
derivative block belongs on the VJP artifact itself, which is exactly where it
lives when you follow vjp_exec to it.
scatter_policy: a contract, never a mechanism#
scatter_policy declares the contract of the block’s terminal write — what the
committed result means — and never the mechanism EAGLE uses to commit it.
unique_write says: every (target, slot) is written by exactly one source per
step. The commit is therefore a plain store, the result is deterministic with zero
atomics, and bit-exact comparison is a legitimate gate mode.
accumulate is the second value: every (target, slot) may be written by more
than one source per step, and the committed result is the carried base plus an
order-unspecified sum of contributions. Its gate mode is band-tolerance, never
bit-exact — unique_write is the only value bit-exact comparison applies to.
Mechanism — atomic add, block-reduce-then-atomic, segmented reduce — is EAGLE’s own
business, chosen on evidence and invisible to the wire. Several of those are simply
different ways to compute the same accumulate contract, faster or more
deterministically; promoting them to schema vocabulary would turn every performance
experiment into a schema event. Changing executors is never a schema change. A third
value will be minted only when a contract genuinely differs again from both of these,
and it will be named after that contract, together with the executor that can
actually run it.
The arg_spec role vocabulary#
arg_spec is an ordered list of two-element [role, name] arrays — role
first. role is one of these 12 canonical arg-spec roles (schema v1).
This set is the single source of truth, enumerated identically in
plugin/roles.h::kPluginArgRoles (C++) and eagle.roles.ROLES (Python),
and pinned by the cross-check test python/tests/test_roles_vocab.py. The
first 9 were present at the schema-v1 freeze; wide_in, wide_out and
accum_out were added later, additively — admitting a role is not a wire
shape change, so none of the three needed a schema_version bump.
Role |
Meaning |
|---|---|
|
The plugin’s output sink (the vector |
|
A named vector input (from |
|
A matrix input to a pure or vector kernel (see the note below;
vector twin — a returning kind’s |
|
A named per-sample scalar input (from |
|
A read-only lookup buffer, bound by value as a flat handle (from
|
|
A writable per-sample state slot — the pure output (from |
|
The per-sample terminated/active mask. |
|
A uniform scalar parameter (from |
|
The sample count. |
|
A named “wide” buffer input, bound by value as a flat handle — the
same wire shape as |
|
A named wide-buffer output: the VJP scatter-gradient counterpart to
|
|
The cross-sample accumulate plane a |
The per-kind subsets (vector = 6, pure = 8 roles) union to exactly this set of 12.
Note
mat_in is a valid schema-v1 role recognized by every loader; the
reference C++ host inject() (plugin/host_registry.h, CPU) packs it
identically to vec_in (a matrix GRef is a width-R*C vector GRef, so
it rides the same GRefMirror), so a matrix pure or vector kernel
host-round-trips today (test_matrix_pure_to_host_matches_direct
/ test_matrix_vector_to_host_matches_direct). The device
PluginRegistry::inject() (plugin/plugin_registry/registry.h, CUDA)
packs it the same way via its own bind_matrix — see the
mattrace fixture in eagle’s C++ embedding example
(eagle’s own userguide/examples page) for a device-side,
driver-loaded matrix input exercised end to end.
C++ reference-consumer scope boundary#
The two C++ PluginRegistry reference consumers
(plugin/plugin_registry/registry.h, the CUDA path; plugin/host_registry.h,
the CPU path) implement the schema-v1 binding surface completely and
correctly, but they do not parse every field the producer emits and the Python
loader reads. Concretely, C++ never reads:
vec_widths— the per-name vector width. The C++ registries never read it, and the underlying binding primitive is not width-limited to begin with:GRefMirroris documented as width-independent (gref_abi.h: “the struct is width-independent… width only changes in in-kernel indexing”), andbind_vectortakes a raw device pointer plus a sample count — no width parameter at all. Buffer sizing is therefore entirely the embedding host’s contract, invisible to the registry. A mis-sized binding (e.g. handing a width-4 buffer to a kernel compiled expecting width-3) is a silent out-of-bounds device read — nothing insidecar.h/manifest.h/ the registry can catch it, because the one piece of information that would (vec_widths) is never consulted. The shipped demos (plugin/plugin_host.cpp,plugin/graph_inject/inject_demo.cu) happen to allocate width-3(3, N)buffers, but that is a property of the demo code, not a limit of the registry or the ABI. A width-generic embedding must size its own buffers correctly; the registry will not warn it if it gets this wrong.accumulate/sink— the vector-pattern accumulate-sink metadata. Both are Python-only; the C++ vector registry always accumulates.the base convenience arrays
vector_inputs/params/per_sample— C++ derives everything it needs to bind purely fromarg_spec(role + name, in order), which is a strict superset of the same information keyed by role. This is not a functional gap:sidecar.h’s own header comment explains the design (“we only need two fields… kernel name and arg_spec”).
In short: the C++ registries bind purely from arg_spec plus raw pointers
and trust the embedding host for buffer sizing (a silent-corruption hazard,
not a missing capability — see the vec_widths note above), and their
binding surface is double-typed only (a separate, genuine capability
limit — see float32 on the C++ registry: a named deferred capability below). Treat them as the reference
implementation of the binding protocol, not as something that validates
every hint the producer is free to emit.
float32 on the C++ registry: a named deferred capability#
eagle::cuda::PluginRegistry’s entire binding surface is double-typed:
the bound-buffer struct’s data pointer (VecBinding::data), bind_vector,
bind_uniform, and consolidate all take double* / double
(plugin/plugin_registry/registry.h). PluginRegistry::from_manifest
therefore rejects any sidecar with scalar_type == "float32" before it ever
loads the artifact. That rejection is correct fail-fast, not a bug — the
guard has been present since eagle’s initial commit (fa6bf85) and is not a
regression.
It is deliberately a different shape of gap from the guard sitting right next
to it in the same function, which admits scalar_type == "softdouble": a
SoftDouble PTX is bit-identical IEEE-754 float64
(aether’s SoftDouble carries a static_assert(sizeof(SoftDouble) == 8),
“bit-compatible with IEEE 754 double”; arrays, uniforms, and the GRef ABI
all stay float64-shaped end to end), so it loads and launches through this same
double-typed surface with zero widening — confirmed by a load-and-run
probe (eagle/tests/test_PluginRegistryDtype.cu,
PluginRegistryDtypeTest.SoftdoubleLoadsAndLaunches). float32 is genuinely
different: an f32 artifact needs float-typed buffers and 4-byte uniform
packing, which changes the layout the registry hands the kernel — a real
capability gap, not a wording problem. Do not read the two rejections as the
same kind of gap: softdouble’s guard was removed because the wire-layout
assumption behind it was wrong; float32’s guard stays because that assumption
is right.
Nothing here blocks float32 users today — eagle’s Python/torch device path
(eagle.LoadedVector / eagle.LoadedPure) already supports
float32 and is the current answer for an f32 consumer. What is deferred is
specifically this C++ registry’s own binding surface.
Trigger. Build float-typed buffers and 4-byte uniform packing for
eagle::cuda::PluginRegistry when the first genuine C++ consumer of an f32
artifact appears — foreseeably a future neural-engine deployment that reaches
the C++ embedding path rather than the Python/torch one. Until that trigger
fires this stays a recorded, deliberate deferral, not an oversight: the
trigger above is the condition for closing it.
Rough size. A float-typed twin of the double-typed binding surface —
the bound-buffer struct, bind_vector / bind_uniform / consolidate,
and 4-byte uniform packing through the per-plugin argument-packing path in
inject() — plus tests. Small-to-medium: on the order of the existing
double surface it would mirror (that surface, header comments and
inject() included, is a few hundred lines), not a one-line relaxation,
because the packed-argument layout inject() builds per plugin is
width-specific end to end, not just pointer-typed.
Compatibility policy#
Validation happens at load time, before any launch, and follows two rules:
- Backward-lenient.
An artifact with no
schema_versionis treated as v1, so every pre-freeze manifest / sidecar stays loadable. For a manifest, the legacyversionkey is honored as a fallback. (Python:eagle.roles.check_schema_version(); C++: thekPluginSchemaVersiongate parsing an absent tag as0⇒ v1.)- Forward-strict.
A loader rejects:
any
arg_specrole outside the 9-role vocabulary (eagle.roles.validate_roles()/is_valid_role);any
schema_versiongreater than the loader supports;any
scalar_typeoutside{"float64", "float32", "softdouble"}(Python: theSCALAR_TYPESenum check ineagle.loaded.LoadedKernel; C++:validate_sidecarinplugin/sidecar.h). An absentscalar_typestays backward-lenient (a pre-dtype sidecar, treated as"float64") — only a recognized value is required, not a present one; andany
patternvalue outside{"vector", "pure"}, at both the manifest and sidecar level (“the discriminant rule”:patternselects a schema variant, so every loader that reads it — even one that does not dispatch on it — must refuse a value it cannot name, rather than silently bind an unfamiliar variant as an ordinary kernel). Python:registry.py::load_manifest(manifest) andeagle.loaded.LoadedKernel._require_pattern()(sidecar, via eachLoaded*subclass). C++:PluginRegistry::from_manifestin bothplugin/plugin_registry/registry.h(device) andplugin/host_registry.h(host) for the manifest level; the sharedvalidate_sidecarinplugin/sidecar.hfor the sidecar level (so the publicadd_pluginentry point — which bypasses manifests entirely — is covered too); andany
aether_abithat is absent (empty), not just one that is present but mismatched — a compatibility gate, not an optional hint, so presence is required at every entry point that binds a sidecar or manifest: the two manifest-level loads (PluginRegistry::from_manifestinplugin/plugin_registry/registry.handhost_registry.h::load) and the sidecar-level, manifest-bypassingadd_plugin(plugin/host_registry.h) all reject an emptyaether_abithe same way they reject a mismatched one. The Python peer ofadd_plugin—eagle.host_launch.HostPluginLibrary— was the last absence-lenient door in either language (a private inline check, noteagle.abi.check_aether_abi()); harmonized to presence-required too, so this is now true of every loader in both languages, not just the C++ ones named above.
Both loaders enforce the same vocabulary and the same version gate. The C++
half is plugin/roles.h and the Python half is eagle/roles.py; the
cross-check test python/tests/test_roles_vocab.py asserts the two role lists
are set-equal (and the same for the RECOGNIZED / LAUNCH-CERTIFIED
pattern sets, the buffer-kind vocabulary, and the manifest-entry format
vocabulary), the two schema_version constants agree, and every committed
fixture uses only canonical roles. eagle/python/tests/test_schema_hardening.py
and eagle/tests/test_SchemaHardening.cpp fix the exact set of forward-strict
fields — now nine: schema_version and aether_abi (class gate);
kernel and arg_spec role (class structural); pattern,
scalar_type, buffer kind, derivative.kind, and manifest-entry
format (class discriminant) — so the strict/lenient split cannot erode
silently, and a CLASS mismatch (a loader implementing a private literal compare
for a classed field outside the shared validation layer) is caught too. See the
frozen-field-list note in each file for the version-bump rule and the full CLASS
definitions.
pattern’s absence handling is the one place C++ and Python genuinely
diverge, by design, not a gap:
At the manifest level, Python’s
load_manifestis absence-strict (an absentpatternis rejected, same as an unknown one) — it must dispatch toLoadedVectororLoadedPureand has no third option. The two C++ registries stay absence-lenient (a pre-freeze manifest that never stamped one still loads): neither registry dispatches a variant class offpattern, binding is driven entirely byarg_spec, so there is nothing an absent value would fail to gate. Pinned mechanically by the shared conformance corpus (a manifest with nopatternkey, deliberately also carrying a duplicate plugin id: C++ passes the lenient gate and rejects on the duplicate id instead; Python’s strict gate rejects first) — theoverridesmechanism ineagle/tests/conformance/row13_manifest_pattern_absence_asymmetry.*encodes both outcomes for one fixture, so a future change that accidentally harmonizes the two languages fails a test rather than drifting unnoticed.At the sidecar level, C++’s
validate_sidecaris likewise absence-lenient, for the same arg_spec-driven reason (Python’s shared validator agrees — both languages accept apattern-less sidecar at this layer, pinned byeagle/tests/conformance/row14_sidecar_pattern_absence_lenient.*). Layered on top, Python is per-Loaded*-class:LoadedVectordefaults an absent value to"vector"(lenient), whileLoadedPurerequires an exact match (no default) — this further split has no C++ counterpart to diverge from (C++ has no per-kind sidecar loader at all), and pinning it needs a real device module load (_require_patternruns afterLoadedKernel.__init__’scupy.RawModuleload), which the corpus’s GPU-free design cannot accommodate. It is instead pinned directly —test_loaded_vector_defaults_absent_pattern_to_vector/test_loaded_pure_rejects_absent_patternineagle/python/tests/test_schema_hardening.py— against the committedgravity/bumpPTX fixtures. The asymmetry itself remains locked behaviour, not a defect.
This is narrower than an earlier state: previously C++ did not even reject an
unrecognized pattern value (a manifest-only fix would also have left the
sidecar-level, manifest-bypassing add_plugin entry point open — see
eagle/tests/test_SchemaHardening.cpp and test_PluginRegistryManifest.cu
for the registry-level rejection tests, and test_HostPlugin.cpp-adjacent
tests in test_SchemaHardening.cpp for the add_plugin path).
The aether_abi tag is separate#
aether_abi ("aether-abi/1") is a distinct, ABI-stable tag — the by-value
binary-layout contract of the GRef / HandleT POD mirrors the artifact was
built against. It is orthogonal to
schema_version: the schema version governs the JSON shape and the role
vocabulary, whereas aether_abi governs the binary struct layout a loader binds
by value. A loader rejects an aether_abi mismatch before binding any by-value
struct (Python: eagle.abi.check_aether_abi(); C++: the EAGLE_AETHER_ABI
static-assert gates). The value is historical and must not change without a
coordinated C++/Python bump — pre-split artifacts remain loadable precisely
because it is stable.
The v2 execution axis (aether-abi/2) and eagle’s Python plan/exec API#
Schema version 2 adds one new axis to the manifest: which execution
structure(s) — a single device kernel launch, an OpenMP team of CPU
threads, an MPI rank partition, and (planned) an NCCL device group —
may drive a plugin body, and how that body is allowed to touch data outside
its own sample. A schema-v2 manifest tags itself "schema_version": 2 and
"aether_abi": "aether-abi/2", and MUST additionally carry three new
top-level, flat keys (absence is a load refusal, never a silent fallback to
v1 behaviour):
Key |
Meaning |
|---|---|
|
A non-empty list drawn from |
|
How the body touches data: one of |
|
The reduction operator ( |
A v2 artifact’s entry points additionally take an explicit partition
triple {base, count, nSamples} after their ordinary role arguments —
eagle, never the plugin, decides how a run is split and drives the same
body once per partition (a device kernel launch, or one CPU thread team’s
tile). The accum_out role (see The arg_spec role vocabulary) names the cross-sample
accumulate plane a cross_sample_write body writes into.
Legacy v1 artifacts are unaffected. A schema_version: 1 /
aether_abi: "aether-abi/1" plugin carries none of the three keys above
(carrying one is itself a strict-key violation) and is always run
whole-view, on a single structure — precisely today’s behaviour, byte for
byte. eagle.abi.check_aether_abi accepts both tags; only a v2-tagged
artifact may be split across more than one partition.
Placement legality is a function of exec_access and the chosen
structure alone, and is refused at plan time, naming the rule, never
guessed or discovered mid-launch: sample_local / cross_sample_read /
mapreduce run under any implemented structure; cross_sample_write
runs on a single-device structure only (splitting it across partitions would
need a partial-accumulate combine step not yet specified) — the same
plain-language rule eagle.exec.check_placement() enforces mechanically.
Driving a v2 body from Python — eagle.plan / eagle.exec#
eagle.exec is the execution-structure vocabulary: the
DeviceKernel / HostTeam /
RankPartition / DeviceGroup
singletons, plus check_placement() and
check_layout_sizes() (a self-check: a v2 artifact’s
exported eagle_layout_sizes must agree with this build’s own POD sizes,
or the load is refused naming the disagreeing field — a matching
aether_abi tag alone cannot catch a stale or wrong-arch binding).
eagle.plan is where a run is PLANNED: eagle.plan.plan() maps a
plugin, a chosen structure, and an explicit partitioning to a Plan, and
Plan.run(**kw) drives it:
import eagle.exec as eexec
from eagle import plan as eplan
# Selection is always EXPLICIT (no residency-driven auto-choice yet) --
# the caller names the structure every time.
whole = eplan.plan(my_v2_kernel, structure=eexec.DeviceKernel).run(x=x, a=2.0, b=-1.5)
# Splitting a `sample_local` body into partitions never changes the
# per-sample answer -- eagle drives the identical body per partition.
split = eplan.plan(
my_v2_kernel,
structure=eexec.DeviceKernel,
partitions=[(0, 32, 64), (32, 32, 64)],
).run(x=x, a=2.0, b=-1.5)
# The same body under HostTeam (OpenMP) instead of DeviceKernel (CUDA):
host = eplan.plan(my_v2_kernel, structure=eexec.HostTeam).run(x=x, a=2.0, b=-1.5)
# An illegal placement is refused at plan() time, naming the rule:
eplan.plan(
my_v2_kernel, structure=eexec.DeviceKernel,
partitions=[(0, 32, 64), (32, 32, 64)], exec_access="cross_sample_write",
) # -> ValueError: "... exec_access 'cross_sample_write' is legal on
# single-device structures only ..."
partitions= is a list of explicit (base, count, n_samples) triples
(contiguous, covering [0, n_samples)); omit it (with npartitions=
left at its default of 1) for a deferred whole-view run, whose size is
resolved from the arrays passed to .run(). See
eagle/python/tests/test_exec_contract_rows.py for the certification rows
this API is held to (bit-exact partition identity for a sample_local
body, a ruled host/device tolerance band, and the two refusal cases above).
Running across MPI ranks — RankPartition#
RankPartition is the multi-process structure. Rank r of
R runs the r-th contiguous sub-partition of one whole partition
through an inner structure — HostTeam or DeviceKernel — carrying
its own {base, count, nSamples} triple, with nSamples unchanged: a
rank is one more cut of the same run, not a smaller run. Inputs are
replicated (which is why a cross_sample_read body is legal here: the
tables it reads are whole on every rank), and outputs are gathered, so
every rank ends holding the whole plane, bit-identical to a single whole run.
cross_sample_write stays refused — two ranks accumulating into one shared
target would need a partial-accumulate combine that is not specified.
# The plan's partitions stay WHOLE partitions: the rank cut happens INSIDE
# the structure, so the same plan describes the run at any world size.
whole = eplan.plan(
my_v2_kernel, structure=eexec.RankPartition, inner=eexec.HostTeam,
).run(x=x, a=2.0, b=-1.5) # every rank gets the WHOLE output plane
rank, size = eexec.RankPartition.world() # (0, 1) outside a launcher
mine = eexec.RankPartition.local(eexec.Partition.whole(len(x)))
A mapreduce body folds its own rank’s partials and then combines the R
partials in a fixed rank order
(allgather_fold()) — never an
MPI_Allreduce, whose combine order is the implementation’s business, so
two ranks could disagree in the last bits.
This structure lives in a separate, optional compiled extension
(eagle._mpi, built with -DEAGLE_PYTHON_MPI=ON; AUTO, the default,
builds it wherever CMake finds an MPI). eagle._core therefore never links
libmpi, and import eagle costs nothing to a user with no MPI
installed; where the extension is absent, every RankPartition call
refuses with a RuntimeError naming that build switch. Note that
initialising MPI is not a passive act — Open MPI installs process-wide memory
hooks — so drive this structure from a process dedicated to it (an mpirun
world, or a short-lived child), not from inside a long-running session that
is also using CUDA. The certification rows are
eagle/python/tests/test_rank_partition.py (single process) and the
two-rank bed eagle/python/tests/mpi/ (run through its own gate,
tests/mpi/check_rank_bed.sh).
Reserved capability fields#
The sidecar schema reserves an optional launch object for future
launch-configuration metadata:
"launch": {
"block": 256,
"shmem_bytes": 0,
"grid": {}
}
block— int (reserved: thread-block size hint)shmem_bytes— int (reserved: dynamic shared-memory bytes hint)grid— object (reserved: freeform grid-configuration hint)
These fields are reserved in v1: unpopulated, and the current loaders ignore
them. They may be populated in a future minor revision without a
``schema_version`` bump. The schema keeps additionalProperties permissive on
both the launch object and the top-level object so unknown/future keys do not
fail validation of the reserved surface.
This ignore-unknown guarantee — for launch specifically, and for any
unrecognized top-level sidecar/manifest key generally — is exercised by an
executable test on both loaders, not just asserted in prose:
eagle/python/tests/test_schema_hardening.py and
eagle/tests/test_SchemaHardening.cpp. No launch hint is populated by the
producer today; adding one is a deliberate design decision, not a casual
addition.
Note
Do not confuse this reserved launch sidecar field with eagle’s own,
unrelated idealBlockSize concept in the native CUDA-graph machinery
(eagle/cuda/{Graph.h, CapturedGraph.h, Launcher.h, ComputeBlocks.h}) — a
per-captured-kernel blockDim-reconciliation cap used when composing native
eagle graph nodes. The two share a naming convention, not a design link.
The schema files#
manifest-v1.schema.json
{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"$id": "https://github.com/amasat01/eagle/raw/main/docs/schemas/manifest-v1.schema.json",
"title": "EAGLE plugin deployment manifest (schema v1)",
"description": "Frozen schema v1 for the EAGLE deployment manifest emitted by the code generator's deploy pipeline (deploy/bundle.py). A manifest describes one deployment set: an ordered list of plugin artifacts (+ their sidecars) that the C++ PluginRegistry (or the Python loaders) consume as a self-contained unit. The list order is the injection order. Validators are backward-lenient (an absent schema_version is treated as v1) and forward-strict (a schema_version newer than the loader is rejected). Unknown/future top-level keys are tolerated so a future minor revision may add keys without a schema_version bump; additionalProperties is therefore true.",
"type": "object",
"required": ["aether_abi", "plugins"],
"additionalProperties": true,
"properties": {
"schema_version": {
"const": 1,
"description": "Plugin-schema version this manifest conforms to. Single-sourced from eagle.roles.SCHEMA_VERSION / plugin/roles.h::kPluginSchemaVersion. Absent ⇒ treated as v1 (backward-lenient); a value greater than the loader supports ⇒ rejected (forward-strict). Orthogonal to aether_abi."
},
"version": {
"type": "integer",
"description": "Legacy manifest-version alias (always 1). Retained for pre-freeze readers; superseded by schema_version, which the loaders read. A loader falls back to this key when schema_version is absent."
},
"pattern": {
"type": "string",
"description": "Tag naming the plugin family of the bundle. The code generator's Bundle builder emits one of the canonical values \"vector\" | \"pure\" (a Bundle is pattern-homogeneous). The Python load_manifest / C++ PluginRegistry are STRICT: an unknown OR absent pattern is rejected. It is typed as a string (not enum-constrained) so a future family can be added without a schema bump, but the current loaders accept only \"vector\" | \"pure\"."
},
"aether_abi": {
"const": "aether-abi/1",
"description": "Binary-ABI tag of the by-value GRef / HandleT POD mirrors this set was built against. Separate from and orthogonal to schema_version: it is the ABI-stable binary-layout contract. A loader rejects a mismatch before binding any by-value struct."
},
"plugins": {
"type": "array",
"description": "The set's plugin artifacts in declared (injection) order.",
"items": { "$ref": "#/$defs/pluginEntry" }
}
},
"$defs": {
"pluginEntry": {
"type": "object",
"description": "One artifact in the set: its stable id, injection order, enabled flag, artifact + sidecar filenames (relative to the manifest directory), and the artifact format.",
"required": ["id", "order", "enabled", "artifact", "sidecar", "format"],
"additionalProperties": true,
"properties": {
"id": {
"type": "string",
"description": "Stable, unique plugin id (from the kernel function name; duplicates get a numeric suffix)."
},
"order": {
"type": "integer",
"minimum": 0,
"description": "Injection order == list index (the deterministic sequence in which a vector set accumulates into outVec)."
},
"enabled": {
"type": "boolean",
"description": "Whether the registry loads this plugin."
},
"artifact": {
"type": "string",
"description": "Artifact filename (<id>.<ext>), relative to the manifest directory."
},
"sidecar": {
"type": "string",
"description": "Sidecar filename (<id>.json), relative to the manifest directory."
},
"format": {
"enum": ["ptx", "cubin", "fatbin"],
"description": "Artifact format: arch-portable PTX, single-arch cubin, or multi-arch fatbin."
}
}
}
}
}
sidecar-v1.schema.json
{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"$id": "https://github.com/amasat01/eagle/raw/main/docs/schemas/sidecar-v1.schema.json",
"title": "EAGLE plugin sidecar (schema v1)",
"description": "Frozen schema v1 for the per-artifact JSON sidecar emitted by the code generator's deploy pipeline (deploy/compile.py). A sidecar describes one plugin's launch signature: a shared base plus one of two pattern-specific variants (vector | pure). arg_spec is the ordered list of [role, name] pairs a loader binds; every role must be one of the 9 canonical schema-v1 roles enumerated inline (plugin/roles.h::kPluginArgRoles ≡ eagle.roles.ROLES). Validators are backward-lenient (absent schema_version ⇒ v1) and forward-strict (an unknown role, or a schema_version newer than the loader, ⇒ rejected). Unknown/future top-level keys are tolerated (additionalProperties true) so a future minor revision may add keys — including populating the reserved launch surface — without a schema_version bump.",
"type": "object",
"additionalProperties": true,
"required": [
"format",
"pattern",
"aether_abi",
"scalar_type",
"kernel",
"vector_inputs",
"params",
"per_sample",
"arg_spec"
],
"properties": {
"schema_version": {
"const": 1,
"description": "Plugin-schema version this sidecar conforms to. Single-sourced from eagle.roles.SCHEMA_VERSION. Absent ⇒ treated as v1 (backward-lenient); newer than the loader ⇒ rejected (forward-strict). Orthogonal to aether_abi."
},
"format": {
"enum": ["ptx", "cubin", "fatbin"],
"description": "Artifact format the sidecar accompanies: arch-portable PTX, single-arch cubin, or multi-arch fatbin."
},
"pattern": {
"enum": ["vector", "pure"],
"description": "Plugin family — selects the variant-specific required fields (see the conditional blocks). A loader refuses the wrong artifact family; an unknown OR absent pattern is a hard error (strict)."
},
"aether_abi": {
"const": "aether-abi/1",
"description": "Binary-ABI tag of the by-value GRef / HandleT POD mirrors the artifact was built against. Separate from and orthogonal to schema_version: the ABI-stable binary-layout contract. A loader rejects a mismatch before binding any by-value struct."
},
"scalar_type": {
"enum": ["float64", "float32", "softdouble"],
"description": "Real scalar type the kernel was compiled for."
},
"kernel": {
"type": "string",
"description": "The extern-\"C\" entry-point symbol the loader resolves in the artifact: always \"raptor_kernel\", the name raptor's schema fixes for every producer, hand-written kernels included."
},
"vector_inputs": {
"type": "array",
"items": { "type": "string" },
"description": "Named vector inputs, in binding order (mirrors the vec_in arg_spec roles)."
},
"params": {
"type": "array",
"items": { "type": "string" },
"description": "Named uniform scalar parameters (bound via the uniform role)."
},
"per_sample": {
"type": "array",
"items": { "type": "string" },
"description": "Named per-sample scalar inputs (bound via the per_sample role)."
},
"arg_spec": {
"$ref": "#/$defs/argSpec"
},
"vec_widths": {
"$ref": "#/$defs/vecWidths"
},
"launch": {
"$ref": "#/$defs/launch"
},
"accumulate": {
"type": "boolean",
"description": "Vector only: whether the sink (out) uses the outVec `+=` accumulate idiom."
},
"sink": {
"type": "string",
"description": "Vector only: the accumulate sink contract name (e.g. \"outVec-scratch\") — a per-kernel deterministic-reduce scratch slot, not the final total."
},
"buffers": {
"$ref": "#/$defs/buffers"
},
"mutables": {
"$ref": "#/$defs/mutables"
},
"matrix_inputs": {
"type": "array",
"items": { "type": "string" },
"description": "Pure only, and only when the kernel has matrix inputs: the mat_in binding order (mirrors vector_inputs)."
},
"mat_shapes": {
"$ref": "#/$defs/matShapes"
},
"derivative": {
"type": "object",
"description": "Optional, additive schema-v1 field (no version bump): present only for a VJP/JVP derivative artifact (autodiff, generated or custom — structurally indistinguishable). Metadata only; the artifact still launches as an ordinary pure kernel. Recompute-only in v1, so residual_policy is 'recompute' and residuals is []; a populated residuals list is a future revision that bumps schema_version and is rejected by a v1 loader.",
"properties": {
"kind": { "enum": ["vjp", "jvp"] },
"primal": { "type": "string", "description": "the primal kernel's logical id" },
"wrt": { "type": "array", "items": { "type": "string" }, "description": "the differentiated input names" },
"residual_policy": { "const": "recompute" },
"residuals": { "type": "array", "maxItems": 0, "description": "always empty (recompute-only)" }
},
"required": ["kind", "wrt", "residual_policy", "residuals"],
"additionalProperties": true
}
},
"allOf": [
{
"if": {
"required": ["pattern"],
"properties": { "pattern": { "const": "vector" } }
},
"then": {
"required": ["accumulate", "sink", "vec_widths", "buffers"]
}
},
{
"if": {
"required": ["pattern"],
"properties": { "pattern": { "const": "pure" } }
},
"then": {
"required": ["mutables", "vec_widths", "buffers"]
}
}
],
"$defs": {
"role": {
"description": "The 9 canonical schema-v1 arg-spec roles. Single source of truth: plugin/roles.h::kPluginArgRoles ≡ eagle.roles.ROLES (cross-checked by python/tests/test_roles_vocab.py). A loader rejects any role outside this set at LOAD time (forward-strict). Note: mat_in is a valid role recognized by every loader, but the reference C++ host inject() does not pack matrix arguments (a CUDA-backend capability).",
"enum": [
"out",
"vec_in",
"mat_in",
"per_sample",
"lookup",
"mutable",
"terminated",
"uniform",
"nsamples"
]
},
"argSpec": {
"type": "array",
"description": "Ordered list of [role, name] pairs (role FIRST) the loader binds in sequence.",
"items": {
"type": "array",
"minItems": 2,
"maxItems": 2,
"prefixItems": [
{ "$ref": "#/$defs/role" },
{ "type": "string", "description": "The argument name (a vector_input / param / per_sample / mutable / buffer name, or a fixed slot name such as \"out\"/\"terminated\"/\"nsamples\")." }
],
"items": false
}
},
"vecWidths": {
"type": "object",
"description": "Per-name vector width (component count) so a host can allocate (W, N) buffers; the GRef struct itself is width-independent. Present on both variants.",
"additionalProperties": { "type": "integer" }
},
"buffers": {
"type": "array",
"description": "Declared read-only lookup buffers (the consolidation contract). Present on both variants. May be empty.",
"items": { "$ref": "#/$defs/buffer" }
},
"buffer": {
"type": "object",
"description": "One read-only lookup buffer the registry uploads once at consolidation and binds by value as a flat handle.",
"required": ["name", "kind", "dtype", "shape", "count"],
"additionalProperties": true,
"properties": {
"name": { "type": "string" },
"kind": { "const": "lookup" },
"dtype": { "const": "float" },
"shape": {
"type": "array",
"items": { "type": "integer" },
"description": "Buffer dimensions."
},
"count": {
"type": "integer",
"description": "Flat element count == prod(shape)."
}
}
},
"mutables": {
"type": "array",
"description": "Pure only: the writable per-sample state that IS the pure output.",
"items": { "$ref": "#/$defs/mutable" }
},
"mutable": {
"type": "object",
"description": "One writable per-sample slot. dtype/width let a loader rebuild which Mutables ride the (W, N) GRef ABI vs a flat scalar handle; the optional default fills an omitted Mutable. A matrix slot additionally carries its (R, C) shape.",
"required": ["name", "dtype", "width"],
"additionalProperties": true,
"properties": {
"name": { "type": "string" },
"dtype": {
"enum": ["float", "int", "vector", "matrix"],
"description": "Element kind of the mutable."
},
"width": {
"type": "integer",
"description": "Flat component count (1 for a scalar)."
},
"default": {
"type": ["number", "null"],
"description": "Scalar fill value for an omitted mutable (null when none)."
},
"shape": {
"type": "array",
"minItems": 2,
"maxItems": 2,
"items": { "type": "integer" },
"description": "The (R, C) shape — present ONLY when dtype == \"matrix\" (the flat width alone cannot be un-flattened)."
}
},
"if": {
"properties": { "dtype": { "const": "matrix" } },
"required": ["dtype"]
},
"then": {
"required": ["shape"]
}
},
"matShapes": {
"type": "object",
"description": "Pure only, and only when the kernel has matrix inputs: the (R, C) shape per matrix-input name.",
"additionalProperties": {
"type": "array",
"minItems": 2,
"maxItems": 2,
"items": { "type": "integer" }
}
},
"launch": {
"type": "object",
"description": "RESERVED in v1 (unpopulated; loaders ignore). Populated in a future minor revision without a schema_version bump. additionalProperties is true so unknown/future launch keys do not fail validation of the reserved surface.",
"additionalProperties": true,
"properties": {
"block": {
"type": "integer",
"description": "Reserved: thread-block size hint."
},
"shmem_bytes": {
"type": "integer",
"description": "Reserved: dynamic shared-memory bytes hint."
},
"grid": {
"type": "object",
"additionalProperties": true,
"description": "Reserved: freeform grid-configuration hint (shape TBD)."
}
}
}
}
}