Foundational interoperability protocols#
This page tells you which array libraries can hand data to RAPTOR, and whether that copies.
The family is built in layers: raptor owns the contracts and depends on nothing; hawk authors and compiles kernels; eagle launches, captures and marshals them (marshals = converts an array into exactly the layout a compiled call expects); downstream packages wire kernels into trainable programs. Each works alone — adding a companion buys speed or a deployment option, never a new ability the family did not already have.
DLPack is the contract#
DLPack is a standard, framework-neutral way for one library to hand
another an array — CPU or GPU — without copying it. Every crossing on this
page is, mechanically, a DLPack exchange: cupy, torch and NVIDIA Warp
arrays all expose __dlpack__/__dlpack_device__, and the family reads or
writes through that protocol rather than through a framework-specific code
path. That is what lets the same kernel accept a numpy, cupy, torch or
Warp argument and route to the right provider without the caller naming
which one.
For a runnable walkthrough rather than a table, see the
zero-copy interop example: a pointer-stable
host round-trip through numpy, exercised for real, plus a worked read of
what the cupy/torch device rows below certify. A CUDA-capable torch
build exchanges device tensors the identical way, through the same
__dlpack__ protocol cupy uses — nothing about the crossing changes
between the two frameworks on the device side.
Roadmap
jax and tensorflow/keras are both DLPack-capable and both foreseen —
see Certified frameworks below. Neither is
certified today. Treat any claim that one of them works
now as a documentation bug, not a feature: this page cites a row for every
capability it claims, and no row names either framework. NVIDIA Warp is no
longer in this category: see the certified rows below.
The rule this document obeys#
Every normative claim below either cites an executable certification row, or is explicitly labelled ROADMAP or UNCERTIFIED. There are no uncited capability claims.
The rows are declared once, in raptor.conformance.interop, and each is
implemented as a real test in the repository on the other side of the
crossing it certifies. That declaration — not this page — is the authority. A doc-sync test asserts
that the row identifiers on this page and the declared set are identical in both
directions, so the two cannot drift; if they ever disagree, the declaration
wins. Read this document as the human-readable face of a machine-checked matrix,
never as a substitute for it.
Dispatch is by input type#
A compiled kernel is called directly, and the framework of the arguments
selects the execution path: a numpy argument routes to the native host
provider, a cupy argument routes to the nvcc + eagle device path. The
caller writes one call and gets host or device execution according to what
it passed.
The host path is genuinely CUDA-free: no nvcc (the CUDA compiler), no
cupy, no eagle anywhere on it. It works by compiling the kernel once,
caching the compiled result by a hash of its source, dlopen-ing that
compiled library (loading it at runtime the way an OS loads any shared
library), and marshalling the call’s arguments (converting them into the
exact layout the compiled function expects) before calling in.
The cuda-free half of that claim is certified by running a subprocess with
nvcc removed from PATH and with cupy and eagle blocked at import,
compiling and executing a kernel there, and asserting inside that subprocess
that neither module was ever imported. The provider lookup is a precedence
chain — the ordered list dispatch checks before picking a provider — so
further providers can be admitted without changing the call surface; no row
certifies any provider beyond these two, and none is claimed.
Backends: the execution substrate#
A wired training program — a training loop assembled once from a backend and then run many times, as opposed to one rebuilt from scratch on every call — binds to a backend by name.
Three backends exist:
eagle— compiled kernels and captured graphs (a GPU’s sequence of launches recorded once, then replayed without re-issuing each one); the performance design point.reference— the reference wiring engine: a plain, uncompiled implementation of the same wiring, run over the host JIT (the CPU kernel path that compiles once and caches the result, with no GPU orcupyinvolved), used as the correctness baseline the other backends are checked against.torch— described next.
TorchBackend is composition, not interpretation. It routes on the device of the input tensors: CPU tensors go to the reference wiring engine over the host JIT (which is genuinely cuda-free), and CUDA tensors go to the eagle compiled path.
Any other torch device raises CapabilityError naming the device and saying
ROADMAP — this covers MPS (Apple’s GPU backend), ROCm (AMD’s GPU compute
stack) and XLA (Google’s whole-program compiler). There is deliberately no
silent .cpu() fallback: falling back would report success for a device
the family cannot actually accelerate. T-DEVICE-FAILLOUD certifies that
behaviour, asserting both the exception type and that its message names the
offending device.
The CPU input crossing is certified zero-copy by T-BACKEND-CPU-IN-ALIAS and
the CUDA input crossing by T-BACKEND-CUDA-IN-ALIAS.
Forward and reverse-mode run as one program-level autograd step — the
backward pass flows through the backend boundary exactly as it would through
plain torch ops, so a caller’s loss.backward() works unchanged on either
leg. The end-to-end eagle-less training loop — forward, backward, and a
parameter update on CPU tensors with eagle, cupy and nvcc all unavailable —
is gated by a named test in the downstream package’s own suite
(test_torch_backend_training.py::test_eagle_less_torch_backend_training_cuda_free)
rather than by a matrix row; it is executable and it runs, but it is a suite gate,
not a row-level certificate, and this page does not pretend otherwise.
Forward-mode differentiation (JVP) through the backend protocol is ROADMAP. No row certifies it.
Certified frameworks, and what “roadmap” means mechanically#
The certified set is numpy, cupy, torch and NVIDIA Warp. jax and tensorflow/keras are ROADMAP, absent by declaration. The family’s stated horizon includes them; no row certifies them and none pretends to.
“Absent by declaration” is enforced in both directions, and neither direction is ever a skip:
a certified framework missing from the gate environment is a failure, raised by
certified_framework— the matrix must never skip its way to green;a roadmap framework that turns out to be installed is also a failure, raised by
roadmap_framework— because the moment it is present, the claim that nothing exercises it stops being checkable.
A skipped test returns success to the runner. A conformance matrix that skips a framework and reports green certifies nothing, so the matrix rows are built to be structurally incapable of skipping: a scanner discovers every file containing a registered row and rejects skip machinery — skip tokens and availability markers alike — on the registered functions themselves.
Reading the matrix, in miniature#
Before the full 23-row table below, here is how to read any one row. These three all share one crossing class, ALIAS — pointer identity plus write-through: the two sides report the same address, and a write on one is visible on the other:
row |
framework |
owner |
class |
certifies |
|---|---|---|---|---|
|
torch |
eagle |
ALIAS |
a CUDA tensor becomes a zero-copy device view |
|
warp |
eagle |
ALIAS |
a |
|
torch |
downstream |
ALIAS |
the CPU backend’s input crossing is a zero-copy view |
Four more classes appear in the full matrix: COPY (pointers differ, and mutating either side afterwards leaves the other unchanged), STREAM (certifies ordering, not placement), ENV (certifies a property of the environment — what is absent, or what is refused), and INVARIANCE (the same eagle execution body agrees across different execution structures — no framework on either side, because there is no crossing). The full certification matrix below defines each in full and lists every one of the 23 rows; the acronyms (ALIAS/COPY/STREAM/ENV/INVARIANCE) are each defined again, in depth, in The two transfer laws further down this page.
The certification matrix#
Twenty-three rows. framework is what the row bridges to; owner is the
repository whose suite carries the assertion; class is how data crosses, in
the vocabulary defined in the next section.
Eight rows bridge to no framework — their framework column reads —.
Four of those are the execution_* rows, each an eagle-internal invariant of
the heterogeneous execution contract (the rule that one plugin body must
give equivalent results under every eagle execution structure, not only the
one it happened to be written against), not a crossing between the family
and an array framework.
The other four are the HAWK-* rows: rows that exist only because hawk
exists — a round-trip row and two reference-comparison rows (below).
Every row on this page is declared before the code it certifies is
built: the execution_* rows are declared here — like CP-CAPTURE-REJECT
and every other eagle-owned row above them — ahead of the eagle-side work
(exec/, eagle.plan, the loader) that will implement and collect them, and
the HAWK-* rows the same way, owner=hawk, ahead of hawk’s own consumer
implementation. That is the family’s declare-before-build convention: a
row may describe a capability that does not exist yet.
HostTeam, DeviceKernel and RankPartition below are eagle’s names for
its three execution structures: a team of CPU threads, a single GPU, and a
multi-process partition, respectively — the same body can run under any of
them.
row |
framework |
owner |
class |
certifies |
|---|---|---|---|---|
|
cupy |
eagle |
ENV |
graph capture rejects a non-device input instead of silently allocating during capture |
|
torch |
eagle |
ALIAS |
a CUDA tensor becomes a zero-copy device view: pointer identity, device-ordinal agreement, and write-through |
|
torch |
eagle |
ALIAS |
a caller-provided CUDA output tensor is filled in place, and the engine’s result exports back pointer-equal and write-through |
|
torch |
eagle |
COPY |
a CPU tensor round-trips as a torch tensor via a copy, write-isolated in both directions |
|
torch |
downstream |
COPY |
the export bridge (the path that hands a result array out as a tensor) copies before handing out a tensor: the export never aliases the live replay buffer (the buffer a captured graph reuses across replays), and a later forward cannot mutate it |
|
torch |
downstream |
ALIAS |
the CPU backend’s input crossing is a zero-copy, write-through view of the tensor’s storage |
|
torch |
downstream |
ALIAS |
the CUDA backend’s forward inputs and reverse-mode cotangents alias the device engine’s view |
|
torch |
downstream |
ENV |
an unsupported torch device raises |
|
torch |
eagle |
STREAM |
work enqueued on a non-default torch stream before a launch is seen by that launch |
|
cupy |
eagle |
STREAM |
a device write enqueued before a graph replay, with no synchronisation, is seen by the replay — on a caller stream and on the legacy null stream (CUDA’s original default stream, used when no stream is named) alike |
|
torch |
eagle |
STREAM |
a torch operation consuming the result immediately, with no synchronisation, reads the finished value |
|
torch |
eagle |
STREAM |
inside the launch context, the device library’s current stream is torch’s current stream |
|
warp |
eagle |
ALIAS |
a |
|
warp |
eagle |
ALIAS |
the reverse crossing, |
|
warp |
eagle |
STREAM |
a device write enqueued on a dedicated |
|
— |
eagle |
INVARIANCE |
the same |
|
— |
eagle |
INVARIANCE |
the same body’s |
|
— |
eagle |
ENV |
a plan placing a body on a structure its declared |
|
— |
eagle |
ENV |
a loader refuses a plugin whose ABI tag agrees but whose exported layout sizes disagree (a self-check), and refuses an outright tag mismatch |
|
— |
hawk |
INVARIANCE |
a two-wire compound quantity (one value carried as two linked slots) binds distinct slots and both wires’ values round-trip on the host path — the positive twin of the matching refusal check |
|
— |
hawk |
ENV |
a HAWK artifact loads under |
|
— |
hawk |
INVARIANCE |
every fixture kernel under |
|
— |
hawk |
INVARIANCE |
the |
The two transfer laws#
Every crossing between a framework and the family is declared, and the declaration is asserted positively rather than inferred:
ALIAS means pointer identity plus write-through. The row asserts that the two sides report the same address and, where the crossing is bidirectional, that a write on one side is visible on the other with nothing in between.
COPY means distinctness plus write-isolation. The row asserts that the pointers differ and that mutating either side afterwards leaves the other unchanged. A pointer check alone would pass on an accidental copy; the isolation half is what makes the row a contract.
Three further classes appear in the table. STREAM rows certify ordering
rather than placement. ENV rows certify a property of the environment —
what is absent, or what is refused. INVARIANCE rows certify that the same
plugin body produces equivalent results under different eagle execution
structures — bit-exact for sample_local/cross_sample_read bodies, or
within a fixed tolerance band otherwise; unlike ALIAS/COPY, an INVARIANCE row
has no framework on either side of the crossing, because there is no
crossing — both sides are eagle.
The output direction through the export bridge is always a COPY. The export
path copies unconditionally before wrapping a result as a tensor, whether the
underlying array is a scratch buffer or a live replay view. This is deliberate
and it is a feature: a program’s buffers are reused across replays, so an export
that aliased them would be silently corrupted by the next forward pass. That
contract is certified by T-BRIDGE-COPY, which asserts the semantic half — a
previously exported tensor survives a later forward with different inputs
unchanged — and not merely the pointer half.
One consequence is worth stating rather than leaving as a surprise: no backend row can be an alias in the output direction, on either leg. The two backend alias rows name their direction in their own identifiers for exactly this reason: the claim only ever covered the input crossing, never the output one.
The stream contract#
Stream correctness is certified behaviourally. Rows STREAM-TORCH-PRODUCER-ORDER,
STREAM-CUPY-PRODUCER-ORDER, STREAM-TORCH-CONSUMER-ORDER and STREAM-IDENTITY
manufacture an adversarial schedule — a long occupying kernel, then a data write
queued behind it, then the launch, with no synchronisation anywhere — so that a
launch which failed to ride the caller’s stream would read stale data with
certainty rather than by luck. Ordering is then asserted on the result. The
ordering guarantee itself lives in the graph pipeline’s launch path, which
records and waits on events so that a replay observes work already enqueued on
the caller’s stream and on the legacy null stream alike.
__cuda_stream__ is ROADMAP. Nothing in the pinned stack produces or
consumes it; eagle exchanges streams as a raw handle instead. Certifying the
presence of a protocol attribute that nothing calls would have certified nothing,
so the family certifies the observable ordering instead.
The full rules (stream codes, event fencing, ownership and access flags, and the named refusals) are in eagle’s interoperability contract.
Manifest key order: strict writer, tolerant reader#
Top-level key order in a kernel manifest is contractual for writers. The Python validator enforces it, and the golden artifacts (committed reference manifests, used as fixtures across the family’s tests) carry it byte-stably, so a manifest produced through the family is reproducible.
The C++ reader is deliberately order-agnostic, and that asymmetry is Postel’s law applied on purpose, not drift (Postel’s law: be strict in what you write, lenient in what you read). Tightening the reader would buy no capability: a wrong-order manifest can only come from a producer that bypassed validation, and the bypass is the actual defect. The tolerance is pinned by a test that feeds the reader a manifest with shuffled top-level keys and asserts it parses, so any future tightening has to be a conscious change of contract rather than an accident.
Schema v2: the execution axis#
schema_version 1 and 2 both load today — a transition bridge, not a
replacement. A v1 manifest keeps aether_abi: "aether-abi/1" and must not carry
any exec_* key; it plans as legacy — whole-view, single-device, exactly
today’s behaviour. A v2 manifest carries aether_abi: "aether-abi/2" plus three
new flat top-level keys, inserted between aether_abi and plugins in the
writer’s key order:
exec_targets— a non-empty list drawn fromdevice/host, no duplicates: which execution structures may run the body.exec_access— one ofsample_local,cross_sample_read,cross_sample_write,mapreduce: how the body touches sample-local vs. cross-sample data, declared by the emitter (hawk’s code-generation stage, which writes the manifest) from the trace (hawk’s recorded representation of the kernel), never guessed.exec_op— one ofsum,times,max,land: required whenexec_accessismapreduce, forbidden otherwise.
Absence of a required key is a load refused, not a default guessed. The keys
are flat rather than one nested execution block because the manifest is a
strict-key, order-contractual document a C++ text scanner reads — flat keys are
the shape this schema already uses everywhere else.
What each install gets you#
The degradation law — dropping an optional dependency may cost speed or a deployment option, but never the ability to run the same program — governs backends. A kernel provider is different — it is the program’s content, not its substrate, so provider presence sits outside the law.
install |
what runs |
|---|---|
|
Validate kernel manifests ( |
|
Author and compile a kernel for a host and/or device target, and publish its manifest. Still does not run it. |
|
Load the compiled artifact and run it on CPU threads ( |
|
Capture and replay the kernel as a CUDA graph on the GPU. |
|
The |
Torch devices other than CPU and CUDA are ROADMAP on every row of that table.
The torch CPU and CUDA crossings are certified by T-BACKEND-CPU-IN-ALIAS and T-BACKEND-CUDA-IN-ALIAS; the refusal
to fake the rest by T-DEVICE-FAILLOUD.
Going deeper (optional)#
Historical context, an edge-case hazard, and the completeness gate’s own fine print — none of it needed to read the matrix above, all of it real.
Row ownership, in transition#
The numpy/cupy/torch crossing rows hawk will eventually own stay under their current owner (above) until that migration happens, then are re-owned in one move — nothing duplicated here for them in the meantime.
A reverted design, stated plainly#
The kernel-provider protocol has exactly one facet today: a compiled kernel described by a schema-versioned manifest.
An earlier design added a second, symbolic facet — an evaluable trace that any array namespace could interpret, which would have given the code generator a standalone interpreter and downstream packages a device-agnostic torch backend for free. It was reverted on measured evidence: compiling and dlopening a kernel ran measurably faster than interpreting it on numpy, even for a trivial kernel, so the interpreter would have been a slow path carrying a second semantics to keep in sync. What ships instead is the native host provider described above, and its orphaned protocol declaration was deleted rather than left standing as a promise nothing implemented.
That mechanism may return, scoped to what only it can do: an interpretive basis for torch devices other than CPU and CUDA. That is ROADMAP.
A sidecar nesting hazard#
A separate and unrelated hazard lives in the sidecar reader (the sidecar is the small per-kernel JSON file describing how to call it): a neural-block descriptor embeds nested execution references whose keys share names with top-level keys, so a nested object appearing before its parent’s own key could be read as the parent’s value. That is a nesting hazard, not a key-order contract; it is handled by bounding and masking the nested objects before the top-level scan, and the conformance corpus is written in the adversarial order so the case is executed rather than reasoned about.
What the completeness gate does not certify#
Each owning repository asserts that every row declared for it has a registered test. Read the scope of that guarantee carefully, because it is narrower than it looks.
Registration happens when the test module is imported, not when the test is executed. The completeness gate therefore certifies row presence, not row execution. Two consequences follow, and both are load-bearing:
Explicitly deselecting the matrix by marker in a CI job is safe — the rows still register, so the gate does not go falsely red — and it is visible in the job definition rather than hidden in a runtime skip. That is the only sanctioned way a row may fail to run.
Because deselection is safe, at least one job must run the matrix without deselecting it, or every job is green and nothing is certified. Any pipeline that adopts this matrix owes that job.
One further constraint: the collected-row registry is per-session, so a partial selection of tests will fail the completeness assertion for the right reason but at the wrong time. Suites containing the completeness assertion must therefore be run whole, in a single process, and not distributed across worker processes.
Everything else on this page is a row, and every row is a test.