Test map#
Reach for this page when you are about to touch a piece of eagle and want to
know what already tests it — “I’m changing how graph capture handles a
dependent node” or “does anything check the CPU plugin registry’s
shape-mismatch rejection?” — without reading through both tests/ and
python/tests/ end to end. eagle ships two independent, dual-language test
suites that mirror each other feature-for-feature (the C++ core and its
Python bindings), so this page is organized by capability, with both
languages’ coverage listed side by side wherever a feature has both.
How the suites are counted, and why that matters here#
Unlike a plain pytest/ctest invocation, the C++ suite’s CI gate
(tests/check_gate.sh, invoked as tests/check_gate.sh <mode> build/tests/eagle_tests from CI and make test) does not trust a raw pass/fail
exit code — it diffs the actual --gtest_list_tests output and the shuffled
run’s summary line against a committed name manifest
(tests/expected_tests_cuda.txt / tests/expected_tests_cpp.txt), so a
silently-skipped or silently-renamed test fails the gate even if every test
that did run passed. The two manifests currently commit 274 CUDA-mode
tests and 163 CPP_MODE (OpenMP host) tests — CPP_MODE is smaller because
several GPU-only suites (graph capture internals, device error handling)
have no CPU equivalent to run. tests/check_gate.sh’s own header explains the reasoning behind this
design; the short version is: a test runner that reports
“0 tests, exit 0” on a typo’d filter is not a defect in the runner, it is a
defect in whichever gate trusted a bare exit code, so the gate here pins the
whole named set instead. The Python suite (python/tests, run as pytest python/tests -q in CI) does not yet have an equivalent named-set
gate.
Two tests exist purely to check that other tests are testing the right
thing: python/tests/test_conformance_corpus.py and its C++ twin
tests/test_ConformanceCorpus.cpp both replay the same language-neutral
fixture corpus (tests/conformance/*.json), so a C++/Python behavioral
divergence shows up as a mechanical test failure rather than something a
human has to notice by comparing two docs pages.
Graph capture, lifetime, and ownership (C++, CUDA-only)#
Test file |
Suite(s) |
Capability |
|---|---|---|
|
|
Core |
|
|
A Launcher stays valid after the |
|
|
Ownership/ref-counting semantics of Graph/Launcher handles. |
|
|
Per-kernel |
|
|
|
|
|
Device-side native graph node behavior. |
|
|
|
|
|
The CPU parity twin of |
|
— |
|
|
— |
Bound-core |
|
— |
The interim |
Launch policy and block sizing (dual-mode)#
Block size is how many GPU threads execute together in one scheduling unit (a thread block); picking it well affects occupancy — how much of the GPU’s parallel capacity a kernel actually keeps busy.
Test file |
Suite(s) |
Capability |
|---|---|---|
|
|
|
|
|
|
|
— |
Gate A — the |
|
— |
Gate B — a GPU end-to-end value twin against a real captured graph. |
|
— |
Gate C — the nonvacuity ritual: proves the gate can actually fail. |
Device properties (dual-mode)#
Test file |
Suite(s) |
Capability |
|---|---|---|
|
|
The CUDA-free half: |
|
|
The hardware half: |
Conditional / branch execution (dual-mode)#
Test file |
Suite(s) |
Capability |
|---|---|---|
|
|
|
|
|
The CPU parity twin of the above, compiled into both build modes. |
|
— |
|
|
— |
|
|
— |
|
|
— |
|
|
— |
The host branch of |
|
— |
|
|
— |
|
|
— |
|
|
— |
The physical reorder of |
Device error handling (C++, CUDA-only)#
Test file |
Suite(s) |
Capability |
|---|---|---|
|
|
|
|
— |
Isolated-process helper for the one leg that needs a fresh process after CUDA init. |
|
|
The host arm of the DLPack layer ( |
|
|
The DLPack stream contract on a device: export fence and re-fence order a slow producer before the consumer; the RED twin without a fence reads the stale value; |
|
|
Every host-API call site that used to be guarded by the deleted, release-mode-silent |
Reductions, scans, and filtering (dual-mode)#
Test file |
Suite(s) |
Capability |
|---|---|---|
|
|
Device and CPU/OpenMP multi-level reductions. |
|
|
Scan and scatter-compaction, device and host, across small/large regimes. |
|
|
Host-graph chained scratch reuse across reduce/scan/filter nodes. |
|
|
Active-set compaction over raw buffers ( |
|
|
The physical reorder ( |
Plugin registry and host/device parity (dual-mode)#
The place where a plugin — a compiled kernel handed to eagle at run time, rather than one eagle compiled itself — gets loaded, bound, and launched, on both the CUDA and CPU registries.
Test file |
Suite(s) |
Capability |
|---|---|---|
|
|
CPU |
|
|
The device registry’s deliberate float32 rejection, and SoftDouble load-and-launch. |
|
|
Manifest-level duplicate-plugin-id rejection on the CUDA registry. |
|
|
|
|
|
The device registry’s additive int64 uniform path ( |
|
|
The CPU host registry’s twin of the int64 uniform path above — same binding vocabulary and refusal wording as the device registry. |
|
— |
The Python half of the int64 broadcast uniform. |
|
— |
An independent multi-plugin registry injecting a plugin set into one graph. |
|
— |
The registered pattern-loader hook (registry inversion). |
|
— |
The precompiled C++/CUDA plugin host loading and running an eagle PTX plugin. |
|
— |
A dlopen’d CPU plugin driven over torch CPU tensors. |
|
— |
A deployed (PTX) launchable as a capturable CUDA-graph node. |
wide_in / wide_out / accum_out roles (Python, additive schema-v1)#
The three arg-spec roles added after the schema-v1 freeze (see Plugin schema (manifest and sidecar)’s role-vocabulary table) — a named “wide” buffer input/output and the cross-sample accumulate plane, all three bound as a plain by-value handle.
Test file |
Capability |
|---|---|
|
|
|
The wide roles on the capturable launch path — |
|
The device half: |
The v2 execution axis: structures, placement, partitioning (dual-mode)#
Schema-v2’s exec_targets/exec_access/exec_op manifest keys and the Python
eagle.exec/eagle.plan API that drives them (see Plugin schema (manifest and sidecar)’s
execution-axis section) — single-device, host-team, and rank-partitioned runs of the
same body, and the placement/layout self-checks that refuse a mismatch at plan time.
Test file |
Suite(s) |
Capability |
|---|---|---|
|
|
The host/CUDA-free execution-contract rows: manifest parsing of the v2 keys, placement legality by |
|
|
The device rows: the PTX layout self-check, the device launch triple, partition identity under |
|
— |
RED-first coverage of eagle’s manifest load path under schema v2 — the new execution axis through the loader, as opposed to |
|
— |
eagle’s four execution-structure interop-certification rows: partition identity, host/device twin, illegal-placement refusal, layout-selfcheck refusal. |
|
— |
The single-process face of |
Schema, sidecar, and manifest validation (dual-mode, GPU-free)#
Test file |
Suite(s) |
Capability |
|---|---|---|
|
|
Forward-strict field pinning and unknown-key tolerance across gate/structural/discriminant classes. |
|
|
The optional |
|
|
The shared C++/Python conformance corpus driver (see above). |
|
— |
The Python-side twin of |
|
— |
The Python-side twin of |
|
— |
The shared corpus driver, Python side. |
|
— |
The |
|
— |
eagle’s schema constants re-point at raptor’s declared values. |
|
— |
The reserved-word absence test on eagle’s own vocabulary. |
ABI and interop (Python)#
Test file |
Capability |
|---|---|
|
The aether-ABI version stamp stays in sync between eagle and the C++ host. |
|
Compile-time gate: eagle’s |
|
Compile-time gate: eagle’s |
|
Framework-free DLPack detection and origin-selection logic. |
|
The generic buffer layer: zero-copy import of numpy, torch, cupy and Warp producers, access/owner/producer reporting, re-export, named refusals and the stream contract (fence, RED twin, re-fence). |
|
eagle’s own 8 declared rows of the cross-repo interop certification matrix. |
|
Completeness gate and no-skip self-test over the matrix above. |
|
The NVIDIA Warp interop certification rows ( |
|
The shared role-to-ABI-shape classifier. |
Python backend seam (eagle._core / CUDA plugin boundary)#
The eagle-backend/1 seam between the always-importable g++ core (eagle._core) and
the optional CUDA plugin (libeagle_cuda.so) — a missing, mismatched or partial
backend is a typed refusal, never a crash.
Test file |
Capability |
|---|---|
|
Every row runs in a fresh subprocess pointed at a plugin through |
|
Reads the built artifacts directly: |
|
The compiled extensions load on a machine with no NVIDIA driver present — only device execution needs one. |
|
The CUDA plugin reaches CUDA through the driver library alone — no undefined CUDA symbol, and no CUDA runtime linked in. |
|
Behaviour that crosses the seam on a real device: a Python callback the backend calls back into surfaces as its own exception; a handle consumed by another call refuses reuse. |
PyTorch bridge (Python)#
Test file |
Capability |
|---|---|
|
|
Golden conformance and the neural-block descriptor#
Test file |
Capability |
|---|---|
|
eagle’s validator accepts raptor’s own golden fixtures. |
|
The |
|
Road equality between raptor’s entry point and eagle’s local one for the same clause. |
Host dispatch internals (C++, CPU)#
Test file |
Suite(s) |
Capability |
|---|---|---|
|
|
|
|
|
|
Docs and utility#
Test file |
Capability |
|---|---|
|
Every doc-included code snippet actually compiles and matches its source. |
|
Re-renders every committed |
|
|
|
|
Assessment binaries (not gated, informative only)#
File |
What it reports |
|---|---|
|
Host-graph executor overhead vs. direct calls, arena flatness, dlopen’d-plugin vs. native OpenMP scaling. Prints a table; not part of |
|
A compile probe (no |