API reference — expression templates#

Every leaf and composite node in aether derives from Expression; the individual node types (Sum, CWiseScale, Cross, QuatRotate, …) live in aether::detail and are reached through the operators documented in Expression templates, not cited individually here. aether::eval is the dual-mode runtime evaluator that walks a tree at assignment time.

aether::Expression#

template<class Derived, class T>
class Expression#

CRTP base for every expression-template leaf/node. A tag base (element_type alias only — no state) that anchors the aether_expression concept via std::derived_from, plus the geometric/reduction/quaternion/slicing member-function surface described in the file docstring above. Adding members here does not change any already-derived type’s ABI (still an empty base — every new member is either state-free or template-only).

Subclassed by aether::View< T, extents< dyn >, layout_right, false, true >, aether::View< T, Extents, Layout, Volatile, ReadOnly >, aether::detail::Sum< L, R, subtract >

Public Types

using element_type = T#

The STORAGE scalar type (matches Derived::element_type).

using working_type = working_type_t<T>#

The on-register carrier every node’s eval() returns. == element_type for every native dtype.

Public Functions

template<class E>
inline constexpr element_type dot(const Expression<E, typename E::element_type> &other) const#

Dot product with another rank-1 expression of the same shape.

inline constexpr element_type squaredNorm() const#

Squared L2 norm — sum of squares of components.

inline constexpr working_type squaredNormWorking_() const#

Squared L2 norm left IN THE WORKING CARRIER - the shared body the four norm accessors below fold on top of, so a norm() or a rCubedNorm() pays ONE encode at its own return rather than one per intermediate. Internal (trailing underscore); the public spelling is squaredNorm() directly above.

inline constexpr element_type norm() const#

L2 norm.

inline constexpr element_type rNorm() const#

Reciprocal of the L2 norm (accurate division, not a fast rsqrt approximation).

inline constexpr element_type rSquaredNorm() const#

Reciprocal of the squared L2 norm.

inline constexpr element_type cubedNorm() const#

Cubed L2 norm (‖v‖³).

inline constexpr element_type rCubedNorm() const#

Reciprocal of the cubed L2 norm (1/‖v‖³).

inline constexpr element_type maxNorm() const#

L-infinity (max-abs) norm.

inline constexpr element_type sum() const#

Sum of all components.

template<class E>
inline constexpr auto cross(const Expression<E, typename E::element_type> &other) const#

Cross product with another 3-vector expression; lazy (evaluate by assigning to a 3-shaped target).

inline constexpr auto unitVector() const#

Unit-vector (normalized) view of this expression; lazy.

template<class E>
inline constexpr auto quatMul(const Expression<E, typename E::element_type> &other) const#

Quaternion (Hamilton) product *this ⊗ other; lazy.

inline constexpr auto quatConj() const#

Quaternion conjugate (negate the vector part); lazy.

inline constexpr auto quatReciprocal() const#

Quaternion reciprocal (q⁻¹ = q* / ‖q‖²); lazy, reuses quatConj() and rSquaredNorm().

template<class E>
inline constexpr auto quatRotate(const Expression<E, typename E::element_type> &vec) const#

Rodrigues-formula rotation of 3-vector vec by this unit quaternion; lazy.

inline constexpr auto asPureQuaternion() const#

View this 3-vector as a pure quaternion [0,x,y,z]; lazy.

inline constexpr auto asBack3DVector() const#

View this quaternion’s vector part (drop the scalar, index 0); reuses tail<3>.

template<std::size_t Off, std::size_t Len>
inline constexpr auto segment() const#

Read-only view of Len consecutive components starting at Off.

template<std::size_t N>
inline constexpr auto head() const#

Read-only view of the first N components.

template<std::size_t N>
inline constexpr auto tail() const#

Read-only view of the last N components.

inline constexpr auto transpose() const#

Read-only transpose of this rank-2 (matrix) expression; lazy.

template<std::size_t R>
inline constexpr auto row() const#

Read-only view of row R of this rank-2 expression (a rank-1/vector expression); lazy.

template<std::size_t C>
inline constexpr auto col() const#

Read-only view of column C of this rank-2 expression (a rank-1/vector expression); lazy.

template<std::size_t R0, std::size_t C0, std::size_t BR, std::size_t BC>
inline constexpr auto block() const#

Read-only view of a BR x BC block starting at (R0, C0) of this rank-2 expression; lazy.

template<class E>
inline constexpr auto cwiseMul(const Expression<E, typename E::element_type> &other) const#

Hadamard (component-wise) product with another expression of the same element_extents; lazy.

template<class E>
inline constexpr element_type matDot(const Expression<E, typename E::element_type> &other) const#

Frobenius inner product sum_{r,c} this(r,c) * other(r,c) of two rank-2 expressions; evaluates eagerly (sample-free, SampleIndex::make(0), matching dot()’s convention).

inline constexpr element_type trace() const#

Trace (sum of the diagonal) of a square 2x2 or 3x3 rank-2 expression.

inline constexpr element_type det() const#

Determinant of a square 2x2 or 3x3 rank-2 expression (closed-form cofactor expansion).

inline constexpr auto inverse() const#

Closed-form inverse of a square 2x2 or 3x3 rank-2 expression; lazy — each component is an INDEPENDENT recompute of the source expression’s components plus the determinant (matches Cross’s own per-component recompute convention, aether/expr/nodes/Geometric.h’s docstring), so it is safe to assign into a batched destination (each eval<R,C>(i) uses the i it is actually called with, unlike trace()/det() above).

Namespace aether::eval#

namespace eval#

Enums

enum class RuntimeOp : unsigned char#

Runtime-evaluator op tags — see the file docstring.

Values:

enumerator Assign#
enumerator Add#
enumerator Sub#
enumerator Scale#
enumerator AddScaled#
enumerator SubScaled#

Functions

inline void runtimeEval(RuntimeOp op, const RuntimeView &a, const RuntimeView &b, const RuntimeView &out, double scalar = 0.0)#

Elementwise runtime evaluator entry point — host-callable. Validates a/b/out (shape + dtype) and throws aether::Error on mismatch BEFORE dispatching; the dispatch itself either runs on the host or launches detail::evalKernel on out.device.

b is read only by Add/Sub/AddScaled/SubScaled — pass a default-constructed RuntimeView{} for the unary ops (Assign/Scale). scalar is read only by Scale/AddScaled/SubScaled (ignored, but harmless, otherwise).

Assign is the ONE op allowed to see a.dtype != out.dtype (that is precisely what makes it also serve as a double<->float cast); every other op requires every operand it actually reads to share out’s dtype.

Throws:

aether::Error – on a shape mismatch (a/b vs out), an unsupported dtype (v1: only double/float), a dtype disagreement not covered by the Assign cast allowance, or (via detail::dispatch) a flat element count too large to address.

namespace detail#

Functions

template<class TIn, class TOut>
inline void evalElement(const RuntimeView &a, const RuntimeView &b, const RuntimeView &out, RuntimeOp op, TIn scalar, offset_t linear)#

One element of the generic strided elementwise loop — THE device-legal core both drivers below call. Decomposes linear (row-major over out.rank/out.extents) into a multi-index, addresses a/b/out through EACH OPERAND’S OWN strides, and applies op.

b’s pointer/strides are simply never DEREFERENCED for the ops that do not use it (Assign/Scale) — callers pass a default RuntimeView{} for b then (arithmetic on its strides is harmless: no read of b.data ever happens on those paths).

inline std::size_t flatSizeWide(const RuntimeView &v)#

Product of v.extents[0..v.rank) at FULL std::size_t width — mirrors extents::extentWide()’s “the guard’s input stays wide” convention (aether/index/Offset.h). Host-only.

inline bool sameShape(const RuntimeView &x, const RuntimeView &y)#
inline bool isFloat32(const DType &dt)#
inline bool isFloat64(const DType &dt)#
template<class TIn, class TOut>
void dispatch(RuntimeOp op, const RuntimeView &a, const RuntimeView &b, const RuntimeView &out, double scalar)#

Host entry point for one resolved <TIn, TOut> pair: validates the flat element count against the addressable cap, then either launches evalKernel (CUDA device targets) or runs the SAME evalElement core in a serial host loop.

Variables

template<class...>
constexpr bool device_compiler_v = (AETHER_DEVICE_COMPILER != 0)#

AETHER_DEVICE_COMPILER as a DEPENDENT constant.

static_assert(AETHER_DEVICE_COMPILER, ...) written directly in a template body is a NON-dependent condition: the compiler evaluates it while merely PARSING the template, so a host TU that only includes this header would be rejected. Routing the same constant through a variable template makes it dependent, so the diagnostic fires exactly when the launch path is INSTANTIATED — including aether from a host TU stays free.