Device math#
Call one
aether::mathfunction the same way from host code or a GPU kernel – same name, same source line, either way.
Time: ~6 min · Runs on: CPU, GPU switch · You need: One array, two machines
running on: cpu
aether::math wraps the host and device math library behind one name
per function – the same call compiles unchanged whether this file
builds as host C++ or as a .cu. sincos is one of them:
print(run_cpp(r"""
double s = 0.0, c = 0.0;
aether::math::sincos(0.5, &s, &c);
std::printf("sin=%.6f cos=%.6f\n", s, c);
"""))
sin=0.479426 cos=0.877583
Same call, same source line, on a host build or a device one.
No dedicated “CDF” function exists – the standard normal CDF
(cumulative distribution function) is just erfc composed with ordinary
arithmetic:
print(run_cpp(r"""
auto phi = [](double x) { return 0.5 * aether::math::erfc(-x / std::sqrt(2.0)); };
std::printf("phi(0.0) = %.6f\n", phi(0.0));
""", includes=("<cmath>",)))
phi(0.0) = 0.500000
0.5 * erfc(-x / sqrt(2)) is the whole function – no bespoke
“statistics” module anywhere in the library.
What you will build (the picture)#
The curve the cell above just computed, point by point:
The exact same call compiles unchanged with nvcc on a GPU machine –
same name, same source line, either way.
What just happened#
Every
aether::math::*name dispatches by argument type (floatvsdouble), and in CUDA mode by compile pass too (device intrinsic vsstd::) – same call, same source line, either way.sincoswrites two outputs through pointers; the standard normal CDF is just0.5 * erfc(-x / sqrt(2)), ordinary arithmetic over one library call.hypot,expand the rest ofaether::mathall follow the same one-call, host-or-device pattern.
Try this#
Change phi(0.0) in the cell above to phi(1.0) (or any other x) and
re-run – compare the printed value against the curve below.
Next#
Your own device function and kernel –
write a function once and call it from a CPU loop or a GPU kernel.
Deeper: C++ API reference for the full
aether::math surface.
Going deeper (optional): two honest differences from NumPy#
aether::math::sign and round/rint are documented to depart from
NumPy on purpose: sign(NaN) is 0 here (NumPy’s np.sign(nan) is
nan), and aether::math::round is C’s round-half-away-from-zero
(NumPy’s np.round is rint, round-half-to-even). aether’s own test
suite checks host against device: exact functions match bit for bit, but
transcendentals – functions like erf/sincos computed by a series
or algorithm, not one hardware instruction – only have to land within
each platform’s own documented accuracy, not identical digits between
host and device.
print(run_cpp(r"""
std::printf("sign(nan) = %.1f round(2.5) = %.1f rint(2.5) = %.1f\n",
aether::math::sign(std::numeric_limits<double>::quiet_NaN()),
aether::math::round(2.5), aether::math::rint(2.5));
""", includes=("<limits>",)))
sign(nan) = 0.0 round(2.5) = 3.0 rint(2.5) = 2.0