Compile options#
Two inputs control every hawk compile beyond the kernel itself: how hard
the compiler optimises (opt_level), and which GPU machine code is
built for (device_arch). Both are resolved by one small function
each, both feed the compile cache’s own key, and both
are plain keyword arguments to hawk.build/hawk.artifact.build.
opt_level: how hard the compiler optimises#
opt_level picks one of "O0"–"O3"
(hawk.compile.OPT_LEVELS); the default is "O3". Resolution order:
the compile call’s own opt_level= > $HAWK_OPT_LEVEL > "O3" —
hawk.compile.opt_level() is the resolver every build door calls:
from hawk.compile import OPT_LEVELS, opt_level
print("levels:", OPT_LEVELS)
print("default:", opt_level())
print("explicit O0:", opt_level("O0"))
levels: ('O0', 'O1', 'O2', 'O3')
default: O3
explicit O0: O0
It feeds -O<n> for the host build, and BOTH -O<n> (nvcc’s own
code generation) and -Xptxas -O<n> (the device assembler) for an AOT
nvcc device build. NVRTC — the pip-only device compiler hawk falls
back to when no nvcc is on $PATH — always optimises and takes no
-O option at all, so opt_level has no effect there.
A different opt_level is a different compile#
opt_level is part of the compile cache’s key, so a build at "O0"
never gets served the "O3" entry, or the other way round — both are
fresh compiles, never a stale hit of each other:
import pathlib, tempfile
import hawk
from hawk import Mutable, Param, Scalar
@hawk.kernel
def scale(x: Scalar, a: Param, y: Mutable[Scalar]):
y = a * x
cache_dir = tempfile.mkdtemp()
default = hawk.build([scale], pathlib.Path(tempfile.mkdtemp()),
targets=("host",), cache_dir=cache_dir)
o0 = hawk.build([scale], pathlib.Path(tempfile.mkdtemp()),
targets=("host",), cache_dir=cache_dir, opt_level="O0")
print("default (O3) cache hit:", default.artifacts[0].entries["host"].hit)
print("O0 cache hit:", o0.artifacts[0].entries["host"].hit)
default (O3) cache hit: False
O0 cache hit: False
Both print False: each opt_level got its own cache slot, compiled
once, rather than one serving the other’s (wrong) binary.
device_arch: which GPU code is built for#
A device build targets one CUDA arch ("sm_80", …), resolved by
hawk.artifact.arch in this order: an explicit device_arch= >
$HAWK_CUDA_ARCH > the GPU actually in this machine (probed once per
process) > "sm_61" (Pascal, the oldest arch hawk targets):
import os
from hawk.artifact import arch
os.environ.pop("HAWK_CUDA_ARCH", None)
print("no override anywhere (falls to sm_61 off a GPU):", arch())
print("an explicit device_arch= always wins:", arch("sm_90"))
os.environ["HAWK_CUDA_ARCH"] = "sm_80"
print("$HAWK_CUDA_ARCH, no explicit device_arch=:", arch())
del os.environ["HAWK_CUDA_ARCH"]
no override anywhere (falls to sm_61 off a GPU): sm_61
an explicit device_arch= always wins: sm_90
$HAWK_CUDA_ARCH, no explicit device_arch=: sm_80
device_arch is part of the compile cache’s key too, but only when
"cuda" is actually one of the build’s targets — resolving it may reach
the device probe above, which a host-only build must never pay for.
Performance numbers in these docs#
Every performance number elsewhere in hawk’s and eagle’s docs is
measured at opt_level="O3", the default — "O0"/"O1"/"O2" are for
a smaller or faster-to-compile debug build, never for comparing wall
time against them.