Compile options#

Two inputs control every hawk compile beyond the kernel itself: how hard the compiler optimises (opt_level), and which GPU machine code is built for (device_arch). Both are resolved by one small function each, both feed the compile cache’s own key, and both are plain keyword arguments to hawk.build/hawk.artifact.build.

opt_level: how hard the compiler optimises#

opt_level picks one of "O0"–"O3" (hawk.compile.OPT_LEVELS); the default is "O3". Resolution order: the compile call’s own opt_level= > $HAWK_OPT_LEVEL > "O3" — hawk.compile.opt_level() is the resolver every build door calls:

from hawk.compile import OPT_LEVELS, opt_level

print("levels:", OPT_LEVELS)
print("default:", opt_level())
print("explicit O0:", opt_level("O0"))
levels: ('O0', 'O1', 'O2', 'O3')
default: O3
explicit O0: O0

It feeds -O<n> for the host build, and BOTH -O<n> (nvcc’s own code generation) and -Xptxas -O<n> (the device assembler) for an AOT nvcc device build. NVRTC — the pip-only device compiler hawk falls back to when no nvcc is on $PATH — always optimises and takes no -O option at all, so opt_level has no effect there.

A different opt_level is a different compile#

opt_level is part of the compile cache’s key, so a build at "O0" never gets served the "O3" entry, or the other way round — both are fresh compiles, never a stale hit of each other:

import pathlib, tempfile

import hawk
from hawk import Mutable, Param, Scalar

@hawk.kernel
def scale(x: Scalar, a: Param, y: Mutable[Scalar]):
    y = a * x

cache_dir = tempfile.mkdtemp()
default = hawk.build([scale], pathlib.Path(tempfile.mkdtemp()),
                     targets=("host",), cache_dir=cache_dir)
o0 = hawk.build([scale], pathlib.Path(tempfile.mkdtemp()),
               targets=("host",), cache_dir=cache_dir, opt_level="O0")
print("default (O3) cache hit:", default.artifacts[0].entries["host"].hit)
print("O0           cache hit:", o0.artifacts[0].entries["host"].hit)
default (O3) cache hit: False
O0           cache hit: False

Both print False: each opt_level got its own cache slot, compiled once, rather than one serving the other’s (wrong) binary.

device_arch: which GPU code is built for#

A device build targets one CUDA arch ("sm_80", …), resolved by hawk.artifact.arch in this order: an explicit device_arch= > $HAWK_CUDA_ARCH > the GPU actually in this machine (probed once per process) > "sm_61" (Pascal, the oldest arch hawk targets):

import os

from hawk.artifact import arch

os.environ.pop("HAWK_CUDA_ARCH", None)
print("no override anywhere (falls to sm_61 off a GPU):", arch())
print("an explicit device_arch= always wins:", arch("sm_90"))

os.environ["HAWK_CUDA_ARCH"] = "sm_80"
print("$HAWK_CUDA_ARCH, no explicit device_arch=:", arch())
del os.environ["HAWK_CUDA_ARCH"]
no override anywhere (falls to sm_61 off a GPU): sm_61
an explicit device_arch= always wins: sm_90
$HAWK_CUDA_ARCH, no explicit device_arch=: sm_80

device_arch is part of the compile cache’s key too, but only when "cuda" is actually one of the build’s targets — resolving it may reach the device probe above, which a host-only build must never pay for.

Performance numbers in these docs#

Every performance number elsewhere in hawk’s and eagle’s docs is measured at opt_level="O3", the default — "O0"/"O1"/"O2" are for a smaller or faster-to-compile debug build, never for comparing wall time against them.