Heuristics#
Heuristics rank the Operators returned by
get_operators() by estimated performance. Discovery
first filters candidates for correctness; a heuristic then orders and may
prune those candidates. A candidate’s rank is its position in the returned
list. Ranking is an estimate, not a guarantee of the fastest Operator.
For a worked GEMM example, see the heuristics tutorial.
Selecting a heuristic#
Pass a Heuristic instance as heuristic= to
get_operators(). For example, given
GemmArguments named args:
import cutlass.operators as ops
from cutlass.operators.heuristics import NvMatmulHeuristics
heuristic = NvMatmulHeuristics(gpu="B200")
# Equivalent construction through the registry:
heuristic = ops.get_heuristic("nvmatmul")(gpu="B200")
operators = ops.get_operators(
args, target_sm="100a", heuristic=heuristic, limit=5
)
get_heuristic() returns a class; instantiate it before
passing it to get_operators. limit must be positive when provided.
It is forwarded to the heuristic as a hint, and get_operators truncates
the ranked result afterward. Pruning may leave fewer than limit results.
Omitting heuristic preserves discovery order. Ranking errors propagate
to the caller.
Built-in nvMatmul heuristic#
NvMatmulHeuristics supports SM100 non-blockscaled dense GEMM. Its gpu
argument selects the modeled GPU SKU: "B200" (the default),
"GB200_NVL", or "GB300_NVL". Other values raise ValueError.
The model is selected explicitly, without detecting the current GPU.
Only Operators designed for the modeled compute capability and matching a
recommended configuration are returned. Other generations and unmatched
Operators are excluded, including kernels designed for older generations
that can run on the target. target_sm filters discovery for compatibility;
it does not change the heuristic’s GPU model.
With non-empty candidates, unsupported arguments or layouts, missing dependencies, and failure to match any recommendation raise errors. Empty candidates return an empty list.
Install the optional dependency with
pip install 'nvidia-cutlass-operators[heuristics]'. Registration does not
imply that this dependency is available: use
is_available() to check package
compatibility. This check does not validate a particular GEMM or guarantee
that a candidate will match.
- class cutlass.operators.heuristics.NvMatmulHeuristics(
- gpu: Literal['B200', 'GB200_NVL', 'GB300_NVL'] = 'B200',
Bases:
HeuristicRank Operators using NVIDIA’s nvMatmulHeuristics analytical model.
Currently supports only SM100 non-blockscaled dense GEMM (
GemmArguments).Queries nvMMH for recommended configs, ranks Operators that match the recommended configs, and prunes away Operators that do not match any recommended config or that were designed for a different compute capability than the GPU selected by
gpu.Missing optional dependencies raise
ImportError. Non-GEMMargsraiseTypeError.Select the nvMMH device this heuristic instance models.
- Parameters:
gpu (Literal["B200", "GB200_NVL", "GB300_NVL"]) – An
NvMatmulHeuristicsNvidiaGpumember name. Currently, this must be a Blackwell GPU SKU, namely"B200","GB200_NVL", or"GB300_NVL".- Raises:
ValueError – If
gpuisn’t a device nvmatmul supports ranking for (only SM100 devices are supported today).
- rank(
- args: RuntimeArguments | None,
- operators: list[Operator],
- *,
- target_sm: TargetSm | str | None = None,
- limit: int | None = None,
Order
operatorsbest-first forargsusing nvMatmulHeuristics.See
cutlass.operators.Heuristic.rank().- Parameters:
args (RuntimeArguments | None) – The problem to rank for. Must be
GemmArgumentswhenoperatorsis non-empty.operators (list[Operator]) – Filtered candidate Operators.
target_sm (TargetSm | str | None) – Accepted for
rank()compatibility; unused.limit (int | None) – When set, the caller will keep only the first
limitOperators from the ranked result.
- Returns:
The subset of
operatorsmatching a recommended config for this instance’s modeled hardware (seeself.modeled_cc), best-first. Empty only ifoperatorswas empty.- Return type:
list[Operator]
- Raises:
ImportError – If
nvMatmulHeuristicsis not installed. The caller asked for this heuristic explicitly, so a missing install is not silently ignored.TypeError – If
operatorsis non-empty andargsis not aGemmArguments, or if operands are not dense (DenseTensor).ValueError – If
operatorsis non-empty and no candidate matches any heuristic config.
- cutlass.operators.heuristics.nvmatmul.is_available() bool#
Return whether a compatible
nvidia-matmul-heuristicsinstall exists.Mirrors the availability signalling of
cutlass.operators.available_providers: the heuristic is always registered, but this reports whether it can actually run in this environment (importable, at leastMIN_NVMMH_VERSION, and compatible with the 0.1.0.27 constructor API shape).- Returns:
Trueif a call torank()can use nvMMH,Falseif it would raiseImportError.- Return type:
bool
- cutlass.operators.heuristics.nvmatmul.MIN_NVMMH_VERSION: str#
Minimum supported version of the optional
nvidia-matmul-heuristicspackage. Availability also requires a compatible package API.
Custom heuristics and registry#
Subclass Heuristic and implement rank. Return
an ordered subset of the input Operator objects; pruning is allowed.
get_operators raises RuntimeError if ranking introduces an Operator
or duplicates one beyond its count in the input. limit is a hint to the
ranker; get_operators enforces the final result limit.
Pass an instance directly, or register the class for lookup with
get_heuristic(). Lookup returns the class itself and
raises KeyError for an unknown name.
The symbols below are also exported by cutlass.operators.heuristics.
The nvMatmul _mapping and _provider modules are implementation details;
their query, configuration, and matching helpers are not public APIs.
- class cutlass.operators.Heuristic#
Orders already-filtered Operators best-first for a problem.
A concrete heuristic implements
rank(), which ranks the given Operators for the given problem by estimated performance. The returned Operators are an ordered subset of the candidates – a heuristic may prune some suboptimal candidates, but must not invent new Operators.Turning a heuristic on or off changes candidate order and possibly which candidates are present, never their correctness.
Heuristics classes are addressable by name through a small registry (
cutlass.operators.available_heuristics).- abstract rank(
- args: RuntimeArguments | None,
- operators: list[Operator],
- *,
- target_sm: TargetSm | str | None = None,
- limit: int | None = None,
Return
operatorsreordered best-first forargs.- Parameters:
args (RuntimeArguments | None) – Runtime arguments describing the problem being ranked for (e.g.
GemmArguments). May beNoneif the caller did not provide arguments.operators (list[Operator]) – The already-filtered candidate Operators to order. Every element is known to support
args.target_sm (TargetSm | str | None) – Optional compute capability the operators are being ranked for.
limit (int | None) – When set, the caller will keep only the first
limitOperators from the ranked result.
- Returns:
A subset of
operators(possibly all of them), reordered best-first. Must not contain an Operator absent fromoperators, or duplicate one beyond its count there.- Return type:
list[Operator]
- cutlass.operators.register_heuristic(
- name: str,
Return a decorator that registers a
Heuristicsubclass undername.Mirrors
cutlass.operators.providers.register_provider().
- cutlass.operators.get_heuristic(
- name: str,
Return the registered
Heuristicsubclass forname.- Parameters:
name (str) – The name the heuristic was registered under.
- Returns:
The registered heuristic class.
- Return type:
type[Heuristic]
- Raises:
KeyError – If no heuristic is registered under
name.