TileLang command cheatsheet: compile, analyze, debug

Configuration & UsagePublished 2026-10-01Author: DeepSeek Plugin Market
TileLangtilelang commandscompile_onlyTILELANG_PASS_DIFFAnalyzerenvironment variablesAscend NPU
TileLang command cheatsheet: compile_only without a GPU, Analyzer.analysis for FLOPs, plot_layout and autodd for debugging, and the ~/.tilelang/cache cache.

TileLang commands fall into four families: compile and build, performance analysis, pass debugging, and build-time environment variables. Two matter most: python -I -m tilelang.tools.compile_only --target cuda --output_file out.cu example.py, which compiles a kernel without a GPU, and tilelang.tools.Analyzer.analysis(tir, device), which estimates cost statically. Compiled artifacts live in ~/.tilelang/cache.

This is a reference page, organized as compile and build → performance and debug → environment variables and cache, with the command, its purpose and its expected output for each entry. Getting the environment set up is in Install TileLang.

What are TileLang's compile and build commands?

TileLang's compile commands split in two: compile_only, which turns a kernel script into backend source (python -I -m tilelang.tools.compile_only --target cuda --output_file out.cu example.py), and the pip or script commands that build TileLang itself from source (source).

  1. Compile a kernel offline → compile_only. Run python -I -m tilelang.tools.compile_only --target cuda --output_file out.cu example.py. Expect generated CUDA source in out.cu; -I isolates environment variables, and --target must be a concrete backend such as cuda or rocm because auto is rejected (source).
  2. Make failures stop a pipeline. The same command exits with code 1 when compilation fails. Expect a script or CI job to gate on that exit code, so the check drops straight into automation.
  3. Build TileLang itself (NVIDIA). In the source tree set USE_CUDA=ON as needed, then produce a wheel with pip or python -m build -w. Expect a wheel in dist/ whose build switches match your toolchain (source).
  4. Build TileLang itself (Ascend). Run ./build_wheel_ascend.sh [--enable-llvm], or bash install_ascend.sh to install in place. Expect a wheel in dist/ in the first case and a completed in-place install in the second (source).

How do I use TileLang's performance and debug commands?

TileLang offers Analyzer for cost estimates without running a kernel, plus several development-time debug entries: plot_layout for layout visualization, python -m tilelang.autodd for automatic parallelization, TL_LOWER_TRACE for lowering traces and TILELANG_PASS_DIFF for pass diffs (source).

  1. Estimate FLOPs and traffic → Analyzer.analysis. Call tilelang.tools.Analyzer.analysis(tir, device). Expect FLOPs, global memory traffic and roofline time for that kernel, enough to tell whether it is compute- or bandwidth-bound (source).
  2. See layouts and timing → plot_layout. Call tilelang.tools.plot_layout. Expect a view of the tiling and memory hierarchy that confirms blocks land where you intended.
  3. Tune automatic parallelization → python -m tilelang.autodd. Run the module. Expect printed reasoning behind automatic parallelization decisions, useful when a schedule is not what you expected (source).
  4. Watch the whole lowering → TL_LOWER_TRACE. Set that variable and recompile. Expect step-by-step TIR lowering logs; to diff a single pass, use TILELANG_PASS_DIFF instead.

These are development-time tools: Analyzer sets direction before tuning, while plot_layout, autodd and TL_LOWER_TRACE locate compile-time problems. None of them belongs in a production inference path.

TileLang build-time environment variables and the compile cache

TileLang's build-time variables split by route: on NVIDIA, USE_CUDA enables the CUDA backend, USE_ROCM and USE_METAL cover the others, USE_LLVM and TVM_ROOT set the compiler stack and TVM source, and WITH_PIP_CUDA_TOOLCHAIN plus NO_VERSION_LABEL control toolchain source and version labels; on Ascend, ASCEND_HOME_PATH locates Ascend libraries. Artifacts cache in ~/.tilelang/cache (source).

Variable / pathPurposeRoute
USE_CUDAEnable the CUDA backendNVIDIA
USE_ROCMEnable the ROCm backendAMD
USE_METALEnable the Metal backendApple
USE_LLVMEnable the LLVM backendGeneral / Ascend builds
TVM_ROOTPoint at the TVM sourceSource builds
WITH_PIP_CUDA_TOOLCHAINUse the pip-provided CUDA toolchainNVIDIA
NO_VERSION_LABELDisable the wheel version labelPackaging
ASCEND_HOME_PATHPoint at the Ascend toolkit rootAscend
~/.tilelang/cacheCompiled artifact cacheBoth routes
  1. Rebuild after changing settings. Build-time variables are read at compile time. Expect nothing to change until you rebuild; exporting alone leaves the installed artifacts as they were.
  2. Delete the cache directory, not just files. Run rm -rf ~/.tilelang/cache. Expect the next compile to regenerate artifacts; this is the first step when a stale result survives a code change or upgrade.
  3. Let CANN supply the Ascend variable. ASCEND_HOME_PATH is normally set by set_env.sh. Expect echo $ASCEND_HOME_PATH to return a value, without which the compiler cannot find Ascend libraries.

TileLang command cheatsheet

GoalCommandExpected output
Compile a kernel offlinepython -I -m tilelang.tools.compile_only --target cuda --output_file out.cu example.pyout.cu written; exit code 1 on failure
Estimate FLOPs / traffictilelang.tools.Analyzer.analysis(tir, device)FLOPs, global traffic, roofline time
Inspect layouttilelang.tools.plot_layoutTiling and memory hierarchy view
Debug auto-parallelizationpython -m tilelang.autoddParallelization decisions
Lowering traceTL_LOWER_TRACE=1 (environment variable)Step-by-step TIR lowering logs
Pass diffTILELANG_PASS_DIFF=1 (environment variable)Pass difference output
Upgradepip install -U tilelangNewer release installed
Install nightlypip install tilelang -f https://tile-ai.github.io/whl/nightlyPreview build of the day installed
Uninstallpip uninstall tilelangPackage removed (use tilelang-ascend on Ascend)

TileLang command notes

  1. --target does not accept auto. compile_only wants an explicit backend such as cuda or rocm.
  2. Keep debug variables to debugging. TL_LOWER_TRACE and TILELANG_PASS_DIFF emit heavy logs, so unset them once you are done.
  3. Manage versions and cache together. Before upgrading, read Update TileLang, which covers version pinning and when to clear ~/.tilelang/cache.
  4. Do not uninstall dependencies by accident. The boundaries of pip uninstall and leftover cleanup are in Uninstall TileLang.

Sources: TileLang Tools, TileLang compile_only tool, TileLang Analyzer, TileLang Installation Guide

FAQ

Can TileLang compile a kernel without a GPU? How does compile_only work?

TileLang can compile without a GPU present: run python -I -m tilelang.tools.compile_only --target cuda --output_file out.cu example.py and it writes the generated CUDA source without executing anything. The target must be a concrete backend such as cuda or rocm, since auto is rejected, and a failed compile exits with code 1 so CI can gate on it.

How do I estimate the FLOPs and memory traffic of a TileLang kernel?

TileLang's performance entry point is tilelang.tools.Analyzer.analysis(tir, device), which statically estimates FLOPs, global memory traffic and roofline time without running the kernel, so you can tell whether you are compute-bound or bandwidth-bound before tuning. Pass the target device as device, and it works on both Ascend and NVIDIA.

What should I watch when debugging TileLang compilation, and what are autodd and TL_LOWER_TRACE?

TileLang debugging works in two layers: tilelang.tools.plot_layout draws the layout view and python -m tilelang.autodd reports automatic-parallelization decisions, while TL_LOWER_TRACE turns on per-step TIR lowering logs and TILELANG_PASS_DIFF shows the difference a pass makes. The first two explain why a schedule looks wrong; the last two trace the compile pipeline.

Where does TileLang keep its compile cache, and how do I make recompiles fast?

TileLang caches compiled artifacts in ~/.tilelang/cache by default, so compiling the same kernel again reuses them and skips the rebuild. When a code change still hits a stale artifact, delete that directory to force a rebuild; clearing it after a version upgrade also prevents old and new artifacts from mixing.

Which environment variables matter when building TileLang, and how do USE_CUDA and ASCEND_HOME_PATH differ?

TileLang build-time variables split by route: on NVIDIA, USE_CUDA enables the CUDA backend, USE_ROCM and USE_METAL cover ROCm and Metal, TVM_ROOT points at the TVM source, and WITH_PIP_CUDA_TOOLCHAIN plus NO_VERSION_LABEL control toolchain source and version labels; on Ascend, CANN's set_env.sh sets ASCEND_HOME_PATH so the compiler can find Ascend libraries.

Related Terms

compile_only
compile_only is TileLang's offline compile tool (python -m tilelang.tools.compile_only) that turns a TileLang script into target-backend source written to a file, requiring no GPU, and is typically used in CI and code review.— TileLang compile_only tool
Analyzer
Analyzer is TileLang's static performance interface: tilelang.tools.Analyzer.analysis(tir, device) estimates FLOPs, global memory traffic and roofline time without running the kernel.— TileLang Analyzer
TL_LOWER_TRACE
TL_LOWER_TRACE is a TileLang debug environment variable that, when enabled, prints trace logs for each TIR transformation during lowering, helping locate compile-time problems.— TileLang Tools
compile cache
TileLang's compile cache lives in ~/.tilelang/cache and stores compiled kernel artifacts; recompiling the same kernel reuses them, and it should be cleared after code changes or version upgrades.— TileLang Installation Guide

Sources