TileLang command cheatsheet: compile, analyze, debug
TileLang commands fall into four families: compile and build, performance analysis, pass debugging, and build-time environment variables. Two matter most: python -I -m tilelang.tools.compile_only --target cuda --output_file out.cu example.py, which compiles a kernel without a GPU, and tilelang.tools.Analyzer.analysis(tir, device), which estimates cost statically. Compiled artifacts live in ~/.tilelang/cache.
This is a reference page, organized as compile and build → performance and debug → environment variables and cache, with the command, its purpose and its expected output for each entry. Getting the environment set up is in Install TileLang.
What are TileLang's compile and build commands?
TileLang's compile commands split in two: compile_only, which turns a kernel script into backend source (python -I -m tilelang.tools.compile_only --target cuda --output_file out.cu example.py), and the pip or script commands that build TileLang itself from source (source).
- Compile a kernel offline → compile_only. Run
python -I -m tilelang.tools.compile_only --target cuda --output_file out.cu example.py. Expect generated CUDA source inout.cu;-Iisolates environment variables, and--targetmust be a concrete backend such ascudaorrocmbecauseautois rejected (source). - Make failures stop a pipeline. The same command exits with code 1 when compilation fails. Expect a script or CI job to gate on that exit code, so the check drops straight into automation.
- Build TileLang itself (NVIDIA). In the source tree set
USE_CUDA=ONas needed, then produce a wheel with pip orpython -m build -w. Expect a wheel indist/whose build switches match your toolchain (source). - Build TileLang itself (Ascend). Run
./build_wheel_ascend.sh [--enable-llvm], orbash install_ascend.shto install in place. Expect a wheel indist/in the first case and a completed in-place install in the second (source).
How do I use TileLang's performance and debug commands?
TileLang offers Analyzer for cost estimates without running a kernel, plus several development-time debug entries: plot_layout for layout visualization, python -m tilelang.autodd for automatic parallelization, TL_LOWER_TRACE for lowering traces and TILELANG_PASS_DIFF for pass diffs (source).
- Estimate FLOPs and traffic → Analyzer.analysis. Call
tilelang.tools.Analyzer.analysis(tir, device). Expect FLOPs, global memory traffic and roofline time for that kernel, enough to tell whether it is compute- or bandwidth-bound (source). - See layouts and timing → plot_layout. Call
tilelang.tools.plot_layout. Expect a view of the tiling and memory hierarchy that confirms blocks land where you intended. - Tune automatic parallelization → python -m tilelang.autodd. Run the module. Expect printed reasoning behind automatic parallelization decisions, useful when a schedule is not what you expected (source).
- Watch the whole lowering → TL_LOWER_TRACE. Set that variable and recompile. Expect step-by-step TIR lowering logs; to diff a single pass, use
TILELANG_PASS_DIFFinstead.
These are development-time tools: Analyzer sets direction before tuning, while plot_layout, autodd and TL_LOWER_TRACE locate compile-time problems. None of them belongs in a production inference path.
TileLang build-time environment variables and the compile cache
TileLang's build-time variables split by route: on NVIDIA, USE_CUDA enables the CUDA backend, USE_ROCM and USE_METAL cover the others, USE_LLVM and TVM_ROOT set the compiler stack and TVM source, and WITH_PIP_CUDA_TOOLCHAIN plus NO_VERSION_LABEL control toolchain source and version labels; on Ascend, ASCEND_HOME_PATH locates Ascend libraries. Artifacts cache in ~/.tilelang/cache (source).
| Variable / path | Purpose | Route |
|---|---|---|
USE_CUDA | Enable the CUDA backend | NVIDIA |
USE_ROCM | Enable the ROCm backend | AMD |
USE_METAL | Enable the Metal backend | Apple |
USE_LLVM | Enable the LLVM backend | General / Ascend builds |
TVM_ROOT | Point at the TVM source | Source builds |
WITH_PIP_CUDA_TOOLCHAIN | Use the pip-provided CUDA toolchain | NVIDIA |
NO_VERSION_LABEL | Disable the wheel version label | Packaging |
ASCEND_HOME_PATH | Point at the Ascend toolkit root | Ascend |
~/.tilelang/cache | Compiled artifact cache | Both routes |
- Rebuild after changing settings. Build-time variables are read at compile time. Expect nothing to change until you rebuild; exporting alone leaves the installed artifacts as they were.
- Delete the cache directory, not just files. Run
rm -rf ~/.tilelang/cache. Expect the next compile to regenerate artifacts; this is the first step when a stale result survives a code change or upgrade. - Let CANN supply the Ascend variable.
ASCEND_HOME_PATHis normally set byset_env.sh. Expectecho $ASCEND_HOME_PATHto return a value, without which the compiler cannot find Ascend libraries.
TileLang command cheatsheet
| Goal | Command | Expected output |
|---|---|---|
| Compile a kernel offline | python -I -m tilelang.tools.compile_only --target cuda --output_file out.cu example.py | out.cu written; exit code 1 on failure |
| Estimate FLOPs / traffic | tilelang.tools.Analyzer.analysis(tir, device) | FLOPs, global traffic, roofline time |
| Inspect layout | tilelang.tools.plot_layout | Tiling and memory hierarchy view |
| Debug auto-parallelization | python -m tilelang.autodd | Parallelization decisions |
| Lowering trace | TL_LOWER_TRACE=1 (environment variable) | Step-by-step TIR lowering logs |
| Pass diff | TILELANG_PASS_DIFF=1 (environment variable) | Pass difference output |
| Upgrade | pip install -U tilelang | Newer release installed |
| Install nightly | pip install tilelang -f https://tile-ai.github.io/whl/nightly | Preview build of the day installed |
| Uninstall | pip uninstall tilelang | Package removed (use tilelang-ascend on Ascend) |
TileLang command notes
--targetdoes not acceptauto. compile_only wants an explicit backend such ascudaorrocm.- Keep debug variables to debugging.
TL_LOWER_TRACEandTILELANG_PASS_DIFFemit heavy logs, so unset them once you are done. - Manage versions and cache together. Before upgrading, read Update TileLang, which covers version pinning and when to clear
~/.tilelang/cache. - Do not uninstall dependencies by accident. The boundaries of
pip uninstalland leftover cleanup are in Uninstall TileLang.
Sources: TileLang Tools, TileLang compile_only tool, TileLang Analyzer, TileLang Installation Guide
FAQ
TileLang can compile without a GPU present: run python -I -m tilelang.tools.compile_only --target cuda --output_file out.cu example.py and it writes the generated CUDA source without executing anything. The target must be a concrete backend such as cuda or rocm, since auto is rejected, and a failed compile exits with code 1 so CI can gate on it.
TileLang's performance entry point is tilelang.tools.Analyzer.analysis(tir, device), which statically estimates FLOPs, global memory traffic and roofline time without running the kernel, so you can tell whether you are compute-bound or bandwidth-bound before tuning. Pass the target device as device, and it works on both Ascend and NVIDIA.
TileLang debugging works in two layers: tilelang.tools.plot_layout draws the layout view and python -m tilelang.autodd reports automatic-parallelization decisions, while TL_LOWER_TRACE turns on per-step TIR lowering logs and TILELANG_PASS_DIFF shows the difference a pass makes. The first two explain why a schedule looks wrong; the last two trace the compile pipeline.
TileLang caches compiled artifacts in ~/.tilelang/cache by default, so compiling the same kernel again reuses them and skips the rebuild. When a code change still hits a stale artifact, delete that directory to force a rebuild; clearing it after a version upgrade also prevents old and new artifacts from mixing.
TileLang build-time variables split by route: on NVIDIA, USE_CUDA enables the CUDA backend, USE_ROCM and USE_METAL cover ROCm and Metal, TVM_ROOT points at the TVM source, and WITH_PIP_CUDA_TOOLCHAIN plus NO_VERSION_LABEL control toolchain source and version labels; on Ascend, CANN's set_env.sh sets ASCEND_HOME_PATH so the compiler can find Ascend libraries.
Related Terms
- compile_only
- compile_only is TileLang's offline compile tool (python -m tilelang.tools.compile_only) that turns a TileLang script into target-backend source written to a file, requiring no GPU, and is typically used in CI and code review.— TileLang compile_only tool
- Analyzer
- Analyzer is TileLang's static performance interface: tilelang.tools.Analyzer.analysis(tir, device) estimates FLOPs, global memory traffic and roofline time without running the kernel.— TileLang Analyzer
- TL_LOWER_TRACE
- TL_LOWER_TRACE is a TileLang debug environment variable that, when enabled, prints trace logs for each TIR transformation during lowering, helping locate compile-time problems.— TileLang Tools
- compile cache
- TileLang's compile cache lives in ~/.tilelang/cache and stores compiled kernel artifacts; recompiling the same kernel reuses them, and it should be cleared after code changes or version upgrades.— TileLang Installation Guide
Sources
- TileLang Tools· TileLang (tile-ai)
- TileLang compile_only tool· TileLang (tile-ai)
- TileLang Analyzer· TileLang (tile-ai)
- TileLang Installation Guide· TileLang (tile-ai)