1========================= 2Compiling CUDA with clang 3========================= 4 5.. contents:: 6 :local: 7 8Introduction 9============ 10 11This document describes how to compile CUDA code with clang, and gives some 12details about LLVM and clang's CUDA implementations. 13 14This document assumes a basic familiarity with CUDA. Information about CUDA 15programming can be found in the 16`CUDA programming guide 17<http://docs.nvidia.com/cuda/cuda-c-programming-guide/index.html>`_. 18 19Compiling CUDA Code 20=================== 21 22Prerequisites 23------------- 24 25CUDA is supported in llvm 3.9, but it's still in active development, so we 26recommend you `compile clang/LLVM from HEAD 27<http://llvm.org/docs/GettingStarted.html>`_. 28 29Before you build CUDA code, you'll need to have installed the appropriate 30driver for your nvidia GPU and the CUDA SDK. See `NVIDIA's CUDA installation 31guide <https://docs.nvidia.com/cuda/cuda-installation-guide-linux/index.html>`_ 32for details. Note that clang `does not support 33<https://llvm.org/bugs/show_bug.cgi?id=26966>`_ the CUDA toolkit as installed 34by many Linux package managers; you probably need to install nvidia's package. 35 36You will need CUDA 7.0 or 7.5 to compile with clang. CUDA 8 support is in the 37works. 38 39Invoking clang 40-------------- 41 42Invoking clang for CUDA compilation works similarly to compiling regular C++. 43You just need to be aware of a few additional flags. 44 45You can use `this <https://gist.github.com/855e277884eb6b388cd2f00d956c2fd4>`_ 46program as a toy example. Save it as ``axpy.cu``. (Clang detects that you're 47compiling CUDA code by noticing that your filename ends with ``.cu``. 48Alternatively, you can pass ``-x cuda``.) 49 50To build and run, run the following commands, filling in the parts in angle 51brackets as described below: 52 53.. code-block:: console 54 55 $ clang++ axpy.cu -o axpy --cuda-gpu-arch=<GPU arch> \ 56 -L<CUDA install path>/<lib64 or lib> \ 57 -lcudart_static -ldl -lrt -pthread 58 $ ./axpy 59 y[0] = 2 60 y[1] = 4 61 y[2] = 6 62 y[3] = 8 63 64* ``<CUDA install path>`` -- the directory where you installed CUDA SDK. 65 Typically, ``/usr/local/cuda``. 66 67 Pass e.g. ``-L/usr/local/cuda/lib64`` if compiling in 64-bit mode; otherwise, 68 pass e.g. ``-L/usr/local/cuda/lib``. (In CUDA, the device code and host code 69 always have the same pointer widths, so if you're compiling 64-bit code for 70 the host, you're also compiling 64-bit code for the device.) 71 72* ``<GPU arch>`` -- the `compute capability 73 <https://developer.nvidia.com/cuda-gpus>`_ of your GPU. For example, if you 74 want to run your program on a GPU with compute capability of 3.5, specify 75 ``--cuda-gpu-arch=sm_35``. 76 77 Note: You cannot pass ``compute_XX`` as an argument to ``--cuda-gpu-arch``; 78 only ``sm_XX`` is currently supported. However, clang always includes PTX in 79 its binaries, so e.g. a binary compiled with ``--cuda-gpu-arch=sm_30`` would be 80 forwards-compatible with e.g. ``sm_35`` GPUs. 81 82 You can pass ``--cuda-gpu-arch`` multiple times to compile for multiple archs. 83 84The `-L` and `-l` flags only need to be passed when linking. When compiling, 85you may also need to pass ``--cuda-path=/path/to/cuda`` if you didn't install 86the CUDA SDK into ``/usr/local/cuda``, ``/usr/local/cuda-7.0``, or 87``/usr/local/cuda-7.5``. 88 89Flags that control numerical code 90--------------------------------- 91 92If you're using GPUs, you probably care about making numerical code run fast. 93GPU hardware allows for more control over numerical operations than most CPUs, 94but this results in more compiler options for you to juggle. 95 96Flags you may wish to tweak include: 97 98* ``-ffp-contract={on,off,fast}`` (defaults to ``fast`` on host and device when 99 compiling CUDA) Controls whether the compiler emits fused multiply-add 100 operations. 101 102 * ``off``: never emit fma operations, and prevent ptxas from fusing multiply 103 and add instructions. 104 * ``on``: fuse multiplies and adds within a single statement, but never 105 across statements (C11 semantics). Prevent ptxas from fusing other 106 multiplies and adds. 107 * ``fast``: fuse multiplies and adds wherever profitable, even across 108 statements. Doesn't prevent ptxas from fusing additional multiplies and 109 adds. 110 111 Fused multiply-add instructions can be much faster than the unfused 112 equivalents, but because the intermediate result in an fma is not rounded, 113 this flag can affect numerical code. 114 115* ``-fcuda-flush-denormals-to-zero`` (default: off) When this is enabled, 116 floating point operations may flush `denormal 117 <https://en.wikipedia.org/wiki/Denormal_number>`_ inputs and/or outputs to 0. 118 Operations on denormal numbers are often much slower than the same operations 119 on normal numbers. 120 121* ``-fcuda-approx-transcendentals`` (default: off) When this is enabled, the 122 compiler may emit calls to faster, approximate versions of transcendental 123 functions, instead of using the slower, fully IEEE-compliant versions. For 124 example, this flag allows clang to emit the ptx ``sin.approx.f32`` 125 instruction. 126 127 This is implied by ``-ffast-math``. 128 129Detecting clang vs NVCC from code 130================================= 131 132Although clang's CUDA implementation is largely compatible with NVCC's, you may 133still want to detect when you're compiling CUDA code specifically with clang. 134 135This is tricky, because NVCC may invoke clang as part of its own compilation 136process! For example, NVCC uses the host compiler's preprocessor when 137compiling for device code, and that host compiler may in fact be clang. 138 139When clang is actually compiling CUDA code -- rather than being used as a 140subtool of NVCC's -- it defines the ``__CUDA__`` macro. ``__CUDA_ARCH__`` is 141defined only in device mode (but will be defined if NVCC is using clang as a 142preprocessor). So you can use the following incantations to detect clang CUDA 143compilation, in host and device modes: 144 145.. code-block:: c++ 146 147 #if defined(__clang__) && defined(__CUDA__) && !defined(__CUDA_ARCH__) 148 // clang compiling CUDA code, host mode. 149 #endif 150 151 #if defined(__clang__) && defined(__CUDA__) && defined(__CUDA_ARCH__) 152 // clang compiling CUDA code, device mode. 153 #endif 154 155Both clang and nvcc define ``__CUDACC__`` during CUDA compilation. You can 156detect NVCC specifically by looking for ``__NVCC__``. 157 158Optimizations 159============= 160 161Modern CPUs and GPUs are architecturally quite different, so code that's fast 162on a CPU isn't necessarily fast on a GPU. We've made a number of changes to 163LLVM to make it generate good GPU code. Among these changes are: 164 165* `Straight-line scalar optimizations <https://goo.gl/4Rb9As>`_ -- These 166 reduce redundancy within straight-line code. 167 168* `Aggressive speculative execution 169 <http://llvm.org/docs/doxygen/html/SpeculativeExecution_8cpp_source.html>`_ 170 -- This is mainly for promoting straight-line scalar optimizations, which are 171 most effective on code along dominator paths. 172 173* `Memory space inference 174 <http://llvm.org/doxygen/NVPTXInferAddressSpaces_8cpp_source.html>`_ -- 175 In PTX, we can operate on pointers that are in a paricular "address space" 176 (global, shared, constant, or local), or we can operate on pointers in the 177 "generic" address space, which can point to anything. Operations in a 178 non-generic address space are faster, but pointers in CUDA are not explicitly 179 annotated with their address space, so it's up to LLVM to infer it where 180 possible. 181 182* `Bypassing 64-bit divides 183 <http://llvm.org/docs/doxygen/html/BypassSlowDivision_8cpp_source.html>`_ -- 184 This was an existing optimization that we enabled for the PTX backend. 185 186 64-bit integer divides are much slower than 32-bit ones on NVIDIA GPUs. 187 Many of the 64-bit divides in our benchmarks have a divisor and dividend 188 which fit in 32-bits at runtime. This optimization provides a fast path for 189 this common case. 190 191* Aggressive loop unrooling and function inlining -- Loop unrolling and 192 function inlining need to be more aggressive for GPUs than for CPUs because 193 control flow transfer in GPU is more expensive. More aggressive unrolling and 194 inlining also promote other optimizations, such as constant propagation and 195 SROA, which sometimes speed up code by over 10x. 196 197 (Programmers can force unrolling and inline using clang's `loop unrolling pragmas 198 <http://clang.llvm.org/docs/AttributeReference.html#pragma-unroll-pragma-nounroll>`_ 199 and ``__attribute__((always_inline))``.) 200 201Publication 202=========== 203 204The team at Google published a paper in CGO 2016 detailing the optimizations 205they'd made to clang/LLVM. Note that "gpucc" is no longer a meaningful name: 206The relevant tools are now just vanilla clang/LLVM. 207 208| `gpucc: An Open-Source GPGPU Compiler <http://dl.acm.org/citation.cfm?id=2854041>`_ 209| Jingyue Wu, Artem Belevich, Eli Bendersky, Mark Heffernan, Chris Leary, Jacques Pienaar, Bjarke Roune, Rob Springer, Xuetian Weng, Robert Hundt 210| *Proceedings of the 2016 International Symposium on Code Generation and Optimization (CGO 2016)* 211| 212| `Slides from the CGO talk <http://wujingyue.com/docs/gpucc-talk.pdf>`_ 213| 214| `Tutorial given at CGO <http://wujingyue.com/docs/gpucc-tutorial.pdf>`_ 215 216Obtaining Help 217============== 218 219To obtain help on LLVM in general and its CUDA support, see `the LLVM 220community <http://llvm.org/docs/#mailing-lists>`_. 221