1============================= 2User Guide for AMDGPU Backend 3============================= 4 5.. contents:: 6 :local: 7 8.. toctree:: 9 :hidden: 10 11 AMDGPU/AMDGPUAsmGFX7 12 AMDGPU/AMDGPUAsmGFX8 13 AMDGPU/AMDGPUAsmGFX9 14 AMDGPU/AMDGPUAsmGFX900 15 AMDGPU/AMDGPUAsmGFX904 16 AMDGPU/AMDGPUAsmGFX906 17 AMDGPU/AMDGPUAsmGFX908 18 AMDGPU/AMDGPUAsmGFX10 19 AMDGPU/AMDGPUAsmGFX1011 20 AMDGPUModifierSyntax 21 AMDGPUOperandSyntax 22 AMDGPUInstructionSyntax 23 AMDGPUInstructionNotation 24 AMDGPUDwarfExtensionsForHeterogeneousDebugging 25 26Introduction 27============ 28 29The AMDGPU backend provides ISA code generation for AMD GPUs, starting with the 30R600 family up until the current GCN families. It lives in the 31``llvm/lib/Target/AMDGPU`` directory. 32 33LLVM 34==== 35 36.. _amdgpu-target-triples: 37 38Target Triples 39-------------- 40 41Use the ``clang -target <Architecture>-<Vendor>-<OS>-<Environment>`` option to 42specify the target triple: 43 44 .. table:: AMDGPU Architectures 45 :name: amdgpu-architecture-table 46 47 ============ ============================================================== 48 Architecture Description 49 ============ ============================================================== 50 ``r600`` AMD GPUs HD2XXX-HD6XXX for graphics and compute shaders. 51 ``amdgcn`` AMD GPUs GCN GFX6 onwards for graphics and compute shaders. 52 ============ ============================================================== 53 54 .. table:: AMDGPU Vendors 55 :name: amdgpu-vendor-table 56 57 ============ ============================================================== 58 Vendor Description 59 ============ ============================================================== 60 ``amd`` Can be used for all AMD GPU usage. 61 ``mesa3d`` Can be used if the OS is ``mesa3d``. 62 ============ ============================================================== 63 64 .. table:: AMDGPU Operating Systems 65 :name: amdgpu-os-table 66 67 ============== ============================================================ 68 OS Description 69 ============== ============================================================ 70 *<empty>* Defaults to the *unknown* OS. 71 ``amdhsa`` Compute kernels executed on HSA [HSA]_ compatible runtimes 72 such as AMD's ROCm [AMD-ROCm]_. 73 ``amdpal`` Graphic shaders and compute kernels executed on AMD PAL 74 runtime. 75 ``mesa3d`` Graphic shaders and compute kernels executed on Mesa 3D 76 runtime. 77 ============== ============================================================ 78 79 .. table:: AMDGPU Environments 80 :name: amdgpu-environment-table 81 82 ============ ============================================================== 83 Environment Description 84 ============ ============================================================== 85 *<empty>* Default. 86 ============ ============================================================== 87 88.. _amdgpu-processors: 89 90Processors 91---------- 92 93Use the ``clang -mcpu <Processor>`` option to specify the AMDGPU processor. The 94names from both the *Processor* and *Alternative Processor* can be used. 95 96 .. table:: AMDGPU Processors 97 :name: amdgpu-processor-table 98 99 =========== =============== ============ ===== ================= ======= ====================== 100 Processor Alternative Target dGPU/ Target ROCm Example 101 Processor Triple APU Features Support Products 102 Architecture Supported 103 [Default] 104 =========== =============== ============ ===== ================= ======= ====================== 105 **Radeon HD 2000/3000 Series (R600)** [AMD-RADEON-HD-2000-3000]_ 106 ----------------------------------------------------------------------------------------------- 107 ``r600`` ``r600`` dGPU 108 ``r630`` ``r600`` dGPU 109 ``rs880`` ``r600`` dGPU 110 ``rv670`` ``r600`` dGPU 111 **Radeon HD 4000 Series (R700)** [AMD-RADEON-HD-4000]_ 112 ----------------------------------------------------------------------------------------------- 113 ``rv710`` ``r600`` dGPU 114 ``rv730`` ``r600`` dGPU 115 ``rv770`` ``r600`` dGPU 116 **Radeon HD 5000 Series (Evergreen)** [AMD-RADEON-HD-5000]_ 117 ----------------------------------------------------------------------------------------------- 118 ``cedar`` ``r600`` dGPU 119 ``cypress`` ``r600`` dGPU 120 ``juniper`` ``r600`` dGPU 121 ``redwood`` ``r600`` dGPU 122 ``sumo`` ``r600`` dGPU 123 **Radeon HD 6000 Series (Northern Islands)** [AMD-RADEON-HD-6000]_ 124 ----------------------------------------------------------------------------------------------- 125 ``barts`` ``r600`` dGPU 126 ``caicos`` ``r600`` dGPU 127 ``cayman`` ``r600`` dGPU 128 ``turks`` ``r600`` dGPU 129 **GCN GFX6 (Southern Islands (SI))** [AMD-GCN-GFX6]_ 130 ----------------------------------------------------------------------------------------------- 131 ``gfx600`` - ``tahiti`` ``amdgcn`` dGPU 132 ``gfx601`` - ``hainan`` ``amdgcn`` dGPU 133 - ``oland`` 134 - ``pitcairn`` 135 - ``verde`` 136 **GCN GFX7 (Sea Islands (CI))** [AMD-GCN-GFX7]_ 137 ----------------------------------------------------------------------------------------------- 138 ``gfx700`` - ``kaveri`` ``amdgcn`` APU - A6-7000 139 - A6 Pro-7050B 140 - A8-7100 141 - A8 Pro-7150B 142 - A10-7300 143 - A10 Pro-7350B 144 - FX-7500 145 - A8-7200P 146 - A10-7400P 147 - FX-7600P 148 ``gfx701`` - ``hawaii`` ``amdgcn`` dGPU ROCm - FirePro W8100 149 - FirePro W9100 150 - FirePro S9150 151 - FirePro S9170 152 ``gfx702`` ``amdgcn`` dGPU ROCm - Radeon R9 290 153 - Radeon R9 290x 154 - Radeon R390 155 - Radeon R390x 156 ``gfx703`` - ``kabini`` ``amdgcn`` APU - E1-2100 157 - ``mullins`` - E1-2200 158 - E1-2500 159 - E2-3000 160 - E2-3800 161 - A4-5000 162 - A4-5100 163 - A6-5200 164 - A4 Pro-3340B 165 ``gfx704`` - ``bonaire`` ``amdgcn`` dGPU - Radeon HD 7790 166 - Radeon HD 8770 167 - R7 260 168 - R7 260X 169 **GCN GFX8 (Volcanic Islands (VI))** [AMD-GCN-GFX8]_ 170 ----------------------------------------------------------------------------------------------- 171 ``gfx801`` - ``carrizo`` ``amdgcn`` APU - xnack - A6-8500P 172 [on] - Pro A6-8500B 173 - A8-8600P 174 - Pro A8-8600B 175 - FX-8800P 176 - Pro A12-8800B 177 \ ``amdgcn`` APU - xnack ROCm - A10-8700P 178 [on] - Pro A10-8700B 179 - A10-8780P 180 \ ``amdgcn`` APU - xnack - A10-9600P 181 [on] - A10-9630P 182 - A12-9700P 183 - A12-9730P 184 - FX-9800P 185 - FX-9830P 186 \ ``amdgcn`` APU - xnack - E2-9010 187 [on] - A6-9210 188 - A9-9410 189 ``gfx802`` - ``iceland`` ``amdgcn`` dGPU - xnack ROCm - FirePro S7150 190 - ``tonga`` [off] - FirePro S7100 191 - FirePro W7100 192 - Radeon R285 193 - Radeon R9 380 194 - Radeon R9 385 195 - Mobile FirePro 196 M7170 197 ``gfx803`` - ``fiji`` ``amdgcn`` dGPU - xnack ROCm - Radeon R9 Nano 198 [off] - Radeon R9 Fury 199 - Radeon R9 FuryX 200 - Radeon Pro Duo 201 - FirePro S9300x2 202 - Radeon Instinct MI8 203 \ - ``polaris10`` ``amdgcn`` dGPU - xnack ROCm - Radeon RX 470 204 [off] - Radeon RX 480 205 - Radeon Instinct MI6 206 \ - ``polaris11`` ``amdgcn`` dGPU - xnack ROCm - Radeon RX 460 207 [off] 208 ``gfx810`` - ``stoney`` ``amdgcn`` APU - xnack 209 [on] 210 **GCN GFX9** [AMD-GCN-GFX9]_ 211 ----------------------------------------------------------------------------------------------- 212 ``gfx900`` ``amdgcn`` dGPU - xnack ROCm - Radeon Vega 213 [off] Frontier Edition 214 - Radeon RX Vega 56 215 - Radeon RX Vega 64 216 - Radeon RX Vega 64 217 Liquid 218 - Radeon Instinct MI25 219 ``gfx902`` ``amdgcn`` APU - xnack - Ryzen 3 2200G 220 [on] - Ryzen 5 2400G 221 ``gfx904`` ``amdgcn`` dGPU - xnack *TBA* 222 [off] 223 .. TODO:: 224 Add product 225 names. 226 ``gfx906`` ``amdgcn`` dGPU - xnack - Radeon Instinct MI50 227 [off] - Radeon Instinct MI60 228 - Radeon VII 229 - Radeon Pro VII 230 ``gfx908`` ``amdgcn`` dGPU - xnack *TBA* 231 [off] 232 sram-ecc 233 [on] 234 .. TODO:: 235 Add product 236 names. 237 ``gfx909`` ``amdgcn`` APU - xnack *TBA* 238 [on] 239 .. TODO:: 240 Add product 241 names. 242 **GCN GFX10** [AMD-GCN-GFX10]_ 243 ----------------------------------------------------------------------------------------------- 244 ``gfx1010`` ``amdgcn`` dGPU - xnack - Radeon RX 5700 245 [off] - Radeon RX 5700 XT 246 - wavefrontsize64 - Radeon Pro 5600 XT 247 [off] 248 - cumode 249 [off] 250 ``gfx1011`` ``amdgcn`` dGPU - xnack - Radeon Pro 5600M 251 [off] 252 - wavefrontsize64 253 [off] 254 - cumode 255 [off] 256 ``gfx1012`` ``amdgcn`` dGPU - xnack - Radeon RX 5500 257 [off] - Radeon RX 5500 XT 258 - wavefrontsize64 259 [off] 260 - cumode 261 [off] 262 ``gfx1030`` ``amdgcn`` dGPU - wavefrontsize64 *TBA* 263 [off] 264 - cumode 265 [off] 266 .. TODO 267 Add product 268 names. 269 =========== =============== ============ ===== ================= ======= ====================== 270 271.. _amdgpu-target-features: 272 273Target Features 274--------------- 275 276Target features control how code is generated to support certain 277processor specific features. Not all target features are supported by 278all processors. The runtime must ensure that the features supported by 279the device used to execute the code match the features enabled when 280generating the code. A mismatch of features may result in incorrect 281execution, or a reduction in performance. 282 283The target features supported by each processor, and the default value 284used if not specified explicitly, is listed in 285:ref:`amdgpu-processor-table`. 286 287Use the ``clang -m[no-]<TargetFeature>`` option to specify the AMDGPU 288target features. 289 290For example: 291 292``-mxnack`` 293 Enable the ``xnack`` feature. 294``-mno-xnack`` 295 Disable the ``xnack`` feature. 296 297 .. table:: AMDGPU Target Features 298 :name: amdgpu-target-feature-table 299 300 ====================== ================================================== 301 Target Feature Description 302 ====================== ================================================== 303 -m[no-]xnack Enable/disable generating code that has 304 memory clauses that are compatible with 305 having XNACK replay enabled. 306 307 This is used for demand paging and page 308 migration. If XNACK replay is enabled in 309 the device, then if a page fault occurs 310 the code may execute incorrectly if the 311 ``xnack`` feature is not enabled. Executing 312 code that has the feature enabled on a 313 device that does not have XNACK replay 314 enabled will execute correctly but may 315 be less performant than code with the 316 feature disabled. 317 318 -m[no-]sram-ecc Enable/disable generating code that assumes SRAM 319 ECC is enabled/disabled. 320 321 -m[no-]wavefrontsize64 Control the default wavefront size used when 322 generating code for kernels. When disabled 323 native wavefront size 32 is used, when enabled 324 wavefront size 64 is used. 325 326 -m[no-]cumode Control the default wavefront execution mode used 327 when generating code for kernels. When disabled 328 native WGP wavefront execution mode is used, 329 when enabled CU wavefront execution mode is used 330 (see :ref:`amdgpu-amdhsa-memory-model`). 331 ====================== ================================================== 332 333.. _amdgpu-address-spaces: 334 335Address Spaces 336-------------- 337 338The AMDGPU architecture supports a number of memory address spaces. The address 339space names use the OpenCL standard names, with some additions. 340 341The AMDGPU address spaces correspond to target architecture specific LLVM 342address space numbers used in LLVM IR. 343 344The AMDGPU address spaces are described in 345:ref:`amdgpu-address-spaces-table`. Only 64-bit process address spaces are 346supported for the ``amdgcn`` target. 347 348 .. table:: AMDGPU Address Spaces 349 :name: amdgpu-address-spaces-table 350 351 ================================= =============== =========== ================ ======= ============================ 352 .. 64-Bit Process Address Space 353 --------------------------------- --------------- ----------- ---------------- ------------------------------------ 354 Address Space Name LLVM IR Address HSA Segment Hardware Address NULL Value 355 Space Number Name Name Size 356 ================================= =============== =========== ================ ======= ============================ 357 Generic 0 flat flat 64 0x0000000000000000 358 Global 1 global global 64 0x0000000000000000 359 Region 2 N/A GDS 32 *not implemented for AMDHSA* 360 Local 3 group LDS 32 0xFFFFFFFF 361 Constant 4 constant *same as global* 64 0x0000000000000000 362 Private 5 private scratch 32 0xFFFFFFFF 363 Constant 32-bit 6 *TODO* 0x00000000 364 Buffer Fat Pointer (experimental) 7 *TODO* 365 ================================= =============== =========== ================ ======= ============================ 366 367**Generic** 368 The generic address space uses the hardware flat address support available in 369 GFX7-GFX10. This uses two fixed ranges of virtual addresses (the private and 370 local apertures), that are outside the range of addressable global memory, to 371 map from a flat address to a private or local address. 372 373 FLAT instructions can take a flat address and access global, private 374 (scratch), and group (LDS) memory depending on if the address is within one 375 of the aperture ranges. Flat access to scratch requires hardware aperture 376 setup and setup in the kernel prologue (see 377 :ref:`amdgpu-amdhsa-kernel-prolog-flat-scratch`). Flat access to LDS requires 378 hardware aperture setup and M0 (GFX7-GFX8) register setup (see 379 :ref:`amdgpu-amdhsa-kernel-prolog-m0`). 380 381 To convert between a private or group address space address (termed a segment 382 address) and a flat address the base address of the corresponding aperture 383 can be used. For GFX7-GFX8 these are available in the 384 :ref:`amdgpu-amdhsa-hsa-aql-queue` the address of which can be obtained with 385 Queue Ptr SGPR (see :ref:`amdgpu-amdhsa-initial-kernel-execution-state`). For 386 GFX9-GFX10 the aperture base addresses are directly available as inline 387 constant registers ``SRC_SHARED_BASE/LIMIT`` and ``SRC_PRIVATE_BASE/LIMIT``. 388 In 64-bit address mode the aperture sizes are 2^32 bytes and the base is 389 aligned to 2^32 which makes it easier to convert from flat to segment or 390 segment to flat. 391 392 A global address space address has the same value when used as a flat address 393 so no conversion is needed. 394 395**Global and Constant** 396 The global and constant address spaces both use global virtual addresses, 397 which are the same virtual address space used by the CPU. However, some 398 virtual addresses may only be accessible to the CPU, some only accessible 399 by the GPU, and some by both. 400 401 Using the constant address space indicates that the data will not change 402 during the execution of the kernel. This allows scalar read instructions to 403 be used. The vector and scalar L1 caches are invalidated of volatile data 404 before each kernel dispatch execution to allow constant memory to change 405 values between kernel dispatches. 406 407**Region** 408 The region address space uses the hardware Global Data Store (GDS). All 409 wavefronts executing on the same device will access the same memory for any 410 given region address. However, the same region address accessed by wavefronts 411 executing on different devices will access different memory. It is higher 412 performance than global memory. It is allocated by the runtime. The data 413 store (DS) instructions can be used to access it. 414 415**Local** 416 The local address space uses the hardware Local Data Store (LDS) which is 417 automatically allocated when the hardware creates the wavefronts of a 418 work-group, and freed when all the wavefronts of a work-group have 419 terminated. All wavefronts belonging to the same work-group will access the 420 same memory for any given local address. However, the same local address 421 accessed by wavefronts belonging to different work-groups will access 422 different memory. It is higher performance than global memory. The data store 423 (DS) instructions can be used to access it. 424 425**Private** 426 The private address space uses the hardware scratch memory support which 427 automatically allocates memory when it creates a wavefront and frees it when 428 a wavefronts terminates. The memory accessed by a lane of a wavefront for any 429 given private address will be different to the memory accessed by another lane 430 of the same or different wavefront for the same private address. 431 432 If a kernel dispatch uses scratch, then the hardware allocates memory from a 433 pool of backing memory allocated by the runtime for each wavefront. The lanes 434 of the wavefront access this using dword (4 byte) interleaving. The mapping 435 used from private address to backing memory address is: 436 437 ``wavefront-scratch-base + 438 ((private-address / 4) * wavefront-size * 4) + 439 (wavefront-lane-id * 4) + (private-address % 4)`` 440 441 If each lane of a wavefront accesses the same private address, the 442 interleaving results in adjacent dwords being accessed and hence requires 443 fewer cache lines to be fetched. 444 445 There are different ways that the wavefront scratch base address is 446 determined by a wavefront (see 447 :ref:`amdgpu-amdhsa-initial-kernel-execution-state`). 448 449 Scratch memory can be accessed in an interleaved manner using buffer 450 instructions with the scratch buffer descriptor and per wavefront scratch 451 offset, by the scratch instructions, or by flat instructions. Multi-dword 452 access is not supported except by flat and scratch instructions in 453 GFX9-GFX10. 454 455**Constant 32-bit** 456 *TODO* 457 458**Buffer Fat Pointer** 459 The buffer fat pointer is an experimental address space that is currently 460 unsupported in the backend. It exposes a non-integral pointer that is in 461 the future intended to support the modelling of 128-bit buffer descriptors 462 plus a 32-bit offset into the buffer (in total encapsulating a 160-bit 463 *pointer*), allowing normal LLVM load/store/atomic operations to be used to 464 model the buffer descriptors used heavily in graphics workloads targeting 465 the backend. 466 467.. _amdgpu-memory-scopes: 468 469Memory Scopes 470------------- 471 472This section provides LLVM memory synchronization scopes supported by the AMDGPU 473backend memory model when the target triple OS is ``amdhsa`` (see 474:ref:`amdgpu-amdhsa-memory-model` and :ref:`amdgpu-target-triples`). 475 476The memory model supported is based on the HSA memory model [HSA]_ which is 477based in turn on HRF-indirect with scope inclusion [HRF]_. The happens-before 478relation is transitive over the synchronizes-with relation independent of scope 479and synchronizes-with allows the memory scope instances to be inclusive (see 480table :ref:`amdgpu-amdhsa-llvm-sync-scopes-table`). 481 482This is different to the OpenCL [OpenCL]_ memory model which does not have scope 483inclusion and requires the memory scopes to exactly match. However, this 484is conservatively correct for OpenCL. 485 486 .. table:: AMDHSA LLVM Sync Scopes 487 :name: amdgpu-amdhsa-llvm-sync-scopes-table 488 489 ======================= =================================================== 490 LLVM Sync Scope Description 491 ======================= =================================================== 492 *none* The default: ``system``. 493 494 Synchronizes with, and participates in modification 495 and seq_cst total orderings with, other operations 496 (except image operations) for all address spaces 497 (except private, or generic that accesses private) 498 provided the other operation's sync scope is: 499 500 - ``system``. 501 - ``agent`` and executed by a thread on the same 502 agent. 503 - ``workgroup`` and executed by a thread in the 504 same work-group. 505 - ``wavefront`` and executed by a thread in the 506 same wavefront. 507 508 ``agent`` Synchronizes with, and participates in modification 509 and seq_cst total orderings with, other operations 510 (except image operations) for all address spaces 511 (except private, or generic that accesses private) 512 provided the other operation's sync scope is: 513 514 - ``system`` or ``agent`` and executed by a thread 515 on the same agent. 516 - ``workgroup`` and executed by a thread in the 517 same work-group. 518 - ``wavefront`` and executed by a thread in the 519 same wavefront. 520 521 ``workgroup`` Synchronizes with, and participates in modification 522 and seq_cst total orderings with, other operations 523 (except image operations) for all address spaces 524 (except private, or generic that accesses private) 525 provided the other operation's sync scope is: 526 527 - ``system``, ``agent`` or ``workgroup`` and 528 executed by a thread in the same work-group. 529 - ``wavefront`` and executed by a thread in the 530 same wavefront. 531 532 ``wavefront`` Synchronizes with, and participates in modification 533 and seq_cst total orderings with, other operations 534 (except image operations) for all address spaces 535 (except private, or generic that accesses private) 536 provided the other operation's sync scope is: 537 538 - ``system``, ``agent``, ``workgroup`` or 539 ``wavefront`` and executed by a thread in the 540 same wavefront. 541 542 ``singlethread`` Only synchronizes with and participates in 543 modification and seq_cst total orderings with, 544 other operations (except image operations) running 545 in the same thread for all address spaces (for 546 example, in signal handlers). 547 548 ``one-as`` Same as ``system`` but only synchronizes with other 549 operations within the same address space. 550 551 ``agent-one-as`` Same as ``agent`` but only synchronizes with other 552 operations within the same address space. 553 554 ``workgroup-one-as`` Same as ``workgroup`` but only synchronizes with 555 other operations within the same address space. 556 557 ``wavefront-one-as`` Same as ``wavefront`` but only synchronizes with 558 other operations within the same address space. 559 560 ``singlethread-one-as`` Same as ``singlethread`` but only synchronizes with 561 other operations within the same address space. 562 ======================= =================================================== 563 564LLVM IR Intrinsics 565------------------ 566 567The AMDGPU backend implements the following LLVM IR intrinsics. 568 569*This section is WIP.* 570 571.. TODO:: 572 573 List AMDGPU intrinsics. 574 575LLVM IR Attributes 576------------------ 577 578The AMDGPU backend supports the following LLVM IR attributes. 579 580 .. table:: AMDGPU LLVM IR Attributes 581 :name: amdgpu-llvm-ir-attributes-table 582 583 ======================================= ========================================================== 584 LLVM Attribute Description 585 ======================================= ========================================================== 586 "amdgpu-flat-work-group-size"="min,max" Specify the minimum and maximum flat work group sizes that 587 will be specified when the kernel is dispatched. Generated 588 by the ``amdgpu_flat_work_group_size`` CLANG attribute [CLANG-ATTR]_. 589 "amdgpu-implicitarg-num-bytes"="n" Number of kernel argument bytes to add to the kernel 590 argument block size for the implicit arguments. This 591 varies by OS and language (for OpenCL see 592 :ref:`opencl-kernel-implicit-arguments-appended-for-amdhsa-os-table`). 593 "amdgpu-num-sgpr"="n" Specifies the number of SGPRs to use. Generated by 594 the ``amdgpu_num_sgpr`` CLANG attribute [CLANG-ATTR]_. 595 "amdgpu-num-vgpr"="n" Specifies the number of VGPRs to use. Generated by the 596 ``amdgpu_num_vgpr`` CLANG attribute [CLANG-ATTR]_. 597 "amdgpu-waves-per-eu"="m,n" Specify the minimum and maximum number of waves per 598 execution unit. Generated by the ``amdgpu_waves_per_eu`` 599 CLANG attribute [CLANG-ATTR]_. 600 "amdgpu-ieee" true/false. Specify whether the function expects the IEEE field of the 601 mode register to be set on entry. Overrides the default for 602 the calling convention. 603 "amdgpu-dx10-clamp" true/false. Specify whether the function expects the DX10_CLAMP field of 604 the mode register to be set on entry. Overrides the default 605 for the calling convention. 606 ======================================= ========================================================== 607 608.. _amdgpu-elf-code-object: 609 610ELF Code Object 611=============== 612 613The AMDGPU backend generates a standard ELF [ELF]_ relocatable code object that 614can be linked by ``lld`` to produce a standard ELF shared code object which can 615be loaded and executed on an AMDGPU target. 616 617.. _amdgpu-elf-header: 618 619Header 620------ 621 622The AMDGPU backend uses the following ELF header: 623 624 .. table:: AMDGPU ELF Header 625 :name: amdgpu-elf-header-table 626 627 ========================== =============================== 628 Field Value 629 ========================== =============================== 630 ``e_ident[EI_CLASS]`` ``ELFCLASS64`` 631 ``e_ident[EI_DATA]`` ``ELFDATA2LSB`` 632 ``e_ident[EI_OSABI]`` - ``ELFOSABI_NONE`` 633 - ``ELFOSABI_AMDGPU_HSA`` 634 - ``ELFOSABI_AMDGPU_PAL`` 635 - ``ELFOSABI_AMDGPU_MESA3D`` 636 ``e_ident[EI_ABIVERSION]`` - ``ELFABIVERSION_AMDGPU_HSA`` 637 - ``ELFABIVERSION_AMDGPU_PAL`` 638 - ``ELFABIVERSION_AMDGPU_MESA3D`` 639 ``e_type`` - ``ET_REL`` 640 - ``ET_DYN`` 641 ``e_machine`` ``EM_AMDGPU`` 642 ``e_entry`` 0 643 ``e_flags`` See :ref:`amdgpu-elf-header-e_flags-table` 644 ========================== =============================== 645 646.. 647 648 .. table:: AMDGPU ELF Header Enumeration Values 649 :name: amdgpu-elf-header-enumeration-values-table 650 651 =============================== ===== 652 Name Value 653 =============================== ===== 654 ``EM_AMDGPU`` 224 655 ``ELFOSABI_NONE`` 0 656 ``ELFOSABI_AMDGPU_HSA`` 64 657 ``ELFOSABI_AMDGPU_PAL`` 65 658 ``ELFOSABI_AMDGPU_MESA3D`` 66 659 ``ELFABIVERSION_AMDGPU_HSA`` 1 660 ``ELFABIVERSION_AMDGPU_PAL`` 0 661 ``ELFABIVERSION_AMDGPU_MESA3D`` 0 662 =============================== ===== 663 664``e_ident[EI_CLASS]`` 665 The ELF class is: 666 667 * ``ELFCLASS32`` for ``r600`` architecture. 668 669 * ``ELFCLASS64`` for ``amdgcn`` architecture which only supports 64-bit 670 process address space applications. 671 672``e_ident[EI_DATA]`` 673 All AMDGPU targets use ``ELFDATA2LSB`` for little-endian byte ordering. 674 675``e_ident[EI_OSABI]`` 676 One of the following AMDGPU target architecture specific OS ABIs 677 (see :ref:`amdgpu-os-table`): 678 679 * ``ELFOSABI_NONE`` for *unknown* OS. 680 681 * ``ELFOSABI_AMDGPU_HSA`` for ``amdhsa`` OS. 682 683 * ``ELFOSABI_AMDGPU_PAL`` for ``amdpal`` OS. 684 685 * ``ELFOSABI_AMDGPU_MESA3D`` for ``mesa3D`` OS. 686 687``e_ident[EI_ABIVERSION]`` 688 The ABI version of the AMDGPU target architecture specific OS ABI to which the code 689 object conforms: 690 691 * ``ELFABIVERSION_AMDGPU_HSA`` is used to specify the version of AMD HSA 692 runtime ABI. 693 694 * ``ELFABIVERSION_AMDGPU_PAL`` is used to specify the version of AMD PAL 695 runtime ABI. 696 697 * ``ELFABIVERSION_AMDGPU_MESA3D`` is used to specify the version of AMD MESA 698 3D runtime ABI. 699 700``e_type`` 701 Can be one of the following values: 702 703 704 ``ET_REL`` 705 The type produced by the AMDGPU backend compiler as it is relocatable code 706 object. 707 708 ``ET_DYN`` 709 The type produced by the linker as it is a shared code object. 710 711 The AMD HSA runtime loader requires a ``ET_DYN`` code object. 712 713``e_machine`` 714 The value ``EM_AMDGPU`` is used for the machine for all processors supported 715 by the ``r600`` and ``amdgcn`` architectures (see 716 :ref:`amdgpu-processor-table`). The specific processor is specified in the 717 ``EF_AMDGPU_MACH`` bit field of the ``e_flags`` (see 718 :ref:`amdgpu-elf-header-e_flags-table`). 719 720``e_entry`` 721 The entry point is 0 as the entry points for individual kernels must be 722 selected in order to invoke them through AQL packets. 723 724``e_flags`` 725 The AMDGPU backend uses the following ELF header flags: 726 727 .. table:: AMDGPU ELF Header ``e_flags`` 728 :name: amdgpu-elf-header-e_flags-table 729 730 ================================= ========== ============================= 731 Name Value Description 732 ================================= ========== ============================= 733 **AMDGPU Processor Flag** See :ref:`amdgpu-processor-table`. 734 -------------------------------------------- ----------------------------- 735 ``EF_AMDGPU_MACH`` 0x000000ff AMDGPU processor selection 736 mask for 737 ``EF_AMDGPU_MACH_xxx`` values 738 defined in 739 :ref:`amdgpu-ef-amdgpu-mach-table`. 740 ``EF_AMDGPU_XNACK`` 0x00000100 Indicates if the ``xnack`` 741 target feature is 742 enabled for all code 743 contained in the code object. 744 If the processor 745 does not support the 746 ``xnack`` target 747 feature then must 748 be 0. 749 See 750 :ref:`amdgpu-target-features`. 751 ``EF_AMDGPU_SRAM_ECC`` 0x00000200 Indicates if the ``sram-ecc`` 752 target feature is 753 enabled for all code 754 contained in the code object. 755 If the processor 756 does not support the 757 ``sram-ecc`` target 758 feature then must 759 be 0. 760 See 761 :ref:`amdgpu-target-features`. 762 ================================= ========== ============================= 763 764 .. table:: AMDGPU ``EF_AMDGPU_MACH`` Values 765 :name: amdgpu-ef-amdgpu-mach-table 766 767 ================================= ========== ============================= 768 Name Value Description (see 769 :ref:`amdgpu-processor-table`) 770 ================================= ========== ============================= 771 ``EF_AMDGPU_MACH_NONE`` 0x000 *not specified* 772 ``EF_AMDGPU_MACH_R600_R600`` 0x001 ``r600`` 773 ``EF_AMDGPU_MACH_R600_R630`` 0x002 ``r630`` 774 ``EF_AMDGPU_MACH_R600_RS880`` 0x003 ``rs880`` 775 ``EF_AMDGPU_MACH_R600_RV670`` 0x004 ``rv670`` 776 ``EF_AMDGPU_MACH_R600_RV710`` 0x005 ``rv710`` 777 ``EF_AMDGPU_MACH_R600_RV730`` 0x006 ``rv730`` 778 ``EF_AMDGPU_MACH_R600_RV770`` 0x007 ``rv770`` 779 ``EF_AMDGPU_MACH_R600_CEDAR`` 0x008 ``cedar`` 780 ``EF_AMDGPU_MACH_R600_CYPRESS`` 0x009 ``cypress`` 781 ``EF_AMDGPU_MACH_R600_JUNIPER`` 0x00a ``juniper`` 782 ``EF_AMDGPU_MACH_R600_REDWOOD`` 0x00b ``redwood`` 783 ``EF_AMDGPU_MACH_R600_SUMO`` 0x00c ``sumo`` 784 ``EF_AMDGPU_MACH_R600_BARTS`` 0x00d ``barts`` 785 ``EF_AMDGPU_MACH_R600_CAICOS`` 0x00e ``caicos`` 786 ``EF_AMDGPU_MACH_R600_CAYMAN`` 0x00f ``cayman`` 787 ``EF_AMDGPU_MACH_R600_TURKS`` 0x010 ``turks`` 788 *reserved* 0x011 - Reserved for ``r600`` 789 0x01f architecture processors. 790 ``EF_AMDGPU_MACH_AMDGCN_GFX600`` 0x020 ``gfx600`` 791 ``EF_AMDGPU_MACH_AMDGCN_GFX601`` 0x021 ``gfx601`` 792 ``EF_AMDGPU_MACH_AMDGCN_GFX700`` 0x022 ``gfx700`` 793 ``EF_AMDGPU_MACH_AMDGCN_GFX701`` 0x023 ``gfx701`` 794 ``EF_AMDGPU_MACH_AMDGCN_GFX702`` 0x024 ``gfx702`` 795 ``EF_AMDGPU_MACH_AMDGCN_GFX703`` 0x025 ``gfx703`` 796 ``EF_AMDGPU_MACH_AMDGCN_GFX704`` 0x026 ``gfx704`` 797 *reserved* 0x027 Reserved. 798 ``EF_AMDGPU_MACH_AMDGCN_GFX801`` 0x028 ``gfx801`` 799 ``EF_AMDGPU_MACH_AMDGCN_GFX802`` 0x029 ``gfx802`` 800 ``EF_AMDGPU_MACH_AMDGCN_GFX803`` 0x02a ``gfx803`` 801 ``EF_AMDGPU_MACH_AMDGCN_GFX810`` 0x02b ``gfx810`` 802 ``EF_AMDGPU_MACH_AMDGCN_GFX900`` 0x02c ``gfx900`` 803 ``EF_AMDGPU_MACH_AMDGCN_GFX902`` 0x02d ``gfx902`` 804 ``EF_AMDGPU_MACH_AMDGCN_GFX904`` 0x02e ``gfx904`` 805 ``EF_AMDGPU_MACH_AMDGCN_GFX906`` 0x02f ``gfx906`` 806 ``EF_AMDGPU_MACH_AMDGCN_GFX908`` 0x030 ``gfx908`` 807 ``EF_AMDGPU_MACH_AMDGCN_GFX909`` 0x031 ``gfx909`` 808 *reserved* 0x032 Reserved. 809 ``EF_AMDGPU_MACH_AMDGCN_GFX1010`` 0x033 ``gfx1010`` 810 ``EF_AMDGPU_MACH_AMDGCN_GFX1011`` 0x034 ``gfx1011`` 811 ``EF_AMDGPU_MACH_AMDGCN_GFX1012`` 0x035 ``gfx1012`` 812 ``EF_AMDGPU_MACH_AMDGCN_GFX1030`` 0x036 ``gfx1030`` 813 ================================= ========== ============================= 814 815Sections 816-------- 817 818An AMDGPU target ELF code object has the standard ELF sections which include: 819 820 .. table:: AMDGPU ELF Sections 821 :name: amdgpu-elf-sections-table 822 823 ================== ================ ================================= 824 Name Type Attributes 825 ================== ================ ================================= 826 ``.bss`` ``SHT_NOBITS`` ``SHF_ALLOC`` + ``SHF_WRITE`` 827 ``.data`` ``SHT_PROGBITS`` ``SHF_ALLOC`` + ``SHF_WRITE`` 828 ``.debug_``\ *\** ``SHT_PROGBITS`` *none* 829 ``.dynamic`` ``SHT_DYNAMIC`` ``SHF_ALLOC`` 830 ``.dynstr`` ``SHT_PROGBITS`` ``SHF_ALLOC`` 831 ``.dynsym`` ``SHT_PROGBITS`` ``SHF_ALLOC`` 832 ``.got`` ``SHT_PROGBITS`` ``SHF_ALLOC`` + ``SHF_WRITE`` 833 ``.hash`` ``SHT_HASH`` ``SHF_ALLOC`` 834 ``.note`` ``SHT_NOTE`` *none* 835 ``.rela``\ *name* ``SHT_RELA`` *none* 836 ``.rela.dyn`` ``SHT_RELA`` *none* 837 ``.rodata`` ``SHT_PROGBITS`` ``SHF_ALLOC`` 838 ``.shstrtab`` ``SHT_STRTAB`` *none* 839 ``.strtab`` ``SHT_STRTAB`` *none* 840 ``.symtab`` ``SHT_SYMTAB`` *none* 841 ``.text`` ``SHT_PROGBITS`` ``SHF_ALLOC`` + ``SHF_EXECINSTR`` 842 ================== ================ ================================= 843 844These sections have their standard meanings (see [ELF]_) and are only generated 845if needed. 846 847``.debug``\ *\** 848 The standard DWARF sections. See :ref:`amdgpu-dwarf-debug-information` for 849 information on the DWARF produced by the AMDGPU backend. 850 851``.dynamic``, ``.dynstr``, ``.dynsym``, ``.hash`` 852 The standard sections used by a dynamic loader. 853 854``.note`` 855 See :ref:`amdgpu-note-records` for the note records supported by the AMDGPU 856 backend. 857 858``.rela``\ *name*, ``.rela.dyn`` 859 For relocatable code objects, *name* is the name of the section that the 860 relocation records apply. For example, ``.rela.text`` is the section name for 861 relocation records associated with the ``.text`` section. 862 863 For linked shared code objects, ``.rela.dyn`` contains all the relocation 864 records from each of the relocatable code object's ``.rela``\ *name* sections. 865 866 See :ref:`amdgpu-relocation-records` for the relocation records supported by 867 the AMDGPU backend. 868 869``.text`` 870 The executable machine code for the kernels and functions they call. Generated 871 as position independent code. See :ref:`amdgpu-code-conventions` for 872 information on conventions used in the isa generation. 873 874.. _amdgpu-note-records: 875 876Note Records 877------------ 878 879The AMDGPU backend code object contains ELF note records in the ``.note`` 880section. The set of generated notes and their semantics depend on the code 881object version; see :ref:`amdgpu-note-records-v2` and 882:ref:`amdgpu-note-records-v3`. 883 884As required by ``ELFCLASS32`` and ``ELFCLASS64``, minimal zero-byte padding 885must be generated after the ``name`` field to ensure the ``desc`` field is 4 886byte aligned. In addition, minimal zero-byte padding must be generated to 887ensure the ``desc`` field size is a multiple of 4 bytes. The ``sh_addralign`` 888field of the ``.note`` section must be at least 4 to indicate at least 8 byte 889alignment. 890 891.. _amdgpu-note-records-v2: 892 893Code Object V2 Note Records (-mattr=-code-object-v3) 894~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 895 896.. warning:: Code Object V2 is not the default code object version emitted by 897 this version of LLVM. For a description of the notes generated with the 898 default configuration (Code Object V3) see :ref:`amdgpu-note-records-v3`. 899 900The AMDGPU backend code object uses the following ELF note record in the 901``.note`` section when compiling for Code Object V2 (-mattr=-code-object-v3). 902 903Additional note records may be present, but any which are not documented here 904are deprecated and should not be used. 905 906 .. table:: AMDGPU Code Object V2 ELF Note Records 907 :name: amdgpu-elf-note-records-table-v2 908 909 ===== ============================== ====================================== 910 Name Type Description 911 ===== ============================== ====================================== 912 "AMD" ``NT_AMD_AMDGPU_HSA_METADATA`` <metadata null terminated string> 913 ===== ============================== ====================================== 914 915.. 916 917 .. table:: AMDGPU Code Object V2 ELF Note Record Enumeration Values 918 :name: amdgpu-elf-note-record-enumeration-values-table-v2 919 920 ============================== ===== 921 Name Value 922 ============================== ===== 923 *reserved* 0-9 924 ``NT_AMD_AMDGPU_HSA_METADATA`` 10 925 *reserved* 11 926 ============================== ===== 927 928``NT_AMD_AMDGPU_HSA_METADATA`` 929 Specifies extensible metadata associated with the code objects executed on HSA 930 [HSA]_ compatible runtimes such as AMD's ROCm [AMD-ROCm]_. It is required when 931 the target triple OS is ``amdhsa`` (see :ref:`amdgpu-target-triples`). See 932 :ref:`amdgpu-amdhsa-code-object-metadata-v2` for the syntax of the code 933 object metadata string. 934 935.. _amdgpu-note-records-v3: 936 937Code Object V3 Note Records (-mattr=+code-object-v3) 938~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 939 940The AMDGPU backend code object uses the following ELF note record in the 941``.note`` section when compiling for Code Object V3 (-mattr=+code-object-v3). 942 943Additional note records may be present, but any which are not documented here 944are deprecated and should not be used. 945 946 .. table:: AMDGPU Code Object V3 ELF Note Records 947 :name: amdgpu-elf-note-records-table-v3 948 949 ======== ============================== ====================================== 950 Name Type Description 951 ======== ============================== ====================================== 952 "AMDGPU" ``NT_AMDGPU_METADATA`` Metadata in Message Pack [MsgPack]_ 953 binary format. 954 ======== ============================== ====================================== 955 956.. 957 958 .. table:: AMDGPU Code Object V3 ELF Note Record Enumeration Values 959 :name: amdgpu-elf-note-record-enumeration-values-table-v3 960 961 ============================== ===== 962 Name Value 963 ============================== ===== 964 *reserved* 0-31 965 ``NT_AMDGPU_METADATA`` 32 966 ============================== ===== 967 968``NT_AMDGPU_METADATA`` 969 Specifies extensible metadata associated with an AMDGPU code 970 object. It is encoded as a map in the Message Pack [MsgPack]_ binary 971 data format. See :ref:`amdgpu-amdhsa-code-object-metadata-v3` for the 972 map keys defined for the ``amdhsa`` OS. 973 974.. _amdgpu-symbols: 975 976Symbols 977------- 978 979Symbols include the following: 980 981 .. table:: AMDGPU ELF Symbols 982 :name: amdgpu-elf-symbols-table 983 984 ===================== ================== ================ ================== 985 Name Type Section Description 986 ===================== ================== ================ ================== 987 *link-name* ``STT_OBJECT`` - ``.data`` Global variable 988 - ``.rodata`` 989 - ``.bss`` 990 *link-name*\ ``.kd`` ``STT_OBJECT`` - ``.rodata`` Kernel descriptor 991 *link-name* ``STT_FUNC`` - ``.text`` Kernel entry point 992 *link-name* ``STT_OBJECT`` - SHN_AMDGPU_LDS Global variable in LDS 993 ===================== ================== ================ ================== 994 995Global variable 996 Global variables both used and defined by the compilation unit. 997 998 If the symbol is defined in the compilation unit then it is allocated in the 999 appropriate section according to if it has initialized data or is readonly. 1000 1001 If the symbol is external then its section is ``STN_UNDEF`` and the loader 1002 will resolve relocations using the definition provided by another code object 1003 or explicitly defined by the runtime. 1004 1005 If the symbol resides in local/group memory (LDS) then its section is the 1006 special processor specific section name ``SHN_AMDGPU_LDS``, and the 1007 ``st_value`` field describes alignment requirements as it does for common 1008 symbols. 1009 1010 .. TODO:: 1011 1012 Add description of linked shared object symbols. Seems undefined symbols 1013 are marked as STT_NOTYPE. 1014 1015Kernel descriptor 1016 Every HSA kernel has an associated kernel descriptor. It is the address of the 1017 kernel descriptor that is used in the AQL dispatch packet used to invoke the 1018 kernel, not the kernel entry point. The layout of the HSA kernel descriptor is 1019 defined in :ref:`amdgpu-amdhsa-kernel-descriptor`. 1020 1021Kernel entry point 1022 Every HSA kernel also has a symbol for its machine code entry point. 1023 1024.. _amdgpu-relocation-records: 1025 1026Relocation Records 1027------------------ 1028 1029AMDGPU backend generates ``Elf64_Rela`` relocation records. Supported 1030relocatable fields are: 1031 1032``word32`` 1033 This specifies a 32-bit field occupying 4 bytes with arbitrary byte 1034 alignment. These values use the same byte order as other word values in the 1035 AMDGPU architecture. 1036 1037``word64`` 1038 This specifies a 64-bit field occupying 8 bytes with arbitrary byte 1039 alignment. These values use the same byte order as other word values in the 1040 AMDGPU architecture. 1041 1042Following notations are used for specifying relocation calculations: 1043 1044**A** 1045 Represents the addend used to compute the value of the relocatable field. 1046 1047**G** 1048 Represents the offset into the global offset table at which the relocation 1049 entry's symbol will reside during execution. 1050 1051**GOT** 1052 Represents the address of the global offset table. 1053 1054**P** 1055 Represents the place (section offset for ``et_rel`` or address for ``et_dyn``) 1056 of the storage unit being relocated (computed using ``r_offset``). 1057 1058**S** 1059 Represents the value of the symbol whose index resides in the relocation 1060 entry. Relocations not using this must specify a symbol index of 1061 ``STN_UNDEF``. 1062 1063**B** 1064 Represents the base address of a loaded executable or shared object which is 1065 the difference between the ELF address and the actual load address. 1066 Relocations using this are only valid in executable or shared objects. 1067 1068The following relocation types are supported: 1069 1070 .. table:: AMDGPU ELF Relocation Records 1071 :name: amdgpu-elf-relocation-records-table 1072 1073 ========================== ======= ===== ========== ============================== 1074 Relocation Type Kind Value Field Calculation 1075 ========================== ======= ===== ========== ============================== 1076 ``R_AMDGPU_NONE`` 0 *none* *none* 1077 ``R_AMDGPU_ABS32_LO`` Static, 1 ``word32`` (S + A) & 0xFFFFFFFF 1078 Dynamic 1079 ``R_AMDGPU_ABS32_HI`` Static, 2 ``word32`` (S + A) >> 32 1080 Dynamic 1081 ``R_AMDGPU_ABS64`` Static, 3 ``word64`` S + A 1082 Dynamic 1083 ``R_AMDGPU_REL32`` Static 4 ``word32`` S + A - P 1084 ``R_AMDGPU_REL64`` Static 5 ``word64`` S + A - P 1085 ``R_AMDGPU_ABS32`` Static, 6 ``word32`` S + A 1086 Dynamic 1087 ``R_AMDGPU_GOTPCREL`` Static 7 ``word32`` G + GOT + A - P 1088 ``R_AMDGPU_GOTPCREL32_LO`` Static 8 ``word32`` (G + GOT + A - P) & 0xFFFFFFFF 1089 ``R_AMDGPU_GOTPCREL32_HI`` Static 9 ``word32`` (G + GOT + A - P) >> 32 1090 ``R_AMDGPU_REL32_LO`` Static 10 ``word32`` (S + A - P) & 0xFFFFFFFF 1091 ``R_AMDGPU_REL32_HI`` Static 11 ``word32`` (S + A - P) >> 32 1092 *reserved* 12 1093 ``R_AMDGPU_RELATIVE64`` Dynamic 13 ``word64`` B + A 1094 ========================== ======= ===== ========== ============================== 1095 1096``R_AMDGPU_ABS32_LO`` and ``R_AMDGPU_ABS32_HI`` are only supported by 1097the ``mesa3d`` OS, which does not support ``R_AMDGPU_ABS64``. 1098 1099There is no current OS loader support for 32-bit programs and so 1100``R_AMDGPU_ABS32`` is not used. 1101 1102.. _amdgpu-loaded-code-object-path-uniform-resource-identifier: 1103 1104Loaded Code Object Path Uniform Resource Identifier (URI) 1105--------------------------------------------------------- 1106 1107The AMD GPU code object loader represents the path of the ELF shared object from 1108which the code object was loaded as a textual Unifom Resource Identifier (URI). 1109Note that the code object is the in memory loaded relocated form of the ELF 1110shared object. Multiple code objects may be loaded at different memory 1111addresses in the same process from the same ELF shared object. 1112 1113The loaded code object path URI syntax is defined by the following BNF syntax: 1114 1115.. code:: 1116 1117 code_object_uri ::== file_uri | memory_uri 1118 file_uri ::== "file://" file_path [ range_specifier ] 1119 memory_uri ::== "memory://" process_id range_specifier 1120 range_specifier ::== [ "#" | "?" ] "offset=" number "&" "size=" number 1121 file_path ::== URI_ENCODED_OS_FILE_PATH 1122 process_id ::== DECIMAL_NUMBER 1123 number ::== HEX_NUMBER | DECIMAL_NUMBER | OCTAL_NUMBER 1124 1125**number** 1126 Is a C integral literal where hexadecimal values are prefixed by "0x" or "0X", 1127 and octal values by "0". 1128 1129**file_path** 1130 Is the file's path specified as a URI encoded UTF-8 string. In URI encoding, 1131 every character that is not in the regular expression ``[a-zA-Z0-9/_.~-]`` is 1132 encoded as two uppercase hexidecimal digits proceeded by "%". Directories in 1133 the path are separated by "/". 1134 1135**offset** 1136 Is a 0-based byte offset to the start of the code object. For a file URI, it 1137 is from the start of the file specified by the ``file_path``, and if omitted 1138 defaults to 0. For a memory URI, it is the memory address and is required. 1139 1140**size** 1141 Is the number of bytes in the code object. For a file URI, if omitted it 1142 defaults to the size of the file. It is required for a memory URI. 1143 1144**process_id** 1145 Is the identity of the process owning the memory. For Linux it is the C 1146 unsigned integral decimal literal for the process ID (PID). 1147 1148For example: 1149 1150.. code:: 1151 1152 file:///dir1/dir2/file1 1153 file:///dir3/dir4/file2#offset=0x2000&size=3000 1154 memory://1234#offset=0x20000&size=3000 1155 1156.. _amdgpu-dwarf-debug-information: 1157 1158DWARF Debug Information 1159======================= 1160 1161.. warning:: 1162 1163 This section describes **provisional support** for AMDGPU DWARF [DWARF]_ that 1164 is not currently fully implemented and is subject to change. 1165 1166AMDGPU generates DWARF [DWARF]_ debugging information ELF sections (see 1167:ref:`amdgpu-elf-code-object`) which contain information that maps the code 1168object executable code and data to the source language constructs. It can be 1169used by tools such as debuggers and profilers. It uses features defined in 1170:doc:`AMDGPUDwarfExtensionsForHeterogeneousDebugging` that are made available in 1171DWARF Version 4 and DWARF Version 5 as an LLVM vendor extension. 1172 1173This section defines the AMDGPU target architecture specific DWARF mappings. 1174 1175.. _amdgpu-dwarf-register-identifier: 1176 1177Register Identifier 1178------------------- 1179 1180This section defines the AMDGPU target architecture register numbers used in 1181DWARF operation expressions (see DWARF Version 5 section 2.5 and 1182:ref:`amdgpu-dwarf-operation-expressions`) and Call Frame Information 1183instructions (see DWARF Version 5 section 6.4 and 1184:ref:`amdgpu-dwarf-call-frame-information`). 1185 1186A single code object can contain code for kernels that have different wavefront 1187sizes. The vector registers and some scalar registers are based on the wavefront 1188size. AMDGPU defines distinct DWARF registers for each wavefront size. This 1189simplifies the consumer of the DWARF so that each register has a fixed size, 1190rather than being dynamic according to the wavefront size mode. Similarly, 1191distinct DWARF registers are defined for those registers that vary in size 1192according to the process address size. This allows a consumer to treat a 1193specific AMDGPU processor as a single architecture regardless of how it is 1194configured at run time. The compiler explicitly specifies the DWARF registers 1195that match the mode in which the code it is generating will be executed. 1196 1197DWARF registers are encoded as numbers, which are mapped to architecture 1198registers. The mapping for AMDGPU is defined in 1199:ref:`amdgpu-dwarf-register-mapping-table`. All AMDGPU targets use the same 1200mapping. 1201 1202.. table:: AMDGPU DWARF Register Mapping 1203 :name: amdgpu-dwarf-register-mapping-table 1204 1205 ============== ================= ======== ================================== 1206 DWARF Register AMDGPU Register Bit Size Description 1207 ============== ================= ======== ================================== 1208 0 PC_32 32 Program Counter (PC) when 1209 executing in a 32-bit process 1210 address space. Used in the CFI to 1211 describe the PC of the calling 1212 frame. 1213 1 EXEC_MASK_32 32 Execution Mask Register when 1214 executing in wavefront 32 mode. 1215 2-15 *Reserved* *Reserved for highly accessed 1216 registers using DWARF shortcut.* 1217 16 PC_64 64 Program Counter (PC) when 1218 executing in a 64-bit process 1219 address space. Used in the CFI to 1220 describe the PC of the calling 1221 frame. 1222 17 EXEC_MASK_64 64 Execution Mask Register when 1223 executing in wavefront 64 mode. 1224 18-31 *Reserved* *Reserved for highly accessed 1225 registers using DWARF shortcut.* 1226 32-95 SGPR0-SGPR63 32 Scalar General Purpose 1227 Registers. 1228 96-127 *Reserved* *Reserved for frequently accessed 1229 registers using DWARF 1-byte ULEB.* 1230 128 SCC 32 Scalar Condition Code Register. 1231 129-511 *Reserved* *Reserved for future Scalar 1232 Architectural Registers.* 1233 512 VCC_32 32 Vector Condition Code Register 1234 when executing in wavefront 32 1235 mode. 1236 513-1023 *Reserved* *Reserved for future Vector 1237 Architectural Registers when 1238 executing in wavefront 32 mode.* 1239 768 VCC_64 32 Vector Condition Code Register 1240 when executing in wavefront 64 1241 mode. 1242 769-1023 *Reserved* *Reserved for future Vector 1243 Architectural Registers when 1244 executing in wavefront 64 mode.* 1245 1024-1087 *Reserved* *Reserved for padding.* 1246 1088-1129 SGPR64-SGPR105 32 Scalar General Purpose Registers. 1247 1130-1535 *Reserved* *Reserved for future Scalar 1248 General Purpose Registers.* 1249 1536-1791 VGPR0-VGPR255 32*32 Vector General Purpose Registers 1250 when executing in wavefront 32 1251 mode. 1252 1792-2047 *Reserved* *Reserved for future Vector 1253 General Purpose Registers when 1254 executing in wavefront 32 mode.* 1255 2048-2303 AGPR0-AGPR255 32*32 Vector Accumulation Registers 1256 when executing in wavefront 32 1257 mode. 1258 2304-2559 *Reserved* *Reserved for future Vector 1259 Accumulation Registers when 1260 executing in wavefront 32 mode.* 1261 2560-2815 VGPR0-VGPR255 64*32 Vector General Purpose Registers 1262 when executing in wavefront 64 1263 mode. 1264 2816-3071 *Reserved* *Reserved for future Vector 1265 General Purpose Registers when 1266 executing in wavefront 64 mode.* 1267 3072-3327 AGPR0-AGPR255 64*32 Vector Accumulation Registers 1268 when executing in wavefront 64 1269 mode. 1270 3328-3583 *Reserved* *Reserved for future Vector 1271 Accumulation Registers when 1272 executing in wavefront 64 mode.* 1273 ============== ================= ======== ================================== 1274 1275The vector registers are represented as the full size for the wavefront. They 1276are organized as consecutive dwords (32-bits), one per lane, with the dword at 1277the least significant bit position corresponding to lane 0 and so forth. DWARF 1278location expressions involving the ``DW_OP_LLVM_offset`` and 1279``DW_OP_LLVM_push_lane`` operations are used to select the part of the vector 1280register corresponding to the lane that is executing the current thread of 1281execution in languages that are implemented using a SIMD or SIMT execution 1282model. 1283 1284If the wavefront size is 32 lanes then the wavefront 32 mode register 1285definitions are used. If the wavefront size is 64 lanes then the wavefront 64 1286mode register definitions are used. Some AMDGPU targets support executing in 1287both wavefront 32 and wavefront 64 mode. The register definitions corresponding 1288to the wavefront mode of the generated code will be used. 1289 1290If code is generated to execute in a 32-bit process address space, then the 129132-bit process address space register definitions are used. If code is generated 1292to execute in a 64-bit process address space, then the 64-bit process address 1293space register definitions are used. The ``amdgcn`` target only supports the 129464-bit process address space. 1295 1296.. _amdgpu-dwarf-address-class-identifier: 1297 1298Address Class Identifier 1299------------------------ 1300 1301The DWARF address class represents the source language memory space. See DWARF 1302Version 5 section 2.12 which is updated by the *DWARF Extensions For 1303Heterogeneous Debugging* section :ref:`amdgpu-dwarf-segment_addresses`. 1304 1305The DWARF address class mapping used for AMDGPU is defined in 1306:ref:`amdgpu-dwarf-address-class-mapping-table`. 1307 1308.. table:: AMDGPU DWARF Address Class Mapping 1309 :name: amdgpu-dwarf-address-class-mapping-table 1310 1311 ========================= ====== ================= 1312 DWARF AMDGPU 1313 -------------------------------- ----------------- 1314 Address Class Name Value Address Space 1315 ========================= ====== ================= 1316 ``DW_ADDR_none`` 0x0000 Generic (Flat) 1317 ``DW_ADDR_LLVM_global`` 0x0001 Global 1318 ``DW_ADDR_LLVM_constant`` 0x0002 Global 1319 ``DW_ADDR_LLVM_group`` 0x0003 Local (group/LDS) 1320 ``DW_ADDR_LLVM_private`` 0x0004 Private (Scratch) 1321 ``DW_ADDR_AMDGPU_region`` 0x8000 Region (GDS) 1322 ========================= ====== ================= 1323 1324The DWARF address class values defined in the *DWARF Extensions For 1325Heterogeneous Debugging* section :ref:`amdgpu-dwarf-segment_addresses` are used. 1326 1327In addition, ``DW_ADDR_AMDGPU_region`` is encoded as a vendor extension. This is 1328available for use for the AMD extension for access to the hardware GDS memory 1329which is scratchpad memory allocated per device. 1330 1331For AMDGPU if no ``DW_AT_address_class`` attribute is present, then the default 1332address class of ``DW_ADDR_none`` is used. 1333 1334See :ref:`amdgpu-dwarf-address-space-identifier` for information on the AMDGPU 1335mapping of DWARF address classes to DWARF address spaces, including address size 1336and NULL value. 1337 1338.. _amdgpu-dwarf-address-space-identifier: 1339 1340Address Space Identifier 1341------------------------ 1342 1343DWARF address spaces correspond to target architecture specific linear 1344addressable memory areas. See DWARF Version 5 section 2.12 and *DWARF Extensions 1345For Heterogeneous Debugging* section :ref:`amdgpu-dwarf-segment_addresses`. 1346 1347The DWARF address space mapping used for AMDGPU is defined in 1348:ref:`amdgpu-dwarf-address-space-mapping-table`. 1349 1350.. table:: AMDGPU DWARF Address Space Mapping 1351 :name: amdgpu-dwarf-address-space-mapping-table 1352 1353 ======================================= ===== ======= ======== ================= ======================= 1354 DWARF AMDGPU Notes 1355 --------------------------------------- ----- ---------------- ----------------- ----------------------- 1356 Address Space Name Value Address Bit Size Address Space 1357 --------------------------------------- ----- ------- -------- ----------------- ----------------------- 1358 .. 64-bit 32-bit 1359 process process 1360 address address 1361 space space 1362 ======================================= ===== ======= ======== ================= ======================= 1363 ``DW_ASPACE_none`` 0x00 8 4 Global *default address space* 1364 ``DW_ASPACE_AMDGPU_generic`` 0x01 8 4 Generic (Flat) 1365 ``DW_ASPACE_AMDGPU_region`` 0x02 4 4 Region (GDS) 1366 ``DW_ASPACE_AMDGPU_local`` 0x03 4 4 Local (group/LDS) 1367 *Reserved* 0x04 1368 ``DW_ASPACE_AMDGPU_private_lane`` 0x05 4 4 Private (Scratch) *focused lane* 1369 ``DW_ASPACE_AMDGPU_private_wave`` 0x06 4 4 Private (Scratch) *unswizzled wavefront* 1370 ======================================= ===== ======= ======== ================= ======================= 1371 1372See :ref:`amdgpu-address-spaces` for information on the AMDGPU address spaces 1373including address size and NULL value. 1374 1375The ``DW_ASPACE_none`` address space is the default target architecture address 1376space used in DWARF operations that do not specify an address space. It 1377therefore has to map to the global address space so that the ``DW_OP_addr*`` and 1378related operations can refer to addresses in the program code. 1379 1380The ``DW_ASPACE_AMDGPU_generic`` address space allows location expressions to 1381specify the flat address space. If the address corresponds to an address in the 1382local address space, then it corresponds to the wavefront that is executing the 1383focused thread of execution. If the address corresponds to an address in the 1384private address space, then it corresponds to the lane that is executing the 1385focused thread of execution for languages that are implemented using a SIMD or 1386SIMT execution model. 1387 1388.. note:: 1389 1390 CUDA-like languages such as HIP that do not have address spaces in the 1391 language type system, but do allow variables to be allocated in different 1392 address spaces, need to explicitly specify the ``DW_ASPACE_AMDGPU_generic`` 1393 address space in the DWARF expression operations as the default address space 1394 is the global address space. 1395 1396The ``DW_ASPACE_AMDGPU_local`` address space allows location expressions to 1397specify the local address space corresponding to the wavefront that is executing 1398the focused thread of execution. 1399 1400The ``DW_ASPACE_AMDGPU_private_lane`` address space allows location expressions 1401to specify the private address space corresponding to the lane that is executing 1402the focused thread of execution for languages that are implemented using a SIMD 1403or SIMT execution model. 1404 1405The ``DW_ASPACE_AMDGPU_private_wave`` address space allows location expressions 1406to specify the unswizzled private address space corresponding to the wavefront 1407that is executing the focused thread of execution. The wavefront view of private 1408memory is the per wavefront unswizzled backing memory layout defined in 1409:ref:`amdgpu-address-spaces`, such that address 0 corresponds to the first 1410location for the backing memory of the wavefront (namely the address is not 1411offset by ``wavefront-scratch-base``). The following formula can be used to 1412convert from a ``DW_ASPACE_AMDGPU_private_lane`` address to a 1413``DW_ASPACE_AMDGPU_private_wave`` address: 1414 1415:: 1416 1417 private-address-wavefront = 1418 ((private-address-lane / 4) * wavefront-size * 4) + 1419 (wavefront-lane-id * 4) + (private-address-lane % 4) 1420 1421If the ``DW_ASPACE_AMDGPU_private_lane`` address is dword aligned, and the start 1422of the dwords for each lane starting with lane 0 is required, then this 1423simplifies to: 1424 1425:: 1426 1427 private-address-wavefront = 1428 private-address-lane * wavefront-size 1429 1430A compiler can use the ``DW_ASPACE_AMDGPU_private_wave`` address space to read a 1431complete spilled vector register back into a complete vector register in the 1432CFI. The frame pointer can be a private lane address which is dword aligned, 1433which can be shifted to multiply by the wavefront size, and then used to form a 1434private wavefront address that gives a location for a contiguous set of dwords, 1435one per lane, where the vector register dwords are spilled. The compiler knows 1436the wavefront size since it generates the code. Note that the type of the 1437address may have to be converted as the size of a 1438``DW_ASPACE_AMDGPU_private_lane`` address may be smaller than the size of a 1439``DW_ASPACE_AMDGPU_private_wave`` address. 1440 1441.. _amdgpu-dwarf-lane-identifier: 1442 1443Lane identifier 1444--------------- 1445 1446DWARF lane identifies specify a target architecture lane position for hardware 1447that executes in a SIMD or SIMT manner, and on which a source language maps its 1448threads of execution onto those lanes. The DWARF lane identifier is pushed by 1449the ``DW_OP_LLVM_push_lane`` DWARF expression operation. See DWARF Version 5 1450section 2.5 which is updated by *DWARF Extensions For Heterogeneous Debugging* 1451section :ref:`amdgpu-dwarf-operation-expressions`. 1452 1453For AMDGPU, the lane identifier corresponds to the hardware lane ID of a 1454wavefront. It is numbered from 0 to the wavefront size minus 1. 1455 1456Operation Expressions 1457--------------------- 1458 1459DWARF expressions are used to compute program values and the locations of 1460program objects. See DWARF Version 5 section 2.5 and 1461:ref:`amdgpu-dwarf-operation-expressions`. 1462 1463DWARF location descriptions describe how to access storage which includes memory 1464and registers. When accessing storage on AMDGPU, bytes are ordered with least 1465significant bytes first, and bits are ordered within bytes with least 1466significant bits first. 1467 1468For AMDGPU CFI expressions, ``DW_OP_LLVM_select_bit_piece`` is used to describe 1469unwinding vector registers that are spilled under the execution mask to memory: 1470the zero-single location description is the vector register, and the one-single 1471location description is the spilled memory location description. The 1472``DW_OP_LLVM_form_aspace_address`` is used to specify the address space of the 1473memory location description. 1474 1475In AMDGPU expressions, ``DW_OP_LLVM_select_bit_piece`` is used by the 1476``DW_AT_LLVM_lane_pc`` attribute expression where divergent control flow is 1477controlled by the execution mask. An undefined location description together 1478with ``DW_OP_LLVM_extend`` is used to indicate the lane was not active on entry 1479to the subprogram. See :ref:`amdgpu-dwarf-dw-at-llvm-lane-pc` for an example. 1480 1481Debugger Information Entry Attributes 1482------------------------------------- 1483 1484This section describes how certain debugger information entry attributes are 1485used by AMDGPU. See the sections in DWARF Version 5 section 2 which are updated 1486by *DWARF Extensions For Heterogeneous Debugging* section 1487:ref:`amdgpu-dwarf-debugging-information-entry-attributes`. 1488 1489.. _amdgpu-dwarf-dw-at-llvm-lane-pc: 1490 1491``DW_AT_LLVM_lane_pc`` 1492~~~~~~~~~~~~~~~~~~~~~~ 1493 1494For AMDGPU, the ``DW_AT_LLVM_lane_pc`` attribute is used to specify the program 1495location of the separate lanes of a SIMT thread. 1496 1497If the lane is an active lane then this will be the same as the current program 1498location. 1499 1500If the lane is inactive, but was active on entry to the subprogram, then this is 1501the program location in the subprogram at which execution of the lane is 1502conceptual positioned. 1503 1504If the lane was not active on entry to the subprogram, then this will be the 1505undefined location. A client debugger can check if the lane is part of a valid 1506work-group by checking that the lane is in the range of the associated 1507work-group within the grid, accounting for partial work-groups. If it is not, 1508then the debugger can omit any information for the lane. Otherwise, the debugger 1509may repeatedly unwind the stack and inspect the ``DW_AT_LLVM_lane_pc`` of the 1510calling subprogram until it finds a non-undefined location. Conceptually the 1511lane only has the call frames that it has a non-undefined 1512``DW_AT_LLVM_lane_pc``. 1513 1514The following example illustrates how the AMDGPU backend can generate a DWARF 1515location list expression for the nested ``IF/THEN/ELSE`` structures of the 1516following subprogram pseudo code for a target with 64 lanes per wavefront. 1517 1518.. code:: 1519 :number-lines: 1520 1521 SUBPROGRAM X 1522 BEGIN 1523 a; 1524 IF (c1) THEN 1525 b; 1526 IF (c2) THEN 1527 c; 1528 ELSE 1529 d; 1530 ENDIF 1531 e; 1532 ELSE 1533 f; 1534 ENDIF 1535 g; 1536 END 1537 1538The AMDGPU backend may generate the following pseudo LLVM MIR to manipulate the 1539execution mask (``EXEC``) to linearize the control flow. The condition is 1540evaluated to make a mask of the lanes for which the condition evaluates to true. 1541First the ``THEN`` region is executed by setting the ``EXEC`` mask to the 1542logical ``AND`` of the current ``EXEC`` mask with the condition mask. Then the 1543``ELSE`` region is executed by negating the ``EXEC`` mask and logical ``AND`` of 1544the saved ``EXEC`` mask at the start of the region. After the ``IF/THEN/ELSE`` 1545region the ``EXEC`` mask is restored to the value it had at the beginning of the 1546region. This is shown below. Other approaches are possible, but the basic 1547concept is the same. 1548 1549.. code:: 1550 :number-lines: 1551 1552 $lex_start: 1553 a; 1554 %1 = EXEC 1555 %2 = c1 1556 $lex_1_start: 1557 EXEC = %1 & %2 1558 $if_1_then: 1559 b; 1560 %3 = EXEC 1561 %4 = c2 1562 $lex_1_1_start: 1563 EXEC = %3 & %4 1564 $lex_1_1_then: 1565 c; 1566 EXEC = ~EXEC & %3 1567 $lex_1_1_else: 1568 d; 1569 EXEC = %3 1570 $lex_1_1_end: 1571 e; 1572 EXEC = ~EXEC & %1 1573 $lex_1_else: 1574 f; 1575 EXEC = %1 1576 $lex_1_end: 1577 g; 1578 $lex_end: 1579 1580To create the DWARF location list expression that defines the location 1581description of a vector of lane program locations, the LLVM MIR ``DBG_VALUE`` 1582pseudo instruction can be used to annotate the linearized control flow. This can 1583be done by defining an artificial variable for the lane PC. The DWARF location 1584list expression created for it is used as the value of the 1585``DW_AT_LLVM_lane_pc`` attribute on the subprogram's debugger information entry. 1586 1587A DWARF procedure is defined for each well nested structured control flow region 1588which provides the conceptual lane program location for a lane if it is not 1589active (namely it is divergent). The DWARF operation expression for each region 1590conceptually inherits the value of the immediately enclosing region and modifies 1591it according to the semantics of the region. 1592 1593For an ``IF/THEN/ELSE`` region the divergent program location is at the start of 1594the region for the ``THEN`` region since it is executed first. For the ``ELSE`` 1595region the divergent program location is at the end of the ``IF/THEN/ELSE`` 1596region since the ``THEN`` region has completed. 1597 1598The lane PC artificial variable is assigned at each region transition. It uses 1599the immediately enclosing region's DWARF procedure to compute the program 1600location for each lane assuming they are divergent, and then modifies the result 1601by inserting the current program location for each lane that the ``EXEC`` mask 1602indicates is active. 1603 1604By having separate DWARF procedures for each region, they can be reused to 1605define the value for any nested region. This reduces the total size of the DWARF 1606operation expressions. 1607 1608The following provides an example using pseudo LLVM MIR. 1609 1610.. code:: 1611 :number-lines: 1612 1613 $lex_start: 1614 DEFINE_DWARF %__uint_64 = DW_TAG_base_type[ 1615 DW_AT_name = "__uint64"; 1616 DW_AT_byte_size = 8; 1617 DW_AT_encoding = DW_ATE_unsigned; 1618 ]; 1619 DEFINE_DWARF %__active_lane_pc = DW_TAG_dwarf_procedure[ 1620 DW_AT_name = "__active_lane_pc"; 1621 DW_AT_location = [ 1622 DW_OP_regx PC; 1623 DW_OP_LLVM_extend 64, 64; 1624 DW_OP_regval_type EXEC, %uint_64; 1625 DW_OP_LLVM_select_bit_piece 64, 64; 1626 ]; 1627 ]; 1628 DEFINE_DWARF %__divergent_lane_pc = DW_TAG_dwarf_procedure[ 1629 DW_AT_name = "__divergent_lane_pc"; 1630 DW_AT_location = [ 1631 DW_OP_LLVM_undefined; 1632 DW_OP_LLVM_extend 64, 64; 1633 ]; 1634 ]; 1635 DBG_VALUE $noreg, $noreg, %DW_AT_LLVM_lane_pc, DIExpression[ 1636 DW_OP_call_ref %__divergent_lane_pc; 1637 DW_OP_call_ref %__active_lane_pc; 1638 ]; 1639 a; 1640 %1 = EXEC; 1641 DBG_VALUE %1, $noreg, %__lex_1_save_exec; 1642 %2 = c1; 1643 $lex_1_start: 1644 EXEC = %1 & %2; 1645 $lex_1_then: 1646 DEFINE_DWARF %__divergent_lane_pc_1_then = DW_TAG_dwarf_procedure[ 1647 DW_AT_name = "__divergent_lane_pc_1_then"; 1648 DW_AT_location = DIExpression[ 1649 DW_OP_call_ref %__divergent_lane_pc; 1650 DW_OP_addrx &lex_1_start; 1651 DW_OP_stack_value; 1652 DW_OP_LLVM_extend 64, 64; 1653 DW_OP_call_ref %__lex_1_save_exec; 1654 DW_OP_deref_type 64, %__uint_64; 1655 DW_OP_LLVM_select_bit_piece 64, 64; 1656 ]; 1657 ]; 1658 DBG_VALUE $noreg, $noreg, %DW_AT_LLVM_lane_pc, DIExpression[ 1659 DW_OP_call_ref %__divergent_lane_pc_1_then; 1660 DW_OP_call_ref %__active_lane_pc; 1661 ]; 1662 b; 1663 %3 = EXEC; 1664 DBG_VALUE %3, %__lex_1_1_save_exec; 1665 %4 = c2; 1666 $lex_1_1_start: 1667 EXEC = %3 & %4; 1668 $lex_1_1_then: 1669 DEFINE_DWARF %__divergent_lane_pc_1_1_then = DW_TAG_dwarf_procedure[ 1670 DW_AT_name = "__divergent_lane_pc_1_1_then"; 1671 DW_AT_location = DIExpression[ 1672 DW_OP_call_ref %__divergent_lane_pc_1_then; 1673 DW_OP_addrx &lex_1_1_start; 1674 DW_OP_stack_value; 1675 DW_OP_LLVM_extend 64, 64; 1676 DW_OP_call_ref %__lex_1_1_save_exec; 1677 DW_OP_deref_type 64, %__uint_64; 1678 DW_OP_LLVM_select_bit_piece 64, 64; 1679 ]; 1680 ]; 1681 DBG_VALUE $noreg, $noreg, %DW_AT_LLVM_lane_pc, DIExpression[ 1682 DW_OP_call_ref %__divergent_lane_pc_1_1_then; 1683 DW_OP_call_ref %__active_lane_pc; 1684 ]; 1685 c; 1686 EXEC = ~EXEC & %3; 1687 $lex_1_1_else: 1688 DEFINE_DWARF %__divergent_lane_pc_1_1_else = DW_TAG_dwarf_procedure[ 1689 DW_AT_name = "__divergent_lane_pc_1_1_else"; 1690 DW_AT_location = DIExpression[ 1691 DW_OP_call_ref %__divergent_lane_pc_1_then; 1692 DW_OP_addrx &lex_1_1_end; 1693 DW_OP_stack_value; 1694 DW_OP_LLVM_extend 64, 64; 1695 DW_OP_call_ref %__lex_1_1_save_exec; 1696 DW_OP_deref_type 64, %__uint_64; 1697 DW_OP_LLVM_select_bit_piece 64, 64; 1698 ]; 1699 ]; 1700 DBG_VALUE $noreg, $noreg, %DW_AT_LLVM_lane_pc, DIExpression[ 1701 DW_OP_call_ref %__divergent_lane_pc_1_1_else; 1702 DW_OP_call_ref %__active_lane_pc; 1703 ]; 1704 d; 1705 EXEC = %3; 1706 $lex_1_1_end: 1707 DBG_VALUE $noreg, $noreg, %DW_AT_LLVM_lane_pc, DIExpression[ 1708 DW_OP_call_ref %__divergent_lane_pc; 1709 DW_OP_call_ref %__active_lane_pc; 1710 ]; 1711 e; 1712 EXEC = ~EXEC & %1; 1713 $lex_1_else: 1714 DEFINE_DWARF %__divergent_lane_pc_1_else = DW_TAG_dwarf_procedure[ 1715 DW_AT_name = "__divergent_lane_pc_1_else"; 1716 DW_AT_location = DIExpression[ 1717 DW_OP_call_ref %__divergent_lane_pc; 1718 DW_OP_addrx &lex_1_end; 1719 DW_OP_stack_value; 1720 DW_OP_LLVM_extend 64, 64; 1721 DW_OP_call_ref %__lex_1_save_exec; 1722 DW_OP_deref_type 64, %__uint_64; 1723 DW_OP_LLVM_select_bit_piece 64, 64; 1724 ]; 1725 ]; 1726 DBG_VALUE $noreg, $noreg, %DW_AT_LLVM_lane_pc, DIExpression[ 1727 DW_OP_call_ref %__divergent_lane_pc_1_else; 1728 DW_OP_call_ref %__active_lane_pc; 1729 ]; 1730 f; 1731 EXEC = %1; 1732 $lex_1_end: 1733 DBG_VALUE $noreg, $noreg, %DW_AT_LLVM_lane_pc DIExpression[ 1734 DW_OP_call_ref %__divergent_lane_pc; 1735 DW_OP_call_ref %__active_lane_pc; 1736 ]; 1737 g; 1738 $lex_end: 1739 1740The DWARF procedure ``%__active_lane_pc`` is used to update the lane pc elements 1741that are active, with the current program location. 1742 1743Artificial variables %__lex_1_save_exec and %__lex_1_1_save_exec are created for 1744the execution masks saved on entry to a region. Using the ``DBG_VALUE`` pseudo 1745instruction, location list entries will be created that describe where the 1746artificial variables are allocated at any given program location. The compiler 1747may allocate them to registers or spill them to memory. 1748 1749The DWARF procedures for each region use the values of the saved execution mask 1750artificial variables to only update the lanes that are active on entry to the 1751region. All other lanes retain the value of the enclosing region where they were 1752last active. If they were not active on entry to the subprogram, then will have 1753the undefined location description. 1754 1755Other structured control flow regions can be handled similarly. For example, 1756loops would set the divergent program location for the region at the end of the 1757loop. Any lanes active will be in the loop, and any lanes not active must have 1758exited the loop. 1759 1760An ``IF/THEN/ELSEIF/ELSEIF/...`` region can be treated as a nest of 1761``IF/THEN/ELSE`` regions. 1762 1763The DWARF procedures can use the active lane artificial variable described in 1764:ref:`amdgpu-dwarf-amdgpu-dw-at-llvm-active-lane` rather than the actual 1765``EXEC`` mask in order to support whole or quad wavefront mode. 1766 1767.. _amdgpu-dwarf-amdgpu-dw-at-llvm-active-lane: 1768 1769``DW_AT_LLVM_active_lane`` 1770~~~~~~~~~~~~~~~~~~~~~~~~~~ 1771 1772The ``DW_AT_LLVM_active_lane`` attribute on a subprogram debugger information 1773entry is used to specify the lanes that are conceptually active for a SIMT 1774thread. 1775 1776The execution mask may be modified to implement whole or quad wavefront mode 1777operations. For example, all lanes may need to temporarily be made active to 1778execute a whole wavefront operation. Such regions would save the ``EXEC`` mask, 1779update it to enable the necessary lanes, perform the operations, and then 1780restore the ``EXEC`` mask from the saved value. While executing the whole 1781wavefront region, the conceptual execution mask is the saved value, not the 1782``EXEC`` value. 1783 1784This is handled by defining an artificial variable for the active lane mask. The 1785active lane mask artificial variable would be the actual ``EXEC`` mask for 1786normal regions, and the saved execution mask for regions where the mask is 1787temporarily updated. The location list expression created for this artificial 1788variable is used to define the value of the ``DW_AT_LLVM_active_lane`` 1789attribute. 1790 1791``DW_AT_LLVM_augmentation`` 1792~~~~~~~~~~~~~~~~~~~~~~~~~~~ 1793 1794For AMDGPU, the ``DW_AT_LLVM_augmentation`` attribute of a compilation unit 1795debugger information entry has the following value for the augmentation string: 1796 1797:: 1798 1799 [amdgpu:v0.0] 1800 1801The "vX.Y" specifies the major X and minor Y version number of the AMDGPU 1802extensions used in the DWARF of the compilation unit. The version number 1803conforms to [SEMVER]_. 1804 1805Call Frame Information 1806---------------------- 1807 1808DWARF Call Frame Information (CFI) describes how a consumer can virtually 1809*unwind* call frames in a running process or core dump. See DWARF Version 5 1810section 6.4 and :ref:`amdgpu-dwarf-call-frame-information`. 1811 1812For AMDGPU, the Common Information Entry (CIE) fields have the following values: 1813 18141. ``augmentation`` string contains the following null-terminated UTF-8 string: 1815 1816 :: 1817 1818 [amd:v0.0] 1819 1820 The ``vX.Y`` specifies the major X and minor Y version number of the AMDGPU 1821 extensions used in this CIE or to the FDEs that use it. The version number 1822 conforms to [SEMVER]_. 1823 18242. ``address_size`` for the ``Global`` address space is defined in 1825 :ref:`amdgpu-dwarf-address-space-identifier`. 1826 18273. ``segment_selector_size`` is 0 as AMDGPU does not use a segment selector. 1828 18294. ``code_alignment_factor`` is 4 bytes. 1830 1831 .. TODO:: 1832 1833 Add to :ref:`amdgpu-processor-table` table. 1834 18355. ``data_alignment_factor`` is 4 bytes. 1836 1837 .. TODO:: 1838 1839 Add to :ref:`amdgpu-processor-table` table. 1840 18416. ``return_address_register`` is ``PC_32`` for 32-bit processes and ``PC_64`` 1842 for 64-bit processes defined in :ref:`amdgpu-dwarf-register-identifier`. 1843 18447. ``initial_instructions`` Since a subprogram X with fewer registers can be 1845 called from subprogram Y that has more allocated, X will not change any of 1846 the extra registers as it cannot access them. Therefore, the default rule 1847 for all columns is ``same value``. 1848 1849For AMDGPU the register number follows the numbering defined in 1850:ref:`amdgpu-dwarf-register-identifier`. 1851 1852For AMDGPU the instructions are variable size. A consumer can subtract 1 from 1853the return address to get the address of a byte within the call site 1854instructions. See DWARF Version 5 section 6.4.4. 1855 1856Accelerated Access 1857------------------ 1858 1859See DWARF Version 5 section 6.1. 1860 1861Lookup By Name Section Header 1862~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 1863 1864See DWARF Version 5 section 6.1.1.4.1 and :ref:`amdgpu-dwarf-lookup-by-name`. 1865 1866For AMDGPU the lookup by name section header table: 1867 1868``augmentation_string_size`` (uword) 1869 1870 Set to the length of the ``augmentation_string`` value which is always a 1871 multiple of 4. 1872 1873``augmentation_string`` (sequence of UTF-8 characters) 1874 1875 Contains the following UTF-8 string null padded to a multiple of 4 bytes: 1876 1877 :: 1878 1879 [amdgpu:v0.0] 1880 1881 The "vX.Y" specifies the major X and minor Y version number of the AMDGPU 1882 extensions used in the DWARF of this index. The version number conforms to 1883 [SEMVER]_. 1884 1885 .. note:: 1886 1887 This is different to the DWARF Version 5 definition that requires the first 1888 4 characters to be the vendor ID. But this is consistent with the other 1889 augmentation strings and does allow multiple vendor contributions. However, 1890 backwards compatibility may be more desirable. 1891 1892Lookup By Address Section Header 1893~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 1894 1895See DWARF Version 5 section 6.1.2. 1896 1897For AMDGPU the lookup by address section header table: 1898 1899``address_size`` (ubyte) 1900 1901 Match the address size for the ``Global`` address space defined in 1902 :ref:`amdgpu-dwarf-address-space-identifier`. 1903 1904``segment_selector_size`` (ubyte) 1905 1906 AMDGPU does not use a segment selector so this is 0. The entries in the 1907 ``.debug_aranges`` do not have a segment selector. 1908 1909Line Number Information 1910----------------------- 1911 1912See DWARF Version 5 section 6.2 and :ref:`amdgpu-dwarf-line-number-information`. 1913 1914AMDGPU does not use the ``isa`` state machine registers and always sets it to 0. 1915The instruction set must be obtained from the ELF file header ``e_flags`` field 1916in the ``EF_AMDGPU_MACH`` bit position (see :ref:`ELF Header 1917<amdgpu-elf-header>`). See DWARF Version 5 section 6.2.2. 1918 1919.. TODO:: 1920 1921 Should the ``isa`` state machine register be used to indicate if the code is 1922 in wavefront32 or wavefront64 mode? Or used to specify the architecture ISA? 1923 1924For AMDGPU the line number program header fields have the following values (see 1925DWARF Version 5 section 6.2.4): 1926 1927``address_size`` (ubyte) 1928 Matches the address size for the ``Global`` address space defined in 1929 :ref:`amdgpu-dwarf-address-space-identifier`. 1930 1931``segment_selector_size`` (ubyte) 1932 AMDGPU does not use a segment selector so this is 0. 1933 1934``minimum_instruction_length`` (ubyte) 1935 For GFX9-GFX10 this is 4. 1936 1937``maximum_operations_per_instruction`` (ubyte) 1938 For GFX9-GFX10 this is 1. 1939 1940Source text for online-compiled programs (for example, those compiled by the 1941OpenCL language runtime) may be embedded into the DWARF Version 5 line table. 1942See DWARF Version 5 section 6.2.4.1 which is updated by *DWARF Extensions For 1943Heterogeneous Debugging* section :ref:`DW_LNCT_LLVM_source 1944<amdgpu-dwarf-line-number-information-dw-lnct-llvm-source>`. 1945 1946The Clang option used to control source embedding in AMDGPU is defined in 1947:ref:`amdgpu-clang-debug-options-table`. 1948 1949 .. table:: AMDGPU Clang Debug Options 1950 :name: amdgpu-clang-debug-options-table 1951 1952 ==================== ================================================== 1953 Debug Flag Description 1954 ==================== ================================================== 1955 -g[no-]embed-source Enable/disable embedding source text in DWARF 1956 debug sections. Useful for environments where 1957 source cannot be written to disk, such as 1958 when performing online compilation. 1959 ==================== ================================================== 1960 1961For example: 1962 1963``-gembed-source`` 1964 Enable the embedded source. 1965 1966``-gno-embed-source`` 1967 Disable the embedded source. 1968 196932-Bit and 64-Bit DWARF Formats 1970------------------------------- 1971 1972See DWARF Version 5 section 7.4 and 1973:ref:`amdgpu-dwarf-32-bit-and-64-bit-dwarf-formats`. 1974 1975For AMDGPU: 1976 1977* For the ``amdgcn`` target architecture only the 64-bit process address space 1978 is supported. 1979 1980* The producer can generate either 32-bit or 64-bit DWARF format. LLVM generates 1981 the 32-bit DWARF format. 1982 1983Unit Headers 1984------------ 1985 1986For AMDGPU the following values apply for each of the unit headers described in 1987DWARF Version 5 sections 7.5.1.1, 7.5.1.2, and 7.5.1.3: 1988 1989``address_size`` (ubyte) 1990 Matches the address size for the ``Global`` address space defined in 1991 :ref:`amdgpu-dwarf-address-space-identifier`. 1992 1993.. _amdgpu-code-conventions: 1994 1995Code Conventions 1996================ 1997 1998This section provides code conventions used for each supported target triple OS 1999(see :ref:`amdgpu-target-triples`). 2000 2001AMDHSA 2002------ 2003 2004This section provides code conventions used when the target triple OS is 2005``amdhsa`` (see :ref:`amdgpu-target-triples`). 2006 2007.. _amdgpu-amdhsa-code-object-target-identification: 2008 2009Code Object Target Identification 2010~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 2011 2012The AMDHSA OS uses the following syntax to specify the code object 2013target as a single string: 2014 2015 ``<Architecture>-<Vendor>-<OS>-<Environment>-<Processor><Target Features>`` 2016 2017Where: 2018 2019 - ``<Architecture>``, ``<Vendor>``, ``<OS>`` and ``<Environment>`` 2020 are the same as the *Target Triple* (see 2021 :ref:`amdgpu-target-triples`). 2022 2023 - ``<Processor>`` is the same as the *Processor* (see 2024 :ref:`amdgpu-processors`). 2025 2026 - ``<Target Features>`` is a list of the enabled *Target Features* 2027 (see :ref:`amdgpu-target-features`), each prefixed by a plus, that 2028 apply to *Processor*. The list must be in the same order as listed 2029 in the table :ref:`amdgpu-target-feature-table`. Note that *Target 2030 Features* must be included in the list if they are enabled even if 2031 that is the default for *Processor*. 2032 2033For example: 2034 2035 ``"amdgcn-amd-amdhsa--gfx902+xnack"`` 2036 2037.. _amdgpu-amdhsa-code-object-metadata: 2038 2039Code Object Metadata 2040~~~~~~~~~~~~~~~~~~~~ 2041 2042The code object metadata specifies extensible metadata associated with the code 2043objects executed on HSA [HSA]_ compatible runtimes such as AMD's ROCm 2044[AMD-ROCm]_. The encoding and semantics of this metadata depends on the code 2045object version; see :ref:`amdgpu-amdhsa-code-object-metadata-v2` and 2046:ref:`amdgpu-amdhsa-code-object-metadata-v3`. 2047 2048Code object metadata is specified in a note record (see 2049:ref:`amdgpu-note-records`) and is required when the target triple OS is 2050``amdhsa`` (see :ref:`amdgpu-target-triples`). It must contain the minimum 2051information necessary to support the ROCM kernel queries. For example, the 2052segment sizes needed in a dispatch packet. In addition, a high-level language 2053runtime may require other information to be included. For example, the AMD 2054OpenCL runtime records kernel argument information. 2055 2056.. _amdgpu-amdhsa-code-object-metadata-v2: 2057 2058Code Object V2 Metadata (-mattr=-code-object-v3) 2059++++++++++++++++++++++++++++++++++++++++++++++++ 2060 2061.. warning:: Code Object V2 is not the default code object version emitted by 2062 this version of LLVM. For a description of the metadata generated with the 2063 default configuration (Code Object V3) see 2064 :ref:`amdgpu-amdhsa-code-object-metadata-v3`. 2065 2066Code object V2 metadata is specified by the ``NT_AMD_AMDGPU_METADATA`` note 2067record (see :ref:`amdgpu-note-records-v2`). 2068 2069The metadata is specified as a YAML formatted string (see [YAML]_ and 2070:doc:`YamlIO`). 2071 2072.. TODO:: 2073 2074 Is the string null terminated? It probably should not if YAML allows it to 2075 contain null characters, otherwise it should be. 2076 2077The metadata is represented as a single YAML document comprised of the mapping 2078defined in table :ref:`amdgpu-amdhsa-code-object-metadata-map-table-v2` and 2079referenced tables. 2080 2081For boolean values, the string values of ``false`` and ``true`` are used for 2082false and true respectively. 2083 2084Additional information can be added to the mappings. To avoid conflicts, any 2085non-AMD key names should be prefixed by "*vendor-name*.". 2086 2087 .. table:: AMDHSA Code Object V2 Metadata Map 2088 :name: amdgpu-amdhsa-code-object-metadata-map-table-v2 2089 2090 ========== ============== ========= ======================================= 2091 String Key Value Type Required? Description 2092 ========== ============== ========= ======================================= 2093 "Version" sequence of Required - The first integer is the major 2094 2 integers version. Currently 1. 2095 - The second integer is the minor 2096 version. Currently 0. 2097 "Printf" sequence of Each string is encoded information 2098 strings about a printf function call. The 2099 encoded information is organized as 2100 fields separated by colon (':'): 2101 2102 ``ID:N:S[0]:S[1]:...:S[N-1]:FormatString`` 2103 2104 where: 2105 2106 ``ID`` 2107 A 32-bit integer as a unique id for 2108 each printf function call 2109 2110 ``N`` 2111 A 32-bit integer equal to the number 2112 of arguments of printf function call 2113 minus 1 2114 2115 ``S[i]`` (where i = 0, 1, ... , N-1) 2116 32-bit integers for the size in bytes 2117 of the i-th FormatString argument of 2118 the printf function call 2119 2120 FormatString 2121 The format string passed to the 2122 printf function call. 2123 "Kernels" sequence of Required Sequence of the mappings for each 2124 mapping kernel in the code object. See 2125 :ref:`amdgpu-amdhsa-code-object-kernel-metadata-map-table-v2` 2126 for the definition of the mapping. 2127 ========== ============== ========= ======================================= 2128 2129.. 2130 2131 .. table:: AMDHSA Code Object V2 Kernel Metadata Map 2132 :name: amdgpu-amdhsa-code-object-kernel-metadata-map-table-v2 2133 2134 ================= ============== ========= ================================ 2135 String Key Value Type Required? Description 2136 ================= ============== ========= ================================ 2137 "Name" string Required Source name of the kernel. 2138 "SymbolName" string Required Name of the kernel 2139 descriptor ELF symbol. 2140 "Language" string Source language of the kernel. 2141 Values include: 2142 2143 - "OpenCL C" 2144 - "OpenCL C++" 2145 - "HCC" 2146 - "OpenMP" 2147 2148 "LanguageVersion" sequence of - The first integer is the major 2149 2 integers version. 2150 - The second integer is the 2151 minor version. 2152 "Attrs" mapping Mapping of kernel attributes. 2153 See 2154 :ref:`amdgpu-amdhsa-code-object-kernel-attribute-metadata-map-table-v2` 2155 for the mapping definition. 2156 "Args" sequence of Sequence of mappings of the 2157 mapping kernel arguments. See 2158 :ref:`amdgpu-amdhsa-code-object-kernel-argument-metadata-map-table-v2` 2159 for the definition of the mapping. 2160 "CodeProps" mapping Mapping of properties related to 2161 the kernel code. See 2162 :ref:`amdgpu-amdhsa-code-object-kernel-code-properties-metadata-map-table-v2` 2163 for the mapping definition. 2164 ================= ============== ========= ================================ 2165 2166.. 2167 2168 .. table:: AMDHSA Code Object V2 Kernel Attribute Metadata Map 2169 :name: amdgpu-amdhsa-code-object-kernel-attribute-metadata-map-table-v2 2170 2171 =================== ============== ========= ============================== 2172 String Key Value Type Required? Description 2173 =================== ============== ========= ============================== 2174 "ReqdWorkGroupSize" sequence of If not 0, 0, 0 then all values 2175 3 integers must be >=1 and the dispatch 2176 work-group size X, Y, Z must 2177 correspond to the specified 2178 values. Defaults to 0, 0, 0. 2179 2180 Corresponds to the OpenCL 2181 ``reqd_work_group_size`` 2182 attribute. 2183 "WorkGroupSizeHint" sequence of The dispatch work-group size 2184 3 integers X, Y, Z is likely to be the 2185 specified values. 2186 2187 Corresponds to the OpenCL 2188 ``work_group_size_hint`` 2189 attribute. 2190 "VecTypeHint" string The name of a scalar or vector 2191 type. 2192 2193 Corresponds to the OpenCL 2194 ``vec_type_hint`` attribute. 2195 2196 "RuntimeHandle" string The external symbol name 2197 associated with a kernel. 2198 OpenCL runtime allocates a 2199 global buffer for the symbol 2200 and saves the kernel's address 2201 to it, which is used for 2202 device side enqueueing. Only 2203 available for device side 2204 enqueued kernels. 2205 =================== ============== ========= ============================== 2206 2207.. 2208 2209 .. table:: AMDHSA Code Object V2 Kernel Argument Metadata Map 2210 :name: amdgpu-amdhsa-code-object-kernel-argument-metadata-map-table-v2 2211 2212 ================= ============== ========= ================================ 2213 String Key Value Type Required? Description 2214 ================= ============== ========= ================================ 2215 "Name" string Kernel argument name. 2216 "TypeName" string Kernel argument type name. 2217 "Size" integer Required Kernel argument size in bytes. 2218 "Align" integer Required Kernel argument alignment in 2219 bytes. Must be a power of two. 2220 "ValueKind" string Required Kernel argument kind that 2221 specifies how to set up the 2222 corresponding argument. 2223 Values include: 2224 2225 "ByValue" 2226 The argument is copied 2227 directly into the kernarg. 2228 2229 "GlobalBuffer" 2230 A global address space pointer 2231 to the buffer data is passed 2232 in the kernarg. 2233 2234 "DynamicSharedPointer" 2235 A group address space pointer 2236 to dynamically allocated LDS 2237 is passed in the kernarg. 2238 2239 "Sampler" 2240 A global address space 2241 pointer to a S# is passed in 2242 the kernarg. 2243 2244 "Image" 2245 A global address space 2246 pointer to a T# is passed in 2247 the kernarg. 2248 2249 "Pipe" 2250 A global address space pointer 2251 to an OpenCL pipe is passed in 2252 the kernarg. 2253 2254 "Queue" 2255 A global address space pointer 2256 to an OpenCL device enqueue 2257 queue is passed in the 2258 kernarg. 2259 2260 "HiddenGlobalOffsetX" 2261 The OpenCL grid dispatch 2262 global offset for the X 2263 dimension is passed in the 2264 kernarg. 2265 2266 "HiddenGlobalOffsetY" 2267 The OpenCL grid dispatch 2268 global offset for the Y 2269 dimension is passed in the 2270 kernarg. 2271 2272 "HiddenGlobalOffsetZ" 2273 The OpenCL grid dispatch 2274 global offset for the Z 2275 dimension is passed in the 2276 kernarg. 2277 2278 "HiddenNone" 2279 An argument that is not used 2280 by the kernel. Space needs to 2281 be left for it, but it does 2282 not need to be set up. 2283 2284 "HiddenPrintfBuffer" 2285 A global address space pointer 2286 to the runtime printf buffer 2287 is passed in kernarg. 2288 2289 "HiddenHostcallBuffer" 2290 A global address space pointer 2291 to the runtime hostcall buffer 2292 is passed in kernarg. 2293 2294 "HiddenDefaultQueue" 2295 A global address space pointer 2296 to the OpenCL device enqueue 2297 queue that should be used by 2298 the kernel by default is 2299 passed in the kernarg. 2300 2301 "HiddenCompletionAction" 2302 A global address space pointer 2303 to help link enqueued kernels into 2304 the ancestor tree for determining 2305 when the parent kernel has finished. 2306 2307 "HiddenMultiGridSyncArg" 2308 A global address space pointer for 2309 multi-grid synchronization is 2310 passed in the kernarg. 2311 2312 "ValueType" string Unused and deprecated. This should no longer 2313 be emitted, but is accepted for compatibility. 2314 2315 2316 "PointeeAlign" integer Alignment in bytes of pointee 2317 type for pointer type kernel 2318 argument. Must be a power 2319 of 2. Only present if 2320 "ValueKind" is 2321 "DynamicSharedPointer". 2322 "AddrSpaceQual" string Kernel argument address space 2323 qualifier. Only present if 2324 "ValueKind" is "GlobalBuffer" or 2325 "DynamicSharedPointer". Values 2326 are: 2327 2328 - "Private" 2329 - "Global" 2330 - "Constant" 2331 - "Local" 2332 - "Generic" 2333 - "Region" 2334 2335 .. TODO:: 2336 Is GlobalBuffer only Global 2337 or Constant? Is 2338 DynamicSharedPointer always 2339 Local? Can HCC allow Generic? 2340 How can Private or Region 2341 ever happen? 2342 "AccQual" string Kernel argument access 2343 qualifier. Only present if 2344 "ValueKind" is "Image" or 2345 "Pipe". Values 2346 are: 2347 2348 - "ReadOnly" 2349 - "WriteOnly" 2350 - "ReadWrite" 2351 2352 .. TODO:: 2353 Does this apply to 2354 GlobalBuffer? 2355 "ActualAccQual" string The actual memory accesses 2356 performed by the kernel on the 2357 kernel argument. Only present if 2358 "ValueKind" is "GlobalBuffer", 2359 "Image", or "Pipe". This may be 2360 more restrictive than indicated 2361 by "AccQual" to reflect what the 2362 kernel actual does. If not 2363 present then the runtime must 2364 assume what is implied by 2365 "AccQual" and "IsConst". Values 2366 are: 2367 2368 - "ReadOnly" 2369 - "WriteOnly" 2370 - "ReadWrite" 2371 2372 "IsConst" boolean Indicates if the kernel argument 2373 is const qualified. Only present 2374 if "ValueKind" is 2375 "GlobalBuffer". 2376 2377 "IsRestrict" boolean Indicates if the kernel argument 2378 is restrict qualified. Only 2379 present if "ValueKind" is 2380 "GlobalBuffer". 2381 2382 "IsVolatile" boolean Indicates if the kernel argument 2383 is volatile qualified. Only 2384 present if "ValueKind" is 2385 "GlobalBuffer". 2386 2387 "IsPipe" boolean Indicates if the kernel argument 2388 is pipe qualified. Only present 2389 if "ValueKind" is "Pipe". 2390 2391 .. TODO:: 2392 Can GlobalBuffer be pipe 2393 qualified? 2394 ================= ============== ========= ================================ 2395 2396.. 2397 2398 .. table:: AMDHSA Code Object V2 Kernel Code Properties Metadata Map 2399 :name: amdgpu-amdhsa-code-object-kernel-code-properties-metadata-map-table-v2 2400 2401 ============================ ============== ========= ===================== 2402 String Key Value Type Required? Description 2403 ============================ ============== ========= ===================== 2404 "KernargSegmentSize" integer Required The size in bytes of 2405 the kernarg segment 2406 that holds the values 2407 of the arguments to 2408 the kernel. 2409 "GroupSegmentFixedSize" integer Required The amount of group 2410 segment memory 2411 required by a 2412 work-group in 2413 bytes. This does not 2414 include any 2415 dynamically allocated 2416 group segment memory 2417 that may be added 2418 when the kernel is 2419 dispatched. 2420 "PrivateSegmentFixedSize" integer Required The amount of fixed 2421 private address space 2422 memory required for a 2423 work-item in 2424 bytes. If the kernel 2425 uses a dynamic call 2426 stack then additional 2427 space must be added 2428 to this value for the 2429 call stack. 2430 "KernargSegmentAlign" integer Required The maximum byte 2431 alignment of 2432 arguments in the 2433 kernarg segment. Must 2434 be a power of 2. 2435 "WavefrontSize" integer Required Wavefront size. Must 2436 be a power of 2. 2437 "NumSGPRs" integer Required Number of scalar 2438 registers used by a 2439 wavefront for 2440 GFX6-GFX10. This 2441 includes the special 2442 SGPRs for VCC, Flat 2443 Scratch (GFX7-GFX10) 2444 and XNACK (for 2445 GFX8-GFX10). It does 2446 not include the 16 2447 SGPR added if a trap 2448 handler is 2449 enabled. It is not 2450 rounded up to the 2451 allocation 2452 granularity. 2453 "NumVGPRs" integer Required Number of vector 2454 registers used by 2455 each work-item for 2456 GFX6-GFX10 2457 "MaxFlatWorkGroupSize" integer Required Maximum flat 2458 work-group size 2459 supported by the 2460 kernel in work-items. 2461 Must be >=1 and 2462 consistent with 2463 ReqdWorkGroupSize if 2464 not 0, 0, 0. 2465 "NumSpilledSGPRs" integer Number of stores from 2466 a scalar register to 2467 a register allocator 2468 created spill 2469 location. 2470 "NumSpilledVGPRs" integer Number of stores from 2471 a vector register to 2472 a register allocator 2473 created spill 2474 location. 2475 ============================ ============== ========= ===================== 2476 2477.. _amdgpu-amdhsa-code-object-metadata-v3: 2478 2479Code Object V3 Metadata (-mattr=+code-object-v3) 2480++++++++++++++++++++++++++++++++++++++++++++++++ 2481 2482Code object V3 metadata is specified by the ``NT_AMDGPU_METADATA`` note record 2483(see :ref:`amdgpu-note-records-v3`). 2484 2485The metadata is represented as Message Pack formatted binary data (see 2486[MsgPack]_). The top level is a Message Pack map that includes the 2487keys defined in table 2488:ref:`amdgpu-amdhsa-code-object-metadata-map-table-v3` and referenced 2489tables. 2490 2491Additional information can be added to the maps. To avoid conflicts, 2492any key names should be prefixed by "*vendor-name*." where 2493``vendor-name`` can be the name of the vendor and specific vendor 2494tool that generates the information. The prefix is abbreviated to 2495simply "." when it appears within a map that has been added by the 2496same *vendor-name*. 2497 2498 .. table:: AMDHSA Code Object V3 Metadata Map 2499 :name: amdgpu-amdhsa-code-object-metadata-map-table-v3 2500 2501 ================= ============== ========= ======================================= 2502 String Key Value Type Required? Description 2503 ================= ============== ========= ======================================= 2504 "amdhsa.version" sequence of Required - The first integer is the major 2505 2 integers version. Currently 1. 2506 - The second integer is the minor 2507 version. Currently 0. 2508 "amdhsa.printf" sequence of Each string is encoded information 2509 strings about a printf function call. The 2510 encoded information is organized as 2511 fields separated by colon (':'): 2512 2513 ``ID:N:S[0]:S[1]:...:S[N-1]:FormatString`` 2514 2515 where: 2516 2517 ``ID`` 2518 A 32-bit integer as a unique id for 2519 each printf function call 2520 2521 ``N`` 2522 A 32-bit integer equal to the number 2523 of arguments of printf function call 2524 minus 1 2525 2526 ``S[i]`` (where i = 0, 1, ... , N-1) 2527 32-bit integers for the size in bytes 2528 of the i-th FormatString argument of 2529 the printf function call 2530 2531 FormatString 2532 The format string passed to the 2533 printf function call. 2534 "amdhsa.kernels" sequence of Required Sequence of the maps for each 2535 map kernel in the code object. See 2536 :ref:`amdgpu-amdhsa-code-object-kernel-metadata-map-table-v3` 2537 for the definition of the keys included 2538 in that map. 2539 ================= ============== ========= ======================================= 2540 2541.. 2542 2543 .. table:: AMDHSA Code Object V3 Kernel Metadata Map 2544 :name: amdgpu-amdhsa-code-object-kernel-metadata-map-table-v3 2545 2546 =================================== ============== ========= ================================ 2547 String Key Value Type Required? Description 2548 =================================== ============== ========= ================================ 2549 ".name" string Required Source name of the kernel. 2550 ".symbol" string Required Name of the kernel 2551 descriptor ELF symbol. 2552 ".language" string Source language of the kernel. 2553 Values include: 2554 2555 - "OpenCL C" 2556 - "OpenCL C++" 2557 - "HCC" 2558 - "HIP" 2559 - "OpenMP" 2560 - "Assembler" 2561 2562 ".language_version" sequence of - The first integer is the major 2563 2 integers version. 2564 - The second integer is the 2565 minor version. 2566 ".args" sequence of Sequence of maps of the 2567 map kernel arguments. See 2568 :ref:`amdgpu-amdhsa-code-object-kernel-argument-metadata-map-table-v3` 2569 for the definition of the keys 2570 included in that map. 2571 ".reqd_workgroup_size" sequence of If not 0, 0, 0 then all values 2572 3 integers must be >=1 and the dispatch 2573 work-group size X, Y, Z must 2574 correspond to the specified 2575 values. Defaults to 0, 0, 0. 2576 2577 Corresponds to the OpenCL 2578 ``reqd_work_group_size`` 2579 attribute. 2580 ".workgroup_size_hint" sequence of The dispatch work-group size 2581 3 integers X, Y, Z is likely to be the 2582 specified values. 2583 2584 Corresponds to the OpenCL 2585 ``work_group_size_hint`` 2586 attribute. 2587 ".vec_type_hint" string The name of a scalar or vector 2588 type. 2589 2590 Corresponds to the OpenCL 2591 ``vec_type_hint`` attribute. 2592 2593 ".device_enqueue_symbol" string The external symbol name 2594 associated with a kernel. 2595 OpenCL runtime allocates a 2596 global buffer for the symbol 2597 and saves the kernel's address 2598 to it, which is used for 2599 device side enqueueing. Only 2600 available for device side 2601 enqueued kernels. 2602 ".kernarg_segment_size" integer Required The size in bytes of 2603 the kernarg segment 2604 that holds the values 2605 of the arguments to 2606 the kernel. 2607 ".group_segment_fixed_size" integer Required The amount of group 2608 segment memory 2609 required by a 2610 work-group in 2611 bytes. This does not 2612 include any 2613 dynamically allocated 2614 group segment memory 2615 that may be added 2616 when the kernel is 2617 dispatched. 2618 ".private_segment_fixed_size" integer Required The amount of fixed 2619 private address space 2620 memory required for a 2621 work-item in 2622 bytes. If the kernel 2623 uses a dynamic call 2624 stack then additional 2625 space must be added 2626 to this value for the 2627 call stack. 2628 ".kernarg_segment_align" integer Required The maximum byte 2629 alignment of 2630 arguments in the 2631 kernarg segment. Must 2632 be a power of 2. 2633 ".wavefront_size" integer Required Wavefront size. Must 2634 be a power of 2. 2635 ".sgpr_count" integer Required Number of scalar 2636 registers required by a 2637 wavefront for 2638 GFX6-GFX9. A register 2639 is required if it is 2640 used explicitly, or 2641 if a higher numbered 2642 register is used 2643 explicitly. This 2644 includes the special 2645 SGPRs for VCC, Flat 2646 Scratch (GFX7-GFX9) 2647 and XNACK (for 2648 GFX8-GFX9). It does 2649 not include the 16 2650 SGPR added if a trap 2651 handler is 2652 enabled. It is not 2653 rounded up to the 2654 allocation 2655 granularity. 2656 ".vgpr_count" integer Required Number of vector 2657 registers required by 2658 each work-item for 2659 GFX6-GFX9. A register 2660 is required if it is 2661 used explicitly, or 2662 if a higher numbered 2663 register is used 2664 explicitly. 2665 ".max_flat_workgroup_size" integer Required Maximum flat 2666 work-group size 2667 supported by the 2668 kernel in work-items. 2669 Must be >=1 and 2670 consistent with 2671 ReqdWorkGroupSize if 2672 not 0, 0, 0. 2673 ".sgpr_spill_count" integer Number of stores from 2674 a scalar register to 2675 a register allocator 2676 created spill 2677 location. 2678 ".vgpr_spill_count" integer Number of stores from 2679 a vector register to 2680 a register allocator 2681 created spill 2682 location. 2683 =================================== ============== ========= ================================ 2684 2685.. 2686 2687 .. table:: AMDHSA Code Object V3 Kernel Argument Metadata Map 2688 :name: amdgpu-amdhsa-code-object-kernel-argument-metadata-map-table-v3 2689 2690 ====================== ============== ========= ================================ 2691 String Key Value Type Required? Description 2692 ====================== ============== ========= ================================ 2693 ".name" string Kernel argument name. 2694 ".type_name" string Kernel argument type name. 2695 ".size" integer Required Kernel argument size in bytes. 2696 ".offset" integer Required Kernel argument offset in 2697 bytes. The offset must be a 2698 multiple of the alignment 2699 required by the argument. 2700 ".value_kind" string Required Kernel argument kind that 2701 specifies how to set up the 2702 corresponding argument. 2703 Values include: 2704 2705 "by_value" 2706 The argument is copied 2707 directly into the kernarg. 2708 2709 "global_buffer" 2710 A global address space pointer 2711 to the buffer data is passed 2712 in the kernarg. 2713 2714 "dynamic_shared_pointer" 2715 A group address space pointer 2716 to dynamically allocated LDS 2717 is passed in the kernarg. 2718 2719 "sampler" 2720 A global address space 2721 pointer to a S# is passed in 2722 the kernarg. 2723 2724 "image" 2725 A global address space 2726 pointer to a T# is passed in 2727 the kernarg. 2728 2729 "pipe" 2730 A global address space pointer 2731 to an OpenCL pipe is passed in 2732 the kernarg. 2733 2734 "queue" 2735 A global address space pointer 2736 to an OpenCL device enqueue 2737 queue is passed in the 2738 kernarg. 2739 2740 "hidden_global_offset_x" 2741 The OpenCL grid dispatch 2742 global offset for the X 2743 dimension is passed in the 2744 kernarg. 2745 2746 "hidden_global_offset_y" 2747 The OpenCL grid dispatch 2748 global offset for the Y 2749 dimension is passed in the 2750 kernarg. 2751 2752 "hidden_global_offset_z" 2753 The OpenCL grid dispatch 2754 global offset for the Z 2755 dimension is passed in the 2756 kernarg. 2757 2758 "hidden_none" 2759 An argument that is not used 2760 by the kernel. Space needs to 2761 be left for it, but it does 2762 not need to be set up. 2763 2764 "hidden_printf_buffer" 2765 A global address space pointer 2766 to the runtime printf buffer 2767 is passed in kernarg. 2768 2769 "hidden_hostcall_buffer" 2770 A global address space pointer 2771 to the runtime hostcall buffer 2772 is passed in kernarg. 2773 2774 "hidden_default_queue" 2775 A global address space pointer 2776 to the OpenCL device enqueue 2777 queue that should be used by 2778 the kernel by default is 2779 passed in the kernarg. 2780 2781 "hidden_completion_action" 2782 A global address space pointer 2783 to help link enqueued kernels into 2784 the ancestor tree for determining 2785 when the parent kernel has finished. 2786 2787 "hidden_multigrid_sync_arg" 2788 A global address space pointer for 2789 multi-grid synchronization is 2790 passed in the kernarg. 2791 2792 ".value_type" string Unused and deprecated. This should no longer 2793 be emitted, but is accepted for compatibility. 2794 2795 ".pointee_align" integer Alignment in bytes of pointee 2796 type for pointer type kernel 2797 argument. Must be a power 2798 of 2. Only present if 2799 ".value_kind" is 2800 "dynamic_shared_pointer". 2801 ".address_space" string Kernel argument address space 2802 qualifier. Only present if 2803 ".value_kind" is "global_buffer" or 2804 "dynamic_shared_pointer". Values 2805 are: 2806 2807 - "private" 2808 - "global" 2809 - "constant" 2810 - "local" 2811 - "generic" 2812 - "region" 2813 2814 .. TODO:: 2815 Is "global_buffer" only "global" 2816 or "constant"? Is 2817 "dynamic_shared_pointer" always 2818 "local"? Can HCC allow "generic"? 2819 How can "private" or "region" 2820 ever happen? 2821 ".access" string Kernel argument access 2822 qualifier. Only present if 2823 ".value_kind" is "image" or 2824 "pipe". Values 2825 are: 2826 2827 - "read_only" 2828 - "write_only" 2829 - "read_write" 2830 2831 .. TODO:: 2832 Does this apply to 2833 "global_buffer"? 2834 ".actual_access" string The actual memory accesses 2835 performed by the kernel on the 2836 kernel argument. Only present if 2837 ".value_kind" is "global_buffer", 2838 "image", or "pipe". This may be 2839 more restrictive than indicated 2840 by ".access" to reflect what the 2841 kernel actual does. If not 2842 present then the runtime must 2843 assume what is implied by 2844 ".access" and ".is_const" . Values 2845 are: 2846 2847 - "read_only" 2848 - "write_only" 2849 - "read_write" 2850 2851 ".is_const" boolean Indicates if the kernel argument 2852 is const qualified. Only present 2853 if ".value_kind" is 2854 "global_buffer". 2855 2856 ".is_restrict" boolean Indicates if the kernel argument 2857 is restrict qualified. Only 2858 present if ".value_kind" is 2859 "global_buffer". 2860 2861 ".is_volatile" boolean Indicates if the kernel argument 2862 is volatile qualified. Only 2863 present if ".value_kind" is 2864 "global_buffer". 2865 2866 ".is_pipe" boolean Indicates if the kernel argument 2867 is pipe qualified. Only present 2868 if ".value_kind" is "pipe". 2869 2870 .. TODO:: 2871 Can "global_buffer" be pipe 2872 qualified? 2873 ====================== ============== ========= ================================ 2874 2875.. 2876 2877Kernel Dispatch 2878~~~~~~~~~~~~~~~ 2879 2880The HSA architected queuing language (AQL) defines a user space memory 2881interface that can be used to control the dispatch of kernels, in an agent 2882independent way. An agent can have zero or more AQL queues created for it using 2883the ROCm runtime, in which AQL packets (all of which are 64 bytes) can be 2884placed. See the *HSA Platform System Architecture Specification* [HSA]_ for the 2885AQL queue mechanics and packet layouts. 2886 2887The packet processor of a kernel agent is responsible for detecting and 2888dispatching HSA kernels from the AQL queues associated with it. For AMD GPUs the 2889packet processor is implemented by the hardware command processor (CP), 2890asynchronous dispatch controller (ADC) and shader processor input controller 2891(SPI). 2892 2893The ROCm runtime can be used to allocate an AQL queue object. It uses the kernel 2894mode driver to initialize and register the AQL queue with CP. 2895 2896To dispatch a kernel the following actions are performed. This can occur in the 2897CPU host program, or from an HSA kernel executing on a GPU. 2898 28991. A pointer to an AQL queue for the kernel agent on which the kernel is to be 2900 executed is obtained. 29012. A pointer to the kernel descriptor (see 2902 :ref:`amdgpu-amdhsa-kernel-descriptor`) of the kernel to execute is obtained. 2903 It must be for a kernel that is contained in a code object that that was 2904 loaded by the ROCm runtime on the kernel agent with which the AQL queue is 2905 associated. 29063. Space is allocated for the kernel arguments using the ROCm runtime allocator 2907 for a memory region with the kernarg property for the kernel agent that will 2908 execute the kernel. It must be at least 16-byte aligned. 29094. Kernel argument values are assigned to the kernel argument memory 2910 allocation. The layout is defined in the *HSA Programmer's Language 2911 Reference* [HSA]_. For AMDGPU the kernel execution directly accesses the 2912 kernel argument memory in the same way constant memory is accessed. (Note 2913 that the HSA specification allows an implementation to copy the kernel 2914 argument contents to another location that is accessed by the kernel.) 29155. An AQL kernel dispatch packet is created on the AQL queue. The ROCm runtime 2916 api uses 64-bit atomic operations to reserve space in the AQL queue for the 2917 packet. The packet must be set up, and the final write must use an atomic 2918 store release to set the packet kind to ensure the packet contents are 2919 visible to the kernel agent. AQL defines a doorbell signal mechanism to 2920 notify the kernel agent that the AQL queue has been updated. These rules, and 2921 the layout of the AQL queue and kernel dispatch packet is defined in the *HSA 2922 System Architecture Specification* [HSA]_. 29236. A kernel dispatch packet includes information about the actual dispatch, 2924 such as grid and work-group size, together with information from the code 2925 object about the kernel, such as segment sizes. The ROCm runtime queries on 2926 the kernel symbol can be used to obtain the code object values which are 2927 recorded in the :ref:`amdgpu-amdhsa-code-object-metadata`. 29287. CP executes micro-code and is responsible for detecting and setting up the 2929 GPU to execute the wavefronts of a kernel dispatch. 29308. CP ensures that when the a wavefront starts executing the kernel machine 2931 code, the scalar general purpose registers (SGPR) and vector general purpose 2932 registers (VGPR) are set up as required by the machine code. The required 2933 setup is defined in the :ref:`amdgpu-amdhsa-kernel-descriptor`. The initial 2934 register state is defined in 2935 :ref:`amdgpu-amdhsa-initial-kernel-execution-state`. 29369. The prolog of the kernel machine code (see 2937 :ref:`amdgpu-amdhsa-kernel-prolog`) sets up the machine state as necessary 2938 before continuing executing the machine code that corresponds to the kernel. 293910. When the kernel dispatch has completed execution, CP signals the completion 2940 signal specified in the kernel dispatch packet if not 0. 2941 2942Image and Samplers 2943~~~~~~~~~~~~~~~~~~ 2944 2945Image and sample handles created by the ROCm runtime are 64-bit addresses of a 2946hardware 32-byte V# and 48 byte S# object respectively. In order to support the 2947HSA ``query_sampler`` operations two extra dwords are used to store the HSA BRIG 2948enumeration values for the queries that are not trivially deducible from the S# 2949representation. 2950 2951HSA Signals 2952~~~~~~~~~~~ 2953 2954HSA signal handles created by the ROCm runtime are 64-bit addresses of a 2955structure allocated in memory accessible from both the CPU and GPU. The 2956structure is defined by the ROCm runtime and subject to change between releases 2957(see [AMD-ROCm-github]_). 2958 2959.. _amdgpu-amdhsa-hsa-aql-queue: 2960 2961HSA AQL Queue 2962~~~~~~~~~~~~~ 2963 2964The HSA AQL queue structure is defined by the ROCm runtime and subject to change 2965between releases (see [AMD-ROCm-github]_). For some processors it contains 2966fields needed to implement certain language features such as the flat address 2967aperture bases. It also contains fields used by CP such as managing the 2968allocation of scratch memory. 2969 2970.. _amdgpu-amdhsa-kernel-descriptor: 2971 2972Kernel Descriptor 2973~~~~~~~~~~~~~~~~~ 2974 2975A kernel descriptor consists of the information needed by CP to initiate the 2976execution of a kernel, including the entry point address of the machine code 2977that implements the kernel. 2978 2979Kernel Descriptor for GFX6-GFX10 2980++++++++++++++++++++++++++++++++ 2981 2982CP microcode requires the Kernel descriptor to be allocated on 64-byte 2983alignment. 2984 2985 .. table:: Kernel Descriptor for GFX6-GFX10 2986 :name: amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table 2987 2988 ======= ======= =============================== ============================ 2989 Bits Size Field Name Description 2990 ======= ======= =============================== ============================ 2991 31:0 4 bytes GROUP_SEGMENT_FIXED_SIZE The amount of fixed local 2992 address space memory 2993 required for a work-group 2994 in bytes. This does not 2995 include any dynamically 2996 allocated local address 2997 space memory that may be 2998 added when the kernel is 2999 dispatched. 3000 63:32 4 bytes PRIVATE_SEGMENT_FIXED_SIZE The amount of fixed 3001 private address space 3002 memory required for a 3003 work-item in bytes. If 3004 is_dynamic_callstack is 1 3005 then additional space must 3006 be added to this value for 3007 the call stack. 3008 127:64 8 bytes Reserved, must be 0. 3009 191:128 8 bytes KERNEL_CODE_ENTRY_BYTE_OFFSET Byte offset (possibly 3010 negative) from base 3011 address of kernel 3012 descriptor to kernel's 3013 entry point instruction 3014 which must be 256 byte 3015 aligned. 3016 351:272 20 Reserved, must be 0. 3017 bytes 3018 383:352 4 bytes COMPUTE_PGM_RSRC3 GFX6-9 3019 Reserved, must be 0. 3020 GFX10 3021 Compute Shader (CS) 3022 program settings used by 3023 CP to set up 3024 ``COMPUTE_PGM_RSRC3`` 3025 configuration 3026 register. See 3027 :ref:`amdgpu-amdhsa-compute_pgm_rsrc3-gfx10-table`. 3028 415:384 4 bytes COMPUTE_PGM_RSRC1 Compute Shader (CS) 3029 program settings used by 3030 CP to set up 3031 ``COMPUTE_PGM_RSRC1`` 3032 configuration 3033 register. See 3034 :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`. 3035 447:416 4 bytes COMPUTE_PGM_RSRC2 Compute Shader (CS) 3036 program settings used by 3037 CP to set up 3038 ``COMPUTE_PGM_RSRC2`` 3039 configuration 3040 register. See 3041 :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`. 3042 448 1 bit ENABLE_SGPR_PRIVATE_SEGMENT Enable the setup of the 3043 _BUFFER SGPR user data registers 3044 (see 3045 :ref:`amdgpu-amdhsa-initial-kernel-execution-state`). 3046 3047 The total number of SGPR 3048 user data registers 3049 requested must not exceed 3050 16 and match value in 3051 ``compute_pgm_rsrc2.user_sgpr.user_sgpr_count``. 3052 Any requests beyond 16 3053 will be ignored. 3054 449 1 bit ENABLE_SGPR_DISPATCH_PTR *see above* 3055 450 1 bit ENABLE_SGPR_QUEUE_PTR *see above* 3056 451 1 bit ENABLE_SGPR_KERNARG_SEGMENT_PTR *see above* 3057 452 1 bit ENABLE_SGPR_DISPATCH_ID *see above* 3058 453 1 bit ENABLE_SGPR_FLAT_SCRATCH_INIT *see above* 3059 454 1 bit ENABLE_SGPR_PRIVATE_SEGMENT *see above* 3060 _SIZE 3061 457:455 3 bits Reserved, must be 0. 3062 458 1 bit ENABLE_WAVEFRONT_SIZE32 GFX6-9 3063 Reserved, must be 0. 3064 GFX10 3065 - If 0 execute in 3066 wavefront size 64 mode. 3067 - If 1 execute in 3068 native wavefront size 3069 32 mode. 3070 463:459 5 bits Reserved, must be 0. 3071 511:464 6 bytes Reserved, must be 0. 3072 512 **Total size 64 bytes.** 3073 ======= ==================================================================== 3074 3075.. 3076 3077 .. table:: compute_pgm_rsrc1 for GFX6-GFX10 3078 :name: amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table 3079 3080 ======= ======= =============================== =========================================================================== 3081 Bits Size Field Name Description 3082 ======= ======= =============================== =========================================================================== 3083 5:0 6 bits GRANULATED_WORKITEM_VGPR_COUNT Number of vector register 3084 blocks used by each work-item; 3085 granularity is device 3086 specific: 3087 3088 GFX6-GFX9 3089 - vgprs_used 0..256 3090 - max(0, ceil(vgprs_used / 4) - 1) 3091 GFX10 (wavefront size 64) 3092 - max_vgpr 1..256 3093 - max(0, ceil(vgprs_used / 4) - 1) 3094 GFX10 (wavefront size 32) 3095 - max_vgpr 1..256 3096 - max(0, ceil(vgprs_used / 8) - 1) 3097 3098 Where vgprs_used is defined 3099 as the highest VGPR number 3100 explicitly referenced plus 3101 one. 3102 3103 Used by CP to set up 3104 ``COMPUTE_PGM_RSRC1.VGPRS``. 3105 3106 The 3107 :ref:`amdgpu-assembler` 3108 calculates this 3109 automatically for the 3110 selected processor from 3111 values provided to the 3112 `.amdhsa_kernel` directive 3113 by the 3114 `.amdhsa_next_free_vgpr` 3115 nested directive (see 3116 :ref:`amdhsa-kernel-directives-table`). 3117 9:6 4 bits GRANULATED_WAVEFRONT_SGPR_COUNT Number of scalar register 3118 blocks used by a wavefront; 3119 granularity is device 3120 specific: 3121 3122 GFX6-GFX8 3123 - sgprs_used 0..112 3124 - max(0, ceil(sgprs_used / 8) - 1) 3125 GFX9 3126 - sgprs_used 0..112 3127 - 2 * max(0, ceil(sgprs_used / 16) - 1) 3128 GFX10 3129 Reserved, must be 0. 3130 (128 SGPRs always 3131 allocated.) 3132 3133 Where sgprs_used is 3134 defined as the highest 3135 SGPR number explicitly 3136 referenced plus one, plus 3137 a target specific number 3138 of additional special 3139 SGPRs for VCC, 3140 FLAT_SCRATCH (GFX7+) and 3141 XNACK_MASK (GFX8+), and 3142 any additional 3143 target specific 3144 limitations. It does not 3145 include the 16 SGPRs added 3146 if a trap handler is 3147 enabled. 3148 3149 The target specific 3150 limitations and special 3151 SGPR layout are defined in 3152 the hardware 3153 documentation, which can 3154 be found in the 3155 :ref:`amdgpu-processors` 3156 table. 3157 3158 Used by CP to set up 3159 ``COMPUTE_PGM_RSRC1.SGPRS``. 3160 3161 The 3162 :ref:`amdgpu-assembler` 3163 calculates this 3164 automatically for the 3165 selected processor from 3166 values provided to the 3167 `.amdhsa_kernel` directive 3168 by the 3169 `.amdhsa_next_free_sgpr` 3170 and `.amdhsa_reserve_*` 3171 nested directives (see 3172 :ref:`amdhsa-kernel-directives-table`). 3173 11:10 2 bits PRIORITY Must be 0. 3174 3175 Start executing wavefront 3176 at the specified priority. 3177 3178 CP is responsible for 3179 filling in 3180 ``COMPUTE_PGM_RSRC1.PRIORITY``. 3181 13:12 2 bits FLOAT_ROUND_MODE_32 Wavefront starts execution 3182 with specified rounding 3183 mode for single (32 3184 bit) floating point 3185 precision floating point 3186 operations. 3187 3188 Floating point rounding 3189 mode values are defined in 3190 :ref:`amdgpu-amdhsa-floating-point-rounding-mode-enumeration-values-table`. 3191 3192 Used by CP to set up 3193 ``COMPUTE_PGM_RSRC1.FLOAT_MODE``. 3194 15:14 2 bits FLOAT_ROUND_MODE_16_64 Wavefront starts execution 3195 with specified rounding 3196 denorm mode for half/double (16 3197 and 64-bit) floating point 3198 precision floating point 3199 operations. 3200 3201 Floating point rounding 3202 mode values are defined in 3203 :ref:`amdgpu-amdhsa-floating-point-rounding-mode-enumeration-values-table`. 3204 3205 Used by CP to set up 3206 ``COMPUTE_PGM_RSRC1.FLOAT_MODE``. 3207 17:16 2 bits FLOAT_DENORM_MODE_32 Wavefront starts execution 3208 with specified denorm mode 3209 for single (32 3210 bit) floating point 3211 precision floating point 3212 operations. 3213 3214 Floating point denorm mode 3215 values are defined in 3216 :ref:`amdgpu-amdhsa-floating-point-denorm-mode-enumeration-values-table`. 3217 3218 Used by CP to set up 3219 ``COMPUTE_PGM_RSRC1.FLOAT_MODE``. 3220 19:18 2 bits FLOAT_DENORM_MODE_16_64 Wavefront starts execution 3221 with specified denorm mode 3222 for half/double (16 3223 and 64-bit) floating point 3224 precision floating point 3225 operations. 3226 3227 Floating point denorm mode 3228 values are defined in 3229 :ref:`amdgpu-amdhsa-floating-point-denorm-mode-enumeration-values-table`. 3230 3231 Used by CP to set up 3232 ``COMPUTE_PGM_RSRC1.FLOAT_MODE``. 3233 20 1 bit PRIV Must be 0. 3234 3235 Start executing wavefront 3236 in privilege trap handler 3237 mode. 3238 3239 CP is responsible for 3240 filling in 3241 ``COMPUTE_PGM_RSRC1.PRIV``. 3242 21 1 bit ENABLE_DX10_CLAMP Wavefront starts execution 3243 with DX10 clamp mode 3244 enabled. Used by the vector 3245 ALU to force DX10 style 3246 treatment of NaN's (when 3247 set, clamp NaN to zero, 3248 otherwise pass NaN 3249 through). 3250 3251 Used by CP to set up 3252 ``COMPUTE_PGM_RSRC1.DX10_CLAMP``. 3253 22 1 bit DEBUG_MODE Must be 0. 3254 3255 Start executing wavefront 3256 in single step mode. 3257 3258 CP is responsible for 3259 filling in 3260 ``COMPUTE_PGM_RSRC1.DEBUG_MODE``. 3261 23 1 bit ENABLE_IEEE_MODE Wavefront starts execution 3262 with IEEE mode 3263 enabled. Floating point 3264 opcodes that support 3265 exception flag gathering 3266 will quiet and propagate 3267 signaling-NaN inputs per 3268 IEEE 754-2008. Min_dx10 and 3269 max_dx10 become IEEE 3270 754-2008 compliant due to 3271 signaling-NaN propagation 3272 and quieting. 3273 3274 Used by CP to set up 3275 ``COMPUTE_PGM_RSRC1.IEEE_MODE``. 3276 24 1 bit BULKY Must be 0. 3277 3278 Only one work-group allowed 3279 to execute on a compute 3280 unit. 3281 3282 CP is responsible for 3283 filling in 3284 ``COMPUTE_PGM_RSRC1.BULKY``. 3285 25 1 bit CDBG_USER Must be 0. 3286 3287 Flag that can be used to 3288 control debugging code. 3289 3290 CP is responsible for 3291 filling in 3292 ``COMPUTE_PGM_RSRC1.CDBG_USER``. 3293 26 1 bit FP16_OVFL GFX6-GFX8 3294 Reserved, must be 0. 3295 GFX9-GFX10 3296 Wavefront starts execution 3297 with specified fp16 overflow 3298 mode. 3299 3300 - If 0, fp16 overflow generates 3301 +/-INF values. 3302 - If 1, fp16 overflow that is the 3303 result of an +/-INF input value 3304 or divide by 0 produces a +/-INF, 3305 otherwise clamps computed 3306 overflow to +/-MAX_FP16 as 3307 appropriate. 3308 3309 Used by CP to set up 3310 ``COMPUTE_PGM_RSRC1.FP16_OVFL``. 3311 28:27 2 bits Reserved, must be 0. 3312 29 1 bit WGP_MODE GFX6-GFX9 3313 Reserved, must be 0. 3314 GFX10 3315 - If 0 execute work-groups in 3316 CU wavefront execution mode. 3317 - If 1 execute work-groups on 3318 in WGP wavefront execution mode. 3319 3320 See :ref:`amdgpu-amdhsa-memory-model`. 3321 3322 Used by CP to set up 3323 ``COMPUTE_PGM_RSRC1.WGP_MODE``. 3324 30 1 bit MEM_ORDERED GFX6-9 3325 Reserved, must be 0. 3326 GFX10 3327 Controls the behavior of the 3328 waitcnt's vmcnt and vscnt 3329 counters. 3330 3331 - If 0 vmcnt reports completion 3332 of load and atomic with return 3333 out of order with sample 3334 instructions, and the vscnt 3335 reports the completion of 3336 store and atomic without 3337 return in order. 3338 - If 1 vmcnt reports completion 3339 of load, atomic with return 3340 and sample instructions in 3341 order, and the vscnt reports 3342 the completion of store and 3343 atomic without return in order. 3344 3345 Used by CP to set up 3346 ``COMPUTE_PGM_RSRC1.MEM_ORDERED``. 3347 31 1 bit FWD_PROGRESS GFX6-9 3348 Reserved, must be 0. 3349 GFX10 3350 - If 0 execute SIMD wavefronts 3351 using oldest first policy. 3352 - If 1 execute SIMD wavefronts to 3353 ensure wavefronts will make some 3354 forward progress. 3355 3356 Used by CP to set up 3357 ``COMPUTE_PGM_RSRC1.FWD_PROGRESS``. 3358 32 **Total size 4 bytes** 3359 ======= =================================================================================================================== 3360 3361.. 3362 3363 .. table:: compute_pgm_rsrc2 for GFX6-GFX10 3364 :name: amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table 3365 3366 ======= ======= =============================== =========================================================================== 3367 Bits Size Field Name Description 3368 ======= ======= =============================== =========================================================================== 3369 0 1 bit ENABLE_SGPR_PRIVATE_SEGMENT Enable the setup of the 3370 _WAVEFRONT_OFFSET SGPR wavefront scratch offset 3371 system register (see 3372 :ref:`amdgpu-amdhsa-initial-kernel-execution-state`). 3373 3374 Used by CP to set up 3375 ``COMPUTE_PGM_RSRC2.SCRATCH_EN``. 3376 5:1 5 bits USER_SGPR_COUNT The total number of SGPR 3377 user data registers 3378 requested. This number must 3379 match the number of user 3380 data registers enabled. 3381 3382 Used by CP to set up 3383 ``COMPUTE_PGM_RSRC2.USER_SGPR``. 3384 6 1 bit ENABLE_TRAP_HANDLER Must be 0. 3385 3386 This bit represents 3387 ``COMPUTE_PGM_RSRC2.TRAP_PRESENT``, 3388 which is set by the CP if 3389 the runtime has installed a 3390 trap handler. 3391 7 1 bit ENABLE_SGPR_WORKGROUP_ID_X Enable the setup of the 3392 system SGPR register for 3393 the work-group id in the X 3394 dimension (see 3395 :ref:`amdgpu-amdhsa-initial-kernel-execution-state`). 3396 3397 Used by CP to set up 3398 ``COMPUTE_PGM_RSRC2.TGID_X_EN``. 3399 8 1 bit ENABLE_SGPR_WORKGROUP_ID_Y Enable the setup of the 3400 system SGPR register for 3401 the work-group id in the Y 3402 dimension (see 3403 :ref:`amdgpu-amdhsa-initial-kernel-execution-state`). 3404 3405 Used by CP to set up 3406 ``COMPUTE_PGM_RSRC2.TGID_Y_EN``. 3407 9 1 bit ENABLE_SGPR_WORKGROUP_ID_Z Enable the setup of the 3408 system SGPR register for 3409 the work-group id in the Z 3410 dimension (see 3411 :ref:`amdgpu-amdhsa-initial-kernel-execution-state`). 3412 3413 Used by CP to set up 3414 ``COMPUTE_PGM_RSRC2.TGID_Z_EN``. 3415 10 1 bit ENABLE_SGPR_WORKGROUP_INFO Enable the setup of the 3416 system SGPR register for 3417 work-group information (see 3418 :ref:`amdgpu-amdhsa-initial-kernel-execution-state`). 3419 3420 Used by CP to set up 3421 ``COMPUTE_PGM_RSRC2.TGID_SIZE_EN``. 3422 12:11 2 bits ENABLE_VGPR_WORKITEM_ID Enable the setup of the 3423 VGPR system registers used 3424 for the work-item ID. 3425 :ref:`amdgpu-amdhsa-system-vgpr-work-item-id-enumeration-values-table` 3426 defines the values. 3427 3428 Used by CP to set up 3429 ``COMPUTE_PGM_RSRC2.TIDIG_CMP_CNT``. 3430 13 1 bit ENABLE_EXCEPTION_ADDRESS_WATCH Must be 0. 3431 3432 Wavefront starts execution 3433 with address watch 3434 exceptions enabled which 3435 are generated when L1 has 3436 witnessed a thread access 3437 an *address of 3438 interest*. 3439 3440 CP is responsible for 3441 filling in the address 3442 watch bit in 3443 ``COMPUTE_PGM_RSRC2.EXCP_EN_MSB`` 3444 according to what the 3445 runtime requests. 3446 14 1 bit ENABLE_EXCEPTION_MEMORY Must be 0. 3447 3448 Wavefront starts execution 3449 with memory violation 3450 exceptions exceptions 3451 enabled which are generated 3452 when a memory violation has 3453 occurred for this wavefront from 3454 L1 or LDS 3455 (write-to-read-only-memory, 3456 mis-aligned atomic, LDS 3457 address out of range, 3458 illegal address, etc.). 3459 3460 CP sets the memory 3461 violation bit in 3462 ``COMPUTE_PGM_RSRC2.EXCP_EN_MSB`` 3463 according to what the 3464 runtime requests. 3465 23:15 9 bits GRANULATED_LDS_SIZE Must be 0. 3466 3467 CP uses the rounded value 3468 from the dispatch packet, 3469 not this value, as the 3470 dispatch may contain 3471 dynamically allocated group 3472 segment memory. CP writes 3473 directly to 3474 ``COMPUTE_PGM_RSRC2.LDS_SIZE``. 3475 3476 Amount of group segment 3477 (LDS) to allocate for each 3478 work-group. Granularity is 3479 device specific: 3480 3481 GFX6: 3482 roundup(lds-size / (64 * 4)) 3483 GFX7-GFX10: 3484 roundup(lds-size / (128 * 4)) 3485 3486 24 1 bit ENABLE_EXCEPTION_IEEE_754_FP Wavefront starts execution 3487 _INVALID_OPERATION with specified exceptions 3488 enabled. 3489 3490 Used by CP to set up 3491 ``COMPUTE_PGM_RSRC2.EXCP_EN`` 3492 (set from bits 0..6). 3493 3494 IEEE 754 FP Invalid 3495 Operation 3496 25 1 bit ENABLE_EXCEPTION_FP_DENORMAL FP Denormal one or more 3497 _SOURCE input operands is a 3498 denormal number 3499 26 1 bit ENABLE_EXCEPTION_IEEE_754_FP IEEE 754 FP Division by 3500 _DIVISION_BY_ZERO Zero 3501 27 1 bit ENABLE_EXCEPTION_IEEE_754_FP IEEE 754 FP FP Overflow 3502 _OVERFLOW 3503 28 1 bit ENABLE_EXCEPTION_IEEE_754_FP IEEE 754 FP Underflow 3504 _UNDERFLOW 3505 29 1 bit ENABLE_EXCEPTION_IEEE_754_FP IEEE 754 FP Inexact 3506 _INEXACT 3507 30 1 bit ENABLE_EXCEPTION_INT_DIVIDE_BY Integer Division by Zero 3508 _ZERO (rcp_iflag_f32 instruction 3509 only) 3510 31 1 bit Reserved, must be 0. 3511 32 **Total size 4 bytes.** 3512 ======= =================================================================================================================== 3513 3514.. 3515 3516 .. table:: compute_pgm_rsrc3 for GFX10 3517 :name: amdgpu-amdhsa-compute_pgm_rsrc3-gfx10-table 3518 3519 ======= ======= =============================== =========================================================================== 3520 Bits Size Field Name Description 3521 ======= ======= =============================== =========================================================================== 3522 3:0 4 bits SHARED_VGPR_COUNT Number of shared VGPRs for wavefront size 64. Granularity 8. Value 0-120. 3523 compute_pgm_rsrc1.vgprs + shared_vgpr_cnt cannot exceed 64. 3524 31:4 28 Reserved, must be 0. 3525 bits 3526 32 **Total size 4 bytes.** 3527 ======= =================================================================================================================== 3528 3529.. 3530 3531 .. table:: Floating Point Rounding Mode Enumeration Values 3532 :name: amdgpu-amdhsa-floating-point-rounding-mode-enumeration-values-table 3533 3534 ====================================== ===== ============================== 3535 Enumeration Name Value Description 3536 ====================================== ===== ============================== 3537 FLOAT_ROUND_MODE_NEAR_EVEN 0 Round Ties To Even 3538 FLOAT_ROUND_MODE_PLUS_INFINITY 1 Round Toward +infinity 3539 FLOAT_ROUND_MODE_MINUS_INFINITY 2 Round Toward -infinity 3540 FLOAT_ROUND_MODE_ZERO 3 Round Toward 0 3541 ====================================== ===== ============================== 3542 3543.. 3544 3545 .. table:: Floating Point Denorm Mode Enumeration Values 3546 :name: amdgpu-amdhsa-floating-point-denorm-mode-enumeration-values-table 3547 3548 ====================================== ===== ============================== 3549 Enumeration Name Value Description 3550 ====================================== ===== ============================== 3551 FLOAT_DENORM_MODE_FLUSH_SRC_DST 0 Flush Source and Destination 3552 Denorms 3553 FLOAT_DENORM_MODE_FLUSH_DST 1 Flush Output Denorms 3554 FLOAT_DENORM_MODE_FLUSH_SRC 2 Flush Source Denorms 3555 FLOAT_DENORM_MODE_FLUSH_NONE 3 No Flush 3556 ====================================== ===== ============================== 3557 3558.. 3559 3560 .. table:: System VGPR Work-Item ID Enumeration Values 3561 :name: amdgpu-amdhsa-system-vgpr-work-item-id-enumeration-values-table 3562 3563 ======================================== ===== ============================ 3564 Enumeration Name Value Description 3565 ======================================== ===== ============================ 3566 SYSTEM_VGPR_WORKITEM_ID_X 0 Set work-item X dimension 3567 ID. 3568 SYSTEM_VGPR_WORKITEM_ID_X_Y 1 Set work-item X and Y 3569 dimensions ID. 3570 SYSTEM_VGPR_WORKITEM_ID_X_Y_Z 2 Set work-item X, Y and Z 3571 dimensions ID. 3572 SYSTEM_VGPR_WORKITEM_ID_UNDEFINED 3 Undefined. 3573 ======================================== ===== ============================ 3574 3575.. _amdgpu-amdhsa-initial-kernel-execution-state: 3576 3577Initial Kernel Execution State 3578~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 3579 3580This section defines the register state that will be set up by the packet 3581processor prior to the start of execution of every wavefront. This is limited by 3582the constraints of the hardware controllers of CP/ADC/SPI. 3583 3584The order of the SGPR registers is defined, but the compiler can specify which 3585ones are actually setup in the kernel descriptor using the ``enable_sgpr_*`` bit 3586fields (see :ref:`amdgpu-amdhsa-kernel-descriptor`). The register numbers used 3587for enabled registers are dense starting at SGPR0: the first enabled register is 3588SGPR0, the next enabled register is SGPR1 etc.; disabled registers do not have 3589an SGPR number. 3590 3591The initial SGPRs comprise up to 16 User SRGPs that are set by CP and apply to 3592all wavefronts of the grid. It is possible to specify more than 16 User SGPRs 3593using the ``enable_sgpr_*`` bit fields, in which case only the first 16 are 3594actually initialized. These are then immediately followed by the System SGPRs 3595that are set up by ADC/SPI and can have different values for each wavefront of 3596the grid dispatch. 3597 3598SGPR register initial state is defined in 3599:ref:`amdgpu-amdhsa-sgpr-register-set-up-order-table`. 3600 3601 .. table:: SGPR Register Set Up Order 3602 :name: amdgpu-amdhsa-sgpr-register-set-up-order-table 3603 3604 ========== ========================== ====== ============================== 3605 SGPR Order Name Number Description 3606 (kernel descriptor enable of 3607 field) SGPRs 3608 ========== ========================== ====== ============================== 3609 First Private Segment Buffer 4 V# that can be used, together 3610 (enable_sgpr_private with Scratch Wavefront Offset 3611 _segment_buffer) as an offset, to access the 3612 private address space using a 3613 segment address. 3614 3615 CP uses the value provided by 3616 the runtime. 3617 then Dispatch Ptr 2 64-bit address of AQL dispatch 3618 (enable_sgpr_dispatch_ptr) packet for kernel dispatch 3619 actually executing. 3620 then Queue Ptr 2 64-bit address of amd_queue_t 3621 (enable_sgpr_queue_ptr) object for AQL queue on which 3622 the dispatch packet was 3623 queued. 3624 then Kernarg Segment Ptr 2 64-bit address of Kernarg 3625 (enable_sgpr_kernarg segment. This is directly 3626 _segment_ptr) copied from the 3627 kernarg_address in the kernel 3628 dispatch packet. 3629 3630 Having CP load it once avoids 3631 loading it at the beginning of 3632 every wavefront. 3633 then Dispatch Id 2 64-bit Dispatch ID of the 3634 (enable_sgpr_dispatch_id) dispatch packet being 3635 executed. 3636 then Flat Scratch Init 2 This is 2 SGPRs: 3637 (enable_sgpr_flat_scratch 3638 _init) GFX6 3639 Not supported. 3640 GFX7-GFX8 3641 The first SGPR is a 32-bit 3642 byte offset from 3643 ``SH_HIDDEN_PRIVATE_BASE_VIMID`` 3644 to per SPI base of memory 3645 for scratch for the queue 3646 executing the kernel 3647 dispatch. CP obtains this 3648 from the runtime. (The 3649 Scratch Segment Buffer base 3650 address is 3651 ``SH_HIDDEN_PRIVATE_BASE_VIMID`` 3652 plus this offset.) The value 3653 of Scratch Wavefront Offset must 3654 be added to this offset by 3655 the kernel machine code, 3656 right shifted by 8, and 3657 moved to the FLAT_SCRATCH_HI 3658 SGPR register. 3659 FLAT_SCRATCH_HI corresponds 3660 to SGPRn-4 on GFX7, and 3661 SGPRn-6 on GFX8 (where SGPRn 3662 is the highest numbered SGPR 3663 allocated to the wavefront). 3664 FLAT_SCRATCH_HI is 3665 multiplied by 256 (as it is 3666 in units of 256 bytes) and 3667 added to 3668 ``SH_HIDDEN_PRIVATE_BASE_VIMID`` 3669 to calculate the per wavefront 3670 FLAT SCRATCH BASE in flat 3671 memory instructions that 3672 access the scratch 3673 aperture. 3674 3675 The second SGPR is 32-bit 3676 byte size of a single 3677 work-item's scratch memory 3678 usage. CP obtains this from 3679 the runtime, and it is 3680 always a multiple of DWORD. 3681 CP checks that the value in 3682 the kernel dispatch packet 3683 Private Segment Byte Size is 3684 not larger and requests the 3685 runtime to increase the 3686 queue's scratch size if 3687 necessary. The kernel code 3688 must move it to 3689 FLAT_SCRATCH_LO which is 3690 SGPRn-3 on GFX7 and SGPRn-5 3691 on GFX8. FLAT_SCRATCH_LO is 3692 used as the FLAT SCRATCH 3693 SIZE in flat memory 3694 instructions. Having CP load 3695 it once avoids loading it at 3696 the beginning of every 3697 wavefront. 3698 GFX9-GFX10 3699 This is the 3700 64-bit base address of the 3701 per SPI scratch backing 3702 memory managed by SPI for 3703 the queue executing the 3704 kernel dispatch. CP obtains 3705 this from the runtime (and 3706 divides it if there are 3707 multiple Shader Arrays each 3708 with its own SPI). The value 3709 of Scratch Wavefront Offset must 3710 be added by the kernel 3711 machine code and the result 3712 moved to the FLAT_SCRATCH 3713 SGPR which is SGPRn-6 and 3714 SGPRn-5. It is used as the 3715 FLAT SCRATCH BASE in flat 3716 memory instructions. 3717 then Private Segment Size 1 The 32-bit byte size of a 3718 (enable_sgpr_private single 3719 work-item's 3720 scratch_segment_size) memory 3721 allocation. This is the 3722 value from the kernel 3723 dispatch packet Private 3724 Segment Byte Size rounded up 3725 by CP to a multiple of 3726 DWORD. 3727 3728 Having CP load it once avoids 3729 loading it at the beginning of 3730 every wavefront. 3731 3732 This is not used for 3733 GFX7-GFX8 since it is the same 3734 value as the second SGPR of 3735 Flat Scratch Init. However, it 3736 may be needed for GFX9-GFX10 which 3737 changes the meaning of the 3738 Flat Scratch Init value. 3739 then Grid Work-Group Count X 1 32-bit count of the number of 3740 (enable_sgpr_grid work-groups in the X dimension 3741 _workgroup_count_X) for the grid being 3742 executed. Computed from the 3743 fields in the kernel dispatch 3744 packet as ((grid_size.x + 3745 workgroup_size.x - 1) / 3746 workgroup_size.x). 3747 then Grid Work-Group Count Y 1 32-bit count of the number of 3748 (enable_sgpr_grid work-groups in the Y dimension 3749 _workgroup_count_Y && for the grid being 3750 less than 16 previous executed. Computed from the 3751 SGPRs) fields in the kernel dispatch 3752 packet as ((grid_size.y + 3753 workgroup_size.y - 1) / 3754 workgroupSize.y). 3755 3756 Only initialized if <16 3757 previous SGPRs initialized. 3758 then Grid Work-Group Count Z 1 32-bit count of the number of 3759 (enable_sgpr_grid work-groups in the Z dimension 3760 _workgroup_count_Z && for the grid being 3761 less than 16 previous executed. Computed from the 3762 SGPRs) fields in the kernel dispatch 3763 packet as ((grid_size.z + 3764 workgroup_size.z - 1) / 3765 workgroupSize.z). 3766 3767 Only initialized if <16 3768 previous SGPRs initialized. 3769 then Work-Group Id X 1 32-bit work-group id in X 3770 (enable_sgpr_workgroup_id dimension of grid for 3771 _X) wavefront. 3772 then Work-Group Id Y 1 32-bit work-group id in Y 3773 (enable_sgpr_workgroup_id dimension of grid for 3774 _Y) wavefront. 3775 then Work-Group Id Z 1 32-bit work-group id in Z 3776 (enable_sgpr_workgroup_id dimension of grid for 3777 _Z) wavefront. 3778 then Work-Group Info 1 {first_wavefront, 14'b0000, 3779 (enable_sgpr_workgroup ordered_append_term[10:0], 3780 _info) threadgroup_size_in_wavefronts[5:0]} 3781 then Scratch Wavefront Offset 1 32-bit byte offset from base 3782 (enable_sgpr_private of scratch base of queue 3783 _segment_wavefront_offset) executing the kernel 3784 dispatch. Must be used as an 3785 offset with Private 3786 segment address when using 3787 Scratch Segment Buffer. It 3788 must be used to set up FLAT 3789 SCRATCH for flat addressing 3790 (see 3791 :ref:`amdgpu-amdhsa-kernel-prolog-flat-scratch`). 3792 ========== ========================== ====== ============================== 3793 3794The order of the VGPR registers is defined, but the compiler can specify which 3795ones are actually setup in the kernel descriptor using the ``enable_vgpr*`` bit 3796fields (see :ref:`amdgpu-amdhsa-kernel-descriptor`). The register numbers used 3797for enabled registers are dense starting at VGPR0: the first enabled register is 3798VGPR0, the next enabled register is VGPR1 etc.; disabled registers do not have a 3799VGPR number. 3800 3801VGPR register initial state is defined in 3802:ref:`amdgpu-amdhsa-vgpr-register-set-up-order-table`. 3803 3804 .. table:: VGPR Register Set Up Order 3805 :name: amdgpu-amdhsa-vgpr-register-set-up-order-table 3806 3807 ========== ========================== ====== ============================== 3808 VGPR Order Name Number Description 3809 (kernel descriptor enable of 3810 field) VGPRs 3811 ========== ========================== ====== ============================== 3812 First Work-Item Id X 1 32-bit work item id in X 3813 (Always initialized) dimension of work-group for 3814 wavefront lane. 3815 then Work-Item Id Y 1 32-bit work item id in Y 3816 (enable_vgpr_workitem_id dimension of work-group for 3817 > 0) wavefront lane. 3818 then Work-Item Id Z 1 32-bit work item id in Z 3819 (enable_vgpr_workitem_id dimension of work-group for 3820 > 1) wavefront lane. 3821 ========== ========================== ====== ============================== 3822 3823The setting of registers is done by GPU CP/ADC/SPI hardware as follows: 3824 38251. SGPRs before the Work-Group Ids are set by CP using the 16 User Data 3826 registers. 38272. Work-group Id registers X, Y, Z are set by ADC which supports any 3828 combination including none. 38293. Scratch Wavefront Offset is set by SPI in a per wavefront basis which is why 3830 its value cannot be included with the flat scratch init value which is per 3831 queue. 38324. The VGPRs are set by SPI which only supports specifying either (X), (X, Y) 3833 or (X, Y, Z). 3834 3835Flat Scratch register pair are adjacent SGPRs so they can be moved as a 64-bit 3836value to the hardware required SGPRn-3 and SGPRn-4 respectively. 3837 3838The global segment can be accessed either using buffer instructions (GFX6 which 3839has V# 64-bit address support), flat instructions (GFX7-GFX10), or global 3840instructions (GFX9-GFX10). 3841 3842If buffer operations are used, then the compiler can generate a V# with the 3843following properties: 3844 3845* base address of 0 3846* no swizzle 3847* ATC: 1 if IOMMU present (such as APU) 3848* ptr64: 1 3849* MTYPE set to support memory coherence that matches the runtime (such as CC for 3850 APU and NC for dGPU). 3851 3852.. _amdgpu-amdhsa-kernel-prolog: 3853 3854Kernel Prolog 3855~~~~~~~~~~~~~ 3856 3857The compiler performs initialization in the kernel prologue depending on the 3858target and information about things like stack usage in the kernel and called 3859functions. Some of this initialization requires the compiler to request certain 3860User and System SGPRs be present in the 3861:ref:`amdgpu-amdhsa-initial-kernel-execution-state` via the 3862:ref:`amdgpu-amdhsa-kernel-descriptor`. 3863 3864.. _amdgpu-amdhsa-kernel-prolog-cfi: 3865 3866CFI 3867+++ 3868 38691. The CFI return address is undefined. 3870 38712. The CFI CFA is defined using an expression which evaluates to a location 3872 description that comprises one memory location description for the 3873 ``DW_ASPACE_AMDGPU_private_lane`` address space address ``0``. 3874 3875.. _amdgpu-amdhsa-kernel-prolog-m0: 3876 3877M0 3878++ 3879 3880GFX6-GFX8 3881 The M0 register must be initialized with a value at least the total LDS size 3882 if the kernel may access LDS via DS or flat operations. Total LDS size is 3883 available in dispatch packet. For M0, it is also possible to use maximum 3884 possible value of LDS for given target (0x7FFF for GFX6 and 0xFFFF for 3885 GFX7-GFX8). 3886GFX9-GFX10 3887 The M0 register is not used for range checking LDS accesses and so does not 3888 need to be initialized in the prolog. 3889 3890.. _amdgpu-amdhsa-kernel-prolog-stack-pointer: 3891 3892Stack Pointer 3893+++++++++++++ 3894 3895If the kernel has function calls it must set up the ABI stack pointer described 3896in :ref:`amdgpu-amdhsa-function-call-convention-non-kernel-functions` by setting 3897SGPR32 to the unswizzled scratch offset of the address past the last local 3898allocation. 3899 3900.. _amdgpu-amdhsa-kernel-prolog-frame-pointer: 3901 3902Frame Pointer 3903+++++++++++++ 3904 3905If the kernel needs a frame pointer for the reasons defined in 3906``SIFrameLowering`` then SGPR33 is used and is always set to ``0`` in the 3907kernel prolog. If a frame pointer is not required then all uses of the frame 3908pointer are replaced with immediate ``0`` offsets. 3909 3910.. _amdgpu-amdhsa-kernel-prolog-flat-scratch: 3911 3912Flat Scratch 3913++++++++++++ 3914 3915If the kernel or any function it calls may use flat operations to access 3916scratch memory, the prolog code must set up the FLAT_SCRATCH register pair 3917(FLAT_SCRATCH_LO/FLAT_SCRATCH_HI which are in SGPRn-4/SGPRn-3). Initialization 3918uses Flat Scratch Init and Scratch Wavefront Offset SGPR registers (see 3919:ref:`amdgpu-amdhsa-initial-kernel-execution-state`): 3920 3921GFX6 3922 Flat scratch is not supported. 3923 3924GFX7-GFX8 3925 3926 1. The low word of Flat Scratch Init is 32-bit byte offset from 3927 ``SH_HIDDEN_PRIVATE_BASE_VIMID`` to the base of scratch backing memory 3928 being managed by SPI for the queue executing the kernel dispatch. This is 3929 the same value used in the Scratch Segment Buffer V# base address. The 3930 prolog must add the value of Scratch Wavefront Offset to get the 3931 wavefront's byte scratch backing memory offset from 3932 ``SH_HIDDEN_PRIVATE_BASE_VIMID``. Since FLAT_SCRATCH_LO is in units of 256 3933 bytes, the offset must be right shifted by 8 before moving into 3934 FLAT_SCRATCH_LO. 3935 2. The second word of Flat Scratch Init is 32-bit byte size of a single 3936 work-items scratch memory usage. This is directly loaded from the kernel 3937 dispatch packet Private Segment Byte Size and rounded up to a multiple of 3938 DWORD. Having CP load it once avoids loading it at the beginning of every 3939 wavefront. The prolog must move it to FLAT_SCRATCH_LO for use as FLAT 3940 SCRATCH SIZE. 3941 3942GFX9-GFX10 3943 The Flat Scratch Init is the 64-bit address of the base of scratch backing 3944 memory being managed by SPI for the queue executing the kernel dispatch. The 3945 prolog must add the value of Scratch Wavefront Offset and moved to the 3946 FLAT_SCRATCH pair for use as the flat scratch base in flat memory 3947 instructions. 3948 3949.. _amdgpu-amdhsa-kernel-prolog-private-segment-buffer: 3950 3951Private Segment Buffer 3952++++++++++++++++++++++ 3953 3954A set of four SGPRs beginning at a four-aligned SGPR index are always selected 3955to serve as the scratch V# for the kernel as follows: 3956 3957 - If it is known during instruction selection that there is stack usage, 3958 SGPR0-3 is reserved for use as the scratch V#. Stack usage is assumed if 3959 optimizations are disabled (``-O0``), if stack objects already exist (for 3960 locals, etc.), or if there are any function calls. 3961 3962 - Otherwise, four high numbered SGPRs beginning at a four-aligned SGPR index 3963 are reserved for the tentative scratch V#. These will be used if it is 3964 determined that spilling is needed. 3965 3966 - If no use is made of the tentative scratch V#, then it is unreserved, 3967 and the register count is determined ignoring it. 3968 - If use is made of the tentative scratch V#, then its register numbers 3969 are shifted to the first four-aligned SGPR index after the highest one 3970 allocated by the register allocator, and all uses are updated. The 3971 register count includes them in the shifted location. 3972 - In either case, if the processor has the SGPR allocation bug, the 3973 tentative allocation is not shifted or unreserved in order to ensure 3974 the register count is higher to workaround the bug. 3975 3976 .. note:: 3977 3978 This approach of using a tentative scratch V# and shifting the register 3979 numbers if used avoids having to perform register allocation a second 3980 time if the tentative V# is eliminated. This is more efficient and 3981 avoids the problem that the second register allocation may perform 3982 spilling which will fail as there is no longer a scratch V#. 3983 3984When the kernel prolog code is being emitted it is known whether the scratch V# 3985described above is actually used. If it is, the prolog code must set it up by 3986copying the Private Segment Buffer to the scratch V# registers and then adding 3987the Private Segment Wavefront Offset to the queue base address in the V#. The 3988result is a V# with a base address pointing to the beginning of the wavefront 3989scratch backing memory. 3990 3991The Private Segment Buffer is always requested, but the Private Segment 3992Wavefront Offset is only requested if it is used (see 3993:ref:`amdgpu-amdhsa-initial-kernel-execution-state`). 3994 3995.. _amdgpu-amdhsa-memory-model: 3996 3997Memory Model 3998~~~~~~~~~~~~ 3999 4000This section describes the mapping of LLVM memory model onto AMDGPU machine code 4001(see :ref:`memmodel`). 4002 4003The AMDGPU backend supports the memory synchronization scopes specified in 4004:ref:`amdgpu-memory-scopes`. 4005 4006The code sequences used to implement the memory model are defined in table 4007:ref:`amdgpu-amdhsa-memory-model-code-sequences-gfx6-gfx10-table`. 4008 4009The sequences specify the order of instructions that a single thread must 4010execute. The ``s_waitcnt`` and ``buffer_wbinvl1_vol`` are defined with respect 4011to other memory instructions executed by the same thread. This allows them to be 4012moved earlier or later which can allow them to be combined with other instances 4013of the same instruction, or hoisted/sunk out of loops to improve 4014performance. Only the instructions related to the memory model are given; 4015additional ``s_waitcnt`` instructions are required to ensure registers are 4016defined before being used. These may be able to be combined with the memory 4017model ``s_waitcnt`` instructions as described above. 4018 4019The AMDGPU backend supports the following memory models: 4020 4021 HSA Memory Model [HSA]_ 4022 The HSA memory model uses a single happens-before relation for all address 4023 spaces (see :ref:`amdgpu-address-spaces`). 4024 OpenCL Memory Model [OpenCL]_ 4025 The OpenCL memory model which has separate happens-before relations for the 4026 global and local address spaces. Only a fence specifying both global and 4027 local address space, and seq_cst instructions join the relationships. Since 4028 the LLVM ``memfence`` instruction does not allow an address space to be 4029 specified the OpenCL fence has to conservatively assume both local and 4030 global address space was specified. However, optimizations can often be 4031 done to eliminate the additional ``s_waitcnt`` instructions when there are 4032 no intervening memory instructions which access the corresponding address 4033 space. The code sequences in the table indicate what can be omitted for the 4034 OpenCL memory. The target triple environment is used to determine if the 4035 source language is OpenCL (see :ref:`amdgpu-opencl`). 4036 4037``ds/flat_load/store/atomic`` instructions to local memory are termed LDS 4038operations. 4039 4040``buffer/global/flat_load/store/atomic`` instructions to global memory are 4041termed vector memory operations. 4042 4043For GFX6-GFX9: 4044 4045* Each agent has multiple shader arrays (SA). 4046* Each SA has multiple compute units (CU). 4047* Each CU has multiple SIMDs that execute wavefronts. 4048* The wavefronts for a single work-group are executed in the same CU but may be 4049 executed by different SIMDs. 4050* Each CU has a single LDS memory shared by the wavefronts of the work-groups 4051 executing on it. 4052* All LDS operations of a CU are performed as wavefront wide operations in a 4053 global order and involve no caching. Completion is reported to a wavefront in 4054 execution order. 4055* The LDS memory has multiple request queues shared by the SIMDs of a 4056 CU. Therefore, the LDS operations performed by different wavefronts of a 4057 work-group can be reordered relative to each other, which can result in 4058 reordering the visibility of vector memory operations with respect to LDS 4059 operations of other wavefronts in the same work-group. A ``s_waitcnt 4060 lgkmcnt(0)`` is required to ensure synchronization between LDS operations and 4061 vector memory operations between wavefronts of a work-group, but not between 4062 operations performed by the same wavefront. 4063* The vector memory operations are performed as wavefront wide operations and 4064 completion is reported to a wavefront in execution order. The exception is 4065 that for GFX7-GFX9 ``flat_load/store/atomic`` instructions can report out of 4066 vector memory order if they access LDS memory, and out of LDS operation order 4067 if they access global memory. 4068* The vector memory operations access a single vector L1 cache shared by all 4069 SIMDs a CU. Therefore, no special action is required for coherence between the 4070 lanes of a single wavefront, or for coherence between wavefronts in the same 4071 work-group. A ``buffer_wbinvl1_vol`` is required for coherence between 4072 wavefronts executing in different work-groups as they may be executing on 4073 different CUs. 4074* The scalar memory operations access a scalar L1 cache shared by all wavefronts 4075 on a group of CUs. The scalar and vector L1 caches are not coherent. However, 4076 scalar operations are used in a restricted way so do not impact the memory 4077 model. See :ref:`amdgpu-address-spaces`. 4078* The vector and scalar memory operations use an L2 cache shared by all CUs on 4079 the same agent. 4080* The L2 cache has independent channels to service disjoint ranges of virtual 4081 addresses. 4082* Each CU has a separate request queue per channel. Therefore, the vector and 4083 scalar memory operations performed by wavefronts executing in different 4084 work-groups (which may be executing on different CUs) of an agent can be 4085 reordered relative to each other. A ``s_waitcnt vmcnt(0)`` is required to 4086 ensure synchronization between vector memory operations of different CUs. It 4087 ensures a previous vector memory operation has completed before executing a 4088 subsequent vector memory or LDS operation and so can be used to meet the 4089 requirements of acquire and release. 4090* The L2 cache can be kept coherent with other agents on some targets, or ranges 4091 of virtual addresses can be set up to bypass it to ensure system coherence. 4092 4093For GFX10: 4094 4095* Each agent has multiple shader arrays (SA). 4096* Each SA has multiple work-group processors (WGP). 4097* Each WGP has multiple compute units (CU). 4098* Each CU has multiple SIMDs that execute wavefronts. 4099* The wavefronts for a single work-group are executed in the same 4100 WGP. In CU wavefront execution mode the wavefronts may be executed by 4101 different SIMDs in the same CU. In WGP wavefront execution mode the 4102 wavefronts may be executed by different SIMDs in different CUs in the same 4103 WGP. 4104* Each WGP has a single LDS memory shared by the wavefronts of the work-groups 4105 executing on it. 4106* All LDS operations of a WGP are performed as wavefront wide operations in a 4107 global order and involve no caching. Completion is reported to a wavefront in 4108 execution order. 4109* The LDS memory has multiple request queues shared by the SIMDs of a 4110 WGP. Therefore, the LDS operations performed by different wavefronts of a 4111 work-group can be reordered relative to each other, which can result in 4112 reordering the visibility of vector memory operations with respect to LDS 4113 operations of other wavefronts in the same work-group. A ``s_waitcnt 4114 lgkmcnt(0)`` is required to ensure synchronization between LDS operations and 4115 vector memory operations between wavefronts of a work-group, but not between 4116 operations performed by the same wavefront. 4117* The vector memory operations are performed as wavefront wide operations. 4118 Completion of load/store/sample operations are reported to a wavefront in 4119 execution order of other load/store/sample operations performed by that 4120 wavefront. 4121* The vector memory operations access a vector L0 cache. There is a single L0 4122 cache per CU. Each SIMD of a CU accesses the same L0 cache. Therefore, no 4123 special action is required for coherence between the lanes of a single 4124 wavefront. However, a ``BUFFER_GL0_INV`` is required for coherence between 4125 wavefronts executing in the same work-group as they may be executing on SIMDs 4126 of different CUs that access different L0s. A ``BUFFER_GL0_INV`` is also 4127 required for coherence between wavefronts executing in different work-groups 4128 as they may be executing on different WGPs. 4129* The scalar memory operations access a scalar L0 cache shared by all wavefronts 4130 on a WGP. The scalar and vector L0 caches are not coherent. However, scalar 4131 operations are used in a restricted way so do not impact the memory model. See 4132 :ref:`amdgpu-address-spaces`. 4133* The vector and scalar memory L0 caches use an L1 cache shared by all WGPs on 4134 the same SA. Therefore, no special action is required for coherence between 4135 the wavefronts of a single work-group. However, a ``BUFFER_GL1_INV`` is 4136 required for coherence between wavefronts executing in different work-groups 4137 as they may be executing on different SAs that access different L1s. 4138* The L1 caches have independent quadrants to service disjoint ranges of virtual 4139 addresses. 4140* Each L0 cache has a separate request queue per L1 quadrant. Therefore, the 4141 vector and scalar memory operations performed by different wavefronts, whether 4142 executing in the same or different work-groups (which may be executing on 4143 different CUs accessing different L0s), can be reordered relative to each 4144 other. A ``s_waitcnt vmcnt(0) & vscnt(0)`` is required to ensure 4145 synchronization between vector memory operations of different wavefronts. It 4146 ensures a previous vector memory operation has completed before executing a 4147 subsequent vector memory or LDS operation and so can be used to meet the 4148 requirements of acquire, release and sequential consistency. 4149* The L1 caches use an L2 cache shared by all SAs on the same agent. 4150* The L2 cache has independent channels to service disjoint ranges of virtual 4151 addresses. 4152* Each L1 quadrant of a single SA accesses a different L2 channel. Each L1 4153 quadrant has a separate request queue per L2 channel. Therefore, the vector 4154 and scalar memory operations performed by wavefronts executing in different 4155 work-groups (which may be executing on different SAs) of an agent can be 4156 reordered relative to each other. A ``s_waitcnt vmcnt(0) & vscnt(0)`` is 4157 required to ensure synchronization between vector memory operations of 4158 different SAs. It ensures a previous vector memory operation has completed 4159 before executing a subsequent vector memory and so can be used to meet the 4160 requirements of acquire, release and sequential consistency. 4161* The L2 cache can be kept coherent with other agents on some targets, or ranges 4162 of virtual addresses can be set up to bypass it to ensure system coherence. 4163 4164Private address space uses ``buffer_load/store`` using the scratch V# 4165(GFX6-GFX8), or ``scratch_load/store`` (GFX9-GFX10). Since only a single thread 4166is accessing the memory, atomic memory orderings are not meaningful, and all 4167accesses are treated as non-atomic. 4168 4169Constant address space uses ``buffer/global_load`` instructions (or equivalent 4170scalar memory instructions). Since the constant address space contents do not 4171change during the execution of a kernel dispatch it is not legal to perform 4172stores, and atomic memory orderings are not meaningful, and all access are 4173treated as non-atomic. 4174 4175A memory synchronization scope wider than work-group is not meaningful for the 4176group (LDS) address space and is treated as work-group. 4177 4178The memory model does not support the region address space which is treated as 4179non-atomic. 4180 4181Acquire memory ordering is not meaningful on store atomic instructions and is 4182treated as non-atomic. 4183 4184Release memory ordering is not meaningful on load atomic instructions and is 4185treated a non-atomic. 4186 4187Acquire-release memory ordering is not meaningful on load or store atomic 4188instructions and is treated as acquire and release respectively. 4189 4190AMDGPU backend only uses scalar memory operations to access memory that is 4191proven to not change during the execution of the kernel dispatch. This includes 4192constant address space and global address space for program scope const 4193variables. Therefore, the kernel machine code does not have to maintain the 4194scalar L1 cache to ensure it is coherent with the vector L1 cache. The scalar 4195and vector L1 caches are invalidated between kernel dispatches by CP since 4196constant address space data may change between kernel dispatch executions. See 4197:ref:`amdgpu-address-spaces`. 4198 4199The one exception is if scalar writes are used to spill SGPR registers. In this 4200case the AMDGPU backend ensures the memory location used to spill is never 4201accessed by vector memory operations at the same time. If scalar writes are used 4202then a ``s_dcache_wb`` is inserted before the ``s_endpgm`` and before a function 4203return since the locations may be used for vector memory instructions by a 4204future wavefront that uses the same scratch area, or a function call that 4205creates a frame at the same address, respectively. There is no need for a 4206``s_dcache_inv`` as all scalar writes are write-before-read in the same thread. 4207 4208For GFX6-GFX9, scratch backing memory (which is used for the private address 4209space) is accessed with MTYPE NC_NV (non-coherent non-volatile). Since the 4210private address space is only accessed by a single thread, and is always 4211write-before-read, there is never a need to invalidate these entries from the L1 4212cache. Hence all cache invalidates are done as ``*_vol`` to only invalidate the 4213volatile cache lines. 4214 4215For GFX10, scratch backing memory (which is used for the private address space) 4216is accessed with MTYPE NC (non-coherent). Since the private address space is 4217only accessed by a single thread, and is always write-before-read, there is 4218never a need to invalidate these entries from the L0 or L1 caches. 4219 4220For GFX10, wavefronts are executed in native mode with in-order reporting of 4221loads and sample instructions. In this mode vmcnt reports completion of load, 4222atomic with return and sample instructions in order, and the vscnt reports the 4223completion of store and atomic without return in order. See ``MEM_ORDERED`` 4224field in :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`. 4225 4226In GFX10, wavefronts can be executed in WGP or CU wavefront execution mode: 4227 4228* In WGP wavefront execution mode the wavefronts of a work-group are executed 4229 on the SIMDs of both CUs of the WGP. Therefore, explicit management of the per 4230 CU L0 caches is required for work-group synchronization. Also accesses to L1 4231 at work-group scope need to be explicitly ordered as the accesses from 4232 different CUs are not ordered. 4233* In CU wavefront execution mode the wavefronts of a work-group are executed on 4234 the SIMDs of a single CU of the WGP. Therefore, all global memory access by 4235 the work-group access the same L0 which in turn ensures L1 accesses are 4236 ordered and so do not require explicit management of the caches for 4237 work-group synchronization. 4238 4239See ``WGP_MODE`` field in 4240:ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table` and 4241:ref:`amdgpu-target-features`. 4242 4243On dGPU the kernarg backing memory is accessed as UC (uncached) to avoid needing 4244to invalidate the L2 cache. For GFX6-GFX9, this also causes it to be treated as 4245non-volatile and so is not invalidated by ``*_vol``. On APU it is accessed as CC 4246(cache coherent) and so the L2 cache will be coherent with the CPU and other 4247agents. 4248 4249 .. table:: AMDHSA Memory Model Code Sequences GFX6-GFX10 4250 :name: amdgpu-amdhsa-memory-model-code-sequences-gfx6-gfx10-table 4251 4252 ============ ============ ============== ========== =============================== ================================== 4253 LLVM Instr LLVM Memory LLVM Memory AMDGPU AMDGPU Machine Code AMDGPU Machine Code 4254 Ordering Sync Scope Address GFX6-9 GFX10 4255 Space 4256 ============ ============ ============== ========== =============================== ================================== 4257 **Non-Atomic** 4258 ---------------------------------------------------------------------------------------------------------------------- 4259 load *none* *none* - global - !volatile & !nontemporal - !volatile & !nontemporal 4260 - generic 4261 - private 1. buffer/global/flat_load 1. buffer/global/flat_load 4262 - constant 4263 - volatile & !nontemporal - volatile & !nontemporal 4264 4265 1. buffer/global/flat_load 1. buffer/global/flat_load 4266 glc=1 glc=1 dlc=1 4267 4268 - nontemporal - nontemporal 4269 4270 1. buffer/global/flat_load 1. buffer/global/flat_load 4271 glc=1 slc=1 slc=1 4272 4273 load *none* *none* - local 1. ds_load 1. ds_load 4274 store *none* *none* - global - !nontemporal - !nontemporal 4275 - generic 4276 - private 1. buffer/global/flat_store 1. buffer/global/flat_store 4277 - constant 4278 - nontemporal - nontemporal 4279 4280 1. buffer/global/flat_store 1. buffer/global/flat_store 4281 glc=1 slc=1 slc=1 4282 4283 store *none* *none* - local 1. ds_store 1. ds_store 4284 **Unordered Atomic** 4285 ---------------------------------------------------------------------------------------------------------------------- 4286 load atomic unordered *any* *any* *Same as non-atomic*. *Same as non-atomic*. 4287 store atomic unordered *any* *any* *Same as non-atomic*. *Same as non-atomic*. 4288 atomicrmw unordered *any* *any* *Same as monotonic *Same as monotonic 4289 atomic*. atomic*. 4290 **Monotonic Atomic** 4291 ---------------------------------------------------------------------------------------------------------------------- 4292 load atomic monotonic - singlethread - global 1. buffer/global/flat_load 1. buffer/global/flat_load 4293 - wavefront - generic 4294 load atomic monotonic - workgroup - global 1. buffer/global/flat_load 1. buffer/global/flat_load 4295 - generic glc=1 4296 4297 - If CU wavefront execution mode, omit glc=1. 4298 4299 load atomic monotonic - singlethread - local 1. ds_load 1. ds_load 4300 - wavefront 4301 - workgroup 4302 load atomic monotonic - agent - global 1. buffer/global/flat_load 1. buffer/global/flat_load 4303 - system - generic glc=1 glc=1 dlc=1 4304 store atomic monotonic - singlethread - global 1. buffer/global/flat_store 1. buffer/global/flat_store 4305 - wavefront - generic 4306 - workgroup 4307 - agent 4308 - system 4309 store atomic monotonic - singlethread - local 1. ds_store 1. ds_store 4310 - wavefront 4311 - workgroup 4312 atomicrmw monotonic - singlethread - global 1. buffer/global/flat_atomic 1. buffer/global/flat_atomic 4313 - wavefront - generic 4314 - workgroup 4315 - agent 4316 - system 4317 atomicrmw monotonic - singlethread - local 1. ds_atomic 1. ds_atomic 4318 - wavefront 4319 - workgroup 4320 **Acquire Atomic** 4321 ---------------------------------------------------------------------------------------------------------------------- 4322 load atomic acquire - singlethread - global 1. buffer/global/ds/flat_load 1. buffer/global/ds/flat_load 4323 - wavefront - local 4324 - generic 4325 load atomic acquire - workgroup - global 1. buffer/global/flat_load 1. buffer/global_load glc=1 4326 4327 - If CU wavefront execution mode, omit glc=1. 4328 4329 2. s_waitcnt vmcnt(0) 4330 4331 - If CU wavefront execution mode, omit. 4332 - Must happen before 4333 the following buffer_gl0_inv 4334 and before any following 4335 global/generic 4336 load/load 4337 atomic/store/store 4338 atomic/atomicrmw. 4339 4340 3. buffer_gl0_inv 4341 4342 - If CU wavefront execution mode, omit. 4343 - Ensures that 4344 following 4345 loads will not see 4346 stale data. 4347 4348 load atomic acquire - workgroup - local 1. ds_load 1. ds_load 4349 2. s_waitcnt lgkmcnt(0) 2. s_waitcnt lgkmcnt(0) 4350 4351 - If OpenCL, omit. - If OpenCL, omit. 4352 - Must happen before - Must happen before 4353 any following the following buffer_gl0_inv 4354 global/generic and before any following 4355 load/load global/generic load/load 4356 atomic/store/store atomic/store/store 4357 atomic/atomicrmw. atomic/atomicrmw. 4358 - Ensures any - Ensures any 4359 following global following global 4360 data read is no data read is no 4361 older than the load older than the load 4362 atomic value being atomic value being 4363 acquired. acquired. 4364 4365 3. buffer_gl0_inv 4366 4367 - If CU wavefront execution mode, omit. 4368 - If OpenCL, omit. 4369 - Ensures that 4370 following 4371 loads will not see 4372 stale data. 4373 4374 load atomic acquire - workgroup - generic 1. flat_load 1. flat_load glc=1 4375 4376 - If CU wavefront execution mode, omit glc=1. 4377 4378 2. s_waitcnt lgkmcnt(0) 2. s_waitcnt lgkmcnt(0) & 4379 vmcnt(0) 4380 4381 - If CU wavefront execution mode, omit vmcnt. 4382 - If OpenCL, omit. - If OpenCL, omit 4383 lgkmcnt(0). 4384 - Must happen before - Must happen before 4385 any following the following 4386 global/generic buffer_gl0_inv and any 4387 load/load following global/generic 4388 atomic/store/store load/load 4389 atomic/atomicrmw. atomic/store/store 4390 atomic/atomicrmw. 4391 - Ensures any - Ensures any 4392 following global following global 4393 data read is no data read is no 4394 older than the load older than the load 4395 atomic value being atomic value being 4396 acquired. acquired. 4397 4398 3. buffer_gl0_inv 4399 4400 - If CU wavefront execution mode, omit. 4401 - Ensures that 4402 following 4403 loads will not see 4404 stale data. 4405 4406 load atomic acquire - agent - global 1. buffer/global/flat_load 1. buffer/global_load 4407 - system glc=1 glc=1 dlc=1 4408 2. s_waitcnt vmcnt(0) 2. s_waitcnt vmcnt(0) 4409 4410 - Must happen before - Must happen before 4411 following following 4412 buffer_wbinvl1_vol. buffer_gl*_inv. 4413 - Ensures the load - Ensures the load 4414 has completed has completed 4415 before invalidating before invalidating 4416 the cache. the caches. 4417 4418 3. buffer_wbinvl1_vol 3. buffer_gl0_inv; 4419 buffer_gl1_inv 4420 4421 - Must happen before - Must happen before 4422 any following any following 4423 global/generic global/generic 4424 load/load load/load 4425 atomic/atomicrmw. atomic/atomicrmw. 4426 - Ensures that - Ensures that 4427 following following 4428 loads will not see loads will not see 4429 stale global data. stale global data. 4430 4431 load atomic acquire - agent - generic 1. flat_load glc=1 1. flat_load glc=1 dlc=1 4432 - system 2. s_waitcnt vmcnt(0) & 2. s_waitcnt vmcnt(0) & 4433 lgkmcnt(0) lgkmcnt(0) 4434 4435 - If OpenCL omit - If OpenCL omit 4436 lgkmcnt(0). lgkmcnt(0). 4437 - Must happen before - Must happen before 4438 following following 4439 buffer_wbinvl1_vol. buffer_gl*_invl. 4440 - Ensures the flat_load - Ensures the flat_load 4441 has completed has completed 4442 before invalidating before invalidating 4443 the cache. the caches. 4444 4445 3. buffer_wbinvl1_vol 3. buffer_gl0_inv; 4446 buffer_gl1_inv 4447 4448 - Must happen before - Must happen before 4449 any following any following 4450 global/generic global/generic 4451 load/load load/load 4452 atomic/atomicrmw. atomic/atomicrmw. 4453 - Ensures that - Ensures that 4454 following loads following loads 4455 will not see stale will not see stale 4456 global data. global data. 4457 4458 atomicrmw acquire - singlethread - global 1. buffer/global/ds/flat_atomic 1. buffer/global/ds/flat_atomic 4459 - wavefront - local 4460 - generic 4461 atomicrmw acquire - workgroup - global 1. buffer/global/flat_atomic 1. buffer/global_atomic 4462 2. s_waitcnt vm/vscnt(0) 4463 4464 - If CU wavefront execution mode, omit. 4465 - Use vmcnt if atomic with 4466 return and vscnt if atomic 4467 with no-return. 4468 - Must happen before 4469 the following buffer_gl0_inv 4470 and before any following 4471 global/generic 4472 load/load 4473 atomic/store/store 4474 atomic/atomicrmw. 4475 4476 3. buffer_gl0_inv 4477 4478 - If CU wavefront execution mode, omit. 4479 - Ensures that 4480 following 4481 loads will not see 4482 stale data. 4483 4484 atomicrmw acquire - workgroup - local 1. ds_atomic 1. ds_atomic 4485 2. waitcnt lgkmcnt(0) 2. waitcnt lgkmcnt(0) 4486 4487 - If OpenCL, omit. - If OpenCL, omit. 4488 - Must happen before - Must happen before 4489 any following the following 4490 global/generic buffer_gl0_inv. 4491 load/load 4492 atomic/store/store 4493 atomic/atomicrmw. 4494 - Ensures any - Ensures any 4495 following global following global 4496 data read is no data read is no 4497 older than the older than the 4498 atomicrmw value atomicrmw value 4499 being acquired. being acquired. 4500 4501 3. buffer_gl0_inv 4502 4503 - If OpenCL omit. 4504 - Ensures that 4505 following 4506 loads will not see 4507 stale data. 4508 4509 atomicrmw acquire - workgroup - generic 1. flat_atomic 1. flat_atomic 4510 2. waitcnt lgkmcnt(0) 2. waitcnt lgkmcnt(0) & 4511 vm/vscnt(0) 4512 4513 - If CU wavefront execution mode, omit vm/vscnt. 4514 - If OpenCL, omit. - If OpenCL, omit 4515 waitcnt lgkmcnt(0).. 4516 - Use vmcnt if atomic with 4517 return and vscnt if atomic 4518 with no-return. 4519 waitcnt lgkmcnt(0). 4520 - Must happen before - Must happen before 4521 any following the following 4522 global/generic buffer_gl0_inv. 4523 load/load 4524 atomic/store/store 4525 atomic/atomicrmw. 4526 - Ensures any - Ensures any 4527 following global following global 4528 data read is no data read is no 4529 older than the older than the 4530 atomicrmw value atomicrmw value 4531 being acquired. being acquired. 4532 4533 3. buffer_gl0_inv 4534 4535 - If CU wavefront execution mode, omit. 4536 - Ensures that 4537 following 4538 loads will not see 4539 stale data. 4540 4541 atomicrmw acquire - agent - global 1. buffer/global/flat_atomic 1. buffer/global_atomic 4542 - system 2. s_waitcnt vmcnt(0) 2. s_waitcnt vm/vscnt(0) 4543 4544 - Use vmcnt if atomic with 4545 return and vscnt if atomic 4546 with no-return. 4547 waitcnt lgkmcnt(0). 4548 - Must happen before - Must happen before 4549 following following 4550 buffer_wbinvl1_vol. buffer_gl*_inv. 4551 - Ensures the - Ensures the 4552 atomicrmw has atomicrmw has 4553 completed before completed before 4554 invalidating the invalidating the 4555 cache. caches. 4556 4557 3. buffer_wbinvl1_vol 3. buffer_gl0_inv; 4558 buffer_gl1_inv 4559 4560 - Must happen before - Must happen before 4561 any following any following 4562 global/generic global/generic 4563 load/load load/load 4564 atomic/atomicrmw. atomic/atomicrmw. 4565 - Ensures that - Ensures that 4566 following loads following loads 4567 will not see stale will not see stale 4568 global data. global data. 4569 4570 atomicrmw acquire - agent - generic 1. flat_atomic 1. flat_atomic 4571 - system 2. s_waitcnt vmcnt(0) & 2. s_waitcnt vm/vscnt(0) & 4572 lgkmcnt(0) lgkmcnt(0) 4573 4574 - If OpenCL, omit - If OpenCL, omit 4575 lgkmcnt(0). lgkmcnt(0). 4576 - Use vmcnt if atomic with 4577 return and vscnt if atomic 4578 with no-return. 4579 - Must happen before - Must happen before 4580 following following 4581 buffer_wbinvl1_vol. buffer_gl*_inv. 4582 - Ensures the - Ensures the 4583 atomicrmw has atomicrmw has 4584 completed before completed before 4585 invalidating the invalidating the 4586 cache. caches. 4587 4588 3. buffer_wbinvl1_vol 3. buffer_gl0_inv; 4589 buffer_gl1_inv 4590 4591 - Must happen before - Must happen before 4592 any following any following 4593 global/generic global/generic 4594 load/load load/load 4595 atomic/atomicrmw. atomic/atomicrmw. 4596 - Ensures that - Ensures that 4597 following loads following loads 4598 will not see stale will not see stale 4599 global data. global data. 4600 4601 fence acquire - singlethread *none* *none* *none* 4602 - wavefront 4603 fence acquire - workgroup *none* 1. s_waitcnt lgkmcnt(0) 1. s_waitcnt lgkmcnt(0) & 4604 vmcnt(0) & vscnt(0) 4605 4606 - If CU wavefront execution mode, omit vmcnt and 4607 vscnt. 4608 - If OpenCL and - If OpenCL and 4609 address space is address space is 4610 not generic, omit. not generic, omit 4611 lgkmcnt(0). 4612 - If OpenCL and 4613 address space is 4614 local, omit 4615 vmcnt(0) and vscnt(0). 4616 - However, since LLVM - However, since LLVM 4617 currently has no currently has no 4618 address space on address space on 4619 the fence need to the fence need to 4620 conservatively conservatively 4621 always generate. If always generate. If 4622 fence had an fence had an 4623 address space then address space then 4624 set to address set to address 4625 space of OpenCL space of OpenCL 4626 fence flag, or to fence flag, or to 4627 generic if both generic if both 4628 local and global local and global 4629 flags are flags are 4630 specified. specified. 4631 - Must happen after 4632 any preceding 4633 local/generic load 4634 atomic/atomicrmw 4635 with an equal or 4636 wider sync scope 4637 and memory ordering 4638 stronger than 4639 unordered (this is 4640 termed the 4641 fence-paired-atomic). 4642 - Must happen before 4643 any following 4644 global/generic 4645 load/load 4646 atomic/store/store 4647 atomic/atomicrmw. 4648 - Ensures any 4649 following global 4650 data read is no 4651 older than the 4652 value read by the 4653 fence-paired-atomic. 4654 - Could be split into 4655 separate s_waitcnt 4656 vmcnt(0), s_waitcnt 4657 vscnt(0) and s_waitcnt 4658 lgkmcnt(0) to allow 4659 them to be 4660 independently moved 4661 according to the 4662 following rules. 4663 - s_waitcnt vmcnt(0) 4664 must happen after 4665 any preceding 4666 global/generic load 4667 atomic/ 4668 atomicrmw-with-return-value 4669 with an equal or 4670 wider sync scope 4671 and memory ordering 4672 stronger than 4673 unordered (this is 4674 termed the 4675 fence-paired-atomic). 4676 - s_waitcnt vscnt(0) 4677 must happen after 4678 any preceding 4679 global/generic 4680 atomicrmw-no-return-value 4681 with an equal or 4682 wider sync scope 4683 and memory ordering 4684 stronger than 4685 unordered (this is 4686 termed the 4687 fence-paired-atomic). 4688 - s_waitcnt lgkmcnt(0) 4689 must happen after 4690 any preceding 4691 local/generic load 4692 atomic/atomicrmw 4693 with an equal or 4694 wider sync scope 4695 and memory ordering 4696 stronger than 4697 unordered (this is 4698 termed the 4699 fence-paired-atomic). 4700 - Must happen before 4701 the following 4702 buffer_gl0_inv. 4703 - Ensures that the 4704 fence-paired atomic 4705 has completed 4706 before invalidating 4707 the 4708 cache. Therefore 4709 any following 4710 locations read must 4711 be no older than 4712 the value read by 4713 the 4714 fence-paired-atomic. 4715 4716 3. buffer_gl0_inv 4717 4718 - If CU wavefront execution mode, omit. 4719 - Ensures that 4720 following 4721 loads will not see 4722 stale data. 4723 4724 fence acquire - agent *none* 1. s_waitcnt lgkmcnt(0) & 1. s_waitcnt lgkmcnt(0) & 4725 - system vmcnt(0) vmcnt(0) & vscnt(0) 4726 4727 - If OpenCL and - If OpenCL and 4728 address space is address space is 4729 not generic, omit not generic, omit 4730 lgkmcnt(0). lgkmcnt(0). 4731 - If OpenCL and 4732 address space is 4733 local, omit 4734 vmcnt(0) and vscnt(0). 4735 - However, since LLVM - However, since LLVM 4736 currently has no currently has no 4737 address space on address space on 4738 the fence need to the fence need to 4739 conservatively conservatively 4740 always generate always generate 4741 (see comment for (see comment for 4742 previous fence). previous fence). 4743 - Could be split into 4744 separate s_waitcnt 4745 vmcnt(0) and 4746 s_waitcnt 4747 lgkmcnt(0) to allow 4748 them to be 4749 independently moved 4750 according to the 4751 following rules. 4752 - s_waitcnt vmcnt(0) 4753 must happen after 4754 any preceding 4755 global/generic load 4756 atomic/atomicrmw 4757 with an equal or 4758 wider sync scope 4759 and memory ordering 4760 stronger than 4761 unordered (this is 4762 termed the 4763 fence-paired-atomic). 4764 - s_waitcnt lgkmcnt(0) 4765 must happen after 4766 any preceding 4767 local/generic load 4768 atomic/atomicrmw 4769 with an equal or 4770 wider sync scope 4771 and memory ordering 4772 stronger than 4773 unordered (this is 4774 termed the 4775 fence-paired-atomic). 4776 - Must happen before 4777 the following 4778 buffer_wbinvl1_vol. 4779 - Ensures that the 4780 fence-paired atomic 4781 has completed 4782 before invalidating 4783 the 4784 cache. Therefore 4785 any following 4786 locations read must 4787 be no older than 4788 the value read by 4789 the 4790 fence-paired-atomic. 4791 - Could be split into 4792 separate s_waitcnt 4793 vmcnt(0), s_waitcnt 4794 vscnt(0) and s_waitcnt 4795 lgkmcnt(0) to allow 4796 them to be 4797 independently moved 4798 according to the 4799 following rules. 4800 - s_waitcnt vmcnt(0) 4801 must happen after 4802 any preceding 4803 global/generic load 4804 atomic/ 4805 atomicrmw-with-return-value 4806 with an equal or 4807 wider sync scope 4808 and memory ordering 4809 stronger than 4810 unordered (this is 4811 termed the 4812 fence-paired-atomic). 4813 - s_waitcnt vscnt(0) 4814 must happen after 4815 any preceding 4816 global/generic 4817 atomicrmw-no-return-value 4818 with an equal or 4819 wider sync scope 4820 and memory ordering 4821 stronger than 4822 unordered (this is 4823 termed the 4824 fence-paired-atomic). 4825 - s_waitcnt lgkmcnt(0) 4826 must happen after 4827 any preceding 4828 local/generic load 4829 atomic/atomicrmw 4830 with an equal or 4831 wider sync scope 4832 and memory ordering 4833 stronger than 4834 unordered (this is 4835 termed the 4836 fence-paired-atomic). 4837 - Must happen before 4838 the following 4839 buffer_gl*_inv. 4840 - Ensures that the 4841 fence-paired atomic 4842 has completed 4843 before invalidating 4844 the 4845 caches. Therefore 4846 any following 4847 locations read must 4848 be no older than 4849 the value read by 4850 the 4851 fence-paired-atomic. 4852 4853 2. buffer_wbinvl1_vol 2. buffer_gl0_inv; 4854 buffer_gl1_inv 4855 4856 - Must happen before any - Must happen before any 4857 following global/generic following global/generic 4858 load/load load/load 4859 atomic/store/store atomic/store/store 4860 atomic/atomicrmw. atomic/atomicrmw. 4861 - Ensures that - Ensures that 4862 following loads following loads 4863 will not see stale will not see stale 4864 global data. global data. 4865 4866 **Release Atomic** 4867 ---------------------------------------------------------------------------------------------------------------------- 4868 store atomic release - singlethread - global 1. buffer/global/ds/flat_store 1. buffer/global/ds/flat_store 4869 - wavefront - local 4870 - generic 4871 store atomic release - workgroup - global 1. s_waitcnt lgkmcnt(0) 1. s_waitcnt lgkmcnt(0) & 4872 vmcnt(0) & vscnt(0) 4873 4874 - If CU wavefront execution mode, omit vmcnt and 4875 vscnt. 4876 - If OpenCL, omit. - If OpenCL, omit 4877 lgkmcnt(0). 4878 - Must happen after 4879 any preceding 4880 local/generic 4881 load/store/load 4882 atomic/store 4883 atomic/atomicrmw. 4884 - Could be split into 4885 separate s_waitcnt 4886 vmcnt(0), s_waitcnt 4887 vscnt(0) and s_waitcnt 4888 lgkmcnt(0) to allow 4889 them to be 4890 independently moved 4891 according to the 4892 following rules. 4893 - s_waitcnt vmcnt(0) 4894 must happen after 4895 any preceding 4896 global/generic load/load 4897 atomic/ 4898 atomicrmw-with-return-value. 4899 - s_waitcnt vscnt(0) 4900 must happen after 4901 any preceding 4902 global/generic 4903 store/store 4904 atomic/ 4905 atomicrmw-no-return-value. 4906 - s_waitcnt lgkmcnt(0) 4907 must happen after 4908 any preceding 4909 local/generic 4910 load/store/load 4911 atomic/store 4912 atomic/atomicrmw. 4913 - Must happen before - Must happen before 4914 the following the following 4915 store. store. 4916 - Ensures that all - Ensures that all 4917 memory operations memory operations 4918 to local have have 4919 completed before completed before 4920 performing the performing the 4921 store that is being store that is being 4922 released. released. 4923 4924 2. buffer/global/flat_store 2. buffer/global_store 4925 store atomic release - workgroup - local 1. waitcnt vmcnt(0) & vscnt(0) 4926 4927 - If CU wavefront execution mode, omit. 4928 - If OpenCL, omit. 4929 - Could be split into 4930 separate s_waitcnt 4931 vmcnt(0) and s_waitcnt 4932 vscnt(0) to allow 4933 them to be 4934 independently moved 4935 according to the 4936 following rules. 4937 - s_waitcnt vmcnt(0) 4938 must happen after 4939 any preceding 4940 global/generic load/load 4941 atomic/ 4942 atomicrmw-with-return-value. 4943 - s_waitcnt vscnt(0) 4944 must happen after 4945 any preceding 4946 global/generic 4947 store/store atomic/ 4948 atomicrmw-no-return-value. 4949 - Must happen before 4950 the following 4951 store. 4952 - Ensures that all 4953 global memory 4954 operations have 4955 completed before 4956 performing the 4957 store that is being 4958 released. 4959 4960 1. ds_store 2. ds_store 4961 store atomic release - workgroup - generic 1. s_waitcnt lgkmcnt(0) 1. s_waitcnt lgkmcnt(0) & 4962 vmcnt(0) & vscnt(0) 4963 4964 - If CU wavefront execution mode, omit vmcnt and 4965 vscnt. 4966 - If OpenCL, omit. - If OpenCL, omit 4967 lgkmcnt(0). 4968 - Must happen after 4969 any preceding 4970 local/generic 4971 load/store/load 4972 atomic/store 4973 atomic/atomicrmw. 4974 - Could be split into 4975 separate s_waitcnt 4976 vmcnt(0), s_waitcnt 4977 vscnt(0) and s_waitcnt 4978 lgkmcnt(0) to allow 4979 them to be 4980 independently moved 4981 according to the 4982 following rules. 4983 - s_waitcnt vmcnt(0) 4984 must happen after 4985 any preceding 4986 global/generic load/load 4987 atomic/ 4988 atomicrmw-with-return-value. 4989 - s_waitcnt vscnt(0) 4990 must happen after 4991 any preceding 4992 global/generic 4993 store/store 4994 atomic/ 4995 atomicrmw-no-return-value. 4996 - s_waitcnt lgkmcnt(0) 4997 must happen after 4998 any preceding 4999 local/generic load/store/load 5000 atomic/store atomic/atomicrmw. 5001 - Must happen before - Must happen before 5002 the following the following 5003 store. store. 5004 - Ensures that all - Ensures that all 5005 memory operations memory operations 5006 to local have have 5007 completed before completed before 5008 performing the performing the 5009 store that is being store that is being 5010 released. released. 5011 5012 2. flat_store 2. flat_store 5013 store atomic release - agent - global 1. s_waitcnt lgkmcnt(0) & 1. s_waitcnt lgkmcnt(0) & 5014 - system - generic vmcnt(0) vmcnt(0) & vscnt(0) 5015 5016 - If OpenCL, omit - If OpenCL, omit 5017 lgkmcnt(0). lgkmcnt(0). 5018 - Could be split into - Could be split into 5019 separate s_waitcnt separate s_waitcnt 5020 vmcnt(0) and vmcnt(0), s_waitcnt vscnt(0) 5021 s_waitcnt and s_waitcnt 5022 lgkmcnt(0) to allow lgkmcnt(0) to allow 5023 them to be them to be 5024 independently moved independently moved 5025 according to the according to the 5026 following rules. following rules. 5027 - s_waitcnt vmcnt(0) - s_waitcnt vmcnt(0) 5028 must happen after must happen after 5029 any preceding any preceding 5030 global/generic global/generic 5031 load/store/load load/load 5032 atomic/store atomic/ 5033 atomic/atomicrmw. atomicrmw-with-return-value. 5034 - s_waitcnt vscnt(0) 5035 must happen after 5036 any preceding 5037 global/generic 5038 store/store atomic/ 5039 atomicrmw-no-return-value. 5040 - s_waitcnt lgkmcnt(0) - s_waitcnt lgkmcnt(0) 5041 must happen after must happen after 5042 any preceding any preceding 5043 local/generic local/generic 5044 load/store/load load/store/load 5045 atomic/store atomic/store 5046 atomic/atomicrmw. atomic/atomicrmw. 5047 - Must happen before - Must happen before 5048 the following the following 5049 store. store. 5050 - Ensures that all - Ensures that all 5051 memory operations memory operations 5052 to memory have to memory have 5053 completed before completed before 5054 performing the performing the 5055 store that is being store that is being 5056 released. released. 5057 5058 2. buffer/global/ds/flat_store 2. buffer/global/ds/flat_store 5059 atomicrmw release - singlethread - global 1. buffer/global/ds/flat_atomic 1. buffer/global/ds/flat_atomic 5060 - wavefront - local 5061 - generic 5062 atomicrmw release - workgroup - global 1. s_waitcnt lgkmcnt(0) 1. s_waitcnt lgkmcnt(0) & 5063 vmcnt(0) & vscnt(0) 5064 5065 - If CU wavefront execution mode, omit vmcnt and 5066 vscnt. 5067 - If OpenCL, omit. 5068 5069 - Must happen after 5070 any preceding 5071 local/generic 5072 load/store/load 5073 atomic/store 5074 atomic/atomicrmw. 5075 - Could be split into 5076 separate s_waitcnt 5077 vmcnt(0), s_waitcnt 5078 vscnt(0) and s_waitcnt 5079 lgkmcnt(0) to allow 5080 them to be 5081 independently moved 5082 according to the 5083 following rules. 5084 - s_waitcnt vmcnt(0) 5085 must happen after 5086 any preceding 5087 global/generic load/load 5088 atomic/ 5089 atomicrmw-with-return-value. 5090 - s_waitcnt vscnt(0) 5091 must happen after 5092 any preceding 5093 global/generic 5094 store/store 5095 atomic/ 5096 atomicrmw-no-return-value. 5097 - s_waitcnt lgkmcnt(0) 5098 must happen after 5099 any preceding 5100 local/generic 5101 load/store/load 5102 atomic/store 5103 atomic/atomicrmw. 5104 - Must happen before - Must happen before 5105 the following the following 5106 atomicrmw. atomicrmw. 5107 - Ensures that all - Ensures that all 5108 memory operations memory operations 5109 to local have have 5110 completed before completed before 5111 performing the performing the 5112 atomicrmw that is atomicrmw that is 5113 being released. being released. 5114 5115 2. buffer/global/flat_atomic 2. buffer/global_atomic 5116 atomicrmw release - workgroup - local 1. waitcnt vmcnt(0) & vscnt(0) 5117 5118 - If CU wavefront execution mode, omit. 5119 - If OpenCL, omit. 5120 - Could be split into 5121 separate s_waitcnt 5122 vmcnt(0) and s_waitcnt 5123 vscnt(0) to allow 5124 them to be 5125 independently moved 5126 according to the 5127 following rules. 5128 - s_waitcnt vmcnt(0) 5129 must happen after 5130 any preceding 5131 global/generic load/load 5132 atomic/ 5133 atomicrmw-with-return-value. 5134 - s_waitcnt vscnt(0) 5135 must happen after 5136 any preceding 5137 global/generic 5138 store/store atomic/ 5139 atomicrmw-no-return-value. 5140 - Must happen before 5141 the following 5142 store. 5143 - Ensures that all 5144 global memory 5145 operations have 5146 completed before 5147 performing the 5148 store that is being 5149 released. 5150 5151 1. ds_atomic 2. ds_atomic 5152 atomicrmw release - workgroup - generic 1. s_waitcnt lgkmcnt(0) 1. s_waitcnt lgkmcnt(0) & 5153 vmcnt(0) & vscnt(0) 5154 5155 - If CU wavefront execution mode, omit vmcnt and 5156 vscnt. 5157 - If OpenCL, omit. - If OpenCL, omit 5158 waitcnt lgkmcnt(0). 5159 - Must happen after 5160 any preceding 5161 local/generic 5162 load/store/load 5163 atomic/store 5164 atomic/atomicrmw. 5165 - Could be split into 5166 separate s_waitcnt 5167 vmcnt(0), s_waitcnt 5168 vscnt(0) and s_waitcnt 5169 lgkmcnt(0) to allow 5170 them to be 5171 independently moved 5172 according to the 5173 following rules. 5174 - s_waitcnt vmcnt(0) 5175 must happen after 5176 any preceding 5177 global/generic load/load 5178 atomic/ 5179 atomicrmw-with-return-value. 5180 - s_waitcnt vscnt(0) 5181 must happen after 5182 any preceding 5183 global/generic 5184 store/store 5185 atomic/ 5186 atomicrmw-no-return-value. 5187 - s_waitcnt lgkmcnt(0) 5188 must happen after 5189 any preceding 5190 local/generic load/store/load 5191 atomic/store atomic/atomicrmw. 5192 - Must happen before - Must happen before 5193 the following the following 5194 atomicrmw. atomicrmw. 5195 - Ensures that all - Ensures that all 5196 memory operations memory operations 5197 to local have have 5198 completed before completed before 5199 performing the performing the 5200 atomicrmw that is atomicrmw that is 5201 being released. being released. 5202 5203 2. flat_atomic 2. flat_atomic 5204 atomicrmw release - agent - global 1. s_waitcnt lgkmcnt(0) & 1. s_waitcnt lkkmcnt(0) & 5205 - system - generic vmcnt(0) vmcnt(0) & vscnt(0) 5206 5207 - If OpenCL, omit - If OpenCL, omit 5208 lgkmcnt(0). lgkmcnt(0). 5209 - Could be split into - Could be split into 5210 separate s_waitcnt separate s_waitcnt 5211 vmcnt(0) and vmcnt(0), s_waitcnt 5212 s_waitcnt vscnt(0) and s_waitcnt 5213 lgkmcnt(0) to allow lgkmcnt(0) to allow 5214 them to be them to be 5215 independently moved independently moved 5216 according to the according to the 5217 following rules. following rules. 5218 - s_waitcnt vmcnt(0) - s_waitcnt vmcnt(0) 5219 must happen after must happen after 5220 any preceding any preceding 5221 global/generic global/generic 5222 load/store/load load/load atomic/ 5223 atomic/store atomicrmw-with-return-value. 5224 atomic/atomicrmw. 5225 - s_waitcnt vscnt(0) 5226 must happen after 5227 any preceding 5228 global/generic 5229 store/store atomic/ 5230 atomicrmw-no-return-value. 5231 - s_waitcnt lgkmcnt(0) - s_waitcnt lgkmcnt(0) 5232 must happen after must happen after 5233 any preceding any preceding 5234 local/generic local/generic 5235 load/store/load load/store/load 5236 atomic/store atomic/store 5237 atomic/atomicrmw. atomic/atomicrmw. 5238 - Must happen before - Must happen before 5239 the following the following 5240 atomicrmw. atomicrmw. 5241 - Ensures that all - Ensures that all 5242 memory operations memory operations 5243 to global and local to global and local 5244 have completed have completed 5245 before performing before performing 5246 the atomicrmw that the atomicrmw that 5247 is being released. is being released. 5248 5249 2. buffer/global/ds/flat_atomic 2. buffer/global/ds/flat_atomic 5250 fence release - singlethread *none* *none* *none* 5251 - wavefront 5252 fence release - workgroup *none* 1. s_waitcnt lgkmcnt(0) 1. s_waitcnt lgkmcnt(0) & 5253 vmcnt(0) & vscnt(0) 5254 5255 - If CU wavefront execution mode, omit vmcnt and 5256 vscnt. 5257 - If OpenCL and - If OpenCL and 5258 address space is address space is 5259 not generic, omit. not generic, omit 5260 lgkmcnt(0). 5261 - If OpenCL and 5262 address space is 5263 local, omit 5264 vmcnt(0) and vscnt(0). 5265 - However, since LLVM - However, since LLVM 5266 currently has no currently has no 5267 address space on address space on 5268 the fence need to the fence need to 5269 conservatively conservatively 5270 always generate. If always generate. If 5271 fence had an fence had an 5272 address space then address space then 5273 set to address set to address 5274 space of OpenCL space of OpenCL 5275 fence flag, or to fence flag, or to 5276 generic if both generic if both 5277 local and global local and global 5278 flags are flags are 5279 specified. specified. 5280 - Must happen after 5281 any preceding 5282 local/generic 5283 load/load 5284 atomic/store/store 5285 atomic/atomicrmw. 5286 - Could be split into 5287 separate s_waitcnt 5288 vmcnt(0), s_waitcnt 5289 vscnt(0) and s_waitcnt 5290 lgkmcnt(0) to allow 5291 them to be 5292 independently moved 5293 according to the 5294 following rules. 5295 - s_waitcnt vmcnt(0) 5296 must happen after 5297 any preceding 5298 global/generic 5299 load/load 5300 atomic/ 5301 atomicrmw-with-return-value. 5302 - s_waitcnt vscnt(0) 5303 must happen after 5304 any preceding 5305 global/generic 5306 store/store atomic/ 5307 atomicrmw-no-return-value. 5308 - s_waitcnt lgkmcnt(0) 5309 must happen after 5310 any preceding 5311 local/generic 5312 load/store/load 5313 atomic/store atomic/ 5314 atomicrmw. 5315 - Must happen before - Must happen before 5316 any following store any following store 5317 atomic/atomicrmw atomic/atomicrmw 5318 with an equal or with an equal or 5319 wider sync scope wider sync scope 5320 and memory ordering and memory ordering 5321 stronger than stronger than 5322 unordered (this is unordered (this is 5323 termed the termed the 5324 fence-paired-atomic). fence-paired-atomic). 5325 - Ensures that all - Ensures that all 5326 memory operations memory operations 5327 to local have have 5328 completed before completed before 5329 performing the performing the 5330 following following 5331 fence-paired-atomic. fence-paired-atomic. 5332 5333 fence release - agent *none* 1. s_waitcnt lgkmcnt(0) & 1. s_waitcnt lgkmcnt(0) & 5334 - system vmcnt(0) vmcnt(0) & vscnt(0) 5335 5336 - If OpenCL and - If OpenCL and 5337 address space is address space is 5338 not generic, omit not generic, omit 5339 lgkmcnt(0). lgkmcnt(0). 5340 - If OpenCL and - If OpenCL and 5341 address space is address space is 5342 local, omit local, omit 5343 vmcnt(0). vmcnt(0) and vscnt(0). 5344 - However, since LLVM - However, since LLVM 5345 currently has no currently has no 5346 address space on address space on 5347 the fence need to the fence need to 5348 conservatively conservatively 5349 always generate. If always generate. If 5350 fence had an fence had an 5351 address space then address space then 5352 set to address set to address 5353 space of OpenCL space of OpenCL 5354 fence flag, or to fence flag, or to 5355 generic if both generic if both 5356 local and global local and global 5357 flags are flags are 5358 specified. specified. 5359 - Could be split into - Could be split into 5360 separate s_waitcnt separate s_waitcnt 5361 vmcnt(0) and vmcnt(0), s_waitcnt 5362 s_waitcnt vscnt(0) and s_waitcnt 5363 lgkmcnt(0) to allow lgkmcnt(0) to allow 5364 them to be them to be 5365 independently moved independently moved 5366 according to the according to the 5367 following rules. following rules. 5368 - s_waitcnt vmcnt(0) - s_waitcnt vmcnt(0) 5369 must happen after must happen after 5370 any preceding any preceding 5371 global/generic global/generic 5372 load/store/load load/load atomic/ 5373 atomic/store atomicrmw-with-return-value. 5374 atomic/atomicrmw. 5375 - s_waitcnt vscnt(0) 5376 must happen after 5377 any preceding 5378 global/generic 5379 store/store atomic/ 5380 atomicrmw-no-return-value. 5381 - s_waitcnt lgkmcnt(0) - s_waitcnt lgkmcnt(0) 5382 must happen after must happen after 5383 any preceding any preceding 5384 local/generic local/generic 5385 load/store/load load/store/load 5386 atomic/store atomic/store 5387 atomic/atomicrmw. atomic/atomicrmw. 5388 - Must happen before - Must happen before 5389 any following store any following store 5390 atomic/atomicrmw atomic/atomicrmw 5391 with an equal or with an equal or 5392 wider sync scope wider sync scope 5393 and memory ordering and memory ordering 5394 stronger than stronger than 5395 unordered (this is unordered (this is 5396 termed the termed the 5397 fence-paired-atomic). fence-paired-atomic). 5398 - Ensures that all - Ensures that all 5399 memory operations memory operations 5400 have have 5401 completed before completed before 5402 performing the performing the 5403 following following 5404 fence-paired-atomic. fence-paired-atomic. 5405 5406 **Acquire-Release Atomic** 5407 ---------------------------------------------------------------------------------------------------------------------- 5408 atomicrmw acq_rel - singlethread - global 1. buffer/global/ds/flat_atomic 1. buffer/global/ds/flat_atomic 5409 - wavefront - local 5410 - generic 5411 atomicrmw acq_rel - workgroup - global 1. s_waitcnt lgkmcnt(0) 1. s_waitcnt lgkmcnt(0) & 5412 vmcnt(0) & vscnt(0) 5413 5414 - If CU wavefront execution mode, omit vmcnt and 5415 vscnt. 5416 - If OpenCL, omit. - If OpenCL, omit 5417 s_waitcnt lgkmcnt(0). 5418 - Must happen after - Must happen after 5419 any preceding any preceding 5420 local/generic local/generic 5421 load/store/load load/store/load 5422 atomic/store atomic/store 5423 atomic/atomicrmw. atomic/atomicrmw. 5424 - Could be split into 5425 separate s_waitcnt 5426 vmcnt(0), s_waitcnt 5427 vscnt(0) and s_waitcnt 5428 lgkmcnt(0) to allow 5429 them to be 5430 independently moved 5431 according to the 5432 following rules. 5433 - s_waitcnt vmcnt(0) 5434 must happen after 5435 any preceding 5436 global/generic load/load 5437 atomic/ 5438 atomicrmw-with-return-value. 5439 - s_waitcnt vscnt(0) 5440 must happen after 5441 any preceding 5442 global/generic 5443 store/store 5444 atomic/ 5445 atomicrmw-no-return-value. 5446 - s_waitcnt lgkmcnt(0) 5447 must happen after 5448 any preceding 5449 local/generic load/store/load 5450 atomic/store atomic/atomicrmw. 5451 - Must happen before - Must happen before 5452 the following the following 5453 atomicrmw. atomicrmw. 5454 - Ensures that all - Ensures that all 5455 memory operations memory operations 5456 to local have have 5457 completed before completed before 5458 performing the performing the 5459 atomicrmw that is atomicrmw that is 5460 being released. being released. 5461 5462 2. buffer/global/flat_atomic 2. buffer/global_atomic 5463 3. s_waitcnt vm/vscnt(0) 5464 5465 - If CU wavefront execution mode, omit vm/vscnt. 5466 - Use vmcnt if atomic with 5467 return and vscnt if atomic 5468 with no-return. 5469 waitcnt lgkmcnt(0). 5470 - Must happen before 5471 the following 5472 buffer_gl0_inv. 5473 - Ensures any 5474 following global 5475 data read is no 5476 older than the 5477 atomicrmw value 5478 being acquired. 5479 5480 4. buffer_gl0_inv 5481 5482 - If CU wavefront execution mode, omit. 5483 - Ensures that 5484 following 5485 loads will not see 5486 stale data. 5487 5488 atomicrmw acq_rel - workgroup - local 1. waitcnt vmcnt(0) & vscnt(0) 5489 5490 - If CU wavefront execution mode, omit. 5491 - If OpenCL, omit. 5492 - Could be split into 5493 separate s_waitcnt 5494 vmcnt(0) and s_waitcnt 5495 vscnt(0) to allow 5496 them to be 5497 independently moved 5498 according to the 5499 following rules. 5500 - s_waitcnt vmcnt(0) 5501 must happen after 5502 any preceding 5503 global/generic load/load 5504 atomic/ 5505 atomicrmw-with-return-value. 5506 - s_waitcnt vscnt(0) 5507 must happen after 5508 any preceding 5509 global/generic 5510 store/store atomic/ 5511 atomicrmw-no-return-value. 5512 - Must happen before 5513 the following 5514 store. 5515 - Ensures that all 5516 global memory 5517 operations have 5518 completed before 5519 performing the 5520 store that is being 5521 released. 5522 5523 1. ds_atomic 2. ds_atomic 5524 2. s_waitcnt lgkmcnt(0) 3. s_waitcnt lgkmcnt(0) 5525 5526 - If OpenCL, omit. - If OpenCL, omit. 5527 - Must happen before - Must happen before 5528 any following the following 5529 global/generic buffer_gl0_inv. 5530 load/load 5531 atomic/store/store 5532 atomic/atomicrmw. 5533 - Ensures any - Ensures any 5534 following global following global 5535 data read is no data read is no 5536 older than the load older than the load 5537 atomic value being atomic value being 5538 acquired. acquired. 5539 5540 4. buffer_gl0_inv 5541 5542 - If CU wavefront execution mode, omit. 5543 - If OpenCL omit. 5544 - Ensures that 5545 following 5546 loads will not see 5547 stale data. 5548 5549 atomicrmw acq_rel - workgroup - generic 1. s_waitcnt lgkmcnt(0) 1. s_waitcnt lgkmcnt(0) & 5550 vmcnt(0) & vscnt(0) 5551 5552 - If CU wavefront execution mode, omit vmcnt and 5553 vscnt. 5554 - If OpenCL, omit. - If OpenCL, omit 5555 waitcnt lgkmcnt(0). 5556 - Must happen after 5557 any preceding 5558 local/generic 5559 load/store/load 5560 atomic/store 5561 atomic/atomicrmw. 5562 - Could be split into 5563 separate s_waitcnt 5564 vmcnt(0), s_waitcnt 5565 vscnt(0) and s_waitcnt 5566 lgkmcnt(0) to allow 5567 them to be 5568 independently moved 5569 according to the 5570 following rules. 5571 - s_waitcnt vmcnt(0) 5572 must happen after 5573 any preceding 5574 global/generic load/load 5575 atomic/ 5576 atomicrmw-with-return-value. 5577 - s_waitcnt vscnt(0) 5578 must happen after 5579 any preceding 5580 global/generic 5581 store/store 5582 atomic/ 5583 atomicrmw-no-return-value. 5584 - s_waitcnt lgkmcnt(0) 5585 must happen after 5586 any preceding 5587 local/generic load/store/load 5588 atomic/store atomic/atomicrmw. 5589 - Must happen before - Must happen before 5590 the following the following 5591 atomicrmw. atomicrmw. 5592 - Ensures that all - Ensures that all 5593 memory operations memory operations 5594 to local have have 5595 completed before completed before 5596 performing the performing the 5597 atomicrmw that is atomicrmw that is 5598 being released. being released. 5599 5600 2. flat_atomic 2. flat_atomic 5601 3. s_waitcnt lgkmcnt(0) 3. s_waitcnt lgkmcnt(0) & 5602 vm/vscnt(0) 5603 5604 - If CU wavefront execution mode, omit vm/vscnt. 5605 - If OpenCL, omit. - If OpenCL, omit 5606 waitcnt lgkmcnt(0). 5607 - Must happen before - Must happen before 5608 any following the following 5609 global/generic buffer_gl0_inv. 5610 load/load 5611 atomic/store/store 5612 atomic/atomicrmw. 5613 - Ensures any - Ensures any 5614 following global following global 5615 data read is no data read is no 5616 older than the load older than the load 5617 atomic value being atomic value being 5618 acquired. acquired. 5619 5620 3. buffer_gl0_inv 5621 5622 - If CU wavefront execution mode, omit. 5623 - Ensures that 5624 following 5625 loads will not see 5626 stale data. 5627 5628 atomicrmw acq_rel - agent - global 1. s_waitcnt lgkmcnt(0) & 1. s_waitcnt lgkmcnt(0) & 5629 - system vmcnt(0) vmcnt(0) & vscnt(0) 5630 5631 - If OpenCL, omit - If OpenCL, omit 5632 lgkmcnt(0). lgkmcnt(0). 5633 - Could be split into - Could be split into 5634 separate s_waitcnt separate s_waitcnt 5635 vmcnt(0) and vmcnt(0), s_waitcnt 5636 s_waitcnt vscnt(0) and s_waitcnt 5637 lgkmcnt(0) to allow lgkmcnt(0) to allow 5638 them to be them to be 5639 independently moved independently moved 5640 according to the according to the 5641 following rules. following rules. 5642 - s_waitcnt vmcnt(0) - s_waitcnt vmcnt(0) 5643 must happen after must happen after 5644 any preceding any preceding 5645 global/generic global/generic 5646 load/store/load load/load atomic/ 5647 atomic/store atomicrmw-with-return-value. 5648 atomic/atomicrmw. 5649 - s_waitcnt vscnt(0) 5650 must happen after 5651 any preceding 5652 global/generic 5653 store/store atomic/ 5654 atomicrmw-no-return-value. 5655 - s_waitcnt lgkmcnt(0) - s_waitcnt lgkmcnt(0) 5656 must happen after must happen after 5657 any preceding any preceding 5658 local/generic local/generic 5659 load/store/load load/store/load 5660 atomic/store atomic/store 5661 atomic/atomicrmw. atomic/atomicrmw. 5662 - Must happen before - Must happen before 5663 the following the following 5664 atomicrmw. atomicrmw. 5665 - Ensures that all - Ensures that all 5666 memory operations memory operations 5667 to global have to global have 5668 completed before completed before 5669 performing the performing the 5670 atomicrmw that is atomicrmw that is 5671 being released. being released. 5672 5673 2. buffer/global/flat_atomic 2. buffer/global_atomic 5674 3. s_waitcnt vmcnt(0) 3. s_waitcnt vm/vscnt(0) 5675 5676 - Use vmcnt if atomic with 5677 return and vscnt if atomic 5678 with no-return. 5679 waitcnt lgkmcnt(0). 5680 - Must happen before - Must happen before 5681 following following 5682 buffer_wbinvl1_vol. buffer_gl*_inv. 5683 - Ensures the - Ensures the 5684 atomicrmw has atomicrmw has 5685 completed before completed before 5686 invalidating the invalidating the 5687 cache. caches. 5688 5689 4. buffer_wbinvl1_vol 4. buffer_gl0_inv; 5690 buffer_gl1_inv 5691 5692 - Must happen before - Must happen before 5693 any following any following 5694 global/generic global/generic 5695 load/load load/load 5696 atomic/atomicrmw. atomic/atomicrmw. 5697 - Ensures that - Ensures that 5698 following loads following loads 5699 will not see stale will not see stale 5700 global data. global data. 5701 5702 atomicrmw acq_rel - agent - generic 1. s_waitcnt lgkmcnt(0) & 1. s_waitcnt lgkmcnt(0) & 5703 - system vmcnt(0) vmcnt(0) & vscnt(0) 5704 5705 - If OpenCL, omit - If OpenCL, omit 5706 lgkmcnt(0). lgkmcnt(0). 5707 - Could be split into - Could be split into 5708 separate s_waitcnt separate s_waitcnt 5709 vmcnt(0) and vmcnt(0), s_waitcnt 5710 s_waitcnt vscnt(0) and s_waitcnt 5711 lgkmcnt(0) to allow lgkmcnt(0) to allow 5712 them to be them to be 5713 independently moved independently moved 5714 according to the according to the 5715 following rules. following rules. 5716 - s_waitcnt vmcnt(0) - s_waitcnt vmcnt(0) 5717 must happen after must happen after 5718 any preceding any preceding 5719 global/generic global/generic 5720 load/store/load load/load atomic 5721 atomic/store atomicrmw-with-return-value. 5722 atomic/atomicrmw. 5723 - s_waitcnt vscnt(0) 5724 must happen after 5725 any preceding 5726 global/generic 5727 store/store atomic/ 5728 atomicrmw-no-return-value. 5729 - s_waitcnt lgkmcnt(0) - s_waitcnt lgkmcnt(0) 5730 must happen after must happen after 5731 any preceding any preceding 5732 local/generic local/generic 5733 load/store/load load/store/load 5734 atomic/store atomic/store 5735 atomic/atomicrmw. atomic/atomicrmw. 5736 - Must happen before - Must happen before 5737 the following the following 5738 atomicrmw. atomicrmw. 5739 - Ensures that all - Ensures that all 5740 memory operations memory operations 5741 to global have have 5742 completed before completed before 5743 performing the performing the 5744 atomicrmw that is atomicrmw that is 5745 being released. being released. 5746 5747 2. flat_atomic 2. flat_atomic 5748 3. s_waitcnt vmcnt(0) & 3. s_waitcnt vm/vscnt(0) & 5749 lgkmcnt(0) lgkmcnt(0) 5750 5751 - If OpenCL, omit - If OpenCL, omit 5752 lgkmcnt(0). lgkmcnt(0). 5753 - Use vmcnt if atomic with 5754 return and vscnt if atomic 5755 with no-return. 5756 - Must happen before - Must happen before 5757 following following 5758 buffer_wbinvl1_vol. buffer_gl*_inv. 5759 - Ensures the - Ensures the 5760 atomicrmw has atomicrmw has 5761 completed before completed before 5762 invalidating the invalidating the 5763 cache. caches. 5764 5765 4. buffer_wbinvl1_vol 4. buffer_gl0_inv; 5766 buffer_gl1_inv 5767 5768 - Must happen before - Must happen before 5769 any following any following 5770 global/generic global/generic 5771 load/load load/load 5772 atomic/atomicrmw. atomic/atomicrmw. 5773 - Ensures that - Ensures that 5774 following loads following loads 5775 will not see stale will not see stale 5776 global data. global data. 5777 5778 fence acq_rel - singlethread *none* *none* *none* 5779 - wavefront 5780 fence acq_rel - workgroup *none* 1. s_waitcnt lgkmcnt(0) 1. s_waitcnt lgkmcnt(0) & 5781 vmcnt(0) & vscnt(0) 5782 5783 - If CU wavefront execution mode, omit vmcnt and 5784 vscnt. 5785 - If OpenCL and - If OpenCL and 5786 address space is address space is 5787 not generic, omit. not generic, omit 5788 lgkmcnt(0). 5789 - If OpenCL and 5790 address space is 5791 local, omit 5792 vmcnt(0) and vscnt(0). 5793 - However, - However, 5794 since LLVM since LLVM 5795 currently has no currently has no 5796 address space on address space on 5797 the fence need to the fence need to 5798 conservatively conservatively 5799 always generate always generate 5800 (see comment for (see comment for 5801 previous fence). previous fence). 5802 - Must happen after 5803 any preceding 5804 local/generic 5805 load/load 5806 atomic/store/store 5807 atomic/atomicrmw. 5808 - Could be split into 5809 separate s_waitcnt 5810 vmcnt(0), s_waitcnt 5811 vscnt(0) and s_waitcnt 5812 lgkmcnt(0) to allow 5813 them to be 5814 independently moved 5815 according to the 5816 following rules. 5817 - s_waitcnt vmcnt(0) 5818 must happen after 5819 any preceding 5820 global/generic 5821 load/load 5822 atomic/ 5823 atomicrmw-with-return-value. 5824 - s_waitcnt vscnt(0) 5825 must happen after 5826 any preceding 5827 global/generic 5828 store/store atomic/ 5829 atomicrmw-no-return-value. 5830 - s_waitcnt lgkmcnt(0) 5831 must happen after 5832 any preceding 5833 local/generic 5834 load/store/load 5835 atomic/store atomic/ 5836 atomicrmw. 5837 - Must happen before - Must happen before 5838 any following any following 5839 global/generic global/generic 5840 load/load load/load 5841 atomic/store/store atomic/store/store 5842 atomic/atomicrmw. atomic/atomicrmw. 5843 - Ensures that all - Ensures that all 5844 memory operations memory operations 5845 to local have have 5846 completed before completed before 5847 performing any performing any 5848 following global following global 5849 memory operations. memory operations. 5850 - Ensures that the - Ensures that the 5851 preceding preceding 5852 local/generic load local/generic load 5853 atomic/atomicrmw atomic/atomicrmw 5854 with an equal or with an equal or 5855 wider sync scope wider sync scope 5856 and memory ordering and memory ordering 5857 stronger than stronger than 5858 unordered (this is unordered (this is 5859 termed the termed the 5860 acquire-fence-paired-atomic acquire-fence-paired-atomic 5861 ) has completed ) has completed 5862 before following before following 5863 global memory global memory 5864 operations. This operations. This 5865 satisfies the satisfies the 5866 requirements of requirements of 5867 acquire. acquire. 5868 - Ensures that all - Ensures that all 5869 previous memory previous memory 5870 operations have operations have 5871 completed before a completed before a 5872 following following 5873 local/generic store local/generic store 5874 atomic/atomicrmw atomic/atomicrmw 5875 with an equal or with an equal or 5876 wider sync scope wider sync scope 5877 and memory ordering and memory ordering 5878 stronger than stronger than 5879 unordered (this is unordered (this is 5880 termed the termed the 5881 release-fence-paired-atomic release-fence-paired-atomic 5882 ). This satisfies the ). This satisfies the 5883 requirements of requirements of 5884 release. release. 5885 - Must happen before 5886 the following 5887 buffer_gl0_inv. 5888 - Ensures that the 5889 acquire-fence-paired 5890 atomic has completed 5891 before invalidating 5892 the 5893 cache. Therefore 5894 any following 5895 locations read must 5896 be no older than 5897 the value read by 5898 the 5899 acquire-fence-paired-atomic. 5900 5901 3. buffer_gl0_inv 5902 5903 - If CU wavefront execution mode, omit. 5904 - Ensures that 5905 following 5906 loads will not see 5907 stale data. 5908 5909 fence acq_rel - agent *none* 1. s_waitcnt lgkmcnt(0) & 1. s_waitcnt lgkmcnt(0) & 5910 - system vmcnt(0) vmcnt(0) & vscnt(0) 5911 5912 - If OpenCL and - If OpenCL and 5913 address space is address space is 5914 not generic, omit not generic, omit 5915 lgkmcnt(0). lgkmcnt(0). 5916 - If OpenCL and 5917 address space is 5918 local, omit 5919 vmcnt(0) and vscnt(0). 5920 - However, since LLVM - However, since LLVM 5921 currently has no currently has no 5922 address space on address space on 5923 the fence need to the fence need to 5924 conservatively conservatively 5925 always generate always generate 5926 (see comment for (see comment for 5927 previous fence). previous fence). 5928 - Could be split into - Could be split into 5929 separate s_waitcnt separate s_waitcnt 5930 vmcnt(0) and vmcnt(0), s_waitcnt 5931 s_waitcnt vscnt(0) and s_waitcnt 5932 lgkmcnt(0) to allow lgkmcnt(0) to allow 5933 them to be them to be 5934 independently moved independently moved 5935 according to the according to the 5936 following rules. following rules. 5937 - s_waitcnt vmcnt(0) - s_waitcnt vmcnt(0) 5938 must happen after must happen after 5939 any preceding any preceding 5940 global/generic global/generic 5941 load/store/load load/load 5942 atomic/store atomic/ 5943 atomic/atomicrmw. atomicrmw-with-return-value. 5944 - s_waitcnt vscnt(0) 5945 must happen after 5946 any preceding 5947 global/generic 5948 store/store atomic/ 5949 atomicrmw-no-return-value. 5950 - s_waitcnt lgkmcnt(0) - s_waitcnt lgkmcnt(0) 5951 must happen after must happen after 5952 any preceding any preceding 5953 local/generic local/generic 5954 load/store/load load/store/load 5955 atomic/store atomic/store 5956 atomic/atomicrmw. atomic/atomicrmw. 5957 - Must happen before - Must happen before 5958 the following the following 5959 buffer_wbinvl1_vol. buffer_gl*_inv. 5960 - Ensures that the - Ensures that the 5961 preceding preceding 5962 global/local/generic global/local/generic 5963 load load 5964 atomic/atomicrmw atomic/atomicrmw 5965 with an equal or with an equal or 5966 wider sync scope wider sync scope 5967 and memory ordering and memory ordering 5968 stronger than stronger than 5969 unordered (this is unordered (this is 5970 termed the termed the 5971 acquire-fence-paired-atomic acquire-fence-paired-atomic 5972 ) has completed ) has completed 5973 before invalidating before invalidating 5974 the cache. This the caches. This 5975 satisfies the satisfies the 5976 requirements of requirements of 5977 acquire. acquire. 5978 - Ensures that all - Ensures that all 5979 previous memory previous memory 5980 operations have operations have 5981 completed before a completed before a 5982 following following 5983 global/local/generic global/local/generic 5984 store store 5985 atomic/atomicrmw atomic/atomicrmw 5986 with an equal or with an equal or 5987 wider sync scope wider sync scope 5988 and memory ordering and memory ordering 5989 stronger than stronger than 5990 unordered (this is unordered (this is 5991 termed the termed the 5992 release-fence-paired-atomic release-fence-paired-atomic 5993 ). This satisfies the ). This satisfies the 5994 requirements of requirements of 5995 release. release. 5996 5997 2. buffer_wbinvl1_vol 2. buffer_gl0_inv; 5998 buffer_gl1_inv 5999 6000 - Must happen before - Must happen before 6001 any following any following 6002 global/generic global/generic 6003 load/load load/load 6004 atomic/store/store atomic/store/store 6005 atomic/atomicrmw. atomic/atomicrmw. 6006 - Ensures that - Ensures that 6007 following loads following loads 6008 will not see stale will not see stale 6009 global data. This global data. This 6010 satisfies the satisfies the 6011 requirements of requirements of 6012 acquire. acquire. 6013 6014 **Sequential Consistent Atomic** 6015 ---------------------------------------------------------------------------------------------------------------------- 6016 load atomic seq_cst - singlethread - global *Same as corresponding *Same as corresponding 6017 - wavefront - local load atomic acquire, load atomic acquire, 6018 - generic except must generated except must generated 6019 all instructions even all instructions even 6020 for OpenCL.* for OpenCL.* 6021 load atomic seq_cst - workgroup - global 1. s_waitcnt lgkmcnt(0) 1. s_waitcnt lgkmcnt(0) & 6022 - generic vmcnt(0) & vscnt(0) 6023 6024 - If CU wavefront execution mode, omit vmcnt and 6025 vscnt. 6026 - Could be split into 6027 separate s_waitcnt 6028 vmcnt(0), s_waitcnt 6029 vscnt(0) and s_waitcnt 6030 lgkmcnt(0) to allow 6031 them to be 6032 independently moved 6033 according to the 6034 following rules. 6035 - Must - waitcnt lgkmcnt(0) must 6036 happen after happen after 6037 preceding preceding 6038 global/generic load local load 6039 atomic/store atomic/store 6040 atomic/atomicrmw atomic/atomicrmw 6041 with memory with memory 6042 ordering of seq_cst ordering of seq_cst 6043 and with equal or and with equal or 6044 wider sync scope. wider sync scope. 6045 (Note that seq_cst (Note that seq_cst 6046 fences have their fences have their 6047 own s_waitcnt own s_waitcnt 6048 lgkmcnt(0) and so do lgkmcnt(0) and so do 6049 not need to be not need to be 6050 considered.) considered.) 6051 - waitcnt vmcnt(0) 6052 Must happen after 6053 preceding 6054 global/generic load 6055 atomic/ 6056 atomicrmw-with-return-value 6057 with memory 6058 ordering of seq_cst 6059 and with equal or 6060 wider sync scope. 6061 (Note that seq_cst 6062 fences have their 6063 own s_waitcnt 6064 vmcnt(0) and so do 6065 not need to be 6066 considered.) 6067 - waitcnt vscnt(0) 6068 Must happen after 6069 preceding 6070 global/generic store 6071 atomic/ 6072 atomicrmw-no-return-value 6073 with memory 6074 ordering of seq_cst 6075 and with equal or 6076 wider sync scope. 6077 (Note that seq_cst 6078 fences have their 6079 own s_waitcnt 6080 vscnt(0) and so do 6081 not need to be 6082 considered.) 6083 - Ensures any - Ensures any 6084 preceding preceding 6085 sequential sequential 6086 consistent local consistent global/local 6087 memory instructions memory instructions 6088 have completed have completed 6089 before executing before executing 6090 this sequentially this sequentially 6091 consistent consistent 6092 instruction. This instruction. This 6093 prevents reordering prevents reordering 6094 a seq_cst store a seq_cst store 6095 followed by a followed by a 6096 seq_cst load. (Note seq_cst load. (Note 6097 that seq_cst is that seq_cst is 6098 stronger than stronger than 6099 acquire/release as acquire/release as 6100 the reordering of the reordering of 6101 load acquire load acquire 6102 followed by a store followed by a store 6103 release is release is 6104 prevented by the prevented by the 6105 waitcnt of waitcnt of 6106 the release, but the release, but 6107 there is nothing there is nothing 6108 preventing a store preventing a store 6109 release followed by release followed by 6110 load acquire from load acquire from 6111 competing out of competing out of 6112 order.) order.) 6113 6114 2. *Following 2. *Following 6115 instructions same as instructions same as 6116 corresponding load corresponding load 6117 atomic acquire, atomic acquire, 6118 except must generated except must generated 6119 all instructions even all instructions even 6120 for OpenCL.* for OpenCL.* 6121 load atomic seq_cst - workgroup - local *Same as corresponding 6122 load atomic acquire, 6123 except must generated 6124 all instructions even 6125 for OpenCL.* 6126 6127 1. s_waitcnt vmcnt(0) & vscnt(0) 6128 6129 - If CU wavefront execution mode, omit. 6130 - Could be split into 6131 separate s_waitcnt 6132 vmcnt(0) and s_waitcnt 6133 vscnt(0) to allow 6134 them to be 6135 independently moved 6136 according to the 6137 following rules. 6138 - waitcnt vmcnt(0) 6139 Must happen after 6140 preceding 6141 global/generic load 6142 atomic/ 6143 atomicrmw-with-return-value 6144 with memory 6145 ordering of seq_cst 6146 and with equal or 6147 wider sync scope. 6148 (Note that seq_cst 6149 fences have their 6150 own s_waitcnt 6151 vmcnt(0) and so do 6152 not need to be 6153 considered.) 6154 - waitcnt vscnt(0) 6155 Must happen after 6156 preceding 6157 global/generic store 6158 atomic/ 6159 atomicrmw-no-return-value 6160 with memory 6161 ordering of seq_cst 6162 and with equal or 6163 wider sync scope. 6164 (Note that seq_cst 6165 fences have their 6166 own s_waitcnt 6167 vscnt(0) and so do 6168 not need to be 6169 considered.) 6170 - Ensures any 6171 preceding 6172 sequential 6173 consistent global 6174 memory instructions 6175 have completed 6176 before executing 6177 this sequentially 6178 consistent 6179 instruction. This 6180 prevents reordering 6181 a seq_cst store 6182 followed by a 6183 seq_cst load. (Note 6184 that seq_cst is 6185 stronger than 6186 acquire/release as 6187 the reordering of 6188 load acquire 6189 followed by a store 6190 release is 6191 prevented by the 6192 waitcnt of 6193 the release, but 6194 there is nothing 6195 preventing a store 6196 release followed by 6197 load acquire from 6198 competing out of 6199 order.) 6200 6201 2. *Following 6202 instructions same as 6203 corresponding load 6204 atomic acquire, 6205 except must generated 6206 all instructions even 6207 for OpenCL.* 6208 6209 load atomic seq_cst - agent - global 1. s_waitcnt lgkmcnt(0) & 1. s_waitcnt lgkmcnt(0) & 6210 - system - generic vmcnt(0) vmcnt(0) & vscnt(0) 6211 6212 - Could be split into - Could be split into 6213 separate s_waitcnt separate s_waitcnt 6214 vmcnt(0) vmcnt(0), s_waitcnt 6215 and s_waitcnt vscnt(0) and s_waitcnt 6216 lgkmcnt(0) to allow lgkmcnt(0) to allow 6217 them to be them to be 6218 independently moved independently moved 6219 according to the according to the 6220 following rules. following rules. 6221 - waitcnt lgkmcnt(0) - waitcnt lgkmcnt(0) 6222 must happen after must happen after 6223 preceding preceding 6224 global/generic load local load 6225 atomic/store atomic/store 6226 atomic/atomicrmw atomic/atomicrmw 6227 with memory with memory 6228 ordering of seq_cst ordering of seq_cst 6229 and with equal or and with equal or 6230 wider sync scope. wider sync scope. 6231 (Note that seq_cst (Note that seq_cst 6232 fences have their fences have their 6233 own s_waitcnt own s_waitcnt 6234 lgkmcnt(0) and so do lgkmcnt(0) and so do 6235 not need to be not need to be 6236 considered.) considered.) 6237 - waitcnt vmcnt(0) - waitcnt vmcnt(0) 6238 must happen after must happen after 6239 preceding preceding 6240 global/generic load global/generic load 6241 atomic/store atomic/ 6242 atomic/atomicrmw atomicrmw-with-return-value 6243 with memory with memory 6244 ordering of seq_cst ordering of seq_cst 6245 and with equal or and with equal or 6246 wider sync scope. wider sync scope. 6247 (Note that seq_cst (Note that seq_cst 6248 fences have their fences have their 6249 own s_waitcnt own s_waitcnt 6250 vmcnt(0) and so do vmcnt(0) and so do 6251 not need to be not need to be 6252 considered.) considered.) 6253 - waitcnt vscnt(0) 6254 Must happen after 6255 preceding 6256 global/generic store 6257 atomic/ 6258 atomicrmw-no-return-value 6259 with memory 6260 ordering of seq_cst 6261 and with equal or 6262 wider sync scope. 6263 (Note that seq_cst 6264 fences have their 6265 own s_waitcnt 6266 vscnt(0) and so do 6267 not need to be 6268 considered.) 6269 - Ensures any - Ensures any 6270 preceding preceding 6271 sequential sequential 6272 consistent global consistent global 6273 memory instructions memory instructions 6274 have completed have completed 6275 before executing before executing 6276 this sequentially this sequentially 6277 consistent consistent 6278 instruction. This instruction. This 6279 prevents reordering prevents reordering 6280 a seq_cst store a seq_cst store 6281 followed by a followed by a 6282 seq_cst load. (Note seq_cst load. (Note 6283 that seq_cst is that seq_cst is 6284 stronger than stronger than 6285 acquire/release as acquire/release as 6286 the reordering of the reordering of 6287 load acquire load acquire 6288 followed by a store followed by a store 6289 release is release is 6290 prevented by the prevented by the 6291 waitcnt of waitcnt of 6292 the release, but the release, but 6293 there is nothing there is nothing 6294 preventing a store preventing a store 6295 release followed by release followed by 6296 load acquire from load acquire from 6297 competing out of competing out of 6298 order.) order.) 6299 6300 2. *Following 2. *Following 6301 instructions same as instructions same as 6302 corresponding load corresponding load 6303 atomic acquire, atomic acquire, 6304 except must generated except must generated 6305 all instructions even all instructions even 6306 for OpenCL.* for OpenCL.* 6307 store atomic seq_cst - singlethread - global *Same as corresponding *Same as corresponding 6308 - wavefront - local store atomic release, store atomic release, 6309 - workgroup - generic except must generated except must generated 6310 all instructions even all instructions even 6311 for OpenCL.* for OpenCL.* 6312 store atomic seq_cst - agent - global *Same as corresponding *Same as corresponding 6313 - system - generic store atomic release, store atomic release, 6314 except must generated except must generated 6315 all instructions even all instructions even 6316 for OpenCL.* for OpenCL.* 6317 atomicrmw seq_cst - singlethread - global *Same as corresponding *Same as corresponding 6318 - wavefront - local atomicrmw acq_rel, atomicrmw acq_rel, 6319 - workgroup - generic except must generated except must generated 6320 all instructions even all instructions even 6321 for OpenCL.* for OpenCL.* 6322 atomicrmw seq_cst - agent - global *Same as corresponding *Same as corresponding 6323 - system - generic atomicrmw acq_rel, atomicrmw acq_rel, 6324 except must generated except must generated 6325 all instructions even all instructions even 6326 for OpenCL.* for OpenCL.* 6327 fence seq_cst - singlethread *none* *Same as corresponding *Same as corresponding 6328 - wavefront fence acq_rel, fence acq_rel, 6329 - workgroup except must generated except must generated 6330 - agent all instructions even all instructions even 6331 - system for OpenCL.* for OpenCL.* 6332 ============ ============ ============== ========== =============================== ================================== 6333 6334The memory order also adds the single thread optimization constrains defined in 6335table 6336:ref:`amdgpu-amdhsa-memory-model-single-thread-optimization-constraints-gfx6-gfx10-table`. 6337 6338 .. table:: AMDHSA Memory Model Single Thread Optimization Constraints GFX6-GFX10 6339 :name: amdgpu-amdhsa-memory-model-single-thread-optimization-constraints-gfx6-gfx10-table 6340 6341 ============ ============================================================== 6342 LLVM Memory Optimization Constraints 6343 Ordering 6344 ============ ============================================================== 6345 unordered *none* 6346 monotonic *none* 6347 acquire - If a load atomic/atomicrmw then no following load/load 6348 atomic/store/ store atomic/atomicrmw/fence instruction can 6349 be moved before the acquire. 6350 - If a fence then same as load atomic, plus no preceding 6351 associated fence-paired-atomic can be moved after the fence. 6352 release - If a store atomic/atomicrmw then no preceding load/load 6353 atomic/store/ store atomic/atomicrmw/fence instruction can 6354 be moved after the release. 6355 - If a fence then same as store atomic, plus no following 6356 associated fence-paired-atomic can be moved before the 6357 fence. 6358 acq_rel Same constraints as both acquire and release. 6359 seq_cst - If a load atomic then same constraints as acquire, plus no 6360 preceding sequentially consistent load atomic/store 6361 atomic/atomicrmw/fence instruction can be moved after the 6362 seq_cst. 6363 - If a store atomic then the same constraints as release, plus 6364 no following sequentially consistent load atomic/store 6365 atomic/atomicrmw/fence instruction can be moved before the 6366 seq_cst. 6367 - If an atomicrmw/fence then same constraints as acq_rel. 6368 ============ ============================================================== 6369 6370Trap Handler ABI 6371~~~~~~~~~~~~~~~~ 6372 6373For code objects generated by AMDGPU backend for HSA [HSA]_ compatible runtimes 6374(such as ROCm [AMD-ROCm]_), the runtime installs a trap handler that supports 6375the ``s_trap`` instruction with the following usage: 6376 6377 .. table:: AMDGPU Trap Handler for AMDHSA OS 6378 :name: amdgpu-trap-handler-for-amdhsa-os-table 6379 6380 =================== =============== =============== ======================= 6381 Usage Code Sequence Trap Handler Description 6382 Inputs 6383 =================== =============== =============== ======================= 6384 reserved ``s_trap 0x00`` Reserved by hardware. 6385 ``debugtrap(arg)`` ``s_trap 0x01`` ``SGPR0-1``: Reserved for HSA 6386 ``queue_ptr`` ``debugtrap`` 6387 ``VGPR0``: intrinsic (not 6388 ``arg`` implemented). 6389 ``llvm.trap`` ``s_trap 0x02`` ``SGPR0-1``: Causes dispatch to be 6390 ``queue_ptr`` terminated and its 6391 associated queue put 6392 into the error state. 6393 ``llvm.debugtrap`` ``s_trap 0x03`` - If debugger not 6394 installed then 6395 behaves as a 6396 no-operation. The 6397 trap handler is 6398 entered and 6399 immediately returns 6400 to continue 6401 execution of the 6402 wavefront. 6403 - If the debugger is 6404 installed, causes 6405 the debug trap to be 6406 reported by the 6407 debugger and the 6408 wavefront is put in 6409 the halt state until 6410 resumed by the 6411 debugger. 6412 reserved ``s_trap 0x04`` Reserved. 6413 reserved ``s_trap 0x05`` Reserved. 6414 reserved ``s_trap 0x06`` Reserved. 6415 debugger breakpoint ``s_trap 0x07`` Reserved for debugger 6416 breakpoints. 6417 reserved ``s_trap 0x08`` Reserved. 6418 reserved ``s_trap 0xfe`` Reserved. 6419 reserved ``s_trap 0xff`` Reserved. 6420 =================== =============== =============== ======================= 6421 6422.. _amdgpu-amdhsa-function-call-convention: 6423 6424Call Convention 6425~~~~~~~~~~~~~~~ 6426 6427.. note:: 6428 6429 This section is currently incomplete and has inakkuracies. It is WIP that will 6430 be updated as information is determined. 6431 6432See :ref:`amdgpu-dwarf-address-space-identifier` for information on swizzled 6433addresses. Unswizzled addresses are normal linear addresses. 6434 6435.. _amdgpu-amdhsa-function-call-convention-kernel-functions: 6436 6437Kernel Functions 6438++++++++++++++++ 6439 6440This section describes the call convention ABI for the outer kernel function. 6441 6442See :ref:`amdgpu-amdhsa-initial-kernel-execution-state` for the kernel call 6443convention. 6444 6445The following is not part of the AMDGPU kernel calling convention but describes 6446how the AMDGPU implements function calls: 6447 64481. Clang decides the kernarg layout to match the *HSA Programmer's Language 6449 Reference* [HSA]_. 6450 6451 - All structs are passed directly. 6452 - Lambda values are passed *TBA*. 6453 6454 .. TODO:: 6455 6456 - Does this really follow HSA rules? Or are structs >16 bytes passed 6457 by-value struct? 6458 - What is ABI for lambda values? 6459 64604. The kernel performs certain setup in its prolog, as described in 6461 :ref:`amdgpu-amdhsa-kernel-prolog`. 6462 6463.. _amdgpu-amdhsa-function-call-convention-non-kernel-functions: 6464 6465Non-Kernel Functions 6466++++++++++++++++++++ 6467 6468This section describes the call convention ABI for functions other than the 6469outer kernel function. 6470 6471If a kernel has function calls then scratch is always allocated and used for 6472the call stack which grows from low address to high address using the swizzled 6473scratch address space. 6474 6475On entry to a function: 6476 64771. SGPR0-3 contain a V# with the following properties (see 6478 :ref:`amdgpu-amdhsa-kernel-prolog-private-segment-buffer`): 6479 6480 * Base address pointing to the beginning of the wavefront scratch backing 6481 memory. 6482 * Swizzled with dword element size and stride of wavefront size elements. 6483 64842. The FLAT_SCRATCH register pair is setup. See 6485 :ref:`amdgpu-amdhsa-kernel-prolog-flat-scratch`. 64863. GFX6-8: M0 register set to the size of LDS in bytes. See 6487 :ref:`amdgpu-amdhsa-kernel-prolog-m0`. 64884. The EXEC register is set to the lanes active on entry to the function. 64895. MODE register: *TBD* 64906. VGPR0-31 and SGPR4-29 are used to pass function input arguments as described 6491 below. 64927. SGPR30-31 return address (RA). The code address that the function must 6493 return to when it completes. The value is undefined if the function is *no 6494 return*. 64958. SGPR32 is used for the stack pointer (SP). It is an unswizzled scratch 6496 offset relative to the beginning of the wavefront scratch backing memory. 6497 6498 The unswizzled SP can be used with buffer instructions as an unswizzled SGPR 6499 offset with the scratch V# in SGPR0-3 to access the stack in a swizzled 6500 manner. 6501 6502 The unswizzled SP value can be converted into the swizzled SP value by: 6503 6504 | swizzled SP = unswizzled SP / wavefront size 6505 6506 This may be used to obtain the private address space address of stack 6507 objects and to convert this address to a flat address by adding the flat 6508 scratch aperture base address. 6509 6510 The swizzled SP value is always 4 bytes aligned for the ``r600`` 6511 architecture and 16 byte aligned for the ``amdgcn`` architecture. 6512 6513 .. note:: 6514 6515 The ``amdgcn`` value is selected to avoid dynamic stack alignment for the 6516 OpenCL language which has the largest base type defined as 16 bytes. 6517 6518 On entry, the swizzled SP value is the address of the first function 6519 argument passed on the stack. Other stack passed arguments are positive 6520 offsets from the entry swizzled SP value. 6521 6522 The function may use positive offsets beyond the last stack passed argument 6523 for stack allocated local variables and register spill slots. If necessary, 6524 the function may align these to greater alignment than 16 bytes. After these 6525 the function may dynamically allocate space for such things as runtime sized 6526 ``alloca`` local allocations. 6527 6528 If the function calls another function, it will place any stack allocated 6529 arguments after the last local allocation and adjust SGPR32 to the address 6530 after the last local allocation. 6531 65329. All other registers are unspecified. 653310. Any necessary ``waitcnt`` has been performed to ensure memory is available 6534 to the function. 6535 6536On exit from a function: 6537 65381. VGPR0-31 and SGPR4-29 are used to pass function result arguments as 6539 described below. Any registers used are considered clobbered registers. 65402. The following registers are preserved and have the same value as on entry: 6541 6542 * FLAT_SCRATCH 6543 * EXEC 6544 * GFX6-8: M0 6545 * All SGPR registers except the clobbered registers of SGPR4-31. 6546 * VGPR40-47 6547 VGPR56-63 6548 VGPR72-79 6549 VGPR88-95 6550 VGPR104-111 6551 VGPR120-127 6552 VGPR136-143 6553 VGPR152-159 6554 VGPR168-175 6555 VGPR184-191 6556 VGPR200-207 6557 VGPR216-223 6558 VGPR232-239 6559 VGPR248-255 6560 6561 *Except the argument registers, the VGPR cloberred and the preserved 6562 registers are intermixed at regular intervals in order to 6563 get a better occupancy.* 6564 6565 For the AMDGPU backend, an inter-procedural register allocation (IPRA) 6566 optimization may mark some of clobbered SGPR and VGPR registers as 6567 preserved if it can be determined that the called function does not change 6568 their value. 6569 65702. The PC is set to the RA provided on entry. 65713. MODE register: *TBD*. 65724. All other registers are clobbered. 65735. Any necessary ``waitcnt`` has been performed to ensure memory accessed by 6574 function is available to the caller. 6575 6576.. TODO:: 6577 6578 - On gfx908 are all ACC registers clobbered? 6579 6580 - How are function results returned? The address of structured types is passed 6581 by reference, but what about other types? 6582 6583The function input arguments are made up of the formal arguments explicitly 6584declared by the source language function plus the implicit input arguments used 6585by the implementation. 6586 6587The source language input arguments are: 6588 65891. Any source language implicit ``this`` or ``self`` argument comes first as a 6590 pointer type. 65912. Followed by the function formal arguments in left to right source order. 6592 6593The source language result arguments are: 6594 65951. The function result argument. 6596 6597The source language input or result struct type arguments that are less than or 6598equal to 16 bytes, are decomposed recursively into their base type fields, and 6599each field is passed as if a separate argument. For input arguments, if the 6600called function requires the struct to be in memory, for example because its 6601address is taken, then the function body is responsible for allocating a stack 6602location and copying the field arguments into it. Clang terms this *direct 6603struct*. 6604 6605The source language input struct type arguments that are greater than 16 bytes, 6606are passed by reference. The caller is responsible for allocating a stack 6607location to make a copy of the struct value and pass the address as the input 6608argument. The called function is responsible to perform the dereference when 6609accessing the input argument. Clang terms this *by-value struct*. 6610 6611A source language result struct type argument that is greater than 16 bytes, is 6612returned by reference. The caller is responsible for allocating a stack location 6613to hold the result value and passes the address as the last input argument 6614(before the implicit input arguments). In this case there are no result 6615arguments. The called function is responsible to perform the dereference when 6616storing the result value. Clang terms this *structured return (sret)*. 6617 6618*TODO: correct the ``sret`` definition.* 6619 6620.. TODO:: 6621 6622 Is this definition correct? Or is ``sret`` only used if passing in registers, and 6623 pass as non-decomposed struct as stack argument? Or something else? Is the 6624 memory location in the caller stack frame, or a stack memory argument and so 6625 no address is passed as the caller can directly write to the argument stack 6626 location? But then the stack location is still live after return. If an 6627 argument stack location is it the first stack argument or the last one? 6628 6629Lambda argument types are treated as struct types with an implementation defined 6630set of fields. 6631 6632.. TODO:: 6633 6634 Need to specify the ABI for lambda types for AMDGPU. 6635 6636For AMDGPU backend all source language arguments (including the decomposed 6637struct type arguments) are passed in VGPRs unless marked ``inreg`` in which case 6638they are passed in SGPRs. 6639 6640The AMDGPU backend walks the function call graph from the leaves to determine 6641which implicit input arguments are used, propagating to each caller of the 6642function. The used implicit arguments are appended to the function arguments 6643after the source language arguments in the following order: 6644 6645.. TODO:: 6646 6647 Is recursion or external functions supported? 6648 66491. Work-Item ID (1 VGPR) 6650 6651 The X, Y and Z work-item ID are packed into a single VGRP with the following 6652 layout. Only fields actually used by the function are set. The other bits 6653 are undefined. 6654 6655 The values come from the initial kernel execution state. See 6656 :ref:`amdgpu-amdhsa-vgpr-register-set-up-order-table`. 6657 6658 .. table:: Work-item implicit argument layout 6659 :name: amdgpu-amdhsa-workitem-implicit-argument-layout-table 6660 6661 ======= ======= ============== 6662 Bits Size Field Name 6663 ======= ======= ============== 6664 9:0 10 bits X Work-Item ID 6665 19:10 10 bits Y Work-Item ID 6666 29:20 10 bits Z Work-Item ID 6667 31:30 2 bits Unused 6668 ======= ======= ============== 6669 66702. Dispatch Ptr (2 SGPRs) 6671 6672 The value comes from the initial kernel execution state. See 6673 :ref:`amdgpu-amdhsa-sgpr-register-set-up-order-table`. 6674 66753. Queue Ptr (2 SGPRs) 6676 6677 The value comes from the initial kernel execution state. See 6678 :ref:`amdgpu-amdhsa-sgpr-register-set-up-order-table`. 6679 66804. Kernarg Segment Ptr (2 SGPRs) 6681 6682 The value comes from the initial kernel execution state. See 6683 :ref:`amdgpu-amdhsa-sgpr-register-set-up-order-table`. 6684 66855. Dispatch id (2 SGPRs) 6686 6687 The value comes from the initial kernel execution state. See 6688 :ref:`amdgpu-amdhsa-sgpr-register-set-up-order-table`. 6689 66906. Work-Group ID X (1 SGPR) 6691 6692 The value comes from the initial kernel execution state. See 6693 :ref:`amdgpu-amdhsa-sgpr-register-set-up-order-table`. 6694 66957. Work-Group ID Y (1 SGPR) 6696 6697 The value comes from the initial kernel execution state. See 6698 :ref:`amdgpu-amdhsa-sgpr-register-set-up-order-table`. 6699 67008. Work-Group ID Z (1 SGPR) 6701 6702 The value comes from the initial kernel execution state. See 6703 :ref:`amdgpu-amdhsa-sgpr-register-set-up-order-table`. 6704 67059. Implicit Argument Ptr (2 SGPRs) 6706 6707 The value is computed by adding an offset to Kernarg Segment Ptr to get the 6708 global address space pointer to the first kernarg implicit argument. 6709 6710The input and result arguments are assigned in order in the following manner: 6711 6712.. note:: 6713 6714 There are likely some errors and omissions in the following description that 6715 need correction. 6716 6717 .. TODO:: 6718 6719 Check the clang source code to decipher how function arguments and return 6720 results are handled. Also see the AMDGPU specific values used. 6721 6722* VGPR arguments are assigned to consecutive VGPRs starting at VGPR0 up to 6723 VGPR31. 6724 6725 If there are more arguments than will fit in these registers, the remaining 6726 arguments are allocated on the stack in order on naturally aligned 6727 addresses. 6728 6729 .. TODO:: 6730 6731 How are overly aligned structures allocated on the stack? 6732 6733* SGPR arguments are assigned to consecutive SGPRs starting at SGPR0 up to 6734 SGPR29. 6735 6736 If there are more arguments than will fit in these registers, the remaining 6737 arguments are allocated on the stack in order on naturally aligned 6738 addresses. 6739 6740Note that decomposed struct type arguments may have some fields passed in 6741registers and some in memory. 6742 6743.. TODO:: 6744 6745 So, a struct which can pass some fields as decomposed register arguments, will 6746 pass the rest as decomposed stack elements? But an argument that will not start 6747 in registers will not be decomposed and will be passed as a non-decomposed 6748 stack value? 6749 6750The following is not part of the AMDGPU function calling convention but 6751describes how the AMDGPU implements function calls: 6752 67531. SGPR33 is used as a frame pointer (FP) if necessary. Like the SP it is an 6754 unswizzled scratch address. It is only needed if runtime sized ``alloca`` 6755 are used, or for the reasons defined in ``SIFrameLowering``. 67562. Runtime stack alignment is supported. SGPR34 is used as a base pointer (BP) 6757 to access the incoming stack arguments in the function. The BP is needed 6758 only when the function requires the runtime stack alignment. 6759 67603. Allocating SGPR arguments on the stack are not supported. 6761 67624. No CFI is currently generated. See 6763 :ref:`amdgpu-dwarf-call-frame-information`. 6764 6765 .. note:: 6766 6767 CFI will be generated that defines the CFA as the unswizzled address 6768 relative to the wave scratch base in the unswizzled private address space 6769 of the lowest address stack allocated local variable. 6770 6771 ``DW_AT_frame_base`` will be defined as the swizzled address in the 6772 swizzled private address space by dividing the CFA by the wavefront size 6773 (since CFA is always at least dword aligned which matches the scratch 6774 swizzle element size). 6775 6776 If no dynamic stack alignment was performed, the stack allocated arguments 6777 are accessed as negative offsets relative to ``DW_AT_frame_base``, and the 6778 local variables and register spill slots are accessed as positive offsets 6779 relative to ``DW_AT_frame_base``. 6780 67815. Function argument passing is implemented by copying the input physical 6782 registers to virtual registers on entry. The register allocator can spill if 6783 necessary. These are copied back to physical registers at call sites. The 6784 net effect is that each function call can have these values in entirely 6785 distinct locations. The IPRA can help avoid shuffling argument registers. 67866. Call sites are implemented by setting up the arguments at positive offsets 6787 from SP. Then SP is incremented to account for the known frame size before 6788 the call and decremented after the call. 6789 6790 .. note:: 6791 6792 The CFI will reflect the changed calculation needed to compute the CFA 6793 from SP. 6794 67957. 4 byte spill slots are used in the stack frame. One slot is allocated for an 6796 emergency spill slot. Buffer instructions are used for stack accesses and 6797 not the ``flat_scratch`` instruction. 6798 6799 .. TODO:: 6800 6801 Explain when the emergency spill slot is used. 6802 6803.. TODO:: 6804 6805 Possible broken issues: 6806 6807 - Stack arguments must be aligned to required alignment. 6808 - Stack is aligned to max(16, max formal argument alignment) 6809 - Direct argument < 64 bits should check register budget. 6810 - Register budget calculation should respect ``inreg`` for SGPR. 6811 - SGPR overflow is not handled. 6812 - struct with 1 member unpeeling is not checking size of member. 6813 - ``sret`` is after ``this`` pointer. 6814 - Caller is not implementing stack realignment: need an extra pointer. 6815 - Should say AMDGPU passes FP rather than SP. 6816 - Should CFI define CFA as address of locals or arguments. Difference is 6817 apparent when have implemented dynamic alignment. 6818 - If ``SCRATCH`` instruction could allow negative offsets, then can make FP be 6819 highest address of stack frame and use negative offset for locals. Would 6820 allow SP to be the same as FP and could support signal-handler-like as now 6821 have a real SP for the top of the stack. 6822 - How is ``sret`` passed on the stack? In argument stack area? Can it overlay 6823 arguments? 6824 6825AMDPAL 6826------ 6827 6828This section provides code conventions used when the target triple OS is 6829``amdpal`` (see :ref:`amdgpu-target-triples`) for passing runtime parameters 6830from the application/runtime to each invocation of a hardware shader. These 6831parameters include both generic, application-controlled parameters called 6832*user data* as well as system-generated parameters that are a product of the 6833draw or dispatch execution. 6834 6835User Data 6836~~~~~~~~~ 6837 6838Each hardware stage has a set of 32-bit *user data registers* which can be 6839written from a command buffer and then loaded into SGPRs when waves are launched 6840via a subsequent dispatch or draw operation. This is the way most arguments are 6841passed from the application/runtime to a hardware shader. 6842 6843Compute User Data 6844~~~~~~~~~~~~~~~~~ 6845 6846Compute shader user data mappings are simpler than graphics shaders and have a 6847fixed mapping. 6848 6849Note that there are always 10 available *user data entries* in registers - 6850entries beyond that limit must be fetched from memory (via the spill table 6851pointer) by the shader. 6852 6853 .. table:: PAL Compute Shader User Data Registers 6854 :name: pal-compute-user-data-registers 6855 6856 ============= ================================ 6857 User Register Description 6858 ============= ================================ 6859 0 Global Internal Table (32-bit pointer) 6860 1 Per-Shader Internal Table (32-bit pointer) 6861 2 - 11 Application-Controlled User Data (10 32-bit values) 6862 12 Spill Table (32-bit pointer) 6863 13 - 14 Thread Group Count (64-bit pointer) 6864 15 GDS Range 6865 ============= ================================ 6866 6867Graphics User Data 6868~~~~~~~~~~~~~~~~~~ 6869 6870Graphics pipelines support a much more flexible user data mapping: 6871 6872 .. table:: PAL Graphics Shader User Data Registers 6873 :name: pal-graphics-user-data-registers 6874 6875 ============= ================================ 6876 User Register Description 6877 ============= ================================ 6878 0 Global Internal Table (32-bit pointer) 6879 + Per-Shader Internal Table (32-bit pointer) 6880 + 1-15 Application Controlled User Data 6881 (1-15 Contiguous 32-bit Values in Registers) 6882 + Spill Table (32-bit pointer) 6883 + Draw Index (First Stage Only) 6884 + Vertex Offset (First Stage Only) 6885 + Instance Offset (First Stage Only) 6886 ============= ================================ 6887 6888 The placement of the global internal table remains fixed in the first *user 6889 data SGPR register*. Otherwise all parameters are optional, and can be mapped 6890 to any desired *user data SGPR register*, with the following restrictions: 6891 6892 * Draw Index, Vertex Offset, and Instance Offset can only be used by the first 6893 active hardware stage in a graphics pipeline (i.e. where the API vertex 6894 shader runs). 6895 6896 * Application-controlled user data must be mapped into a contiguous range of 6897 user data registers. 6898 6899 * The application-controlled user data range supports compaction remapping, so 6900 only *entries* that are actually consumed by the shader must be assigned to 6901 corresponding *registers*. Note that in order to support an efficient runtime 6902 implementation, the remapping must pack *registers* in the same order as 6903 *entries*, with unused *entries* removed. 6904 6905.. _pal_global_internal_table: 6906 6907Global Internal Table 6908~~~~~~~~~~~~~~~~~~~~~ 6909 6910The global internal table is a table of *shader resource descriptors* (SRDs) 6911that define how certain engine-wide, runtime-managed resources should be 6912accessed from a shader. The majority of these resources have HW-defined formats, 6913and it is up to the compiler to write/read data as required by the target 6914hardware. 6915 6916The following table illustrates the required format: 6917 6918 .. table:: PAL Global Internal Table 6919 :name: pal-git-table 6920 6921 ============= ================================ 6922 Offset Description 6923 ============= ================================ 6924 0-3 Graphics Scratch SRD 6925 4-7 Compute Scratch SRD 6926 8-11 ES/GS Ring Output SRD 6927 12-15 ES/GS Ring Input SRD 6928 16-19 GS/VS Ring Output #0 6929 20-23 GS/VS Ring Output #1 6930 24-27 GS/VS Ring Output #2 6931 28-31 GS/VS Ring Output #3 6932 32-35 GS/VS Ring Input SRD 6933 36-39 Tessellation Factor Buffer SRD 6934 40-43 Off-Chip LDS Buffer SRD 6935 44-47 Off-Chip Param Cache Buffer SRD 6936 48-51 Sample Position Buffer SRD 6937 52 vaRange::ShadowDescriptorTable High Bits 6938 ============= ================================ 6939 6940 The pointer to the global internal table passed to the shader as user data 6941 is a 32-bit pointer. The top 32 bits should be assumed to be the same as 6942 the top 32 bits of the pipeline, so the shader may use the program 6943 counter's top 32 bits. 6944 6945Unspecified OS 6946-------------- 6947 6948This section provides code conventions used when the target triple OS is 6949empty (see :ref:`amdgpu-target-triples`). 6950 6951Trap Handler ABI 6952~~~~~~~~~~~~~~~~ 6953 6954For code objects generated by AMDGPU backend for non-amdhsa OS, the runtime does 6955not install a trap handler. The ``llvm.trap`` and ``llvm.debugtrap`` 6956instructions are handled as follows: 6957 6958 .. table:: AMDGPU Trap Handler for Non-AMDHSA OS 6959 :name: amdgpu-trap-handler-for-non-amdhsa-os-table 6960 6961 =============== =============== =========================================== 6962 Usage Code Sequence Description 6963 =============== =============== =========================================== 6964 llvm.trap s_endpgm Causes wavefront to be terminated. 6965 llvm.debugtrap *none* Compiler warning given that there is no 6966 trap handler installed. 6967 =============== =============== =========================================== 6968 6969Source Languages 6970================ 6971 6972.. _amdgpu-opencl: 6973 6974OpenCL 6975------ 6976 6977When the language is OpenCL the following differences occur: 6978 69791. The OpenCL memory model is used (see :ref:`amdgpu-amdhsa-memory-model`). 69802. The AMDGPU backend appends additional arguments to the kernel's explicit 6981 arguments for the AMDHSA OS (see 6982 :ref:`opencl-kernel-implicit-arguments-appended-for-amdhsa-os-table`). 69833. Additional metadata is generated 6984 (see :ref:`amdgpu-amdhsa-code-object-metadata`). 6985 6986 .. table:: OpenCL kernel implicit arguments appended for AMDHSA OS 6987 :name: opencl-kernel-implicit-arguments-appended-for-amdhsa-os-table 6988 6989 ======== ==== ========= =========================================== 6990 Position Byte Byte Description 6991 Size Alignment 6992 ======== ==== ========= =========================================== 6993 1 8 8 OpenCL Global Offset X 6994 2 8 8 OpenCL Global Offset Y 6995 3 8 8 OpenCL Global Offset Z 6996 4 8 8 OpenCL address of printf buffer 6997 5 8 8 OpenCL address of virtual queue used by 6998 enqueue_kernel. 6999 6 8 8 OpenCL address of AqlWrap struct used by 7000 enqueue_kernel. 7001 7 8 8 Pointer argument used for Multi-gird 7002 synchronization. 7003 ======== ==== ========= =========================================== 7004 7005.. _amdgpu-hcc: 7006 7007HCC 7008--- 7009 7010When the language is HCC the following differences occur: 7011 70121. The HSA memory model is used (see :ref:`amdgpu-amdhsa-memory-model`). 7013 7014.. _amdgpu-assembler: 7015 7016Assembler 7017--------- 7018 7019AMDGPU backend has LLVM-MC based assembler which is currently in development. 7020It supports AMDGCN GFX6-GFX10. 7021 7022This section describes general syntax for instructions and operands. 7023 7024Instructions 7025~~~~~~~~~~~~ 7026 7027An instruction has the following :doc:`syntax<AMDGPUInstructionSyntax>`: 7028 7029 | ``<``\ *opcode*\ ``> <``\ *operand0*\ ``>, <``\ *operand1*\ ``>,... 7030 <``\ *modifier0*\ ``> <``\ *modifier1*\ ``>...`` 7031 7032:doc:`Operands<AMDGPUOperandSyntax>` are comma-separated while 7033:doc:`modifiers<AMDGPUModifierSyntax>` are space-separated. 7034 7035The order of operands and modifiers is fixed. 7036Most modifiers are optional and may be omitted. 7037 7038Links to detailed instruction syntax description may be found in the following 7039table. Note that features under development are not included 7040in this description. 7041 7042 =================================== ======================================= 7043 Core ISA ISA Extensions 7044 =================================== ======================================= 7045 :doc:`GFX7<AMDGPU/AMDGPUAsmGFX7>` \- 7046 :doc:`GFX8<AMDGPU/AMDGPUAsmGFX8>` \- 7047 :doc:`GFX9<AMDGPU/AMDGPUAsmGFX9>` :doc:`gfx900<AMDGPU/AMDGPUAsmGFX900>` 7048 7049 :doc:`gfx902<AMDGPU/AMDGPUAsmGFX900>` 7050 7051 :doc:`gfx904<AMDGPU/AMDGPUAsmGFX904>` 7052 7053 :doc:`gfx906<AMDGPU/AMDGPUAsmGFX906>` 7054 7055 :doc:`gfx908<AMDGPU/AMDGPUAsmGFX908>` 7056 7057 :doc:`gfx909<AMDGPU/AMDGPUAsmGFX900>` 7058 7059 :doc:`GFX10<AMDGPU/AMDGPUAsmGFX10>` :doc:`gfx1011<AMDGPU/AMDGPUAsmGFX1011>` 7060 7061 :doc:`gfx1012<AMDGPU/AMDGPUAsmGFX1011>` 7062 =================================== ======================================= 7063 7064For more information about instructions, their semantics and supported 7065combinations of operands, refer to one of instruction set architecture manuals 7066[AMD-GCN-GFX6]_, [AMD-GCN-GFX7]_, [AMD-GCN-GFX8]_, [AMD-GCN-GFX9]_ and 7067[AMD-GCN-GFX10]_. 7068 7069Operands 7070~~~~~~~~ 7071 7072Detailed description of operands may be found :doc:`here<AMDGPUOperandSyntax>`. 7073 7074Modifiers 7075~~~~~~~~~ 7076 7077Detailed description of modifiers may be found 7078:doc:`here<AMDGPUModifierSyntax>`. 7079 7080Instruction Examples 7081~~~~~~~~~~~~~~~~~~~~ 7082 7083DS 7084++ 7085 7086.. code-block:: nasm 7087 7088 ds_add_u32 v2, v4 offset:16 7089 ds_write_src2_b64 v2 offset0:4 offset1:8 7090 ds_cmpst_f32 v2, v4, v6 7091 ds_min_rtn_f64 v[8:9], v2, v[4:5] 7092 7093For full list of supported instructions, refer to "LDS/GDS instructions" in ISA 7094Manual. 7095 7096FLAT 7097++++ 7098 7099.. code-block:: nasm 7100 7101 flat_load_dword v1, v[3:4] 7102 flat_store_dwordx3 v[3:4], v[5:7] 7103 flat_atomic_swap v1, v[3:4], v5 glc 7104 flat_atomic_cmpswap v1, v[3:4], v[5:6] glc slc 7105 flat_atomic_fmax_x2 v[1:2], v[3:4], v[5:6] glc 7106 7107For full list of supported instructions, refer to "FLAT instructions" in ISA 7108Manual. 7109 7110MUBUF 7111+++++ 7112 7113.. code-block:: nasm 7114 7115 buffer_load_dword v1, off, s[4:7], s1 7116 buffer_store_dwordx4 v[1:4], v2, ttmp[4:7], s1 offen offset:4 glc tfe 7117 buffer_store_format_xy v[1:2], off, s[4:7], s1 7118 buffer_wbinvl1 7119 buffer_atomic_inc v1, v2, s[8:11], s4 idxen offset:4 slc 7120 7121For full list of supported instructions, refer to "MUBUF Instructions" in ISA 7122Manual. 7123 7124SMRD/SMEM 7125+++++++++ 7126 7127.. code-block:: nasm 7128 7129 s_load_dword s1, s[2:3], 0xfc 7130 s_load_dwordx8 s[8:15], s[2:3], s4 7131 s_load_dwordx16 s[88:103], s[2:3], s4 7132 s_dcache_inv_vol 7133 s_memtime s[4:5] 7134 7135For full list of supported instructions, refer to "Scalar Memory Operations" in 7136ISA Manual. 7137 7138SOP1 7139++++ 7140 7141.. code-block:: nasm 7142 7143 s_mov_b32 s1, s2 7144 s_mov_b64 s[0:1], 0x80000000 7145 s_cmov_b32 s1, 200 7146 s_wqm_b64 s[2:3], s[4:5] 7147 s_bcnt0_i32_b64 s1, s[2:3] 7148 s_swappc_b64 s[2:3], s[4:5] 7149 s_cbranch_join s[4:5] 7150 7151For full list of supported instructions, refer to "SOP1 Instructions" in ISA 7152Manual. 7153 7154SOP2 7155++++ 7156 7157.. code-block:: nasm 7158 7159 s_add_u32 s1, s2, s3 7160 s_and_b64 s[2:3], s[4:5], s[6:7] 7161 s_cselect_b32 s1, s2, s3 7162 s_andn2_b32 s2, s4, s6 7163 s_lshr_b64 s[2:3], s[4:5], s6 7164 s_ashr_i32 s2, s4, s6 7165 s_bfm_b64 s[2:3], s4, s6 7166 s_bfe_i64 s[2:3], s[4:5], s6 7167 s_cbranch_g_fork s[4:5], s[6:7] 7168 7169For full list of supported instructions, refer to "SOP2 Instructions" in ISA 7170Manual. 7171 7172SOPC 7173++++ 7174 7175.. code-block:: nasm 7176 7177 s_cmp_eq_i32 s1, s2 7178 s_bitcmp1_b32 s1, s2 7179 s_bitcmp0_b64 s[2:3], s4 7180 s_setvskip s3, s5 7181 7182For full list of supported instructions, refer to "SOPC Instructions" in ISA 7183Manual. 7184 7185SOPP 7186++++ 7187 7188.. code-block:: nasm 7189 7190 s_barrier 7191 s_nop 2 7192 s_endpgm 7193 s_waitcnt 0 ; Wait for all counters to be 0 7194 s_waitcnt vmcnt(0) & expcnt(0) & lgkmcnt(0) ; Equivalent to above 7195 s_waitcnt vmcnt(1) ; Wait for vmcnt counter to be 1. 7196 s_sethalt 9 7197 s_sleep 10 7198 s_sendmsg 0x1 7199 s_sendmsg sendmsg(MSG_INTERRUPT) 7200 s_trap 1 7201 7202For full list of supported instructions, refer to "SOPP Instructions" in ISA 7203Manual. 7204 7205Unless otherwise mentioned, little verification is performed on the operands 7206of SOPP Instructions, so it is up to the programmer to be familiar with the 7207range or acceptable values. 7208 7209VALU 7210++++ 7211 7212For vector ALU instruction opcodes (VOP1, VOP2, VOP3, VOPC, VOP_DPP, VOP_SDWA), 7213the assembler will automatically use optimal encoding based on its operands. To 7214force specific encoding, one can add a suffix to the opcode of the instruction: 7215 7216* _e32 for 32-bit VOP1/VOP2/VOPC 7217* _e64 for 64-bit VOP3 7218* _dpp for VOP_DPP 7219* _sdwa for VOP_SDWA 7220 7221VOP1/VOP2/VOP3/VOPC examples: 7222 7223.. code-block:: nasm 7224 7225 v_mov_b32 v1, v2 7226 v_mov_b32_e32 v1, v2 7227 v_nop 7228 v_cvt_f64_i32_e32 v[1:2], v2 7229 v_floor_f32_e32 v1, v2 7230 v_bfrev_b32_e32 v1, v2 7231 v_add_f32_e32 v1, v2, v3 7232 v_mul_i32_i24_e64 v1, v2, 3 7233 v_mul_i32_i24_e32 v1, -3, v3 7234 v_mul_i32_i24_e32 v1, -100, v3 7235 v_addc_u32 v1, s[0:1], v2, v3, s[2:3] 7236 v_max_f16_e32 v1, v2, v3 7237 7238VOP_DPP examples: 7239 7240.. code-block:: nasm 7241 7242 v_mov_b32 v0, v0 quad_perm:[0,2,1,1] 7243 v_sin_f32 v0, v0 row_shl:1 row_mask:0xa bank_mask:0x1 bound_ctrl:0 7244 v_mov_b32 v0, v0 wave_shl:1 7245 v_mov_b32 v0, v0 row_mirror 7246 v_mov_b32 v0, v0 row_bcast:31 7247 v_mov_b32 v0, v0 quad_perm:[1,3,0,1] row_mask:0xa bank_mask:0x1 bound_ctrl:0 7248 v_add_f32 v0, v0, |v0| row_shl:1 row_mask:0xa bank_mask:0x1 bound_ctrl:0 7249 v_max_f16 v1, v2, v3 row_shl:1 row_mask:0xa bank_mask:0x1 bound_ctrl:0 7250 7251VOP_SDWA examples: 7252 7253.. code-block:: nasm 7254 7255 v_mov_b32 v1, v2 dst_sel:BYTE_0 dst_unused:UNUSED_PRESERVE src0_sel:DWORD 7256 v_min_u32 v200, v200, v1 dst_sel:WORD_1 dst_unused:UNUSED_PAD src0_sel:BYTE_1 src1_sel:DWORD 7257 v_sin_f32 v0, v0 dst_unused:UNUSED_PAD src0_sel:WORD_1 7258 v_fract_f32 v0, |v0| dst_sel:DWORD dst_unused:UNUSED_PAD src0_sel:WORD_1 7259 v_cmpx_le_u32 vcc, v1, v2 src0_sel:BYTE_2 src1_sel:WORD_0 7260 7261For full list of supported instructions, refer to "Vector ALU instructions". 7262 7263.. TODO:: 7264 7265 Remove once we switch to code object v3 by default. 7266 7267.. _amdgpu-amdhsa-assembler-predefined-symbols-v2: 7268 7269Code Object V2 Predefined Symbols (-mattr=-code-object-v3) 7270~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 7271 7272.. warning:: Code Object V2 is not the default code object version emitted by 7273 this version of LLVM. For a description of the predefined symbols available 7274 with the default configuration (Code Object V3) see 7275 :ref:`amdgpu-amdhsa-assembler-predefined-symbols-v3`. 7276 7277The AMDGPU assembler defines and updates some symbols automatically. These 7278symbols do not affect code generation. 7279 7280.option.machine_version_major 7281+++++++++++++++++++++++++++++ 7282 7283Set to the GFX major generation number of the target being assembled for. For 7284example, when assembling for a "GFX9" target this will be set to the integer 7285value "9". The possible GFX major generation numbers are presented in 7286:ref:`amdgpu-processors`. 7287 7288.option.machine_version_minor 7289+++++++++++++++++++++++++++++ 7290 7291Set to the GFX minor generation number of the target being assembled for. For 7292example, when assembling for a "GFX810" target this will be set to the integer 7293value "1". The possible GFX minor generation numbers are presented in 7294:ref:`amdgpu-processors`. 7295 7296.option.machine_version_stepping 7297++++++++++++++++++++++++++++++++ 7298 7299Set to the GFX stepping generation number of the target being assembled for. 7300For example, when assembling for a "GFX704" target this will be set to the 7301integer value "4". The possible GFX stepping generation numbers are presented 7302in :ref:`amdgpu-processors`. 7303 7304.kernel.vgpr_count 7305++++++++++++++++++ 7306 7307Set to zero each time a 7308:ref:`amdgpu-amdhsa-assembler-directive-amdgpu_hsa_kernel` directive is 7309encountered. At each instruction, if the current value of this symbol is less 7310than or equal to the maximum VPGR number explicitly referenced within that 7311instruction then the symbol value is updated to equal that VGPR number plus 7312one. 7313 7314.kernel.sgpr_count 7315++++++++++++++++++ 7316 7317Set to zero each time a 7318:ref:`amdgpu-amdhsa-assembler-directive-amdgpu_hsa_kernel` directive is 7319encountered. At each instruction, if the current value of this symbol is less 7320than or equal to the maximum VPGR number explicitly referenced within that 7321instruction then the symbol value is updated to equal that SGPR number plus 7322one. 7323 7324.. _amdgpu-amdhsa-assembler-directives-v2: 7325 7326Code Object V2 Directives (-mattr=-code-object-v3) 7327~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 7328 7329.. warning:: Code Object V2 is not the default code object version emitted by 7330 this version of LLVM. For a description of the directives supported with 7331 the default configuration (Code Object V3) see 7332 :ref:`amdgpu-amdhsa-assembler-directives-v3`. 7333 7334AMDGPU ABI defines auxiliary data in output code object. In assembly source, 7335one can specify them with assembler directives. 7336 7337.hsa_code_object_version major, minor 7338+++++++++++++++++++++++++++++++++++++ 7339 7340*major* and *minor* are integers that specify the version of the HSA code 7341object that will be generated by the assembler. 7342 7343.hsa_code_object_isa [major, minor, stepping, vendor, arch] 7344+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ 7345 7346 7347*major*, *minor*, and *stepping* are all integers that describe the instruction 7348set architecture (ISA) version of the assembly program. 7349 7350*vendor* and *arch* are quoted strings. *vendor* should always be equal to 7351"AMD" and *arch* should always be equal to "AMDGPU". 7352 7353By default, the assembler will derive the ISA version, *vendor*, and *arch* 7354from the value of the -mcpu option that is passed to the assembler. 7355 7356.. _amdgpu-amdhsa-assembler-directive-amdgpu_hsa_kernel: 7357 7358.amdgpu_hsa_kernel (name) 7359+++++++++++++++++++++++++ 7360 7361This directives specifies that the symbol with given name is a kernel entry 7362point (label) and the object should contain corresponding symbol of type 7363STT_AMDGPU_HSA_KERNEL. 7364 7365.amd_kernel_code_t 7366++++++++++++++++++ 7367 7368This directive marks the beginning of a list of key / value pairs that are used 7369to specify the amd_kernel_code_t object that will be emitted by the assembler. 7370The list must be terminated by the *.end_amd_kernel_code_t* directive. For any 7371amd_kernel_code_t values that are unspecified a default value will be used. The 7372default value for all keys is 0, with the following exceptions: 7373 7374- *amd_code_version_major* defaults to 1. 7375- *amd_kernel_code_version_minor* defaults to 2. 7376- *amd_machine_kind* defaults to 1. 7377- *amd_machine_version_major*, *machine_version_minor*, and 7378 *amd_machine_version_stepping* are derived from the value of the -mcpu option 7379 that is passed to the assembler. 7380- *kernel_code_entry_byte_offset* defaults to 256. 7381- *wavefront_size* defaults 6 for all targets before GFX10. For GFX10 onwards 7382 defaults to 6 if target feature ``wavefrontsize64`` is enabled, otherwise 5. 7383 Note that wavefront size is specified as a power of two, so a value of **n** 7384 means a size of 2^ **n**. 7385- *call_convention* defaults to -1. 7386- *kernarg_segment_alignment*, *group_segment_alignment*, and 7387 *private_segment_alignment* default to 4. Note that alignments are specified 7388 as a power of 2, so a value of **n** means an alignment of 2^ **n**. 7389- *enable_wgp_mode* defaults to 1 if target feature ``cumode`` is disabled for 7390 GFX10 onwards. 7391- *enable_mem_ordered* defaults to 1 for GFX10 onwards. 7392 7393The *.amd_kernel_code_t* directive must be placed immediately after the 7394function label and before any instructions. 7395 7396For a full list of amd_kernel_code_t keys, refer to AMDGPU ABI document, 7397comments in lib/Target/AMDGPU/AmdKernelCodeT.h and test/CodeGen/AMDGPU/hsa.s. 7398 7399.. _amdgpu-amdhsa-assembler-example-v2: 7400 7401Code Object V2 Example Source Code (-mattr=-code-object-v3) 7402~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 7403 7404.. warning:: Code Object V2 is not the default code object version emitted by 7405 this version of LLVM. For a description of the directives supported with 7406 the default configuration (Code Object V3) see 7407 :ref:`amdgpu-amdhsa-assembler-example-v3`. 7408 7409Here is an example of a minimal assembly source file, defining one HSA kernel: 7410 7411.. code:: 7412 :number-lines: 7413 7414 .hsa_code_object_version 1,0 7415 .hsa_code_object_isa 7416 7417 .hsatext 7418 .globl hello_world 7419 .p2align 8 7420 .amdgpu_hsa_kernel hello_world 7421 7422 hello_world: 7423 7424 .amd_kernel_code_t 7425 enable_sgpr_kernarg_segment_ptr = 1 7426 is_ptr64 = 1 7427 compute_pgm_rsrc1_vgprs = 0 7428 compute_pgm_rsrc1_sgprs = 0 7429 compute_pgm_rsrc2_user_sgpr = 2 7430 compute_pgm_rsrc1_wgp_mode = 0 7431 compute_pgm_rsrc1_mem_ordered = 0 7432 compute_pgm_rsrc1_fwd_progress = 1 7433 .end_amd_kernel_code_t 7434 7435 s_load_dwordx2 s[0:1], s[0:1] 0x0 7436 v_mov_b32 v0, 3.14159 7437 s_waitcnt lgkmcnt(0) 7438 v_mov_b32 v1, s0 7439 v_mov_b32 v2, s1 7440 flat_store_dword v[1:2], v0 7441 s_endpgm 7442 .Lfunc_end0: 7443 .size hello_world, .Lfunc_end0-hello_world 7444 7445.. _amdgpu-amdhsa-assembler-predefined-symbols-v3: 7446 7447Code Object V3 Predefined Symbols (-mattr=+code-object-v3) 7448~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 7449 7450The AMDGPU assembler defines and updates some symbols automatically. These 7451symbols do not affect code generation. 7452 7453.amdgcn.gfx_generation_number 7454+++++++++++++++++++++++++++++ 7455 7456Set to the GFX major generation number of the target being assembled for. For 7457example, when assembling for a "GFX9" target this will be set to the integer 7458value "9". The possible GFX major generation numbers are presented in 7459:ref:`amdgpu-processors`. 7460 7461.amdgcn.gfx_generation_minor 7462++++++++++++++++++++++++++++ 7463 7464Set to the GFX minor generation number of the target being assembled for. For 7465example, when assembling for a "GFX810" target this will be set to the integer 7466value "1". The possible GFX minor generation numbers are presented in 7467:ref:`amdgpu-processors`. 7468 7469.amdgcn.gfx_generation_stepping 7470+++++++++++++++++++++++++++++++ 7471 7472Set to the GFX stepping generation number of the target being assembled for. 7473For example, when assembling for a "GFX704" target this will be set to the 7474integer value "4". The possible GFX stepping generation numbers are presented 7475in :ref:`amdgpu-processors`. 7476 7477.. _amdgpu-amdhsa-assembler-symbol-next_free_vgpr: 7478 7479.amdgcn.next_free_vgpr 7480++++++++++++++++++++++ 7481 7482Set to zero before assembly begins. At each instruction, if the current value 7483of this symbol is less than or equal to the maximum VGPR number explicitly 7484referenced within that instruction then the symbol value is updated to equal 7485that VGPR number plus one. 7486 7487May be used to set the `.amdhsa_next_free_vpgr` directive in 7488:ref:`amdhsa-kernel-directives-table`. 7489 7490May be set at any time, e.g. manually set to zero at the start of each kernel. 7491 7492.. _amdgpu-amdhsa-assembler-symbol-next_free_sgpr: 7493 7494.amdgcn.next_free_sgpr 7495++++++++++++++++++++++ 7496 7497Set to zero before assembly begins. At each instruction, if the current value 7498of this symbol is less than or equal the maximum SGPR number explicitly 7499referenced within that instruction then the symbol value is updated to equal 7500that SGPR number plus one. 7501 7502May be used to set the `.amdhsa_next_free_spgr` directive in 7503:ref:`amdhsa-kernel-directives-table`. 7504 7505May be set at any time, e.g. manually set to zero at the start of each kernel. 7506 7507.. _amdgpu-amdhsa-assembler-directives-v3: 7508 7509Code Object V3 Directives (-mattr=+code-object-v3) 7510~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 7511 7512Directives which begin with ``.amdgcn`` are valid for all ``amdgcn`` 7513architecture processors, and are not OS-specific. Directives which begin with 7514``.amdhsa`` are specific to ``amdgcn`` architecture processors when the 7515``amdhsa`` OS is specified. See :ref:`amdgpu-target-triples` and 7516:ref:`amdgpu-processors`. 7517 7518.amdgcn_target <target> 7519+++++++++++++++++++++++ 7520 7521Optional directive which declares the target supported by the containing 7522assembler source file. Valid values are described in 7523:ref:`amdgpu-amdhsa-code-object-target-identification`. Used by the assembler 7524to validate command-line options such as ``-triple``, ``-mcpu``, and those 7525which specify target features. 7526 7527.amdhsa_kernel <name> 7528+++++++++++++++++++++ 7529 7530Creates a correctly aligned AMDHSA kernel descriptor and a symbol, 7531``<name>.kd``, in the current location of the current section. Only valid when 7532the OS is ``amdhsa``. ``<name>`` must be a symbol that labels the first 7533instruction to execute, and does not need to be previously defined. 7534 7535Marks the beginning of a list of directives used to generate the bytes of a 7536kernel descriptor, as described in :ref:`amdgpu-amdhsa-kernel-descriptor`. 7537Directives which may appear in this list are described in 7538:ref:`amdhsa-kernel-directives-table`. Directives may appear in any order, must 7539be valid for the target being assembled for, and cannot be repeated. Directives 7540support the range of values specified by the field they reference in 7541:ref:`amdgpu-amdhsa-kernel-descriptor`. If a directive is not specified, it is 7542assumed to have its default value, unless it is marked as "Required", in which 7543case it is an error to omit the directive. This list of directives is 7544terminated by an ``.end_amdhsa_kernel`` directive. 7545 7546 .. table:: AMDHSA Kernel Assembler Directives 7547 :name: amdhsa-kernel-directives-table 7548 7549 ======================================================== =================== ============ =================== 7550 Directive Default Supported On Description 7551 ======================================================== =================== ============ =================== 7552 ``.amdhsa_group_segment_fixed_size`` 0 GFX6-GFX10 Controls GROUP_SEGMENT_FIXED_SIZE in 7553 :ref:`amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table`. 7554 ``.amdhsa_private_segment_fixed_size`` 0 GFX6-GFX10 Controls PRIVATE_SEGMENT_FIXED_SIZE in 7555 :ref:`amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table`. 7556 ``.amdhsa_user_sgpr_private_segment_buffer`` 0 GFX6-GFX10 Controls ENABLE_SGPR_PRIVATE_SEGMENT_BUFFER in 7557 :ref:`amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table`. 7558 ``.amdhsa_user_sgpr_dispatch_ptr`` 0 GFX6-GFX10 Controls ENABLE_SGPR_DISPATCH_PTR in 7559 :ref:`amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table`. 7560 ``.amdhsa_user_sgpr_queue_ptr`` 0 GFX6-GFX10 Controls ENABLE_SGPR_QUEUE_PTR in 7561 :ref:`amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table`. 7562 ``.amdhsa_user_sgpr_kernarg_segment_ptr`` 0 GFX6-GFX10 Controls ENABLE_SGPR_KERNARG_SEGMENT_PTR in 7563 :ref:`amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table`. 7564 ``.amdhsa_user_sgpr_dispatch_id`` 0 GFX6-GFX10 Controls ENABLE_SGPR_DISPATCH_ID in 7565 :ref:`amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table`. 7566 ``.amdhsa_user_sgpr_flat_scratch_init`` 0 GFX6-GFX10 Controls ENABLE_SGPR_FLAT_SCRATCH_INIT in 7567 :ref:`amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table`. 7568 ``.amdhsa_user_sgpr_private_segment_size`` 0 GFX6-GFX10 Controls ENABLE_SGPR_PRIVATE_SEGMENT_SIZE in 7569 :ref:`amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table`. 7570 ``.amdhsa_wavefront_size32`` Target GFX10 Controls ENABLE_WAVEFRONT_SIZE32 in 7571 Feature :ref:`amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table`. 7572 Specific 7573 (-wavefrontsize64) 7574 ``.amdhsa_system_sgpr_private_segment_wavefront_offset`` 0 GFX6-GFX10 Controls ENABLE_SGPR_PRIVATE_SEGMENT_WAVEFRONT_OFFSET in 7575 :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`. 7576 ``.amdhsa_system_sgpr_workgroup_id_x`` 1 GFX6-GFX10 Controls ENABLE_SGPR_WORKGROUP_ID_X in 7577 :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`. 7578 ``.amdhsa_system_sgpr_workgroup_id_y`` 0 GFX6-GFX10 Controls ENABLE_SGPR_WORKGROUP_ID_Y in 7579 :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`. 7580 ``.amdhsa_system_sgpr_workgroup_id_z`` 0 GFX6-GFX10 Controls ENABLE_SGPR_WORKGROUP_ID_Z in 7581 :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`. 7582 ``.amdhsa_system_sgpr_workgroup_info`` 0 GFX6-GFX10 Controls ENABLE_SGPR_WORKGROUP_INFO in 7583 :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`. 7584 ``.amdhsa_system_vgpr_workitem_id`` 0 GFX6-GFX10 Controls ENABLE_VGPR_WORKITEM_ID in 7585 :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`. 7586 Possible values are defined in 7587 :ref:`amdgpu-amdhsa-system-vgpr-work-item-id-enumeration-values-table`. 7588 ``.amdhsa_next_free_vgpr`` Required GFX6-GFX10 Maximum VGPR number explicitly referenced, plus one. 7589 Used to calculate GRANULATED_WORKITEM_VGPR_COUNT in 7590 :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`. 7591 ``.amdhsa_next_free_sgpr`` Required GFX6-GFX10 Maximum SGPR number explicitly referenced, plus one. 7592 Used to calculate GRANULATED_WAVEFRONT_SGPR_COUNT in 7593 :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`. 7594 ``.amdhsa_reserve_vcc`` 1 GFX6-GFX10 Whether the kernel may use the special VCC SGPR. 7595 Used to calculate GRANULATED_WAVEFRONT_SGPR_COUNT in 7596 :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`. 7597 ``.amdhsa_reserve_flat_scratch`` 1 GFX7-GFX10 Whether the kernel may use flat instructions to access 7598 scratch memory. Used to calculate 7599 GRANULATED_WAVEFRONT_SGPR_COUNT in 7600 :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`. 7601 ``.amdhsa_reserve_xnack_mask`` Target GFX8-GFX10 Whether the kernel may trigger XNACK replay. 7602 Feature Used to calculate GRANULATED_WAVEFRONT_SGPR_COUNT in 7603 Specific :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`. 7604 (+xnack) 7605 ``.amdhsa_float_round_mode_32`` 0 GFX6-GFX10 Controls FLOAT_ROUND_MODE_32 in 7606 :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`. 7607 Possible values are defined in 7608 :ref:`amdgpu-amdhsa-floating-point-rounding-mode-enumeration-values-table`. 7609 ``.amdhsa_float_round_mode_16_64`` 0 GFX6-GFX10 Controls FLOAT_ROUND_MODE_16_64 in 7610 :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`. 7611 Possible values are defined in 7612 :ref:`amdgpu-amdhsa-floating-point-rounding-mode-enumeration-values-table`. 7613 ``.amdhsa_float_denorm_mode_32`` 0 GFX6-GFX10 Controls FLOAT_DENORM_MODE_32 in 7614 :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`. 7615 Possible values are defined in 7616 :ref:`amdgpu-amdhsa-floating-point-denorm-mode-enumeration-values-table`. 7617 ``.amdhsa_float_denorm_mode_16_64`` 3 GFX6-GFX10 Controls FLOAT_DENORM_MODE_16_64 in 7618 :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`. 7619 Possible values are defined in 7620 :ref:`amdgpu-amdhsa-floating-point-denorm-mode-enumeration-values-table`. 7621 ``.amdhsa_dx10_clamp`` 1 GFX6-GFX10 Controls ENABLE_DX10_CLAMP in 7622 :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`. 7623 ``.amdhsa_ieee_mode`` 1 GFX6-GFX10 Controls ENABLE_IEEE_MODE in 7624 :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`. 7625 ``.amdhsa_fp16_overflow`` 0 GFX9-GFX10 Controls FP16_OVFL in 7626 :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`. 7627 ``.amdhsa_workgroup_processor_mode`` Target GFX10 Controls ENABLE_WGP_MODE in 7628 Feature :ref:`amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table`. 7629 Specific 7630 (-cumode) 7631 ``.amdhsa_memory_ordered`` 1 GFX10 Controls MEM_ORDERED in 7632 :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`. 7633 ``.amdhsa_forward_progress`` 0 GFX10 Controls FWD_PROGRESS in 7634 :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`. 7635 ``.amdhsa_exception_fp_ieee_invalid_op`` 0 GFX6-GFX10 Controls ENABLE_EXCEPTION_IEEE_754_FP_INVALID_OPERATION in 7636 :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`. 7637 ``.amdhsa_exception_fp_denorm_src`` 0 GFX6-GFX10 Controls ENABLE_EXCEPTION_FP_DENORMAL_SOURCE in 7638 :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`. 7639 ``.amdhsa_exception_fp_ieee_div_zero`` 0 GFX6-GFX10 Controls ENABLE_EXCEPTION_IEEE_754_FP_DIVISION_BY_ZERO in 7640 :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`. 7641 ``.amdhsa_exception_fp_ieee_overflow`` 0 GFX6-GFX10 Controls ENABLE_EXCEPTION_IEEE_754_FP_OVERFLOW in 7642 :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`. 7643 ``.amdhsa_exception_fp_ieee_underflow`` 0 GFX6-GFX10 Controls ENABLE_EXCEPTION_IEEE_754_FP_UNDERFLOW in 7644 :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`. 7645 ``.amdhsa_exception_fp_ieee_inexact`` 0 GFX6-GFX10 Controls ENABLE_EXCEPTION_IEEE_754_FP_INEXACT in 7646 :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`. 7647 ``.amdhsa_exception_int_div_zero`` 0 GFX6-GFX10 Controls ENABLE_EXCEPTION_INT_DIVIDE_BY_ZERO in 7648 :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`. 7649 ======================================================== =================== ============ =================== 7650 7651.amdgpu_metadata 7652++++++++++++++++ 7653 7654Optional directive which declares the contents of the ``NT_AMDGPU_METADATA`` 7655note record (see :ref:`amdgpu-elf-note-records-table-v3`). 7656 7657The contents must be in the [YAML]_ markup format, with the same structure and 7658semantics described in :ref:`amdgpu-amdhsa-code-object-metadata-v3`. 7659 7660This directive is terminated by an ``.end_amdgpu_metadata`` directive. 7661 7662.. _amdgpu-amdhsa-assembler-example-v3: 7663 7664Code Object V3 Example Source Code (-mattr=+code-object-v3) 7665~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 7666 7667Here is an example of a minimal assembly source file, defining one HSA kernel: 7668 7669.. code:: 7670 :number-lines: 7671 7672 .amdgcn_target "amdgcn-amd-amdhsa--gfx900+xnack" // optional 7673 7674 .text 7675 .globl hello_world 7676 .p2align 8 7677 .type hello_world,@function 7678 hello_world: 7679 s_load_dwordx2 s[0:1], s[0:1] 0x0 7680 v_mov_b32 v0, 3.14159 7681 s_waitcnt lgkmcnt(0) 7682 v_mov_b32 v1, s0 7683 v_mov_b32 v2, s1 7684 flat_store_dword v[1:2], v0 7685 s_endpgm 7686 .Lfunc_end0: 7687 .size hello_world, .Lfunc_end0-hello_world 7688 7689 .rodata 7690 .p2align 6 7691 .amdhsa_kernel hello_world 7692 .amdhsa_user_sgpr_kernarg_segment_ptr 1 7693 .amdhsa_next_free_vgpr .amdgcn.next_free_vgpr 7694 .amdhsa_next_free_sgpr .amdgcn.next_free_sgpr 7695 .end_amdhsa_kernel 7696 7697 .amdgpu_metadata 7698 --- 7699 amdhsa.version: 7700 - 1 7701 - 0 7702 amdhsa.kernels: 7703 - .name: hello_world 7704 .symbol: hello_world.kd 7705 .kernarg_segment_size: 48 7706 .group_segment_fixed_size: 0 7707 .private_segment_fixed_size: 0 7708 .kernarg_segment_align: 4 7709 .wavefront_size: 64 7710 .sgpr_count: 2 7711 .vgpr_count: 3 7712 .max_flat_workgroup_size: 256 7713 ... 7714 .end_amdgpu_metadata 7715 7716If an assembly source file contains multiple kernels and/or functions, the 7717:ref:`amdgpu-amdhsa-assembler-symbol-next_free_vgpr` and 7718:ref:`amdgpu-amdhsa-assembler-symbol-next_free_sgpr` symbols may be reset using 7719the ``.set <symbol>, <expression>`` directive. For example, in the case of two 7720kernels, where ``function1`` is only called from ``kernel1`` it is sufficient 7721to group the function with the kernel that calls it and reset the symbols 7722between the two connected components: 7723 7724.. code:: 7725 :number-lines: 7726 7727 .amdgcn_target "amdgcn-amd-amdhsa--gfx900+xnack" // optional 7728 7729 // gpr tracking symbols are implicitly set to zero 7730 7731 .text 7732 .globl kern0 7733 .p2align 8 7734 .type kern0,@function 7735 kern0: 7736 // ... 7737 s_endpgm 7738 .Lkern0_end: 7739 .size kern0, .Lkern0_end-kern0 7740 7741 .rodata 7742 .p2align 6 7743 .amdhsa_kernel kern0 7744 // ... 7745 .amdhsa_next_free_vgpr .amdgcn.next_free_vgpr 7746 .amdhsa_next_free_sgpr .amdgcn.next_free_sgpr 7747 .end_amdhsa_kernel 7748 7749 // reset symbols to begin tracking usage in func1 and kern1 7750 .set .amdgcn.next_free_vgpr, 0 7751 .set .amdgcn.next_free_sgpr, 0 7752 7753 .text 7754 .hidden func1 7755 .global func1 7756 .p2align 2 7757 .type func1,@function 7758 func1: 7759 // ... 7760 s_setpc_b64 s[30:31] 7761 .Lfunc1_end: 7762 .size func1, .Lfunc1_end-func1 7763 7764 .globl kern1 7765 .p2align 8 7766 .type kern1,@function 7767 kern1: 7768 // ... 7769 s_getpc_b64 s[4:5] 7770 s_add_u32 s4, s4, func1@rel32@lo+4 7771 s_addc_u32 s5, s5, func1@rel32@lo+4 7772 s_swappc_b64 s[30:31], s[4:5] 7773 // ... 7774 s_endpgm 7775 .Lkern1_end: 7776 .size kern1, .Lkern1_end-kern1 7777 7778 .rodata 7779 .p2align 6 7780 .amdhsa_kernel kern1 7781 // ... 7782 .amdhsa_next_free_vgpr .amdgcn.next_free_vgpr 7783 .amdhsa_next_free_sgpr .amdgcn.next_free_sgpr 7784 .end_amdhsa_kernel 7785 7786These symbols cannot identify connected components in order to automatically 7787track the usage for each kernel. However, in some cases careful organization of 7788the kernels and functions in the source file means there is minimal additional 7789effort required to accurately calculate GPR usage. 7790 7791Additional Documentation 7792======================== 7793 7794.. [AMD-GCN-GFX6] `AMD Southern Islands Series ISA <http://developer.amd.com/wordpress/media/2012/12/AMD_Southern_Islands_Instruction_Set_Architecture.pdf>`__ 7795.. [AMD-GCN-GFX7] `AMD Sea Islands Series ISA <http://developer.amd.com/wordpress/media/2013/07/AMD_Sea_Islands_Instruction_Set_Architecture.pdf>`_ 7796.. [AMD-GCN-GFX8] `AMD GCN3 Instruction Set Architecture <http://amd-dev.wpengine.netdna-cdn.com/wordpress/media/2013/12/AMD_GCN3_Instruction_Set_Architecture_rev1.1.pdf>`__ 7797.. [AMD-GCN-GFX9] `AMD "Vega" Instruction Set Architecture <http://developer.amd.com/wordpress/media/2013/12/Vega_Shader_ISA_28July2017.pdf>`__ 7798.. [AMD-GCN-GFX10] `AMD "RDNA 1.0" Instruction Set Architecture <https://gpuopen.com/wp-content/uploads/2019/08/RDNA_Shader_ISA_5August2019.pdf>`__ 7799.. [AMD-RADEON-HD-2000-3000] `AMD R6xx shader ISA <http://developer.amd.com/wordpress/media/2012/10/R600_Instruction_Set_Architecture.pdf>`__ 7800.. [AMD-RADEON-HD-4000] `AMD R7xx shader ISA <http://developer.amd.com/wordpress/media/2012/10/R700-Family_Instruction_Set_Architecture.pdf>`__ 7801.. [AMD-RADEON-HD-5000] `AMD Evergreen shader ISA <http://developer.amd.com/wordpress/media/2012/10/AMD_Evergreen-Family_Instruction_Set_Architecture.pdf>`__ 7802.. [AMD-RADEON-HD-6000] `AMD Cayman/Trinity shader ISA <http://developer.amd.com/wordpress/media/2012/10/AMD_HD_6900_Series_Instruction_Set_Architecture.pdf>`__ 7803.. [AMD-ROCm] `AMD ROCm Platform <https://rocm-documentation.readthedocs.io>`__ 7804.. [AMD-ROCm-github] `ROCm github <http://github.com/RadeonOpenCompute>`__ 7805.. [CLANG-ATTR] `Attributes in Clang <https://clang.llvm.org/docs/AttributeReference.html>`__ 7806.. [DWARF] `DWARF Debugging Information Format <http://dwarfstd.org/>`__ 7807.. [ELF] `Executable and Linkable Format (ELF) <http://www.sco.com/developers/gabi/>`__ 7808.. [HRF] `Heterogeneous-race-free Memory Models <http://benedictgaster.org/wp-content/uploads/2014/01/asplos269-FINAL.pdf>`__ 7809.. [HSA] `Heterogeneous System Architecture (HSA) Foundation <http://www.hsafoundation.com/>`__ 7810.. [MsgPack] `Message Pack <http://www.msgpack.org/>`__ 7811.. [OpenCL] `The OpenCL Specification Version 2.0 <http://www.khronos.org/registry/cl/specs/opencl-2.0.pdf>`__ 7812.. [SEMVER] `Semantic Versioning <https://semver.org/>`__ 7813.. [YAML] `YAML Ain't Markup Language (YAML™) Version 1.2 <http://www.yaml.org/spec/1.2/spec.html>`__ 7814