1============================= 2User Guide for AMDGPU Backend 3============================= 4 5.. contents:: 6 :local: 7 8.. toctree:: 9 :hidden: 10 11 AMDGPU/AMDGPUAsmGFX7 12 AMDGPU/AMDGPUAsmGFX8 13 AMDGPU/AMDGPUAsmGFX9 14 AMDGPU/AMDGPUAsmGFX900 15 AMDGPU/AMDGPUAsmGFX904 16 AMDGPU/AMDGPUAsmGFX906 17 AMDGPU/AMDGPUAsmGFX908 18 AMDGPU/AMDGPUAsmGFX10 19 AMDGPU/AMDGPUAsmGFX1011 20 AMDGPUModifierSyntax 21 AMDGPUOperandSyntax 22 AMDGPUInstructionSyntax 23 AMDGPUInstructionNotation 24 AMDGPUDwarfExtensionsForHeterogeneousDebugging 25 26Introduction 27============ 28 29The AMDGPU backend provides ISA code generation for AMD GPUs, starting with the 30R600 family up until the current GCN families. It lives in the 31``llvm/lib/Target/AMDGPU`` directory. 32 33LLVM 34==== 35 36.. _amdgpu-target-triples: 37 38Target Triples 39-------------- 40 41Use the ``clang -target <Architecture>-<Vendor>-<OS>-<Environment>`` option to 42specify the target triple: 43 44 .. table:: AMDGPU Architectures 45 :name: amdgpu-architecture-table 46 47 ============ ============================================================== 48 Architecture Description 49 ============ ============================================================== 50 ``r600`` AMD GPUs HD2XXX-HD6XXX for graphics and compute shaders. 51 ``amdgcn`` AMD GPUs GCN GFX6 onwards for graphics and compute shaders. 52 ============ ============================================================== 53 54 .. table:: AMDGPU Vendors 55 :name: amdgpu-vendor-table 56 57 ============ ============================================================== 58 Vendor Description 59 ============ ============================================================== 60 ``amd`` Can be used for all AMD GPU usage. 61 ``mesa3d`` Can be used if the OS is ``mesa3d``. 62 ============ ============================================================== 63 64 .. table:: AMDGPU Operating Systems 65 :name: amdgpu-os-table 66 67 ============== ============================================================ 68 OS Description 69 ============== ============================================================ 70 *<empty>* Defaults to the *unknown* OS. 71 ``amdhsa`` Compute kernels executed on HSA [HSA]_ compatible runtimes 72 such as AMD's ROCm [AMD-ROCm]_. 73 ``amdpal`` Graphic shaders and compute kernels executed on AMD PAL 74 runtime. 75 ``mesa3d`` Graphic shaders and compute kernels executed on Mesa 3D 76 runtime. 77 ============== ============================================================ 78 79 .. table:: AMDGPU Environments 80 :name: amdgpu-environment-table 81 82 ============ ============================================================== 83 Environment Description 84 ============ ============================================================== 85 *<empty>* Default. 86 ============ ============================================================== 87 88.. _amdgpu-processors: 89 90Processors 91---------- 92 93Use the ``clang -mcpu <Processor>`` option to specify the AMDGPU processor. The 94names from both the *Processor* and *Alternative Processor* can be used. 95 96 .. table:: AMDGPU Processors 97 :name: amdgpu-processor-table 98 99 =========== =============== ============ ===== ================= ======= ====================== 100 Processor Alternative Target dGPU/ Target ROCm Example 101 Processor Triple APU Features Support Products 102 Architecture Supported 103 [Default] 104 =========== =============== ============ ===== ================= ======= ====================== 105 **Radeon HD 2000/3000 Series (R600)** [AMD-RADEON-HD-2000-3000]_ 106 ----------------------------------------------------------------------------------------------- 107 ``r600`` ``r600`` dGPU 108 ``r630`` ``r600`` dGPU 109 ``rs880`` ``r600`` dGPU 110 ``rv670`` ``r600`` dGPU 111 **Radeon HD 4000 Series (R700)** [AMD-RADEON-HD-4000]_ 112 ----------------------------------------------------------------------------------------------- 113 ``rv710`` ``r600`` dGPU 114 ``rv730`` ``r600`` dGPU 115 ``rv770`` ``r600`` dGPU 116 **Radeon HD 5000 Series (Evergreen)** [AMD-RADEON-HD-5000]_ 117 ----------------------------------------------------------------------------------------------- 118 ``cedar`` ``r600`` dGPU 119 ``cypress`` ``r600`` dGPU 120 ``juniper`` ``r600`` dGPU 121 ``redwood`` ``r600`` dGPU 122 ``sumo`` ``r600`` dGPU 123 **Radeon HD 6000 Series (Northern Islands)** [AMD-RADEON-HD-6000]_ 124 ----------------------------------------------------------------------------------------------- 125 ``barts`` ``r600`` dGPU 126 ``caicos`` ``r600`` dGPU 127 ``cayman`` ``r600`` dGPU 128 ``turks`` ``r600`` dGPU 129 **GCN GFX6 (Southern Islands (SI))** [AMD-GCN-GFX6]_ 130 ----------------------------------------------------------------------------------------------- 131 ``gfx600`` - ``tahiti`` ``amdgcn`` dGPU 132 ``gfx601`` - ``hainan`` ``amdgcn`` dGPU 133 - ``oland`` 134 - ``pitcairn`` 135 - ``verde`` 136 **GCN GFX7 (Sea Islands (CI))** [AMD-GCN-GFX7]_ 137 ----------------------------------------------------------------------------------------------- 138 ``gfx700`` - ``kaveri`` ``amdgcn`` APU - A6-7000 139 - A6 Pro-7050B 140 - A8-7100 141 - A8 Pro-7150B 142 - A10-7300 143 - A10 Pro-7350B 144 - FX-7500 145 - A8-7200P 146 - A10-7400P 147 - FX-7600P 148 ``gfx701`` - ``hawaii`` ``amdgcn`` dGPU ROCm - FirePro W8100 149 - FirePro W9100 150 - FirePro S9150 151 - FirePro S9170 152 ``gfx702`` ``amdgcn`` dGPU ROCm - Radeon R9 290 153 - Radeon R9 290x 154 - Radeon R390 155 - Radeon R390x 156 ``gfx703`` - ``kabini`` ``amdgcn`` APU - E1-2100 157 - ``mullins`` - E1-2200 158 - E1-2500 159 - E2-3000 160 - E2-3800 161 - A4-5000 162 - A4-5100 163 - A6-5200 164 - A4 Pro-3340B 165 ``gfx704`` - ``bonaire`` ``amdgcn`` dGPU - Radeon HD 7790 166 - Radeon HD 8770 167 - R7 260 168 - R7 260X 169 **GCN GFX8 (Volcanic Islands (VI))** [AMD-GCN-GFX8]_ 170 ----------------------------------------------------------------------------------------------- 171 ``gfx801`` - ``carrizo`` ``amdgcn`` APU - xnack - A6-8500P 172 [on] - Pro A6-8500B 173 - A8-8600P 174 - Pro A8-8600B 175 - FX-8800P 176 - Pro A12-8800B 177 \ ``amdgcn`` APU - xnack ROCm - A10-8700P 178 [on] - Pro A10-8700B 179 - A10-8780P 180 \ ``amdgcn`` APU - xnack - A10-9600P 181 [on] - A10-9630P 182 - A12-9700P 183 - A12-9730P 184 - FX-9800P 185 - FX-9830P 186 \ ``amdgcn`` APU - xnack - E2-9010 187 [on] - A6-9210 188 - A9-9410 189 ``gfx802`` - ``iceland`` ``amdgcn`` dGPU - xnack ROCm - FirePro S7150 190 - ``tonga`` [off] - FirePro S7100 191 - FirePro W7100 192 - Radeon R285 193 - Radeon R9 380 194 - Radeon R9 385 195 - Mobile FirePro 196 M7170 197 ``gfx803`` - ``fiji`` ``amdgcn`` dGPU - xnack ROCm - Radeon R9 Nano 198 [off] - Radeon R9 Fury 199 - Radeon R9 FuryX 200 - Radeon Pro Duo 201 - FirePro S9300x2 202 - Radeon Instinct MI8 203 \ - ``polaris10`` ``amdgcn`` dGPU - xnack ROCm - Radeon RX 470 204 [off] - Radeon RX 480 205 - Radeon Instinct MI6 206 \ - ``polaris11`` ``amdgcn`` dGPU - xnack ROCm - Radeon RX 460 207 [off] 208 ``gfx810`` - ``stoney`` ``amdgcn`` APU - xnack 209 [on] 210 **GCN GFX9** [AMD-GCN-GFX9]_ 211 ----------------------------------------------------------------------------------------------- 212 ``gfx900`` ``amdgcn`` dGPU - xnack ROCm - Radeon Vega 213 [off] Frontier Edition 214 - Radeon RX Vega 56 215 - Radeon RX Vega 64 216 - Radeon RX Vega 64 217 Liquid 218 - Radeon Instinct MI25 219 ``gfx902`` ``amdgcn`` APU - xnack - Ryzen 3 2200G 220 [on] - Ryzen 5 2400G 221 ``gfx904`` ``amdgcn`` dGPU - xnack *TBA* 222 [off] 223 .. TODO:: 224 Add product 225 names. 226 ``gfx906`` ``amdgcn`` dGPU - xnack - Radeon Instinct MI50 227 [off] - Radeon Instinct MI60 228 - sram-ecc - Radeon VII 229 [off] - Radeon Pro VII 230 ``gfx908`` ``amdgcn`` dGPU - xnack *TBA* 231 [off] 232 - sram-ecc 233 [on] 234 .. TODO:: 235 Add product 236 names. 237 ``gfx909`` ``amdgcn`` APU - xnack *TBA* 238 [on] 239 .. TODO:: 240 Add product 241 names. 242 **GCN GFX10** [AMD-GCN-GFX10]_ 243 ----------------------------------------------------------------------------------------------- 244 ``gfx1010`` ``amdgcn`` dGPU - xnack - Radeon RX 5700 245 [off] - Radeon RX 5700 XT 246 - wavefrontsize64 - Radeon Pro 5600 XT 247 [off] 248 - cumode 249 [off] 250 ``gfx1011`` ``amdgcn`` dGPU - xnack - Radeon Pro 5600M 251 [off] 252 - wavefrontsize64 253 [off] 254 - cumode 255 [off] 256 ``gfx1012`` ``amdgcn`` dGPU - xnack - Radeon RX 5500 257 [off] - Radeon RX 5500 XT 258 - wavefrontsize64 259 [off] 260 - cumode 261 [off] 262 ``gfx1030`` ``amdgcn`` dGPU - wavefrontsize64 *TBA* 263 [off] 264 - cumode 265 [off] 266 .. TODO 267 Add product 268 names. 269 ``gfx1031`` ``amdgcn`` dGPU - xnack *TBA* 270 [off] 271 - wavefrontsize64 272 [off] 273 - cumode 274 [off] 275 .. TODO 276 Add product 277 names. 278 =========== =============== ============ ===== ================= ======= ====================== 279 280.. _amdgpu-target-features: 281 282Target Features 283--------------- 284 285Target features control how code is generated to support certain 286processor specific features. Not all target features are supported by 287all processors. The runtime must ensure that the features supported by 288the device used to execute the code match the features enabled when 289generating the code. A mismatch of features may result in incorrect 290execution, or a reduction in performance. 291 292The target features supported by each processor, and the default value 293used if not specified explicitly, is listed in 294:ref:`amdgpu-processor-table`. 295 296Use the ``clang -m[no-]<TargetFeature>`` option to specify the AMDGPU 297target features. 298 299For example: 300 301``-mxnack`` 302 Enable the ``xnack`` feature. 303``-mno-xnack`` 304 Disable the ``xnack`` feature. 305 306 .. table:: AMDGPU Target Features 307 :name: amdgpu-target-feature-table 308 309 ====================== ================================================== 310 Target Feature Description 311 ====================== ================================================== 312 -m[no-]xnack Enable/disable generating code that has 313 memory clauses that are compatible with 314 having XNACK replay enabled. 315 316 This is used for demand paging and page 317 migration. If XNACK replay is enabled in 318 the device, then if a page fault occurs 319 the code may execute incorrectly if the 320 ``xnack`` feature is not enabled. Executing 321 code that has the feature enabled on a 322 device that does not have XNACK replay 323 enabled will execute correctly but may 324 be less performant than code with the 325 feature disabled. 326 327 -m[no-]sram-ecc Enable/disable generating code that assumes SRAM 328 ECC is enabled/disabled. 329 330 -m[no-]wavefrontsize64 Control the default wavefront size used when 331 generating code for kernels. When disabled 332 native wavefront size 32 is used, when enabled 333 wavefront size 64 is used. 334 335 -m[no-]cumode Control the default wavefront execution mode used 336 when generating code for kernels. When disabled 337 native WGP wavefront execution mode is used, 338 when enabled CU wavefront execution mode is used 339 (see :ref:`amdgpu-amdhsa-memory-model`). 340 ====================== ================================================== 341 342.. _amdgpu-address-spaces: 343 344Address Spaces 345-------------- 346 347The AMDGPU architecture supports a number of memory address spaces. The address 348space names use the OpenCL standard names, with some additions. 349 350The AMDGPU address spaces correspond to target architecture specific LLVM 351address space numbers used in LLVM IR. 352 353The AMDGPU address spaces are described in 354:ref:`amdgpu-address-spaces-table`. Only 64-bit process address spaces are 355supported for the ``amdgcn`` target. 356 357 .. table:: AMDGPU Address Spaces 358 :name: amdgpu-address-spaces-table 359 360 ================================= =============== =========== ================ ======= ============================ 361 .. 64-Bit Process Address Space 362 --------------------------------- --------------- ----------- ---------------- ------------------------------------ 363 Address Space Name LLVM IR Address HSA Segment Hardware Address NULL Value 364 Space Number Name Name Size 365 ================================= =============== =========== ================ ======= ============================ 366 Generic 0 flat flat 64 0x0000000000000000 367 Global 1 global global 64 0x0000000000000000 368 Region 2 N/A GDS 32 *not implemented for AMDHSA* 369 Local 3 group LDS 32 0xFFFFFFFF 370 Constant 4 constant *same as global* 64 0x0000000000000000 371 Private 5 private scratch 32 0xFFFFFFFF 372 Constant 32-bit 6 *TODO* 0x00000000 373 Buffer Fat Pointer (experimental) 7 *TODO* 374 ================================= =============== =========== ================ ======= ============================ 375 376**Generic** 377 The generic address space uses the hardware flat address support available in 378 GFX7-GFX10. This uses two fixed ranges of virtual addresses (the private and 379 local apertures), that are outside the range of addressable global memory, to 380 map from a flat address to a private or local address. 381 382 FLAT instructions can take a flat address and access global, private 383 (scratch), and group (LDS) memory depending on if the address is within one 384 of the aperture ranges. Flat access to scratch requires hardware aperture 385 setup and setup in the kernel prologue (see 386 :ref:`amdgpu-amdhsa-kernel-prolog-flat-scratch`). Flat access to LDS requires 387 hardware aperture setup and M0 (GFX7-GFX8) register setup (see 388 :ref:`amdgpu-amdhsa-kernel-prolog-m0`). 389 390 To convert between a private or group address space address (termed a segment 391 address) and a flat address the base address of the corresponding aperture 392 can be used. For GFX7-GFX8 these are available in the 393 :ref:`amdgpu-amdhsa-hsa-aql-queue` the address of which can be obtained with 394 Queue Ptr SGPR (see :ref:`amdgpu-amdhsa-initial-kernel-execution-state`). For 395 GFX9-GFX10 the aperture base addresses are directly available as inline 396 constant registers ``SRC_SHARED_BASE/LIMIT`` and ``SRC_PRIVATE_BASE/LIMIT``. 397 In 64-bit address mode the aperture sizes are 2^32 bytes and the base is 398 aligned to 2^32 which makes it easier to convert from flat to segment or 399 segment to flat. 400 401 A global address space address has the same value when used as a flat address 402 so no conversion is needed. 403 404**Global and Constant** 405 The global and constant address spaces both use global virtual addresses, 406 which are the same virtual address space used by the CPU. However, some 407 virtual addresses may only be accessible to the CPU, some only accessible 408 by the GPU, and some by both. 409 410 Using the constant address space indicates that the data will not change 411 during the execution of the kernel. This allows scalar read instructions to 412 be used. The vector and scalar L1 caches are invalidated of volatile data 413 before each kernel dispatch execution to allow constant memory to change 414 values between kernel dispatches. 415 416**Region** 417 The region address space uses the hardware Global Data Store (GDS). All 418 wavefronts executing on the same device will access the same memory for any 419 given region address. However, the same region address accessed by wavefronts 420 executing on different devices will access different memory. It is higher 421 performance than global memory. It is allocated by the runtime. The data 422 store (DS) instructions can be used to access it. 423 424**Local** 425 The local address space uses the hardware Local Data Store (LDS) which is 426 automatically allocated when the hardware creates the wavefronts of a 427 work-group, and freed when all the wavefronts of a work-group have 428 terminated. All wavefronts belonging to the same work-group will access the 429 same memory for any given local address. However, the same local address 430 accessed by wavefronts belonging to different work-groups will access 431 different memory. It is higher performance than global memory. The data store 432 (DS) instructions can be used to access it. 433 434**Private** 435 The private address space uses the hardware scratch memory support which 436 automatically allocates memory when it creates a wavefront and frees it when 437 a wavefronts terminates. The memory accessed by a lane of a wavefront for any 438 given private address will be different to the memory accessed by another lane 439 of the same or different wavefront for the same private address. 440 441 If a kernel dispatch uses scratch, then the hardware allocates memory from a 442 pool of backing memory allocated by the runtime for each wavefront. The lanes 443 of the wavefront access this using dword (4 byte) interleaving. The mapping 444 used from private address to backing memory address is: 445 446 ``wavefront-scratch-base + 447 ((private-address / 4) * wavefront-size * 4) + 448 (wavefront-lane-id * 4) + (private-address % 4)`` 449 450 If each lane of a wavefront accesses the same private address, the 451 interleaving results in adjacent dwords being accessed and hence requires 452 fewer cache lines to be fetched. 453 454 There are different ways that the wavefront scratch base address is 455 determined by a wavefront (see 456 :ref:`amdgpu-amdhsa-initial-kernel-execution-state`). 457 458 Scratch memory can be accessed in an interleaved manner using buffer 459 instructions with the scratch buffer descriptor and per wavefront scratch 460 offset, by the scratch instructions, or by flat instructions. Multi-dword 461 access is not supported except by flat and scratch instructions in 462 GFX9-GFX10. 463 464**Constant 32-bit** 465 *TODO* 466 467**Buffer Fat Pointer** 468 The buffer fat pointer is an experimental address space that is currently 469 unsupported in the backend. It exposes a non-integral pointer that is in 470 the future intended to support the modelling of 128-bit buffer descriptors 471 plus a 32-bit offset into the buffer (in total encapsulating a 160-bit 472 *pointer*), allowing normal LLVM load/store/atomic operations to be used to 473 model the buffer descriptors used heavily in graphics workloads targeting 474 the backend. 475 476.. _amdgpu-memory-scopes: 477 478Memory Scopes 479------------- 480 481This section provides LLVM memory synchronization scopes supported by the AMDGPU 482backend memory model when the target triple OS is ``amdhsa`` (see 483:ref:`amdgpu-amdhsa-memory-model` and :ref:`amdgpu-target-triples`). 484 485The memory model supported is based on the HSA memory model [HSA]_ which is 486based in turn on HRF-indirect with scope inclusion [HRF]_. The happens-before 487relation is transitive over the synchronizes-with relation independent of scope 488and synchronizes-with allows the memory scope instances to be inclusive (see 489table :ref:`amdgpu-amdhsa-llvm-sync-scopes-table`). 490 491This is different to the OpenCL [OpenCL]_ memory model which does not have scope 492inclusion and requires the memory scopes to exactly match. However, this 493is conservatively correct for OpenCL. 494 495 .. table:: AMDHSA LLVM Sync Scopes 496 :name: amdgpu-amdhsa-llvm-sync-scopes-table 497 498 ======================= =================================================== 499 LLVM Sync Scope Description 500 ======================= =================================================== 501 *none* The default: ``system``. 502 503 Synchronizes with, and participates in modification 504 and seq_cst total orderings with, other operations 505 (except image operations) for all address spaces 506 (except private, or generic that accesses private) 507 provided the other operation's sync scope is: 508 509 - ``system``. 510 - ``agent`` and executed by a thread on the same 511 agent. 512 - ``workgroup`` and executed by a thread in the 513 same work-group. 514 - ``wavefront`` and executed by a thread in the 515 same wavefront. 516 517 ``agent`` Synchronizes with, and participates in modification 518 and seq_cst total orderings with, other operations 519 (except image operations) for all address spaces 520 (except private, or generic that accesses private) 521 provided the other operation's sync scope is: 522 523 - ``system`` or ``agent`` and executed by a thread 524 on the same agent. 525 - ``workgroup`` and executed by a thread in the 526 same work-group. 527 - ``wavefront`` and executed by a thread in the 528 same wavefront. 529 530 ``workgroup`` Synchronizes with, and participates in modification 531 and seq_cst total orderings with, other operations 532 (except image operations) for all address spaces 533 (except private, or generic that accesses private) 534 provided the other operation's sync scope is: 535 536 - ``system``, ``agent`` or ``workgroup`` and 537 executed by a thread in the same work-group. 538 - ``wavefront`` and executed by a thread in the 539 same wavefront. 540 541 ``wavefront`` Synchronizes with, and participates in modification 542 and seq_cst total orderings with, other operations 543 (except image operations) for all address spaces 544 (except private, or generic that accesses private) 545 provided the other operation's sync scope is: 546 547 - ``system``, ``agent``, ``workgroup`` or 548 ``wavefront`` and executed by a thread in the 549 same wavefront. 550 551 ``singlethread`` Only synchronizes with and participates in 552 modification and seq_cst total orderings with, 553 other operations (except image operations) running 554 in the same thread for all address spaces (for 555 example, in signal handlers). 556 557 ``one-as`` Same as ``system`` but only synchronizes with other 558 operations within the same address space. 559 560 ``agent-one-as`` Same as ``agent`` but only synchronizes with other 561 operations within the same address space. 562 563 ``workgroup-one-as`` Same as ``workgroup`` but only synchronizes with 564 other operations within the same address space. 565 566 ``wavefront-one-as`` Same as ``wavefront`` but only synchronizes with 567 other operations within the same address space. 568 569 ``singlethread-one-as`` Same as ``singlethread`` but only synchronizes with 570 other operations within the same address space. 571 ======================= =================================================== 572 573LLVM IR Intrinsics 574------------------ 575 576The AMDGPU backend implements the following LLVM IR intrinsics. 577 578*This section is WIP.* 579 580.. TODO:: 581 582 List AMDGPU intrinsics. 583 584LLVM IR Attributes 585------------------ 586 587The AMDGPU backend supports the following LLVM IR attributes. 588 589 .. table:: AMDGPU LLVM IR Attributes 590 :name: amdgpu-llvm-ir-attributes-table 591 592 ======================================= ========================================================== 593 LLVM Attribute Description 594 ======================================= ========================================================== 595 "amdgpu-flat-work-group-size"="min,max" Specify the minimum and maximum flat work group sizes that 596 will be specified when the kernel is dispatched. Generated 597 by the ``amdgpu_flat_work_group_size`` CLANG attribute [CLANG-ATTR]_. 598 "amdgpu-implicitarg-num-bytes"="n" Number of kernel argument bytes to add to the kernel 599 argument block size for the implicit arguments. This 600 varies by OS and language (for OpenCL see 601 :ref:`opencl-kernel-implicit-arguments-appended-for-amdhsa-os-table`). 602 "amdgpu-num-sgpr"="n" Specifies the number of SGPRs to use. Generated by 603 the ``amdgpu_num_sgpr`` CLANG attribute [CLANG-ATTR]_. 604 "amdgpu-num-vgpr"="n" Specifies the number of VGPRs to use. Generated by the 605 ``amdgpu_num_vgpr`` CLANG attribute [CLANG-ATTR]_. 606 "amdgpu-waves-per-eu"="m,n" Specify the minimum and maximum number of waves per 607 execution unit. Generated by the ``amdgpu_waves_per_eu`` 608 CLANG attribute [CLANG-ATTR]_. 609 "amdgpu-ieee" true/false. Specify whether the function expects the IEEE field of the 610 mode register to be set on entry. Overrides the default for 611 the calling convention. 612 "amdgpu-dx10-clamp" true/false. Specify whether the function expects the DX10_CLAMP field of 613 the mode register to be set on entry. Overrides the default 614 for the calling convention. 615 ======================================= ========================================================== 616 617.. _amdgpu-elf-code-object: 618 619ELF Code Object 620=============== 621 622The AMDGPU backend generates a standard ELF [ELF]_ relocatable code object that 623can be linked by ``lld`` to produce a standard ELF shared code object which can 624be loaded and executed on an AMDGPU target. 625 626.. _amdgpu-elf-header: 627 628Header 629------ 630 631The AMDGPU backend uses the following ELF header: 632 633 .. table:: AMDGPU ELF Header 634 :name: amdgpu-elf-header-table 635 636 ========================== =============================== 637 Field Value 638 ========================== =============================== 639 ``e_ident[EI_CLASS]`` ``ELFCLASS64`` 640 ``e_ident[EI_DATA]`` ``ELFDATA2LSB`` 641 ``e_ident[EI_OSABI]`` - ``ELFOSABI_NONE`` 642 - ``ELFOSABI_AMDGPU_HSA`` 643 - ``ELFOSABI_AMDGPU_PAL`` 644 - ``ELFOSABI_AMDGPU_MESA3D`` 645 ``e_ident[EI_ABIVERSION]`` - ``ELFABIVERSION_AMDGPU_HSA`` 646 - ``ELFABIVERSION_AMDGPU_PAL`` 647 - ``ELFABIVERSION_AMDGPU_MESA3D`` 648 ``e_type`` - ``ET_REL`` 649 - ``ET_DYN`` 650 ``e_machine`` ``EM_AMDGPU`` 651 ``e_entry`` 0 652 ``e_flags`` See :ref:`amdgpu-elf-header-e_flags-table` 653 ========================== =============================== 654 655.. 656 657 .. table:: AMDGPU ELF Header Enumeration Values 658 :name: amdgpu-elf-header-enumeration-values-table 659 660 =============================== ===== 661 Name Value 662 =============================== ===== 663 ``EM_AMDGPU`` 224 664 ``ELFOSABI_NONE`` 0 665 ``ELFOSABI_AMDGPU_HSA`` 64 666 ``ELFOSABI_AMDGPU_PAL`` 65 667 ``ELFOSABI_AMDGPU_MESA3D`` 66 668 ``ELFABIVERSION_AMDGPU_HSA`` 1 669 ``ELFABIVERSION_AMDGPU_PAL`` 0 670 ``ELFABIVERSION_AMDGPU_MESA3D`` 0 671 =============================== ===== 672 673``e_ident[EI_CLASS]`` 674 The ELF class is: 675 676 * ``ELFCLASS32`` for ``r600`` architecture. 677 678 * ``ELFCLASS64`` for ``amdgcn`` architecture which only supports 64-bit 679 process address space applications. 680 681``e_ident[EI_DATA]`` 682 All AMDGPU targets use ``ELFDATA2LSB`` for little-endian byte ordering. 683 684``e_ident[EI_OSABI]`` 685 One of the following AMDGPU target architecture specific OS ABIs 686 (see :ref:`amdgpu-os-table`): 687 688 * ``ELFOSABI_NONE`` for *unknown* OS. 689 690 * ``ELFOSABI_AMDGPU_HSA`` for ``amdhsa`` OS. 691 692 * ``ELFOSABI_AMDGPU_PAL`` for ``amdpal`` OS. 693 694 * ``ELFOSABI_AMDGPU_MESA3D`` for ``mesa3D`` OS. 695 696``e_ident[EI_ABIVERSION]`` 697 The ABI version of the AMDGPU target architecture specific OS ABI to which the code 698 object conforms: 699 700 * ``ELFABIVERSION_AMDGPU_HSA`` is used to specify the version of AMD HSA 701 runtime ABI. 702 703 * ``ELFABIVERSION_AMDGPU_PAL`` is used to specify the version of AMD PAL 704 runtime ABI. 705 706 * ``ELFABIVERSION_AMDGPU_MESA3D`` is used to specify the version of AMD MESA 707 3D runtime ABI. 708 709``e_type`` 710 Can be one of the following values: 711 712 713 ``ET_REL`` 714 The type produced by the AMDGPU backend compiler as it is relocatable code 715 object. 716 717 ``ET_DYN`` 718 The type produced by the linker as it is a shared code object. 719 720 The AMD HSA runtime loader requires a ``ET_DYN`` code object. 721 722``e_machine`` 723 The value ``EM_AMDGPU`` is used for the machine for all processors supported 724 by the ``r600`` and ``amdgcn`` architectures (see 725 :ref:`amdgpu-processor-table`). The specific processor is specified in the 726 ``EF_AMDGPU_MACH`` bit field of the ``e_flags`` (see 727 :ref:`amdgpu-elf-header-e_flags-table`). 728 729``e_entry`` 730 The entry point is 0 as the entry points for individual kernels must be 731 selected in order to invoke them through AQL packets. 732 733``e_flags`` 734 The AMDGPU backend uses the following ELF header flags: 735 736 .. table:: AMDGPU ELF Header ``e_flags`` 737 :name: amdgpu-elf-header-e_flags-table 738 739 ================================= ========== ============================= 740 Name Value Description 741 ================================= ========== ============================= 742 **AMDGPU Processor Flag** See :ref:`amdgpu-processor-table`. 743 -------------------------------------------- ----------------------------- 744 ``EF_AMDGPU_MACH`` 0x000000ff AMDGPU processor selection 745 mask for 746 ``EF_AMDGPU_MACH_xxx`` values 747 defined in 748 :ref:`amdgpu-ef-amdgpu-mach-table`. 749 ``EF_AMDGPU_XNACK`` 0x00000100 Indicates if the ``xnack`` 750 target feature is 751 enabled for all code 752 contained in the code object. 753 If the processor 754 does not support the 755 ``xnack`` target 756 feature then must 757 be 0. 758 See 759 :ref:`amdgpu-target-features`. 760 ``EF_AMDGPU_SRAM_ECC`` 0x00000200 Indicates if the ``sram-ecc`` 761 target feature is 762 enabled for all code 763 contained in the code object. 764 If the processor 765 does not support the 766 ``sram-ecc`` target 767 feature then must 768 be 0. 769 See 770 :ref:`amdgpu-target-features`. 771 ================================= ========== ============================= 772 773 .. table:: AMDGPU ``EF_AMDGPU_MACH`` Values 774 :name: amdgpu-ef-amdgpu-mach-table 775 776 ================================= ========== ============================= 777 Name Value Description (see 778 :ref:`amdgpu-processor-table`) 779 ================================= ========== ============================= 780 ``EF_AMDGPU_MACH_NONE`` 0x000 *not specified* 781 ``EF_AMDGPU_MACH_R600_R600`` 0x001 ``r600`` 782 ``EF_AMDGPU_MACH_R600_R630`` 0x002 ``r630`` 783 ``EF_AMDGPU_MACH_R600_RS880`` 0x003 ``rs880`` 784 ``EF_AMDGPU_MACH_R600_RV670`` 0x004 ``rv670`` 785 ``EF_AMDGPU_MACH_R600_RV710`` 0x005 ``rv710`` 786 ``EF_AMDGPU_MACH_R600_RV730`` 0x006 ``rv730`` 787 ``EF_AMDGPU_MACH_R600_RV770`` 0x007 ``rv770`` 788 ``EF_AMDGPU_MACH_R600_CEDAR`` 0x008 ``cedar`` 789 ``EF_AMDGPU_MACH_R600_CYPRESS`` 0x009 ``cypress`` 790 ``EF_AMDGPU_MACH_R600_JUNIPER`` 0x00a ``juniper`` 791 ``EF_AMDGPU_MACH_R600_REDWOOD`` 0x00b ``redwood`` 792 ``EF_AMDGPU_MACH_R600_SUMO`` 0x00c ``sumo`` 793 ``EF_AMDGPU_MACH_R600_BARTS`` 0x00d ``barts`` 794 ``EF_AMDGPU_MACH_R600_CAICOS`` 0x00e ``caicos`` 795 ``EF_AMDGPU_MACH_R600_CAYMAN`` 0x00f ``cayman`` 796 ``EF_AMDGPU_MACH_R600_TURKS`` 0x010 ``turks`` 797 *reserved* 0x011 - Reserved for ``r600`` 798 0x01f architecture processors. 799 ``EF_AMDGPU_MACH_AMDGCN_GFX600`` 0x020 ``gfx600`` 800 ``EF_AMDGPU_MACH_AMDGCN_GFX601`` 0x021 ``gfx601`` 801 ``EF_AMDGPU_MACH_AMDGCN_GFX700`` 0x022 ``gfx700`` 802 ``EF_AMDGPU_MACH_AMDGCN_GFX701`` 0x023 ``gfx701`` 803 ``EF_AMDGPU_MACH_AMDGCN_GFX702`` 0x024 ``gfx702`` 804 ``EF_AMDGPU_MACH_AMDGCN_GFX703`` 0x025 ``gfx703`` 805 ``EF_AMDGPU_MACH_AMDGCN_GFX704`` 0x026 ``gfx704`` 806 *reserved* 0x027 Reserved. 807 ``EF_AMDGPU_MACH_AMDGCN_GFX801`` 0x028 ``gfx801`` 808 ``EF_AMDGPU_MACH_AMDGCN_GFX802`` 0x029 ``gfx802`` 809 ``EF_AMDGPU_MACH_AMDGCN_GFX803`` 0x02a ``gfx803`` 810 ``EF_AMDGPU_MACH_AMDGCN_GFX810`` 0x02b ``gfx810`` 811 ``EF_AMDGPU_MACH_AMDGCN_GFX900`` 0x02c ``gfx900`` 812 ``EF_AMDGPU_MACH_AMDGCN_GFX902`` 0x02d ``gfx902`` 813 ``EF_AMDGPU_MACH_AMDGCN_GFX904`` 0x02e ``gfx904`` 814 ``EF_AMDGPU_MACH_AMDGCN_GFX906`` 0x02f ``gfx906`` 815 ``EF_AMDGPU_MACH_AMDGCN_GFX908`` 0x030 ``gfx908`` 816 ``EF_AMDGPU_MACH_AMDGCN_GFX909`` 0x031 ``gfx909`` 817 *reserved* 0x032 Reserved. 818 ``EF_AMDGPU_MACH_AMDGCN_GFX1010`` 0x033 ``gfx1010`` 819 ``EF_AMDGPU_MACH_AMDGCN_GFX1011`` 0x034 ``gfx1011`` 820 ``EF_AMDGPU_MACH_AMDGCN_GFX1012`` 0x035 ``gfx1012`` 821 ``EF_AMDGPU_MACH_AMDGCN_GFX1030`` 0x036 ``gfx1030`` 822 ``EF_AMDGPU_MACH_AMDGCN_GFX1031`` 0x037 ``gfx1031`` 823 ================================= ========== ============================= 824 825Sections 826-------- 827 828An AMDGPU target ELF code object has the standard ELF sections which include: 829 830 .. table:: AMDGPU ELF Sections 831 :name: amdgpu-elf-sections-table 832 833 ================== ================ ================================= 834 Name Type Attributes 835 ================== ================ ================================= 836 ``.bss`` ``SHT_NOBITS`` ``SHF_ALLOC`` + ``SHF_WRITE`` 837 ``.data`` ``SHT_PROGBITS`` ``SHF_ALLOC`` + ``SHF_WRITE`` 838 ``.debug_``\ *\** ``SHT_PROGBITS`` *none* 839 ``.dynamic`` ``SHT_DYNAMIC`` ``SHF_ALLOC`` 840 ``.dynstr`` ``SHT_PROGBITS`` ``SHF_ALLOC`` 841 ``.dynsym`` ``SHT_PROGBITS`` ``SHF_ALLOC`` 842 ``.got`` ``SHT_PROGBITS`` ``SHF_ALLOC`` + ``SHF_WRITE`` 843 ``.hash`` ``SHT_HASH`` ``SHF_ALLOC`` 844 ``.note`` ``SHT_NOTE`` *none* 845 ``.rela``\ *name* ``SHT_RELA`` *none* 846 ``.rela.dyn`` ``SHT_RELA`` *none* 847 ``.rodata`` ``SHT_PROGBITS`` ``SHF_ALLOC`` 848 ``.shstrtab`` ``SHT_STRTAB`` *none* 849 ``.strtab`` ``SHT_STRTAB`` *none* 850 ``.symtab`` ``SHT_SYMTAB`` *none* 851 ``.text`` ``SHT_PROGBITS`` ``SHF_ALLOC`` + ``SHF_EXECINSTR`` 852 ================== ================ ================================= 853 854These sections have their standard meanings (see [ELF]_) and are only generated 855if needed. 856 857``.debug``\ *\** 858 The standard DWARF sections. See :ref:`amdgpu-dwarf-debug-information` for 859 information on the DWARF produced by the AMDGPU backend. 860 861``.dynamic``, ``.dynstr``, ``.dynsym``, ``.hash`` 862 The standard sections used by a dynamic loader. 863 864``.note`` 865 See :ref:`amdgpu-note-records` for the note records supported by the AMDGPU 866 backend. 867 868``.rela``\ *name*, ``.rela.dyn`` 869 For relocatable code objects, *name* is the name of the section that the 870 relocation records apply. For example, ``.rela.text`` is the section name for 871 relocation records associated with the ``.text`` section. 872 873 For linked shared code objects, ``.rela.dyn`` contains all the relocation 874 records from each of the relocatable code object's ``.rela``\ *name* sections. 875 876 See :ref:`amdgpu-relocation-records` for the relocation records supported by 877 the AMDGPU backend. 878 879``.text`` 880 The executable machine code for the kernels and functions they call. Generated 881 as position independent code. See :ref:`amdgpu-code-conventions` for 882 information on conventions used in the isa generation. 883 884.. _amdgpu-note-records: 885 886Note Records 887------------ 888 889The AMDGPU backend code object contains ELF note records in the ``.note`` 890section. The set of generated notes and their semantics depend on the code 891object version; see :ref:`amdgpu-note-records-v2` and 892:ref:`amdgpu-note-records-v3`. 893 894As required by ``ELFCLASS32`` and ``ELFCLASS64``, minimal zero-byte padding 895must be generated after the ``name`` field to ensure the ``desc`` field is 4 896byte aligned. In addition, minimal zero-byte padding must be generated to 897ensure the ``desc`` field size is a multiple of 4 bytes. The ``sh_addralign`` 898field of the ``.note`` section must be at least 4 to indicate at least 8 byte 899alignment. 900 901.. _amdgpu-note-records-v2: 902 903Code Object V2 Note Records (-mattr=-code-object-v3) 904~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 905 906.. warning:: Code Object V2 is not the default code object version emitted by 907 this version of LLVM. For a description of the notes generated with the 908 default configuration (Code Object V3) see :ref:`amdgpu-note-records-v3`. 909 910The AMDGPU backend code object uses the following ELF note record in the 911``.note`` section when compiling for Code Object V2 (-mattr=-code-object-v3). 912 913Additional note records may be present, but any which are not documented here 914are deprecated and should not be used. 915 916 .. table:: AMDGPU Code Object V2 ELF Note Records 917 :name: amdgpu-elf-note-records-table-v2 918 919 ===== ============================== ====================================== 920 Name Type Description 921 ===== ============================== ====================================== 922 "AMD" ``NT_AMD_AMDGPU_HSA_METADATA`` <metadata null terminated string> 923 ===== ============================== ====================================== 924 925.. 926 927 .. table:: AMDGPU Code Object V2 ELF Note Record Enumeration Values 928 :name: amdgpu-elf-note-record-enumeration-values-table-v2 929 930 ============================== ===== 931 Name Value 932 ============================== ===== 933 *reserved* 0-9 934 ``NT_AMD_AMDGPU_HSA_METADATA`` 10 935 *reserved* 11 936 ============================== ===== 937 938``NT_AMD_AMDGPU_HSA_METADATA`` 939 Specifies extensible metadata associated with the code objects executed on HSA 940 [HSA]_ compatible runtimes such as AMD's ROCm [AMD-ROCm]_. It is required when 941 the target triple OS is ``amdhsa`` (see :ref:`amdgpu-target-triples`). See 942 :ref:`amdgpu-amdhsa-code-object-metadata-v2` for the syntax of the code 943 object metadata string. 944 945.. _amdgpu-note-records-v3: 946 947Code Object V3 Note Records (-mattr=+code-object-v3) 948~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 949 950The AMDGPU backend code object uses the following ELF note record in the 951``.note`` section when compiling for Code Object V3 (-mattr=+code-object-v3). 952 953Additional note records may be present, but any which are not documented here 954are deprecated and should not be used. 955 956 .. table:: AMDGPU Code Object V3 ELF Note Records 957 :name: amdgpu-elf-note-records-table-v3 958 959 ======== ============================== ====================================== 960 Name Type Description 961 ======== ============================== ====================================== 962 "AMDGPU" ``NT_AMDGPU_METADATA`` Metadata in Message Pack [MsgPack]_ 963 binary format. 964 ======== ============================== ====================================== 965 966.. 967 968 .. table:: AMDGPU Code Object V3 ELF Note Record Enumeration Values 969 :name: amdgpu-elf-note-record-enumeration-values-table-v3 970 971 ============================== ===== 972 Name Value 973 ============================== ===== 974 *reserved* 0-31 975 ``NT_AMDGPU_METADATA`` 32 976 ============================== ===== 977 978``NT_AMDGPU_METADATA`` 979 Specifies extensible metadata associated with an AMDGPU code 980 object. It is encoded as a map in the Message Pack [MsgPack]_ binary 981 data format. See :ref:`amdgpu-amdhsa-code-object-metadata-v3` for the 982 map keys defined for the ``amdhsa`` OS. 983 984.. _amdgpu-symbols: 985 986Symbols 987------- 988 989Symbols include the following: 990 991 .. table:: AMDGPU ELF Symbols 992 :name: amdgpu-elf-symbols-table 993 994 ===================== ================== ================ ================== 995 Name Type Section Description 996 ===================== ================== ================ ================== 997 *link-name* ``STT_OBJECT`` - ``.data`` Global variable 998 - ``.rodata`` 999 - ``.bss`` 1000 *link-name*\ ``.kd`` ``STT_OBJECT`` - ``.rodata`` Kernel descriptor 1001 *link-name* ``STT_FUNC`` - ``.text`` Kernel entry point 1002 *link-name* ``STT_OBJECT`` - SHN_AMDGPU_LDS Global variable in LDS 1003 ===================== ================== ================ ================== 1004 1005Global variable 1006 Global variables both used and defined by the compilation unit. 1007 1008 If the symbol is defined in the compilation unit then it is allocated in the 1009 appropriate section according to if it has initialized data or is readonly. 1010 1011 If the symbol is external then its section is ``STN_UNDEF`` and the loader 1012 will resolve relocations using the definition provided by another code object 1013 or explicitly defined by the runtime. 1014 1015 If the symbol resides in local/group memory (LDS) then its section is the 1016 special processor specific section name ``SHN_AMDGPU_LDS``, and the 1017 ``st_value`` field describes alignment requirements as it does for common 1018 symbols. 1019 1020 .. TODO:: 1021 1022 Add description of linked shared object symbols. Seems undefined symbols 1023 are marked as STT_NOTYPE. 1024 1025Kernel descriptor 1026 Every HSA kernel has an associated kernel descriptor. It is the address of the 1027 kernel descriptor that is used in the AQL dispatch packet used to invoke the 1028 kernel, not the kernel entry point. The layout of the HSA kernel descriptor is 1029 defined in :ref:`amdgpu-amdhsa-kernel-descriptor`. 1030 1031Kernel entry point 1032 Every HSA kernel also has a symbol for its machine code entry point. 1033 1034.. _amdgpu-relocation-records: 1035 1036Relocation Records 1037------------------ 1038 1039AMDGPU backend generates ``Elf64_Rela`` relocation records. Supported 1040relocatable fields are: 1041 1042``word32`` 1043 This specifies a 32-bit field occupying 4 bytes with arbitrary byte 1044 alignment. These values use the same byte order as other word values in the 1045 AMDGPU architecture. 1046 1047``word64`` 1048 This specifies a 64-bit field occupying 8 bytes with arbitrary byte 1049 alignment. These values use the same byte order as other word values in the 1050 AMDGPU architecture. 1051 1052Following notations are used for specifying relocation calculations: 1053 1054**A** 1055 Represents the addend used to compute the value of the relocatable field. 1056 1057**G** 1058 Represents the offset into the global offset table at which the relocation 1059 entry's symbol will reside during execution. 1060 1061**GOT** 1062 Represents the address of the global offset table. 1063 1064**P** 1065 Represents the place (section offset for ``et_rel`` or address for ``et_dyn``) 1066 of the storage unit being relocated (computed using ``r_offset``). 1067 1068**S** 1069 Represents the value of the symbol whose index resides in the relocation 1070 entry. Relocations not using this must specify a symbol index of 1071 ``STN_UNDEF``. 1072 1073**B** 1074 Represents the base address of a loaded executable or shared object which is 1075 the difference between the ELF address and the actual load address. 1076 Relocations using this are only valid in executable or shared objects. 1077 1078The following relocation types are supported: 1079 1080 .. table:: AMDGPU ELF Relocation Records 1081 :name: amdgpu-elf-relocation-records-table 1082 1083 ========================== ======= ===== ========== ============================== 1084 Relocation Type Kind Value Field Calculation 1085 ========================== ======= ===== ========== ============================== 1086 ``R_AMDGPU_NONE`` 0 *none* *none* 1087 ``R_AMDGPU_ABS32_LO`` Static, 1 ``word32`` (S + A) & 0xFFFFFFFF 1088 Dynamic 1089 ``R_AMDGPU_ABS32_HI`` Static, 2 ``word32`` (S + A) >> 32 1090 Dynamic 1091 ``R_AMDGPU_ABS64`` Static, 3 ``word64`` S + A 1092 Dynamic 1093 ``R_AMDGPU_REL32`` Static 4 ``word32`` S + A - P 1094 ``R_AMDGPU_REL64`` Static 5 ``word64`` S + A - P 1095 ``R_AMDGPU_ABS32`` Static, 6 ``word32`` S + A 1096 Dynamic 1097 ``R_AMDGPU_GOTPCREL`` Static 7 ``word32`` G + GOT + A - P 1098 ``R_AMDGPU_GOTPCREL32_LO`` Static 8 ``word32`` (G + GOT + A - P) & 0xFFFFFFFF 1099 ``R_AMDGPU_GOTPCREL32_HI`` Static 9 ``word32`` (G + GOT + A - P) >> 32 1100 ``R_AMDGPU_REL32_LO`` Static 10 ``word32`` (S + A - P) & 0xFFFFFFFF 1101 ``R_AMDGPU_REL32_HI`` Static 11 ``word32`` (S + A - P) >> 32 1102 *reserved* 12 1103 ``R_AMDGPU_RELATIVE64`` Dynamic 13 ``word64`` B + A 1104 ========================== ======= ===== ========== ============================== 1105 1106``R_AMDGPU_ABS32_LO`` and ``R_AMDGPU_ABS32_HI`` are only supported by 1107the ``mesa3d`` OS, which does not support ``R_AMDGPU_ABS64``. 1108 1109There is no current OS loader support for 32-bit programs and so 1110``R_AMDGPU_ABS32`` is not used. 1111 1112.. _amdgpu-loaded-code-object-path-uniform-resource-identifier: 1113 1114Loaded Code Object Path Uniform Resource Identifier (URI) 1115--------------------------------------------------------- 1116 1117The AMD GPU code object loader represents the path of the ELF shared object from 1118which the code object was loaded as a textual Unifom Resource Identifier (URI). 1119Note that the code object is the in memory loaded relocated form of the ELF 1120shared object. Multiple code objects may be loaded at different memory 1121addresses in the same process from the same ELF shared object. 1122 1123The loaded code object path URI syntax is defined by the following BNF syntax: 1124 1125.. code:: 1126 1127 code_object_uri ::== file_uri | memory_uri 1128 file_uri ::== "file://" file_path [ range_specifier ] 1129 memory_uri ::== "memory://" process_id range_specifier 1130 range_specifier ::== [ "#" | "?" ] "offset=" number "&" "size=" number 1131 file_path ::== URI_ENCODED_OS_FILE_PATH 1132 process_id ::== DECIMAL_NUMBER 1133 number ::== HEX_NUMBER | DECIMAL_NUMBER | OCTAL_NUMBER 1134 1135**number** 1136 Is a C integral literal where hexadecimal values are prefixed by "0x" or "0X", 1137 and octal values by "0". 1138 1139**file_path** 1140 Is the file's path specified as a URI encoded UTF-8 string. In URI encoding, 1141 every character that is not in the regular expression ``[a-zA-Z0-9/_.~-]`` is 1142 encoded as two uppercase hexidecimal digits proceeded by "%". Directories in 1143 the path are separated by "/". 1144 1145**offset** 1146 Is a 0-based byte offset to the start of the code object. For a file URI, it 1147 is from the start of the file specified by the ``file_path``, and if omitted 1148 defaults to 0. For a memory URI, it is the memory address and is required. 1149 1150**size** 1151 Is the number of bytes in the code object. For a file URI, if omitted it 1152 defaults to the size of the file. It is required for a memory URI. 1153 1154**process_id** 1155 Is the identity of the process owning the memory. For Linux it is the C 1156 unsigned integral decimal literal for the process ID (PID). 1157 1158For example: 1159 1160.. code:: 1161 1162 file:///dir1/dir2/file1 1163 file:///dir3/dir4/file2#offset=0x2000&size=3000 1164 memory://1234#offset=0x20000&size=3000 1165 1166.. _amdgpu-dwarf-debug-information: 1167 1168DWARF Debug Information 1169======================= 1170 1171.. warning:: 1172 1173 This section describes **provisional support** for AMDGPU DWARF [DWARF]_ that 1174 is not currently fully implemented and is subject to change. 1175 1176AMDGPU generates DWARF [DWARF]_ debugging information ELF sections (see 1177:ref:`amdgpu-elf-code-object`) which contain information that maps the code 1178object executable code and data to the source language constructs. It can be 1179used by tools such as debuggers and profilers. It uses features defined in 1180:doc:`AMDGPUDwarfExtensionsForHeterogeneousDebugging` that are made available in 1181DWARF Version 4 and DWARF Version 5 as an LLVM vendor extension. 1182 1183This section defines the AMDGPU target architecture specific DWARF mappings. 1184 1185.. _amdgpu-dwarf-register-identifier: 1186 1187Register Identifier 1188------------------- 1189 1190This section defines the AMDGPU target architecture register numbers used in 1191DWARF operation expressions (see DWARF Version 5 section 2.5 and 1192:ref:`amdgpu-dwarf-operation-expressions`) and Call Frame Information 1193instructions (see DWARF Version 5 section 6.4 and 1194:ref:`amdgpu-dwarf-call-frame-information`). 1195 1196A single code object can contain code for kernels that have different wavefront 1197sizes. The vector registers and some scalar registers are based on the wavefront 1198size. AMDGPU defines distinct DWARF registers for each wavefront size. This 1199simplifies the consumer of the DWARF so that each register has a fixed size, 1200rather than being dynamic according to the wavefront size mode. Similarly, 1201distinct DWARF registers are defined for those registers that vary in size 1202according to the process address size. This allows a consumer to treat a 1203specific AMDGPU processor as a single architecture regardless of how it is 1204configured at run time. The compiler explicitly specifies the DWARF registers 1205that match the mode in which the code it is generating will be executed. 1206 1207DWARF registers are encoded as numbers, which are mapped to architecture 1208registers. The mapping for AMDGPU is defined in 1209:ref:`amdgpu-dwarf-register-mapping-table`. All AMDGPU targets use the same 1210mapping. 1211 1212.. table:: AMDGPU DWARF Register Mapping 1213 :name: amdgpu-dwarf-register-mapping-table 1214 1215 ============== ================= ======== ================================== 1216 DWARF Register AMDGPU Register Bit Size Description 1217 ============== ================= ======== ================================== 1218 0 PC_32 32 Program Counter (PC) when 1219 executing in a 32-bit process 1220 address space. Used in the CFI to 1221 describe the PC of the calling 1222 frame. 1223 1 EXEC_MASK_32 32 Execution Mask Register when 1224 executing in wavefront 32 mode. 1225 2-15 *Reserved* *Reserved for highly accessed 1226 registers using DWARF shortcut.* 1227 16 PC_64 64 Program Counter (PC) when 1228 executing in a 64-bit process 1229 address space. Used in the CFI to 1230 describe the PC of the calling 1231 frame. 1232 17 EXEC_MASK_64 64 Execution Mask Register when 1233 executing in wavefront 64 mode. 1234 18-31 *Reserved* *Reserved for highly accessed 1235 registers using DWARF shortcut.* 1236 32-95 SGPR0-SGPR63 32 Scalar General Purpose 1237 Registers. 1238 96-127 *Reserved* *Reserved for frequently accessed 1239 registers using DWARF 1-byte ULEB.* 1240 128 SCC 32 Scalar Condition Code Register. 1241 129-511 *Reserved* *Reserved for future Scalar 1242 Architectural Registers.* 1243 512 VCC_32 32 Vector Condition Code Register 1244 when executing in wavefront 32 1245 mode. 1246 513-1023 *Reserved* *Reserved for future Vector 1247 Architectural Registers when 1248 executing in wavefront 32 mode.* 1249 768 VCC_64 32 Vector Condition Code Register 1250 when executing in wavefront 64 1251 mode. 1252 769-1023 *Reserved* *Reserved for future Vector 1253 Architectural Registers when 1254 executing in wavefront 64 mode.* 1255 1024-1087 *Reserved* *Reserved for padding.* 1256 1088-1129 SGPR64-SGPR105 32 Scalar General Purpose Registers. 1257 1130-1535 *Reserved* *Reserved for future Scalar 1258 General Purpose Registers.* 1259 1536-1791 VGPR0-VGPR255 32*32 Vector General Purpose Registers 1260 when executing in wavefront 32 1261 mode. 1262 1792-2047 *Reserved* *Reserved for future Vector 1263 General Purpose Registers when 1264 executing in wavefront 32 mode.* 1265 2048-2303 AGPR0-AGPR255 32*32 Vector Accumulation Registers 1266 when executing in wavefront 32 1267 mode. 1268 2304-2559 *Reserved* *Reserved for future Vector 1269 Accumulation Registers when 1270 executing in wavefront 32 mode.* 1271 2560-2815 VGPR0-VGPR255 64*32 Vector General Purpose Registers 1272 when executing in wavefront 64 1273 mode. 1274 2816-3071 *Reserved* *Reserved for future Vector 1275 General Purpose Registers when 1276 executing in wavefront 64 mode.* 1277 3072-3327 AGPR0-AGPR255 64*32 Vector Accumulation Registers 1278 when executing in wavefront 64 1279 mode. 1280 3328-3583 *Reserved* *Reserved for future Vector 1281 Accumulation Registers when 1282 executing in wavefront 64 mode.* 1283 ============== ================= ======== ================================== 1284 1285The vector registers are represented as the full size for the wavefront. They 1286are organized as consecutive dwords (32-bits), one per lane, with the dword at 1287the least significant bit position corresponding to lane 0 and so forth. DWARF 1288location expressions involving the ``DW_OP_LLVM_offset`` and 1289``DW_OP_LLVM_push_lane`` operations are used to select the part of the vector 1290register corresponding to the lane that is executing the current thread of 1291execution in languages that are implemented using a SIMD or SIMT execution 1292model. 1293 1294If the wavefront size is 32 lanes then the wavefront 32 mode register 1295definitions are used. If the wavefront size is 64 lanes then the wavefront 64 1296mode register definitions are used. Some AMDGPU targets support executing in 1297both wavefront 32 and wavefront 64 mode. The register definitions corresponding 1298to the wavefront mode of the generated code will be used. 1299 1300If code is generated to execute in a 32-bit process address space, then the 130132-bit process address space register definitions are used. If code is generated 1302to execute in a 64-bit process address space, then the 64-bit process address 1303space register definitions are used. The ``amdgcn`` target only supports the 130464-bit process address space. 1305 1306.. _amdgpu-dwarf-address-class-identifier: 1307 1308Address Class Identifier 1309------------------------ 1310 1311The DWARF address class represents the source language memory space. See DWARF 1312Version 5 section 2.12 which is updated by the *DWARF Extensions For 1313Heterogeneous Debugging* section :ref:`amdgpu-dwarf-segment_addresses`. 1314 1315The DWARF address class mapping used for AMDGPU is defined in 1316:ref:`amdgpu-dwarf-address-class-mapping-table`. 1317 1318.. table:: AMDGPU DWARF Address Class Mapping 1319 :name: amdgpu-dwarf-address-class-mapping-table 1320 1321 ========================= ====== ================= 1322 DWARF AMDGPU 1323 -------------------------------- ----------------- 1324 Address Class Name Value Address Space 1325 ========================= ====== ================= 1326 ``DW_ADDR_none`` 0x0000 Generic (Flat) 1327 ``DW_ADDR_LLVM_global`` 0x0001 Global 1328 ``DW_ADDR_LLVM_constant`` 0x0002 Global 1329 ``DW_ADDR_LLVM_group`` 0x0003 Local (group/LDS) 1330 ``DW_ADDR_LLVM_private`` 0x0004 Private (Scratch) 1331 ``DW_ADDR_AMDGPU_region`` 0x8000 Region (GDS) 1332 ========================= ====== ================= 1333 1334The DWARF address class values defined in the *DWARF Extensions For 1335Heterogeneous Debugging* section :ref:`amdgpu-dwarf-segment_addresses` are used. 1336 1337In addition, ``DW_ADDR_AMDGPU_region`` is encoded as a vendor extension. This is 1338available for use for the AMD extension for access to the hardware GDS memory 1339which is scratchpad memory allocated per device. 1340 1341For AMDGPU if no ``DW_AT_address_class`` attribute is present, then the default 1342address class of ``DW_ADDR_none`` is used. 1343 1344See :ref:`amdgpu-dwarf-address-space-identifier` for information on the AMDGPU 1345mapping of DWARF address classes to DWARF address spaces, including address size 1346and NULL value. 1347 1348.. _amdgpu-dwarf-address-space-identifier: 1349 1350Address Space Identifier 1351------------------------ 1352 1353DWARF address spaces correspond to target architecture specific linear 1354addressable memory areas. See DWARF Version 5 section 2.12 and *DWARF Extensions 1355For Heterogeneous Debugging* section :ref:`amdgpu-dwarf-segment_addresses`. 1356 1357The DWARF address space mapping used for AMDGPU is defined in 1358:ref:`amdgpu-dwarf-address-space-mapping-table`. 1359 1360.. table:: AMDGPU DWARF Address Space Mapping 1361 :name: amdgpu-dwarf-address-space-mapping-table 1362 1363 ======================================= ===== ======= ======== ================= ======================= 1364 DWARF AMDGPU Notes 1365 --------------------------------------- ----- ---------------- ----------------- ----------------------- 1366 Address Space Name Value Address Bit Size Address Space 1367 --------------------------------------- ----- ------- -------- ----------------- ----------------------- 1368 .. 64-bit 32-bit 1369 process process 1370 address address 1371 space space 1372 ======================================= ===== ======= ======== ================= ======================= 1373 ``DW_ASPACE_none`` 0x00 8 4 Global *default address space* 1374 ``DW_ASPACE_AMDGPU_generic`` 0x01 8 4 Generic (Flat) 1375 ``DW_ASPACE_AMDGPU_region`` 0x02 4 4 Region (GDS) 1376 ``DW_ASPACE_AMDGPU_local`` 0x03 4 4 Local (group/LDS) 1377 *Reserved* 0x04 1378 ``DW_ASPACE_AMDGPU_private_lane`` 0x05 4 4 Private (Scratch) *focused lane* 1379 ``DW_ASPACE_AMDGPU_private_wave`` 0x06 4 4 Private (Scratch) *unswizzled wavefront* 1380 ======================================= ===== ======= ======== ================= ======================= 1381 1382See :ref:`amdgpu-address-spaces` for information on the AMDGPU address spaces 1383including address size and NULL value. 1384 1385The ``DW_ASPACE_none`` address space is the default target architecture address 1386space used in DWARF operations that do not specify an address space. It 1387therefore has to map to the global address space so that the ``DW_OP_addr*`` and 1388related operations can refer to addresses in the program code. 1389 1390The ``DW_ASPACE_AMDGPU_generic`` address space allows location expressions to 1391specify the flat address space. If the address corresponds to an address in the 1392local address space, then it corresponds to the wavefront that is executing the 1393focused thread of execution. If the address corresponds to an address in the 1394private address space, then it corresponds to the lane that is executing the 1395focused thread of execution for languages that are implemented using a SIMD or 1396SIMT execution model. 1397 1398.. note:: 1399 1400 CUDA-like languages such as HIP that do not have address spaces in the 1401 language type system, but do allow variables to be allocated in different 1402 address spaces, need to explicitly specify the ``DW_ASPACE_AMDGPU_generic`` 1403 address space in the DWARF expression operations as the default address space 1404 is the global address space. 1405 1406The ``DW_ASPACE_AMDGPU_local`` address space allows location expressions to 1407specify the local address space corresponding to the wavefront that is executing 1408the focused thread of execution. 1409 1410The ``DW_ASPACE_AMDGPU_private_lane`` address space allows location expressions 1411to specify the private address space corresponding to the lane that is executing 1412the focused thread of execution for languages that are implemented using a SIMD 1413or SIMT execution model. 1414 1415The ``DW_ASPACE_AMDGPU_private_wave`` address space allows location expressions 1416to specify the unswizzled private address space corresponding to the wavefront 1417that is executing the focused thread of execution. The wavefront view of private 1418memory is the per wavefront unswizzled backing memory layout defined in 1419:ref:`amdgpu-address-spaces`, such that address 0 corresponds to the first 1420location for the backing memory of the wavefront (namely the address is not 1421offset by ``wavefront-scratch-base``). The following formula can be used to 1422convert from a ``DW_ASPACE_AMDGPU_private_lane`` address to a 1423``DW_ASPACE_AMDGPU_private_wave`` address: 1424 1425:: 1426 1427 private-address-wavefront = 1428 ((private-address-lane / 4) * wavefront-size * 4) + 1429 (wavefront-lane-id * 4) + (private-address-lane % 4) 1430 1431If the ``DW_ASPACE_AMDGPU_private_lane`` address is dword aligned, and the start 1432of the dwords for each lane starting with lane 0 is required, then this 1433simplifies to: 1434 1435:: 1436 1437 private-address-wavefront = 1438 private-address-lane * wavefront-size 1439 1440A compiler can use the ``DW_ASPACE_AMDGPU_private_wave`` address space to read a 1441complete spilled vector register back into a complete vector register in the 1442CFI. The frame pointer can be a private lane address which is dword aligned, 1443which can be shifted to multiply by the wavefront size, and then used to form a 1444private wavefront address that gives a location for a contiguous set of dwords, 1445one per lane, where the vector register dwords are spilled. The compiler knows 1446the wavefront size since it generates the code. Note that the type of the 1447address may have to be converted as the size of a 1448``DW_ASPACE_AMDGPU_private_lane`` address may be smaller than the size of a 1449``DW_ASPACE_AMDGPU_private_wave`` address. 1450 1451.. _amdgpu-dwarf-lane-identifier: 1452 1453Lane identifier 1454--------------- 1455 1456DWARF lane identifies specify a target architecture lane position for hardware 1457that executes in a SIMD or SIMT manner, and on which a source language maps its 1458threads of execution onto those lanes. The DWARF lane identifier is pushed by 1459the ``DW_OP_LLVM_push_lane`` DWARF expression operation. See DWARF Version 5 1460section 2.5 which is updated by *DWARF Extensions For Heterogeneous Debugging* 1461section :ref:`amdgpu-dwarf-operation-expressions`. 1462 1463For AMDGPU, the lane identifier corresponds to the hardware lane ID of a 1464wavefront. It is numbered from 0 to the wavefront size minus 1. 1465 1466Operation Expressions 1467--------------------- 1468 1469DWARF expressions are used to compute program values and the locations of 1470program objects. See DWARF Version 5 section 2.5 and 1471:ref:`amdgpu-dwarf-operation-expressions`. 1472 1473DWARF location descriptions describe how to access storage which includes memory 1474and registers. When accessing storage on AMDGPU, bytes are ordered with least 1475significant bytes first, and bits are ordered within bytes with least 1476significant bits first. 1477 1478For AMDGPU CFI expressions, ``DW_OP_LLVM_select_bit_piece`` is used to describe 1479unwinding vector registers that are spilled under the execution mask to memory: 1480the zero-single location description is the vector register, and the one-single 1481location description is the spilled memory location description. The 1482``DW_OP_LLVM_form_aspace_address`` is used to specify the address space of the 1483memory location description. 1484 1485In AMDGPU expressions, ``DW_OP_LLVM_select_bit_piece`` is used by the 1486``DW_AT_LLVM_lane_pc`` attribute expression where divergent control flow is 1487controlled by the execution mask. An undefined location description together 1488with ``DW_OP_LLVM_extend`` is used to indicate the lane was not active on entry 1489to the subprogram. See :ref:`amdgpu-dwarf-dw-at-llvm-lane-pc` for an example. 1490 1491Debugger Information Entry Attributes 1492------------------------------------- 1493 1494This section describes how certain debugger information entry attributes are 1495used by AMDGPU. See the sections in DWARF Version 5 section 2 which are updated 1496by *DWARF Extensions For Heterogeneous Debugging* section 1497:ref:`amdgpu-dwarf-debugging-information-entry-attributes`. 1498 1499.. _amdgpu-dwarf-dw-at-llvm-lane-pc: 1500 1501``DW_AT_LLVM_lane_pc`` 1502~~~~~~~~~~~~~~~~~~~~~~ 1503 1504For AMDGPU, the ``DW_AT_LLVM_lane_pc`` attribute is used to specify the program 1505location of the separate lanes of a SIMT thread. 1506 1507If the lane is an active lane then this will be the same as the current program 1508location. 1509 1510If the lane is inactive, but was active on entry to the subprogram, then this is 1511the program location in the subprogram at which execution of the lane is 1512conceptual positioned. 1513 1514If the lane was not active on entry to the subprogram, then this will be the 1515undefined location. A client debugger can check if the lane is part of a valid 1516work-group by checking that the lane is in the range of the associated 1517work-group within the grid, accounting for partial work-groups. If it is not, 1518then the debugger can omit any information for the lane. Otherwise, the debugger 1519may repeatedly unwind the stack and inspect the ``DW_AT_LLVM_lane_pc`` of the 1520calling subprogram until it finds a non-undefined location. Conceptually the 1521lane only has the call frames that it has a non-undefined 1522``DW_AT_LLVM_lane_pc``. 1523 1524The following example illustrates how the AMDGPU backend can generate a DWARF 1525location list expression for the nested ``IF/THEN/ELSE`` structures of the 1526following subprogram pseudo code for a target with 64 lanes per wavefront. 1527 1528.. code:: 1529 :number-lines: 1530 1531 SUBPROGRAM X 1532 BEGIN 1533 a; 1534 IF (c1) THEN 1535 b; 1536 IF (c2) THEN 1537 c; 1538 ELSE 1539 d; 1540 ENDIF 1541 e; 1542 ELSE 1543 f; 1544 ENDIF 1545 g; 1546 END 1547 1548The AMDGPU backend may generate the following pseudo LLVM MIR to manipulate the 1549execution mask (``EXEC``) to linearize the control flow. The condition is 1550evaluated to make a mask of the lanes for which the condition evaluates to true. 1551First the ``THEN`` region is executed by setting the ``EXEC`` mask to the 1552logical ``AND`` of the current ``EXEC`` mask with the condition mask. Then the 1553``ELSE`` region is executed by negating the ``EXEC`` mask and logical ``AND`` of 1554the saved ``EXEC`` mask at the start of the region. After the ``IF/THEN/ELSE`` 1555region the ``EXEC`` mask is restored to the value it had at the beginning of the 1556region. This is shown below. Other approaches are possible, but the basic 1557concept is the same. 1558 1559.. code:: 1560 :number-lines: 1561 1562 $lex_start: 1563 a; 1564 %1 = EXEC 1565 %2 = c1 1566 $lex_1_start: 1567 EXEC = %1 & %2 1568 $if_1_then: 1569 b; 1570 %3 = EXEC 1571 %4 = c2 1572 $lex_1_1_start: 1573 EXEC = %3 & %4 1574 $lex_1_1_then: 1575 c; 1576 EXEC = ~EXEC & %3 1577 $lex_1_1_else: 1578 d; 1579 EXEC = %3 1580 $lex_1_1_end: 1581 e; 1582 EXEC = ~EXEC & %1 1583 $lex_1_else: 1584 f; 1585 EXEC = %1 1586 $lex_1_end: 1587 g; 1588 $lex_end: 1589 1590To create the DWARF location list expression that defines the location 1591description of a vector of lane program locations, the LLVM MIR ``DBG_VALUE`` 1592pseudo instruction can be used to annotate the linearized control flow. This can 1593be done by defining an artificial variable for the lane PC. The DWARF location 1594list expression created for it is used as the value of the 1595``DW_AT_LLVM_lane_pc`` attribute on the subprogram's debugger information entry. 1596 1597A DWARF procedure is defined for each well nested structured control flow region 1598which provides the conceptual lane program location for a lane if it is not 1599active (namely it is divergent). The DWARF operation expression for each region 1600conceptually inherits the value of the immediately enclosing region and modifies 1601it according to the semantics of the region. 1602 1603For an ``IF/THEN/ELSE`` region the divergent program location is at the start of 1604the region for the ``THEN`` region since it is executed first. For the ``ELSE`` 1605region the divergent program location is at the end of the ``IF/THEN/ELSE`` 1606region since the ``THEN`` region has completed. 1607 1608The lane PC artificial variable is assigned at each region transition. It uses 1609the immediately enclosing region's DWARF procedure to compute the program 1610location for each lane assuming they are divergent, and then modifies the result 1611by inserting the current program location for each lane that the ``EXEC`` mask 1612indicates is active. 1613 1614By having separate DWARF procedures for each region, they can be reused to 1615define the value for any nested region. This reduces the total size of the DWARF 1616operation expressions. 1617 1618The following provides an example using pseudo LLVM MIR. 1619 1620.. code:: 1621 :number-lines: 1622 1623 $lex_start: 1624 DEFINE_DWARF %__uint_64 = DW_TAG_base_type[ 1625 DW_AT_name = "__uint64"; 1626 DW_AT_byte_size = 8; 1627 DW_AT_encoding = DW_ATE_unsigned; 1628 ]; 1629 DEFINE_DWARF %__active_lane_pc = DW_TAG_dwarf_procedure[ 1630 DW_AT_name = "__active_lane_pc"; 1631 DW_AT_location = [ 1632 DW_OP_regx PC; 1633 DW_OP_LLVM_extend 64, 64; 1634 DW_OP_regval_type EXEC, %uint_64; 1635 DW_OP_LLVM_select_bit_piece 64, 64; 1636 ]; 1637 ]; 1638 DEFINE_DWARF %__divergent_lane_pc = DW_TAG_dwarf_procedure[ 1639 DW_AT_name = "__divergent_lane_pc"; 1640 DW_AT_location = [ 1641 DW_OP_LLVM_undefined; 1642 DW_OP_LLVM_extend 64, 64; 1643 ]; 1644 ]; 1645 DBG_VALUE $noreg, $noreg, %DW_AT_LLVM_lane_pc, DIExpression[ 1646 DW_OP_call_ref %__divergent_lane_pc; 1647 DW_OP_call_ref %__active_lane_pc; 1648 ]; 1649 a; 1650 %1 = EXEC; 1651 DBG_VALUE %1, $noreg, %__lex_1_save_exec; 1652 %2 = c1; 1653 $lex_1_start: 1654 EXEC = %1 & %2; 1655 $lex_1_then: 1656 DEFINE_DWARF %__divergent_lane_pc_1_then = DW_TAG_dwarf_procedure[ 1657 DW_AT_name = "__divergent_lane_pc_1_then"; 1658 DW_AT_location = DIExpression[ 1659 DW_OP_call_ref %__divergent_lane_pc; 1660 DW_OP_addrx &lex_1_start; 1661 DW_OP_stack_value; 1662 DW_OP_LLVM_extend 64, 64; 1663 DW_OP_call_ref %__lex_1_save_exec; 1664 DW_OP_deref_type 64, %__uint_64; 1665 DW_OP_LLVM_select_bit_piece 64, 64; 1666 ]; 1667 ]; 1668 DBG_VALUE $noreg, $noreg, %DW_AT_LLVM_lane_pc, DIExpression[ 1669 DW_OP_call_ref %__divergent_lane_pc_1_then; 1670 DW_OP_call_ref %__active_lane_pc; 1671 ]; 1672 b; 1673 %3 = EXEC; 1674 DBG_VALUE %3, %__lex_1_1_save_exec; 1675 %4 = c2; 1676 $lex_1_1_start: 1677 EXEC = %3 & %4; 1678 $lex_1_1_then: 1679 DEFINE_DWARF %__divergent_lane_pc_1_1_then = DW_TAG_dwarf_procedure[ 1680 DW_AT_name = "__divergent_lane_pc_1_1_then"; 1681 DW_AT_location = DIExpression[ 1682 DW_OP_call_ref %__divergent_lane_pc_1_then; 1683 DW_OP_addrx &lex_1_1_start; 1684 DW_OP_stack_value; 1685 DW_OP_LLVM_extend 64, 64; 1686 DW_OP_call_ref %__lex_1_1_save_exec; 1687 DW_OP_deref_type 64, %__uint_64; 1688 DW_OP_LLVM_select_bit_piece 64, 64; 1689 ]; 1690 ]; 1691 DBG_VALUE $noreg, $noreg, %DW_AT_LLVM_lane_pc, DIExpression[ 1692 DW_OP_call_ref %__divergent_lane_pc_1_1_then; 1693 DW_OP_call_ref %__active_lane_pc; 1694 ]; 1695 c; 1696 EXEC = ~EXEC & %3; 1697 $lex_1_1_else: 1698 DEFINE_DWARF %__divergent_lane_pc_1_1_else = DW_TAG_dwarf_procedure[ 1699 DW_AT_name = "__divergent_lane_pc_1_1_else"; 1700 DW_AT_location = DIExpression[ 1701 DW_OP_call_ref %__divergent_lane_pc_1_then; 1702 DW_OP_addrx &lex_1_1_end; 1703 DW_OP_stack_value; 1704 DW_OP_LLVM_extend 64, 64; 1705 DW_OP_call_ref %__lex_1_1_save_exec; 1706 DW_OP_deref_type 64, %__uint_64; 1707 DW_OP_LLVM_select_bit_piece 64, 64; 1708 ]; 1709 ]; 1710 DBG_VALUE $noreg, $noreg, %DW_AT_LLVM_lane_pc, DIExpression[ 1711 DW_OP_call_ref %__divergent_lane_pc_1_1_else; 1712 DW_OP_call_ref %__active_lane_pc; 1713 ]; 1714 d; 1715 EXEC = %3; 1716 $lex_1_1_end: 1717 DBG_VALUE $noreg, $noreg, %DW_AT_LLVM_lane_pc, DIExpression[ 1718 DW_OP_call_ref %__divergent_lane_pc; 1719 DW_OP_call_ref %__active_lane_pc; 1720 ]; 1721 e; 1722 EXEC = ~EXEC & %1; 1723 $lex_1_else: 1724 DEFINE_DWARF %__divergent_lane_pc_1_else = DW_TAG_dwarf_procedure[ 1725 DW_AT_name = "__divergent_lane_pc_1_else"; 1726 DW_AT_location = DIExpression[ 1727 DW_OP_call_ref %__divergent_lane_pc; 1728 DW_OP_addrx &lex_1_end; 1729 DW_OP_stack_value; 1730 DW_OP_LLVM_extend 64, 64; 1731 DW_OP_call_ref %__lex_1_save_exec; 1732 DW_OP_deref_type 64, %__uint_64; 1733 DW_OP_LLVM_select_bit_piece 64, 64; 1734 ]; 1735 ]; 1736 DBG_VALUE $noreg, $noreg, %DW_AT_LLVM_lane_pc, DIExpression[ 1737 DW_OP_call_ref %__divergent_lane_pc_1_else; 1738 DW_OP_call_ref %__active_lane_pc; 1739 ]; 1740 f; 1741 EXEC = %1; 1742 $lex_1_end: 1743 DBG_VALUE $noreg, $noreg, %DW_AT_LLVM_lane_pc DIExpression[ 1744 DW_OP_call_ref %__divergent_lane_pc; 1745 DW_OP_call_ref %__active_lane_pc; 1746 ]; 1747 g; 1748 $lex_end: 1749 1750The DWARF procedure ``%__active_lane_pc`` is used to update the lane pc elements 1751that are active, with the current program location. 1752 1753Artificial variables %__lex_1_save_exec and %__lex_1_1_save_exec are created for 1754the execution masks saved on entry to a region. Using the ``DBG_VALUE`` pseudo 1755instruction, location list entries will be created that describe where the 1756artificial variables are allocated at any given program location. The compiler 1757may allocate them to registers or spill them to memory. 1758 1759The DWARF procedures for each region use the values of the saved execution mask 1760artificial variables to only update the lanes that are active on entry to the 1761region. All other lanes retain the value of the enclosing region where they were 1762last active. If they were not active on entry to the subprogram, then will have 1763the undefined location description. 1764 1765Other structured control flow regions can be handled similarly. For example, 1766loops would set the divergent program location for the region at the end of the 1767loop. Any lanes active will be in the loop, and any lanes not active must have 1768exited the loop. 1769 1770An ``IF/THEN/ELSEIF/ELSEIF/...`` region can be treated as a nest of 1771``IF/THEN/ELSE`` regions. 1772 1773The DWARF procedures can use the active lane artificial variable described in 1774:ref:`amdgpu-dwarf-amdgpu-dw-at-llvm-active-lane` rather than the actual 1775``EXEC`` mask in order to support whole or quad wavefront mode. 1776 1777.. _amdgpu-dwarf-amdgpu-dw-at-llvm-active-lane: 1778 1779``DW_AT_LLVM_active_lane`` 1780~~~~~~~~~~~~~~~~~~~~~~~~~~ 1781 1782The ``DW_AT_LLVM_active_lane`` attribute on a subprogram debugger information 1783entry is used to specify the lanes that are conceptually active for a SIMT 1784thread. 1785 1786The execution mask may be modified to implement whole or quad wavefront mode 1787operations. For example, all lanes may need to temporarily be made active to 1788execute a whole wavefront operation. Such regions would save the ``EXEC`` mask, 1789update it to enable the necessary lanes, perform the operations, and then 1790restore the ``EXEC`` mask from the saved value. While executing the whole 1791wavefront region, the conceptual execution mask is the saved value, not the 1792``EXEC`` value. 1793 1794This is handled by defining an artificial variable for the active lane mask. The 1795active lane mask artificial variable would be the actual ``EXEC`` mask for 1796normal regions, and the saved execution mask for regions where the mask is 1797temporarily updated. The location list expression created for this artificial 1798variable is used to define the value of the ``DW_AT_LLVM_active_lane`` 1799attribute. 1800 1801``DW_AT_LLVM_augmentation`` 1802~~~~~~~~~~~~~~~~~~~~~~~~~~~ 1803 1804For AMDGPU, the ``DW_AT_LLVM_augmentation`` attribute of a compilation unit 1805debugger information entry has the following value for the augmentation string: 1806 1807:: 1808 1809 [amdgpu:v0.0] 1810 1811The "vX.Y" specifies the major X and minor Y version number of the AMDGPU 1812extensions used in the DWARF of the compilation unit. The version number 1813conforms to [SEMVER]_. 1814 1815Call Frame Information 1816---------------------- 1817 1818DWARF Call Frame Information (CFI) describes how a consumer can virtually 1819*unwind* call frames in a running process or core dump. See DWARF Version 5 1820section 6.4 and :ref:`amdgpu-dwarf-call-frame-information`. 1821 1822For AMDGPU, the Common Information Entry (CIE) fields have the following values: 1823 18241. ``augmentation`` string contains the following null-terminated UTF-8 string: 1825 1826 :: 1827 1828 [amd:v0.0] 1829 1830 The ``vX.Y`` specifies the major X and minor Y version number of the AMDGPU 1831 extensions used in this CIE or to the FDEs that use it. The version number 1832 conforms to [SEMVER]_. 1833 18342. ``address_size`` for the ``Global`` address space is defined in 1835 :ref:`amdgpu-dwarf-address-space-identifier`. 1836 18373. ``segment_selector_size`` is 0 as AMDGPU does not use a segment selector. 1838 18394. ``code_alignment_factor`` is 4 bytes. 1840 1841 .. TODO:: 1842 1843 Add to :ref:`amdgpu-processor-table` table. 1844 18455. ``data_alignment_factor`` is 4 bytes. 1846 1847 .. TODO:: 1848 1849 Add to :ref:`amdgpu-processor-table` table. 1850 18516. ``return_address_register`` is ``PC_32`` for 32-bit processes and ``PC_64`` 1852 for 64-bit processes defined in :ref:`amdgpu-dwarf-register-identifier`. 1853 18547. ``initial_instructions`` Since a subprogram X with fewer registers can be 1855 called from subprogram Y that has more allocated, X will not change any of 1856 the extra registers as it cannot access them. Therefore, the default rule 1857 for all columns is ``same value``. 1858 1859For AMDGPU the register number follows the numbering defined in 1860:ref:`amdgpu-dwarf-register-identifier`. 1861 1862For AMDGPU the instructions are variable size. A consumer can subtract 1 from 1863the return address to get the address of a byte within the call site 1864instructions. See DWARF Version 5 section 6.4.4. 1865 1866Accelerated Access 1867------------------ 1868 1869See DWARF Version 5 section 6.1. 1870 1871Lookup By Name Section Header 1872~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 1873 1874See DWARF Version 5 section 6.1.1.4.1 and :ref:`amdgpu-dwarf-lookup-by-name`. 1875 1876For AMDGPU the lookup by name section header table: 1877 1878``augmentation_string_size`` (uword) 1879 1880 Set to the length of the ``augmentation_string`` value which is always a 1881 multiple of 4. 1882 1883``augmentation_string`` (sequence of UTF-8 characters) 1884 1885 Contains the following UTF-8 string null padded to a multiple of 4 bytes: 1886 1887 :: 1888 1889 [amdgpu:v0.0] 1890 1891 The "vX.Y" specifies the major X and minor Y version number of the AMDGPU 1892 extensions used in the DWARF of this index. The version number conforms to 1893 [SEMVER]_. 1894 1895 .. note:: 1896 1897 This is different to the DWARF Version 5 definition that requires the first 1898 4 characters to be the vendor ID. But this is consistent with the other 1899 augmentation strings and does allow multiple vendor contributions. However, 1900 backwards compatibility may be more desirable. 1901 1902Lookup By Address Section Header 1903~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 1904 1905See DWARF Version 5 section 6.1.2. 1906 1907For AMDGPU the lookup by address section header table: 1908 1909``address_size`` (ubyte) 1910 1911 Match the address size for the ``Global`` address space defined in 1912 :ref:`amdgpu-dwarf-address-space-identifier`. 1913 1914``segment_selector_size`` (ubyte) 1915 1916 AMDGPU does not use a segment selector so this is 0. The entries in the 1917 ``.debug_aranges`` do not have a segment selector. 1918 1919Line Number Information 1920----------------------- 1921 1922See DWARF Version 5 section 6.2 and :ref:`amdgpu-dwarf-line-number-information`. 1923 1924AMDGPU does not use the ``isa`` state machine registers and always sets it to 0. 1925The instruction set must be obtained from the ELF file header ``e_flags`` field 1926in the ``EF_AMDGPU_MACH`` bit position (see :ref:`ELF Header 1927<amdgpu-elf-header>`). See DWARF Version 5 section 6.2.2. 1928 1929.. TODO:: 1930 1931 Should the ``isa`` state machine register be used to indicate if the code is 1932 in wavefront32 or wavefront64 mode? Or used to specify the architecture ISA? 1933 1934For AMDGPU the line number program header fields have the following values (see 1935DWARF Version 5 section 6.2.4): 1936 1937``address_size`` (ubyte) 1938 Matches the address size for the ``Global`` address space defined in 1939 :ref:`amdgpu-dwarf-address-space-identifier`. 1940 1941``segment_selector_size`` (ubyte) 1942 AMDGPU does not use a segment selector so this is 0. 1943 1944``minimum_instruction_length`` (ubyte) 1945 For GFX9-GFX10 this is 4. 1946 1947``maximum_operations_per_instruction`` (ubyte) 1948 For GFX9-GFX10 this is 1. 1949 1950Source text for online-compiled programs (for example, those compiled by the 1951OpenCL language runtime) may be embedded into the DWARF Version 5 line table. 1952See DWARF Version 5 section 6.2.4.1 which is updated by *DWARF Extensions For 1953Heterogeneous Debugging* section :ref:`DW_LNCT_LLVM_source 1954<amdgpu-dwarf-line-number-information-dw-lnct-llvm-source>`. 1955 1956The Clang option used to control source embedding in AMDGPU is defined in 1957:ref:`amdgpu-clang-debug-options-table`. 1958 1959 .. table:: AMDGPU Clang Debug Options 1960 :name: amdgpu-clang-debug-options-table 1961 1962 ==================== ================================================== 1963 Debug Flag Description 1964 ==================== ================================================== 1965 -g[no-]embed-source Enable/disable embedding source text in DWARF 1966 debug sections. Useful for environments where 1967 source cannot be written to disk, such as 1968 when performing online compilation. 1969 ==================== ================================================== 1970 1971For example: 1972 1973``-gembed-source`` 1974 Enable the embedded source. 1975 1976``-gno-embed-source`` 1977 Disable the embedded source. 1978 197932-Bit and 64-Bit DWARF Formats 1980------------------------------- 1981 1982See DWARF Version 5 section 7.4 and 1983:ref:`amdgpu-dwarf-32-bit-and-64-bit-dwarf-formats`. 1984 1985For AMDGPU: 1986 1987* For the ``amdgcn`` target architecture only the 64-bit process address space 1988 is supported. 1989 1990* The producer can generate either 32-bit or 64-bit DWARF format. LLVM generates 1991 the 32-bit DWARF format. 1992 1993Unit Headers 1994------------ 1995 1996For AMDGPU the following values apply for each of the unit headers described in 1997DWARF Version 5 sections 7.5.1.1, 7.5.1.2, and 7.5.1.3: 1998 1999``address_size`` (ubyte) 2000 Matches the address size for the ``Global`` address space defined in 2001 :ref:`amdgpu-dwarf-address-space-identifier`. 2002 2003.. _amdgpu-code-conventions: 2004 2005Code Conventions 2006================ 2007 2008This section provides code conventions used for each supported target triple OS 2009(see :ref:`amdgpu-target-triples`). 2010 2011AMDHSA 2012------ 2013 2014This section provides code conventions used when the target triple OS is 2015``amdhsa`` (see :ref:`amdgpu-target-triples`). 2016 2017.. _amdgpu-amdhsa-code-object-target-identification: 2018 2019Code Object Target Identification 2020~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 2021 2022The AMDHSA OS uses the following syntax to specify the code object 2023target as a single string: 2024 2025 ``<Architecture>-<Vendor>-<OS>-<Environment>-<Processor><Target Features>`` 2026 2027Where: 2028 2029 - ``<Architecture>``, ``<Vendor>``, ``<OS>`` and ``<Environment>`` 2030 are the same as the *Target Triple* (see 2031 :ref:`amdgpu-target-triples`). 2032 2033 - ``<Processor>`` is the same as the *Processor* (see 2034 :ref:`amdgpu-processors`). 2035 2036 - ``<Target Features>`` is a list of the enabled *Target Features* 2037 (see :ref:`amdgpu-target-features`), each prefixed by a plus, that 2038 apply to *Processor*. The list must be in the same order as listed 2039 in the table :ref:`amdgpu-target-feature-table`. Note that *Target 2040 Features* must be included in the list if they are enabled even if 2041 that is the default for *Processor*. 2042 2043For example: 2044 2045 ``"amdgcn-amd-amdhsa--gfx902+xnack"`` 2046 2047.. _amdgpu-amdhsa-code-object-metadata: 2048 2049Code Object Metadata 2050~~~~~~~~~~~~~~~~~~~~ 2051 2052The code object metadata specifies extensible metadata associated with the code 2053objects executed on HSA [HSA]_ compatible runtimes such as AMD's ROCm 2054[AMD-ROCm]_. The encoding and semantics of this metadata depends on the code 2055object version; see :ref:`amdgpu-amdhsa-code-object-metadata-v2` and 2056:ref:`amdgpu-amdhsa-code-object-metadata-v3`. 2057 2058Code object metadata is specified in a note record (see 2059:ref:`amdgpu-note-records`) and is required when the target triple OS is 2060``amdhsa`` (see :ref:`amdgpu-target-triples`). It must contain the minimum 2061information necessary to support the ROCM kernel queries. For example, the 2062segment sizes needed in a dispatch packet. In addition, a high-level language 2063runtime may require other information to be included. For example, the AMD 2064OpenCL runtime records kernel argument information. 2065 2066.. _amdgpu-amdhsa-code-object-metadata-v2: 2067 2068Code Object V2 Metadata (-mattr=-code-object-v3) 2069++++++++++++++++++++++++++++++++++++++++++++++++ 2070 2071.. warning:: Code Object V2 is not the default code object version emitted by 2072 this version of LLVM. For a description of the metadata generated with the 2073 default configuration (Code Object V3) see 2074 :ref:`amdgpu-amdhsa-code-object-metadata-v3`. 2075 2076Code object V2 metadata is specified by the ``NT_AMD_AMDGPU_METADATA`` note 2077record (see :ref:`amdgpu-note-records-v2`). 2078 2079The metadata is specified as a YAML formatted string (see [YAML]_ and 2080:doc:`YamlIO`). 2081 2082.. TODO:: 2083 2084 Is the string null terminated? It probably should not if YAML allows it to 2085 contain null characters, otherwise it should be. 2086 2087The metadata is represented as a single YAML document comprised of the mapping 2088defined in table :ref:`amdgpu-amdhsa-code-object-metadata-map-table-v2` and 2089referenced tables. 2090 2091For boolean values, the string values of ``false`` and ``true`` are used for 2092false and true respectively. 2093 2094Additional information can be added to the mappings. To avoid conflicts, any 2095non-AMD key names should be prefixed by "*vendor-name*.". 2096 2097 .. table:: AMDHSA Code Object V2 Metadata Map 2098 :name: amdgpu-amdhsa-code-object-metadata-map-table-v2 2099 2100 ========== ============== ========= ======================================= 2101 String Key Value Type Required? Description 2102 ========== ============== ========= ======================================= 2103 "Version" sequence of Required - The first integer is the major 2104 2 integers version. Currently 1. 2105 - The second integer is the minor 2106 version. Currently 0. 2107 "Printf" sequence of Each string is encoded information 2108 strings about a printf function call. The 2109 encoded information is organized as 2110 fields separated by colon (':'): 2111 2112 ``ID:N:S[0]:S[1]:...:S[N-1]:FormatString`` 2113 2114 where: 2115 2116 ``ID`` 2117 A 32-bit integer as a unique id for 2118 each printf function call 2119 2120 ``N`` 2121 A 32-bit integer equal to the number 2122 of arguments of printf function call 2123 minus 1 2124 2125 ``S[i]`` (where i = 0, 1, ... , N-1) 2126 32-bit integers for the size in bytes 2127 of the i-th FormatString argument of 2128 the printf function call 2129 2130 FormatString 2131 The format string passed to the 2132 printf function call. 2133 "Kernels" sequence of Required Sequence of the mappings for each 2134 mapping kernel in the code object. See 2135 :ref:`amdgpu-amdhsa-code-object-kernel-metadata-map-table-v2` 2136 for the definition of the mapping. 2137 ========== ============== ========= ======================================= 2138 2139.. 2140 2141 .. table:: AMDHSA Code Object V2 Kernel Metadata Map 2142 :name: amdgpu-amdhsa-code-object-kernel-metadata-map-table-v2 2143 2144 ================= ============== ========= ================================ 2145 String Key Value Type Required? Description 2146 ================= ============== ========= ================================ 2147 "Name" string Required Source name of the kernel. 2148 "SymbolName" string Required Name of the kernel 2149 descriptor ELF symbol. 2150 "Language" string Source language of the kernel. 2151 Values include: 2152 2153 - "OpenCL C" 2154 - "OpenCL C++" 2155 - "HCC" 2156 - "OpenMP" 2157 2158 "LanguageVersion" sequence of - The first integer is the major 2159 2 integers version. 2160 - The second integer is the 2161 minor version. 2162 "Attrs" mapping Mapping of kernel attributes. 2163 See 2164 :ref:`amdgpu-amdhsa-code-object-kernel-attribute-metadata-map-table-v2` 2165 for the mapping definition. 2166 "Args" sequence of Sequence of mappings of the 2167 mapping kernel arguments. See 2168 :ref:`amdgpu-amdhsa-code-object-kernel-argument-metadata-map-table-v2` 2169 for the definition of the mapping. 2170 "CodeProps" mapping Mapping of properties related to 2171 the kernel code. See 2172 :ref:`amdgpu-amdhsa-code-object-kernel-code-properties-metadata-map-table-v2` 2173 for the mapping definition. 2174 ================= ============== ========= ================================ 2175 2176.. 2177 2178 .. table:: AMDHSA Code Object V2 Kernel Attribute Metadata Map 2179 :name: amdgpu-amdhsa-code-object-kernel-attribute-metadata-map-table-v2 2180 2181 =================== ============== ========= ============================== 2182 String Key Value Type Required? Description 2183 =================== ============== ========= ============================== 2184 "ReqdWorkGroupSize" sequence of If not 0, 0, 0 then all values 2185 3 integers must be >=1 and the dispatch 2186 work-group size X, Y, Z must 2187 correspond to the specified 2188 values. Defaults to 0, 0, 0. 2189 2190 Corresponds to the OpenCL 2191 ``reqd_work_group_size`` 2192 attribute. 2193 "WorkGroupSizeHint" sequence of The dispatch work-group size 2194 3 integers X, Y, Z is likely to be the 2195 specified values. 2196 2197 Corresponds to the OpenCL 2198 ``work_group_size_hint`` 2199 attribute. 2200 "VecTypeHint" string The name of a scalar or vector 2201 type. 2202 2203 Corresponds to the OpenCL 2204 ``vec_type_hint`` attribute. 2205 2206 "RuntimeHandle" string The external symbol name 2207 associated with a kernel. 2208 OpenCL runtime allocates a 2209 global buffer for the symbol 2210 and saves the kernel's address 2211 to it, which is used for 2212 device side enqueueing. Only 2213 available for device side 2214 enqueued kernels. 2215 =================== ============== ========= ============================== 2216 2217.. 2218 2219 .. table:: AMDHSA Code Object V2 Kernel Argument Metadata Map 2220 :name: amdgpu-amdhsa-code-object-kernel-argument-metadata-map-table-v2 2221 2222 ================= ============== ========= ================================ 2223 String Key Value Type Required? Description 2224 ================= ============== ========= ================================ 2225 "Name" string Kernel argument name. 2226 "TypeName" string Kernel argument type name. 2227 "Size" integer Required Kernel argument size in bytes. 2228 "Align" integer Required Kernel argument alignment in 2229 bytes. Must be a power of two. 2230 "ValueKind" string Required Kernel argument kind that 2231 specifies how to set up the 2232 corresponding argument. 2233 Values include: 2234 2235 "ByValue" 2236 The argument is copied 2237 directly into the kernarg. 2238 2239 "GlobalBuffer" 2240 A global address space pointer 2241 to the buffer data is passed 2242 in the kernarg. 2243 2244 "DynamicSharedPointer" 2245 A group address space pointer 2246 to dynamically allocated LDS 2247 is passed in the kernarg. 2248 2249 "Sampler" 2250 A global address space 2251 pointer to a S# is passed in 2252 the kernarg. 2253 2254 "Image" 2255 A global address space 2256 pointer to a T# is passed in 2257 the kernarg. 2258 2259 "Pipe" 2260 A global address space pointer 2261 to an OpenCL pipe is passed in 2262 the kernarg. 2263 2264 "Queue" 2265 A global address space pointer 2266 to an OpenCL device enqueue 2267 queue is passed in the 2268 kernarg. 2269 2270 "HiddenGlobalOffsetX" 2271 The OpenCL grid dispatch 2272 global offset for the X 2273 dimension is passed in the 2274 kernarg. 2275 2276 "HiddenGlobalOffsetY" 2277 The OpenCL grid dispatch 2278 global offset for the Y 2279 dimension is passed in the 2280 kernarg. 2281 2282 "HiddenGlobalOffsetZ" 2283 The OpenCL grid dispatch 2284 global offset for the Z 2285 dimension is passed in the 2286 kernarg. 2287 2288 "HiddenNone" 2289 An argument that is not used 2290 by the kernel. Space needs to 2291 be left for it, but it does 2292 not need to be set up. 2293 2294 "HiddenPrintfBuffer" 2295 A global address space pointer 2296 to the runtime printf buffer 2297 is passed in kernarg. 2298 2299 "HiddenHostcallBuffer" 2300 A global address space pointer 2301 to the runtime hostcall buffer 2302 is passed in kernarg. 2303 2304 "HiddenDefaultQueue" 2305 A global address space pointer 2306 to the OpenCL device enqueue 2307 queue that should be used by 2308 the kernel by default is 2309 passed in the kernarg. 2310 2311 "HiddenCompletionAction" 2312 A global address space pointer 2313 to help link enqueued kernels into 2314 the ancestor tree for determining 2315 when the parent kernel has finished. 2316 2317 "HiddenMultiGridSyncArg" 2318 A global address space pointer for 2319 multi-grid synchronization is 2320 passed in the kernarg. 2321 2322 "ValueType" string Unused and deprecated. This should no longer 2323 be emitted, but is accepted for compatibility. 2324 2325 2326 "PointeeAlign" integer Alignment in bytes of pointee 2327 type for pointer type kernel 2328 argument. Must be a power 2329 of 2. Only present if 2330 "ValueKind" is 2331 "DynamicSharedPointer". 2332 "AddrSpaceQual" string Kernel argument address space 2333 qualifier. Only present if 2334 "ValueKind" is "GlobalBuffer" or 2335 "DynamicSharedPointer". Values 2336 are: 2337 2338 - "Private" 2339 - "Global" 2340 - "Constant" 2341 - "Local" 2342 - "Generic" 2343 - "Region" 2344 2345 .. TODO:: 2346 Is GlobalBuffer only Global 2347 or Constant? Is 2348 DynamicSharedPointer always 2349 Local? Can HCC allow Generic? 2350 How can Private or Region 2351 ever happen? 2352 "AccQual" string Kernel argument access 2353 qualifier. Only present if 2354 "ValueKind" is "Image" or 2355 "Pipe". Values 2356 are: 2357 2358 - "ReadOnly" 2359 - "WriteOnly" 2360 - "ReadWrite" 2361 2362 .. TODO:: 2363 Does this apply to 2364 GlobalBuffer? 2365 "ActualAccQual" string The actual memory accesses 2366 performed by the kernel on the 2367 kernel argument. Only present if 2368 "ValueKind" is "GlobalBuffer", 2369 "Image", or "Pipe". This may be 2370 more restrictive than indicated 2371 by "AccQual" to reflect what the 2372 kernel actual does. If not 2373 present then the runtime must 2374 assume what is implied by 2375 "AccQual" and "IsConst". Values 2376 are: 2377 2378 - "ReadOnly" 2379 - "WriteOnly" 2380 - "ReadWrite" 2381 2382 "IsConst" boolean Indicates if the kernel argument 2383 is const qualified. Only present 2384 if "ValueKind" is 2385 "GlobalBuffer". 2386 2387 "IsRestrict" boolean Indicates if the kernel argument 2388 is restrict qualified. Only 2389 present if "ValueKind" is 2390 "GlobalBuffer". 2391 2392 "IsVolatile" boolean Indicates if the kernel argument 2393 is volatile qualified. Only 2394 present if "ValueKind" is 2395 "GlobalBuffer". 2396 2397 "IsPipe" boolean Indicates if the kernel argument 2398 is pipe qualified. Only present 2399 if "ValueKind" is "Pipe". 2400 2401 .. TODO:: 2402 Can GlobalBuffer be pipe 2403 qualified? 2404 ================= ============== ========= ================================ 2405 2406.. 2407 2408 .. table:: AMDHSA Code Object V2 Kernel Code Properties Metadata Map 2409 :name: amdgpu-amdhsa-code-object-kernel-code-properties-metadata-map-table-v2 2410 2411 ============================ ============== ========= ===================== 2412 String Key Value Type Required? Description 2413 ============================ ============== ========= ===================== 2414 "KernargSegmentSize" integer Required The size in bytes of 2415 the kernarg segment 2416 that holds the values 2417 of the arguments to 2418 the kernel. 2419 "GroupSegmentFixedSize" integer Required The amount of group 2420 segment memory 2421 required by a 2422 work-group in 2423 bytes. This does not 2424 include any 2425 dynamically allocated 2426 group segment memory 2427 that may be added 2428 when the kernel is 2429 dispatched. 2430 "PrivateSegmentFixedSize" integer Required The amount of fixed 2431 private address space 2432 memory required for a 2433 work-item in 2434 bytes. If the kernel 2435 uses a dynamic call 2436 stack then additional 2437 space must be added 2438 to this value for the 2439 call stack. 2440 "KernargSegmentAlign" integer Required The maximum byte 2441 alignment of 2442 arguments in the 2443 kernarg segment. Must 2444 be a power of 2. 2445 "WavefrontSize" integer Required Wavefront size. Must 2446 be a power of 2. 2447 "NumSGPRs" integer Required Number of scalar 2448 registers used by a 2449 wavefront for 2450 GFX6-GFX10. This 2451 includes the special 2452 SGPRs for VCC, Flat 2453 Scratch (GFX7-GFX10) 2454 and XNACK (for 2455 GFX8-GFX10). It does 2456 not include the 16 2457 SGPR added if a trap 2458 handler is 2459 enabled. It is not 2460 rounded up to the 2461 allocation 2462 granularity. 2463 "NumVGPRs" integer Required Number of vector 2464 registers used by 2465 each work-item for 2466 GFX6-GFX10 2467 "MaxFlatWorkGroupSize" integer Required Maximum flat 2468 work-group size 2469 supported by the 2470 kernel in work-items. 2471 Must be >=1 and 2472 consistent with 2473 ReqdWorkGroupSize if 2474 not 0, 0, 0. 2475 "NumSpilledSGPRs" integer Number of stores from 2476 a scalar register to 2477 a register allocator 2478 created spill 2479 location. 2480 "NumSpilledVGPRs" integer Number of stores from 2481 a vector register to 2482 a register allocator 2483 created spill 2484 location. 2485 ============================ ============== ========= ===================== 2486 2487.. _amdgpu-amdhsa-code-object-metadata-v3: 2488 2489Code Object V3 Metadata (-mattr=+code-object-v3) 2490++++++++++++++++++++++++++++++++++++++++++++++++ 2491 2492Code object V3 metadata is specified by the ``NT_AMDGPU_METADATA`` note record 2493(see :ref:`amdgpu-note-records-v3`). 2494 2495The metadata is represented as Message Pack formatted binary data (see 2496[MsgPack]_). The top level is a Message Pack map that includes the 2497keys defined in table 2498:ref:`amdgpu-amdhsa-code-object-metadata-map-table-v3` and referenced 2499tables. 2500 2501Additional information can be added to the maps. To avoid conflicts, 2502any key names should be prefixed by "*vendor-name*." where 2503``vendor-name`` can be the name of the vendor and specific vendor 2504tool that generates the information. The prefix is abbreviated to 2505simply "." when it appears within a map that has been added by the 2506same *vendor-name*. 2507 2508 .. table:: AMDHSA Code Object V3 Metadata Map 2509 :name: amdgpu-amdhsa-code-object-metadata-map-table-v3 2510 2511 ================= ============== ========= ======================================= 2512 String Key Value Type Required? Description 2513 ================= ============== ========= ======================================= 2514 "amdhsa.version" sequence of Required - The first integer is the major 2515 2 integers version. Currently 1. 2516 - The second integer is the minor 2517 version. Currently 0. 2518 "amdhsa.printf" sequence of Each string is encoded information 2519 strings about a printf function call. The 2520 encoded information is organized as 2521 fields separated by colon (':'): 2522 2523 ``ID:N:S[0]:S[1]:...:S[N-1]:FormatString`` 2524 2525 where: 2526 2527 ``ID`` 2528 A 32-bit integer as a unique id for 2529 each printf function call 2530 2531 ``N`` 2532 A 32-bit integer equal to the number 2533 of arguments of printf function call 2534 minus 1 2535 2536 ``S[i]`` (where i = 0, 1, ... , N-1) 2537 32-bit integers for the size in bytes 2538 of the i-th FormatString argument of 2539 the printf function call 2540 2541 FormatString 2542 The format string passed to the 2543 printf function call. 2544 "amdhsa.kernels" sequence of Required Sequence of the maps for each 2545 map kernel in the code object. See 2546 :ref:`amdgpu-amdhsa-code-object-kernel-metadata-map-table-v3` 2547 for the definition of the keys included 2548 in that map. 2549 ================= ============== ========= ======================================= 2550 2551.. 2552 2553 .. table:: AMDHSA Code Object V3 Kernel Metadata Map 2554 :name: amdgpu-amdhsa-code-object-kernel-metadata-map-table-v3 2555 2556 =================================== ============== ========= ================================ 2557 String Key Value Type Required? Description 2558 =================================== ============== ========= ================================ 2559 ".name" string Required Source name of the kernel. 2560 ".symbol" string Required Name of the kernel 2561 descriptor ELF symbol. 2562 ".language" string Source language of the kernel. 2563 Values include: 2564 2565 - "OpenCL C" 2566 - "OpenCL C++" 2567 - "HCC" 2568 - "HIP" 2569 - "OpenMP" 2570 - "Assembler" 2571 2572 ".language_version" sequence of - The first integer is the major 2573 2 integers version. 2574 - The second integer is the 2575 minor version. 2576 ".args" sequence of Sequence of maps of the 2577 map kernel arguments. See 2578 :ref:`amdgpu-amdhsa-code-object-kernel-argument-metadata-map-table-v3` 2579 for the definition of the keys 2580 included in that map. 2581 ".reqd_workgroup_size" sequence of If not 0, 0, 0 then all values 2582 3 integers must be >=1 and the dispatch 2583 work-group size X, Y, Z must 2584 correspond to the specified 2585 values. Defaults to 0, 0, 0. 2586 2587 Corresponds to the OpenCL 2588 ``reqd_work_group_size`` 2589 attribute. 2590 ".workgroup_size_hint" sequence of The dispatch work-group size 2591 3 integers X, Y, Z is likely to be the 2592 specified values. 2593 2594 Corresponds to the OpenCL 2595 ``work_group_size_hint`` 2596 attribute. 2597 ".vec_type_hint" string The name of a scalar or vector 2598 type. 2599 2600 Corresponds to the OpenCL 2601 ``vec_type_hint`` attribute. 2602 2603 ".device_enqueue_symbol" string The external symbol name 2604 associated with a kernel. 2605 OpenCL runtime allocates a 2606 global buffer for the symbol 2607 and saves the kernel's address 2608 to it, which is used for 2609 device side enqueueing. Only 2610 available for device side 2611 enqueued kernels. 2612 ".kernarg_segment_size" integer Required The size in bytes of 2613 the kernarg segment 2614 that holds the values 2615 of the arguments to 2616 the kernel. 2617 ".group_segment_fixed_size" integer Required The amount of group 2618 segment memory 2619 required by a 2620 work-group in 2621 bytes. This does not 2622 include any 2623 dynamically allocated 2624 group segment memory 2625 that may be added 2626 when the kernel is 2627 dispatched. 2628 ".private_segment_fixed_size" integer Required The amount of fixed 2629 private address space 2630 memory required for a 2631 work-item in 2632 bytes. If the kernel 2633 uses a dynamic call 2634 stack then additional 2635 space must be added 2636 to this value for the 2637 call stack. 2638 ".kernarg_segment_align" integer Required The maximum byte 2639 alignment of 2640 arguments in the 2641 kernarg segment. Must 2642 be a power of 2. 2643 ".wavefront_size" integer Required Wavefront size. Must 2644 be a power of 2. 2645 ".sgpr_count" integer Required Number of scalar 2646 registers required by a 2647 wavefront for 2648 GFX6-GFX9. A register 2649 is required if it is 2650 used explicitly, or 2651 if a higher numbered 2652 register is used 2653 explicitly. This 2654 includes the special 2655 SGPRs for VCC, Flat 2656 Scratch (GFX7-GFX9) 2657 and XNACK (for 2658 GFX8-GFX9). It does 2659 not include the 16 2660 SGPR added if a trap 2661 handler is 2662 enabled. It is not 2663 rounded up to the 2664 allocation 2665 granularity. 2666 ".vgpr_count" integer Required Number of vector 2667 registers required by 2668 each work-item for 2669 GFX6-GFX9. A register 2670 is required if it is 2671 used explicitly, or 2672 if a higher numbered 2673 register is used 2674 explicitly. 2675 ".max_flat_workgroup_size" integer Required Maximum flat 2676 work-group size 2677 supported by the 2678 kernel in work-items. 2679 Must be >=1 and 2680 consistent with 2681 ReqdWorkGroupSize if 2682 not 0, 0, 0. 2683 ".sgpr_spill_count" integer Number of stores from 2684 a scalar register to 2685 a register allocator 2686 created spill 2687 location. 2688 ".vgpr_spill_count" integer Number of stores from 2689 a vector register to 2690 a register allocator 2691 created spill 2692 location. 2693 =================================== ============== ========= ================================ 2694 2695.. 2696 2697 .. table:: AMDHSA Code Object V3 Kernel Argument Metadata Map 2698 :name: amdgpu-amdhsa-code-object-kernel-argument-metadata-map-table-v3 2699 2700 ====================== ============== ========= ================================ 2701 String Key Value Type Required? Description 2702 ====================== ============== ========= ================================ 2703 ".name" string Kernel argument name. 2704 ".type_name" string Kernel argument type name. 2705 ".size" integer Required Kernel argument size in bytes. 2706 ".offset" integer Required Kernel argument offset in 2707 bytes. The offset must be a 2708 multiple of the alignment 2709 required by the argument. 2710 ".value_kind" string Required Kernel argument kind that 2711 specifies how to set up the 2712 corresponding argument. 2713 Values include: 2714 2715 "by_value" 2716 The argument is copied 2717 directly into the kernarg. 2718 2719 "global_buffer" 2720 A global address space pointer 2721 to the buffer data is passed 2722 in the kernarg. 2723 2724 "dynamic_shared_pointer" 2725 A group address space pointer 2726 to dynamically allocated LDS 2727 is passed in the kernarg. 2728 2729 "sampler" 2730 A global address space 2731 pointer to a S# is passed in 2732 the kernarg. 2733 2734 "image" 2735 A global address space 2736 pointer to a T# is passed in 2737 the kernarg. 2738 2739 "pipe" 2740 A global address space pointer 2741 to an OpenCL pipe is passed in 2742 the kernarg. 2743 2744 "queue" 2745 A global address space pointer 2746 to an OpenCL device enqueue 2747 queue is passed in the 2748 kernarg. 2749 2750 "hidden_global_offset_x" 2751 The OpenCL grid dispatch 2752 global offset for the X 2753 dimension is passed in the 2754 kernarg. 2755 2756 "hidden_global_offset_y" 2757 The OpenCL grid dispatch 2758 global offset for the Y 2759 dimension is passed in the 2760 kernarg. 2761 2762 "hidden_global_offset_z" 2763 The OpenCL grid dispatch 2764 global offset for the Z 2765 dimension is passed in the 2766 kernarg. 2767 2768 "hidden_none" 2769 An argument that is not used 2770 by the kernel. Space needs to 2771 be left for it, but it does 2772 not need to be set up. 2773 2774 "hidden_printf_buffer" 2775 A global address space pointer 2776 to the runtime printf buffer 2777 is passed in kernarg. 2778 2779 "hidden_hostcall_buffer" 2780 A global address space pointer 2781 to the runtime hostcall buffer 2782 is passed in kernarg. 2783 2784 "hidden_default_queue" 2785 A global address space pointer 2786 to the OpenCL device enqueue 2787 queue that should be used by 2788 the kernel by default is 2789 passed in the kernarg. 2790 2791 "hidden_completion_action" 2792 A global address space pointer 2793 to help link enqueued kernels into 2794 the ancestor tree for determining 2795 when the parent kernel has finished. 2796 2797 "hidden_multigrid_sync_arg" 2798 A global address space pointer for 2799 multi-grid synchronization is 2800 passed in the kernarg. 2801 2802 ".value_type" string Unused and deprecated. This should no longer 2803 be emitted, but is accepted for compatibility. 2804 2805 ".pointee_align" integer Alignment in bytes of pointee 2806 type for pointer type kernel 2807 argument. Must be a power 2808 of 2. Only present if 2809 ".value_kind" is 2810 "dynamic_shared_pointer". 2811 ".address_space" string Kernel argument address space 2812 qualifier. Only present if 2813 ".value_kind" is "global_buffer" or 2814 "dynamic_shared_pointer". Values 2815 are: 2816 2817 - "private" 2818 - "global" 2819 - "constant" 2820 - "local" 2821 - "generic" 2822 - "region" 2823 2824 .. TODO:: 2825 Is "global_buffer" only "global" 2826 or "constant"? Is 2827 "dynamic_shared_pointer" always 2828 "local"? Can HCC allow "generic"? 2829 How can "private" or "region" 2830 ever happen? 2831 ".access" string Kernel argument access 2832 qualifier. Only present if 2833 ".value_kind" is "image" or 2834 "pipe". Values 2835 are: 2836 2837 - "read_only" 2838 - "write_only" 2839 - "read_write" 2840 2841 .. TODO:: 2842 Does this apply to 2843 "global_buffer"? 2844 ".actual_access" string The actual memory accesses 2845 performed by the kernel on the 2846 kernel argument. Only present if 2847 ".value_kind" is "global_buffer", 2848 "image", or "pipe". This may be 2849 more restrictive than indicated 2850 by ".access" to reflect what the 2851 kernel actual does. If not 2852 present then the runtime must 2853 assume what is implied by 2854 ".access" and ".is_const" . Values 2855 are: 2856 2857 - "read_only" 2858 - "write_only" 2859 - "read_write" 2860 2861 ".is_const" boolean Indicates if the kernel argument 2862 is const qualified. Only present 2863 if ".value_kind" is 2864 "global_buffer". 2865 2866 ".is_restrict" boolean Indicates if the kernel argument 2867 is restrict qualified. Only 2868 present if ".value_kind" is 2869 "global_buffer". 2870 2871 ".is_volatile" boolean Indicates if the kernel argument 2872 is volatile qualified. Only 2873 present if ".value_kind" is 2874 "global_buffer". 2875 2876 ".is_pipe" boolean Indicates if the kernel argument 2877 is pipe qualified. Only present 2878 if ".value_kind" is "pipe". 2879 2880 .. TODO:: 2881 Can "global_buffer" be pipe 2882 qualified? 2883 ====================== ============== ========= ================================ 2884 2885.. 2886 2887Kernel Dispatch 2888~~~~~~~~~~~~~~~ 2889 2890The HSA architected queuing language (AQL) defines a user space memory 2891interface that can be used to control the dispatch of kernels, in an agent 2892independent way. An agent can have zero or more AQL queues created for it using 2893the ROCm runtime, in which AQL packets (all of which are 64 bytes) can be 2894placed. See the *HSA Platform System Architecture Specification* [HSA]_ for the 2895AQL queue mechanics and packet layouts. 2896 2897The packet processor of a kernel agent is responsible for detecting and 2898dispatching HSA kernels from the AQL queues associated with it. For AMD GPUs the 2899packet processor is implemented by the hardware command processor (CP), 2900asynchronous dispatch controller (ADC) and shader processor input controller 2901(SPI). 2902 2903The ROCm runtime can be used to allocate an AQL queue object. It uses the kernel 2904mode driver to initialize and register the AQL queue with CP. 2905 2906To dispatch a kernel the following actions are performed. This can occur in the 2907CPU host program, or from an HSA kernel executing on a GPU. 2908 29091. A pointer to an AQL queue for the kernel agent on which the kernel is to be 2910 executed is obtained. 29112. A pointer to the kernel descriptor (see 2912 :ref:`amdgpu-amdhsa-kernel-descriptor`) of the kernel to execute is obtained. 2913 It must be for a kernel that is contained in a code object that that was 2914 loaded by the ROCm runtime on the kernel agent with which the AQL queue is 2915 associated. 29163. Space is allocated for the kernel arguments using the ROCm runtime allocator 2917 for a memory region with the kernarg property for the kernel agent that will 2918 execute the kernel. It must be at least 16-byte aligned. 29194. Kernel argument values are assigned to the kernel argument memory 2920 allocation. The layout is defined in the *HSA Programmer's Language 2921 Reference* [HSA]_. For AMDGPU the kernel execution directly accesses the 2922 kernel argument memory in the same way constant memory is accessed. (Note 2923 that the HSA specification allows an implementation to copy the kernel 2924 argument contents to another location that is accessed by the kernel.) 29255. An AQL kernel dispatch packet is created on the AQL queue. The ROCm runtime 2926 api uses 64-bit atomic operations to reserve space in the AQL queue for the 2927 packet. The packet must be set up, and the final write must use an atomic 2928 store release to set the packet kind to ensure the packet contents are 2929 visible to the kernel agent. AQL defines a doorbell signal mechanism to 2930 notify the kernel agent that the AQL queue has been updated. These rules, and 2931 the layout of the AQL queue and kernel dispatch packet is defined in the *HSA 2932 System Architecture Specification* [HSA]_. 29336. A kernel dispatch packet includes information about the actual dispatch, 2934 such as grid and work-group size, together with information from the code 2935 object about the kernel, such as segment sizes. The ROCm runtime queries on 2936 the kernel symbol can be used to obtain the code object values which are 2937 recorded in the :ref:`amdgpu-amdhsa-code-object-metadata`. 29387. CP executes micro-code and is responsible for detecting and setting up the 2939 GPU to execute the wavefronts of a kernel dispatch. 29408. CP ensures that when the a wavefront starts executing the kernel machine 2941 code, the scalar general purpose registers (SGPR) and vector general purpose 2942 registers (VGPR) are set up as required by the machine code. The required 2943 setup is defined in the :ref:`amdgpu-amdhsa-kernel-descriptor`. The initial 2944 register state is defined in 2945 :ref:`amdgpu-amdhsa-initial-kernel-execution-state`. 29469. The prolog of the kernel machine code (see 2947 :ref:`amdgpu-amdhsa-kernel-prolog`) sets up the machine state as necessary 2948 before continuing executing the machine code that corresponds to the kernel. 294910. When the kernel dispatch has completed execution, CP signals the completion 2950 signal specified in the kernel dispatch packet if not 0. 2951 2952Image and Samplers 2953~~~~~~~~~~~~~~~~~~ 2954 2955Image and sample handles created by the ROCm runtime are 64-bit addresses of a 2956hardware 32-byte V# and 48 byte S# object respectively. In order to support the 2957HSA ``query_sampler`` operations two extra dwords are used to store the HSA BRIG 2958enumeration values for the queries that are not trivially deducible from the S# 2959representation. 2960 2961HSA Signals 2962~~~~~~~~~~~ 2963 2964HSA signal handles created by the ROCm runtime are 64-bit addresses of a 2965structure allocated in memory accessible from both the CPU and GPU. The 2966structure is defined by the ROCm runtime and subject to change between releases 2967(see [AMD-ROCm-github]_). 2968 2969.. _amdgpu-amdhsa-hsa-aql-queue: 2970 2971HSA AQL Queue 2972~~~~~~~~~~~~~ 2973 2974The HSA AQL queue structure is defined by the ROCm runtime and subject to change 2975between releases (see [AMD-ROCm-github]_). For some processors it contains 2976fields needed to implement certain language features such as the flat address 2977aperture bases. It also contains fields used by CP such as managing the 2978allocation of scratch memory. 2979 2980.. _amdgpu-amdhsa-kernel-descriptor: 2981 2982Kernel Descriptor 2983~~~~~~~~~~~~~~~~~ 2984 2985A kernel descriptor consists of the information needed by CP to initiate the 2986execution of a kernel, including the entry point address of the machine code 2987that implements the kernel. 2988 2989Kernel Descriptor for GFX6-GFX10 2990++++++++++++++++++++++++++++++++ 2991 2992CP microcode requires the Kernel descriptor to be allocated on 64-byte 2993alignment. 2994 2995 .. table:: Kernel Descriptor for GFX6-GFX10 2996 :name: amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table 2997 2998 ======= ======= =============================== ============================ 2999 Bits Size Field Name Description 3000 ======= ======= =============================== ============================ 3001 31:0 4 bytes GROUP_SEGMENT_FIXED_SIZE The amount of fixed local 3002 address space memory 3003 required for a work-group 3004 in bytes. This does not 3005 include any dynamically 3006 allocated local address 3007 space memory that may be 3008 added when the kernel is 3009 dispatched. 3010 63:32 4 bytes PRIVATE_SEGMENT_FIXED_SIZE The amount of fixed 3011 private address space 3012 memory required for a 3013 work-item in bytes. If 3014 is_dynamic_callstack is 1 3015 then additional space must 3016 be added to this value for 3017 the call stack. 3018 127:64 8 bytes Reserved, must be 0. 3019 191:128 8 bytes KERNEL_CODE_ENTRY_BYTE_OFFSET Byte offset (possibly 3020 negative) from base 3021 address of kernel 3022 descriptor to kernel's 3023 entry point instruction 3024 which must be 256 byte 3025 aligned. 3026 351:272 20 Reserved, must be 0. 3027 bytes 3028 383:352 4 bytes COMPUTE_PGM_RSRC3 GFX6-9 3029 Reserved, must be 0. 3030 GFX10 3031 Compute Shader (CS) 3032 program settings used by 3033 CP to set up 3034 ``COMPUTE_PGM_RSRC3`` 3035 configuration 3036 register. See 3037 :ref:`amdgpu-amdhsa-compute_pgm_rsrc3-gfx10-table`. 3038 415:384 4 bytes COMPUTE_PGM_RSRC1 Compute Shader (CS) 3039 program settings used by 3040 CP to set up 3041 ``COMPUTE_PGM_RSRC1`` 3042 configuration 3043 register. See 3044 :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`. 3045 447:416 4 bytes COMPUTE_PGM_RSRC2 Compute Shader (CS) 3046 program settings used by 3047 CP to set up 3048 ``COMPUTE_PGM_RSRC2`` 3049 configuration 3050 register. See 3051 :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`. 3052 448 1 bit ENABLE_SGPR_PRIVATE_SEGMENT Enable the setup of the 3053 _BUFFER SGPR user data registers 3054 (see 3055 :ref:`amdgpu-amdhsa-initial-kernel-execution-state`). 3056 3057 The total number of SGPR 3058 user data registers 3059 requested must not exceed 3060 16 and match value in 3061 ``compute_pgm_rsrc2.user_sgpr.user_sgpr_count``. 3062 Any requests beyond 16 3063 will be ignored. 3064 449 1 bit ENABLE_SGPR_DISPATCH_PTR *see above* 3065 450 1 bit ENABLE_SGPR_QUEUE_PTR *see above* 3066 451 1 bit ENABLE_SGPR_KERNARG_SEGMENT_PTR *see above* 3067 452 1 bit ENABLE_SGPR_DISPATCH_ID *see above* 3068 453 1 bit ENABLE_SGPR_FLAT_SCRATCH_INIT *see above* 3069 454 1 bit ENABLE_SGPR_PRIVATE_SEGMENT *see above* 3070 _SIZE 3071 457:455 3 bits Reserved, must be 0. 3072 458 1 bit ENABLE_WAVEFRONT_SIZE32 GFX6-9 3073 Reserved, must be 0. 3074 GFX10 3075 - If 0 execute in 3076 wavefront size 64 mode. 3077 - If 1 execute in 3078 native wavefront size 3079 32 mode. 3080 463:459 5 bits Reserved, must be 0. 3081 511:464 6 bytes Reserved, must be 0. 3082 512 **Total size 64 bytes.** 3083 ======= ==================================================================== 3084 3085.. 3086 3087 .. table:: compute_pgm_rsrc1 for GFX6-GFX10 3088 :name: amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table 3089 3090 ======= ======= =============================== =========================================================================== 3091 Bits Size Field Name Description 3092 ======= ======= =============================== =========================================================================== 3093 5:0 6 bits GRANULATED_WORKITEM_VGPR_COUNT Number of vector register 3094 blocks used by each work-item; 3095 granularity is device 3096 specific: 3097 3098 GFX6-GFX9 3099 - vgprs_used 0..256 3100 - max(0, ceil(vgprs_used / 4) - 1) 3101 GFX10 (wavefront size 64) 3102 - max_vgpr 1..256 3103 - max(0, ceil(vgprs_used / 4) - 1) 3104 GFX10 (wavefront size 32) 3105 - max_vgpr 1..256 3106 - max(0, ceil(vgprs_used / 8) - 1) 3107 3108 Where vgprs_used is defined 3109 as the highest VGPR number 3110 explicitly referenced plus 3111 one. 3112 3113 Used by CP to set up 3114 ``COMPUTE_PGM_RSRC1.VGPRS``. 3115 3116 The 3117 :ref:`amdgpu-assembler` 3118 calculates this 3119 automatically for the 3120 selected processor from 3121 values provided to the 3122 `.amdhsa_kernel` directive 3123 by the 3124 `.amdhsa_next_free_vgpr` 3125 nested directive (see 3126 :ref:`amdhsa-kernel-directives-table`). 3127 9:6 4 bits GRANULATED_WAVEFRONT_SGPR_COUNT Number of scalar register 3128 blocks used by a wavefront; 3129 granularity is device 3130 specific: 3131 3132 GFX6-GFX8 3133 - sgprs_used 0..112 3134 - max(0, ceil(sgprs_used / 8) - 1) 3135 GFX9 3136 - sgprs_used 0..112 3137 - 2 * max(0, ceil(sgprs_used / 16) - 1) 3138 GFX10 3139 Reserved, must be 0. 3140 (128 SGPRs always 3141 allocated.) 3142 3143 Where sgprs_used is 3144 defined as the highest 3145 SGPR number explicitly 3146 referenced plus one, plus 3147 a target specific number 3148 of additional special 3149 SGPRs for VCC, 3150 FLAT_SCRATCH (GFX7+) and 3151 XNACK_MASK (GFX8+), and 3152 any additional 3153 target specific 3154 limitations. It does not 3155 include the 16 SGPRs added 3156 if a trap handler is 3157 enabled. 3158 3159 The target specific 3160 limitations and special 3161 SGPR layout are defined in 3162 the hardware 3163 documentation, which can 3164 be found in the 3165 :ref:`amdgpu-processors` 3166 table. 3167 3168 Used by CP to set up 3169 ``COMPUTE_PGM_RSRC1.SGPRS``. 3170 3171 The 3172 :ref:`amdgpu-assembler` 3173 calculates this 3174 automatically for the 3175 selected processor from 3176 values provided to the 3177 `.amdhsa_kernel` directive 3178 by the 3179 `.amdhsa_next_free_sgpr` 3180 and `.amdhsa_reserve_*` 3181 nested directives (see 3182 :ref:`amdhsa-kernel-directives-table`). 3183 11:10 2 bits PRIORITY Must be 0. 3184 3185 Start executing wavefront 3186 at the specified priority. 3187 3188 CP is responsible for 3189 filling in 3190 ``COMPUTE_PGM_RSRC1.PRIORITY``. 3191 13:12 2 bits FLOAT_ROUND_MODE_32 Wavefront starts execution 3192 with specified rounding 3193 mode for single (32 3194 bit) floating point 3195 precision floating point 3196 operations. 3197 3198 Floating point rounding 3199 mode values are defined in 3200 :ref:`amdgpu-amdhsa-floating-point-rounding-mode-enumeration-values-table`. 3201 3202 Used by CP to set up 3203 ``COMPUTE_PGM_RSRC1.FLOAT_MODE``. 3204 15:14 2 bits FLOAT_ROUND_MODE_16_64 Wavefront starts execution 3205 with specified rounding 3206 denorm mode for half/double (16 3207 and 64-bit) floating point 3208 precision floating point 3209 operations. 3210 3211 Floating point rounding 3212 mode values are defined in 3213 :ref:`amdgpu-amdhsa-floating-point-rounding-mode-enumeration-values-table`. 3214 3215 Used by CP to set up 3216 ``COMPUTE_PGM_RSRC1.FLOAT_MODE``. 3217 17:16 2 bits FLOAT_DENORM_MODE_32 Wavefront starts execution 3218 with specified denorm mode 3219 for single (32 3220 bit) floating point 3221 precision floating point 3222 operations. 3223 3224 Floating point denorm mode 3225 values are defined in 3226 :ref:`amdgpu-amdhsa-floating-point-denorm-mode-enumeration-values-table`. 3227 3228 Used by CP to set up 3229 ``COMPUTE_PGM_RSRC1.FLOAT_MODE``. 3230 19:18 2 bits FLOAT_DENORM_MODE_16_64 Wavefront starts execution 3231 with specified denorm mode 3232 for half/double (16 3233 and 64-bit) floating point 3234 precision floating point 3235 operations. 3236 3237 Floating point denorm mode 3238 values are defined in 3239 :ref:`amdgpu-amdhsa-floating-point-denorm-mode-enumeration-values-table`. 3240 3241 Used by CP to set up 3242 ``COMPUTE_PGM_RSRC1.FLOAT_MODE``. 3243 20 1 bit PRIV Must be 0. 3244 3245 Start executing wavefront 3246 in privilege trap handler 3247 mode. 3248 3249 CP is responsible for 3250 filling in 3251 ``COMPUTE_PGM_RSRC1.PRIV``. 3252 21 1 bit ENABLE_DX10_CLAMP Wavefront starts execution 3253 with DX10 clamp mode 3254 enabled. Used by the vector 3255 ALU to force DX10 style 3256 treatment of NaN's (when 3257 set, clamp NaN to zero, 3258 otherwise pass NaN 3259 through). 3260 3261 Used by CP to set up 3262 ``COMPUTE_PGM_RSRC1.DX10_CLAMP``. 3263 22 1 bit DEBUG_MODE Must be 0. 3264 3265 Start executing wavefront 3266 in single step mode. 3267 3268 CP is responsible for 3269 filling in 3270 ``COMPUTE_PGM_RSRC1.DEBUG_MODE``. 3271 23 1 bit ENABLE_IEEE_MODE Wavefront starts execution 3272 with IEEE mode 3273 enabled. Floating point 3274 opcodes that support 3275 exception flag gathering 3276 will quiet and propagate 3277 signaling-NaN inputs per 3278 IEEE 754-2008. Min_dx10 and 3279 max_dx10 become IEEE 3280 754-2008 compliant due to 3281 signaling-NaN propagation 3282 and quieting. 3283 3284 Used by CP to set up 3285 ``COMPUTE_PGM_RSRC1.IEEE_MODE``. 3286 24 1 bit BULKY Must be 0. 3287 3288 Only one work-group allowed 3289 to execute on a compute 3290 unit. 3291 3292 CP is responsible for 3293 filling in 3294 ``COMPUTE_PGM_RSRC1.BULKY``. 3295 25 1 bit CDBG_USER Must be 0. 3296 3297 Flag that can be used to 3298 control debugging code. 3299 3300 CP is responsible for 3301 filling in 3302 ``COMPUTE_PGM_RSRC1.CDBG_USER``. 3303 26 1 bit FP16_OVFL GFX6-GFX8 3304 Reserved, must be 0. 3305 GFX9-GFX10 3306 Wavefront starts execution 3307 with specified fp16 overflow 3308 mode. 3309 3310 - If 0, fp16 overflow generates 3311 +/-INF values. 3312 - If 1, fp16 overflow that is the 3313 result of an +/-INF input value 3314 or divide by 0 produces a +/-INF, 3315 otherwise clamps computed 3316 overflow to +/-MAX_FP16 as 3317 appropriate. 3318 3319 Used by CP to set up 3320 ``COMPUTE_PGM_RSRC1.FP16_OVFL``. 3321 28:27 2 bits Reserved, must be 0. 3322 29 1 bit WGP_MODE GFX6-GFX9 3323 Reserved, must be 0. 3324 GFX10 3325 - If 0 execute work-groups in 3326 CU wavefront execution mode. 3327 - If 1 execute work-groups on 3328 in WGP wavefront execution mode. 3329 3330 See :ref:`amdgpu-amdhsa-memory-model`. 3331 3332 Used by CP to set up 3333 ``COMPUTE_PGM_RSRC1.WGP_MODE``. 3334 30 1 bit MEM_ORDERED GFX6-9 3335 Reserved, must be 0. 3336 GFX10 3337 Controls the behavior of the 3338 waitcnt's vmcnt and vscnt 3339 counters. 3340 3341 - If 0 vmcnt reports completion 3342 of load and atomic with return 3343 out of order with sample 3344 instructions, and the vscnt 3345 reports the completion of 3346 store and atomic without 3347 return in order. 3348 - If 1 vmcnt reports completion 3349 of load, atomic with return 3350 and sample instructions in 3351 order, and the vscnt reports 3352 the completion of store and 3353 atomic without return in order. 3354 3355 Used by CP to set up 3356 ``COMPUTE_PGM_RSRC1.MEM_ORDERED``. 3357 31 1 bit FWD_PROGRESS GFX6-9 3358 Reserved, must be 0. 3359 GFX10 3360 - If 0 execute SIMD wavefronts 3361 using oldest first policy. 3362 - If 1 execute SIMD wavefronts to 3363 ensure wavefronts will make some 3364 forward progress. 3365 3366 Used by CP to set up 3367 ``COMPUTE_PGM_RSRC1.FWD_PROGRESS``. 3368 32 **Total size 4 bytes** 3369 ======= =================================================================================================================== 3370 3371.. 3372 3373 .. table:: compute_pgm_rsrc2 for GFX6-GFX10 3374 :name: amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table 3375 3376 ======= ======= =============================== =========================================================================== 3377 Bits Size Field Name Description 3378 ======= ======= =============================== =========================================================================== 3379 0 1 bit ENABLE_SGPR_PRIVATE_SEGMENT Enable the setup of the 3380 _WAVEFRONT_OFFSET SGPR wavefront scratch offset 3381 system register (see 3382 :ref:`amdgpu-amdhsa-initial-kernel-execution-state`). 3383 3384 Used by CP to set up 3385 ``COMPUTE_PGM_RSRC2.SCRATCH_EN``. 3386 5:1 5 bits USER_SGPR_COUNT The total number of SGPR 3387 user data registers 3388 requested. This number must 3389 match the number of user 3390 data registers enabled. 3391 3392 Used by CP to set up 3393 ``COMPUTE_PGM_RSRC2.USER_SGPR``. 3394 6 1 bit ENABLE_TRAP_HANDLER Must be 0. 3395 3396 This bit represents 3397 ``COMPUTE_PGM_RSRC2.TRAP_PRESENT``, 3398 which is set by the CP if 3399 the runtime has installed a 3400 trap handler. 3401 7 1 bit ENABLE_SGPR_WORKGROUP_ID_X Enable the setup of the 3402 system SGPR register for 3403 the work-group id in the X 3404 dimension (see 3405 :ref:`amdgpu-amdhsa-initial-kernel-execution-state`). 3406 3407 Used by CP to set up 3408 ``COMPUTE_PGM_RSRC2.TGID_X_EN``. 3409 8 1 bit ENABLE_SGPR_WORKGROUP_ID_Y Enable the setup of the 3410 system SGPR register for 3411 the work-group id in the Y 3412 dimension (see 3413 :ref:`amdgpu-amdhsa-initial-kernel-execution-state`). 3414 3415 Used by CP to set up 3416 ``COMPUTE_PGM_RSRC2.TGID_Y_EN``. 3417 9 1 bit ENABLE_SGPR_WORKGROUP_ID_Z Enable the setup of the 3418 system SGPR register for 3419 the work-group id in the Z 3420 dimension (see 3421 :ref:`amdgpu-amdhsa-initial-kernel-execution-state`). 3422 3423 Used by CP to set up 3424 ``COMPUTE_PGM_RSRC2.TGID_Z_EN``. 3425 10 1 bit ENABLE_SGPR_WORKGROUP_INFO Enable the setup of the 3426 system SGPR register for 3427 work-group information (see 3428 :ref:`amdgpu-amdhsa-initial-kernel-execution-state`). 3429 3430 Used by CP to set up 3431 ``COMPUTE_PGM_RSRC2.TGID_SIZE_EN``. 3432 12:11 2 bits ENABLE_VGPR_WORKITEM_ID Enable the setup of the 3433 VGPR system registers used 3434 for the work-item ID. 3435 :ref:`amdgpu-amdhsa-system-vgpr-work-item-id-enumeration-values-table` 3436 defines the values. 3437 3438 Used by CP to set up 3439 ``COMPUTE_PGM_RSRC2.TIDIG_CMP_CNT``. 3440 13 1 bit ENABLE_EXCEPTION_ADDRESS_WATCH Must be 0. 3441 3442 Wavefront starts execution 3443 with address watch 3444 exceptions enabled which 3445 are generated when L1 has 3446 witnessed a thread access 3447 an *address of 3448 interest*. 3449 3450 CP is responsible for 3451 filling in the address 3452 watch bit in 3453 ``COMPUTE_PGM_RSRC2.EXCP_EN_MSB`` 3454 according to what the 3455 runtime requests. 3456 14 1 bit ENABLE_EXCEPTION_MEMORY Must be 0. 3457 3458 Wavefront starts execution 3459 with memory violation 3460 exceptions exceptions 3461 enabled which are generated 3462 when a memory violation has 3463 occurred for this wavefront from 3464 L1 or LDS 3465 (write-to-read-only-memory, 3466 mis-aligned atomic, LDS 3467 address out of range, 3468 illegal address, etc.). 3469 3470 CP sets the memory 3471 violation bit in 3472 ``COMPUTE_PGM_RSRC2.EXCP_EN_MSB`` 3473 according to what the 3474 runtime requests. 3475 23:15 9 bits GRANULATED_LDS_SIZE Must be 0. 3476 3477 CP uses the rounded value 3478 from the dispatch packet, 3479 not this value, as the 3480 dispatch may contain 3481 dynamically allocated group 3482 segment memory. CP writes 3483 directly to 3484 ``COMPUTE_PGM_RSRC2.LDS_SIZE``. 3485 3486 Amount of group segment 3487 (LDS) to allocate for each 3488 work-group. Granularity is 3489 device specific: 3490 3491 GFX6: 3492 roundup(lds-size / (64 * 4)) 3493 GFX7-GFX10: 3494 roundup(lds-size / (128 * 4)) 3495 3496 24 1 bit ENABLE_EXCEPTION_IEEE_754_FP Wavefront starts execution 3497 _INVALID_OPERATION with specified exceptions 3498 enabled. 3499 3500 Used by CP to set up 3501 ``COMPUTE_PGM_RSRC2.EXCP_EN`` 3502 (set from bits 0..6). 3503 3504 IEEE 754 FP Invalid 3505 Operation 3506 25 1 bit ENABLE_EXCEPTION_FP_DENORMAL FP Denormal one or more 3507 _SOURCE input operands is a 3508 denormal number 3509 26 1 bit ENABLE_EXCEPTION_IEEE_754_FP IEEE 754 FP Division by 3510 _DIVISION_BY_ZERO Zero 3511 27 1 bit ENABLE_EXCEPTION_IEEE_754_FP IEEE 754 FP FP Overflow 3512 _OVERFLOW 3513 28 1 bit ENABLE_EXCEPTION_IEEE_754_FP IEEE 754 FP Underflow 3514 _UNDERFLOW 3515 29 1 bit ENABLE_EXCEPTION_IEEE_754_FP IEEE 754 FP Inexact 3516 _INEXACT 3517 30 1 bit ENABLE_EXCEPTION_INT_DIVIDE_BY Integer Division by Zero 3518 _ZERO (rcp_iflag_f32 instruction 3519 only) 3520 31 1 bit Reserved, must be 0. 3521 32 **Total size 4 bytes.** 3522 ======= =================================================================================================================== 3523 3524.. 3525 3526 .. table:: compute_pgm_rsrc3 for GFX10 3527 :name: amdgpu-amdhsa-compute_pgm_rsrc3-gfx10-table 3528 3529 ======= ======= =============================== =========================================================================== 3530 Bits Size Field Name Description 3531 ======= ======= =============================== =========================================================================== 3532 3:0 4 bits SHARED_VGPR_COUNT Number of shared VGPRs for wavefront size 64. Granularity 8. Value 0-120. 3533 compute_pgm_rsrc1.vgprs + shared_vgpr_cnt cannot exceed 64. 3534 31:4 28 Reserved, must be 0. 3535 bits 3536 32 **Total size 4 bytes.** 3537 ======= =================================================================================================================== 3538 3539.. 3540 3541 .. table:: Floating Point Rounding Mode Enumeration Values 3542 :name: amdgpu-amdhsa-floating-point-rounding-mode-enumeration-values-table 3543 3544 ====================================== ===== ============================== 3545 Enumeration Name Value Description 3546 ====================================== ===== ============================== 3547 FLOAT_ROUND_MODE_NEAR_EVEN 0 Round Ties To Even 3548 FLOAT_ROUND_MODE_PLUS_INFINITY 1 Round Toward +infinity 3549 FLOAT_ROUND_MODE_MINUS_INFINITY 2 Round Toward -infinity 3550 FLOAT_ROUND_MODE_ZERO 3 Round Toward 0 3551 ====================================== ===== ============================== 3552 3553.. 3554 3555 .. table:: Floating Point Denorm Mode Enumeration Values 3556 :name: amdgpu-amdhsa-floating-point-denorm-mode-enumeration-values-table 3557 3558 ====================================== ===== ============================== 3559 Enumeration Name Value Description 3560 ====================================== ===== ============================== 3561 FLOAT_DENORM_MODE_FLUSH_SRC_DST 0 Flush Source and Destination 3562 Denorms 3563 FLOAT_DENORM_MODE_FLUSH_DST 1 Flush Output Denorms 3564 FLOAT_DENORM_MODE_FLUSH_SRC 2 Flush Source Denorms 3565 FLOAT_DENORM_MODE_FLUSH_NONE 3 No Flush 3566 ====================================== ===== ============================== 3567 3568.. 3569 3570 .. table:: System VGPR Work-Item ID Enumeration Values 3571 :name: amdgpu-amdhsa-system-vgpr-work-item-id-enumeration-values-table 3572 3573 ======================================== ===== ============================ 3574 Enumeration Name Value Description 3575 ======================================== ===== ============================ 3576 SYSTEM_VGPR_WORKITEM_ID_X 0 Set work-item X dimension 3577 ID. 3578 SYSTEM_VGPR_WORKITEM_ID_X_Y 1 Set work-item X and Y 3579 dimensions ID. 3580 SYSTEM_VGPR_WORKITEM_ID_X_Y_Z 2 Set work-item X, Y and Z 3581 dimensions ID. 3582 SYSTEM_VGPR_WORKITEM_ID_UNDEFINED 3 Undefined. 3583 ======================================== ===== ============================ 3584 3585.. _amdgpu-amdhsa-initial-kernel-execution-state: 3586 3587Initial Kernel Execution State 3588~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 3589 3590This section defines the register state that will be set up by the packet 3591processor prior to the start of execution of every wavefront. This is limited by 3592the constraints of the hardware controllers of CP/ADC/SPI. 3593 3594The order of the SGPR registers is defined, but the compiler can specify which 3595ones are actually setup in the kernel descriptor using the ``enable_sgpr_*`` bit 3596fields (see :ref:`amdgpu-amdhsa-kernel-descriptor`). The register numbers used 3597for enabled registers are dense starting at SGPR0: the first enabled register is 3598SGPR0, the next enabled register is SGPR1 etc.; disabled registers do not have 3599an SGPR number. 3600 3601The initial SGPRs comprise up to 16 User SRGPs that are set by CP and apply to 3602all wavefronts of the grid. It is possible to specify more than 16 User SGPRs 3603using the ``enable_sgpr_*`` bit fields, in which case only the first 16 are 3604actually initialized. These are then immediately followed by the System SGPRs 3605that are set up by ADC/SPI and can have different values for each wavefront of 3606the grid dispatch. 3607 3608SGPR register initial state is defined in 3609:ref:`amdgpu-amdhsa-sgpr-register-set-up-order-table`. 3610 3611 .. table:: SGPR Register Set Up Order 3612 :name: amdgpu-amdhsa-sgpr-register-set-up-order-table 3613 3614 ========== ========================== ====== ============================== 3615 SGPR Order Name Number Description 3616 (kernel descriptor enable of 3617 field) SGPRs 3618 ========== ========================== ====== ============================== 3619 First Private Segment Buffer 4 V# that can be used, together 3620 (enable_sgpr_private with Scratch Wavefront Offset 3621 _segment_buffer) as an offset, to access the 3622 private address space using a 3623 segment address. 3624 3625 CP uses the value provided by 3626 the runtime. 3627 then Dispatch Ptr 2 64-bit address of AQL dispatch 3628 (enable_sgpr_dispatch_ptr) packet for kernel dispatch 3629 actually executing. 3630 then Queue Ptr 2 64-bit address of amd_queue_t 3631 (enable_sgpr_queue_ptr) object for AQL queue on which 3632 the dispatch packet was 3633 queued. 3634 then Kernarg Segment Ptr 2 64-bit address of Kernarg 3635 (enable_sgpr_kernarg segment. This is directly 3636 _segment_ptr) copied from the 3637 kernarg_address in the kernel 3638 dispatch packet. 3639 3640 Having CP load it once avoids 3641 loading it at the beginning of 3642 every wavefront. 3643 then Dispatch Id 2 64-bit Dispatch ID of the 3644 (enable_sgpr_dispatch_id) dispatch packet being 3645 executed. 3646 then Flat Scratch Init 2 This is 2 SGPRs: 3647 (enable_sgpr_flat_scratch 3648 _init) GFX6 3649 Not supported. 3650 GFX7-GFX8 3651 The first SGPR is a 32-bit 3652 byte offset from 3653 ``SH_HIDDEN_PRIVATE_BASE_VIMID`` 3654 to per SPI base of memory 3655 for scratch for the queue 3656 executing the kernel 3657 dispatch. CP obtains this 3658 from the runtime. (The 3659 Scratch Segment Buffer base 3660 address is 3661 ``SH_HIDDEN_PRIVATE_BASE_VIMID`` 3662 plus this offset.) The value 3663 of Scratch Wavefront Offset must 3664 be added to this offset by 3665 the kernel machine code, 3666 right shifted by 8, and 3667 moved to the FLAT_SCRATCH_HI 3668 SGPR register. 3669 FLAT_SCRATCH_HI corresponds 3670 to SGPRn-4 on GFX7, and 3671 SGPRn-6 on GFX8 (where SGPRn 3672 is the highest numbered SGPR 3673 allocated to the wavefront). 3674 FLAT_SCRATCH_HI is 3675 multiplied by 256 (as it is 3676 in units of 256 bytes) and 3677 added to 3678 ``SH_HIDDEN_PRIVATE_BASE_VIMID`` 3679 to calculate the per wavefront 3680 FLAT SCRATCH BASE in flat 3681 memory instructions that 3682 access the scratch 3683 aperture. 3684 3685 The second SGPR is 32-bit 3686 byte size of a single 3687 work-item's scratch memory 3688 usage. CP obtains this from 3689 the runtime, and it is 3690 always a multiple of DWORD. 3691 CP checks that the value in 3692 the kernel dispatch packet 3693 Private Segment Byte Size is 3694 not larger and requests the 3695 runtime to increase the 3696 queue's scratch size if 3697 necessary. The kernel code 3698 must move it to 3699 FLAT_SCRATCH_LO which is 3700 SGPRn-3 on GFX7 and SGPRn-5 3701 on GFX8. FLAT_SCRATCH_LO is 3702 used as the FLAT SCRATCH 3703 SIZE in flat memory 3704 instructions. Having CP load 3705 it once avoids loading it at 3706 the beginning of every 3707 wavefront. 3708 GFX9-GFX10 3709 This is the 3710 64-bit base address of the 3711 per SPI scratch backing 3712 memory managed by SPI for 3713 the queue executing the 3714 kernel dispatch. CP obtains 3715 this from the runtime (and 3716 divides it if there are 3717 multiple Shader Arrays each 3718 with its own SPI). The value 3719 of Scratch Wavefront Offset must 3720 be added by the kernel 3721 machine code and the result 3722 moved to the FLAT_SCRATCH 3723 SGPR which is SGPRn-6 and 3724 SGPRn-5. It is used as the 3725 FLAT SCRATCH BASE in flat 3726 memory instructions. 3727 then Private Segment Size 1 The 32-bit byte size of a 3728 (enable_sgpr_private single 3729 work-item's 3730 scratch_segment_size) memory 3731 allocation. This is the 3732 value from the kernel 3733 dispatch packet Private 3734 Segment Byte Size rounded up 3735 by CP to a multiple of 3736 DWORD. 3737 3738 Having CP load it once avoids 3739 loading it at the beginning of 3740 every wavefront. 3741 3742 This is not used for 3743 GFX7-GFX8 since it is the same 3744 value as the second SGPR of 3745 Flat Scratch Init. However, it 3746 may be needed for GFX9-GFX10 which 3747 changes the meaning of the 3748 Flat Scratch Init value. 3749 then Grid Work-Group Count X 1 32-bit count of the number of 3750 (enable_sgpr_grid work-groups in the X dimension 3751 _workgroup_count_X) for the grid being 3752 executed. Computed from the 3753 fields in the kernel dispatch 3754 packet as ((grid_size.x + 3755 workgroup_size.x - 1) / 3756 workgroup_size.x). 3757 then Grid Work-Group Count Y 1 32-bit count of the number of 3758 (enable_sgpr_grid work-groups in the Y dimension 3759 _workgroup_count_Y && for the grid being 3760 less than 16 previous executed. Computed from the 3761 SGPRs) fields in the kernel dispatch 3762 packet as ((grid_size.y + 3763 workgroup_size.y - 1) / 3764 workgroupSize.y). 3765 3766 Only initialized if <16 3767 previous SGPRs initialized. 3768 then Grid Work-Group Count Z 1 32-bit count of the number of 3769 (enable_sgpr_grid work-groups in the Z dimension 3770 _workgroup_count_Z && for the grid being 3771 less than 16 previous executed. Computed from the 3772 SGPRs) fields in the kernel dispatch 3773 packet as ((grid_size.z + 3774 workgroup_size.z - 1) / 3775 workgroupSize.z). 3776 3777 Only initialized if <16 3778 previous SGPRs initialized. 3779 then Work-Group Id X 1 32-bit work-group id in X 3780 (enable_sgpr_workgroup_id dimension of grid for 3781 _X) wavefront. 3782 then Work-Group Id Y 1 32-bit work-group id in Y 3783 (enable_sgpr_workgroup_id dimension of grid for 3784 _Y) wavefront. 3785 then Work-Group Id Z 1 32-bit work-group id in Z 3786 (enable_sgpr_workgroup_id dimension of grid for 3787 _Z) wavefront. 3788 then Work-Group Info 1 {first_wavefront, 14'b0000, 3789 (enable_sgpr_workgroup ordered_append_term[10:0], 3790 _info) threadgroup_size_in_wavefronts[5:0]} 3791 then Scratch Wavefront Offset 1 32-bit byte offset from base 3792 (enable_sgpr_private of scratch base of queue 3793 _segment_wavefront_offset) executing the kernel 3794 dispatch. Must be used as an 3795 offset with Private 3796 segment address when using 3797 Scratch Segment Buffer. It 3798 must be used to set up FLAT 3799 SCRATCH for flat addressing 3800 (see 3801 :ref:`amdgpu-amdhsa-kernel-prolog-flat-scratch`). 3802 ========== ========================== ====== ============================== 3803 3804The order of the VGPR registers is defined, but the compiler can specify which 3805ones are actually setup in the kernel descriptor using the ``enable_vgpr*`` bit 3806fields (see :ref:`amdgpu-amdhsa-kernel-descriptor`). The register numbers used 3807for enabled registers are dense starting at VGPR0: the first enabled register is 3808VGPR0, the next enabled register is VGPR1 etc.; disabled registers do not have a 3809VGPR number. 3810 3811VGPR register initial state is defined in 3812:ref:`amdgpu-amdhsa-vgpr-register-set-up-order-table`. 3813 3814 .. table:: VGPR Register Set Up Order 3815 :name: amdgpu-amdhsa-vgpr-register-set-up-order-table 3816 3817 ========== ========================== ====== ============================== 3818 VGPR Order Name Number Description 3819 (kernel descriptor enable of 3820 field) VGPRs 3821 ========== ========================== ====== ============================== 3822 First Work-Item Id X 1 32-bit work item id in X 3823 (Always initialized) dimension of work-group for 3824 wavefront lane. 3825 then Work-Item Id Y 1 32-bit work item id in Y 3826 (enable_vgpr_workitem_id dimension of work-group for 3827 > 0) wavefront lane. 3828 then Work-Item Id Z 1 32-bit work item id in Z 3829 (enable_vgpr_workitem_id dimension of work-group for 3830 > 1) wavefront lane. 3831 ========== ========================== ====== ============================== 3832 3833The setting of registers is done by GPU CP/ADC/SPI hardware as follows: 3834 38351. SGPRs before the Work-Group Ids are set by CP using the 16 User Data 3836 registers. 38372. Work-group Id registers X, Y, Z are set by ADC which supports any 3838 combination including none. 38393. Scratch Wavefront Offset is set by SPI in a per wavefront basis which is why 3840 its value cannot be included with the flat scratch init value which is per 3841 queue. 38424. The VGPRs are set by SPI which only supports specifying either (X), (X, Y) 3843 or (X, Y, Z). 3844 3845Flat Scratch register pair are adjacent SGPRs so they can be moved as a 64-bit 3846value to the hardware required SGPRn-3 and SGPRn-4 respectively. 3847 3848The global segment can be accessed either using buffer instructions (GFX6 which 3849has V# 64-bit address support), flat instructions (GFX7-GFX10), or global 3850instructions (GFX9-GFX10). 3851 3852If buffer operations are used, then the compiler can generate a V# with the 3853following properties: 3854 3855* base address of 0 3856* no swizzle 3857* ATC: 1 if IOMMU present (such as APU) 3858* ptr64: 1 3859* MTYPE set to support memory coherence that matches the runtime (such as CC for 3860 APU and NC for dGPU). 3861 3862.. _amdgpu-amdhsa-kernel-prolog: 3863 3864Kernel Prolog 3865~~~~~~~~~~~~~ 3866 3867The compiler performs initialization in the kernel prologue depending on the 3868target and information about things like stack usage in the kernel and called 3869functions. Some of this initialization requires the compiler to request certain 3870User and System SGPRs be present in the 3871:ref:`amdgpu-amdhsa-initial-kernel-execution-state` via the 3872:ref:`amdgpu-amdhsa-kernel-descriptor`. 3873 3874.. _amdgpu-amdhsa-kernel-prolog-cfi: 3875 3876CFI 3877+++ 3878 38791. The CFI return address is undefined. 3880 38812. The CFI CFA is defined using an expression which evaluates to a location 3882 description that comprises one memory location description for the 3883 ``DW_ASPACE_AMDGPU_private_lane`` address space address ``0``. 3884 3885.. _amdgpu-amdhsa-kernel-prolog-m0: 3886 3887M0 3888++ 3889 3890GFX6-GFX8 3891 The M0 register must be initialized with a value at least the total LDS size 3892 if the kernel may access LDS via DS or flat operations. Total LDS size is 3893 available in dispatch packet. For M0, it is also possible to use maximum 3894 possible value of LDS for given target (0x7FFF for GFX6 and 0xFFFF for 3895 GFX7-GFX8). 3896GFX9-GFX10 3897 The M0 register is not used for range checking LDS accesses and so does not 3898 need to be initialized in the prolog. 3899 3900.. _amdgpu-amdhsa-kernel-prolog-stack-pointer: 3901 3902Stack Pointer 3903+++++++++++++ 3904 3905If the kernel has function calls it must set up the ABI stack pointer described 3906in :ref:`amdgpu-amdhsa-function-call-convention-non-kernel-functions` by setting 3907SGPR32 to the unswizzled scratch offset of the address past the last local 3908allocation. 3909 3910.. _amdgpu-amdhsa-kernel-prolog-frame-pointer: 3911 3912Frame Pointer 3913+++++++++++++ 3914 3915If the kernel needs a frame pointer for the reasons defined in 3916``SIFrameLowering`` then SGPR33 is used and is always set to ``0`` in the 3917kernel prolog. If a frame pointer is not required then all uses of the frame 3918pointer are replaced with immediate ``0`` offsets. 3919 3920.. _amdgpu-amdhsa-kernel-prolog-flat-scratch: 3921 3922Flat Scratch 3923++++++++++++ 3924 3925If the kernel or any function it calls may use flat operations to access 3926scratch memory, the prolog code must set up the FLAT_SCRATCH register pair 3927(FLAT_SCRATCH_LO/FLAT_SCRATCH_HI which are in SGPRn-4/SGPRn-3). Initialization 3928uses Flat Scratch Init and Scratch Wavefront Offset SGPR registers (see 3929:ref:`amdgpu-amdhsa-initial-kernel-execution-state`): 3930 3931GFX6 3932 Flat scratch is not supported. 3933 3934GFX7-GFX8 3935 3936 1. The low word of Flat Scratch Init is 32-bit byte offset from 3937 ``SH_HIDDEN_PRIVATE_BASE_VIMID`` to the base of scratch backing memory 3938 being managed by SPI for the queue executing the kernel dispatch. This is 3939 the same value used in the Scratch Segment Buffer V# base address. The 3940 prolog must add the value of Scratch Wavefront Offset to get the 3941 wavefront's byte scratch backing memory offset from 3942 ``SH_HIDDEN_PRIVATE_BASE_VIMID``. Since FLAT_SCRATCH_LO is in units of 256 3943 bytes, the offset must be right shifted by 8 before moving into 3944 FLAT_SCRATCH_LO. 3945 2. The second word of Flat Scratch Init is 32-bit byte size of a single 3946 work-items scratch memory usage. This is directly loaded from the kernel 3947 dispatch packet Private Segment Byte Size and rounded up to a multiple of 3948 DWORD. Having CP load it once avoids loading it at the beginning of every 3949 wavefront. The prolog must move it to FLAT_SCRATCH_LO for use as FLAT 3950 SCRATCH SIZE. 3951 3952GFX9-GFX10 3953 The Flat Scratch Init is the 64-bit address of the base of scratch backing 3954 memory being managed by SPI for the queue executing the kernel dispatch. The 3955 prolog must add the value of Scratch Wavefront Offset and moved to the 3956 FLAT_SCRATCH pair for use as the flat scratch base in flat memory 3957 instructions. 3958 3959.. _amdgpu-amdhsa-kernel-prolog-private-segment-buffer: 3960 3961Private Segment Buffer 3962++++++++++++++++++++++ 3963 3964A set of four SGPRs beginning at a four-aligned SGPR index are always selected 3965to serve as the scratch V# for the kernel as follows: 3966 3967 - If it is known during instruction selection that there is stack usage, 3968 SGPR0-3 is reserved for use as the scratch V#. Stack usage is assumed if 3969 optimizations are disabled (``-O0``), if stack objects already exist (for 3970 locals, etc.), or if there are any function calls. 3971 3972 - Otherwise, four high numbered SGPRs beginning at a four-aligned SGPR index 3973 are reserved for the tentative scratch V#. These will be used if it is 3974 determined that spilling is needed. 3975 3976 - If no use is made of the tentative scratch V#, then it is unreserved, 3977 and the register count is determined ignoring it. 3978 - If use is made of the tentative scratch V#, then its register numbers 3979 are shifted to the first four-aligned SGPR index after the highest one 3980 allocated by the register allocator, and all uses are updated. The 3981 register count includes them in the shifted location. 3982 - In either case, if the processor has the SGPR allocation bug, the 3983 tentative allocation is not shifted or unreserved in order to ensure 3984 the register count is higher to workaround the bug. 3985 3986 .. note:: 3987 3988 This approach of using a tentative scratch V# and shifting the register 3989 numbers if used avoids having to perform register allocation a second 3990 time if the tentative V# is eliminated. This is more efficient and 3991 avoids the problem that the second register allocation may perform 3992 spilling which will fail as there is no longer a scratch V#. 3993 3994When the kernel prolog code is being emitted it is known whether the scratch V# 3995described above is actually used. If it is, the prolog code must set it up by 3996copying the Private Segment Buffer to the scratch V# registers and then adding 3997the Private Segment Wavefront Offset to the queue base address in the V#. The 3998result is a V# with a base address pointing to the beginning of the wavefront 3999scratch backing memory. 4000 4001The Private Segment Buffer is always requested, but the Private Segment 4002Wavefront Offset is only requested if it is used (see 4003:ref:`amdgpu-amdhsa-initial-kernel-execution-state`). 4004 4005.. _amdgpu-amdhsa-memory-model: 4006 4007Memory Model 4008~~~~~~~~~~~~ 4009 4010This section describes the mapping of LLVM memory model onto AMDGPU machine code 4011(see :ref:`memmodel`). 4012 4013The AMDGPU backend supports the memory synchronization scopes specified in 4014:ref:`amdgpu-memory-scopes`. 4015 4016The code sequences used to implement the memory model are defined in table 4017:ref:`amdgpu-amdhsa-memory-model-code-sequences-gfx6-gfx10-table`. 4018 4019The sequences specify the order of instructions that a single thread must 4020execute. The ``s_waitcnt`` and ``buffer_wbinvl1_vol`` are defined with respect 4021to other memory instructions executed by the same thread. This allows them to be 4022moved earlier or later which can allow them to be combined with other instances 4023of the same instruction, or hoisted/sunk out of loops to improve 4024performance. Only the instructions related to the memory model are given; 4025additional ``s_waitcnt`` instructions are required to ensure registers are 4026defined before being used. These may be able to be combined with the memory 4027model ``s_waitcnt`` instructions as described above. 4028 4029The AMDGPU backend supports the following memory models: 4030 4031 HSA Memory Model [HSA]_ 4032 The HSA memory model uses a single happens-before relation for all address 4033 spaces (see :ref:`amdgpu-address-spaces`). 4034 OpenCL Memory Model [OpenCL]_ 4035 The OpenCL memory model which has separate happens-before relations for the 4036 global and local address spaces. Only a fence specifying both global and 4037 local address space, and seq_cst instructions join the relationships. Since 4038 the LLVM ``memfence`` instruction does not allow an address space to be 4039 specified the OpenCL fence has to conservatively assume both local and 4040 global address space was specified. However, optimizations can often be 4041 done to eliminate the additional ``s_waitcnt`` instructions when there are 4042 no intervening memory instructions which access the corresponding address 4043 space. The code sequences in the table indicate what can be omitted for the 4044 OpenCL memory. The target triple environment is used to determine if the 4045 source language is OpenCL (see :ref:`amdgpu-opencl`). 4046 4047``ds/flat_load/store/atomic`` instructions to local memory are termed LDS 4048operations. 4049 4050``buffer/global/flat_load/store/atomic`` instructions to global memory are 4051termed vector memory operations. 4052 4053For GFX6-GFX9: 4054 4055* Each agent has multiple shader arrays (SA). 4056* Each SA has multiple compute units (CU). 4057* Each CU has multiple SIMDs that execute wavefronts. 4058* The wavefronts for a single work-group are executed in the same CU but may be 4059 executed by different SIMDs. 4060* Each CU has a single LDS memory shared by the wavefronts of the work-groups 4061 executing on it. 4062* All LDS operations of a CU are performed as wavefront wide operations in a 4063 global order and involve no caching. Completion is reported to a wavefront in 4064 execution order. 4065* The LDS memory has multiple request queues shared by the SIMDs of a 4066 CU. Therefore, the LDS operations performed by different wavefronts of a 4067 work-group can be reordered relative to each other, which can result in 4068 reordering the visibility of vector memory operations with respect to LDS 4069 operations of other wavefronts in the same work-group. A ``s_waitcnt 4070 lgkmcnt(0)`` is required to ensure synchronization between LDS operations and 4071 vector memory operations between wavefronts of a work-group, but not between 4072 operations performed by the same wavefront. 4073* The vector memory operations are performed as wavefront wide operations and 4074 completion is reported to a wavefront in execution order. The exception is 4075 that for GFX7-GFX9 ``flat_load/store/atomic`` instructions can report out of 4076 vector memory order if they access LDS memory, and out of LDS operation order 4077 if they access global memory. 4078* The vector memory operations access a single vector L1 cache shared by all 4079 SIMDs a CU. Therefore, no special action is required for coherence between the 4080 lanes of a single wavefront, or for coherence between wavefronts in the same 4081 work-group. A ``buffer_wbinvl1_vol`` is required for coherence between 4082 wavefronts executing in different work-groups as they may be executing on 4083 different CUs. 4084* The scalar memory operations access a scalar L1 cache shared by all wavefronts 4085 on a group of CUs. The scalar and vector L1 caches are not coherent. However, 4086 scalar operations are used in a restricted way so do not impact the memory 4087 model. See :ref:`amdgpu-address-spaces`. 4088* The vector and scalar memory operations use an L2 cache shared by all CUs on 4089 the same agent. 4090* The L2 cache has independent channels to service disjoint ranges of virtual 4091 addresses. 4092* Each CU has a separate request queue per channel. Therefore, the vector and 4093 scalar memory operations performed by wavefronts executing in different 4094 work-groups (which may be executing on different CUs) of an agent can be 4095 reordered relative to each other. A ``s_waitcnt vmcnt(0)`` is required to 4096 ensure synchronization between vector memory operations of different CUs. It 4097 ensures a previous vector memory operation has completed before executing a 4098 subsequent vector memory or LDS operation and so can be used to meet the 4099 requirements of acquire and release. 4100* The L2 cache can be kept coherent with other agents on some targets, or ranges 4101 of virtual addresses can be set up to bypass it to ensure system coherence. 4102 4103For GFX10: 4104 4105* Each agent has multiple shader arrays (SA). 4106* Each SA has multiple work-group processors (WGP). 4107* Each WGP has multiple compute units (CU). 4108* Each CU has multiple SIMDs that execute wavefronts. 4109* The wavefronts for a single work-group are executed in the same 4110 WGP. In CU wavefront execution mode the wavefronts may be executed by 4111 different SIMDs in the same CU. In WGP wavefront execution mode the 4112 wavefronts may be executed by different SIMDs in different CUs in the same 4113 WGP. 4114* Each WGP has a single LDS memory shared by the wavefronts of the work-groups 4115 executing on it. 4116* All LDS operations of a WGP are performed as wavefront wide operations in a 4117 global order and involve no caching. Completion is reported to a wavefront in 4118 execution order. 4119* The LDS memory has multiple request queues shared by the SIMDs of a 4120 WGP. Therefore, the LDS operations performed by different wavefronts of a 4121 work-group can be reordered relative to each other, which can result in 4122 reordering the visibility of vector memory operations with respect to LDS 4123 operations of other wavefronts in the same work-group. A ``s_waitcnt 4124 lgkmcnt(0)`` is required to ensure synchronization between LDS operations and 4125 vector memory operations between wavefronts of a work-group, but not between 4126 operations performed by the same wavefront. 4127* The vector memory operations are performed as wavefront wide operations. 4128 Completion of load/store/sample operations are reported to a wavefront in 4129 execution order of other load/store/sample operations performed by that 4130 wavefront. 4131* The vector memory operations access a vector L0 cache. There is a single L0 4132 cache per CU. Each SIMD of a CU accesses the same L0 cache. Therefore, no 4133 special action is required for coherence between the lanes of a single 4134 wavefront. However, a ``BUFFER_GL0_INV`` is required for coherence between 4135 wavefronts executing in the same work-group as they may be executing on SIMDs 4136 of different CUs that access different L0s. A ``BUFFER_GL0_INV`` is also 4137 required for coherence between wavefronts executing in different work-groups 4138 as they may be executing on different WGPs. 4139* The scalar memory operations access a scalar L0 cache shared by all wavefronts 4140 on a WGP. The scalar and vector L0 caches are not coherent. However, scalar 4141 operations are used in a restricted way so do not impact the memory model. See 4142 :ref:`amdgpu-address-spaces`. 4143* The vector and scalar memory L0 caches use an L1 cache shared by all WGPs on 4144 the same SA. Therefore, no special action is required for coherence between 4145 the wavefronts of a single work-group. However, a ``BUFFER_GL1_INV`` is 4146 required for coherence between wavefronts executing in different work-groups 4147 as they may be executing on different SAs that access different L1s. 4148* The L1 caches have independent quadrants to service disjoint ranges of virtual 4149 addresses. 4150* Each L0 cache has a separate request queue per L1 quadrant. Therefore, the 4151 vector and scalar memory operations performed by different wavefronts, whether 4152 executing in the same or different work-groups (which may be executing on 4153 different CUs accessing different L0s), can be reordered relative to each 4154 other. A ``s_waitcnt vmcnt(0) & vscnt(0)`` is required to ensure 4155 synchronization between vector memory operations of different wavefronts. It 4156 ensures a previous vector memory operation has completed before executing a 4157 subsequent vector memory or LDS operation and so can be used to meet the 4158 requirements of acquire, release and sequential consistency. 4159* The L1 caches use an L2 cache shared by all SAs on the same agent. 4160* The L2 cache has independent channels to service disjoint ranges of virtual 4161 addresses. 4162* Each L1 quadrant of a single SA accesses a different L2 channel. Each L1 4163 quadrant has a separate request queue per L2 channel. Therefore, the vector 4164 and scalar memory operations performed by wavefronts executing in different 4165 work-groups (which may be executing on different SAs) of an agent can be 4166 reordered relative to each other. A ``s_waitcnt vmcnt(0) & vscnt(0)`` is 4167 required to ensure synchronization between vector memory operations of 4168 different SAs. It ensures a previous vector memory operation has completed 4169 before executing a subsequent vector memory and so can be used to meet the 4170 requirements of acquire, release and sequential consistency. 4171* The L2 cache can be kept coherent with other agents on some targets, or ranges 4172 of virtual addresses can be set up to bypass it to ensure system coherence. 4173 4174Private address space uses ``buffer_load/store`` using the scratch V# 4175(GFX6-GFX8), or ``scratch_load/store`` (GFX9-GFX10). Since only a single thread 4176is accessing the memory, atomic memory orderings are not meaningful, and all 4177accesses are treated as non-atomic. 4178 4179Constant address space uses ``buffer/global_load`` instructions (or equivalent 4180scalar memory instructions). Since the constant address space contents do not 4181change during the execution of a kernel dispatch it is not legal to perform 4182stores, and atomic memory orderings are not meaningful, and all access are 4183treated as non-atomic. 4184 4185A memory synchronization scope wider than work-group is not meaningful for the 4186group (LDS) address space and is treated as work-group. 4187 4188The memory model does not support the region address space which is treated as 4189non-atomic. 4190 4191Acquire memory ordering is not meaningful on store atomic instructions and is 4192treated as non-atomic. 4193 4194Release memory ordering is not meaningful on load atomic instructions and is 4195treated a non-atomic. 4196 4197Acquire-release memory ordering is not meaningful on load or store atomic 4198instructions and is treated as acquire and release respectively. 4199 4200AMDGPU backend only uses scalar memory operations to access memory that is 4201proven to not change during the execution of the kernel dispatch. This includes 4202constant address space and global address space for program scope const 4203variables. Therefore, the kernel machine code does not have to maintain the 4204scalar L1 cache to ensure it is coherent with the vector L1 cache. The scalar 4205and vector L1 caches are invalidated between kernel dispatches by CP since 4206constant address space data may change between kernel dispatch executions. See 4207:ref:`amdgpu-address-spaces`. 4208 4209The one exception is if scalar writes are used to spill SGPR registers. In this 4210case the AMDGPU backend ensures the memory location used to spill is never 4211accessed by vector memory operations at the same time. If scalar writes are used 4212then a ``s_dcache_wb`` is inserted before the ``s_endpgm`` and before a function 4213return since the locations may be used for vector memory instructions by a 4214future wavefront that uses the same scratch area, or a function call that 4215creates a frame at the same address, respectively. There is no need for a 4216``s_dcache_inv`` as all scalar writes are write-before-read in the same thread. 4217 4218For GFX6-GFX9, scratch backing memory (which is used for the private address 4219space) is accessed with MTYPE NC_NV (non-coherent non-volatile). Since the 4220private address space is only accessed by a single thread, and is always 4221write-before-read, there is never a need to invalidate these entries from the L1 4222cache. Hence all cache invalidates are done as ``*_vol`` to only invalidate the 4223volatile cache lines. 4224 4225For GFX10, scratch backing memory (which is used for the private address space) 4226is accessed with MTYPE NC (non-coherent). Since the private address space is 4227only accessed by a single thread, and is always write-before-read, there is 4228never a need to invalidate these entries from the L0 or L1 caches. 4229 4230For GFX10, wavefronts are executed in native mode with in-order reporting of 4231loads and sample instructions. In this mode vmcnt reports completion of load, 4232atomic with return and sample instructions in order, and the vscnt reports the 4233completion of store and atomic without return in order. See ``MEM_ORDERED`` 4234field in :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`. 4235 4236In GFX10, wavefronts can be executed in WGP or CU wavefront execution mode: 4237 4238* In WGP wavefront execution mode the wavefronts of a work-group are executed 4239 on the SIMDs of both CUs of the WGP. Therefore, explicit management of the per 4240 CU L0 caches is required for work-group synchronization. Also accesses to L1 4241 at work-group scope need to be explicitly ordered as the accesses from 4242 different CUs are not ordered. 4243* In CU wavefront execution mode the wavefronts of a work-group are executed on 4244 the SIMDs of a single CU of the WGP. Therefore, all global memory access by 4245 the work-group access the same L0 which in turn ensures L1 accesses are 4246 ordered and so do not require explicit management of the caches for 4247 work-group synchronization. 4248 4249See ``WGP_MODE`` field in 4250:ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table` and 4251:ref:`amdgpu-target-features`. 4252 4253On dGPU the kernarg backing memory is accessed as UC (uncached) to avoid needing 4254to invalidate the L2 cache. For GFX6-GFX9, this also causes it to be treated as 4255non-volatile and so is not invalidated by ``*_vol``. On APU it is accessed as CC 4256(cache coherent) and so the L2 cache will be coherent with the CPU and other 4257agents. 4258 4259 .. table:: AMDHSA Memory Model Code Sequences GFX6-GFX10 4260 :name: amdgpu-amdhsa-memory-model-code-sequences-gfx6-gfx10-table 4261 4262 ============ ============ ============== ========== =============================== ================================== 4263 LLVM Instr LLVM Memory LLVM Memory AMDGPU AMDGPU Machine Code AMDGPU Machine Code 4264 Ordering Sync Scope Address GFX6-9 GFX10 4265 Space 4266 ============ ============ ============== ========== =============================== ================================== 4267 **Non-Atomic** 4268 ---------------------------------------------------------------------------------------------------------------------- 4269 load *none* *none* - global - !volatile & !nontemporal - !volatile & !nontemporal 4270 - generic 4271 - private 1. buffer/global/flat_load 1. buffer/global/flat_load 4272 - constant 4273 - volatile & !nontemporal - volatile & !nontemporal 4274 4275 1. buffer/global/flat_load 1. buffer/global/flat_load 4276 glc=1 glc=1 dlc=1 4277 4278 - nontemporal - nontemporal 4279 4280 1. buffer/global/flat_load 1. buffer/global/flat_load 4281 glc=1 slc=1 slc=1 4282 4283 load *none* *none* - local 1. ds_load 1. ds_load 4284 store *none* *none* - global - !nontemporal - !nontemporal 4285 - generic 4286 - private 1. buffer/global/flat_store 1. buffer/global/flat_store 4287 - constant 4288 - nontemporal - nontemporal 4289 4290 1. buffer/global/flat_store 1. buffer/global/flat_store 4291 glc=1 slc=1 slc=1 4292 4293 store *none* *none* - local 1. ds_store 1. ds_store 4294 **Unordered Atomic** 4295 ---------------------------------------------------------------------------------------------------------------------- 4296 load atomic unordered *any* *any* *Same as non-atomic*. *Same as non-atomic*. 4297 store atomic unordered *any* *any* *Same as non-atomic*. *Same as non-atomic*. 4298 atomicrmw unordered *any* *any* *Same as monotonic *Same as monotonic 4299 atomic*. atomic*. 4300 **Monotonic Atomic** 4301 ---------------------------------------------------------------------------------------------------------------------- 4302 load atomic monotonic - singlethread - global 1. buffer/global/flat_load 1. buffer/global/flat_load 4303 - wavefront - generic 4304 load atomic monotonic - workgroup - global 1. buffer/global/flat_load 1. buffer/global/flat_load 4305 - generic glc=1 4306 4307 - If CU wavefront execution mode, omit glc=1. 4308 4309 load atomic monotonic - singlethread - local 1. ds_load 1. ds_load 4310 - wavefront 4311 - workgroup 4312 load atomic monotonic - agent - global 1. buffer/global/flat_load 1. buffer/global/flat_load 4313 - system - generic glc=1 glc=1 dlc=1 4314 store atomic monotonic - singlethread - global 1. buffer/global/flat_store 1. buffer/global/flat_store 4315 - wavefront - generic 4316 - workgroup 4317 - agent 4318 - system 4319 store atomic monotonic - singlethread - local 1. ds_store 1. ds_store 4320 - wavefront 4321 - workgroup 4322 atomicrmw monotonic - singlethread - global 1. buffer/global/flat_atomic 1. buffer/global/flat_atomic 4323 - wavefront - generic 4324 - workgroup 4325 - agent 4326 - system 4327 atomicrmw monotonic - singlethread - local 1. ds_atomic 1. ds_atomic 4328 - wavefront 4329 - workgroup 4330 **Acquire Atomic** 4331 ---------------------------------------------------------------------------------------------------------------------- 4332 load atomic acquire - singlethread - global 1. buffer/global/ds/flat_load 1. buffer/global/ds/flat_load 4333 - wavefront - local 4334 - generic 4335 load atomic acquire - workgroup - global 1. buffer/global/flat_load 1. buffer/global_load glc=1 4336 4337 - If CU wavefront execution mode, omit glc=1. 4338 4339 2. s_waitcnt vmcnt(0) 4340 4341 - If CU wavefront execution mode, omit. 4342 - Must happen before 4343 the following buffer_gl0_inv 4344 and before any following 4345 global/generic 4346 load/load 4347 atomic/store/store 4348 atomic/atomicrmw. 4349 4350 3. buffer_gl0_inv 4351 4352 - If CU wavefront execution mode, omit. 4353 - Ensures that 4354 following 4355 loads will not see 4356 stale data. 4357 4358 load atomic acquire - workgroup - local 1. ds_load 1. ds_load 4359 2. s_waitcnt lgkmcnt(0) 2. s_waitcnt lgkmcnt(0) 4360 4361 - If OpenCL, omit. - If OpenCL, omit. 4362 - Must happen before - Must happen before 4363 any following the following buffer_gl0_inv 4364 global/generic and before any following 4365 load/load global/generic load/load 4366 atomic/store/store atomic/store/store 4367 atomic/atomicrmw. atomic/atomicrmw. 4368 - Ensures any - Ensures any 4369 following global following global 4370 data read is no data read is no 4371 older than the load older than the load 4372 atomic value being atomic value being 4373 acquired. acquired. 4374 4375 3. buffer_gl0_inv 4376 4377 - If CU wavefront execution mode, omit. 4378 - If OpenCL, omit. 4379 - Ensures that 4380 following 4381 loads will not see 4382 stale data. 4383 4384 load atomic acquire - workgroup - generic 1. flat_load 1. flat_load glc=1 4385 4386 - If CU wavefront execution mode, omit glc=1. 4387 4388 2. s_waitcnt lgkmcnt(0) 2. s_waitcnt lgkmcnt(0) & 4389 vmcnt(0) 4390 4391 - If CU wavefront execution mode, omit vmcnt. 4392 - If OpenCL, omit. - If OpenCL, omit 4393 lgkmcnt(0). 4394 - Must happen before - Must happen before 4395 any following the following 4396 global/generic buffer_gl0_inv and any 4397 load/load following global/generic 4398 atomic/store/store load/load 4399 atomic/atomicrmw. atomic/store/store 4400 atomic/atomicrmw. 4401 - Ensures any - Ensures any 4402 following global following global 4403 data read is no data read is no 4404 older than the load older than the load 4405 atomic value being atomic value being 4406 acquired. acquired. 4407 4408 3. buffer_gl0_inv 4409 4410 - If CU wavefront execution mode, omit. 4411 - Ensures that 4412 following 4413 loads will not see 4414 stale data. 4415 4416 load atomic acquire - agent - global 1. buffer/global/flat_load 1. buffer/global_load 4417 - system glc=1 glc=1 dlc=1 4418 2. s_waitcnt vmcnt(0) 2. s_waitcnt vmcnt(0) 4419 4420 - Must happen before - Must happen before 4421 following following 4422 buffer_wbinvl1_vol. buffer_gl*_inv. 4423 - Ensures the load - Ensures the load 4424 has completed has completed 4425 before invalidating before invalidating 4426 the cache. the caches. 4427 4428 3. buffer_wbinvl1_vol 3. buffer_gl0_inv; 4429 buffer_gl1_inv 4430 4431 - Must happen before - Must happen before 4432 any following any following 4433 global/generic global/generic 4434 load/load load/load 4435 atomic/atomicrmw. atomic/atomicrmw. 4436 - Ensures that - Ensures that 4437 following following 4438 loads will not see loads will not see 4439 stale global data. stale global data. 4440 4441 load atomic acquire - agent - generic 1. flat_load glc=1 1. flat_load glc=1 dlc=1 4442 - system 2. s_waitcnt vmcnt(0) & 2. s_waitcnt vmcnt(0) & 4443 lgkmcnt(0) lgkmcnt(0) 4444 4445 - If OpenCL omit - If OpenCL omit 4446 lgkmcnt(0). lgkmcnt(0). 4447 - Must happen before - Must happen before 4448 following following 4449 buffer_wbinvl1_vol. buffer_gl*_invl. 4450 - Ensures the flat_load - Ensures the flat_load 4451 has completed has completed 4452 before invalidating before invalidating 4453 the cache. the caches. 4454 4455 3. buffer_wbinvl1_vol 3. buffer_gl0_inv; 4456 buffer_gl1_inv 4457 4458 - Must happen before - Must happen before 4459 any following any following 4460 global/generic global/generic 4461 load/load load/load 4462 atomic/atomicrmw. atomic/atomicrmw. 4463 - Ensures that - Ensures that 4464 following loads following loads 4465 will not see stale will not see stale 4466 global data. global data. 4467 4468 atomicrmw acquire - singlethread - global 1. buffer/global/ds/flat_atomic 1. buffer/global/ds/flat_atomic 4469 - wavefront - local 4470 - generic 4471 atomicrmw acquire - workgroup - global 1. buffer/global/flat_atomic 1. buffer/global_atomic 4472 2. s_waitcnt vm/vscnt(0) 4473 4474 - If CU wavefront execution mode, omit. 4475 - Use vmcnt if atomic with 4476 return and vscnt if atomic 4477 with no-return. 4478 - Must happen before 4479 the following buffer_gl0_inv 4480 and before any following 4481 global/generic 4482 load/load 4483 atomic/store/store 4484 atomic/atomicrmw. 4485 4486 3. buffer_gl0_inv 4487 4488 - If CU wavefront execution mode, omit. 4489 - Ensures that 4490 following 4491 loads will not see 4492 stale data. 4493 4494 atomicrmw acquire - workgroup - local 1. ds_atomic 1. ds_atomic 4495 2. waitcnt lgkmcnt(0) 2. waitcnt lgkmcnt(0) 4496 4497 - If OpenCL, omit. - If OpenCL, omit. 4498 - Must happen before - Must happen before 4499 any following the following 4500 global/generic buffer_gl0_inv. 4501 load/load 4502 atomic/store/store 4503 atomic/atomicrmw. 4504 - Ensures any - Ensures any 4505 following global following global 4506 data read is no data read is no 4507 older than the older than the 4508 atomicrmw value atomicrmw value 4509 being acquired. being acquired. 4510 4511 3. buffer_gl0_inv 4512 4513 - If OpenCL omit. 4514 - Ensures that 4515 following 4516 loads will not see 4517 stale data. 4518 4519 atomicrmw acquire - workgroup - generic 1. flat_atomic 1. flat_atomic 4520 2. waitcnt lgkmcnt(0) 2. waitcnt lgkmcnt(0) & 4521 vm/vscnt(0) 4522 4523 - If CU wavefront execution mode, omit vm/vscnt. 4524 - If OpenCL, omit. - If OpenCL, omit 4525 waitcnt lgkmcnt(0).. 4526 - Use vmcnt if atomic with 4527 return and vscnt if atomic 4528 with no-return. 4529 waitcnt lgkmcnt(0). 4530 - Must happen before - Must happen before 4531 any following the following 4532 global/generic buffer_gl0_inv. 4533 load/load 4534 atomic/store/store 4535 atomic/atomicrmw. 4536 - Ensures any - Ensures any 4537 following global following global 4538 data read is no data read is no 4539 older than the older than the 4540 atomicrmw value atomicrmw value 4541 being acquired. being acquired. 4542 4543 3. buffer_gl0_inv 4544 4545 - If CU wavefront execution mode, omit. 4546 - Ensures that 4547 following 4548 loads will not see 4549 stale data. 4550 4551 atomicrmw acquire - agent - global 1. buffer/global/flat_atomic 1. buffer/global_atomic 4552 - system 2. s_waitcnt vmcnt(0) 2. s_waitcnt vm/vscnt(0) 4553 4554 - Use vmcnt if atomic with 4555 return and vscnt if atomic 4556 with no-return. 4557 waitcnt lgkmcnt(0). 4558 - Must happen before - Must happen before 4559 following following 4560 buffer_wbinvl1_vol. buffer_gl*_inv. 4561 - Ensures the - Ensures the 4562 atomicrmw has atomicrmw has 4563 completed before completed before 4564 invalidating the invalidating the 4565 cache. caches. 4566 4567 3. buffer_wbinvl1_vol 3. buffer_gl0_inv; 4568 buffer_gl1_inv 4569 4570 - Must happen before - Must happen before 4571 any following any following 4572 global/generic global/generic 4573 load/load load/load 4574 atomic/atomicrmw. atomic/atomicrmw. 4575 - Ensures that - Ensures that 4576 following loads following loads 4577 will not see stale will not see stale 4578 global data. global data. 4579 4580 atomicrmw acquire - agent - generic 1. flat_atomic 1. flat_atomic 4581 - system 2. s_waitcnt vmcnt(0) & 2. s_waitcnt vm/vscnt(0) & 4582 lgkmcnt(0) lgkmcnt(0) 4583 4584 - If OpenCL, omit - If OpenCL, omit 4585 lgkmcnt(0). lgkmcnt(0). 4586 - Use vmcnt if atomic with 4587 return and vscnt if atomic 4588 with no-return. 4589 - Must happen before - Must happen before 4590 following following 4591 buffer_wbinvl1_vol. buffer_gl*_inv. 4592 - Ensures the - Ensures the 4593 atomicrmw has atomicrmw has 4594 completed before completed before 4595 invalidating the invalidating the 4596 cache. caches. 4597 4598 3. buffer_wbinvl1_vol 3. buffer_gl0_inv; 4599 buffer_gl1_inv 4600 4601 - Must happen before - Must happen before 4602 any following any following 4603 global/generic global/generic 4604 load/load load/load 4605 atomic/atomicrmw. atomic/atomicrmw. 4606 - Ensures that - Ensures that 4607 following loads following loads 4608 will not see stale will not see stale 4609 global data. global data. 4610 4611 fence acquire - singlethread *none* *none* *none* 4612 - wavefront 4613 fence acquire - workgroup *none* 1. s_waitcnt lgkmcnt(0) 1. s_waitcnt lgkmcnt(0) & 4614 vmcnt(0) & vscnt(0) 4615 4616 - If CU wavefront execution mode, omit vmcnt and 4617 vscnt. 4618 - If OpenCL and - If OpenCL and 4619 address space is address space is 4620 not generic, omit. not generic, omit 4621 lgkmcnt(0). 4622 - If OpenCL and 4623 address space is 4624 local, omit 4625 vmcnt(0) and vscnt(0). 4626 - However, since LLVM - However, since LLVM 4627 currently has no currently has no 4628 address space on address space on 4629 the fence need to the fence need to 4630 conservatively conservatively 4631 always generate. If always generate. If 4632 fence had an fence had an 4633 address space then address space then 4634 set to address set to address 4635 space of OpenCL space of OpenCL 4636 fence flag, or to fence flag, or to 4637 generic if both generic if both 4638 local and global local and global 4639 flags are flags are 4640 specified. specified. 4641 - Must happen after 4642 any preceding 4643 local/generic load 4644 atomic/atomicrmw 4645 with an equal or 4646 wider sync scope 4647 and memory ordering 4648 stronger than 4649 unordered (this is 4650 termed the 4651 fence-paired-atomic). 4652 - Must happen before 4653 any following 4654 global/generic 4655 load/load 4656 atomic/store/store 4657 atomic/atomicrmw. 4658 - Ensures any 4659 following global 4660 data read is no 4661 older than the 4662 value read by the 4663 fence-paired-atomic. 4664 - Could be split into 4665 separate s_waitcnt 4666 vmcnt(0), s_waitcnt 4667 vscnt(0) and s_waitcnt 4668 lgkmcnt(0) to allow 4669 them to be 4670 independently moved 4671 according to the 4672 following rules. 4673 - s_waitcnt vmcnt(0) 4674 must happen after 4675 any preceding 4676 global/generic load 4677 atomic/ 4678 atomicrmw-with-return-value 4679 with an equal or 4680 wider sync scope 4681 and memory ordering 4682 stronger than 4683 unordered (this is 4684 termed the 4685 fence-paired-atomic). 4686 - s_waitcnt vscnt(0) 4687 must happen after 4688 any preceding 4689 global/generic 4690 atomicrmw-no-return-value 4691 with an equal or 4692 wider sync scope 4693 and memory ordering 4694 stronger than 4695 unordered (this is 4696 termed the 4697 fence-paired-atomic). 4698 - s_waitcnt lgkmcnt(0) 4699 must happen after 4700 any preceding 4701 local/generic load 4702 atomic/atomicrmw 4703 with an equal or 4704 wider sync scope 4705 and memory ordering 4706 stronger than 4707 unordered (this is 4708 termed the 4709 fence-paired-atomic). 4710 - Must happen before 4711 the following 4712 buffer_gl0_inv. 4713 - Ensures that the 4714 fence-paired atomic 4715 has completed 4716 before invalidating 4717 the 4718 cache. Therefore 4719 any following 4720 locations read must 4721 be no older than 4722 the value read by 4723 the 4724 fence-paired-atomic. 4725 4726 3. buffer_gl0_inv 4727 4728 - If CU wavefront execution mode, omit. 4729 - Ensures that 4730 following 4731 loads will not see 4732 stale data. 4733 4734 fence acquire - agent *none* 1. s_waitcnt lgkmcnt(0) & 1. s_waitcnt lgkmcnt(0) & 4735 - system vmcnt(0) vmcnt(0) & vscnt(0) 4736 4737 - If OpenCL and - If OpenCL and 4738 address space is address space is 4739 not generic, omit not generic, omit 4740 lgkmcnt(0). lgkmcnt(0). 4741 - If OpenCL and 4742 address space is 4743 local, omit 4744 vmcnt(0) and vscnt(0). 4745 - However, since LLVM - However, since LLVM 4746 currently has no currently has no 4747 address space on address space on 4748 the fence need to the fence need to 4749 conservatively conservatively 4750 always generate always generate 4751 (see comment for (see comment for 4752 previous fence). previous fence). 4753 - Could be split into 4754 separate s_waitcnt 4755 vmcnt(0) and 4756 s_waitcnt 4757 lgkmcnt(0) to allow 4758 them to be 4759 independently moved 4760 according to the 4761 following rules. 4762 - s_waitcnt vmcnt(0) 4763 must happen after 4764 any preceding 4765 global/generic load 4766 atomic/atomicrmw 4767 with an equal or 4768 wider sync scope 4769 and memory ordering 4770 stronger than 4771 unordered (this is 4772 termed the 4773 fence-paired-atomic). 4774 - s_waitcnt lgkmcnt(0) 4775 must happen after 4776 any preceding 4777 local/generic load 4778 atomic/atomicrmw 4779 with an equal or 4780 wider sync scope 4781 and memory ordering 4782 stronger than 4783 unordered (this is 4784 termed the 4785 fence-paired-atomic). 4786 - Must happen before 4787 the following 4788 buffer_wbinvl1_vol. 4789 - Ensures that the 4790 fence-paired atomic 4791 has completed 4792 before invalidating 4793 the 4794 cache. Therefore 4795 any following 4796 locations read must 4797 be no older than 4798 the value read by 4799 the 4800 fence-paired-atomic. 4801 - Could be split into 4802 separate s_waitcnt 4803 vmcnt(0), s_waitcnt 4804 vscnt(0) and s_waitcnt 4805 lgkmcnt(0) to allow 4806 them to be 4807 independently moved 4808 according to the 4809 following rules. 4810 - s_waitcnt vmcnt(0) 4811 must happen after 4812 any preceding 4813 global/generic load 4814 atomic/ 4815 atomicrmw-with-return-value 4816 with an equal or 4817 wider sync scope 4818 and memory ordering 4819 stronger than 4820 unordered (this is 4821 termed the 4822 fence-paired-atomic). 4823 - s_waitcnt vscnt(0) 4824 must happen after 4825 any preceding 4826 global/generic 4827 atomicrmw-no-return-value 4828 with an equal or 4829 wider sync scope 4830 and memory ordering 4831 stronger than 4832 unordered (this is 4833 termed the 4834 fence-paired-atomic). 4835 - s_waitcnt lgkmcnt(0) 4836 must happen after 4837 any preceding 4838 local/generic load 4839 atomic/atomicrmw 4840 with an equal or 4841 wider sync scope 4842 and memory ordering 4843 stronger than 4844 unordered (this is 4845 termed the 4846 fence-paired-atomic). 4847 - Must happen before 4848 the following 4849 buffer_gl*_inv. 4850 - Ensures that the 4851 fence-paired atomic 4852 has completed 4853 before invalidating 4854 the 4855 caches. Therefore 4856 any following 4857 locations read must 4858 be no older than 4859 the value read by 4860 the 4861 fence-paired-atomic. 4862 4863 2. buffer_wbinvl1_vol 2. buffer_gl0_inv; 4864 buffer_gl1_inv 4865 4866 - Must happen before any - Must happen before any 4867 following global/generic following global/generic 4868 load/load load/load 4869 atomic/store/store atomic/store/store 4870 atomic/atomicrmw. atomic/atomicrmw. 4871 - Ensures that - Ensures that 4872 following loads following loads 4873 will not see stale will not see stale 4874 global data. global data. 4875 4876 **Release Atomic** 4877 ---------------------------------------------------------------------------------------------------------------------- 4878 store atomic release - singlethread - global 1. buffer/global/ds/flat_store 1. buffer/global/ds/flat_store 4879 - wavefront - local 4880 - generic 4881 store atomic release - workgroup - global 1. s_waitcnt lgkmcnt(0) 1. s_waitcnt lgkmcnt(0) & 4882 vmcnt(0) & vscnt(0) 4883 4884 - If CU wavefront execution mode, omit vmcnt and 4885 vscnt. 4886 - If OpenCL, omit. - If OpenCL, omit 4887 lgkmcnt(0). 4888 - Must happen after 4889 any preceding 4890 local/generic 4891 load/store/load 4892 atomic/store 4893 atomic/atomicrmw. 4894 - Could be split into 4895 separate s_waitcnt 4896 vmcnt(0), s_waitcnt 4897 vscnt(0) and s_waitcnt 4898 lgkmcnt(0) to allow 4899 them to be 4900 independently moved 4901 according to the 4902 following rules. 4903 - s_waitcnt vmcnt(0) 4904 must happen after 4905 any preceding 4906 global/generic load/load 4907 atomic/ 4908 atomicrmw-with-return-value. 4909 - s_waitcnt vscnt(0) 4910 must happen after 4911 any preceding 4912 global/generic 4913 store/store 4914 atomic/ 4915 atomicrmw-no-return-value. 4916 - s_waitcnt lgkmcnt(0) 4917 must happen after 4918 any preceding 4919 local/generic 4920 load/store/load 4921 atomic/store 4922 atomic/atomicrmw. 4923 - Must happen before - Must happen before 4924 the following the following 4925 store. store. 4926 - Ensures that all - Ensures that all 4927 memory operations memory operations 4928 to local have have 4929 completed before completed before 4930 performing the performing the 4931 store that is being store that is being 4932 released. released. 4933 4934 2. buffer/global/flat_store 2. buffer/global_store 4935 store atomic release - workgroup - local 1. waitcnt vmcnt(0) & vscnt(0) 4936 4937 - If CU wavefront execution mode, omit. 4938 - If OpenCL, omit. 4939 - Could be split into 4940 separate s_waitcnt 4941 vmcnt(0) and s_waitcnt 4942 vscnt(0) to allow 4943 them to be 4944 independently moved 4945 according to the 4946 following rules. 4947 - s_waitcnt vmcnt(0) 4948 must happen after 4949 any preceding 4950 global/generic load/load 4951 atomic/ 4952 atomicrmw-with-return-value. 4953 - s_waitcnt vscnt(0) 4954 must happen after 4955 any preceding 4956 global/generic 4957 store/store atomic/ 4958 atomicrmw-no-return-value. 4959 - Must happen before 4960 the following 4961 store. 4962 - Ensures that all 4963 global memory 4964 operations have 4965 completed before 4966 performing the 4967 store that is being 4968 released. 4969 4970 1. ds_store 2. ds_store 4971 store atomic release - workgroup - generic 1. s_waitcnt lgkmcnt(0) 1. s_waitcnt lgkmcnt(0) & 4972 vmcnt(0) & vscnt(0) 4973 4974 - If CU wavefront execution mode, omit vmcnt and 4975 vscnt. 4976 - If OpenCL, omit. - If OpenCL, omit 4977 lgkmcnt(0). 4978 - Must happen after 4979 any preceding 4980 local/generic 4981 load/store/load 4982 atomic/store 4983 atomic/atomicrmw. 4984 - Could be split into 4985 separate s_waitcnt 4986 vmcnt(0), s_waitcnt 4987 vscnt(0) and s_waitcnt 4988 lgkmcnt(0) to allow 4989 them to be 4990 independently moved 4991 according to the 4992 following rules. 4993 - s_waitcnt vmcnt(0) 4994 must happen after 4995 any preceding 4996 global/generic load/load 4997 atomic/ 4998 atomicrmw-with-return-value. 4999 - s_waitcnt vscnt(0) 5000 must happen after 5001 any preceding 5002 global/generic 5003 store/store 5004 atomic/ 5005 atomicrmw-no-return-value. 5006 - s_waitcnt lgkmcnt(0) 5007 must happen after 5008 any preceding 5009 local/generic load/store/load 5010 atomic/store atomic/atomicrmw. 5011 - Must happen before - Must happen before 5012 the following the following 5013 store. store. 5014 - Ensures that all - Ensures that all 5015 memory operations memory operations 5016 to local have have 5017 completed before completed before 5018 performing the performing the 5019 store that is being store that is being 5020 released. released. 5021 5022 2. flat_store 2. flat_store 5023 store atomic release - agent - global 1. s_waitcnt lgkmcnt(0) & 1. s_waitcnt lgkmcnt(0) & 5024 - system - generic vmcnt(0) vmcnt(0) & vscnt(0) 5025 5026 - If OpenCL, omit - If OpenCL, omit 5027 lgkmcnt(0). lgkmcnt(0). 5028 - Could be split into - Could be split into 5029 separate s_waitcnt separate s_waitcnt 5030 vmcnt(0) and vmcnt(0), s_waitcnt vscnt(0) 5031 s_waitcnt and s_waitcnt 5032 lgkmcnt(0) to allow lgkmcnt(0) to allow 5033 them to be them to be 5034 independently moved independently moved 5035 according to the according to the 5036 following rules. following rules. 5037 - s_waitcnt vmcnt(0) - s_waitcnt vmcnt(0) 5038 must happen after must happen after 5039 any preceding any preceding 5040 global/generic global/generic 5041 load/store/load load/load 5042 atomic/store atomic/ 5043 atomic/atomicrmw. atomicrmw-with-return-value. 5044 - s_waitcnt vscnt(0) 5045 must happen after 5046 any preceding 5047 global/generic 5048 store/store atomic/ 5049 atomicrmw-no-return-value. 5050 - s_waitcnt lgkmcnt(0) - s_waitcnt lgkmcnt(0) 5051 must happen after must happen after 5052 any preceding any preceding 5053 local/generic local/generic 5054 load/store/load load/store/load 5055 atomic/store atomic/store 5056 atomic/atomicrmw. atomic/atomicrmw. 5057 - Must happen before - Must happen before 5058 the following the following 5059 store. store. 5060 - Ensures that all - Ensures that all 5061 memory operations memory operations 5062 to memory have to memory have 5063 completed before completed before 5064 performing the performing the 5065 store that is being store that is being 5066 released. released. 5067 5068 2. buffer/global/ds/flat_store 2. buffer/global/ds/flat_store 5069 atomicrmw release - singlethread - global 1. buffer/global/ds/flat_atomic 1. buffer/global/ds/flat_atomic 5070 - wavefront - local 5071 - generic 5072 atomicrmw release - workgroup - global 1. s_waitcnt lgkmcnt(0) 1. s_waitcnt lgkmcnt(0) & 5073 vmcnt(0) & vscnt(0) 5074 5075 - If CU wavefront execution mode, omit vmcnt and 5076 vscnt. 5077 - If OpenCL, omit. 5078 5079 - Must happen after 5080 any preceding 5081 local/generic 5082 load/store/load 5083 atomic/store 5084 atomic/atomicrmw. 5085 - Could be split into 5086 separate s_waitcnt 5087 vmcnt(0), s_waitcnt 5088 vscnt(0) and s_waitcnt 5089 lgkmcnt(0) to allow 5090 them to be 5091 independently moved 5092 according to the 5093 following rules. 5094 - s_waitcnt vmcnt(0) 5095 must happen after 5096 any preceding 5097 global/generic load/load 5098 atomic/ 5099 atomicrmw-with-return-value. 5100 - s_waitcnt vscnt(0) 5101 must happen after 5102 any preceding 5103 global/generic 5104 store/store 5105 atomic/ 5106 atomicrmw-no-return-value. 5107 - s_waitcnt lgkmcnt(0) 5108 must happen after 5109 any preceding 5110 local/generic 5111 load/store/load 5112 atomic/store 5113 atomic/atomicrmw. 5114 - Must happen before - Must happen before 5115 the following the following 5116 atomicrmw. atomicrmw. 5117 - Ensures that all - Ensures that all 5118 memory operations memory operations 5119 to local have have 5120 completed before completed before 5121 performing the performing the 5122 atomicrmw that is atomicrmw that is 5123 being released. being released. 5124 5125 2. buffer/global/flat_atomic 2. buffer/global_atomic 5126 atomicrmw release - workgroup - local 1. waitcnt vmcnt(0) & vscnt(0) 5127 5128 - If CU wavefront execution mode, omit. 5129 - If OpenCL, omit. 5130 - Could be split into 5131 separate s_waitcnt 5132 vmcnt(0) and s_waitcnt 5133 vscnt(0) to allow 5134 them to be 5135 independently moved 5136 according to the 5137 following rules. 5138 - s_waitcnt vmcnt(0) 5139 must happen after 5140 any preceding 5141 global/generic load/load 5142 atomic/ 5143 atomicrmw-with-return-value. 5144 - s_waitcnt vscnt(0) 5145 must happen after 5146 any preceding 5147 global/generic 5148 store/store atomic/ 5149 atomicrmw-no-return-value. 5150 - Must happen before 5151 the following 5152 store. 5153 - Ensures that all 5154 global memory 5155 operations have 5156 completed before 5157 performing the 5158 store that is being 5159 released. 5160 5161 1. ds_atomic 2. ds_atomic 5162 atomicrmw release - workgroup - generic 1. s_waitcnt lgkmcnt(0) 1. s_waitcnt lgkmcnt(0) & 5163 vmcnt(0) & vscnt(0) 5164 5165 - If CU wavefront execution mode, omit vmcnt and 5166 vscnt. 5167 - If OpenCL, omit. - If OpenCL, omit 5168 waitcnt lgkmcnt(0). 5169 - Must happen after 5170 any preceding 5171 local/generic 5172 load/store/load 5173 atomic/store 5174 atomic/atomicrmw. 5175 - Could be split into 5176 separate s_waitcnt 5177 vmcnt(0), s_waitcnt 5178 vscnt(0) and s_waitcnt 5179 lgkmcnt(0) to allow 5180 them to be 5181 independently moved 5182 according to the 5183 following rules. 5184 - s_waitcnt vmcnt(0) 5185 must happen after 5186 any preceding 5187 global/generic load/load 5188 atomic/ 5189 atomicrmw-with-return-value. 5190 - s_waitcnt vscnt(0) 5191 must happen after 5192 any preceding 5193 global/generic 5194 store/store 5195 atomic/ 5196 atomicrmw-no-return-value. 5197 - s_waitcnt lgkmcnt(0) 5198 must happen after 5199 any preceding 5200 local/generic load/store/load 5201 atomic/store atomic/atomicrmw. 5202 - Must happen before - Must happen before 5203 the following the following 5204 atomicrmw. atomicrmw. 5205 - Ensures that all - Ensures that all 5206 memory operations memory operations 5207 to local have have 5208 completed before completed before 5209 performing the performing the 5210 atomicrmw that is atomicrmw that is 5211 being released. being released. 5212 5213 2. flat_atomic 2. flat_atomic 5214 atomicrmw release - agent - global 1. s_waitcnt lgkmcnt(0) & 1. s_waitcnt lkkmcnt(0) & 5215 - system - generic vmcnt(0) vmcnt(0) & vscnt(0) 5216 5217 - If OpenCL, omit - If OpenCL, omit 5218 lgkmcnt(0). lgkmcnt(0). 5219 - Could be split into - Could be split into 5220 separate s_waitcnt separate s_waitcnt 5221 vmcnt(0) and vmcnt(0), s_waitcnt 5222 s_waitcnt vscnt(0) and s_waitcnt 5223 lgkmcnt(0) to allow lgkmcnt(0) to allow 5224 them to be them to be 5225 independently moved independently moved 5226 according to the according to the 5227 following rules. following rules. 5228 - s_waitcnt vmcnt(0) - s_waitcnt vmcnt(0) 5229 must happen after must happen after 5230 any preceding any preceding 5231 global/generic global/generic 5232 load/store/load load/load atomic/ 5233 atomic/store atomicrmw-with-return-value. 5234 atomic/atomicrmw. 5235 - s_waitcnt vscnt(0) 5236 must happen after 5237 any preceding 5238 global/generic 5239 store/store atomic/ 5240 atomicrmw-no-return-value. 5241 - s_waitcnt lgkmcnt(0) - s_waitcnt lgkmcnt(0) 5242 must happen after must happen after 5243 any preceding any preceding 5244 local/generic local/generic 5245 load/store/load load/store/load 5246 atomic/store atomic/store 5247 atomic/atomicrmw. atomic/atomicrmw. 5248 - Must happen before - Must happen before 5249 the following the following 5250 atomicrmw. atomicrmw. 5251 - Ensures that all - Ensures that all 5252 memory operations memory operations 5253 to global and local to global and local 5254 have completed have completed 5255 before performing before performing 5256 the atomicrmw that the atomicrmw that 5257 is being released. is being released. 5258 5259 2. buffer/global/ds/flat_atomic 2. buffer/global/ds/flat_atomic 5260 fence release - singlethread *none* *none* *none* 5261 - wavefront 5262 fence release - workgroup *none* 1. s_waitcnt lgkmcnt(0) 1. s_waitcnt lgkmcnt(0) & 5263 vmcnt(0) & vscnt(0) 5264 5265 - If CU wavefront execution mode, omit vmcnt and 5266 vscnt. 5267 - If OpenCL and - If OpenCL and 5268 address space is address space is 5269 not generic, omit. not generic, omit 5270 lgkmcnt(0). 5271 - If OpenCL and 5272 address space is 5273 local, omit 5274 vmcnt(0) and vscnt(0). 5275 - However, since LLVM - However, since LLVM 5276 currently has no currently has no 5277 address space on address space on 5278 the fence need to the fence need to 5279 conservatively conservatively 5280 always generate. If always generate. If 5281 fence had an fence had an 5282 address space then address space then 5283 set to address set to address 5284 space of OpenCL space of OpenCL 5285 fence flag, or to fence flag, or to 5286 generic if both generic if both 5287 local and global local and global 5288 flags are flags are 5289 specified. specified. 5290 - Must happen after 5291 any preceding 5292 local/generic 5293 load/load 5294 atomic/store/store 5295 atomic/atomicrmw. 5296 - Could be split into 5297 separate s_waitcnt 5298 vmcnt(0), s_waitcnt 5299 vscnt(0) and s_waitcnt 5300 lgkmcnt(0) to allow 5301 them to be 5302 independently moved 5303 according to the 5304 following rules. 5305 - s_waitcnt vmcnt(0) 5306 must happen after 5307 any preceding 5308 global/generic 5309 load/load 5310 atomic/ 5311 atomicrmw-with-return-value. 5312 - s_waitcnt vscnt(0) 5313 must happen after 5314 any preceding 5315 global/generic 5316 store/store atomic/ 5317 atomicrmw-no-return-value. 5318 - s_waitcnt lgkmcnt(0) 5319 must happen after 5320 any preceding 5321 local/generic 5322 load/store/load 5323 atomic/store atomic/ 5324 atomicrmw. 5325 - Must happen before - Must happen before 5326 any following store any following store 5327 atomic/atomicrmw atomic/atomicrmw 5328 with an equal or with an equal or 5329 wider sync scope wider sync scope 5330 and memory ordering and memory ordering 5331 stronger than stronger than 5332 unordered (this is unordered (this is 5333 termed the termed the 5334 fence-paired-atomic). fence-paired-atomic). 5335 - Ensures that all - Ensures that all 5336 memory operations memory operations 5337 to local have have 5338 completed before completed before 5339 performing the performing the 5340 following following 5341 fence-paired-atomic. fence-paired-atomic. 5342 5343 fence release - agent *none* 1. s_waitcnt lgkmcnt(0) & 1. s_waitcnt lgkmcnt(0) & 5344 - system vmcnt(0) vmcnt(0) & vscnt(0) 5345 5346 - If OpenCL and - If OpenCL and 5347 address space is address space is 5348 not generic, omit not generic, omit 5349 lgkmcnt(0). lgkmcnt(0). 5350 - If OpenCL and - If OpenCL and 5351 address space is address space is 5352 local, omit local, omit 5353 vmcnt(0). vmcnt(0) and vscnt(0). 5354 - However, since LLVM - However, since LLVM 5355 currently has no currently has no 5356 address space on address space on 5357 the fence need to the fence need to 5358 conservatively conservatively 5359 always generate. If always generate. If 5360 fence had an fence had an 5361 address space then address space then 5362 set to address set to address 5363 space of OpenCL space of OpenCL 5364 fence flag, or to fence flag, or to 5365 generic if both generic if both 5366 local and global local and global 5367 flags are flags are 5368 specified. specified. 5369 - Could be split into - Could be split into 5370 separate s_waitcnt separate s_waitcnt 5371 vmcnt(0) and vmcnt(0), s_waitcnt 5372 s_waitcnt vscnt(0) and s_waitcnt 5373 lgkmcnt(0) to allow lgkmcnt(0) to allow 5374 them to be them to be 5375 independently moved independently moved 5376 according to the according to the 5377 following rules. following rules. 5378 - s_waitcnt vmcnt(0) - s_waitcnt vmcnt(0) 5379 must happen after must happen after 5380 any preceding any preceding 5381 global/generic global/generic 5382 load/store/load load/load atomic/ 5383 atomic/store atomicrmw-with-return-value. 5384 atomic/atomicrmw. 5385 - s_waitcnt vscnt(0) 5386 must happen after 5387 any preceding 5388 global/generic 5389 store/store atomic/ 5390 atomicrmw-no-return-value. 5391 - s_waitcnt lgkmcnt(0) - s_waitcnt lgkmcnt(0) 5392 must happen after must happen after 5393 any preceding any preceding 5394 local/generic local/generic 5395 load/store/load load/store/load 5396 atomic/store atomic/store 5397 atomic/atomicrmw. atomic/atomicrmw. 5398 - Must happen before - Must happen before 5399 any following store any following store 5400 atomic/atomicrmw atomic/atomicrmw 5401 with an equal or with an equal or 5402 wider sync scope wider sync scope 5403 and memory ordering and memory ordering 5404 stronger than stronger than 5405 unordered (this is unordered (this is 5406 termed the termed the 5407 fence-paired-atomic). fence-paired-atomic). 5408 - Ensures that all - Ensures that all 5409 memory operations memory operations 5410 have have 5411 completed before completed before 5412 performing the performing the 5413 following following 5414 fence-paired-atomic. fence-paired-atomic. 5415 5416 **Acquire-Release Atomic** 5417 ---------------------------------------------------------------------------------------------------------------------- 5418 atomicrmw acq_rel - singlethread - global 1. buffer/global/ds/flat_atomic 1. buffer/global/ds/flat_atomic 5419 - wavefront - local 5420 - generic 5421 atomicrmw acq_rel - workgroup - global 1. s_waitcnt lgkmcnt(0) 1. s_waitcnt lgkmcnt(0) & 5422 vmcnt(0) & vscnt(0) 5423 5424 - If CU wavefront execution mode, omit vmcnt and 5425 vscnt. 5426 - If OpenCL, omit. - If OpenCL, omit 5427 s_waitcnt lgkmcnt(0). 5428 - Must happen after - Must happen after 5429 any preceding any preceding 5430 local/generic local/generic 5431 load/store/load load/store/load 5432 atomic/store atomic/store 5433 atomic/atomicrmw. atomic/atomicrmw. 5434 - Could be split into 5435 separate s_waitcnt 5436 vmcnt(0), s_waitcnt 5437 vscnt(0) and s_waitcnt 5438 lgkmcnt(0) to allow 5439 them to be 5440 independently moved 5441 according to the 5442 following rules. 5443 - s_waitcnt vmcnt(0) 5444 must happen after 5445 any preceding 5446 global/generic load/load 5447 atomic/ 5448 atomicrmw-with-return-value. 5449 - s_waitcnt vscnt(0) 5450 must happen after 5451 any preceding 5452 global/generic 5453 store/store 5454 atomic/ 5455 atomicrmw-no-return-value. 5456 - s_waitcnt lgkmcnt(0) 5457 must happen after 5458 any preceding 5459 local/generic load/store/load 5460 atomic/store atomic/atomicrmw. 5461 - Must happen before - Must happen before 5462 the following the following 5463 atomicrmw. atomicrmw. 5464 - Ensures that all - Ensures that all 5465 memory operations memory operations 5466 to local have have 5467 completed before completed before 5468 performing the performing the 5469 atomicrmw that is atomicrmw that is 5470 being released. being released. 5471 5472 2. buffer/global/flat_atomic 2. buffer/global_atomic 5473 3. s_waitcnt vm/vscnt(0) 5474 5475 - If CU wavefront execution mode, omit vm/vscnt. 5476 - Use vmcnt if atomic with 5477 return and vscnt if atomic 5478 with no-return. 5479 waitcnt lgkmcnt(0). 5480 - Must happen before 5481 the following 5482 buffer_gl0_inv. 5483 - Ensures any 5484 following global 5485 data read is no 5486 older than the 5487 atomicrmw value 5488 being acquired. 5489 5490 4. buffer_gl0_inv 5491 5492 - If CU wavefront execution mode, omit. 5493 - Ensures that 5494 following 5495 loads will not see 5496 stale data. 5497 5498 atomicrmw acq_rel - workgroup - local 1. waitcnt vmcnt(0) & vscnt(0) 5499 5500 - If CU wavefront execution mode, omit. 5501 - If OpenCL, omit. 5502 - Could be split into 5503 separate s_waitcnt 5504 vmcnt(0) and s_waitcnt 5505 vscnt(0) to allow 5506 them to be 5507 independently moved 5508 according to the 5509 following rules. 5510 - s_waitcnt vmcnt(0) 5511 must happen after 5512 any preceding 5513 global/generic load/load 5514 atomic/ 5515 atomicrmw-with-return-value. 5516 - s_waitcnt vscnt(0) 5517 must happen after 5518 any preceding 5519 global/generic 5520 store/store atomic/ 5521 atomicrmw-no-return-value. 5522 - Must happen before 5523 the following 5524 store. 5525 - Ensures that all 5526 global memory 5527 operations have 5528 completed before 5529 performing the 5530 store that is being 5531 released. 5532 5533 1. ds_atomic 2. ds_atomic 5534 2. s_waitcnt lgkmcnt(0) 3. s_waitcnt lgkmcnt(0) 5535 5536 - If OpenCL, omit. - If OpenCL, omit. 5537 - Must happen before - Must happen before 5538 any following the following 5539 global/generic buffer_gl0_inv. 5540 load/load 5541 atomic/store/store 5542 atomic/atomicrmw. 5543 - Ensures any - Ensures any 5544 following global following global 5545 data read is no data read is no 5546 older than the load older than the load 5547 atomic value being atomic value being 5548 acquired. acquired. 5549 5550 4. buffer_gl0_inv 5551 5552 - If CU wavefront execution mode, omit. 5553 - If OpenCL omit. 5554 - Ensures that 5555 following 5556 loads will not see 5557 stale data. 5558 5559 atomicrmw acq_rel - workgroup - generic 1. s_waitcnt lgkmcnt(0) 1. s_waitcnt lgkmcnt(0) & 5560 vmcnt(0) & vscnt(0) 5561 5562 - If CU wavefront execution mode, omit vmcnt and 5563 vscnt. 5564 - If OpenCL, omit. - If OpenCL, omit 5565 waitcnt lgkmcnt(0). 5566 - Must happen after 5567 any preceding 5568 local/generic 5569 load/store/load 5570 atomic/store 5571 atomic/atomicrmw. 5572 - Could be split into 5573 separate s_waitcnt 5574 vmcnt(0), s_waitcnt 5575 vscnt(0) and s_waitcnt 5576 lgkmcnt(0) to allow 5577 them to be 5578 independently moved 5579 according to the 5580 following rules. 5581 - s_waitcnt vmcnt(0) 5582 must happen after 5583 any preceding 5584 global/generic load/load 5585 atomic/ 5586 atomicrmw-with-return-value. 5587 - s_waitcnt vscnt(0) 5588 must happen after 5589 any preceding 5590 global/generic 5591 store/store 5592 atomic/ 5593 atomicrmw-no-return-value. 5594 - s_waitcnt lgkmcnt(0) 5595 must happen after 5596 any preceding 5597 local/generic load/store/load 5598 atomic/store atomic/atomicrmw. 5599 - Must happen before - Must happen before 5600 the following the following 5601 atomicrmw. atomicrmw. 5602 - Ensures that all - Ensures that all 5603 memory operations memory operations 5604 to local have have 5605 completed before completed before 5606 performing the performing the 5607 atomicrmw that is atomicrmw that is 5608 being released. being released. 5609 5610 2. flat_atomic 2. flat_atomic 5611 3. s_waitcnt lgkmcnt(0) 3. s_waitcnt lgkmcnt(0) & 5612 vm/vscnt(0) 5613 5614 - If CU wavefront execution mode, omit vm/vscnt. 5615 - If OpenCL, omit. - If OpenCL, omit 5616 waitcnt lgkmcnt(0). 5617 - Must happen before - Must happen before 5618 any following the following 5619 global/generic buffer_gl0_inv. 5620 load/load 5621 atomic/store/store 5622 atomic/atomicrmw. 5623 - Ensures any - Ensures any 5624 following global following global 5625 data read is no data read is no 5626 older than the load older than the load 5627 atomic value being atomic value being 5628 acquired. acquired. 5629 5630 3. buffer_gl0_inv 5631 5632 - If CU wavefront execution mode, omit. 5633 - Ensures that 5634 following 5635 loads will not see 5636 stale data. 5637 5638 atomicrmw acq_rel - agent - global 1. s_waitcnt lgkmcnt(0) & 1. s_waitcnt lgkmcnt(0) & 5639 - system vmcnt(0) vmcnt(0) & vscnt(0) 5640 5641 - If OpenCL, omit - If OpenCL, omit 5642 lgkmcnt(0). lgkmcnt(0). 5643 - Could be split into - Could be split into 5644 separate s_waitcnt separate s_waitcnt 5645 vmcnt(0) and vmcnt(0), s_waitcnt 5646 s_waitcnt vscnt(0) and s_waitcnt 5647 lgkmcnt(0) to allow lgkmcnt(0) to allow 5648 them to be them to be 5649 independently moved independently moved 5650 according to the according to the 5651 following rules. following rules. 5652 - s_waitcnt vmcnt(0) - s_waitcnt vmcnt(0) 5653 must happen after must happen after 5654 any preceding any preceding 5655 global/generic global/generic 5656 load/store/load load/load atomic/ 5657 atomic/store atomicrmw-with-return-value. 5658 atomic/atomicrmw. 5659 - s_waitcnt vscnt(0) 5660 must happen after 5661 any preceding 5662 global/generic 5663 store/store atomic/ 5664 atomicrmw-no-return-value. 5665 - s_waitcnt lgkmcnt(0) - s_waitcnt lgkmcnt(0) 5666 must happen after must happen after 5667 any preceding any preceding 5668 local/generic local/generic 5669 load/store/load load/store/load 5670 atomic/store atomic/store 5671 atomic/atomicrmw. atomic/atomicrmw. 5672 - Must happen before - Must happen before 5673 the following the following 5674 atomicrmw. atomicrmw. 5675 - Ensures that all - Ensures that all 5676 memory operations memory operations 5677 to global have to global have 5678 completed before completed before 5679 performing the performing the 5680 atomicrmw that is atomicrmw that is 5681 being released. being released. 5682 5683 2. buffer/global/flat_atomic 2. buffer/global_atomic 5684 3. s_waitcnt vmcnt(0) 3. s_waitcnt vm/vscnt(0) 5685 5686 - Use vmcnt if atomic with 5687 return and vscnt if atomic 5688 with no-return. 5689 waitcnt lgkmcnt(0). 5690 - Must happen before - Must happen before 5691 following following 5692 buffer_wbinvl1_vol. buffer_gl*_inv. 5693 - Ensures the - Ensures the 5694 atomicrmw has atomicrmw has 5695 completed before completed before 5696 invalidating the invalidating the 5697 cache. caches. 5698 5699 4. buffer_wbinvl1_vol 4. buffer_gl0_inv; 5700 buffer_gl1_inv 5701 5702 - Must happen before - Must happen before 5703 any following any following 5704 global/generic global/generic 5705 load/load load/load 5706 atomic/atomicrmw. atomic/atomicrmw. 5707 - Ensures that - Ensures that 5708 following loads following loads 5709 will not see stale will not see stale 5710 global data. global data. 5711 5712 atomicrmw acq_rel - agent - generic 1. s_waitcnt lgkmcnt(0) & 1. s_waitcnt lgkmcnt(0) & 5713 - system vmcnt(0) vmcnt(0) & vscnt(0) 5714 5715 - If OpenCL, omit - If OpenCL, omit 5716 lgkmcnt(0). lgkmcnt(0). 5717 - Could be split into - Could be split into 5718 separate s_waitcnt separate s_waitcnt 5719 vmcnt(0) and vmcnt(0), s_waitcnt 5720 s_waitcnt vscnt(0) and s_waitcnt 5721 lgkmcnt(0) to allow lgkmcnt(0) to allow 5722 them to be them to be 5723 independently moved independently moved 5724 according to the according to the 5725 following rules. following rules. 5726 - s_waitcnt vmcnt(0) - s_waitcnt vmcnt(0) 5727 must happen after must happen after 5728 any preceding any preceding 5729 global/generic global/generic 5730 load/store/load load/load atomic 5731 atomic/store atomicrmw-with-return-value. 5732 atomic/atomicrmw. 5733 - s_waitcnt vscnt(0) 5734 must happen after 5735 any preceding 5736 global/generic 5737 store/store atomic/ 5738 atomicrmw-no-return-value. 5739 - s_waitcnt lgkmcnt(0) - s_waitcnt lgkmcnt(0) 5740 must happen after must happen after 5741 any preceding any preceding 5742 local/generic local/generic 5743 load/store/load load/store/load 5744 atomic/store atomic/store 5745 atomic/atomicrmw. atomic/atomicrmw. 5746 - Must happen before - Must happen before 5747 the following the following 5748 atomicrmw. atomicrmw. 5749 - Ensures that all - Ensures that all 5750 memory operations memory operations 5751 to global have have 5752 completed before completed before 5753 performing the performing the 5754 atomicrmw that is atomicrmw that is 5755 being released. being released. 5756 5757 2. flat_atomic 2. flat_atomic 5758 3. s_waitcnt vmcnt(0) & 3. s_waitcnt vm/vscnt(0) & 5759 lgkmcnt(0) lgkmcnt(0) 5760 5761 - If OpenCL, omit - If OpenCL, omit 5762 lgkmcnt(0). lgkmcnt(0). 5763 - Use vmcnt if atomic with 5764 return and vscnt if atomic 5765 with no-return. 5766 - Must happen before - Must happen before 5767 following following 5768 buffer_wbinvl1_vol. buffer_gl*_inv. 5769 - Ensures the - Ensures the 5770 atomicrmw has atomicrmw has 5771 completed before completed before 5772 invalidating the invalidating the 5773 cache. caches. 5774 5775 4. buffer_wbinvl1_vol 4. buffer_gl0_inv; 5776 buffer_gl1_inv 5777 5778 - Must happen before - Must happen before 5779 any following any following 5780 global/generic global/generic 5781 load/load load/load 5782 atomic/atomicrmw. atomic/atomicrmw. 5783 - Ensures that - Ensures that 5784 following loads following loads 5785 will not see stale will not see stale 5786 global data. global data. 5787 5788 fence acq_rel - singlethread *none* *none* *none* 5789 - wavefront 5790 fence acq_rel - workgroup *none* 1. s_waitcnt lgkmcnt(0) 1. s_waitcnt lgkmcnt(0) & 5791 vmcnt(0) & vscnt(0) 5792 5793 - If CU wavefront execution mode, omit vmcnt and 5794 vscnt. 5795 - If OpenCL and - If OpenCL and 5796 address space is address space is 5797 not generic, omit. not generic, omit 5798 lgkmcnt(0). 5799 - If OpenCL and 5800 address space is 5801 local, omit 5802 vmcnt(0) and vscnt(0). 5803 - However, - However, 5804 since LLVM since LLVM 5805 currently has no currently has no 5806 address space on address space on 5807 the fence need to the fence need to 5808 conservatively conservatively 5809 always generate always generate 5810 (see comment for (see comment for 5811 previous fence). previous fence). 5812 - Must happen after 5813 any preceding 5814 local/generic 5815 load/load 5816 atomic/store/store 5817 atomic/atomicrmw. 5818 - Could be split into 5819 separate s_waitcnt 5820 vmcnt(0), s_waitcnt 5821 vscnt(0) and s_waitcnt 5822 lgkmcnt(0) to allow 5823 them to be 5824 independently moved 5825 according to the 5826 following rules. 5827 - s_waitcnt vmcnt(0) 5828 must happen after 5829 any preceding 5830 global/generic 5831 load/load 5832 atomic/ 5833 atomicrmw-with-return-value. 5834 - s_waitcnt vscnt(0) 5835 must happen after 5836 any preceding 5837 global/generic 5838 store/store atomic/ 5839 atomicrmw-no-return-value. 5840 - s_waitcnt lgkmcnt(0) 5841 must happen after 5842 any preceding 5843 local/generic 5844 load/store/load 5845 atomic/store atomic/ 5846 atomicrmw. 5847 - Must happen before - Must happen before 5848 any following any following 5849 global/generic global/generic 5850 load/load load/load 5851 atomic/store/store atomic/store/store 5852 atomic/atomicrmw. atomic/atomicrmw. 5853 - Ensures that all - Ensures that all 5854 memory operations memory operations 5855 to local have have 5856 completed before completed before 5857 performing any performing any 5858 following global following global 5859 memory operations. memory operations. 5860 - Ensures that the - Ensures that the 5861 preceding preceding 5862 local/generic load local/generic load 5863 atomic/atomicrmw atomic/atomicrmw 5864 with an equal or with an equal or 5865 wider sync scope wider sync scope 5866 and memory ordering and memory ordering 5867 stronger than stronger than 5868 unordered (this is unordered (this is 5869 termed the termed the 5870 acquire-fence-paired-atomic acquire-fence-paired-atomic 5871 ) has completed ) has completed 5872 before following before following 5873 global memory global memory 5874 operations. This operations. This 5875 satisfies the satisfies the 5876 requirements of requirements of 5877 acquire. acquire. 5878 - Ensures that all - Ensures that all 5879 previous memory previous memory 5880 operations have operations have 5881 completed before a completed before a 5882 following following 5883 local/generic store local/generic store 5884 atomic/atomicrmw atomic/atomicrmw 5885 with an equal or with an equal or 5886 wider sync scope wider sync scope 5887 and memory ordering and memory ordering 5888 stronger than stronger than 5889 unordered (this is unordered (this is 5890 termed the termed the 5891 release-fence-paired-atomic release-fence-paired-atomic 5892 ). This satisfies the ). This satisfies the 5893 requirements of requirements of 5894 release. release. 5895 - Must happen before 5896 the following 5897 buffer_gl0_inv. 5898 - Ensures that the 5899 acquire-fence-paired 5900 atomic has completed 5901 before invalidating 5902 the 5903 cache. Therefore 5904 any following 5905 locations read must 5906 be no older than 5907 the value read by 5908 the 5909 acquire-fence-paired-atomic. 5910 5911 3. buffer_gl0_inv 5912 5913 - If CU wavefront execution mode, omit. 5914 - Ensures that 5915 following 5916 loads will not see 5917 stale data. 5918 5919 fence acq_rel - agent *none* 1. s_waitcnt lgkmcnt(0) & 1. s_waitcnt lgkmcnt(0) & 5920 - system vmcnt(0) vmcnt(0) & vscnt(0) 5921 5922 - If OpenCL and - If OpenCL and 5923 address space is address space is 5924 not generic, omit not generic, omit 5925 lgkmcnt(0). lgkmcnt(0). 5926 - If OpenCL and 5927 address space is 5928 local, omit 5929 vmcnt(0) and vscnt(0). 5930 - However, since LLVM - However, since LLVM 5931 currently has no currently has no 5932 address space on address space on 5933 the fence need to the fence need to 5934 conservatively conservatively 5935 always generate always generate 5936 (see comment for (see comment for 5937 previous fence). previous fence). 5938 - Could be split into - Could be split into 5939 separate s_waitcnt separate s_waitcnt 5940 vmcnt(0) and vmcnt(0), s_waitcnt 5941 s_waitcnt vscnt(0) and s_waitcnt 5942 lgkmcnt(0) to allow lgkmcnt(0) to allow 5943 them to be them to be 5944 independently moved independently moved 5945 according to the according to the 5946 following rules. following rules. 5947 - s_waitcnt vmcnt(0) - s_waitcnt vmcnt(0) 5948 must happen after must happen after 5949 any preceding any preceding 5950 global/generic global/generic 5951 load/store/load load/load 5952 atomic/store atomic/ 5953 atomic/atomicrmw. atomicrmw-with-return-value. 5954 - s_waitcnt vscnt(0) 5955 must happen after 5956 any preceding 5957 global/generic 5958 store/store atomic/ 5959 atomicrmw-no-return-value. 5960 - s_waitcnt lgkmcnt(0) - s_waitcnt lgkmcnt(0) 5961 must happen after must happen after 5962 any preceding any preceding 5963 local/generic local/generic 5964 load/store/load load/store/load 5965 atomic/store atomic/store 5966 atomic/atomicrmw. atomic/atomicrmw. 5967 - Must happen before - Must happen before 5968 the following the following 5969 buffer_wbinvl1_vol. buffer_gl*_inv. 5970 - Ensures that the - Ensures that the 5971 preceding preceding 5972 global/local/generic global/local/generic 5973 load load 5974 atomic/atomicrmw atomic/atomicrmw 5975 with an equal or with an equal or 5976 wider sync scope wider sync scope 5977 and memory ordering and memory ordering 5978 stronger than stronger than 5979 unordered (this is unordered (this is 5980 termed the termed the 5981 acquire-fence-paired-atomic acquire-fence-paired-atomic 5982 ) has completed ) has completed 5983 before invalidating before invalidating 5984 the cache. This the caches. This 5985 satisfies the satisfies the 5986 requirements of requirements of 5987 acquire. acquire. 5988 - Ensures that all - Ensures that all 5989 previous memory previous memory 5990 operations have operations have 5991 completed before a completed before a 5992 following following 5993 global/local/generic global/local/generic 5994 store store 5995 atomic/atomicrmw atomic/atomicrmw 5996 with an equal or with an equal or 5997 wider sync scope wider sync scope 5998 and memory ordering and memory ordering 5999 stronger than stronger than 6000 unordered (this is unordered (this is 6001 termed the termed the 6002 release-fence-paired-atomic release-fence-paired-atomic 6003 ). This satisfies the ). This satisfies the 6004 requirements of requirements of 6005 release. release. 6006 6007 2. buffer_wbinvl1_vol 2. buffer_gl0_inv; 6008 buffer_gl1_inv 6009 6010 - Must happen before - Must happen before 6011 any following any following 6012 global/generic global/generic 6013 load/load load/load 6014 atomic/store/store atomic/store/store 6015 atomic/atomicrmw. atomic/atomicrmw. 6016 - Ensures that - Ensures that 6017 following loads following loads 6018 will not see stale will not see stale 6019 global data. This global data. This 6020 satisfies the satisfies the 6021 requirements of requirements of 6022 acquire. acquire. 6023 6024 **Sequential Consistent Atomic** 6025 ---------------------------------------------------------------------------------------------------------------------- 6026 load atomic seq_cst - singlethread - global *Same as corresponding *Same as corresponding 6027 - wavefront - local load atomic acquire, load atomic acquire, 6028 - generic except must generated except must generated 6029 all instructions even all instructions even 6030 for OpenCL.* for OpenCL.* 6031 load atomic seq_cst - workgroup - global 1. s_waitcnt lgkmcnt(0) 1. s_waitcnt lgkmcnt(0) & 6032 - generic vmcnt(0) & vscnt(0) 6033 6034 - If CU wavefront execution mode, omit vmcnt and 6035 vscnt. 6036 - Could be split into 6037 separate s_waitcnt 6038 vmcnt(0), s_waitcnt 6039 vscnt(0) and s_waitcnt 6040 lgkmcnt(0) to allow 6041 them to be 6042 independently moved 6043 according to the 6044 following rules. 6045 - Must - waitcnt lgkmcnt(0) must 6046 happen after happen after 6047 preceding preceding 6048 global/generic load local load 6049 atomic/store atomic/store 6050 atomic/atomicrmw atomic/atomicrmw 6051 with memory with memory 6052 ordering of seq_cst ordering of seq_cst 6053 and with equal or and with equal or 6054 wider sync scope. wider sync scope. 6055 (Note that seq_cst (Note that seq_cst 6056 fences have their fences have their 6057 own s_waitcnt own s_waitcnt 6058 lgkmcnt(0) and so do lgkmcnt(0) and so do 6059 not need to be not need to be 6060 considered.) considered.) 6061 - waitcnt vmcnt(0) 6062 Must happen after 6063 preceding 6064 global/generic load 6065 atomic/ 6066 atomicrmw-with-return-value 6067 with memory 6068 ordering of seq_cst 6069 and with equal or 6070 wider sync scope. 6071 (Note that seq_cst 6072 fences have their 6073 own s_waitcnt 6074 vmcnt(0) and so do 6075 not need to be 6076 considered.) 6077 - waitcnt vscnt(0) 6078 Must happen after 6079 preceding 6080 global/generic store 6081 atomic/ 6082 atomicrmw-no-return-value 6083 with memory 6084 ordering of seq_cst 6085 and with equal or 6086 wider sync scope. 6087 (Note that seq_cst 6088 fences have their 6089 own s_waitcnt 6090 vscnt(0) and so do 6091 not need to be 6092 considered.) 6093 - Ensures any - Ensures any 6094 preceding preceding 6095 sequential sequential 6096 consistent local consistent global/local 6097 memory instructions memory instructions 6098 have completed have completed 6099 before executing before executing 6100 this sequentially this sequentially 6101 consistent consistent 6102 instruction. This instruction. This 6103 prevents reordering prevents reordering 6104 a seq_cst store a seq_cst store 6105 followed by a followed by a 6106 seq_cst load. (Note seq_cst load. (Note 6107 that seq_cst is that seq_cst is 6108 stronger than stronger than 6109 acquire/release as acquire/release as 6110 the reordering of the reordering of 6111 load acquire load acquire 6112 followed by a store followed by a store 6113 release is release is 6114 prevented by the prevented by the 6115 waitcnt of waitcnt of 6116 the release, but the release, but 6117 there is nothing there is nothing 6118 preventing a store preventing a store 6119 release followed by release followed by 6120 load acquire from load acquire from 6121 competing out of competing out of 6122 order.) order.) 6123 6124 2. *Following 2. *Following 6125 instructions same as instructions same as 6126 corresponding load corresponding load 6127 atomic acquire, atomic acquire, 6128 except must generated except must generated 6129 all instructions even all instructions even 6130 for OpenCL.* for OpenCL.* 6131 load atomic seq_cst - workgroup - local *Same as corresponding 6132 load atomic acquire, 6133 except must generated 6134 all instructions even 6135 for OpenCL.* 6136 6137 1. s_waitcnt vmcnt(0) & vscnt(0) 6138 6139 - If CU wavefront execution mode, omit. 6140 - Could be split into 6141 separate s_waitcnt 6142 vmcnt(0) and s_waitcnt 6143 vscnt(0) to allow 6144 them to be 6145 independently moved 6146 according to the 6147 following rules. 6148 - waitcnt vmcnt(0) 6149 Must happen after 6150 preceding 6151 global/generic load 6152 atomic/ 6153 atomicrmw-with-return-value 6154 with memory 6155 ordering of seq_cst 6156 and with equal or 6157 wider sync scope. 6158 (Note that seq_cst 6159 fences have their 6160 own s_waitcnt 6161 vmcnt(0) and so do 6162 not need to be 6163 considered.) 6164 - waitcnt vscnt(0) 6165 Must happen after 6166 preceding 6167 global/generic store 6168 atomic/ 6169 atomicrmw-no-return-value 6170 with memory 6171 ordering of seq_cst 6172 and with equal or 6173 wider sync scope. 6174 (Note that seq_cst 6175 fences have their 6176 own s_waitcnt 6177 vscnt(0) and so do 6178 not need to be 6179 considered.) 6180 - Ensures any 6181 preceding 6182 sequential 6183 consistent global 6184 memory instructions 6185 have completed 6186 before executing 6187 this sequentially 6188 consistent 6189 instruction. This 6190 prevents reordering 6191 a seq_cst store 6192 followed by a 6193 seq_cst load. (Note 6194 that seq_cst is 6195 stronger than 6196 acquire/release as 6197 the reordering of 6198 load acquire 6199 followed by a store 6200 release is 6201 prevented by the 6202 waitcnt of 6203 the release, but 6204 there is nothing 6205 preventing a store 6206 release followed by 6207 load acquire from 6208 competing out of 6209 order.) 6210 6211 2. *Following 6212 instructions same as 6213 corresponding load 6214 atomic acquire, 6215 except must generated 6216 all instructions even 6217 for OpenCL.* 6218 6219 load atomic seq_cst - agent - global 1. s_waitcnt lgkmcnt(0) & 1. s_waitcnt lgkmcnt(0) & 6220 - system - generic vmcnt(0) vmcnt(0) & vscnt(0) 6221 6222 - Could be split into - Could be split into 6223 separate s_waitcnt separate s_waitcnt 6224 vmcnt(0) vmcnt(0), s_waitcnt 6225 and s_waitcnt vscnt(0) and s_waitcnt 6226 lgkmcnt(0) to allow lgkmcnt(0) to allow 6227 them to be them to be 6228 independently moved independently moved 6229 according to the according to the 6230 following rules. following rules. 6231 - waitcnt lgkmcnt(0) - waitcnt lgkmcnt(0) 6232 must happen after must happen after 6233 preceding preceding 6234 global/generic load local load 6235 atomic/store atomic/store 6236 atomic/atomicrmw atomic/atomicrmw 6237 with memory with memory 6238 ordering of seq_cst ordering of seq_cst 6239 and with equal or and with equal or 6240 wider sync scope. wider sync scope. 6241 (Note that seq_cst (Note that seq_cst 6242 fences have their fences have their 6243 own s_waitcnt own s_waitcnt 6244 lgkmcnt(0) and so do lgkmcnt(0) and so do 6245 not need to be not need to be 6246 considered.) considered.) 6247 - waitcnt vmcnt(0) - waitcnt vmcnt(0) 6248 must happen after must happen after 6249 preceding preceding 6250 global/generic load global/generic load 6251 atomic/store atomic/ 6252 atomic/atomicrmw atomicrmw-with-return-value 6253 with memory with memory 6254 ordering of seq_cst ordering of seq_cst 6255 and with equal or and with equal or 6256 wider sync scope. wider sync scope. 6257 (Note that seq_cst (Note that seq_cst 6258 fences have their fences have their 6259 own s_waitcnt own s_waitcnt 6260 vmcnt(0) and so do vmcnt(0) and so do 6261 not need to be not need to be 6262 considered.) considered.) 6263 - waitcnt vscnt(0) 6264 Must happen after 6265 preceding 6266 global/generic store 6267 atomic/ 6268 atomicrmw-no-return-value 6269 with memory 6270 ordering of seq_cst 6271 and with equal or 6272 wider sync scope. 6273 (Note that seq_cst 6274 fences have their 6275 own s_waitcnt 6276 vscnt(0) and so do 6277 not need to be 6278 considered.) 6279 - Ensures any - Ensures any 6280 preceding preceding 6281 sequential sequential 6282 consistent global consistent global 6283 memory instructions memory instructions 6284 have completed have completed 6285 before executing before executing 6286 this sequentially this sequentially 6287 consistent consistent 6288 instruction. This instruction. This 6289 prevents reordering prevents reordering 6290 a seq_cst store a seq_cst store 6291 followed by a followed by a 6292 seq_cst load. (Note seq_cst load. (Note 6293 that seq_cst is that seq_cst is 6294 stronger than stronger than 6295 acquire/release as acquire/release as 6296 the reordering of the reordering of 6297 load acquire load acquire 6298 followed by a store followed by a store 6299 release is release is 6300 prevented by the prevented by the 6301 waitcnt of waitcnt of 6302 the release, but the release, but 6303 there is nothing there is nothing 6304 preventing a store preventing a store 6305 release followed by release followed by 6306 load acquire from load acquire from 6307 competing out of competing out of 6308 order.) order.) 6309 6310 2. *Following 2. *Following 6311 instructions same as instructions same as 6312 corresponding load corresponding load 6313 atomic acquire, atomic acquire, 6314 except must generated except must generated 6315 all instructions even all instructions even 6316 for OpenCL.* for OpenCL.* 6317 store atomic seq_cst - singlethread - global *Same as corresponding *Same as corresponding 6318 - wavefront - local store atomic release, store atomic release, 6319 - workgroup - generic except must generated except must generated 6320 all instructions even all instructions even 6321 for OpenCL.* for OpenCL.* 6322 store atomic seq_cst - agent - global *Same as corresponding *Same as corresponding 6323 - system - generic store atomic release, store atomic release, 6324 except must generated except must generated 6325 all instructions even all instructions even 6326 for OpenCL.* for OpenCL.* 6327 atomicrmw seq_cst - singlethread - global *Same as corresponding *Same as corresponding 6328 - wavefront - local atomicrmw acq_rel, atomicrmw acq_rel, 6329 - workgroup - generic except must generated except must generated 6330 all instructions even all instructions even 6331 for OpenCL.* for OpenCL.* 6332 atomicrmw seq_cst - agent - global *Same as corresponding *Same as corresponding 6333 - system - generic atomicrmw acq_rel, atomicrmw acq_rel, 6334 except must generated except must generated 6335 all instructions even all instructions even 6336 for OpenCL.* for OpenCL.* 6337 fence seq_cst - singlethread *none* *Same as corresponding *Same as corresponding 6338 - wavefront fence acq_rel, fence acq_rel, 6339 - workgroup except must generated except must generated 6340 - agent all instructions even all instructions even 6341 - system for OpenCL.* for OpenCL.* 6342 ============ ============ ============== ========== =============================== ================================== 6343 6344The memory order also adds the single thread optimization constrains defined in 6345table 6346:ref:`amdgpu-amdhsa-memory-model-single-thread-optimization-constraints-gfx6-gfx10-table`. 6347 6348 .. table:: AMDHSA Memory Model Single Thread Optimization Constraints GFX6-GFX10 6349 :name: amdgpu-amdhsa-memory-model-single-thread-optimization-constraints-gfx6-gfx10-table 6350 6351 ============ ============================================================== 6352 LLVM Memory Optimization Constraints 6353 Ordering 6354 ============ ============================================================== 6355 unordered *none* 6356 monotonic *none* 6357 acquire - If a load atomic/atomicrmw then no following load/load 6358 atomic/store/ store atomic/atomicrmw/fence instruction can 6359 be moved before the acquire. 6360 - If a fence then same as load atomic, plus no preceding 6361 associated fence-paired-atomic can be moved after the fence. 6362 release - If a store atomic/atomicrmw then no preceding load/load 6363 atomic/store/ store atomic/atomicrmw/fence instruction can 6364 be moved after the release. 6365 - If a fence then same as store atomic, plus no following 6366 associated fence-paired-atomic can be moved before the 6367 fence. 6368 acq_rel Same constraints as both acquire and release. 6369 seq_cst - If a load atomic then same constraints as acquire, plus no 6370 preceding sequentially consistent load atomic/store 6371 atomic/atomicrmw/fence instruction can be moved after the 6372 seq_cst. 6373 - If a store atomic then the same constraints as release, plus 6374 no following sequentially consistent load atomic/store 6375 atomic/atomicrmw/fence instruction can be moved before the 6376 seq_cst. 6377 - If an atomicrmw/fence then same constraints as acq_rel. 6378 ============ ============================================================== 6379 6380Trap Handler ABI 6381~~~~~~~~~~~~~~~~ 6382 6383For code objects generated by AMDGPU backend for HSA [HSA]_ compatible runtimes 6384(such as ROCm [AMD-ROCm]_), the runtime installs a trap handler that supports 6385the ``s_trap`` instruction with the following usage: 6386 6387 .. table:: AMDGPU Trap Handler for AMDHSA OS 6388 :name: amdgpu-trap-handler-for-amdhsa-os-table 6389 6390 =================== =============== =============== ======================= 6391 Usage Code Sequence Trap Handler Description 6392 Inputs 6393 =================== =============== =============== ======================= 6394 reserved ``s_trap 0x00`` Reserved by hardware. 6395 ``debugtrap(arg)`` ``s_trap 0x01`` ``SGPR0-1``: Reserved for HSA 6396 ``queue_ptr`` ``debugtrap`` 6397 ``VGPR0``: intrinsic (not 6398 ``arg`` implemented). 6399 ``llvm.trap`` ``s_trap 0x02`` ``SGPR0-1``: Causes dispatch to be 6400 ``queue_ptr`` terminated and its 6401 associated queue put 6402 into the error state. 6403 ``llvm.debugtrap`` ``s_trap 0x03`` - If debugger not 6404 installed then 6405 behaves as a 6406 no-operation. The 6407 trap handler is 6408 entered and 6409 immediately returns 6410 to continue 6411 execution of the 6412 wavefront. 6413 - If the debugger is 6414 installed, causes 6415 the debug trap to be 6416 reported by the 6417 debugger and the 6418 wavefront is put in 6419 the halt state until 6420 resumed by the 6421 debugger. 6422 reserved ``s_trap 0x04`` Reserved. 6423 reserved ``s_trap 0x05`` Reserved. 6424 reserved ``s_trap 0x06`` Reserved. 6425 debugger breakpoint ``s_trap 0x07`` Reserved for debugger 6426 breakpoints. 6427 reserved ``s_trap 0x08`` Reserved. 6428 reserved ``s_trap 0xfe`` Reserved. 6429 reserved ``s_trap 0xff`` Reserved. 6430 =================== =============== =============== ======================= 6431 6432.. _amdgpu-amdhsa-function-call-convention: 6433 6434Call Convention 6435~~~~~~~~~~~~~~~ 6436 6437.. note:: 6438 6439 This section is currently incomplete and has inakkuracies. It is WIP that will 6440 be updated as information is determined. 6441 6442See :ref:`amdgpu-dwarf-address-space-identifier` for information on swizzled 6443addresses. Unswizzled addresses are normal linear addresses. 6444 6445.. _amdgpu-amdhsa-function-call-convention-kernel-functions: 6446 6447Kernel Functions 6448++++++++++++++++ 6449 6450This section describes the call convention ABI for the outer kernel function. 6451 6452See :ref:`amdgpu-amdhsa-initial-kernel-execution-state` for the kernel call 6453convention. 6454 6455The following is not part of the AMDGPU kernel calling convention but describes 6456how the AMDGPU implements function calls: 6457 64581. Clang decides the kernarg layout to match the *HSA Programmer's Language 6459 Reference* [HSA]_. 6460 6461 - All structs are passed directly. 6462 - Lambda values are passed *TBA*. 6463 6464 .. TODO:: 6465 6466 - Does this really follow HSA rules? Or are structs >16 bytes passed 6467 by-value struct? 6468 - What is ABI for lambda values? 6469 64704. The kernel performs certain setup in its prolog, as described in 6471 :ref:`amdgpu-amdhsa-kernel-prolog`. 6472 6473.. _amdgpu-amdhsa-function-call-convention-non-kernel-functions: 6474 6475Non-Kernel Functions 6476++++++++++++++++++++ 6477 6478This section describes the call convention ABI for functions other than the 6479outer kernel function. 6480 6481If a kernel has function calls then scratch is always allocated and used for 6482the call stack which grows from low address to high address using the swizzled 6483scratch address space. 6484 6485On entry to a function: 6486 64871. SGPR0-3 contain a V# with the following properties (see 6488 :ref:`amdgpu-amdhsa-kernel-prolog-private-segment-buffer`): 6489 6490 * Base address pointing to the beginning of the wavefront scratch backing 6491 memory. 6492 * Swizzled with dword element size and stride of wavefront size elements. 6493 64942. The FLAT_SCRATCH register pair is setup. See 6495 :ref:`amdgpu-amdhsa-kernel-prolog-flat-scratch`. 64963. GFX6-8: M0 register set to the size of LDS in bytes. See 6497 :ref:`amdgpu-amdhsa-kernel-prolog-m0`. 64984. The EXEC register is set to the lanes active on entry to the function. 64995. MODE register: *TBD* 65006. VGPR0-31 and SGPR4-29 are used to pass function input arguments as described 6501 below. 65027. SGPR30-31 return address (RA). The code address that the function must 6503 return to when it completes. The value is undefined if the function is *no 6504 return*. 65058. SGPR32 is used for the stack pointer (SP). It is an unswizzled scratch 6506 offset relative to the beginning of the wavefront scratch backing memory. 6507 6508 The unswizzled SP can be used with buffer instructions as an unswizzled SGPR 6509 offset with the scratch V# in SGPR0-3 to access the stack in a swizzled 6510 manner. 6511 6512 The unswizzled SP value can be converted into the swizzled SP value by: 6513 6514 | swizzled SP = unswizzled SP / wavefront size 6515 6516 This may be used to obtain the private address space address of stack 6517 objects and to convert this address to a flat address by adding the flat 6518 scratch aperture base address. 6519 6520 The swizzled SP value is always 4 bytes aligned for the ``r600`` 6521 architecture and 16 byte aligned for the ``amdgcn`` architecture. 6522 6523 .. note:: 6524 6525 The ``amdgcn`` value is selected to avoid dynamic stack alignment for the 6526 OpenCL language which has the largest base type defined as 16 bytes. 6527 6528 On entry, the swizzled SP value is the address of the first function 6529 argument passed on the stack. Other stack passed arguments are positive 6530 offsets from the entry swizzled SP value. 6531 6532 The function may use positive offsets beyond the last stack passed argument 6533 for stack allocated local variables and register spill slots. If necessary, 6534 the function may align these to greater alignment than 16 bytes. After these 6535 the function may dynamically allocate space for such things as runtime sized 6536 ``alloca`` local allocations. 6537 6538 If the function calls another function, it will place any stack allocated 6539 arguments after the last local allocation and adjust SGPR32 to the address 6540 after the last local allocation. 6541 65429. All other registers are unspecified. 654310. Any necessary ``waitcnt`` has been performed to ensure memory is available 6544 to the function. 6545 6546On exit from a function: 6547 65481. VGPR0-31 and SGPR4-29 are used to pass function result arguments as 6549 described below. Any registers used are considered clobbered registers. 65502. The following registers are preserved and have the same value as on entry: 6551 6552 * FLAT_SCRATCH 6553 * EXEC 6554 * GFX6-8: M0 6555 * All SGPR registers except the clobbered registers of SGPR4-31. 6556 * VGPR40-47 6557 VGPR56-63 6558 VGPR72-79 6559 VGPR88-95 6560 VGPR104-111 6561 VGPR120-127 6562 VGPR136-143 6563 VGPR152-159 6564 VGPR168-175 6565 VGPR184-191 6566 VGPR200-207 6567 VGPR216-223 6568 VGPR232-239 6569 VGPR248-255 6570 6571 *Except the argument registers, the VGPR cloberred and the preserved 6572 registers are intermixed at regular intervals in order to 6573 get a better occupancy.* 6574 6575 For the AMDGPU backend, an inter-procedural register allocation (IPRA) 6576 optimization may mark some of clobbered SGPR and VGPR registers as 6577 preserved if it can be determined that the called function does not change 6578 their value. 6579 65802. The PC is set to the RA provided on entry. 65813. MODE register: *TBD*. 65824. All other registers are clobbered. 65835. Any necessary ``waitcnt`` has been performed to ensure memory accessed by 6584 function is available to the caller. 6585 6586.. TODO:: 6587 6588 - On gfx908 are all ACC registers clobbered? 6589 6590 - How are function results returned? The address of structured types is passed 6591 by reference, but what about other types? 6592 6593The function input arguments are made up of the formal arguments explicitly 6594declared by the source language function plus the implicit input arguments used 6595by the implementation. 6596 6597The source language input arguments are: 6598 65991. Any source language implicit ``this`` or ``self`` argument comes first as a 6600 pointer type. 66012. Followed by the function formal arguments in left to right source order. 6602 6603The source language result arguments are: 6604 66051. The function result argument. 6606 6607The source language input or result struct type arguments that are less than or 6608equal to 16 bytes, are decomposed recursively into their base type fields, and 6609each field is passed as if a separate argument. For input arguments, if the 6610called function requires the struct to be in memory, for example because its 6611address is taken, then the function body is responsible for allocating a stack 6612location and copying the field arguments into it. Clang terms this *direct 6613struct*. 6614 6615The source language input struct type arguments that are greater than 16 bytes, 6616are passed by reference. The caller is responsible for allocating a stack 6617location to make a copy of the struct value and pass the address as the input 6618argument. The called function is responsible to perform the dereference when 6619accessing the input argument. Clang terms this *by-value struct*. 6620 6621A source language result struct type argument that is greater than 16 bytes, is 6622returned by reference. The caller is responsible for allocating a stack location 6623to hold the result value and passes the address as the last input argument 6624(before the implicit input arguments). In this case there are no result 6625arguments. The called function is responsible to perform the dereference when 6626storing the result value. Clang terms this *structured return (sret)*. 6627 6628*TODO: correct the ``sret`` definition.* 6629 6630.. TODO:: 6631 6632 Is this definition correct? Or is ``sret`` only used if passing in registers, and 6633 pass as non-decomposed struct as stack argument? Or something else? Is the 6634 memory location in the caller stack frame, or a stack memory argument and so 6635 no address is passed as the caller can directly write to the argument stack 6636 location? But then the stack location is still live after return. If an 6637 argument stack location is it the first stack argument or the last one? 6638 6639Lambda argument types are treated as struct types with an implementation defined 6640set of fields. 6641 6642.. TODO:: 6643 6644 Need to specify the ABI for lambda types for AMDGPU. 6645 6646For AMDGPU backend all source language arguments (including the decomposed 6647struct type arguments) are passed in VGPRs unless marked ``inreg`` in which case 6648they are passed in SGPRs. 6649 6650The AMDGPU backend walks the function call graph from the leaves to determine 6651which implicit input arguments are used, propagating to each caller of the 6652function. The used implicit arguments are appended to the function arguments 6653after the source language arguments in the following order: 6654 6655.. TODO:: 6656 6657 Is recursion or external functions supported? 6658 66591. Work-Item ID (1 VGPR) 6660 6661 The X, Y and Z work-item ID are packed into a single VGRP with the following 6662 layout. Only fields actually used by the function are set. The other bits 6663 are undefined. 6664 6665 The values come from the initial kernel execution state. See 6666 :ref:`amdgpu-amdhsa-vgpr-register-set-up-order-table`. 6667 6668 .. table:: Work-item implicit argument layout 6669 :name: amdgpu-amdhsa-workitem-implicit-argument-layout-table 6670 6671 ======= ======= ============== 6672 Bits Size Field Name 6673 ======= ======= ============== 6674 9:0 10 bits X Work-Item ID 6675 19:10 10 bits Y Work-Item ID 6676 29:20 10 bits Z Work-Item ID 6677 31:30 2 bits Unused 6678 ======= ======= ============== 6679 66802. Dispatch Ptr (2 SGPRs) 6681 6682 The value comes from the initial kernel execution state. See 6683 :ref:`amdgpu-amdhsa-sgpr-register-set-up-order-table`. 6684 66853. Queue Ptr (2 SGPRs) 6686 6687 The value comes from the initial kernel execution state. See 6688 :ref:`amdgpu-amdhsa-sgpr-register-set-up-order-table`. 6689 66904. Kernarg Segment Ptr (2 SGPRs) 6691 6692 The value comes from the initial kernel execution state. See 6693 :ref:`amdgpu-amdhsa-sgpr-register-set-up-order-table`. 6694 66955. Dispatch id (2 SGPRs) 6696 6697 The value comes from the initial kernel execution state. See 6698 :ref:`amdgpu-amdhsa-sgpr-register-set-up-order-table`. 6699 67006. Work-Group ID X (1 SGPR) 6701 6702 The value comes from the initial kernel execution state. See 6703 :ref:`amdgpu-amdhsa-sgpr-register-set-up-order-table`. 6704 67057. Work-Group ID Y (1 SGPR) 6706 6707 The value comes from the initial kernel execution state. See 6708 :ref:`amdgpu-amdhsa-sgpr-register-set-up-order-table`. 6709 67108. Work-Group ID Z (1 SGPR) 6711 6712 The value comes from the initial kernel execution state. See 6713 :ref:`amdgpu-amdhsa-sgpr-register-set-up-order-table`. 6714 67159. Implicit Argument Ptr (2 SGPRs) 6716 6717 The value is computed by adding an offset to Kernarg Segment Ptr to get the 6718 global address space pointer to the first kernarg implicit argument. 6719 6720The input and result arguments are assigned in order in the following manner: 6721 6722.. note:: 6723 6724 There are likely some errors and omissions in the following description that 6725 need correction. 6726 6727 .. TODO:: 6728 6729 Check the clang source code to decipher how function arguments and return 6730 results are handled. Also see the AMDGPU specific values used. 6731 6732* VGPR arguments are assigned to consecutive VGPRs starting at VGPR0 up to 6733 VGPR31. 6734 6735 If there are more arguments than will fit in these registers, the remaining 6736 arguments are allocated on the stack in order on naturally aligned 6737 addresses. 6738 6739 .. TODO:: 6740 6741 How are overly aligned structures allocated on the stack? 6742 6743* SGPR arguments are assigned to consecutive SGPRs starting at SGPR0 up to 6744 SGPR29. 6745 6746 If there are more arguments than will fit in these registers, the remaining 6747 arguments are allocated on the stack in order on naturally aligned 6748 addresses. 6749 6750Note that decomposed struct type arguments may have some fields passed in 6751registers and some in memory. 6752 6753.. TODO:: 6754 6755 So, a struct which can pass some fields as decomposed register arguments, will 6756 pass the rest as decomposed stack elements? But an argument that will not start 6757 in registers will not be decomposed and will be passed as a non-decomposed 6758 stack value? 6759 6760The following is not part of the AMDGPU function calling convention but 6761describes how the AMDGPU implements function calls: 6762 67631. SGPR33 is used as a frame pointer (FP) if necessary. Like the SP it is an 6764 unswizzled scratch address. It is only needed if runtime sized ``alloca`` 6765 are used, or for the reasons defined in ``SIFrameLowering``. 67662. Runtime stack alignment is supported. SGPR34 is used as a base pointer (BP) 6767 to access the incoming stack arguments in the function. The BP is needed 6768 only when the function requires the runtime stack alignment. 6769 67703. Allocating SGPR arguments on the stack are not supported. 6771 67724. No CFI is currently generated. See 6773 :ref:`amdgpu-dwarf-call-frame-information`. 6774 6775 .. note:: 6776 6777 CFI will be generated that defines the CFA as the unswizzled address 6778 relative to the wave scratch base in the unswizzled private address space 6779 of the lowest address stack allocated local variable. 6780 6781 ``DW_AT_frame_base`` will be defined as the swizzled address in the 6782 swizzled private address space by dividing the CFA by the wavefront size 6783 (since CFA is always at least dword aligned which matches the scratch 6784 swizzle element size). 6785 6786 If no dynamic stack alignment was performed, the stack allocated arguments 6787 are accessed as negative offsets relative to ``DW_AT_frame_base``, and the 6788 local variables and register spill slots are accessed as positive offsets 6789 relative to ``DW_AT_frame_base``. 6790 67915. Function argument passing is implemented by copying the input physical 6792 registers to virtual registers on entry. The register allocator can spill if 6793 necessary. These are copied back to physical registers at call sites. The 6794 net effect is that each function call can have these values in entirely 6795 distinct locations. The IPRA can help avoid shuffling argument registers. 67966. Call sites are implemented by setting up the arguments at positive offsets 6797 from SP. Then SP is incremented to account for the known frame size before 6798 the call and decremented after the call. 6799 6800 .. note:: 6801 6802 The CFI will reflect the changed calculation needed to compute the CFA 6803 from SP. 6804 68057. 4 byte spill slots are used in the stack frame. One slot is allocated for an 6806 emergency spill slot. Buffer instructions are used for stack accesses and 6807 not the ``flat_scratch`` instruction. 6808 6809 .. TODO:: 6810 6811 Explain when the emergency spill slot is used. 6812 6813.. TODO:: 6814 6815 Possible broken issues: 6816 6817 - Stack arguments must be aligned to required alignment. 6818 - Stack is aligned to max(16, max formal argument alignment) 6819 - Direct argument < 64 bits should check register budget. 6820 - Register budget calculation should respect ``inreg`` for SGPR. 6821 - SGPR overflow is not handled. 6822 - struct with 1 member unpeeling is not checking size of member. 6823 - ``sret`` is after ``this`` pointer. 6824 - Caller is not implementing stack realignment: need an extra pointer. 6825 - Should say AMDGPU passes FP rather than SP. 6826 - Should CFI define CFA as address of locals or arguments. Difference is 6827 apparent when have implemented dynamic alignment. 6828 - If ``SCRATCH`` instruction could allow negative offsets, then can make FP be 6829 highest address of stack frame and use negative offset for locals. Would 6830 allow SP to be the same as FP and could support signal-handler-like as now 6831 have a real SP for the top of the stack. 6832 - How is ``sret`` passed on the stack? In argument stack area? Can it overlay 6833 arguments? 6834 6835AMDPAL 6836------ 6837 6838This section provides code conventions used when the target triple OS is 6839``amdpal`` (see :ref:`amdgpu-target-triples`) for passing runtime parameters 6840from the application/runtime to each invocation of a hardware shader. These 6841parameters include both generic, application-controlled parameters called 6842*user data* as well as system-generated parameters that are a product of the 6843draw or dispatch execution. 6844 6845User Data 6846~~~~~~~~~ 6847 6848Each hardware stage has a set of 32-bit *user data registers* which can be 6849written from a command buffer and then loaded into SGPRs when waves are launched 6850via a subsequent dispatch or draw operation. This is the way most arguments are 6851passed from the application/runtime to a hardware shader. 6852 6853Compute User Data 6854~~~~~~~~~~~~~~~~~ 6855 6856Compute shader user data mappings are simpler than graphics shaders and have a 6857fixed mapping. 6858 6859Note that there are always 10 available *user data entries* in registers - 6860entries beyond that limit must be fetched from memory (via the spill table 6861pointer) by the shader. 6862 6863 .. table:: PAL Compute Shader User Data Registers 6864 :name: pal-compute-user-data-registers 6865 6866 ============= ================================ 6867 User Register Description 6868 ============= ================================ 6869 0 Global Internal Table (32-bit pointer) 6870 1 Per-Shader Internal Table (32-bit pointer) 6871 2 - 11 Application-Controlled User Data (10 32-bit values) 6872 12 Spill Table (32-bit pointer) 6873 13 - 14 Thread Group Count (64-bit pointer) 6874 15 GDS Range 6875 ============= ================================ 6876 6877Graphics User Data 6878~~~~~~~~~~~~~~~~~~ 6879 6880Graphics pipelines support a much more flexible user data mapping: 6881 6882 .. table:: PAL Graphics Shader User Data Registers 6883 :name: pal-graphics-user-data-registers 6884 6885 ============= ================================ 6886 User Register Description 6887 ============= ================================ 6888 0 Global Internal Table (32-bit pointer) 6889 + Per-Shader Internal Table (32-bit pointer) 6890 + 1-15 Application Controlled User Data 6891 (1-15 Contiguous 32-bit Values in Registers) 6892 + Spill Table (32-bit pointer) 6893 + Draw Index (First Stage Only) 6894 + Vertex Offset (First Stage Only) 6895 + Instance Offset (First Stage Only) 6896 ============= ================================ 6897 6898 The placement of the global internal table remains fixed in the first *user 6899 data SGPR register*. Otherwise all parameters are optional, and can be mapped 6900 to any desired *user data SGPR register*, with the following restrictions: 6901 6902 * Draw Index, Vertex Offset, and Instance Offset can only be used by the first 6903 active hardware stage in a graphics pipeline (i.e. where the API vertex 6904 shader runs). 6905 6906 * Application-controlled user data must be mapped into a contiguous range of 6907 user data registers. 6908 6909 * The application-controlled user data range supports compaction remapping, so 6910 only *entries* that are actually consumed by the shader must be assigned to 6911 corresponding *registers*. Note that in order to support an efficient runtime 6912 implementation, the remapping must pack *registers* in the same order as 6913 *entries*, with unused *entries* removed. 6914 6915.. _pal_global_internal_table: 6916 6917Global Internal Table 6918~~~~~~~~~~~~~~~~~~~~~ 6919 6920The global internal table is a table of *shader resource descriptors* (SRDs) 6921that define how certain engine-wide, runtime-managed resources should be 6922accessed from a shader. The majority of these resources have HW-defined formats, 6923and it is up to the compiler to write/read data as required by the target 6924hardware. 6925 6926The following table illustrates the required format: 6927 6928 .. table:: PAL Global Internal Table 6929 :name: pal-git-table 6930 6931 ============= ================================ 6932 Offset Description 6933 ============= ================================ 6934 0-3 Graphics Scratch SRD 6935 4-7 Compute Scratch SRD 6936 8-11 ES/GS Ring Output SRD 6937 12-15 ES/GS Ring Input SRD 6938 16-19 GS/VS Ring Output #0 6939 20-23 GS/VS Ring Output #1 6940 24-27 GS/VS Ring Output #2 6941 28-31 GS/VS Ring Output #3 6942 32-35 GS/VS Ring Input SRD 6943 36-39 Tessellation Factor Buffer SRD 6944 40-43 Off-Chip LDS Buffer SRD 6945 44-47 Off-Chip Param Cache Buffer SRD 6946 48-51 Sample Position Buffer SRD 6947 52 vaRange::ShadowDescriptorTable High Bits 6948 ============= ================================ 6949 6950 The pointer to the global internal table passed to the shader as user data 6951 is a 32-bit pointer. The top 32 bits should be assumed to be the same as 6952 the top 32 bits of the pipeline, so the shader may use the program 6953 counter's top 32 bits. 6954 6955Unspecified OS 6956-------------- 6957 6958This section provides code conventions used when the target triple OS is 6959empty (see :ref:`amdgpu-target-triples`). 6960 6961Trap Handler ABI 6962~~~~~~~~~~~~~~~~ 6963 6964For code objects generated by AMDGPU backend for non-amdhsa OS, the runtime does 6965not install a trap handler. The ``llvm.trap`` and ``llvm.debugtrap`` 6966instructions are handled as follows: 6967 6968 .. table:: AMDGPU Trap Handler for Non-AMDHSA OS 6969 :name: amdgpu-trap-handler-for-non-amdhsa-os-table 6970 6971 =============== =============== =========================================== 6972 Usage Code Sequence Description 6973 =============== =============== =========================================== 6974 llvm.trap s_endpgm Causes wavefront to be terminated. 6975 llvm.debugtrap *none* Compiler warning given that there is no 6976 trap handler installed. 6977 =============== =============== =========================================== 6978 6979Source Languages 6980================ 6981 6982.. _amdgpu-opencl: 6983 6984OpenCL 6985------ 6986 6987When the language is OpenCL the following differences occur: 6988 69891. The OpenCL memory model is used (see :ref:`amdgpu-amdhsa-memory-model`). 69902. The AMDGPU backend appends additional arguments to the kernel's explicit 6991 arguments for the AMDHSA OS (see 6992 :ref:`opencl-kernel-implicit-arguments-appended-for-amdhsa-os-table`). 69933. Additional metadata is generated 6994 (see :ref:`amdgpu-amdhsa-code-object-metadata`). 6995 6996 .. table:: OpenCL kernel implicit arguments appended for AMDHSA OS 6997 :name: opencl-kernel-implicit-arguments-appended-for-amdhsa-os-table 6998 6999 ======== ==== ========= =========================================== 7000 Position Byte Byte Description 7001 Size Alignment 7002 ======== ==== ========= =========================================== 7003 1 8 8 OpenCL Global Offset X 7004 2 8 8 OpenCL Global Offset Y 7005 3 8 8 OpenCL Global Offset Z 7006 4 8 8 OpenCL address of printf buffer 7007 5 8 8 OpenCL address of virtual queue used by 7008 enqueue_kernel. 7009 6 8 8 OpenCL address of AqlWrap struct used by 7010 enqueue_kernel. 7011 7 8 8 Pointer argument used for Multi-gird 7012 synchronization. 7013 ======== ==== ========= =========================================== 7014 7015.. _amdgpu-hcc: 7016 7017HCC 7018--- 7019 7020When the language is HCC the following differences occur: 7021 70221. The HSA memory model is used (see :ref:`amdgpu-amdhsa-memory-model`). 7023 7024.. _amdgpu-assembler: 7025 7026Assembler 7027--------- 7028 7029AMDGPU backend has LLVM-MC based assembler which is currently in development. 7030It supports AMDGCN GFX6-GFX10. 7031 7032This section describes general syntax for instructions and operands. 7033 7034Instructions 7035~~~~~~~~~~~~ 7036 7037An instruction has the following :doc:`syntax<AMDGPUInstructionSyntax>`: 7038 7039 | ``<``\ *opcode*\ ``> <``\ *operand0*\ ``>, <``\ *operand1*\ ``>,... 7040 <``\ *modifier0*\ ``> <``\ *modifier1*\ ``>...`` 7041 7042:doc:`Operands<AMDGPUOperandSyntax>` are comma-separated while 7043:doc:`modifiers<AMDGPUModifierSyntax>` are space-separated. 7044 7045The order of operands and modifiers is fixed. 7046Most modifiers are optional and may be omitted. 7047 7048Links to detailed instruction syntax description may be found in the following 7049table. Note that features under development are not included 7050in this description. 7051 7052 =================================== ======================================= 7053 Core ISA ISA Extensions 7054 =================================== ======================================= 7055 :doc:`GFX7<AMDGPU/AMDGPUAsmGFX7>` \- 7056 :doc:`GFX8<AMDGPU/AMDGPUAsmGFX8>` \- 7057 :doc:`GFX9<AMDGPU/AMDGPUAsmGFX9>` :doc:`gfx900<AMDGPU/AMDGPUAsmGFX900>` 7058 7059 :doc:`gfx902<AMDGPU/AMDGPUAsmGFX900>` 7060 7061 :doc:`gfx904<AMDGPU/AMDGPUAsmGFX904>` 7062 7063 :doc:`gfx906<AMDGPU/AMDGPUAsmGFX906>` 7064 7065 :doc:`gfx908<AMDGPU/AMDGPUAsmGFX908>` 7066 7067 :doc:`gfx909<AMDGPU/AMDGPUAsmGFX900>` 7068 7069 :doc:`GFX10<AMDGPU/AMDGPUAsmGFX10>` :doc:`gfx1011<AMDGPU/AMDGPUAsmGFX1011>` 7070 7071 :doc:`gfx1012<AMDGPU/AMDGPUAsmGFX1011>` 7072 =================================== ======================================= 7073 7074For more information about instructions, their semantics and supported 7075combinations of operands, refer to one of instruction set architecture manuals 7076[AMD-GCN-GFX6]_, [AMD-GCN-GFX7]_, [AMD-GCN-GFX8]_, [AMD-GCN-GFX9]_ and 7077[AMD-GCN-GFX10]_. 7078 7079Operands 7080~~~~~~~~ 7081 7082Detailed description of operands may be found :doc:`here<AMDGPUOperandSyntax>`. 7083 7084Modifiers 7085~~~~~~~~~ 7086 7087Detailed description of modifiers may be found 7088:doc:`here<AMDGPUModifierSyntax>`. 7089 7090Instruction Examples 7091~~~~~~~~~~~~~~~~~~~~ 7092 7093DS 7094++ 7095 7096.. code-block:: nasm 7097 7098 ds_add_u32 v2, v4 offset:16 7099 ds_write_src2_b64 v2 offset0:4 offset1:8 7100 ds_cmpst_f32 v2, v4, v6 7101 ds_min_rtn_f64 v[8:9], v2, v[4:5] 7102 7103For full list of supported instructions, refer to "LDS/GDS instructions" in ISA 7104Manual. 7105 7106FLAT 7107++++ 7108 7109.. code-block:: nasm 7110 7111 flat_load_dword v1, v[3:4] 7112 flat_store_dwordx3 v[3:4], v[5:7] 7113 flat_atomic_swap v1, v[3:4], v5 glc 7114 flat_atomic_cmpswap v1, v[3:4], v[5:6] glc slc 7115 flat_atomic_fmax_x2 v[1:2], v[3:4], v[5:6] glc 7116 7117For full list of supported instructions, refer to "FLAT instructions" in ISA 7118Manual. 7119 7120MUBUF 7121+++++ 7122 7123.. code-block:: nasm 7124 7125 buffer_load_dword v1, off, s[4:7], s1 7126 buffer_store_dwordx4 v[1:4], v2, ttmp[4:7], s1 offen offset:4 glc tfe 7127 buffer_store_format_xy v[1:2], off, s[4:7], s1 7128 buffer_wbinvl1 7129 buffer_atomic_inc v1, v2, s[8:11], s4 idxen offset:4 slc 7130 7131For full list of supported instructions, refer to "MUBUF Instructions" in ISA 7132Manual. 7133 7134SMRD/SMEM 7135+++++++++ 7136 7137.. code-block:: nasm 7138 7139 s_load_dword s1, s[2:3], 0xfc 7140 s_load_dwordx8 s[8:15], s[2:3], s4 7141 s_load_dwordx16 s[88:103], s[2:3], s4 7142 s_dcache_inv_vol 7143 s_memtime s[4:5] 7144 7145For full list of supported instructions, refer to "Scalar Memory Operations" in 7146ISA Manual. 7147 7148SOP1 7149++++ 7150 7151.. code-block:: nasm 7152 7153 s_mov_b32 s1, s2 7154 s_mov_b64 s[0:1], 0x80000000 7155 s_cmov_b32 s1, 200 7156 s_wqm_b64 s[2:3], s[4:5] 7157 s_bcnt0_i32_b64 s1, s[2:3] 7158 s_swappc_b64 s[2:3], s[4:5] 7159 s_cbranch_join s[4:5] 7160 7161For full list of supported instructions, refer to "SOP1 Instructions" in ISA 7162Manual. 7163 7164SOP2 7165++++ 7166 7167.. code-block:: nasm 7168 7169 s_add_u32 s1, s2, s3 7170 s_and_b64 s[2:3], s[4:5], s[6:7] 7171 s_cselect_b32 s1, s2, s3 7172 s_andn2_b32 s2, s4, s6 7173 s_lshr_b64 s[2:3], s[4:5], s6 7174 s_ashr_i32 s2, s4, s6 7175 s_bfm_b64 s[2:3], s4, s6 7176 s_bfe_i64 s[2:3], s[4:5], s6 7177 s_cbranch_g_fork s[4:5], s[6:7] 7178 7179For full list of supported instructions, refer to "SOP2 Instructions" in ISA 7180Manual. 7181 7182SOPC 7183++++ 7184 7185.. code-block:: nasm 7186 7187 s_cmp_eq_i32 s1, s2 7188 s_bitcmp1_b32 s1, s2 7189 s_bitcmp0_b64 s[2:3], s4 7190 s_setvskip s3, s5 7191 7192For full list of supported instructions, refer to "SOPC Instructions" in ISA 7193Manual. 7194 7195SOPP 7196++++ 7197 7198.. code-block:: nasm 7199 7200 s_barrier 7201 s_nop 2 7202 s_endpgm 7203 s_waitcnt 0 ; Wait for all counters to be 0 7204 s_waitcnt vmcnt(0) & expcnt(0) & lgkmcnt(0) ; Equivalent to above 7205 s_waitcnt vmcnt(1) ; Wait for vmcnt counter to be 1. 7206 s_sethalt 9 7207 s_sleep 10 7208 s_sendmsg 0x1 7209 s_sendmsg sendmsg(MSG_INTERRUPT) 7210 s_trap 1 7211 7212For full list of supported instructions, refer to "SOPP Instructions" in ISA 7213Manual. 7214 7215Unless otherwise mentioned, little verification is performed on the operands 7216of SOPP Instructions, so it is up to the programmer to be familiar with the 7217range or acceptable values. 7218 7219VALU 7220++++ 7221 7222For vector ALU instruction opcodes (VOP1, VOP2, VOP3, VOPC, VOP_DPP, VOP_SDWA), 7223the assembler will automatically use optimal encoding based on its operands. To 7224force specific encoding, one can add a suffix to the opcode of the instruction: 7225 7226* _e32 for 32-bit VOP1/VOP2/VOPC 7227* _e64 for 64-bit VOP3 7228* _dpp for VOP_DPP 7229* _sdwa for VOP_SDWA 7230 7231VOP1/VOP2/VOP3/VOPC examples: 7232 7233.. code-block:: nasm 7234 7235 v_mov_b32 v1, v2 7236 v_mov_b32_e32 v1, v2 7237 v_nop 7238 v_cvt_f64_i32_e32 v[1:2], v2 7239 v_floor_f32_e32 v1, v2 7240 v_bfrev_b32_e32 v1, v2 7241 v_add_f32_e32 v1, v2, v3 7242 v_mul_i32_i24_e64 v1, v2, 3 7243 v_mul_i32_i24_e32 v1, -3, v3 7244 v_mul_i32_i24_e32 v1, -100, v3 7245 v_addc_u32 v1, s[0:1], v2, v3, s[2:3] 7246 v_max_f16_e32 v1, v2, v3 7247 7248VOP_DPP examples: 7249 7250.. code-block:: nasm 7251 7252 v_mov_b32 v0, v0 quad_perm:[0,2,1,1] 7253 v_sin_f32 v0, v0 row_shl:1 row_mask:0xa bank_mask:0x1 bound_ctrl:0 7254 v_mov_b32 v0, v0 wave_shl:1 7255 v_mov_b32 v0, v0 row_mirror 7256 v_mov_b32 v0, v0 row_bcast:31 7257 v_mov_b32 v0, v0 quad_perm:[1,3,0,1] row_mask:0xa bank_mask:0x1 bound_ctrl:0 7258 v_add_f32 v0, v0, |v0| row_shl:1 row_mask:0xa bank_mask:0x1 bound_ctrl:0 7259 v_max_f16 v1, v2, v3 row_shl:1 row_mask:0xa bank_mask:0x1 bound_ctrl:0 7260 7261VOP_SDWA examples: 7262 7263.. code-block:: nasm 7264 7265 v_mov_b32 v1, v2 dst_sel:BYTE_0 dst_unused:UNUSED_PRESERVE src0_sel:DWORD 7266 v_min_u32 v200, v200, v1 dst_sel:WORD_1 dst_unused:UNUSED_PAD src0_sel:BYTE_1 src1_sel:DWORD 7267 v_sin_f32 v0, v0 dst_unused:UNUSED_PAD src0_sel:WORD_1 7268 v_fract_f32 v0, |v0| dst_sel:DWORD dst_unused:UNUSED_PAD src0_sel:WORD_1 7269 v_cmpx_le_u32 vcc, v1, v2 src0_sel:BYTE_2 src1_sel:WORD_0 7270 7271For full list of supported instructions, refer to "Vector ALU instructions". 7272 7273.. TODO:: 7274 7275 Remove once we switch to code object v3 by default. 7276 7277.. _amdgpu-amdhsa-assembler-predefined-symbols-v2: 7278 7279Code Object V2 Predefined Symbols (-mattr=-code-object-v3) 7280~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 7281 7282.. warning:: Code Object V2 is not the default code object version emitted by 7283 this version of LLVM. For a description of the predefined symbols available 7284 with the default configuration (Code Object V3) see 7285 :ref:`amdgpu-amdhsa-assembler-predefined-symbols-v3`. 7286 7287The AMDGPU assembler defines and updates some symbols automatically. These 7288symbols do not affect code generation. 7289 7290.option.machine_version_major 7291+++++++++++++++++++++++++++++ 7292 7293Set to the GFX major generation number of the target being assembled for. For 7294example, when assembling for a "GFX9" target this will be set to the integer 7295value "9". The possible GFX major generation numbers are presented in 7296:ref:`amdgpu-processors`. 7297 7298.option.machine_version_minor 7299+++++++++++++++++++++++++++++ 7300 7301Set to the GFX minor generation number of the target being assembled for. For 7302example, when assembling for a "GFX810" target this will be set to the integer 7303value "1". The possible GFX minor generation numbers are presented in 7304:ref:`amdgpu-processors`. 7305 7306.option.machine_version_stepping 7307++++++++++++++++++++++++++++++++ 7308 7309Set to the GFX stepping generation number of the target being assembled for. 7310For example, when assembling for a "GFX704" target this will be set to the 7311integer value "4". The possible GFX stepping generation numbers are presented 7312in :ref:`amdgpu-processors`. 7313 7314.kernel.vgpr_count 7315++++++++++++++++++ 7316 7317Set to zero each time a 7318:ref:`amdgpu-amdhsa-assembler-directive-amdgpu_hsa_kernel` directive is 7319encountered. At each instruction, if the current value of this symbol is less 7320than or equal to the maximum VPGR number explicitly referenced within that 7321instruction then the symbol value is updated to equal that VGPR number plus 7322one. 7323 7324.kernel.sgpr_count 7325++++++++++++++++++ 7326 7327Set to zero each time a 7328:ref:`amdgpu-amdhsa-assembler-directive-amdgpu_hsa_kernel` directive is 7329encountered. At each instruction, if the current value of this symbol is less 7330than or equal to the maximum VPGR number explicitly referenced within that 7331instruction then the symbol value is updated to equal that SGPR number plus 7332one. 7333 7334.. _amdgpu-amdhsa-assembler-directives-v2: 7335 7336Code Object V2 Directives (-mattr=-code-object-v3) 7337~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 7338 7339.. warning:: Code Object V2 is not the default code object version emitted by 7340 this version of LLVM. For a description of the directives supported with 7341 the default configuration (Code Object V3) see 7342 :ref:`amdgpu-amdhsa-assembler-directives-v3`. 7343 7344AMDGPU ABI defines auxiliary data in output code object. In assembly source, 7345one can specify them with assembler directives. 7346 7347.hsa_code_object_version major, minor 7348+++++++++++++++++++++++++++++++++++++ 7349 7350*major* and *minor* are integers that specify the version of the HSA code 7351object that will be generated by the assembler. 7352 7353.hsa_code_object_isa [major, minor, stepping, vendor, arch] 7354+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ 7355 7356 7357*major*, *minor*, and *stepping* are all integers that describe the instruction 7358set architecture (ISA) version of the assembly program. 7359 7360*vendor* and *arch* are quoted strings. *vendor* should always be equal to 7361"AMD" and *arch* should always be equal to "AMDGPU". 7362 7363By default, the assembler will derive the ISA version, *vendor*, and *arch* 7364from the value of the -mcpu option that is passed to the assembler. 7365 7366.. _amdgpu-amdhsa-assembler-directive-amdgpu_hsa_kernel: 7367 7368.amdgpu_hsa_kernel (name) 7369+++++++++++++++++++++++++ 7370 7371This directives specifies that the symbol with given name is a kernel entry 7372point (label) and the object should contain corresponding symbol of type 7373STT_AMDGPU_HSA_KERNEL. 7374 7375.amd_kernel_code_t 7376++++++++++++++++++ 7377 7378This directive marks the beginning of a list of key / value pairs that are used 7379to specify the amd_kernel_code_t object that will be emitted by the assembler. 7380The list must be terminated by the *.end_amd_kernel_code_t* directive. For any 7381amd_kernel_code_t values that are unspecified a default value will be used. The 7382default value for all keys is 0, with the following exceptions: 7383 7384- *amd_code_version_major* defaults to 1. 7385- *amd_kernel_code_version_minor* defaults to 2. 7386- *amd_machine_kind* defaults to 1. 7387- *amd_machine_version_major*, *machine_version_minor*, and 7388 *amd_machine_version_stepping* are derived from the value of the -mcpu option 7389 that is passed to the assembler. 7390- *kernel_code_entry_byte_offset* defaults to 256. 7391- *wavefront_size* defaults 6 for all targets before GFX10. For GFX10 onwards 7392 defaults to 6 if target feature ``wavefrontsize64`` is enabled, otherwise 5. 7393 Note that wavefront size is specified as a power of two, so a value of **n** 7394 means a size of 2^ **n**. 7395- *call_convention* defaults to -1. 7396- *kernarg_segment_alignment*, *group_segment_alignment*, and 7397 *private_segment_alignment* default to 4. Note that alignments are specified 7398 as a power of 2, so a value of **n** means an alignment of 2^ **n**. 7399- *enable_wgp_mode* defaults to 1 if target feature ``cumode`` is disabled for 7400 GFX10 onwards. 7401- *enable_mem_ordered* defaults to 1 for GFX10 onwards. 7402 7403The *.amd_kernel_code_t* directive must be placed immediately after the 7404function label and before any instructions. 7405 7406For a full list of amd_kernel_code_t keys, refer to AMDGPU ABI document, 7407comments in lib/Target/AMDGPU/AmdKernelCodeT.h and test/CodeGen/AMDGPU/hsa.s. 7408 7409.. _amdgpu-amdhsa-assembler-example-v2: 7410 7411Code Object V2 Example Source Code (-mattr=-code-object-v3) 7412~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 7413 7414.. warning:: Code Object V2 is not the default code object version emitted by 7415 this version of LLVM. For a description of the directives supported with 7416 the default configuration (Code Object V3) see 7417 :ref:`amdgpu-amdhsa-assembler-example-v3`. 7418 7419Here is an example of a minimal assembly source file, defining one HSA kernel: 7420 7421.. code:: 7422 :number-lines: 7423 7424 .hsa_code_object_version 1,0 7425 .hsa_code_object_isa 7426 7427 .hsatext 7428 .globl hello_world 7429 .p2align 8 7430 .amdgpu_hsa_kernel hello_world 7431 7432 hello_world: 7433 7434 .amd_kernel_code_t 7435 enable_sgpr_kernarg_segment_ptr = 1 7436 is_ptr64 = 1 7437 compute_pgm_rsrc1_vgprs = 0 7438 compute_pgm_rsrc1_sgprs = 0 7439 compute_pgm_rsrc2_user_sgpr = 2 7440 compute_pgm_rsrc1_wgp_mode = 0 7441 compute_pgm_rsrc1_mem_ordered = 0 7442 compute_pgm_rsrc1_fwd_progress = 1 7443 .end_amd_kernel_code_t 7444 7445 s_load_dwordx2 s[0:1], s[0:1] 0x0 7446 v_mov_b32 v0, 3.14159 7447 s_waitcnt lgkmcnt(0) 7448 v_mov_b32 v1, s0 7449 v_mov_b32 v2, s1 7450 flat_store_dword v[1:2], v0 7451 s_endpgm 7452 .Lfunc_end0: 7453 .size hello_world, .Lfunc_end0-hello_world 7454 7455.. _amdgpu-amdhsa-assembler-predefined-symbols-v3: 7456 7457Code Object V3 Predefined Symbols (-mattr=+code-object-v3) 7458~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 7459 7460The AMDGPU assembler defines and updates some symbols automatically. These 7461symbols do not affect code generation. 7462 7463.amdgcn.gfx_generation_number 7464+++++++++++++++++++++++++++++ 7465 7466Set to the GFX major generation number of the target being assembled for. For 7467example, when assembling for a "GFX9" target this will be set to the integer 7468value "9". The possible GFX major generation numbers are presented in 7469:ref:`amdgpu-processors`. 7470 7471.amdgcn.gfx_generation_minor 7472++++++++++++++++++++++++++++ 7473 7474Set to the GFX minor generation number of the target being assembled for. For 7475example, when assembling for a "GFX810" target this will be set to the integer 7476value "1". The possible GFX minor generation numbers are presented in 7477:ref:`amdgpu-processors`. 7478 7479.amdgcn.gfx_generation_stepping 7480+++++++++++++++++++++++++++++++ 7481 7482Set to the GFX stepping generation number of the target being assembled for. 7483For example, when assembling for a "GFX704" target this will be set to the 7484integer value "4". The possible GFX stepping generation numbers are presented 7485in :ref:`amdgpu-processors`. 7486 7487.. _amdgpu-amdhsa-assembler-symbol-next_free_vgpr: 7488 7489.amdgcn.next_free_vgpr 7490++++++++++++++++++++++ 7491 7492Set to zero before assembly begins. At each instruction, if the current value 7493of this symbol is less than or equal to the maximum VGPR number explicitly 7494referenced within that instruction then the symbol value is updated to equal 7495that VGPR number plus one. 7496 7497May be used to set the `.amdhsa_next_free_vpgr` directive in 7498:ref:`amdhsa-kernel-directives-table`. 7499 7500May be set at any time, e.g. manually set to zero at the start of each kernel. 7501 7502.. _amdgpu-amdhsa-assembler-symbol-next_free_sgpr: 7503 7504.amdgcn.next_free_sgpr 7505++++++++++++++++++++++ 7506 7507Set to zero before assembly begins. At each instruction, if the current value 7508of this symbol is less than or equal the maximum SGPR number explicitly 7509referenced within that instruction then the symbol value is updated to equal 7510that SGPR number plus one. 7511 7512May be used to set the `.amdhsa_next_free_spgr` directive in 7513:ref:`amdhsa-kernel-directives-table`. 7514 7515May be set at any time, e.g. manually set to zero at the start of each kernel. 7516 7517.. _amdgpu-amdhsa-assembler-directives-v3: 7518 7519Code Object V3 Directives (-mattr=+code-object-v3) 7520~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 7521 7522Directives which begin with ``.amdgcn`` are valid for all ``amdgcn`` 7523architecture processors, and are not OS-specific. Directives which begin with 7524``.amdhsa`` are specific to ``amdgcn`` architecture processors when the 7525``amdhsa`` OS is specified. See :ref:`amdgpu-target-triples` and 7526:ref:`amdgpu-processors`. 7527 7528.amdgcn_target <target> 7529+++++++++++++++++++++++ 7530 7531Optional directive which declares the target supported by the containing 7532assembler source file. Valid values are described in 7533:ref:`amdgpu-amdhsa-code-object-target-identification`. Used by the assembler 7534to validate command-line options such as ``-triple``, ``-mcpu``, and those 7535which specify target features. 7536 7537.amdhsa_kernel <name> 7538+++++++++++++++++++++ 7539 7540Creates a correctly aligned AMDHSA kernel descriptor and a symbol, 7541``<name>.kd``, in the current location of the current section. Only valid when 7542the OS is ``amdhsa``. ``<name>`` must be a symbol that labels the first 7543instruction to execute, and does not need to be previously defined. 7544 7545Marks the beginning of a list of directives used to generate the bytes of a 7546kernel descriptor, as described in :ref:`amdgpu-amdhsa-kernel-descriptor`. 7547Directives which may appear in this list are described in 7548:ref:`amdhsa-kernel-directives-table`. Directives may appear in any order, must 7549be valid for the target being assembled for, and cannot be repeated. Directives 7550support the range of values specified by the field they reference in 7551:ref:`amdgpu-amdhsa-kernel-descriptor`. If a directive is not specified, it is 7552assumed to have its default value, unless it is marked as "Required", in which 7553case it is an error to omit the directive. This list of directives is 7554terminated by an ``.end_amdhsa_kernel`` directive. 7555 7556 .. table:: AMDHSA Kernel Assembler Directives 7557 :name: amdhsa-kernel-directives-table 7558 7559 ======================================================== =================== ============ =================== 7560 Directive Default Supported On Description 7561 ======================================================== =================== ============ =================== 7562 ``.amdhsa_group_segment_fixed_size`` 0 GFX6-GFX10 Controls GROUP_SEGMENT_FIXED_SIZE in 7563 :ref:`amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table`. 7564 ``.amdhsa_private_segment_fixed_size`` 0 GFX6-GFX10 Controls PRIVATE_SEGMENT_FIXED_SIZE in 7565 :ref:`amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table`. 7566 ``.amdhsa_user_sgpr_private_segment_buffer`` 0 GFX6-GFX10 Controls ENABLE_SGPR_PRIVATE_SEGMENT_BUFFER in 7567 :ref:`amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table`. 7568 ``.amdhsa_user_sgpr_dispatch_ptr`` 0 GFX6-GFX10 Controls ENABLE_SGPR_DISPATCH_PTR in 7569 :ref:`amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table`. 7570 ``.amdhsa_user_sgpr_queue_ptr`` 0 GFX6-GFX10 Controls ENABLE_SGPR_QUEUE_PTR in 7571 :ref:`amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table`. 7572 ``.amdhsa_user_sgpr_kernarg_segment_ptr`` 0 GFX6-GFX10 Controls ENABLE_SGPR_KERNARG_SEGMENT_PTR in 7573 :ref:`amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table`. 7574 ``.amdhsa_user_sgpr_dispatch_id`` 0 GFX6-GFX10 Controls ENABLE_SGPR_DISPATCH_ID in 7575 :ref:`amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table`. 7576 ``.amdhsa_user_sgpr_flat_scratch_init`` 0 GFX6-GFX10 Controls ENABLE_SGPR_FLAT_SCRATCH_INIT in 7577 :ref:`amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table`. 7578 ``.amdhsa_user_sgpr_private_segment_size`` 0 GFX6-GFX10 Controls ENABLE_SGPR_PRIVATE_SEGMENT_SIZE in 7579 :ref:`amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table`. 7580 ``.amdhsa_wavefront_size32`` Target GFX10 Controls ENABLE_WAVEFRONT_SIZE32 in 7581 Feature :ref:`amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table`. 7582 Specific 7583 (-wavefrontsize64) 7584 ``.amdhsa_system_sgpr_private_segment_wavefront_offset`` 0 GFX6-GFX10 Controls ENABLE_SGPR_PRIVATE_SEGMENT_WAVEFRONT_OFFSET in 7585 :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`. 7586 ``.amdhsa_system_sgpr_workgroup_id_x`` 1 GFX6-GFX10 Controls ENABLE_SGPR_WORKGROUP_ID_X in 7587 :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`. 7588 ``.amdhsa_system_sgpr_workgroup_id_y`` 0 GFX6-GFX10 Controls ENABLE_SGPR_WORKGROUP_ID_Y in 7589 :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`. 7590 ``.amdhsa_system_sgpr_workgroup_id_z`` 0 GFX6-GFX10 Controls ENABLE_SGPR_WORKGROUP_ID_Z in 7591 :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`. 7592 ``.amdhsa_system_sgpr_workgroup_info`` 0 GFX6-GFX10 Controls ENABLE_SGPR_WORKGROUP_INFO in 7593 :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`. 7594 ``.amdhsa_system_vgpr_workitem_id`` 0 GFX6-GFX10 Controls ENABLE_VGPR_WORKITEM_ID in 7595 :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`. 7596 Possible values are defined in 7597 :ref:`amdgpu-amdhsa-system-vgpr-work-item-id-enumeration-values-table`. 7598 ``.amdhsa_next_free_vgpr`` Required GFX6-GFX10 Maximum VGPR number explicitly referenced, plus one. 7599 Used to calculate GRANULATED_WORKITEM_VGPR_COUNT in 7600 :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`. 7601 ``.amdhsa_next_free_sgpr`` Required GFX6-GFX10 Maximum SGPR number explicitly referenced, plus one. 7602 Used to calculate GRANULATED_WAVEFRONT_SGPR_COUNT in 7603 :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`. 7604 ``.amdhsa_reserve_vcc`` 1 GFX6-GFX10 Whether the kernel may use the special VCC SGPR. 7605 Used to calculate GRANULATED_WAVEFRONT_SGPR_COUNT in 7606 :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`. 7607 ``.amdhsa_reserve_flat_scratch`` 1 GFX7-GFX10 Whether the kernel may use flat instructions to access 7608 scratch memory. Used to calculate 7609 GRANULATED_WAVEFRONT_SGPR_COUNT in 7610 :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`. 7611 ``.amdhsa_reserve_xnack_mask`` Target GFX8-GFX10 Whether the kernel may trigger XNACK replay. 7612 Feature Used to calculate GRANULATED_WAVEFRONT_SGPR_COUNT in 7613 Specific :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`. 7614 (+xnack) 7615 ``.amdhsa_float_round_mode_32`` 0 GFX6-GFX10 Controls FLOAT_ROUND_MODE_32 in 7616 :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`. 7617 Possible values are defined in 7618 :ref:`amdgpu-amdhsa-floating-point-rounding-mode-enumeration-values-table`. 7619 ``.amdhsa_float_round_mode_16_64`` 0 GFX6-GFX10 Controls FLOAT_ROUND_MODE_16_64 in 7620 :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`. 7621 Possible values are defined in 7622 :ref:`amdgpu-amdhsa-floating-point-rounding-mode-enumeration-values-table`. 7623 ``.amdhsa_float_denorm_mode_32`` 0 GFX6-GFX10 Controls FLOAT_DENORM_MODE_32 in 7624 :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`. 7625 Possible values are defined in 7626 :ref:`amdgpu-amdhsa-floating-point-denorm-mode-enumeration-values-table`. 7627 ``.amdhsa_float_denorm_mode_16_64`` 3 GFX6-GFX10 Controls FLOAT_DENORM_MODE_16_64 in 7628 :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`. 7629 Possible values are defined in 7630 :ref:`amdgpu-amdhsa-floating-point-denorm-mode-enumeration-values-table`. 7631 ``.amdhsa_dx10_clamp`` 1 GFX6-GFX10 Controls ENABLE_DX10_CLAMP in 7632 :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`. 7633 ``.amdhsa_ieee_mode`` 1 GFX6-GFX10 Controls ENABLE_IEEE_MODE in 7634 :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`. 7635 ``.amdhsa_fp16_overflow`` 0 GFX9-GFX10 Controls FP16_OVFL in 7636 :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`. 7637 ``.amdhsa_workgroup_processor_mode`` Target GFX10 Controls ENABLE_WGP_MODE in 7638 Feature :ref:`amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table`. 7639 Specific 7640 (-cumode) 7641 ``.amdhsa_memory_ordered`` 1 GFX10 Controls MEM_ORDERED in 7642 :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`. 7643 ``.amdhsa_forward_progress`` 0 GFX10 Controls FWD_PROGRESS in 7644 :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`. 7645 ``.amdhsa_exception_fp_ieee_invalid_op`` 0 GFX6-GFX10 Controls ENABLE_EXCEPTION_IEEE_754_FP_INVALID_OPERATION in 7646 :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`. 7647 ``.amdhsa_exception_fp_denorm_src`` 0 GFX6-GFX10 Controls ENABLE_EXCEPTION_FP_DENORMAL_SOURCE in 7648 :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`. 7649 ``.amdhsa_exception_fp_ieee_div_zero`` 0 GFX6-GFX10 Controls ENABLE_EXCEPTION_IEEE_754_FP_DIVISION_BY_ZERO in 7650 :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`. 7651 ``.amdhsa_exception_fp_ieee_overflow`` 0 GFX6-GFX10 Controls ENABLE_EXCEPTION_IEEE_754_FP_OVERFLOW in 7652 :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`. 7653 ``.amdhsa_exception_fp_ieee_underflow`` 0 GFX6-GFX10 Controls ENABLE_EXCEPTION_IEEE_754_FP_UNDERFLOW in 7654 :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`. 7655 ``.amdhsa_exception_fp_ieee_inexact`` 0 GFX6-GFX10 Controls ENABLE_EXCEPTION_IEEE_754_FP_INEXACT in 7656 :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`. 7657 ``.amdhsa_exception_int_div_zero`` 0 GFX6-GFX10 Controls ENABLE_EXCEPTION_INT_DIVIDE_BY_ZERO in 7658 :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`. 7659 ======================================================== =================== ============ =================== 7660 7661.amdgpu_metadata 7662++++++++++++++++ 7663 7664Optional directive which declares the contents of the ``NT_AMDGPU_METADATA`` 7665note record (see :ref:`amdgpu-elf-note-records-table-v3`). 7666 7667The contents must be in the [YAML]_ markup format, with the same structure and 7668semantics described in :ref:`amdgpu-amdhsa-code-object-metadata-v3`. 7669 7670This directive is terminated by an ``.end_amdgpu_metadata`` directive. 7671 7672.. _amdgpu-amdhsa-assembler-example-v3: 7673 7674Code Object V3 Example Source Code (-mattr=+code-object-v3) 7675~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 7676 7677Here is an example of a minimal assembly source file, defining one HSA kernel: 7678 7679.. code:: 7680 :number-lines: 7681 7682 .amdgcn_target "amdgcn-amd-amdhsa--gfx900+xnack" // optional 7683 7684 .text 7685 .globl hello_world 7686 .p2align 8 7687 .type hello_world,@function 7688 hello_world: 7689 s_load_dwordx2 s[0:1], s[0:1] 0x0 7690 v_mov_b32 v0, 3.14159 7691 s_waitcnt lgkmcnt(0) 7692 v_mov_b32 v1, s0 7693 v_mov_b32 v2, s1 7694 flat_store_dword v[1:2], v0 7695 s_endpgm 7696 .Lfunc_end0: 7697 .size hello_world, .Lfunc_end0-hello_world 7698 7699 .rodata 7700 .p2align 6 7701 .amdhsa_kernel hello_world 7702 .amdhsa_user_sgpr_kernarg_segment_ptr 1 7703 .amdhsa_next_free_vgpr .amdgcn.next_free_vgpr 7704 .amdhsa_next_free_sgpr .amdgcn.next_free_sgpr 7705 .end_amdhsa_kernel 7706 7707 .amdgpu_metadata 7708 --- 7709 amdhsa.version: 7710 - 1 7711 - 0 7712 amdhsa.kernels: 7713 - .name: hello_world 7714 .symbol: hello_world.kd 7715 .kernarg_segment_size: 48 7716 .group_segment_fixed_size: 0 7717 .private_segment_fixed_size: 0 7718 .kernarg_segment_align: 4 7719 .wavefront_size: 64 7720 .sgpr_count: 2 7721 .vgpr_count: 3 7722 .max_flat_workgroup_size: 256 7723 ... 7724 .end_amdgpu_metadata 7725 7726If an assembly source file contains multiple kernels and/or functions, the 7727:ref:`amdgpu-amdhsa-assembler-symbol-next_free_vgpr` and 7728:ref:`amdgpu-amdhsa-assembler-symbol-next_free_sgpr` symbols may be reset using 7729the ``.set <symbol>, <expression>`` directive. For example, in the case of two 7730kernels, where ``function1`` is only called from ``kernel1`` it is sufficient 7731to group the function with the kernel that calls it and reset the symbols 7732between the two connected components: 7733 7734.. code:: 7735 :number-lines: 7736 7737 .amdgcn_target "amdgcn-amd-amdhsa--gfx900+xnack" // optional 7738 7739 // gpr tracking symbols are implicitly set to zero 7740 7741 .text 7742 .globl kern0 7743 .p2align 8 7744 .type kern0,@function 7745 kern0: 7746 // ... 7747 s_endpgm 7748 .Lkern0_end: 7749 .size kern0, .Lkern0_end-kern0 7750 7751 .rodata 7752 .p2align 6 7753 .amdhsa_kernel kern0 7754 // ... 7755 .amdhsa_next_free_vgpr .amdgcn.next_free_vgpr 7756 .amdhsa_next_free_sgpr .amdgcn.next_free_sgpr 7757 .end_amdhsa_kernel 7758 7759 // reset symbols to begin tracking usage in func1 and kern1 7760 .set .amdgcn.next_free_vgpr, 0 7761 .set .amdgcn.next_free_sgpr, 0 7762 7763 .text 7764 .hidden func1 7765 .global func1 7766 .p2align 2 7767 .type func1,@function 7768 func1: 7769 // ... 7770 s_setpc_b64 s[30:31] 7771 .Lfunc1_end: 7772 .size func1, .Lfunc1_end-func1 7773 7774 .globl kern1 7775 .p2align 8 7776 .type kern1,@function 7777 kern1: 7778 // ... 7779 s_getpc_b64 s[4:5] 7780 s_add_u32 s4, s4, func1@rel32@lo+4 7781 s_addc_u32 s5, s5, func1@rel32@lo+4 7782 s_swappc_b64 s[30:31], s[4:5] 7783 // ... 7784 s_endpgm 7785 .Lkern1_end: 7786 .size kern1, .Lkern1_end-kern1 7787 7788 .rodata 7789 .p2align 6 7790 .amdhsa_kernel kern1 7791 // ... 7792 .amdhsa_next_free_vgpr .amdgcn.next_free_vgpr 7793 .amdhsa_next_free_sgpr .amdgcn.next_free_sgpr 7794 .end_amdhsa_kernel 7795 7796These symbols cannot identify connected components in order to automatically 7797track the usage for each kernel. However, in some cases careful organization of 7798the kernels and functions in the source file means there is minimal additional 7799effort required to accurately calculate GPR usage. 7800 7801Additional Documentation 7802======================== 7803 7804.. [AMD-GCN-GFX6] `AMD Southern Islands Series ISA <http://developer.amd.com/wordpress/media/2012/12/AMD_Southern_Islands_Instruction_Set_Architecture.pdf>`__ 7805.. [AMD-GCN-GFX7] `AMD Sea Islands Series ISA <http://developer.amd.com/wordpress/media/2013/07/AMD_Sea_Islands_Instruction_Set_Architecture.pdf>`_ 7806.. [AMD-GCN-GFX8] `AMD GCN3 Instruction Set Architecture <http://amd-dev.wpengine.netdna-cdn.com/wordpress/media/2013/12/AMD_GCN3_Instruction_Set_Architecture_rev1.1.pdf>`__ 7807.. [AMD-GCN-GFX9] `AMD "Vega" Instruction Set Architecture <http://developer.amd.com/wordpress/media/2013/12/Vega_Shader_ISA_28July2017.pdf>`__ 7808.. [AMD-GCN-GFX10] `AMD "RDNA 1.0" Instruction Set Architecture <https://gpuopen.com/wp-content/uploads/2019/08/RDNA_Shader_ISA_5August2019.pdf>`__ 7809.. [AMD-RADEON-HD-2000-3000] `AMD R6xx shader ISA <http://developer.amd.com/wordpress/media/2012/10/R600_Instruction_Set_Architecture.pdf>`__ 7810.. [AMD-RADEON-HD-4000] `AMD R7xx shader ISA <http://developer.amd.com/wordpress/media/2012/10/R700-Family_Instruction_Set_Architecture.pdf>`__ 7811.. [AMD-RADEON-HD-5000] `AMD Evergreen shader ISA <http://developer.amd.com/wordpress/media/2012/10/AMD_Evergreen-Family_Instruction_Set_Architecture.pdf>`__ 7812.. [AMD-RADEON-HD-6000] `AMD Cayman/Trinity shader ISA <http://developer.amd.com/wordpress/media/2012/10/AMD_HD_6900_Series_Instruction_Set_Architecture.pdf>`__ 7813.. [AMD-ROCm] `AMD ROCm Platform <https://rocm-documentation.readthedocs.io>`__ 7814.. [AMD-ROCm-github] `ROCm github <http://github.com/RadeonOpenCompute>`__ 7815.. [CLANG-ATTR] `Attributes in Clang <https://clang.llvm.org/docs/AttributeReference.html>`__ 7816.. [DWARF] `DWARF Debugging Information Format <http://dwarfstd.org/>`__ 7817.. [ELF] `Executable and Linkable Format (ELF) <http://www.sco.com/developers/gabi/>`__ 7818.. [HRF] `Heterogeneous-race-free Memory Models <http://benedictgaster.org/wp-content/uploads/2014/01/asplos269-FINAL.pdf>`__ 7819.. [HSA] `Heterogeneous System Architecture (HSA) Foundation <http://www.hsafoundation.com/>`__ 7820.. [MsgPack] `Message Pack <http://www.msgpack.org/>`__ 7821.. [OpenCL] `The OpenCL Specification Version 2.0 <http://www.khronos.org/registry/cl/specs/opencl-2.0.pdf>`__ 7822.. [SEMVER] `Semantic Versioning <https://semver.org/>`__ 7823.. [YAML] `YAML Ain't Markup Language (YAML™) Version 1.2 <http://www.yaml.org/spec/1.2/spec.html>`__ 7824