1============================= 2User Guide for AMDGPU Backend 3============================= 4 5.. contents:: 6 :local: 7 8Introduction 9============ 10 11The AMDGPU backend provides ISA code generation for AMD GPUs, starting with the 12R600 family up until the current GCN families. It lives in the 13``llvm/lib/Target/AMDGPU`` directory. 14 15LLVM 16==== 17 18.. _amdgpu-target-triples: 19 20Target Triples 21-------------- 22 23Use the ``clang -target <Architecture>-<Vendor>-<OS>-<Environment>`` option to 24specify the target triple: 25 26 .. table:: AMDGPU Architectures 27 :name: amdgpu-architecture-table 28 29 ============ ============================================================== 30 Architecture Description 31 ============ ============================================================== 32 ``r600`` AMD GPUs HD2XXX-HD6XXX for graphics and compute shaders. 33 ``amdgcn`` AMD GPUs GCN GFX6 onwards for graphics and compute shaders. 34 ============ ============================================================== 35 36 .. table:: AMDGPU Vendors 37 :name: amdgpu-vendor-table 38 39 ============ ============================================================== 40 Vendor Description 41 ============ ============================================================== 42 ``amd`` Can be used for all AMD GPU usage. 43 ``mesa3d`` Can be used if the OS is ``mesa3d``. 44 ============ ============================================================== 45 46 .. table:: AMDGPU Operating Systems 47 :name: amdgpu-os-table 48 49 ============== ============================================================ 50 OS Description 51 ============== ============================================================ 52 *<empty>* Defaults to the *unknown* OS. 53 ``amdhsa`` Compute kernels executed on HSA [HSA]_ compatible runtimes 54 such as AMD's ROCm [AMD-ROCm]_. 55 ``amdpal`` Graphic shaders and compute kernels executed on AMD PAL 56 runtime. 57 ``mesa3d`` Graphic shaders and compute kernels executed on Mesa 3D 58 runtime. 59 ============== ============================================================ 60 61 .. table:: AMDGPU Environments 62 :name: amdgpu-environment-table 63 64 ============ ============================================================== 65 Environment Description 66 ============ ============================================================== 67 *<empty>* Default. 68 ============ ============================================================== 69 70.. _amdgpu-processors: 71 72Processors 73---------- 74 75Use the ``clang -mcpu <Processor>`` option to specify the AMDGPU processor. The 76names from both the *Processor* and *Alternative Processor* can be used. 77 78 .. table:: AMDGPU Processors 79 :name: amdgpu-processor-table 80 81 =========== =============== ============ ===== ================= ======= ====================== 82 Processor Alternative Target dGPU/ Target ROCm Example 83 Processor Triple APU Features Support Products 84 Architecture Supported 85 [Default] 86 =========== =============== ============ ===== ================= ======= ====================== 87 **Radeon HD 2000/3000 Series (R600)** [AMD-RADEON-HD-2000-3000]_ 88 ----------------------------------------------------------------------------------------------- 89 ``r600`` ``r600`` dGPU 90 ``r630`` ``r600`` dGPU 91 ``rs880`` ``r600`` dGPU 92 ``rv670`` ``r600`` dGPU 93 **Radeon HD 4000 Series (R700)** [AMD-RADEON-HD-4000]_ 94 ----------------------------------------------------------------------------------------------- 95 ``rv710`` ``r600`` dGPU 96 ``rv730`` ``r600`` dGPU 97 ``rv770`` ``r600`` dGPU 98 **Radeon HD 5000 Series (Evergreen)** [AMD-RADEON-HD-5000]_ 99 ----------------------------------------------------------------------------------------------- 100 ``cedar`` ``r600`` dGPU 101 ``cypress`` ``r600`` dGPU 102 ``juniper`` ``r600`` dGPU 103 ``redwood`` ``r600`` dGPU 104 ``sumo`` ``r600`` dGPU 105 **Radeon HD 6000 Series (Northern Islands)** [AMD-RADEON-HD-6000]_ 106 ----------------------------------------------------------------------------------------------- 107 ``barts`` ``r600`` dGPU 108 ``caicos`` ``r600`` dGPU 109 ``cayman`` ``r600`` dGPU 110 ``turks`` ``r600`` dGPU 111 **GCN GFX6 (Southern Islands (SI))** [AMD-GCN-GFX6]_ 112 ----------------------------------------------------------------------------------------------- 113 ``gfx600`` - ``tahiti`` ``amdgcn`` dGPU 114 ``gfx601`` - ``hainan`` ``amdgcn`` dGPU 115 - ``oland`` 116 - ``pitcairn`` 117 - ``verde`` 118 **GCN GFX7 (Sea Islands (CI))** [AMD-GCN-GFX7]_ 119 ----------------------------------------------------------------------------------------------- 120 ``gfx700`` - ``kaveri`` ``amdgcn`` APU - A6-7000 121 - A6 Pro-7050B 122 - A8-7100 123 - A8 Pro-7150B 124 - A10-7300 125 - A10 Pro-7350B 126 - FX-7500 127 - A8-7200P 128 - A10-7400P 129 - FX-7600P 130 ``gfx701`` - ``hawaii`` ``amdgcn`` dGPU ROCm - FirePro W8100 131 - FirePro W9100 132 - FirePro S9150 133 - FirePro S9170 134 ``gfx702`` ``amdgcn`` dGPU ROCm - Radeon R9 290 135 - Radeon R9 290x 136 - Radeon R390 137 - Radeon R390x 138 ``gfx703`` - ``kabini`` ``amdgcn`` APU - E1-2100 139 - ``mullins`` - E1-2200 140 - E1-2500 141 - E2-3000 142 - E2-3800 143 - A4-5000 144 - A4-5100 145 - A6-5200 146 - A4 Pro-3340B 147 ``gfx704`` - ``bonaire`` ``amdgcn`` dGPU - Radeon HD 7790 148 - Radeon HD 8770 149 - R7 260 150 - R7 260X 151 **GCN GFX8 (Volcanic Islands (VI))** [AMD-GCN-GFX8]_ 152 ----------------------------------------------------------------------------------------------- 153 ``gfx801`` - ``carrizo`` ``amdgcn`` APU - xnack - A6-8500P 154 [on] - Pro A6-8500B 155 - A8-8600P 156 - Pro A8-8600B 157 - FX-8800P 158 - Pro A12-8800B 159 \ ``amdgcn`` APU - xnack ROCm - A10-8700P 160 [on] - Pro A10-8700B 161 - A10-8780P 162 \ ``amdgcn`` APU - xnack - A10-9600P 163 [on] - A10-9630P 164 - A12-9700P 165 - A12-9730P 166 - FX-9800P 167 - FX-9830P 168 \ ``amdgcn`` APU - xnack - E2-9010 169 [on] - A6-9210 170 - A9-9410 171 ``gfx802`` - ``iceland`` ``amdgcn`` dGPU - xnack ROCm - FirePro S7150 172 - ``tonga`` [off] - FirePro S7100 173 - FirePro W7100 174 - Radeon R285 175 - Radeon R9 380 176 - Radeon R9 385 177 - Mobile FirePro 178 M7170 179 ``gfx803`` - ``fiji`` ``amdgcn`` dGPU - xnack ROCm - Radeon R9 Nano 180 [off] - Radeon R9 Fury 181 - Radeon R9 FuryX 182 - Radeon Pro Duo 183 - FirePro S9300x2 184 - Radeon Instinct MI8 185 \ - ``polaris10`` ``amdgcn`` dGPU - xnack ROCm - Radeon RX 470 186 [off] - Radeon RX 480 187 - Radeon Instinct MI6 188 \ - ``polaris11`` ``amdgcn`` dGPU - xnack ROCm - Radeon RX 460 189 [off] 190 ``gfx810`` - ``stoney`` ``amdgcn`` APU - xnack 191 [on] 192 **GCN GFX9** [AMD-GCN-GFX9]_ 193 ----------------------------------------------------------------------------------------------- 194 ``gfx900`` ``amdgcn`` dGPU - xnack ROCm - Radeon Vega 195 [off] Frontier Edition 196 - Radeon RX Vega 56 197 - Radeon RX Vega 64 198 - Radeon RX Vega 64 199 Liquid 200 - Radeon Instinct MI25 201 ``gfx902`` ``amdgcn`` APU - xnack - Ryzen 3 2200G 202 [on] - Ryzen 5 2400G 203 ``gfx904`` ``amdgcn`` dGPU - xnack *TBA* 204 [off] 205 .. TODO:: 206 Add product 207 names. 208 ``gfx906`` ``amdgcn`` dGPU - xnack - Radeon Instinct MI50 209 [off] - Radeon Instinct MI60 210 ``gfx908`` ``amdgcn`` dGPU - xnack *TBA* 211 [off] 212 sram-ecc 213 [on] 214 ``gfx909`` ``amdgcn`` APU - xnack *TBA* (Raven Ridge 2) 215 [on] 216 .. TODO:: 217 Add product 218 names. 219 **GCN GFX10** [AMD-GCN-GFX10]_ 220 ----------------------------------------------------------------------------------------------- 221 ``gfx1010`` ``amdgcn`` dGPU - xnack *TBA* 222 [off] 223 - wavefrontsize64 224 [off] 225 - cumode 226 [off] 227 .. TODO:: 228 Add product 229 names. 230 ``gfx1011`` ``amdgcn`` dGPU - xnack *TBA* 231 [off] 232 - wavefrontsize64 233 [off] 234 - cumode 235 [off] 236 .. TODO:: 237 Add product 238 names. 239 ``gfx1012`` ``amdgcn`` dGPU - xnack *TBA* 240 [off] 241 - wavefrontsize64 242 [off] 243 - cumode 244 [off] 245 .. TODO:: 246 Add product 247 names. 248 =========== =============== ============ ===== ================= ======= ====================== 249 250.. _amdgpu-target-features: 251 252Target Features 253--------------- 254 255Target features control how code is generated to support certain 256processor specific features. Not all target features are supported by 257all processors. The runtime must ensure that the features supported by 258the device used to execute the code match the features enabled when 259generating the code. A mismatch of features may result in incorrect 260execution, or a reduction in performance. 261 262The target features supported by each processor, and the default value 263used if not specified explicitly, is listed in 264:ref:`amdgpu-processor-table`. 265 266Use the ``clang -m[no-]<TargetFeature>`` option to specify the AMDGPU 267target features. 268 269For example: 270 271``-mxnack`` 272 Enable the ``xnack`` feature. 273``-mno-xnack`` 274 Disable the ``xnack`` feature. 275 276 .. table:: AMDGPU Target Features 277 :name: amdgpu-target-feature-table 278 279 ====================== ================================================== 280 Target Feature Description 281 ====================== ================================================== 282 -m[no-]xnack Enable/disable generating code that has 283 memory clauses that are compatible with 284 having XNACK replay enabled. 285 286 This is used for demand paging and page 287 migration. If XNACK replay is enabled in 288 the device, then if a page fault occurs 289 the code may execute incorrectly if the 290 ``xnack`` feature is not enabled. Executing 291 code that has the feature enabled on a 292 device that does not have XNACK replay 293 enabled will execute correctly, but may 294 be less performant than code with the 295 feature disabled. 296 297 -m[no-]sram-ecc Enable/disable generating code that assumes SRAM 298 ECC is enabled/disabled. 299 300 -m[no-]wavefrontsize64 Control the default wavefront size used when 301 generating code for kernels. When disabled 302 native wavefront size 32 is used, when enabled 303 wavefront size 64 is used. 304 305 -m[no-]cumode Control the default wavefront execution mode used 306 when generating code for kernels. When disabled 307 native WGP wavefront execution mode is used, 308 when enabled CU wavefront execution mode is used 309 (see :ref:`amdgpu-amdhsa-memory-model`). 310 ====================== ================================================== 311 312.. _amdgpu-address-spaces: 313 314Address Spaces 315-------------- 316 317The AMDGPU architecture supports a number of memory address spaces. The address 318space names use the OpenCL standard names, with some additions. 319 320The AMDGPU address spaces correspond to architecture-specific LLVM address 321space numbers used in LLVM IR. 322 323The AMDGPU address spaces are described in 324:ref:`amdgpu-address-spaces-table`. Only 64-bit process address spaces are 325supported for the ``amdgcn`` target. 326 327 .. table:: AMDGPU Address Spaces 328 :name: amdgpu-address-spaces-table 329 330 ================================= =============== =========== ================ ======= ============================ 331 .. 64-Bit Process Address Space 332 --------------------------------- --------------- ----------- ---------------- ------------------------------------ 333 Address Space Name LLVM IR Address HSA Segment Hardware Address NULL Value 334 Space Number Name Name Size 335 ================================= =============== =========== ================ ======= ============================ 336 Generic 0 flat flat 64 0x0000000000000000 337 Global 1 global global 64 0x0000000000000000 338 Region 2 N/A GDS 32 *not implemented for AMDHSA* 339 Local 3 group LDS 32 0xFFFFFFFF 340 Constant 4 constant *same as global* 64 0x0000000000000000 341 Private 5 private scratch 32 0x00000000 342 Constant 32-bit 6 *TODO* 343 Buffer Fat Pointer (experimental) 7 *TODO* 344 ================================= =============== =========== ================ ======= ============================ 345 346**Generic** 347 The generic address space uses the hardware flat address support available in 348 GFX7-GFX10. This uses two fixed ranges of virtual addresses (the private and 349 local apertures), that are outside the range of addressable global memory, to 350 map from a flat address to a private or local address. 351 352 FLAT instructions can take a flat address and access global, private 353 (scratch), and group (LDS) memory depending on if the address is within one 354 of the aperture ranges. Flat access to scratch requires hardware aperture 355 setup and setup in the kernel prologue (see 356 :ref:`amdgpu-amdhsa-kernel-prolog-flat-scratch`). Flat access to LDS requires 357 hardware aperture setup and M0 (GFX7-GFX8) register setup (see 358 :ref:`amdgpu-amdhsa-kernel-prolog-m0`). 359 360 To convert between a private or group address space address (termed a segment 361 address) and a flat address the base address of the corresponding aperture 362 can be used. For GFX7-GFX8 these are available in the 363 :ref:`amdgpu-amdhsa-hsa-aql-queue` the address of which can be obtained with 364 Queue Ptr SGPR (see :ref:`amdgpu-amdhsa-initial-kernel-execution-state`). For 365 GFX9-GFX10 the aperture base addresses are directly available as inline 366 constant registers ``SRC_SHARED_BASE/LIMIT`` and ``SRC_PRIVATE_BASE/LIMIT``. 367 In 64-bit address mode the aperture sizes are 2^32 bytes and the base is 368 aligned to 2^32 which makes it easier to convert from flat to segment or 369 segment to flat. 370 371 A global address space address has the same value when used as a flat address 372 so no conversion is needed. 373 374**Global and Constant** 375 The global and constant address spaces both use global virtual addresses, 376 which are the same virtual address space used by the CPU. However, some 377 virtual addresses may only be accessible to the CPU, some only accessible 378 by the GPU, and some by both. 379 380 Using the constant address space indicates that the data will not change 381 during the execution of the kernel. This allows scalar read instructions to 382 be used. The vector and scalar L1 caches are invalidated of volatile data 383 before each kernel dispatch execution to allow constant memory to change 384 values between kernel dispatches. 385 386**Region** 387 The region address space uses the hardware Global Data Store (GDS). All 388 wavefronts executing on the same device will access the same memory for any 389 given region address. However, the same region address accessed by wavefronts 390 executing on different devices will access different memory. It is higher 391 performance than global memory. It is allocated by the runtime. The data 392 store (DS) instructions can be used to access it. 393 394**Local** 395 The local address space uses the hardware Local Data Store (LDS) which is 396 automatically allocated when the hardware creates the wavefronts of a 397 work-group, and freed when all the wavefronts of a work-group have 398 terminated. All wavefronts belonging to the same work-group will access the 399 same memory for any given local address. However, the same local address 400 accessed by wavefronts belonging to different work-groups will access 401 different memory. It is higher performance than global memory. The data store 402 (DS) instructions can be used to access it. 403 404**Private** 405 The private address space uses the hardware scratch memory support which 406 automatically allocates memory when it creates a wavefront, and frees it when 407 a wavefronts terminates. The memory accessed by a lane of a wavefront for any 408 given private address will be different to the memory accessed by another lane 409 of the same or different wavefront for the same private address. 410 411 If a kernel dispatch uses scratch, then the hardware allocates memory from a 412 pool of backing memory allocated by the runtime for each wavefront. The lanes 413 of the wavefront access this using dword (4 byte) interleaving. The mapping 414 used from private address to backing memory address is: 415 416 ``wavefront-scratch-base + 417 ((private-address / 4) * wavefront-size * 4) + 418 (wavefront-lane-id * 4) + (private-address % 4)`` 419 420 If each lane of a wavefront accesses the same private address, the 421 interleaving results in adjacent dwords being accessed and hence requires 422 fewer cache lines to be fetched. 423 424 There are different ways that the wavefront scratch base address is 425 determined by a wavefront (see 426 :ref:`amdgpu-amdhsa-initial-kernel-execution-state`). 427 428 Scratch memory can be accessed in an interleaved manner using buffer 429 instructions with the scratch buffer descriptor and per wavefront scratch 430 offset, by the scratch instructions, or by flat instructions. Multi-dword 431 access is not supported except by flat and scratch instructions in 432 GFX9-GFX10. 433 434**Constant 32-bit** 435 *TODO* 436 437**Buffer Fat Pointer** 438 The buffer fat pointer is an experimental address space that is currently 439 unsupported in the backend. It exposes a non-integral pointer that is in 440 the future intended to support the modelling of 128-bit buffer descriptors 441 plus a 32-bit offset into the buffer (in total encapsulating a 160-bit 442 *pointer*), allowing normal LLVM load/store/atomic operations to be used to 443 model the buffer descriptors used heavily in graphics workloads targeting 444 the backend. 445 446.. _amdgpu-memory-scopes: 447 448Memory Scopes 449------------- 450 451This section provides LLVM memory synchronization scopes supported by the AMDGPU 452backend memory model when the target triple OS is ``amdhsa`` (see 453:ref:`amdgpu-amdhsa-memory-model` and :ref:`amdgpu-target-triples`). 454 455The memory model supported is based on the HSA memory model [HSA]_ which is 456based in turn on HRF-indirect with scope inclusion [HRF]_. The happens-before 457relation is transitive over the synchronizes-with relation independent of scope, 458and synchronizes-with allows the memory scope instances to be inclusive (see 459table :ref:`amdgpu-amdhsa-llvm-sync-scopes-table`). 460 461This is different to the OpenCL [OpenCL]_ memory model which does not have scope 462inclusion and requires the memory scopes to exactly match. However, this 463is conservatively correct for OpenCL. 464 465 .. table:: AMDHSA LLVM Sync Scopes 466 :name: amdgpu-amdhsa-llvm-sync-scopes-table 467 468 ======================= =================================================== 469 LLVM Sync Scope Description 470 ======================= =================================================== 471 *none* The default: ``system``. 472 473 Synchronizes with, and participates in modification 474 and seq_cst total orderings with, other operations 475 (except image operations) for all address spaces 476 (except private, or generic that accesses private) 477 provided the other operation's sync scope is: 478 479 - ``system``. 480 - ``agent`` and executed by a thread on the same 481 agent. 482 - ``workgroup`` and executed by a thread in the 483 same work-group. 484 - ``wavefront`` and executed by a thread in the 485 same wavefront. 486 487 ``agent`` Synchronizes with, and participates in modification 488 and seq_cst total orderings with, other operations 489 (except image operations) for all address spaces 490 (except private, or generic that accesses private) 491 provided the other operation's sync scope is: 492 493 - ``system`` or ``agent`` and executed by a thread 494 on the same agent. 495 - ``workgroup`` and executed by a thread in the 496 same work-group. 497 - ``wavefront`` and executed by a thread in the 498 same wavefront. 499 500 ``workgroup`` Synchronizes with, and participates in modification 501 and seq_cst total orderings with, other operations 502 (except image operations) for all address spaces 503 (except private, or generic that accesses private) 504 provided the other operation's sync scope is: 505 506 - ``system``, ``agent`` or ``workgroup`` and 507 executed by a thread in the same work-group. 508 - ``wavefront`` and executed by a thread in the 509 same wavefront. 510 511 ``wavefront`` Synchronizes with, and participates in modification 512 and seq_cst total orderings with, other operations 513 (except image operations) for all address spaces 514 (except private, or generic that accesses private) 515 provided the other operation's sync scope is: 516 517 - ``system``, ``agent``, ``workgroup`` or 518 ``wavefront`` and executed by a thread in the 519 same wavefront. 520 521 ``singlethread`` Only synchronizes with, and participates in 522 modification and seq_cst total orderings with, 523 other operations (except image operations) running 524 in the same thread for all address spaces (for 525 example, in signal handlers). 526 527 ``one-as`` Same as ``system`` but only synchronizes with other 528 operations within the same address space. 529 530 ``agent-one-as`` Same as ``agent`` but only synchronizes with other 531 operations within the same address space. 532 533 ``workgroup-one-as`` Same as ``workgroup`` but only synchronizes with 534 other operations within the same address space. 535 536 ``wavefront-one-as`` Same as ``wavefront`` but only synchronizes with 537 other operations within the same address space. 538 539 ``singlethread-one-as`` Same as ``singlethread`` but only synchronizes with 540 other operations within the same address space. 541 ======================= =================================================== 542 543AMDGPU Intrinsics 544----------------- 545 546The AMDGPU backend implements the following LLVM IR intrinsics. 547 548*This section is WIP.* 549 550.. TODO:: 551 552 List AMDGPU intrinsics. 553 554AMDGPU Attributes 555----------------- 556 557The AMDGPU backend supports the following LLVM IR attributes. 558 559 .. table:: AMDGPU LLVM IR Attributes 560 :name: amdgpu-llvm-ir-attributes-table 561 562 ======================================= ========================================================== 563 LLVM Attribute Description 564 ======================================= ========================================================== 565 "amdgpu-flat-work-group-size"="min,max" Specify the minimum and maximum flat work group sizes that 566 will be specified when the kernel is dispatched. Generated 567 by the ``amdgpu_flat_work_group_size`` CLANG attribute [CLANG-ATTR]_. 568 "amdgpu-implicitarg-num-bytes"="n" Number of kernel argument bytes to add to the kernel 569 argument block size for the implicit arguments. This 570 varies by OS and language (for OpenCL see 571 :ref:`opencl-kernel-implicit-arguments-appended-for-amdhsa-os-table`). 572 "amdgpu-num-sgpr"="n" Specifies the number of SGPRs to use. Generated by 573 the ``amdgpu_num_sgpr`` CLANG attribute [CLANG-ATTR]_. 574 "amdgpu-num-vgpr"="n" Specifies the number of VGPRs to use. Generated by the 575 ``amdgpu_num_vgpr`` CLANG attribute [CLANG-ATTR]_. 576 "amdgpu-waves-per-eu"="m,n" Specify the minimum and maximum number of waves per 577 execution unit. Generated by the ``amdgpu_waves_per_eu`` 578 CLANG attribute [CLANG-ATTR]_. 579 "amdgpu-ieee" true/false. Specify whether the function expects the IEEE field of the 580 mode register to be set on entry. Overrides the default for 581 the calling convention. 582 "amdgpu-dx10-clamp" true/false. Specify whether the function expects the DX10_CLAMP field of 583 the mode register to be set on entry. Overrides the default 584 for the calling convention. 585 ======================================= ========================================================== 586 587Code Object 588=========== 589 590The AMDGPU backend generates a standard ELF [ELF]_ relocatable code object that 591can be linked by ``lld`` to produce a standard ELF shared code object which can 592be loaded and executed on an AMDGPU target. 593 594Header 595------ 596 597The AMDGPU backend uses the following ELF header: 598 599 .. table:: AMDGPU ELF Header 600 :name: amdgpu-elf-header-table 601 602 ========================== =============================== 603 Field Value 604 ========================== =============================== 605 ``e_ident[EI_CLASS]`` ``ELFCLASS64`` 606 ``e_ident[EI_DATA]`` ``ELFDATA2LSB`` 607 ``e_ident[EI_OSABI]`` - ``ELFOSABI_NONE`` 608 - ``ELFOSABI_AMDGPU_HSA`` 609 - ``ELFOSABI_AMDGPU_PAL`` 610 - ``ELFOSABI_AMDGPU_MESA3D`` 611 ``e_ident[EI_ABIVERSION]`` - ``ELFABIVERSION_AMDGPU_HSA`` 612 - ``ELFABIVERSION_AMDGPU_PAL`` 613 - ``ELFABIVERSION_AMDGPU_MESA3D`` 614 ``e_type`` - ``ET_REL`` 615 - ``ET_DYN`` 616 ``e_machine`` ``EM_AMDGPU`` 617 ``e_entry`` 0 618 ``e_flags`` See :ref:`amdgpu-elf-header-e_flags-table` 619 ========================== =============================== 620 621.. 622 623 .. table:: AMDGPU ELF Header Enumeration Values 624 :name: amdgpu-elf-header-enumeration-values-table 625 626 =============================== ===== 627 Name Value 628 =============================== ===== 629 ``EM_AMDGPU`` 224 630 ``ELFOSABI_NONE`` 0 631 ``ELFOSABI_AMDGPU_HSA`` 64 632 ``ELFOSABI_AMDGPU_PAL`` 65 633 ``ELFOSABI_AMDGPU_MESA3D`` 66 634 ``ELFABIVERSION_AMDGPU_HSA`` 1 635 ``ELFABIVERSION_AMDGPU_PAL`` 0 636 ``ELFABIVERSION_AMDGPU_MESA3D`` 0 637 =============================== ===== 638 639``e_ident[EI_CLASS]`` 640 The ELF class is: 641 642 * ``ELFCLASS32`` for ``r600`` architecture. 643 644 * ``ELFCLASS64`` for ``amdgcn`` architecture which only supports 64-bit 645 process address space applications. 646 647``e_ident[EI_DATA]`` 648 All AMDGPU targets use ``ELFDATA2LSB`` for little-endian byte ordering. 649 650``e_ident[EI_OSABI]`` 651 One of the following AMDGPU architecture specific OS ABIs 652 (see :ref:`amdgpu-os-table`): 653 654 * ``ELFOSABI_NONE`` for *unknown* OS. 655 656 * ``ELFOSABI_AMDGPU_HSA`` for ``amdhsa`` OS. 657 658 * ``ELFOSABI_AMDGPU_PAL`` for ``amdpal`` OS. 659 660 * ``ELFOSABI_AMDGPU_MESA3D`` for ``mesa3D`` OS. 661 662``e_ident[EI_ABIVERSION]`` 663 The ABI version of the AMDGPU architecture specific OS ABI to which the code 664 object conforms: 665 666 * ``ELFABIVERSION_AMDGPU_HSA`` is used to specify the version of AMD HSA 667 runtime ABI. 668 669 * ``ELFABIVERSION_AMDGPU_PAL`` is used to specify the version of AMD PAL 670 runtime ABI. 671 672 * ``ELFABIVERSION_AMDGPU_MESA3D`` is used to specify the version of AMD MESA 673 3D runtime ABI. 674 675``e_type`` 676 Can be one of the following values: 677 678 679 ``ET_REL`` 680 The type produced by the AMDGPU backend compiler as it is relocatable code 681 object. 682 683 ``ET_DYN`` 684 The type produced by the linker as it is a shared code object. 685 686 The AMD HSA runtime loader requires a ``ET_DYN`` code object. 687 688``e_machine`` 689 The value ``EM_AMDGPU`` is used for the machine for all processors supported 690 by the ``r600`` and ``amdgcn`` architectures (see 691 :ref:`amdgpu-processor-table`). The specific processor is specified in the 692 ``EF_AMDGPU_MACH`` bit field of the ``e_flags`` (see 693 :ref:`amdgpu-elf-header-e_flags-table`). 694 695``e_entry`` 696 The entry point is 0 as the entry points for individual kernels must be 697 selected in order to invoke them through AQL packets. 698 699``e_flags`` 700 The AMDGPU backend uses the following ELF header flags: 701 702 .. table:: AMDGPU ELF Header ``e_flags`` 703 :name: amdgpu-elf-header-e_flags-table 704 705 ================================= ========== ============================= 706 Name Value Description 707 ================================= ========== ============================= 708 **AMDGPU Processor Flag** See :ref:`amdgpu-processor-table`. 709 -------------------------------------------- ----------------------------- 710 ``EF_AMDGPU_MACH`` 0x000000ff AMDGPU processor selection 711 mask for 712 ``EF_AMDGPU_MACH_xxx`` values 713 defined in 714 :ref:`amdgpu-ef-amdgpu-mach-table`. 715 ``EF_AMDGPU_XNACK`` 0x00000100 Indicates if the ``xnack`` 716 target feature is 717 enabled for all code 718 contained in the code object. 719 If the processor 720 does not support the 721 ``xnack`` target 722 feature then must 723 be 0. 724 See 725 :ref:`amdgpu-target-features`. 726 ``EF_AMDGPU_SRAM_ECC`` 0x00000200 Indicates if the ``sram-ecc`` 727 target feature is 728 enabled for all code 729 contained in the code object. 730 If the processor 731 does not support the 732 ``sram-ecc`` target 733 feature then must 734 be 0. 735 See 736 :ref:`amdgpu-target-features`. 737 ================================= ========== ============================= 738 739 .. table:: AMDGPU ``EF_AMDGPU_MACH`` Values 740 :name: amdgpu-ef-amdgpu-mach-table 741 742 ================================= ========== ============================= 743 Name Value Description (see 744 :ref:`amdgpu-processor-table`) 745 ================================= ========== ============================= 746 ``EF_AMDGPU_MACH_NONE`` 0x000 *not specified* 747 ``EF_AMDGPU_MACH_R600_R600`` 0x001 ``r600`` 748 ``EF_AMDGPU_MACH_R600_R630`` 0x002 ``r630`` 749 ``EF_AMDGPU_MACH_R600_RS880`` 0x003 ``rs880`` 750 ``EF_AMDGPU_MACH_R600_RV670`` 0x004 ``rv670`` 751 ``EF_AMDGPU_MACH_R600_RV710`` 0x005 ``rv710`` 752 ``EF_AMDGPU_MACH_R600_RV730`` 0x006 ``rv730`` 753 ``EF_AMDGPU_MACH_R600_RV770`` 0x007 ``rv770`` 754 ``EF_AMDGPU_MACH_R600_CEDAR`` 0x008 ``cedar`` 755 ``EF_AMDGPU_MACH_R600_CYPRESS`` 0x009 ``cypress`` 756 ``EF_AMDGPU_MACH_R600_JUNIPER`` 0x00a ``juniper`` 757 ``EF_AMDGPU_MACH_R600_REDWOOD`` 0x00b ``redwood`` 758 ``EF_AMDGPU_MACH_R600_SUMO`` 0x00c ``sumo`` 759 ``EF_AMDGPU_MACH_R600_BARTS`` 0x00d ``barts`` 760 ``EF_AMDGPU_MACH_R600_CAICOS`` 0x00e ``caicos`` 761 ``EF_AMDGPU_MACH_R600_CAYMAN`` 0x00f ``cayman`` 762 ``EF_AMDGPU_MACH_R600_TURKS`` 0x010 ``turks`` 763 *reserved* 0x011 - Reserved for ``r600`` 764 0x01f architecture processors. 765 ``EF_AMDGPU_MACH_AMDGCN_GFX600`` 0x020 ``gfx600`` 766 ``EF_AMDGPU_MACH_AMDGCN_GFX601`` 0x021 ``gfx601`` 767 ``EF_AMDGPU_MACH_AMDGCN_GFX700`` 0x022 ``gfx700`` 768 ``EF_AMDGPU_MACH_AMDGCN_GFX701`` 0x023 ``gfx701`` 769 ``EF_AMDGPU_MACH_AMDGCN_GFX702`` 0x024 ``gfx702`` 770 ``EF_AMDGPU_MACH_AMDGCN_GFX703`` 0x025 ``gfx703`` 771 ``EF_AMDGPU_MACH_AMDGCN_GFX704`` 0x026 ``gfx704`` 772 *reserved* 0x027 Reserved. 773 ``EF_AMDGPU_MACH_AMDGCN_GFX801`` 0x028 ``gfx801`` 774 ``EF_AMDGPU_MACH_AMDGCN_GFX802`` 0x029 ``gfx802`` 775 ``EF_AMDGPU_MACH_AMDGCN_GFX803`` 0x02a ``gfx803`` 776 ``EF_AMDGPU_MACH_AMDGCN_GFX810`` 0x02b ``gfx810`` 777 ``EF_AMDGPU_MACH_AMDGCN_GFX900`` 0x02c ``gfx900`` 778 ``EF_AMDGPU_MACH_AMDGCN_GFX902`` 0x02d ``gfx902`` 779 ``EF_AMDGPU_MACH_AMDGCN_GFX904`` 0x02e ``gfx904`` 780 ``EF_AMDGPU_MACH_AMDGCN_GFX906`` 0x02f ``gfx906`` 781 ``EF_AMDGPU_MACH_AMDGCN_GFX908`` 0x030 ``gfx908`` 782 ``EF_AMDGPU_MACH_AMDGCN_GFX909`` 0x031 ``gfx909`` 783 *reserved* 0x032 Reserved. 784 ``EF_AMDGPU_MACH_AMDGCN_GFX1010`` 0x033 ``gfx1010`` 785 ``EF_AMDGPU_MACH_AMDGCN_GFX1011`` 0x034 ``gfx1011`` 786 ``EF_AMDGPU_MACH_AMDGCN_GFX1012`` 0x035 ``gfx1012`` 787 ================================= ========== ============================= 788 789Sections 790-------- 791 792An AMDGPU target ELF code object has the standard ELF sections which include: 793 794 .. table:: AMDGPU ELF Sections 795 :name: amdgpu-elf-sections-table 796 797 ================== ================ ================================= 798 Name Type Attributes 799 ================== ================ ================================= 800 ``.bss`` ``SHT_NOBITS`` ``SHF_ALLOC`` + ``SHF_WRITE`` 801 ``.data`` ``SHT_PROGBITS`` ``SHF_ALLOC`` + ``SHF_WRITE`` 802 ``.debug_``\ *\** ``SHT_PROGBITS`` *none* 803 ``.dynamic`` ``SHT_DYNAMIC`` ``SHF_ALLOC`` 804 ``.dynstr`` ``SHT_PROGBITS`` ``SHF_ALLOC`` 805 ``.dynsym`` ``SHT_PROGBITS`` ``SHF_ALLOC`` 806 ``.got`` ``SHT_PROGBITS`` ``SHF_ALLOC`` + ``SHF_WRITE`` 807 ``.hash`` ``SHT_HASH`` ``SHF_ALLOC`` 808 ``.note`` ``SHT_NOTE`` *none* 809 ``.rela``\ *name* ``SHT_RELA`` *none* 810 ``.rela.dyn`` ``SHT_RELA`` *none* 811 ``.rodata`` ``SHT_PROGBITS`` ``SHF_ALLOC`` 812 ``.shstrtab`` ``SHT_STRTAB`` *none* 813 ``.strtab`` ``SHT_STRTAB`` *none* 814 ``.symtab`` ``SHT_SYMTAB`` *none* 815 ``.text`` ``SHT_PROGBITS`` ``SHF_ALLOC`` + ``SHF_EXECINSTR`` 816 ================== ================ ================================= 817 818These sections have their standard meanings (see [ELF]_) and are only generated 819if needed. 820 821``.debug``\ *\** 822 The standard DWARF sections. See :ref:`amdgpu-dwarf` for information on the 823 DWARF produced by the AMDGPU backend. 824 825``.dynamic``, ``.dynstr``, ``.dynsym``, ``.hash`` 826 The standard sections used by a dynamic loader. 827 828``.note`` 829 See :ref:`amdgpu-note-records` for the note records supported by the AMDGPU 830 backend. 831 832``.rela``\ *name*, ``.rela.dyn`` 833 For relocatable code objects, *name* is the name of the section that the 834 relocation records apply. For example, ``.rela.text`` is the section name for 835 relocation records associated with the ``.text`` section. 836 837 For linked shared code objects, ``.rela.dyn`` contains all the relocation 838 records from each of the relocatable code object's ``.rela``\ *name* sections. 839 840 See :ref:`amdgpu-relocation-records` for the relocation records supported by 841 the AMDGPU backend. 842 843``.text`` 844 The executable machine code for the kernels and functions they call. Generated 845 as position independent code. See :ref:`amdgpu-code-conventions` for 846 information on conventions used in the isa generation. 847 848.. _amdgpu-note-records: 849 850Note Records 851------------ 852 853The AMDGPU backend code object contains ELF note records in the ``.note`` 854section. The set of generated notes and their semantics depend on the code 855object version; see :ref:`amdgpu-note-records-v2` and 856:ref:`amdgpu-note-records-v3`. 857 858As required by ``ELFCLASS32`` and ``ELFCLASS64``, minimal zero byte padding 859must be generated after the ``name`` field to ensure the ``desc`` field is 4 860byte aligned. In addition, minimal zero byte padding must be generated to 861ensure the ``desc`` field size is a multiple of 4 bytes. The ``sh_addralign`` 862field of the ``.note`` section must be at least 4 to indicate at least 8 byte 863alignment. 864 865.. _amdgpu-note-records-v2: 866 867Code Object V2 Note Records (-mattr=-code-object-v3) 868~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 869 870.. warning:: Code Object V2 is not the default code object version emitted by 871 this version of LLVM. For a description of the notes generated with the 872 default configuration (Code Object V3) see :ref:`amdgpu-note-records-v3`. 873 874The AMDGPU backend code object uses the following ELF note record in the 875``.note`` section when compiling for Code Object V2 (-mattr=-code-object-v3). 876 877Additional note records may be present, but any which are not documented here 878are deprecated and should not be used. 879 880 .. table:: AMDGPU Code Object V2 ELF Note Records 881 :name: amdgpu-elf-note-records-table-v2 882 883 ===== ============================== ====================================== 884 Name Type Description 885 ===== ============================== ====================================== 886 "AMD" ``NT_AMD_AMDGPU_HSA_METADATA`` <metadata null terminated string> 887 ===== ============================== ====================================== 888 889.. 890 891 .. table:: AMDGPU Code Object V2 ELF Note Record Enumeration Values 892 :name: amdgpu-elf-note-record-enumeration-values-table-v2 893 894 ============================== ===== 895 Name Value 896 ============================== ===== 897 *reserved* 0-9 898 ``NT_AMD_AMDGPU_HSA_METADATA`` 10 899 *reserved* 11 900 ============================== ===== 901 902``NT_AMD_AMDGPU_HSA_METADATA`` 903 Specifies extensible metadata associated with the code objects executed on HSA 904 [HSA]_ compatible runtimes such as AMD's ROCm [AMD-ROCm]_. It is required when 905 the target triple OS is ``amdhsa`` (see :ref:`amdgpu-target-triples`). See 906 :ref:`amdgpu-amdhsa-code-object-metadata-v2` for the syntax of the code 907 object metadata string. 908 909.. _amdgpu-note-records-v3: 910 911Code Object V3 Note Records (-mattr=+code-object-v3) 912~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 913 914The AMDGPU backend code object uses the following ELF note record in the 915``.note`` section when compiling for Code Object V3 (-mattr=+code-object-v3). 916 917Additional note records may be present, but any which are not documented here 918are deprecated and should not be used. 919 920 .. table:: AMDGPU Code Object V3 ELF Note Records 921 :name: amdgpu-elf-note-records-table-v3 922 923 ======== ============================== ====================================== 924 Name Type Description 925 ======== ============================== ====================================== 926 "AMDGPU" ``NT_AMDGPU_METADATA`` Metadata in Message Pack [MsgPack]_ 927 binary format. 928 ======== ============================== ====================================== 929 930.. 931 932 .. table:: AMDGPU Code Object V3 ELF Note Record Enumeration Values 933 :name: amdgpu-elf-note-record-enumeration-values-table-v3 934 935 ============================== ===== 936 Name Value 937 ============================== ===== 938 *reserved* 0-31 939 ``NT_AMDGPU_METADATA`` 32 940 ============================== ===== 941 942``NT_AMDGPU_METADATA`` 943 Specifies extensible metadata associated with an AMDGPU code 944 object. It is encoded as a map in the Message Pack [MsgPack]_ binary 945 data format. See :ref:`amdgpu-amdhsa-code-object-metadata-v3` for the 946 map keys defined for the ``amdhsa`` OS. 947 948.. _amdgpu-symbols: 949 950Symbols 951------- 952 953Symbols include the following: 954 955 .. table:: AMDGPU ELF Symbols 956 :name: amdgpu-elf-symbols-table 957 958 ===================== ================== ================ ================== 959 Name Type Section Description 960 ===================== ================== ================ ================== 961 *link-name* ``STT_OBJECT`` - ``.data`` Global variable 962 - ``.rodata`` 963 - ``.bss`` 964 *link-name*\ ``.kd`` ``STT_OBJECT`` - ``.rodata`` Kernel descriptor 965 *link-name* ``STT_FUNC`` - ``.text`` Kernel entry point 966 *link-name* ``STT_OBJECT`` - SHN_AMDGPU_LDS Global variable in LDS 967 ===================== ================== ================ ================== 968 969Global variable 970 Global variables both used and defined by the compilation unit. 971 972 If the symbol is defined in the compilation unit then it is allocated in the 973 appropriate section according to if it has initialized data or is readonly. 974 975 If the symbol is external then its section is ``STN_UNDEF`` and the loader 976 will resolve relocations using the definition provided by another code object 977 or explicitly defined by the runtime. 978 979 If the symbol resides in local/group memory (LDS) then its section is the 980 special processor-specific section name ``SHN_AMDGPU_LDS``, and the 981 ``st_value`` field describes alignment requirements as it does for common 982 symbols. 983 984 .. TODO:: 985 986 Add description of linked shared object symbols. Seems undefined symbols 987 are marked as STT_NOTYPE. 988 989Kernel descriptor 990 Every HSA kernel has an associated kernel descriptor. It is the address of the 991 kernel descriptor that is used in the AQL dispatch packet used to invoke the 992 kernel, not the kernel entry point. The layout of the HSA kernel descriptor is 993 defined in :ref:`amdgpu-amdhsa-kernel-descriptor`. 994 995Kernel entry point 996 Every HSA kernel also has a symbol for its machine code entry point. 997 998.. _amdgpu-relocation-records: 999 1000Relocation Records 1001------------------ 1002 1003AMDGPU backend generates ``Elf64_Rela`` relocation records. Supported 1004relocatable fields are: 1005 1006``word32`` 1007 This specifies a 32-bit field occupying 4 bytes with arbitrary byte 1008 alignment. These values use the same byte order as other word values in the 1009 AMDGPU architecture. 1010 1011``word64`` 1012 This specifies a 64-bit field occupying 8 bytes with arbitrary byte 1013 alignment. These values use the same byte order as other word values in the 1014 AMDGPU architecture. 1015 1016Following notations are used for specifying relocation calculations: 1017 1018**A** 1019 Represents the addend used to compute the value of the relocatable field. 1020 1021**G** 1022 Represents the offset into the global offset table at which the relocation 1023 entry's symbol will reside during execution. 1024 1025**GOT** 1026 Represents the address of the global offset table. 1027 1028**P** 1029 Represents the place (section offset for ``et_rel`` or address for ``et_dyn``) 1030 of the storage unit being relocated (computed using ``r_offset``). 1031 1032**S** 1033 Represents the value of the symbol whose index resides in the relocation 1034 entry. Relocations not using this must specify a symbol index of 1035 ``STN_UNDEF``. 1036 1037**B** 1038 Represents the base address of a loaded executable or shared object which is 1039 the difference between the ELF address and the actual load address. 1040 Relocations using this are only valid in executable or shared objects. 1041 1042The following relocation types are supported: 1043 1044 .. table:: AMDGPU ELF Relocation Records 1045 :name: amdgpu-elf-relocation-records-table 1046 1047 ========================== ======= ===== ========== ============================== 1048 Relocation Type Kind Value Field Calculation 1049 ========================== ======= ===== ========== ============================== 1050 ``R_AMDGPU_NONE`` 0 *none* *none* 1051 ``R_AMDGPU_ABS32_LO`` Static, 1 ``word32`` (S + A) & 0xFFFFFFFF 1052 Dynamic 1053 ``R_AMDGPU_ABS32_HI`` Static, 2 ``word32`` (S + A) >> 32 1054 Dynamic 1055 ``R_AMDGPU_ABS64`` Static, 3 ``word64`` S + A 1056 Dynamic 1057 ``R_AMDGPU_REL32`` Static 4 ``word32`` S + A - P 1058 ``R_AMDGPU_REL64`` Static 5 ``word64`` S + A - P 1059 ``R_AMDGPU_ABS32`` Static, 6 ``word32`` S + A 1060 Dynamic 1061 ``R_AMDGPU_GOTPCREL`` Static 7 ``word32`` G + GOT + A - P 1062 ``R_AMDGPU_GOTPCREL32_LO`` Static 8 ``word32`` (G + GOT + A - P) & 0xFFFFFFFF 1063 ``R_AMDGPU_GOTPCREL32_HI`` Static 9 ``word32`` (G + GOT + A - P) >> 32 1064 ``R_AMDGPU_REL32_LO`` Static 10 ``word32`` (S + A - P) & 0xFFFFFFFF 1065 ``R_AMDGPU_REL32_HI`` Static 11 ``word32`` (S + A - P) >> 32 1066 *reserved* 12 1067 ``R_AMDGPU_RELATIVE64`` Dynamic 13 ``word64`` B + A 1068 ========================== ======= ===== ========== ============================== 1069 1070``R_AMDGPU_ABS32_LO`` and ``R_AMDGPU_ABS32_HI`` are only supported by 1071the ``mesa3d`` OS, which does not support ``R_AMDGPU_ABS64``. 1072 1073There is no current OS loader support for 32-bit programs and so 1074``R_AMDGPU_ABS32`` is not used. 1075 1076.. _amdgpu-dwarf: 1077 1078DWARF 1079----- 1080 1081.. warning:: 1082 This section describes a **provisional proposal** that is not currently 1083 fully implemented and is subject to change. 1084 1085Standard DWARF [DWARF]_ sections can be generated. These contain information 1086that maps the code object executable code and data to the source language 1087constructs. It can be used by tools such as debuggers and profilers. 1088 1089This section defines the AMDGPU target specific DWARF. It applies to DWARF 1090Version 4 and 5. 1091 1092.. _amdgpu-dwarf-overview: 1093 1094Overview 1095~~~~~~~~ 1096 1097The AMDGPU has several features that require additional DWARF functionality in 1098order to support optimized code. 1099 1100A single code object can contain code for kernels that have different wave 1101sizes. The vector registers and some scalar registers are based on the wave 1102size. AMDGPU defines distinct DWARF registers for each wave size. This 1103simplifies the consumer of the DWARF so that each register has a fixed size, 1104rather than being dynamic according to the wave mode. Similarly, distinct DWARF 1105registers are defined for those registers that vary in size according to the 1106process address size. This allows a consumer to treat a specific AMDGPU target 1107as a single architecture regardless of how it is configured. The compiler 1108explicitly specifies the registers that match the mode of the code it is 1109generating. 1110 1111AMDGPU optimized code may spill vector registers to non-global address space 1112memory, and this spilling may be done only for lanes that are active on entry to 1113the subprogram. To support this, a location description that can be created as a 1114masked select is required. 1115 1116Since the active lane mask may be held in a register, a way to get the value of 1117a register on entry to a subprogram is required. To support this an operation 1118that returns the caller value of a register as specified by the Call Frame 1119Information (see :ref:`amdgpu-call-frame-information`) is required. 1120 1121Current DWARF uses an empty expression to indicate an undefined location 1122description. Since the masked select composite location description operation 1123takes more than one location description, it is necessary to have an explicit 1124way to specify an undefined location description. Otherwise it is not possible 1125to specify that a particular one of the input location descriptions is 1126undefined. 1127 1128CFI describes restoring callee saved registers that are spilled. Currently CFI 1129only allows a location description that is a register, memory address, or 1130implicit location description. AMDGPU optimized code may spill scalar registers 1131into portions of vector registers. This requires extending CFI to allow any 1132location description. 1133 1134The vector registers of the AMDGPU are represented as their full wave size, 1135meaning the wave size times the dword size. This reflects the actual hardware, 1136and allows the compiler to generate DWARF for languages that map a thread to the 1137complete wave. It also allows more efficient DWARF to be generated to describe 1138the CFI as only a single expression is required for the whole vector register, 1139rather than a separate expression for each lane's dword of the vector register. 1140It also allows the compiler to produce DWARF that indexes the vector register if 1141it spills scalar registers into portions of a vector registers. 1142 1143Since DWARF stack value entries have a base type and AMDGPU registers are a 1144vector of dwords, the ability to specify that a base type is a vector is 1145required. 1146 1147If the source language is mapped onto the AMDGPU wavefronts in a SIMT manner, 1148then the variable DWARF location expressions must compute the location for a 1149single lane of the wavefront. Therefore, a DWARF operator is required to denote 1150the current lane, much like ``DW_OP_push_object_address`` denotes the current 1151object. The ``DW_OP_*piece`` operators only allow literal indices. Therefore, a 1152composite location description is required that can take a computed index of a 1153location description (such as a vector register). 1154 1155If the source language is mapped onto the AMDGPU wavefronts in a SIMT manner the 1156compiler can use the AMDGPU execution mask register to control which lanes are 1157active. To describe the conceptual location of non-active lanes a DWARF 1158expression is needed that can compute a per lane PC. For efficiency, this is 1159done for the wave as a whole. This expression benefits by having a masked select 1160composite location description operation. This requires an attribute for source 1161location of each lane. The AMDGPU may update the execution mask for whole wave 1162operations and so needs an attribute that computes the current active lane mask. 1163 1164AMDGPU needs to be able to describe addresses that are in different kinds of 1165memory. Optimized code may need to describe a variable that resides in pieces 1166that are in different kinds of storage which may include parts of registers, 1167memory that is in a mixture of memory kinds, implicit values, or be undefined. 1168DWARF has the concept of segment addresses. However, the segment cannot be 1169specified within a DWARF expression, which is only able to specify the offset 1170portion of a segment address. The segment index is only provided by the entity 1171that species the DWARF expression. Therefore, the segment index is a property 1172that can only be put on complete objects, such as a variable. That makes it only 1173suitable for describing an entity (such as variable or subprogram code) that is 1174in a single kind of memory. Therefore, AMDGPU uses the DWARF concept of address 1175spaces. For example, a variable may be allocated in a register that is partially 1176spilled to the call stack which is in the private address space, and partially 1177spilled to the local address space. 1178 1179DWARF uses the concept of an address in many expression operators but does not 1180define how it relates to address spaces. For example, 1181``DW_OP_push_object_address`` pushes the address of an object. Other contexts 1182implicitly push an address on the stack before evaluating an expression. For 1183example, the ``DW_AT_use_location`` attribute of the 1184``DW_TAG_ptr_to_member_type``. The expression that uses the address needs to do 1185so in a general way and not need to be dependent on the address space of the 1186address. For example, a pointer to member value may want to be applied to an 1187object that may reside in any address space. 1188 1189The number of registers and the cost of memory operations is much higher for 1190AMDGPU than a typical CPU. The compiler attempts to optimize whole variables and 1191arrays into registers. Currently DWARF only allows ``DW_OP_push_object_address`` 1192and related operations to work with a global memory location. To support AMDGPU 1193optimized code it is required to generalize DWARF to allow any location 1194description to be used. This allows registers, or composite location 1195descriptions that may be a mixture of memory, registers, or even implicit 1196values. 1197 1198Allowing a location description to be an entry on the DWARF stack allows them to 1199compose naturally. It allows objects to be located in any kind of memory address 1200space, in registers, be implicit values, be undefined, or a composite of any of 1201these. 1202 1203By extending DWARF carefully, all existing DWARF expressions can retain their 1204current semantic meaning. DWARF has implicit conversions that convert from a 1205value that is treated as an address in the default address space to a memory 1206location description. This can be extended to allow a default address space 1207memory location description to be implicitly converted back to its address 1208value. To allow composition of composite location descriptions, an explicit 1209operator that indicates the end is required. This can be implied if the end of a 1210DWARF expression is reached, allowing current DWARF expressions to remain legal. 1211 1212The ``DW_OP_plus`` and ``DW_OP_minus`` can be defined to operate on a memory 1213location description in the default target architecture address space and a 1214generic type, and produce a memory location description. This allows them to 1215continue to be used to offset an address. To generalize offsetting to any 1216location description, including location descriptions that describe when bytes 1217are in registers, are implicit, or a composite of these, the 1218``DW_OP_LLVM_offset`` and ``DW_OP_LLVM_bit_offset`` operations are added. These 1219do not perform wrapping which would be hard to define for location descriptions 1220of non-memory kinds. This allows ``DW_OP_push_object_address`` to push a 1221location description that may be in a register, or be an implicit value, and the 1222DWARF expression of ``DW_TAG_ptr_to_member_type`` can contain 1223``DW_OP_LLVM_offset`` to offset within it. ``DW_OP_LLVM_bit_offset`` generalizes 1224DWARF to work with bit fields. 1225 1226The DWARF ``DW_OP_xderef*`` operation allows a value to be converted into an 1227address of a specified address space which is then read. But provides no way to 1228create a memory location description for an address in the non-default address 1229space. For example, AMDGPU variables can be allocated in the local address space 1230at a fixed address. It is required to have an operation to create an address in 1231a specific address space that can be used to define the location description of 1232the variable. Defining this operation to produce a location description allows 1233the size of addresses in an address space to be larger than the generic type. 1234 1235If an operation had to produce a value that can be implicitly converted to a 1236memory location description, then it would be limited to the size of the generic 1237type which matches the size of the default address space. Its value would be 1238unspecified and likely not match any value in the actual program. By making the 1239result a location description, it allows a consumer great freedom in how it 1240implements it. The implicit conversion back to a value can be limited only to 1241the default address space to maintain compatibility. 1242 1243Similarly ``DW_OP_breg*`` treats the register as containing an address in the 1244default address space. It is required to be able to specify the address space of 1245the register value. 1246 1247Almost all uses of addresses in DWARF are limited to defining location 1248descriptions, or to be dereferenced to read memory. The exception is 1249``DW_CFA_val_offset`` which uses the address to set the value of a register. By 1250defining the CFA DWARF expression as being a memory location description, it can 1251maintain what address space it is, and that can be used to convert the offset 1252address back to an address in that address space. (An alternative is to defined 1253``DW_CFA_val_offset`` to implicitly use the default address space, and add 1254another operation that specifies the address space.) 1255 1256This approach allows all existing DWARF to have the identical semantics. It 1257allows the compiler to explicitly specify the address space it is using. For 1258example, a compiler could choose to access private memory in a swizzled manner 1259when mapping a source language to a wave in a SIMT manner, or to access it in an 1260unswizzled manner if mapping the same language with the wave being the thread. 1261It also allows the compiler to mix the address space it uses to access private 1262memory. For example, for SIMT it can still spill entire vector registers in an 1263unswizzled manner, while using swizzled for SIMT variable access. This approach 1264allows memory location descriptions for different address spaces to be combined 1265using the regular ``DW_OP_*piece`` operators. 1266 1267Location descriptions are an abstraction of storage, they give freedom to the 1268consumer on how to implement them. They allow the address space to encode lane 1269information so they can be used to read memory with only the memory description 1270and no extra arguments. The same set of operations can operate on locations 1271independent of their kind of storage. The ``DW_OP_deref*`` therefore can be used 1272on any storage kind. ``DW_OP_xderef*`` is unnecessary except to become a more 1273compact way to convert a segment address followed by dereferencing it. 1274 1275Several approaches were considered, and the one proposed appears to be the 1276cleanest and offers the greatest improvement of DWARF's ability to support 1277optimized code. Examining the gdb debugger and LLVM compiler, it appears only to 1278require modest changes as they both already have to support general use of 1279location descriptions. It is anticipated that will be the case for other 1280debuggers and compilers. 1281 1282The following provides the definitions for the additional operators, as well as 1283clarifying how existing expression operators, CFI operators, and attributes 1284behave with respect to generalized location descriptions that support address 1285spaces. It has been defined such that it is backwards compatible with DWARF 5. 1286The definitions are intended to fully define well-formed DWARF in a consistent 1287style. Some sections are organized to mirror the DWARF 5 specification 1288structure, with non-normative text shown in *italics*. 1289 1290.. _amdgpu-dwarf-language-names: 1291 1292Language Names 1293~~~~~~~~~~~~~~ 1294 1295Language codes defined for use with the ``DW_AT_language`` attribute are 1296defined in :ref:`amdgpu-dwarf-language-names-table`. 1297 1298.. table:: AMDGPU DWARF Language Names 1299 :name: amdgpu-dwarf-language-names-table 1300 1301 ==================== ====== =================== ============================= 1302 Language Name Code Default Lower Bound Description 1303 ==================== ====== =================== ============================= 1304 ``DW_LANG_LLVM_HIP`` 0x8100 0 AMD HIP Language. See [HIP]_. 1305 ==================== ====== =================== ============================= 1306 1307The ``DW_LANG_LLVM_HIP`` language can be supported by extending the C++ 1308language. 1309 1310.. _amdgpu-dwarf-register-mapping: 1311 1312Register Mapping 1313~~~~~~~~~~~~~~~~ 1314 1315DWARF registers are encoded as numbers, which are mapped to architecture 1316registers. The mapping for AMDGPU is defined in 1317:ref:`amdgpu-dwarf-register-mapping-table`. 1318 1319.. table:: AMDGPU DWARF Register Mapping 1320 :name: amdgpu-dwarf-register-mapping-table 1321 1322 ============== ================= ======== ================================== 1323 DWARF Register AMDGPU Register Bit Size Description 1324 ============== ================= ======== ================================== 1325 0 PC_32 32 Program Counter (PC) when 1326 executing in a 32-bit process 1327 address space. Used in the CFI to 1328 describe the PC of the calling 1329 frame. 1330 1 EXEC_MASK_32 32 Execution Mask Register when 1331 executing in wave 32 mode. 1332 2-15 *Reserved* 1333 16 PC_64 64 Program Counter (PC) when 1334 executing in a 64-bit process 1335 address space. Used in the CFI to 1336 describe the PC of the calling 1337 frame. 1338 17 EXEC_MASK_64 64 Execution Mask Register when 1339 executing in wave 64 mode. 1340 18-31 *Reserved* 1341 32-95 SGPR0-SGPR63 32 Scalar General Purpose 1342 Registers. 1343 96-127 *Reserved* 1344 128-511 *Reserved* 1345 512-1023 *Reserved* 1346 1024-1087 *Reserved* 1347 1088-1129 SGPR64-SGPR105 32 Scalar General Purpose Registers 1348 1130-1535 *Reserved* 1349 1536-1791 VGPR0-VGPR255 32*32 Vector General Purpose Registers 1350 when executing in wave 32 mode. 1351 1792-2047 *Reserved* 1352 2048-2303 AGPR0-AGPR255 32*32 Vector Accumulation Registers 1353 when executing in wave 32 mode. 1354 2304-2559 *Reserved* 1355 2560-2815 VGPR0-VGPR255 64*32 Vector General Purpose Registers 1356 when executing in wave 64 mode. 1357 2816-3071 *Reserved* 1358 3072-3327 AGPR0-AGPR255 64*32 Vector Accumulation Registers 1359 when executing in wave 64 mode. 1360 3328-3583 *Reserved* 1361 ============== ================= ======== ================================== 1362 1363The vector registers are represented as the full size for the wavefront. They 1364are organized as consecutive dwords (32-bits), one per lane, with the dword at 1365the least significant bit position corresponding to lane 0 and so forth. DWARF 1366location expressions involving the ``DW_OP_LLVM_offset`` and 1367``DW_OP_LLVM_push_lane`` operations are used to select the part of the vector 1368register corresponding to the lane that is executing the current thread of 1369execution in languages that are implemented using a SIMD or SIMT execution 1370model. 1371 1372If the wavefront size is 32 lanes then the wave 32 mode register definitions 1373are used. If the wavefront size is 64 lanes then the wave 64 mode register 1374definitions are used. Some AMDGPU targets support executing in both wave 32 1375and wave 64 mode. The register definitions corresponding to the wave mode 1376of the generated code will be used. 1377 1378If code is generated to execute in a 32-bit process address space then the 137932-bit process address space register definitions are used. If code is 1380generated to execute in a 64-bit process address space then the 64-bit process 1381address space register definitions are used. The ``amdgcn`` target only 1382supports the 64-bit process address space. 1383 1384Address Class Mapping 1385~~~~~~~~~~~~~~~~~~~~~ 1386 1387DWARF address classes are used for languages with the concept of memory address 1388spaces. They are used in the ``DW_AT_address_class`` attribute for pointer type, 1389reference type, subroutine, and subroutine type debugger information entries 1390(DIEs). 1391 1392The address class mapping for AMDGPU is defined in 1393:ref:`amdgpu-dwarf-address-class-mapping-table`. 1394 1395.. table:: AMDGPU DWARF Address Class Mapping 1396 :name: amdgpu-dwarf-address-class-mapping-table 1397 1398 =========================== ===== ================= 1399 DWARF AMDGPU 1400 --------------------------------- ----------------- 1401 Address Class Name Value Address Space 1402 =========================== ===== ================= 1403 ``DW_ADDR_none`` 0x00 Generic (Flat) 1404 ``DW_ADDR_AMDGPU_global`` 0x01 Global 1405 ``DW_ADDR_AMDGPU_region`` 0x02 Region (GDS) 1406 ``DW_ADDR_AMDGPU_local`` 0x03 Local (group/LDS) 1407 ``DW_ADDR_AMDGPU_constant`` 0x04 Global 1408 ``DW_ADDR_AMDGPU_private`` 0x05 Private (Scratch) 1409 =========================== ===== ================= 1410 1411See :ref:`amdgpu-address-spaces` for information on the AMDGPU address spaces 1412including address size and NULL value. 1413 1414For AMDGPU the address class encodes the address class as declared in the 1415source language type. 1416 1417For AMDGPU if no ``DW_AT_address_class`` attribute is present, then the 1418``DW_ADDR_none`` address class is used. 1419 1420.. note:: 1421 1422 The ``DW_ADDR_none`` default was defined as ``Generic`` and not ``Global`` 1423 to match the LLVM address space ordering. This ordering was chosen to better 1424 support CUDA-like languages such as HIP that do not have address spaces in 1425 the language type system, but do allow variables to be allocated in 1426 different address spaces. So effectively all CUDA and HIP source language 1427 addresses are generic. 1428 1429.. note:: 1430 1431 Currently DWARF defines address class values as architecture specific. It 1432 is unclear how language specific address spaces are intended to be 1433 represented in DWARF. 1434 1435 For example, OpenCL defines address spaces for ``global``, ``local``, 1436 ``constant``, and ``private``. These are part of the type system and are 1437 modifies to pointer types. In addition, OpenCL defines ``generic`` pointers 1438 that can reference either the ``global``, ``local``, or ``private`` address 1439 spaces. To support the OpenCL language the debugger would want to support 1440 casting pointers between the ``generic`` and other address spaces, and 1441 possibly using pointer casting to form an address for a specific address 1442 space out of an integral value. 1443 1444 The method to use to dereference a pointer type or reference type value is 1445 defined in DWARF expressions using ``DW_OP_xderef*`` which uses an 1446 architecture specific address space. 1447 1448 DWARF defines the ``DW_AT_address_class`` attribute on pointer types and 1449 reference types. It specifies the method to use to dereference them. Why 1450 is the value of this not the same as the address space value used in 1451 ``DW_OP_xderef*`` since in both cases it is architecture specific and the 1452 architecture presumably will use the same set of methods to dereference 1453 pointers in both cases? 1454 1455 Since ``DW_AT_address_class`` uses an architecture specific value it cannot 1456 in general capture the source language address space type modifier concept. 1457 On some architectures all source language address space modifies may 1458 actually use the same method for dereferencing pointers. 1459 1460 One possibility is for DWARF to add an ``DW_TAG_LLVM_address_class_type`` 1461 type modifier that can be applied to a pointer type and reference type. The 1462 ``DW_AT_address_class`` attribute could be re-defined to not be architecture 1463 specific and instead define generalized language values that will support 1464 OpenCL and other languages using address spaces. The ``DW_AT_address_class`` 1465 could be defined to not be applied to pointer or reference types, but 1466 instead only to the ``DW_TAG_LLVM_address_class_type`` type modifier entry. 1467 1468 If a pointer type or reference type is not modified by 1469 ``DW_TAG_LLVM_address_class_type`` or if ``DW_TAG_LLVM_address_class_type`` 1470 has no ``DW_AT_address_class`` attribute, then the pointer type or reference 1471 type would be defined to use the ``DW_ADDR_none`` address class as 1472 currently. Since modifiers can be chained, it would need to be defined if 1473 multiple ``DW_TAG_LLVM_address_class_type`` modifies was legal, and if so if 1474 the outermost one is the one that takes precedence. 1475 1476 A target implementation that supports multiple address spaces would need to 1477 map ``DW_ADDR_none`` appropriately to support CUDA-like languages 1478 that have no address classes in the type system, but do support variable 1479 allocation in address spaces. See the above note that describes why AMDGPU 1480 choose to make ``DW_ADDR_none`` map to the ``Generic`` AMDGPU address space 1481 and not the ``Global`` address space. 1482 1483 An alternative would be to define ``DW_ADDR_none`` as being the global 1484 address class and then change ``DW_ADDR_global`` to ``DW_ADDR_generic``. 1485 Compilers generating DWARF for CUDA-like languages would then have to define 1486 every CUDA-like language pointer type or reference type using 1487 ``DW_TAG_LLVM_address_class_type`` with a ``DW_AT_address_class`` attribute 1488 of ``DW_ADDR_generic`` to match the language semantics. The AMDGPU 1489 alternative avoids needing to do this and seems to fit better into how CLANG 1490 and LLVM have added support for the CUDA-like languages on top of existing 1491 C++ language support. 1492 1493 A new ``DW_AT_address_space`` attribute could be defined that can be applied 1494 to pointer type, reference type, subroutine, and subroutine type to describe 1495 how objects having the given type are dereferenced or called (the role that 1496 ``DW_AT_address_class`` currently provides). The values of 1497 ``DW_AT_address_space`` would be architecture specific and the same as used 1498 in ``DW_OP_xderef*``. 1499 1500.. _amdgpu-dwarf-address-space-mapping: 1501 1502Address Space Mapping 1503~~~~~~~~~~~~~~~~~~~~~ 1504 1505DWARF address spaces are used in location expressions to describe the memory 1506space where data resides. Address spaces correspond to a target specific memory 1507space and are not tied to any source language concept. 1508 1509The AMDGPU address space mapping is defined in 1510:ref:`amdgpu-dwarf-address-space-mapping-table`. 1511 1512.. table:: AMDGPU DWARF Address Space Mapping 1513 :name: amdgpu-dwarf-address-space-mapping-table 1514 1515 ======================================= ===== ======= ======== ================= ======================= 1516 DWARF AMDGPU Notes 1517 --------------------------------------- ----- ---------------- ----------------- ----------------------- 1518 Address Space Name Value Address Bit Size Address Space 1519 --------------------------------------- ----- ------- -------- ----------------- ----------------------- 1520 .. 64-bit 32-bit 1521 process process 1522 address address 1523 space space 1524 ======================================= ===== ======= ======== ================= ======================= 1525 ``DW_ASPACE_none`` 0x00 8 4 Global *default address space* 1526 ``DW_ASPACE_AMDGPU_generic`` 0x01 8 4 Generic (Flat) 1527 ``DW_ASPACE_AMDGPU_region`` 0x02 4 4 Region (GDS) 1528 ``DW_ASPACE_AMDGPU_local`` 0x03 4 4 Local (group/LDS) 1529 *Reserved* 0x04 1530 ``DW_ASPACE_AMDGPU_private_lane`` 0x05 4 4 Private (Scratch) *focused lane* 1531 ``DW_ASPACE_AMDGPU_private_wave`` 0x06 4 4 Private (Scratch) *unswizzled wave* 1532 *Reserved* 0x07- 1533 0x1F 1534 ``DW_ASPACE_AMDGPU_private_lane<0-63>`` 0x20- 4 4 Private (Scratch) *specific lane* 1535 0x5F 1536 ======================================= ===== ======= ======== ================= ======================= 1537 1538See :ref:`amdgpu-address-spaces` for information on the AMDGPU address spaces 1539including address size and NULL value. 1540 1541The ``DW_ASPACE_none`` address space is the default address space used in DWARF 1542operations that do not specify an address space. It therefore has to map to the 1543global address space so that the ``DW_OP_addr*`` and related operations can 1544refer to addresses in the program code. 1545 1546The ``DW_ASPACE_AMDGPU_generic`` address space allows location expressions to 1547specify the flat address space. If the address corresponds to an address in the 1548local address space then it corresponds to the wave that is executing the 1549focused thread of execution. If the address corresponds to an address in the 1550private address space then it corresponds to the lane that is executing the 1551focused thread of execution for languages that are implemented using a SIMD or 1552SIMT execution model. 1553 1554.. note:: 1555 1556 CUDA-like languages such as HIP that do not have address spaces in the 1557 language type system, but do allow variables to be allocated in different 1558 address spaces, will need to explicitly specify the 1559 ``DW_ASPACE_AMDGPU_generic`` address space in the DWARF operations as the 1560 default address space is the global address space. 1561 1562The ``DW_ASPACE_AMDGPU_local`` address space allows location expressions to 1563specify the local address space corresponding to the wave that is executing the 1564focused thread of execution. 1565 1566The ``DW_ASPACE_AMDGPU_private_lane`` address space allows location expressions 1567to specify the private address space corresponding to the lane that is 1568executing the focused thread of execution for languages that are implemented 1569using a SIMD or SIMT execution model. 1570 1571The ``DW_ASPACE_AMDGPU_private_wave`` address space allows location expressions 1572to specify the unswizzled private address space corresponding to the wave that 1573is executing the focused thread of execution. The wave view of private memory 1574is the per wave unswizzled backing memory layout defined in 1575:ref:`amdgpu-address-spaces`, such that address 0 corresponds to the first 1576location for the backing memory of the wave (namely the address is not offset 1577by ``wavefront-scratch-base``). So to convert from a 1578``DW_ASPACE_AMDGPU_private_lane`` to a ``DW_ASPACE_AMDGPU_private_wave`` 1579segment address perform the following: 1580 1581:: 1582 1583 private-address-wave = 1584 ((private-address-lane / 4) * wavefront-size * 4) + 1585 (wavefront-lane-id * 4) + (private-address-lane % 4) 1586 1587If the ``DW_ASPACE_AMDGPU_private_lane`` segment address is dword aligned and 1588the start of the dwords for each lane starting with lane 0 is required, then 1589this simplifies to: 1590 1591:: 1592 1593 private-address-wave = 1594 private-address-lane * wavefront-size 1595 1596A compiler can use this address space to read a complete spilled vector 1597register back into a complete vector register in the CFI. The frame pointer can 1598be a private lane segment address which is dword aligned, which can be shifted 1599to multiply by the wave size, and then used to form a private wave segment 1600address that gives a location for a contiguous set of dwords, one per lane, 1601where the vector register dwords are spilled. The compiler knows the wave size 1602since it generates the code. Note that the type of the address may have to be 1603converted as the size of a private lane segment address may be smaller than the 1604size of a private wave segment address. 1605 1606The ``DW_ASPACE_AMDGPU_private_lane<n>`` address space allows location 1607expressions to specify the private address space corresponding to a specific 1608lane. For example, this can be used when the compiler spills scalar registers 1609to scratch memory, with each scalar register being saved to a different lane's 1610scratch memory. 1611 1612.. _amdgpu-dwarf-expressions: 1613 1614Expressions 1615~~~~~~~~~~~ 1616 1617The following sections define the new DWARF expression operator used by AMDGPU, 1618as well as clarifying the extensions to already existing DWARF 5 operations. 1619 1620DWARF expressions describe how to compute a value or specify a location 1621description. An expression is encoded as a stream of operations, each consisting 1622of an opcode followed by zero or more literal operands. The number of operands 1623is implied by the opcode. 1624 1625Operations represent a postfix operation on a simple stack machine. They can act 1626on entries on the stack, including adding entries and removing entries. If the 1627kind of a stack entry does not match the kind required by the operation, and is 1628not implicitly convertible to the required kind, then the DWARF expression is 1629ill-formed. 1630 1631Each stack entry can be one of two kinds: a value or a location description. 1632Value stack entries are described in :ref:`amdgpu-value-operations` and 1633location description stack entries are described in 1634:ref:`amdgpu-location-description-operations`. 1635 1636*The evaluation of a DWARF expression can provide the location description of an 1637object, the value of an array bound, the length of a dynamic string, the desired 1638value itself, and so on.* 1639 1640The result of the evaluation of a DWARF expression is defined as: 1641 1642* If evaluation of the DWARF expression is on behalf of a ``DW_OP_call*`` 1643 operation for a ``DW_AT_location`` attribute that belongs to a 1644 ``DW_TAG_dwarf_procedure`` debugging information entry, then all the entries 1645 on the stack are left, and execution of the DWARF expression containing the 1646 ``DW_OP_call*`` operation continues. 1647 1648* If evaluation of the DWARF expression requires a location description, then: 1649 1650 * If the stack is empty, an undefined location description is returned. 1651 1652 * If the top stack entry is a location description, or can be converted to 1653 one, then the, possibly converted, location description is returned. Any 1654 other entries on the stack are discarded. 1655 1656 * Otherwise the DWARF expression is ill-formed. 1657 1658 .. note:: 1659 1660 Could define this case as returning an implicit location description as 1661 if the ``DW_OP_implicit`` operation is performed. 1662 1663* If evaluation of the DWARF expression requires a value, then: 1664 1665 * If the top stack entry is a value, or can be converted to one, then the, 1666 possibly converted, value is returned. Any other entries on the stack are 1667 discarded. 1668 1669 * Otherwise the DWARF expression is ill-formed. 1670 1671.. _amdgpu-stack-operations: 1672 1673Stack Operations 1674++++++++++++++++ 1675 1676The following operations manipulate the DWARF stack. Operations that index 1677the stack assume that the top of the stack (most recently added entry) has index 16780. They allow the stack entries to be either a value or location description. 1679 1680If any stack entry accessed by a stack operation is an incomplete composite 1681location description, then the DWARF expression is ill-formed. 1682 1683.. note:: 1684 1685 These operations now support stack entries that are values and location 1686 descriptions. 1687 1688.. note:: 1689 1690 If it is desired to also make them work with incomplete composite location 1691 descriptions then would need to define that the composite location storage 1692 specified by the incomplete composite location description is also replicated 1693 when a copy is pushed. This ensures that each copy of the incomplete composite 1694 location description can updated the composite location storage they specify 1695 independently. 1696 16971. ``DW_OP_dup`` 1698 1699 ``DW_OP_dup`` duplicates the stack entry at the top of the stack. 1700 17012. ``DW_OP_drop`` 1702 1703 ``DW_OP_drop`` pops the stack entry at the top of the stack and discards it. 1704 17053. ``DW_OP_pick`` 1706 1707 ``DW_OP_pick`` has a single unsigned 1-byte operand that is treated as an 1708 index I. A copy of the stack entry with index I is pushed onto the stack. 1709 17104. ``DW_OP_over`` 1711 1712 ``DW_OP_over`` pushes a copy of the entry entry with index 1. 1713 1714 *This is equivalent to a ``DW_OP_pick 1`` operation.* 1715 17165. ``DW_OP_swap`` 1717 1718 ``DW_OP_swap`` swaps the top two stack entries. The entry at the top of the 1719 stack becomes the second stack entry, and the second stack entry becomes the 1720 top of the stack. 1721 17226. ``DW_OP_rot`` 1723 1724 ``DW_OP_rot`` rotates the first three stack entries. The entry at the top of 1725 the stack becomes the third stack entry, the second entry becomes the top of 1726 the stack, and the third entry becomes the second entry. 1727 1728.. _amdgpu-value-operations: 1729 1730Value Operations 1731++++++++++++++++ 1732 1733Each value stack entry has a type and a value, and can represent a value of 1734any supported base type of the target machine. The base type specifies the size 1735and encoding of the value. 1736 1737.. note:: 1738 1739 It may be better to add an implicit pointer value kind that is produced when 1740 ``DW_OP_deref*`` retrieves the full contents of an implicit pointer location 1741 storage created by the ``DW_OP_implicit_pointer`` or 1742 ``DW_OP_LLVM_aspace_implicit_pointer`` operations. 1743 1744Instead of a base type, value stack entries can have a distinguished generic 1745type, which is an integral type that has the size of an address in the target 1746architecture default address space on the target machine and unspecified 1747signedness. 1748 1749*The generic type is the same as the unspecified type used for stack operations 1750defined in DWARF Version 4 and before.* 1751 1752An integral type is a base type that has an encoding of ``DW_ATE_signed``, 1753``DW_ATE_signed_char``, ``DW_ATE_unsigned``, ``DW_ATE_unsigned_char``, 1754``DW_ATE_boolean``, or any target architecture defined integral encoding in the 1755inclusive range ``DW_ATE_lo_user`` to ``DW_ATE_hi_user``. 1756 1757.. note:: 1758 1759 Unclear if ``DW_ATE_address`` is an integral type. gdb does not seem to 1760 consider as integral. 1761 17621. ``DW_OP_LLVM_push_lane`` *New* 1763 1764 ``DW_OP_LLVM_push_lane`` pushes a value with the generic type that is the 1765 target architecture lane identifier of the thread of execution for which a 1766 user presented expression is currently being evaluated. For languages that 1767 are implemented using a SIMD or SIMT execution model this is the lane number 1768 that corresponds to the source language thread of execution upon which the 1769 user is focused. Otherwise this is the value 0. 1770 1771 For AMDGPU, the lane identifier returned by ``DW_OP_LLVM_push_lane`` 1772 corresponds to the the hardware lane number which is numbered from 0 to the 1773 wavefront size minus 1. 1774 17752. ``DW_OP_entry_value`` 1776 1777 ``DW_OP_entry_value`` pushes the value that the described location held upon 1778 entering the current subprogram. 1779 1780 It has two operands. The first is an unsigned LEB128 integer. The second is 1781 a block of bytes, with a length equal to the first operand, treated as a 1782 DWARF expression E. 1783 1784 E is evaluated as if it had been evaluated upon entering the current 1785 subprogram. E assumes no values are present on the DWARF stack initially and 1786 results in exactly one value being pushed on the DWARF stack when completed. 1787 1788 ``DW_OP_push_object_address`` is not meaningful inside of this DWARF 1789 operation. 1790 1791 If the result of E is a register location description (see 1792 :ref:`amdgpu-register-location-descriptions`), ``DW_OP_entry_value`` pushes 1793 the value that register had upon entering the current subprogram. The value 1794 entry type is the target machine register base type. If the register value 1795 is undefined or the register location description bit offset is not 0, then 1796 the DWARF expression is ill-formed. 1797 1798 *The register location description provides a more compact form for the case 1799 where the value was in a register on entry to the subprogram.* 1800 1801 Otherwise, the expression result is required to be a value, and 1802 ``DW_OP_entry_value`` pushes that value. 1803 1804 *The values needed to evaluate* ``DW_OP_entry_value`` *could be obtained in 1805 several ways. The consumer could suspend execution on entry to the 1806 subprogram, record values needed by* ``DW_OP_entry_value`` *expressions 1807 within the subprogram, and then continue; when evaluating* 1808 ``DW_OP_entry_value``\ *, the consumer would use these recorded values 1809 rather than the current values. Or, when evaluating* ``DW_OP_entry_value``\ 1810 *, the consumer could virtually unwind using the Call Frame Information 1811 (see* :ref:`amdgpu-call-frame-information`\ *) to recover register values 1812 that might have been clobbered since the subprogram entry point.* 1813 1814 .. note:: 1815 1816 Unclear why this operation is defined this way. If the expression is 1817 simply using existing variables then it is just a regular expression. It 1818 is unclear how the compiler instructs the consumer how to create the saved 1819 copies of the variables on entry. Seems only the compiler knows how to do 1820 this. If the main purpose is only to read the entry value of a register 1821 using CFI then would be better to have an operation that explicitly does 1822 just that such as ``DW_OP_LLVM_call_frame_entry_reg``. 1823 1824.. _amdgpu-location-description-operations: 1825 1826Location Description Operations 1827+++++++++++++++++++++++++++++++ 1828 1829Information about the location of program objects is provided by location 1830descriptions. Location descriptions specify the storage that holds the program 1831objects, and a position within the storage. 1832 1833A location storage is a linear stream of bits that can hold values. Each 1834location storage has a size in bits and can be accessed using a zero-based bit 1835offset. The ordering of bits within location storage uses the bit numbering and 1836direction conventions that are appropriate to the current language on the target 1837architecture. 1838 1839.. note:: 1840 1841 For AMDGPU bytes are ordered with least significant bytes first, and bits are 1842 ordered within bytes with least significant bits first. 1843 1844There are five kinds of location storage: undefined, memory, register, implicit, 1845and composite. Memory and register location storage corresponds to the target 1846architecture memory address spaces and registers. Implicit location storage 1847corresponds to fixed values that can only be read. Undefined location storage 1848indicates no value is available and therefore cannot be read or written. 1849Composite location storage allows a mixture of these where some bits come from 1850one kind of location storage and some from another kind of location storage. 1851 1852.. note:: 1853 1854 It may be better to add an implicit pointer location storage kind for 1855 ``DW_OP_implicit_pointer`` or ``DW_OP_LLVM_aspace_implicit_pointer``. 1856 1857Location description stack entries specify a location storage to which they 1858refer, and a bit offset relative to the start of the location storage. 1859 1860General Operations 1861################## 1862 18631. ``DW_OP_LLVM_offset`` *New* 1864 1865 ``DW_OP_LLVM_offset`` pops two stack entries. The first must be an integral 1866 type value that is treated as a byte displacement D. The second must be a 1867 location description L. 1868 1869 It adds the value of D scaled by 8 (the byte size) to the bit offset of L, 1870 and pushes the updated L. 1871 1872 If the updated bit offset of L is less than 0 or greater than or equal to 1873 the size of the location storage specified by L, then the DWARF expression 1874 is ill-formed. 1875 18762. ``DW_OP_LLVM_offset_uconst`` *New* 1877 1878 ``DW_OP_LLVM_offset_uconst`` has a single unsigned LEB128 integer operand 1879 that is treated as a displacement D. 1880 1881 It pops one stack entry that must be a location description L. It adds the 1882 value of D scaled by 8 (the byte size) to the bit offset of L, and pushes 1883 the updated L. 1884 1885 If the updated bit offset of L is less than 0 or greater than or equal to 1886 the size of the location storage specified by L, then the DWARF expression 1887 is ill-formed. 1888 1889 *This operation is supplied specifically to be able to encode more field 1890 displacements in two bytes than can be done with* ``DW_OP_lit<n> 1891 DW_OP_LLVM_offset``\ *.* 1892 18933. ``DW_OP_LLVM_bit_offset`` *New* 1894 1895 ``DW_OP_LLVM_bit_offset`` pops two stack entries. The first must be an 1896 integral type value that is treated as a bit displacement D. The second must 1897 be a location description L. 1898 1899 It adds the value of D to the bit offset of L, and pushes the updated L. 1900 1901 If the updated bit offset of L is less than 0 or greater than or equal to 1902 the size of the location storage specified by L, then the DWARF expression 1903 is ill-formed. 1904 19054. ``DW_OP_deref`` 1906 1907 The ``DW_OP_deref`` operation pops one stack entry that must be a location 1908 description L. 1909 1910 A value of the bit size of the generic type is retrieved from the location 1911 storage specified by L starting at the bit offset specified by L. The 1912 retrieved generic type value V is pushed on the stack. 1913 1914 If any bit of the value is retrieved from the undefined location storage, or 1915 the offset of any bit exceeds the size of the location storage specified by 1916 L, then the DWARF expression is ill-formed. 1917 1918 See :ref:`amdgpu-implicit-location-descriptions` for special rules 1919 concerning implicit location descriptions created by the 1920 ``DW_OP_implicit_pointer`` and ``DW_OP_LLVM_implicit_aspace_pointer`` 1921 operations. 1922 19235. ``DW_OP_deref_size`` 1924 1925 ``DW_OP_deref_size`` has a single 1-byte unsigned integral constant treated 1926 as a byte result size S. 1927 1928 It pops one stack entry that must be a location description L. 1929 1930 A value of S scaled by 8 (the byte size) bits is retrieved from the location 1931 storage specified by L starting at the bit offset specified by L. The value 1932 V retrieved is zero-extended to the bit size of the generic type before 1933 being pushed onto the stack with the generic type. 1934 1935 If S is larger than the byte size of the generic type, if any bit of the 1936 value is retrieved from the undefined location storage, or if the offset of 1937 any bit exceeds the size of the location storage specified by L, then the 1938 DWARF expression is ill-formed. 1939 1940 See :ref:`amdgpu-implicit-location-descriptions` for special rules 1941 concerning implicit location descriptions created by the 1942 ``DW_OP_implicit_pointer`` and ``DW_OP_LLVM_implicit_aspace_pointer`` 1943 operations. 1944 19456. ``DW_OP_deref_type`` 1946 1947 ``DW_OP_deref_type`` has two operands. The first is a 1-byte unsigned 1948 integral constant whose value S is the same as the size of the base type 1949 referenced by the second operand. The second operand is an unsigned LEB128 1950 integer that represents the offset of a debugging information entry E in the 1951 current compilation unit, which must be a ``DW_TAG_base_type`` entry that 1952 provides the type of the result value. 1953 1954 It pops one stack entry that must be a location description L. A value of 1955 the bit size S is retrieved from the location storage specified by L 1956 starting at the bit offset specified by the L. The retrieved result type 1957 value V is pushed on the stack. 1958 1959 If any bit of the value is retrieved from the undefined location storage, or 1960 if the offset of any bit exceeds the size of the specified location storage, 1961 then the DWARF expression is ill-formed. 1962 1963 See :ref:`amdgpu-implicit-location-descriptions` for special rules 1964 concerning implicit location descriptions created by the 1965 ``DW_OP_implicit_pointer`` and ``DW_OP_LLVM_implicit_aspace_pointer`` 1966 operations. 1967 1968 *While the size of the pushed value could be inferred from the base type 1969 definition, it is encoded explicitly into the operation so that the 1970 operation can be parsed easily without reference to the* ``.debug_info`` 1971 *section.* 1972 19737. ``DW_OP_xderef`` *Deprecated* 1974 1975 ``DW_OP_xderef`` pops two stack entries. The first must be an integral type 1976 value that is treated as an address A. The second must be an integral type 1977 value that is treated as an address space identifier AS for those 1978 architectures that support multiple address spaces. 1979 1980 The operation is equivalent to performing ``DW_OP_swap; 1981 DW_OP_LLVM_form_aspace_address; DW_OP_deref``. The retrieved generic type 1982 value V is left on the stack. 1983 19848. ``DW_OP_xderef_size`` *Deprecated* 1985 1986 ``DW_OP_xderef_size`` has a single 1-byte unsigned integral constant treated 1987 as a byte result size S. 1988 1989 It pops two stack entries. The first must be an integral type value that is 1990 treated as an address A. The second must be an integral type value that is 1991 treated as an address space identifier AS for those architectures that 1992 support multiple address spaces. 1993 1994 The operation is equivalent to performing ``DW_OP_swap; 1995 DW_OP_LLVM_form_aspace_address; DW_OP_deref_size S``. The zero-extended 1996 retrieved generic type value V is left on the stack. 1997 19989. ``DW_OP_xderef_type`` *Deprecated* 1999 2000 ``DW_OP_xderef_type`` has two operands. The first is a 1-byte unsigned 2001 integral constant S whose value is the same as the size of the base type 2002 referenced by the second operand. The second operand is an unsigned LEB128 2003 integer R that represents the offset of a debugging information entry E in 2004 the current compilation unit, which must be a ``DW_TAG_base_type`` entry 2005 that provides the type of the result value. 2006 2007 It pops two stack entries. The first must be an integral type value that is 2008 treated as an address A. The second must be an integral type value that is 2009 treated as an address space identifier AS for those architectures that 2010 support multiple address spaces. 2011 2012 The operation is equivalent to performing ``DW_OP_swap; 2013 DW_OP_LLVM_form_aspace_address; DW_OP_deref_type S R``. The retrieved result 2014 type value V is left on the stack. 2015 201610. ``DW_OP_push_object_address`` 2017 2018 ``DW_OP_push_object_address`` pushes the location description L of the 2019 object currently being evaluated as part of evaluation of a user presented 2020 expression. 2021 2022 This object may correspond to an independent variable described by its own 2023 debugging information entry or it may be a component of an array, structure, 2024 or class whose address has been dynamically determined by an earlier step 2025 during user expression evaluation. 2026 2027 *This operator provides explicit functionality (especially for arrays 2028 involving descriptions) that is analogous to the implicit push of the base 2029 address of a structure prior to evaluation of a 2030 ``DW_AT_data_member_location`` to access a data member of a structure.* 2031 203211. ``DW_OP_call2, DW_OP_call4, DW_OP_call_ref`` 2033 2034 ``DW_OP_call2``, ``DW_OP_call4``, and ``DW_OP_call_ref`` perform DWARF 2035 procedure calls during evaluation of a DWARF expression or location 2036 description. 2037 2038 ``DW_OP_call2`` and ``DW_OP_call4``, have one operand that is a 2- or 4-byte 2039 unsigned offset, respectively, of a debugging information entry D in the 2040 current compilation unit. 2041 2042 ``DW_OP_LLVM_call_ref`` has one operand that is a 4-byte unsigned value in 2043 the 32-bit DWARF format, or an 8-byte unsigned value in the 64-bit DWARF 2044 format, that is treated as an offset of a debugging information entry D in a 2045 ``.debug_info`` section, which may be contained in an executable or shared 2046 object file other than that containing the operator. For references from one 2047 executable or shared object file to another, the relocation must be 2048 performed by the consumer. 2049 2050 *Operand interpretation of* ``DW_OP_call2``\ *,* ``DW_OP_call4``\ *, and* 2051 ``DW_OP_call_ref`` *is exactly like that for* ``DW_FORM_ref2``\ *, 2052 ``DW_FORM_ref4``\ *, and* ``DW_FORM_ref_addr``\ *, respectively.* 2053 2054 If D has a ``DW_AT_location`` attribute, then the DWARF expression E 2055 corresponding to the current program location is selected. 2056 2057 .. note:: 2058 2059 To allow ``DW_OP_call*`` to compute the location description for any 2060 variable or formal parameter regardless of whether the producer has 2061 optimized it to a constant, the following rule could be added: 2062 2063 .. note:: 2064 2065 If D has a ``DW_AT_const_value`` attribute, then a DWARF expression E 2066 consisting a ``DW_OP_implicit_value`` operation with the value of the 2067 ``DW_AT_const_value`` attribute is selected. 2068 2069 This would be consistent with ``DW_OP_implicit_pointer``. 2070 2071 Alternatively, could deprecate using ``DW_AT_const_value`` for 2072 ``DW_TAG_variable`` and ``DW_TAG_formal_parameter`` debugger information 2073 entries that are constants and instead use ``DW_AT_location`` with an 2074 implicit location description instead, then this rule would not be 2075 required. 2076 2077 Otherwise, an empty expression E is selected. 2078 2079 If D is a ``DW_TAG_dwarf_procedure`` debugging information entry, then E is 2080 evaluated using the same DWARF expression stack. Any existing stack entries 2081 may be accessed and/or removed in the evaluation of E, and the evaluation of 2082 E may add any new stack entries. 2083 2084 *Values on the stack at the time of the call may be used as parameters by 2085 the called expression and values left on the stack by the called expression 2086 may be used as return values by prior agreement between the calling and 2087 called expressions.* 2088 2089 Otherwise, E is evaluated on a separate DWARF stack and the resulting 2090 location description L is pushed on the ``DW_OP_call*`` operation's stack. 2091 2092 .. note: 2093 2094 In DWARF 5, if D does not have a ``DW_AT_location`` then ``DW_OP_call*`` 2095 is defined to have no effect. It is unclear that this is the right 2096 definition as a producer should be able to rely on using ``DW_OP_call*`` 2097 to get a location description for any non-\ ``DW_TAG_dwarf_procedure`` 2098 debugging information entries, and should not be creating DWARF with 2099 ``DW_OP_call*`` to a ``DW_TAG_dwarf_procedure`` that does not have a 2100 ``DW_AT_location`` attribute. 2101 210212. ``DW_OP_LLVM_call_frame_entry_reg`` *New* 2103 2104 ``DW_OP_LLVM_call_frame_entry_reg`` has a single unsigned LEB128 integer 2105 operand that is treated as a target architecture register number R. 2106 2107 It pushes a location description L that holds the value of register R on 2108 entry to the current subprogram as defined by the Call Frame Information 2109 (see :ref:`amdgpu-call-frame-information`). 2110 2111 *If there is no Call Frame Information defined, then the default rules for 2112 the target architecture are used. If the register rule is* undefined\ *, 2113 then the undefined location description is pushed. If the register rule is* 2114 same value\ *, then a register location description for R is pushed.* 2115 2116Undefined Location Descriptions 2117############################### 2118 2119The undefined location storage represents a piece or all of an object that is 2120present in the source but not in the object code (perhaps due to optimization). 2121Neither reading or writing to the undefined location storage is meaningful. 2122 2123An undefined location description specifies the undefined location storage. 2124There is no concept of the size of the undefined location storage, nor of a bit 2125offset for an undefined location description. The ``DW_OP_LLVM_*offset`` 2126operations leave an undefined location description unchanged. The 2127``DW_OP_*piece`` operations can explicitly or implicitly specify an undefined 2128location description, allowing any size and offset to be specified, and results 2129in a part with all undefined bits. 2130 21311. ``DW_OP_LLVM_undefined`` *New* 2132 2133 ``DW_OP_LLVM_undefined`` pushes an undefined location description L. 2134 2135Memory Location Descriptions 2136############################ 2137 2138There is a memory location storage that corresponds to each of the target 2139architecture linear memory address spaces. The size of each memory location 2140storage corresponds to the range of the addresses in the address space. 2141 2142*It is target architecture defined how address space location storage maps to 2143target architecture physical memory. For example, they may be independent memory 2144or more than one location storage may alias the same physical memory possibly at 2145different offsets and with different interleaving. The mapping may also be 2146dictated by the source language address classes.* 2147 2148A memory location description specifies a memory location storage. The bit 2149offset corresponds to an address in the address space scaled by 8 (the byte 2150size). Bits accessed using a memory location description, access the 2151corresponding target architecture memory starting at the bit offset. 2152 2153``DW_ASPACE_none`` is defined as the target architecture default address space. 2154 2155*The target architecture default address space for AMDGPU is the global address 2156space.* 2157 2158If a stack entry is required to be a location description, but it is a value 2159with the generic type, then it is implicitly convert to a memory location 2160description that specifies memory in the target architecture default address 2161space with a bit offset equal to the value scaled by 8 (the byte size). 2162 2163 .. note:: 2164 2165 If want to allow any integral type value to be implicitly converted to a 2166 memory location description in the target architecture default address 2167 space: 2168 2169 .. note:: 2170 2171 If a stack entry is required to be a location description, but it is a 2172 value with an integral type, then it is implicitly convert to a memory 2173 location description. The stack entry value is zero extended to the size 2174 of the generic type and the least significant generic type size bits are 2175 treated as a twos-complement unsigned value to be used as an address. The 2176 converted memory location description specifies memory location storage 2177 corresponding to the target architecture default address space with a bit 2178 offset equal to the address scaled by 8 (the byte size). 2179 2180 The implicit conversion could also be defined as target specific. For 2181 example, gdb checks if the value is an integral type. If it is not it gives 2182 an error. Otherwise, gdb zero-extends the value to 64 bits. If the gdb 2183 target defines a hook function then it is called and it can modify the 64 2184 bit value, possibly sign extending the original value. Finally, gdb treats 2185 the 64 bit value as a memory location address. 2186 2187If a stack entry is required to be a location description, but it is an implicit 2188pointer value IPV with the target architecture default address space, then it is 2189implicitly convert to the location description specified by IPV. See 2190:ref:`amdgpu-implicit-location-descriptions`. 2191 2192If a stack entry is required to be a value with a generic type, but it is a 2193memory location description in the target architecture default address space 2194with a bit offset that is a multiple of 8, then it is implicitly converted to a 2195value with a generic type that is equal to the bit offset divided by 8 (the byte 2196size). 2197 21981. ``DW_OP_addr`` 2199 2200 ``DW_OP_addr`` has a single byte constant value operand, which has the size 2201 of the generic type, treated as an address A. 2202 2203 It pushes a memory location description L on the stack that specifies the 2204 memory location storage for the target architecture default address space 2205 with a bit offset equal to A scaled by 8 (the byte size). 2206 2207 *If the DWARF is part of a code object, then A may need to be relocated. For 2208 example, in the ELF code object format, A must be adjusted by the difference 2209 between the ELF segment virtual address and the virtual address at which the 2210 segment is loaded.* 2211 22122. ``DW_OP_addrx`` 2213 2214 ``DW_OP_addrx`` has a single unsigned LEB128 integer operand that is treated 2215 as a zero-based index into the ``.debug_addr`` section relative to the value 2216 of the ``DW_AT_addr_base`` attribute of the associated compilation unit. The 2217 address value A in the ``.debug_addr`` section has the size of generic type. 2218 2219 It pushes a memory location description L on the stack that specifies the 2220 memory location storage for the target architecture default address space 2221 with a bit offset equal to A scaled by 8 (the byte size). 2222 2223 *If the DWARF is part of a code object, then A may need to be relocated. For 2224 example, in the ELF code object format, A must be adjusted by the difference 2225 between the ELF segment virtual address and the virtual address at which the 2226 segment is loaded.* 2227 22283. ``DW_OP_LLVM_form_aspace_address`` *New* 2229 2230 ``DW_OP_LLVM_form_aspace_address`` pops top two stack entries. The first 2231 must be an integral type value that is treated as an address space 2232 identifier AS for those architectures that support multiple address spaces. 2233 The second must be an integral type value that is treated as an address A. 2234 2235 The address size S is defined as the address bit size of the target 2236 architecture's address space that corresponds to AS. 2237 2238 A is adjusted by zero extending it to S bits and the least significant S 2239 bits are treated as a twos-complement unsigned value. 2240 2241 ``DW_OP_LLVM_form_aspace_address`` pushes a memory location description L 2242 that specifies the memory location storage that corresponds to AS, with a 2243 bit offset equal to the adjusted A scaled by 8 (the byte size). 2244 2245 If AS is not one of the values defined by the target architecture's 2246 ``DW_ASPACE_*`` values, then the DWARF expression is ill-formed. 2247 2248 See :ref:`amdgpu-implicit-location-descriptions` for special rules 2249 concerning implicit pointer values produced by dereferencing implicit 2250 location descriptions created by the ``DW_OP_implicit_pointer`` and 2251 ``DW_OP_LLVM_implicit_aspace_pointer`` operations. 2252 2253 The AMDGPU address spaces are defined in 2254 :ref:`amdgpu-dwarf-address-space-mapping-table`. 2255 22564. ``DW_OP_form_tls_address`` 2257 2258 ``DW_OP_form_tls_address`` pops one stack entry that must be an integral 2259 type value, and treats it as a thread-local storage address. 2260 2261 ``DW_OP_form_tls_address`` pushes a memory location description L for the 2262 target architecture default address space that corresponds to the 2263 thread-local storage address. 2264 2265 The meaning of the thread-local storage address is defined by the run-time 2266 environment. If the run-time environment supports multiple thread-local 2267 storage blocks for a single thread, then the block corresponding to the 2268 executable or shared library containing this DWARF expression is used. 2269 2270 *Some implementations of C, C++, Fortran, and other languages, support a 2271 thread-local storage class. Variables with this storage class have distinct 2272 values and addresses in distinct threads, much as automatic variables have 2273 distinct values and addresses in each function invocation. Typically, there 2274 is a single block of storage containing all thread-local variables declared 2275 in the main executable, and a separate block for the variables declared in 2276 each shared library. Each thread-local variable can then be accessed in its 2277 block using an identifier. This identifier is typically an offset into the 2278 block and pushed onto the DWARF stack by one of the* ``DW_OP_const<n><x>`` 2279 *operations prior to the* ``DW_OP_form_tls_address`` *operation. Computing 2280 the address of the appropriate block can be complex (in some cases, the 2281 compiler emits a function call to do it), and difficult to describe using 2282 ordinary DWARF location descriptions. Instead of forcing complex 2283 thread-local storage calculations into the DWARF expressions, the* 2284 ``DW_OP_form_tls_address`` *allows the consumer to perform the computation 2285 based on the run-time environment.* 2286 22875. ``DW_OP_call_frame_cfa`` 2288 2289 ``DW_OP_call_frame_cfa`` pushes the memory location description L of the 2290 Canonical Frame Address (CFA) of the current function, obtained from the 2291 Call Frame Information (see :ref:`amdgpu-call-frame-information`). 2292 2293 *Although the value of* ``DW_AT_frame_base`` *can be computed using other 2294 DWARF expression operators, in some cases this would require an extensive 2295 location list because the values of the registers used in computing the CFA 2296 change during a subroutine. If the Call Frame Information is present, then 2297 it already encodes such changes, and it is space efficient to reference 2298 that.* 2299 23006. ``DW_OP_fbreg`` 2301 2302 ``DW_OP_fbreg`` has a single signed LEB128 integer operand that is treated 2303 as a byte displacement D. 2304 2305 The DWARF expression E corresponding to the current program location is 2306 selected from the ``DW_AT_frame_base`` attribute of the current function and 2307 evaluated. The resulting memory location description L's bit offset is 2308 updated as if the ``DW_OP_LLVM_offset D`` operation were applied. The 2309 updated L is pushed. 2310 2311 *This is typically a stack pointer register plus or minus some offset.* 2312 23137. ``DW_OP_breg0, DW_OP_breg1, ..., DW_OP_breg31`` 2314 2315 The ``DW_OP_breg<n>`` operations encode the numbers of up to 32 registers, 2316 numbered from 0 through 31, inclusive. The register number R corresponds to 2317 the ``n`` in the operation name. 2318 2319 They have a single signed LEB128 integer operand that is treated as a byte 2320 displacement D. 2321 2322 The address space identifier AS is defined as the one corresponding to the 2323 target architecture's default address space. 2324 2325 The address size S is defined as the address bit size of the target 2326 architecture's address space corresponding to AS. 2327 2328 The contents of the register specified by R is retrieved as a 2329 twos-complement unsigned value and zero extended to S bits. D is added and 2330 the least significant S bits are treated as a twos-complement unsigned value 2331 to be used as an address A. 2332 2333 They push a memory location description L that specifies the memory location 2334 storage that corresponds to AS, with a bit offset equal to A scaled by 8 2335 (the byte size). 2336 23378. ``DW_OP_bregx`` 2338 2339 ``DW_OP_bregx`` has two operands. The first is an unsigned LEB128 integer 2340 that is treated as a register number R. The second is a signed LEB128 2341 integer that is treated as a byte displacement D. 2342 2343 The action is the same as for ``DW_OP_breg<n>`` except that R is used as the 2344 register number and D is used as the byte displacement. 2345 23469. ``DW_OP_LLVM_aspace_bregx`` *New* 2347 2348 ``DW_OP_LLVM_aspace_bregx`` has two operands. The first is an unsigned 2349 LEB128 integer that is treated as a register number R. The second is a 2350 signed LEB128 integer that is treated as a byte displacement D. It pops one 2351 stack entry that is required to be an integral type value that is treated as 2352 an address space identifier AS for those architectures that support multiple 2353 address spaces. 2354 2355 The action is the same as for ``DW_OP_breg<n>`` except that R is used as the 2356 register number, D is used as the byte displacement, and AS is used as the 2357 address space identifier. 2358 2359 If AS is not one of the values defined by the target architecture's 2360 ``DW_ASPACE_*`` values, then the DWARF expression is ill-formed. 2361 2362 .. note:: 2363 2364 Could also consider adding ``DW_OP_aspace_breg0, DW_OP_aspace_breg1, ..., 2365 DW_OP_aspace_bref31`` which would save encoding size. 2366 2367.. _amdgpu-register-location-descriptions: 2368 2369Register Location Descriptions 2370############################## 2371 2372There is a register location storage that corresponds to each of the target 2373architecture registers. The size of each register location storage corresponds 2374to the size of the corresponding target architecture register. 2375 2376A register location description specifies a register location storage. The bit 2377offset corresponds to a bit position within the register. Bits accessed using a 2378register location description, access the corresponding target architecture 2379register starting at the bit offset. 2380 23811. ``DW_OP_reg0, DW_OP_reg1, ..., DW_OP_reg31`` 2382 2383 ``DW_OP_reg<n>`` operations encode the numbers of up to 32 registers, 2384 numbered from 0 through 31, inclusive. The target architecture register 2385 number R corresponds to the ``n`` in the operation name. 2386 2387 ``DW_OP_reg<n>`` pushes a register location description L that specifies the 2388 register location storage that corresponds to R, with a bit offset of 0. 2389 23902. ``DW_OP_regx`` 2391 2392 ``DW_OP_regx`` has a single unsigned LEB128 integer operand that is treated 2393 as a target architecture register number R. 2394 2395 ``DW_OP_regx`` pushes a register location description L that specifies the 2396 register location storage that corresponds to R, with a bit offset of 0. 2397 2398*These operations name a register location. To fetch the contents of a register, 2399it is necessary to use* ``DW_OP_regval_type``\ *, or one of the register based 2400addressing operations such as* ``DW_OP_bregx``\ *, or using* ``DW_OP_deref*`` 2401*on a register location description.* 2402 2403.. _amdgpu-implicit-location-descriptions: 2404 2405Implicit Location Descriptions 2406############################## 2407 2408Implicit location storage represents a piece or all of an object which has no 2409actual location in the program but whose contents are nonetheless known, either 2410as a constant or can be computed from other locations and values in the program. 2411 2412An implicit location description specifies an implicit location storage. The bit 2413offset corresponds to a bit position within the implicit location storage. Bits 2414accessed using an implicit location description, access the corresponding 2415implicit storage value starting at the bit offset. 2416 24171. ``DW_OP_implicit_value`` 2418 2419 ``DW_OP_implicit_value`` has two operands. The first is an unsigned LEB128 2420 integer treated as a byte size S. The second is a block of bytes with a 2421 length equal to S treated as a literal value V. 2422 2423 An implicit location storage LS is created with the literal value V and a 2424 size of S. An implicit location description L is pushed that specifies LS 2425 with a bit offset of 0. 2426 24272. ``DW_OP_stack_value`` 2428 2429 ``DW_OP_stack_value`` pops one stack entry that must be a value treated as a 2430 literal value V. 2431 2432 An implicit location storage LS is created with the literal value V and a 2433 size equal to V's base type size. An implicit location description L is 2434 pushed that specifies LS with a bit offset of 0. 2435 2436 The ``DW_OP_stack_value`` operation specifies that the object does not exist 2437 in memory but its value is nonetheless known and is at the top of the DWARF 2438 expression stack. In this form of location description, the DWARF expression 2439 represents the actual value of the object, rather than its location. 2440 2441 See :ref:`amdgpu-implicit-location-descriptions` for special rules 2442 concerning implicit pointer values produced by dereferencing implicit 2443 location descriptions created by the ``DW_OP_implicit_pointer`` and 2444 ``DW_OP_LLVM_implicit_aspace_pointer`` operations. 2445 2446 .. note:: 2447 2448 Since location descriptions are allowed on the stack, the 2449 ``DW_OP_stack_value`` operation no longer terminates the DWARF expression. 2450 24513. ``DW_OP_implicit_pointer`` 2452 2453 *An optimizing compiler may eliminate a pointer, while still retaining the 2454 value that the pointer addressed.* ``DW_OP_implicit_pointer`` *allows a 2455 producer to describe this value.* 2456 2457 ``DW_OP_implicit_pointer`` specifies that the object is a pointer to the 2458 target architecture default address space that cannot be represented as a 2459 real pointer, even though the value it would point to can be described. In 2460 this form of location description, the DWARF expression refers to a 2461 debugging information entry that represents the actual location description 2462 of the object to which the pointer would point. Thus, a consumer of the 2463 debug information would be able to access the the dereferenced pointer, even 2464 when it cannot access of the pointer itself. 2465 2466 ``DW_OP_implicit_pointer`` has two operands. The first is a 4-byte unsigned 2467 value in the 32-bit DWARF format, or an 8-byte unsigned value in the 64-bit 2468 DWARF format, that is treated as a debugging information entry reference R. 2469 The second is a signed LEB128 integer that is treated as a byte 2470 displacement D. 2471 2472 R is used as the offset of a debugging information entry E in a 2473 ``.debug_info`` section, which may be contained in an executable or shared 2474 object file other than that containing the operator. For references from one 2475 executable or shared object file to another, the relocation must be 2476 performed by the consumer. 2477 2478 *The first operand interpretation is exactly like that for* 2479 ``DW_FORM_ref_addr``\ *.* 2480 2481 The address space identifier AS is defined as the one corresponding to the 2482 target architecture's default address space. 2483 2484 The address size S is defined as the address bit size of the target 2485 architecture's address space corresponding to AS. 2486 2487 An implicit location storage LS is created that has the bit size of S. An 2488 implicit location description L is pushed that specifies LS and has a bit 2489 offset of 0. 2490 2491 If a ``DW_OP_deref*`` operation pops a location description L' and retrieves 2492 S' bits where some retrieved bits come from LS such that either: 2493 2494 1. L' is an implicit location description that specifies LS with bit offset 2495 0, and S' equals S. 2496 2497 2. L' is a complete composite location description that specifies a 2498 canonical form composite location storage LS'. The bits retrieved all 2499 come from a single part P' of LS'. P' has a bit size of S and has 2500 an implicit location description PL'. PL' specifies LS with a bit offset 2501 of 0. 2502 2503 Then the value V pushed by the ``DW_OP_deref*`` operation is an implicit 2504 pointer value IPV with an address space of AS, a debugging information entry 2505 of E, and a base type of T. If AS is the target architecture default address 2506 space, then T is the generic type. Otherwise, T is an architecture specific 2507 integral type with a bit size equal to S. 2508 2509 Otherwise, if a ``DW_OP_deref*`` operation is applied to a location 2510 description such that some retrieved bits come from LS, then the DWARF 2511 expression is ill-formed. 2512 2513 If IPV is either implicitly converted to a location description (only done 2514 if AS is the target architecture default address space) or used by 2515 ``DW_OP_LLVM_form_aspace_address`` (only done if the address space specified 2516 is AS), then the resulting location description is: 2517 2518 * If E has a ``DW_AT_location`` attribute, the DWARF expression 2519 corresponding to the current program location is selected and evaluated 2520 from the ``DW_AT_location`` attribute. The expression result is the 2521 resulting location description RL. 2522 2523 * If E has a ``DW_AT_const_value`` attribute, then an implicit location 2524 storage RLS is created from the ``DW_AT_const_value`` attribute's value, 2525 with a size matching the size of the ``DW_AT_const_value`` attribute's 2526 value. The resulting implicit location description RL specifies RLS with a 2527 bit offset of 0. 2528 2529 .. note:: 2530 2531 If deprecate using ``DW_AT_const_value`` for variables and formal 2532 parameters and instead use ``DW_AT_location`` with an implicit location 2533 description instead, then this rule would not be required. 2534 2535 * Otherwise the DWARF expression is ill-formed. 2536 2537 The bit offset of RL is updated as if the ``DW_OP_LLVM_offset D`` operation 2538 were applied. 2539 2540 If a ``DW_OP_stack_value`` operation pops a value that is the same as IPV, 2541 then it pushes a location description that is the same as L. 2542 2543 The DWARF expression is ill-formed if it accesses LS or IPV in any other 2544 manner. 2545 2546 *The restrictions on how an implicit pointer location description created by 2547 ``DW_OP_implicit_pointer`` and ``DW_OP_LLVM_aspace_implicit_pointer``, or an 2548 implicit pointer value created by ``DW_OP_deref*``, can be used are to 2549 simplify the DWARF consumer.* 2550 25514. ``DW_OP_LLVM_aspace_implicit_pointer`` *New* 2552 2553 ``DW_OP_LLVM_aspace_implicit_pointer`` has two operands that are the same as 2554 for ``DW_OP_implicit_pointer``. 2555 2556 It pops one stack entry that must be an integral type value that is treated 2557 as an address space identifier AS for those architectures that support 2558 multiple address spaces. 2559 2560 The implicit location description L that is pushed is the same as for 2561 ``DW_OP_implicit_pointer`` except that the address space identifier used is 2562 AS. 2563 2564 If AS is not one of the values defined by the target architecture's 2565 ``DW_ASPACE_*`` values, then the DWARF expression is ill-formed. 2566 2567*The debugging information entry referenced by a* ``DW_OP_implicit_pointer`` or 2568``DW_OP_LLVM_aspace_implicit_pointer`` *operation is typically a* 2569``DW_TAG_variable`` *or* ``DW_TAG_formal_parameter`` *entry whose* 2570``DW_AT_location`` *attribute gives a second DWARF expression or a location list 2571that describes the value of the object, but the referenced entry may be any 2572entry that contains a* ``DW_AT_location`` *or* ``DW_AT_const_value`` *attribute 2573(for example,* ``DW_TAG_dwarf_procedure``\ *). By using the second DWARF 2574expression, a consumer can reconstruct the value of the object when asked to 2575dereference the pointer described by the original DWARF expression containing 2576the* ``DW_OP_implicit_pointer`` or ``DW_OP_LLVM_aspace_implicit_pointer`` 2577*operation.* 2578 2579Composite Location Descriptions 2580############################### 2581 2582A composite location storage represents an object or value which may be 2583contained in part of another location storage, or contained in parts of more 2584than one location storage. 2585 2586Each part has a part location description L and a part bit size S. The bits of 2587the part comprise S contiguous bits from the location storage specified by L, 2588starting at the bit offset specified by L. All the bits must be within the size 2589of the location storage specified by L or the DWARF expression is ill-formed. 2590 2591A composite location storage can have zero or more parts. The parts are 2592contiguous such that the zero-based location storage bit index will range over 2593each part with no gaps between them. Therefore, the size of a composite location 2594storage is the size of its parts. The DWARF expression is ill-formed if the size 2595of the contiguous location storage is larger than the size of the memory 2596location storage corresponding to the target architecture's largest address 2597space. 2598 2599The canonical form of a composite location storage is computed by applying the 2600following steps to a composite location storage: 2601 26021. If any part P has a composite location description L, it is replaced by a 2603 copy of the parts of the composite location storage specified by L that are 2604 selected by the bit size of P starting at the bit offset of L. The location 2605 description of the first copied part has its bit offset updated as 2606 necessary, and the last copied part has its bit size updated as necessary, 2607 to reflect the bits selected by P. This rule is applied repeatedly until no 2608 part has a composite location description. 2609 26102. If the size on any part is zero, it is removed. 2611 26123. If any adjacent parts P\ :sup:`1` to P\ :sup:`n` have location descriptions 2613 that specify the same location storage LS such that the bits selected form a 2614 contiguous portion of LS, then they are replaced by a single new part P'. P' 2615 has a location description L that specifies LS with the same bit offset as 2616 P\ :sup:`1`\ 's location description, and a bit size equal to the sum of the 2617 bit sizes of P\ :sup:`1` to P\ :sup:`n` inclusive. 2618 2619A composite location description specifies the canonical form of a composite 2620location storage and a bit offset. 2621 2622There are operations that push a composite location description that specifies a 2623composite location storage that is created by the operation. 2624 2625There are other operations that allow a composite location storage and a 2626composite location description that specifies it to be created incrementally. 2627Each part is described by a separate operation. There may be one or more 2628operations to create the final composite location storage and associated 2629description. A series of such operations describes the parts of the composite 2630location storage that are in the order that the associated part operations are 2631executed. 2632 2633To support incremental creation, a composite location description can be in an 2634incomplete state. When an incremental operation operates on an incomplete 2635composite location description, it adds a new part, otherwise it creates a new 2636composite location description. The ``DW_OP_LLVM_piece_end`` operation 2637explicitly makes an incomplete composite location description complete. 2638 2639If the top stack entry is an incomplete composite location description after the 2640execution of a DWARF expression has completed, it is converted to a complete 2641composite location description. 2642 2643If a stack entry is required to be a location description, but it is an 2644incomplete composite location description, then the DWARF expression is 2645ill-formed. 2646 2647*Note that a DWARF expression may arbitrarily compose composite location 2648descriptions from any other location description, including other composite 2649location descriptions.* 2650 2651*The incremental composite location description operations are defined to be 2652compatible with the definitions in DWARF 5 and earlier.* 2653 26541. ``DW_OP_piece`` 2655 2656 ``DW_OP_piece`` has a single unsigned LEB128 integer that is treated as a 2657 byte size S. 2658 2659 The action is based on the context: 2660 2661 * If the stack is empty, then an incomplete composite location description 2662 L is pushed that specifies a new composite location storage LS and has a 2663 bit offset of 0. LS has a single part P that specifies the undefined 2664 location description, and has a bit size of S scaled by 8 (the byte size). 2665 2666 * If the top stack entry is an incomplete composite location description L, 2667 then the composite location storage LS that it specifies is updated to 2668 append a part that specifies an undefined location description, and has a 2669 bit size S scaled by 8 (the byte size). 2670 2671 * If the top stack entry is a location description or can be converted to 2672 one, then it is popped and treated as a part location description PL. 2673 Then: 2674 2675 * If the stack is empty or the top stack entry is not an incomplete 2676 composite location description, then an incomplete composite location 2677 description L is pushed that specifies a new composite location storage 2678 LS. LS has a single part that specifies PL, and has a bit size of S 2679 scaled by 8 (the byte size). 2680 2681 * Otherwise, the composite location storage LS specified by the top stack 2682 incomplete composite location description L is updated to append a part 2683 that specifies PL, and has a bit size S scaled by 8 (the byte size). 2684 2685 * Otherwise, the DWARF expression is ill-formed 2686 2687 If LS is not in canonical form it is updated to be in canonical form. 2688 2689 *Many compilers store a single variable in sets of registers, or store a 2690 variable partially in memory and partially in registers.* ``DW_OP_piece`` 2691 *provides a way of describing how large a part of a variable a particular 2692 DWARF location description refers to.* 2693 2694 *If a computed byte displacement is required, the* ``DW_OP_LLVM_offset`` 2695 *can be used to update the part location description.* 2696 26972. ``DW_OP_bit_piece`` 2698 2699 ``DW_OP_bit_piece`` has two operands. The first is an unsigned LEB128 2700 integer that is treated as the part bit size S. The second is an unsigned 2701 LEB128 integer that is treated as a bit displacement D. 2702 2703 The action is the same as for ``DW_OP_piece`` except that any part created 2704 has the bit size S, and the location description of any created part has its 2705 bit offset updated as if the ``DW_OP_LLVM_bit_offset D`` operation were 2706 applied. 2707 2708 *If a computed bit displacement is required, the* ``DW_OP_LLVM_bit_offset`` 2709 *can be used to update the part location description.* 2710 2711 .. note:: 2712 2713 The bit offset operand is not needed as ``DW_OP_LLVM_bit_offset`` can be 2714 used on the part's location description. 2715 27163. ``DW_OP_LLVM_piece_end`` *New* 2717 2718 If the top stack entry is an incomplete composite location description L, 2719 then it is updated to be a complete composite location description with the 2720 same parts. Otherwise, the DWARF expression is ill-formed. 2721 27224. ``DW_OP_LLVM_extend`` *New* 2723 2724 ``DW_OP_LLVM_extend`` has two operands. The first is an unsigned LEB128 2725 integer that is treated as the element bit size S. The second is an unsigned 2726 LEB128 integer that is treated as a count C. 2727 2728 It pops one stack entry that must be a location description and is treated 2729 as the part location description PL. 2730 2731 A complete composite location description L is pushed that comprises C parts 2732 that each specify PL and have a bit size of S. 2733 2734 The DWARF expression is ill-formed if the element bit size or count are 0. 2735 27365. ``DW_OP_LLVM_select_bit_piece`` *New* 2737 2738 ``DW_OP_LLVM_select_bit_piece`` has two operands. The first is an unsigned 2739 LEB128 integer that is treated as the element bit size S. The second is an 2740 unsigned LEB128 integer that is treated as a count C. 2741 2742 It pops three stack entries. The first must be an integral type value that 2743 is treated as a bit mask value M. The second must be a location description 2744 that is treated as the one-location description L1. The third must be a 2745 location description that is treated as the zero-location description L0. 2746 2747 A complete composite location description L is pushed that specifies a new 2748 composite location storage LS. LS comprises C parts that each specify a part 2749 location description PL and have a bit size of S. The PL for part N is 2750 defined as: 2751 2752 1. If the Nth least significant bit of M is a zero then the PL for part N 2753 is the same as L0, otherwise it is the same as L1. 2754 2755 2. The PL for part N is updated as if the ``DW_OP_LLVM_bit_offset N*S`` 2756 operation was applied. 2757 2758 If LS is not in canonical form it is updated to be in canonical form. 2759 2760 The DWARF expression is ill-formed if S or C are 0, or if the bit size of M 2761 is less than C. 2762 2763``DW_OP_bit_piece`` *is used instead of* ``DW_OP_piece`` *when the piece to be 2764assembled into a value or assigned to is not byte-sized or is not at the start 2765of the part location description.* 2766 2767.. note:: 2768 2769 For AMDGPU: 2770 2771 * In CFI expressions ``DW_OP_LLVM_select_bit_piece`` is used to describe 2772 unwinding vector registers that are spilled under the execution mask to 2773 memory: the zero location description is the vector register, and the one 2774 location description is the spilled memory location. The 2775 ``DW_OP_LLVM_form_aspace_address`` is used to specify the address space of 2776 the memory location description. 2777 2778 * ``DW_OP_LLVM_select_bit_piece`` is used by the ``lane_pc`` attribute 2779 expression where divergent control flow is controlled by the execution mask. 2780 An undefined location description together with ``DW_OP_LLVM_extend`` is 2781 used to indicate the lane was not active on entry to the subprogram. 2782 2783Expression Operation Encodings 2784++++++++++++++++++++++++++++++ 2785 2786The following table gives the encoding of the DWARF expression operations added 2787for AMDGPU. 2788 2789.. table:: AMDGPU DWARF Expression Operation Encodings 2790 :name: amdgpu-dwarf-expression-operation-encodings-table 2791 2792 ================================== ===== ======== =============================== 2793 Operation Code Number Notes 2794 of 2795 Operands 2796 ================================== ===== ======== =============================== 2797 DW_OP_LLVM_form_aspace_address 0xe7 0 2798 DW_OP_LLVM_push_lane 0xea 0 2799 DW_OP_LLVM_offset 0xe9 0 2800 DW_OP_LLVM_offset_uconst *TBD* 1 ULEB128 byte displacement 2801 DW_OP_LLVM_bit_offset *TBD* 0 2802 DW_OP_LLVM_call_frame_entry_reg *TBD* 1 ULEB128 register number 2803 DW_OP_LLVM_undefined *TBD* 0 2804 DW_OP_LLVM_aspace_bregx *TBD* 2 ULEB128 register number, 2805 ULEB128 byte displacement 2806 DW_OP_LLVM_aspace_implicit_pointer *TBD* 2 4- or 8-byte offset of DIE, 2807 SLEB128 byte displacement 2808 DW_OP_LLVM_piece_end *TBD* 0 2809 DW_OP_LLVM_extend *TBD* 2 ULEB128 bit size, 2810 ULEB128 count 2811 DW_OP_LLVM_select_bit_piece *TBD* 2 ULEB128 bit size, 2812 ULEB128 count 2813 ================================== ===== ======== =============================== 2814 2815.. _amdgpu-dwarf-debugging-information-entry-attributes: 2816 2817Debugging Information Entry Attributes 2818~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 2819 2820This section provides changes to existing debugger information attributes and 2821defines attributes added by the AMDGPU target. 2822 28231. ``DW_AT_location`` 2824 2825 If the result of the ``DW_AT_location`` DWARF expression is required to be a 2826 location description, then it may have any kind of location description (see 2827 :ref:`amdgpu-location-description-operations`). 2828 28292. ``DW_AT_const_value`` 2830 2831 .. note:: 2832 2833 Could deprecate using the ``DW_AT_const_value`` attribute for 2834 ``DW_TAG_variable`` or ``DW_TAG_formal_parameter`` debugger information 2835 entries that are constants. Instead, ``DW_AT_location`` could be used with 2836 a DWARF expression that produces an implicit location description now that 2837 any location description can be used within a DWARF expression. This 2838 allows the ``DW_OP_call*`` operations to be used to push the location 2839 description of any variable regardless of how it is optimized. 2840 28413. ``DW_AT_frame_base`` 2842 2843 A ``DW_TAG_subprogram`` or ``DW_TAG_entry_point`` debugger information entry 2844 may have a ``DW_AT_frame_base`` attribute, whose value is a DWARF expression 2845 or location list that describes the *frame base* for the subroutine or entry 2846 point. 2847 2848 If the result of the DWARF expression is a register location description, 2849 then the ``DW_OP_deref`` operation is applied to compute the frame base 2850 memory location description in the target architecture default address 2851 space. 2852 2853 .. note:: 2854 2855 This rule could be removed and require the producer to create the 2856 required location descriptor directly using ``DW_OP_call_frame_cfa``, 2857 ``DW_OP_fbreg``, ``DW_OP_breg*``, or ``DW_OP_LLVM-aspace_bregx``. This 2858 would also then allow a target to implement the call frames withing a 2859 large register. 2860 2861 Otherwise, the result of the DWARF expression is required to be a memory 2862 location description in any of the target architecture address spaces which 2863 is the frame base. 2864 28654. ``DW_AT_data_member_location`` 2866 2867 For a ``DW_AT_data_member_location`` attribute there are two cases: 2868 2869 1. If the value is an integer constant, it is the offset in bytes from the 2870 beginning of the containing entity. If the beginning of the containing 2871 entity has a non-zero bit offset then the beginning of the member entry 2872 has that same bit offset as well. 2873 2874 2. Otherwise, the value must be a DWARF expression or location list. The 2875 DWARF expression E corresponding to the current program location is 2876 selected. The location description of the beginning of the containing 2877 entity is pushed on the DWARF stack before E is evaluated. The result of 2878 the evaluation is the location description of the base of the member 2879 entry. 2880 2881 .. note:: 2882 2883 The beginning of the containing entity can now be any location 2884 description and can be bit aligned. 2885 28865. ``DW_AT_use_location`` 2887 2888 The ``DW_TAG_ptr_to_member_type`` debugging information entry has a 2889 ``DW_AT_use_location`` attribute whose value is a DWARF expression or 2890 location list. The DWARF expression E corresponding to the current program 2891 location is selected. It is used to computes the location description of the 2892 member of the class to which the pointer to member entry points 2893 2894 *The method used to find the location description of a given member of a 2895 class or structure is common to any instance of that class or structure and 2896 to any instance of the pointer or member type. The method is thus associated 2897 with the type entry, rather than with each instance of the type.* 2898 2899 The ``DW_AT_use_location`` description is used in conjunction with the 2900 location descriptions for a particular object of the given pointer to member 2901 type and for a particular structure or class instance. 2902 2903 Two values are pushed onto the DWARF expression stack before E is evaluated. 2904 The first value pushed is the value of the pointer to member object itself. 2905 The second value pushed is the location description of the base of the 2906 entire structure or union instance containing the member whose address is 2907 being calculated. 2908 29096. ``DW_AT_data_location`` 2910 2911 The ``DW_AT_data_location`` attribute may be used with any type that 2912 provides one or more levels of hidden indirection and/or run-time parameters 2913 in its representation. Its value is a DWARF expression E which computes the 2914 location description of the data for an object. When this attribute is 2915 omitted, the location description of the data is the same as the location 2916 description of the object. 2917 2918 *E will typically begin with ``DW_OP_push_object_address`` which loads the 2919 location description of the object which can then serve as a descriptor in 2920 subsequent calculation.* 2921 29227. ``DW_AT_vtable_elem_location`` 2923 2924 An entry for a virtual function also has a ``DW_AT_vtable_elem_location`` 2925 attribute whose value is a DWARF expression or location list. The DWARF 2926 expression E corresponding to the current program location is selected. The 2927 location description of the object of the enclosing type is pushed onto the 2928 expression stack before E is evaluated. The resulting location description 2929 is the slot for the function within the virtual function table for the 2930 enclosing class. 2931 29328. ``DW_AT_static_link`` 2933 2934 If a ``DW_TAG_subprogram`` or ``DW_TAG_entry_point`` debugger information 2935 entry is nested, it may have a ``DW_AT_static_link`` attribute, whose value 2936 is a DWARF expression or location list. The DWARF expression E corresponding 2937 to the current program location is selected. The result of evaluating E is 2938 the frame base memory location description of the relevant instance of the 2939 subroutine that immediately encloses the subroutine or entry point. 2940 29419. ``DW_AT_return_addr`` 2942 2943 A ``DW_TAG_subprogram``, ``DW_TAG_inlined_subroutine``, or 2944 ``DW_TAG_entry_point`` debugger information entry may have a 2945 ``DW_AT_return_addr`` attribute, whose value is a DWARF expression or 2946 location list. The DWARF expression E corresponding to the current program 2947 location is selected. The result of evaluating E is the location description 2948 for the place where the return address for the subroutine or entry point is 2949 stored. 2950 2951 .. note:: 2952 2953 It is unclear why ``DW_TAG_inlined_subroutine`` has a 2954 ``DW_AT_return_addr`` attribute but not a ``DW_AT_frame_base`` or 2955 ``DW_AT_static_link`` attribute. Seems it would either have all of them or 2956 none. Since inlined subprograms do not have a frame it seems they would 2957 have none of these attributes. 2958 295910. ``DW_AT_LLVM_lanes`` *New* 2960 2961 For languages that are implemented using a SIMD or SIMT execution model, a 2962 ``DW_TAG_subprogram``, ``DW_TAG_inlined_subroutine``, or 2963 ``DW_TAG_entry_point`` debugger information entry may have a 2964 ``DW_AT_LLVM_lanes`` attribute whose value is an integer constant that is 2965 the number of lanes per thread. 2966 2967 If not present, the default value of 1 is used. 2968 2969 The DWARF is ill-formed if the value is 0. 2970 297111. ``DW_AT_LLVM_lane_pc`` *New* 2972 2973 For languages that are implemented using a SIMD or SIMT execution model, a 2974 ``DW_TAG_subprogram``, ``DW_TAG_inlined_subroutine``, or 2975 ``DW_TAG_entry_point`` debugging information entry may have a 2976 ``DW_AT_LLVM_lane_pc`` attribute whose value is a DWARF expression or 2977 location list. The DWARF expression E corresponding to the current program 2978 location is selected. The result of evaluating E is a location description 2979 that references a wave size vector of generic type elements. Each element 2980 holds the conceptual program location of the corresponding lane, where the 2981 least significant element corresponds to the first target architecture lane 2982 identifier and so forth. If the lane was not active when the subprogram was 2983 called, its element is an undefined location description. 2984 2985 *``DW_AT_LLVM_lane_pc`` allows the compiler to indicate conceptually where 2986 each lane of a SIMT thread is positioned even when it is in divergent 2987 control flow that is not active.* 2988 2989 If not present, the thread is not being used in a SIMT manner, and the 2990 thread's program location is used. 2991 2992 *See* :ref:`amdgpu-dwarf-amdgpu-dw-at-llvm-lane-pc` *for AMDGPU 2993 information.* 2994 299512. ``DW_AT_LLVM_active_lane`` *New* 2996 2997 For languages that are implemented using a SIMD or SIMT execution model, a 2998 ``DW_TAG_subprogram``, ``DW_TAG_inlined_subroutine``, or 2999 ``DW_TAG_entry_point`` debugger information entry may have a 3000 ``DW_AT_LLVM_active_lane`` attribute whose value is a DWARF expression or 3001 location list. The DWARF expression E corresponding to the current program 3002 location is selected. The result of evaluating E is a integral value that is 3003 the mask of active lanes for the current program location. The Nth least 3004 significant bit of the mask corresponds to the Nth lane. If the bit is 1 the 3005 lane is active, otherwise it is inactive. 3006 3007 *Some targets may update the target architecture execution mask for regions 3008 of code that must execute with different sets of lanes than the current 3009 active lanes. For example, some code must execute in whole wave mode. 3010 ``DW_AT_LLVM_active_lane` allows the compiler can provide the means to 3011 determine the actual active lanes.* 3012 3013 If not present and ``DW_AT_LLVM_lanes`` is greater than 1, then the target 3014 architecture execution mask is used. 3015 3016 *See* :ref:`amdgpu-dwarf-amdgpu-dw-at-llvm-active-lane` *for AMDGPU 3017 information.* 3018 301913. ``DW_AT_LLVM_vector_size`` *New* 3020 3021 A base type V may have the ``DW_AT_LLVM_vector_size`` attribute whose value 3022 is an integer constant that is the vector size S. 3023 3024 The representation of a vector base type is as S contiguous elements, each 3025 one having the representation of a base type E that is the same as V without 3026 the ``DW_AT_LLVM_vector_size`` attribute. 3027 3028 If not present, the base type is not a vector. 3029 3030 The DWARF is ill-formed if S not greater than 0. 3031 3032 .. note:: 3033 3034 LLVM has mention of non-upstreamed debugger information entry that is 3035 intended to support vector types. However, that was not for a base type 3036 so would not be suitable as the type of a stack value entry. But perhaps 3037 that could be replaced by using this attribute. 3038 303914. ``DW_AT_LLVM_augmentation`` *New* 3040 3041 A compilation unit may have a ``DW_AT_LLVM_augmentation`` attribute, whose 3042 value is an augmentation string. 3043 3044 *The augmentation string allows users to indicate that there is additional 3045 target-specific information in the debugging information entries. For 3046 example, this might be information about the version of target-specific 3047 extensions that are being used.* 3048 3049 If not present, or if the string is empty, then the compilation unit has no 3050 augmentation string. 3051 3052 .. note:: 3053 3054 For AMDGPU, the augmentation string contains: 3055 3056 :: 3057 3058 [amd:v0.0] 3059 3060 The "vX.Y" specifies the major X and minor Y version number of the AMDGPU 3061 extensions used in the DWARF of the compilation unit. The version number 3062 conforms to [SEMVER]_. 3063 3064Attribute Encodings 3065+++++++++++++++++++ 3066 3067The following table gives the encoding of the debugging information entry 3068attributes added for AMDGPU. 3069 3070.. table:: AMDGPU DWARF Attribute Encodings 3071 :name: amdgpu-dwarf-attribute-encodings-table 3072 3073 ================================== ===== ==================================== 3074 Attribute Name Value Classes 3075 ================================== ===== ==================================== 3076 DW_AT_LLVM_lanes constant 3077 DW_AT_LLVM_lane_pc exprloc, loclist 3078 DW_AT_LLVM_active_lane exprloc, loclist 3079 DW_AT_LLVM_vector_size constant 3080 DW_AT_LLVM_augmentation string 3081 ================================== ===== ==================================== 3082 3083.. _amdgpu-call-frame-information: 3084 3085Call Frame Information 3086~~~~~~~~~~~~~~~~~~~~~~ 3087 3088DWARF Call Frame Information describes how an agent can virtually *unwind* 3089call frames in a running process or core dump. 3090 3091.. note:: 3092 3093 AMDGPU conforms to the DWARF standard with additional support added for 3094 address spaces. Register unwind DWARF expressions are generalized to allow any 3095 location description, including composite and implicit location descriptions. 3096 3097Structure of Call Frame Information 3098+++++++++++++++++++++++++++++++++++ 3099 3100The register rules are: 3101 3102*undefined* 3103 A register that has this rule has no recoverable value in the previous frame. 3104 (By convention, it is not preserved by a callee.) 3105 3106*same value* 3107 This register has not been modified from the previous frame. (By convention, 3108 it is preserved by the callee, but the callee has not modified it.) 3109 3110*offset(N)* 3111 The previous value of this register is saved at the location description 3112 computed as if the ``DW_OP_LLVM_offset N`` operation is applied to the current 3113 CFA memory location description where N is a signed byte offset. 3114 3115*val_offset(N)* 3116 The previous value of this register is the address in the address space of the 3117 memory location description computed as if the ``DW_OP_LLVM_offset N`` 3118 operation is applied to the current CFA memory location description where N is 3119 a signed byte displacement. 3120 3121 If the register size does not match the size of an address in the address 3122 space of the current CFA memory location description, then the DWARF is 3123 ill-formed . 3124 3125*register(R)* 3126 The previous value of this register is stored in another register numbered R. 3127 3128 If the register sizes do not match, then the DWARF is ill-formed. 3129 3130*expression(E)* 3131 The previous value of this register is located at the location description 3132 produced by executing the DWARF expression E (see 3133 :ref:`amdgpu-dwarf-expressions`). 3134 3135*val_expression(E)* 3136 The previous value of this register is the value produced by executing the 3137 DWARF expression E (see :ref:`amdgpu-dwarf-expressions`). 3138 3139 If value type size does not match the register size, then the DWARF is 3140 ill-formed. 3141 3142*architectural* 3143 The rule is defined externally to this specification by the augmenter. 3144 3145A Common Information Entry holds information that is shared among many Frame 3146Description Entries. There is at least one CIE in every non-empty 3147``.debug_frame`` section. A CIE contains the following fields, in order: 3148 31491. ``length`` (initial length) 3150 3151 A constant that gives the number of bytes of the CIE structure, not 3152 including the length field itself. The size of the length field plus the 3153 value of length must be an integral multiple of the address size specified 3154 in the ``address_size`` field. 3155 31562. ``CIE_id`` (4 or 8 bytes, see 3157 :ref:`amdgpu-dwarf-32-bit-and-64-bit-dwarf-formats`) 3158 3159 A constant that is used to distinguish CIEs from FDEs. 3160 3161 In the 32-bit DWARF format, the value of the CIE id in the CIE header is 3162 0xffffffff; in the 64-bit DWARF format, the value is 0xffffffffffffffff. 3163 31643. ``version`` (ubyte) 3165 3166 A version number. This number is specific to the call frame information and 3167 is independent of the DWARF version number. 3168 3169 The value of the CIE version number is 4. 3170 31714. ``augmentation`` (sequence of UTF-8 characters) 3172 3173 A null-terminated UTF-8 string that identifies the augmentation to this CIE 3174 or to the FDEs that use it. If a reader encounters an augmentation string 3175 that is unexpected, then only the following fields can be read: 3176 3177 * CIE: length, CIE_id, version, augmentation 3178 * FDE: length, CIE_pointer, initial_location, address_range 3179 3180 If there is no augmentation, this value is a zero byte. 3181 3182 *The augmentation string allows users to indicate that there is additional 3183 target-specific information in the CIE or FDE which is needed to virtually 3184 unwind a stack frame. For example, this might be information about 3185 dynamically allocated data which needs to be freed on exit from the 3186 routine.* 3187 3188 *Because the .debug_frame section is useful independently of any 3189 ``.debug_info`` section, the augmentation string always uses UTF-8 3190 encoding.* 3191 3192 .. note:: 3193 3194 For AMDGPU, the augmentation string contains: 3195 3196 :: 3197 3198 [amd:v0.0] 3199 3200 The "vX.Y" specifies the major X and minor Y version number of the AMDGPU 3201 extensions used in the DWARF of the compilation unit. The version number 3202 conforms to [SEMVER]_. 3203 32045. ``address_size`` (ubyte) 3205 3206 The size of a target address in this CIE and any FDEs that use it, in bytes. 3207 If a compilation unit exists for this frame, its address size must match the 3208 address size here. 3209 3210 .. note:: 3211 3212 For AMDGPU: 3213 3214 * The address size for the ``Global`` address space defined in 3215 :ref:`amdgpu-dwarf-address-space-mapping-table`. 3216 32176. ``segment_selector_size`` (ubyte) 3218 3219 The size of a segment selector in this CIE and any FDEs that use it, in 3220 bytes. 3221 3222 .. note:: 3223 3224 For AMDGPU: 3225 3226 * Does not use a segment selector so this is 0. 3227 32287. ``code_alignment_factor`` (unsigned LEB128) 3229 3230 A constant that is factored out of all advance location instructions (see 3231 :ref:`amdgpu-dwarf-row-creation-instructions`). The resulting value is 3232 ``(operand * code_alignment_factor)``. 3233 3234 .. note:: 3235 3236 For AMDGPU: 3237 3238 * 4 bytes. 3239 3240 .. TODO:: 3241 3242 Add to :ref:`amdgpu-processor-table` table. 3243 32448. ``data_alignment_factor`` (signed LEB128) 3245 3246 A constant that is factored out of certain offset instructions (see 3247 :ref:`amdgpu-dwarf-cfa-definition-instructions` and 3248 :ref:`amdgpu-dwarf-register-rule-instructions`). The resulting value is 3249 ``(operand * data_alignment_factor)``. 3250 3251 .. note:: 3252 3253 For AMDGPU: 3254 3255 * 4 bytes. 3256 3257 .. TODO:: 3258 3259 Add to :ref:`amdgpu-processor-table` table. 3260 32619. ``return_address_register`` (unsigned LEB128) 3262 3263 An unsigned LEB128 constant that indicates which column in the rule table 3264 represents the return address of the function. Note that this column might 3265 not correspond to an actual machine register. 3266 3267 .. note:: 3268 3269 For AMDGPU: 3270 3271 * ``PC_32`` for 32-bit processes and ``PC_64`` for 3272 64-bit processes defined in :ref:`amdgpu-dwarf-register-mapping`. 3273 327410. ``initial_instructions`` (array of ubyte) 3275 3276 A sequence of rules that are interpreted to create the initial setting of 3277 each column in the table. 3278 3279 The default rule for all columns before interpretation of the initial 3280 instructions is the undefined rule. However, an ABI authoring body or a 3281 compilation system authoring body may specify an alternate default value for 3282 any or all columns. 3283 3284 .. note:: 3285 3286 For AMDGPU: 3287 3288 * Since a subprogram A with fewer registers can be called from subprogram 3289 B that has more allocated, A will not change any of the extra registers 3290 as it cannot access them. Therefore, The default rule for all columns is 3291 ``same value``. 3292 329311. ``padding`` (array of ubyte) 3294 3295 Enough ``DW_CFA_nop`` instructions to make the size of this entry match the 3296 length value above. 3297 3298An FDE contains the following fields, in order: 3299 33001. ``length`` (initial length) 3301 3302 A constant that gives the number of bytes of the header and instruction 3303 stream for this function, not including the length field itself. The size of 3304 the length field plus the value of length must be an integral multiple of 3305 the address size. 3306 33072. ``CIE_pointer`` (4 or 8 bytes, see 3308 :ref:`amdgpu-dwarf-32-bit-and-64-bit-dwarf-formats`) 3309 3310 A constant offset into the ``.debug_frame`` section that denotes the CIE 3311 that is associated with this FDE. 3312 33133. ``initial_location`` (segment selector and target address) 3314 3315 The address of the first location associated with this table entry. If the 3316 segment_selector_size field of this FDE’s CIE is non-zero, the initial 3317 location is preceded by a segment selector of the given length. 3318 33194. ``address_range`` (target address) 3320 3321 The number of bytes of program instructions described by this entry. 3322 33235. ``instructions`` (array of ubyte) 3324 3325 A sequence of table defining instructions that are described in 3326 :ref:`amdgpu-dwarf-call-frame-instructions`. 3327 33286. ``padding`` (array of ubyte) 3329 3330 Enough ``DW_CFA_nop`` instructions to make the size of this entry match the 3331 length value above. 3332 3333.. _amdgpu-dwarf-call-frame-instructions: 3334 3335Call Frame Instructions 3336+++++++++++++++++++++++ 3337 3338Some call frame instructions have operands that are encoded as DWARF expressions 3339E (see :ref:`amdgpu-dwarf-expressions`). The DWARF operators that can be used in 3340E have the following restrictions: 3341 3342* ``DW_OP_addrx``, ``DW_OP_call2``, ``DW_OP_call4``, ``DW_OP_call_ref``, 3343 ``DW_OP_const_type``, ``DW_OP_constx``, ``DW_OP_convert``, 3344 ``DW_OP_deref_type``, ``DW_OP_regval_type``, and ``DW_OP_reinterpret`` 3345 operators are not allowed because the call frame information must not depend 3346 on other debug sections. 3347 3348* ``DW_OP_push_object_address`` is not allowed because there is no object 3349 context to provide a value to push. 3350 3351* ``DW_OP_call_frame_cfa`` and ``DW_OP_entry_value`` are not allowed because 3352 their use would be circular. 3353 3354* ``DW_OP_LLVM_call_frame_entry_reg`` is not allowed if evaluating E causes a 3355 circular dependency between ``DW_OP_LLVM_call_frame_entry_reg`` operators. 3356 3357 *For example, if a register R1 has a* ``DW_CFA_def_cfa_expression`` 3358 *instruction that evaluates a* ``DW_OP_LLVM_call_frame_entry_reg`` *operator 3359 that specifies register R2, and register R2 has a* 3360 ``DW_CFA_def_cfa_expression`` *instruction that that evaluates a* 3361 ``DW_OP_LLVM_call_frame_entry_reg`` *operator that specifies register R1.* 3362 3363*Call frame instructions to which these restrictions apply include* 3364``DW_CFA_def_cfa_expression``\ *,* ``DW_CFA_expression``\ *, and* 3365``DW_CFA_val_expression``\ *.* 3366 3367.. _amdgpu-dwarf-row-creation-instructions: 3368 3369Row Creation Instructions 3370######################### 3371 3372These instructions are the same as in DWARF 5. 3373 3374.. _amdgpu-dwarf-cfa-definition-instructions: 3375 3376CFA Definition Instructions 3377########################### 3378 33791. ``DW_CFA_def_cfa`` 3380 3381 The ``DW_CFA_def_cfa`` instruction takes two unsigned LEB128 operands 3382 representing a register number R and a (non-factored) byte displacement D. 3383 The required action is to define the current CFA rule to be the memory 3384 location description that is the result of evaluating the DWARF expression 3385 ``DW_OP_bregx R, D``. 3386 3387 .. note:: 3388 3389 Could also consider adding ``DW_CFA_def_aspace_cfa`` and 3390 ``DW_CFA_def_aspace_cfa_sf`` which allow a register R, offset D, and 3391 address space AS to be specified. For example, that would save a byte of 3392 encoding over using ``DW_CFA_def_cfa R, D; DW_CFA_LLVM_def_cfa_aspace 3393 AS;``. 3394 33952. ``DW_CFA_def_cfa_sf`` 3396 3397 The ``DW_CFA_def_cfa_sf`` instruction takes two operands: an unsigned LEB128 3398 value representing a register number R and a signed LEB128 factored byte 3399 displacement D. The required action is to define the current CFA rule to be 3400 the memory location description that is the result of evaluating the DWARF 3401 expression ``DW_OP_bregx R, D*data_alignment_factor``. 3402 3403 *The action is the same as ``DW_CFA_def_cfa`` except that the second operand 3404 is signed and factored.* 3405 34063. ``DW_CFA_def_cfa_register`` 3407 3408 The ``DW_CFA_def_cfa_register`` instruction takes a single unsigned LEB128 3409 operand representing a register number R. The required action is to define 3410 the current CFA rule to be the memory location description that is the 3411 result of evaluating the DWARF expression ``DW_OP_constu AS; 3412 DW_OP_aspace_bregx R, D`` where D and AS are the old CFA byte displacement 3413 and address space respectively. 3414 3415 If the subprogram has no current CFA rule, or the rule was defined by a 3416 ``DW_CFA_def_cfa_expression`` instruction, then the DWARF is ill-formed. 3417 34184. ``DW_CFA_def_cfa_offset`` 3419 3420 The ``DW_CFA_def_cfa_offset`` instruction takes a single unsigned LEB128 3421 operand representing a (non-factored) byte displacement D. The required 3422 action is to define the current CFA rule to be the memory location 3423 description that is the result of evaluating the DWARF expression 3424 ``DW_OP_constu AS; DW_OP_aspace_bregx R, D`` where R and AS are the old CFA 3425 register number and address space respectively. 3426 3427 If the subprogram has no current CFA rule, or the rule was defined by a 3428 ``DW_CFA_def_cfa_expression`` instruction, then the DWARF is ill-formed. 3429 34305. ``DW_CFA_def_cfa_offset_sf`` 3431 3432 The ``DW_CFA_def_cfa_offset_sf`` instruction takes a signed LEB128 operand 3433 representing a factored byte displacement D. The required action is to 3434 define the current CFA rule to be the memory location description that is 3435 the result of evaluating the DWARF expression ``DW_OP_constu AS; 3436 DW_OP_aspace_bregx R, D*data_alignment_factor`` where R and AS are the old 3437 CFA register number and address space respectively. 3438 3439 If the subprogram has no current CFA rule, or the rule was defined by a 3440 ``DW_CFA_def_cfa_expression`` instruction, then the DWARF is ill-formed. 3441 3442 *The action is the same as ``DW_CFA_def_cfa_offset`` except that the operand 3443 is signed and factored.* 3444 34456. ``DW_CFA_LLVM_def_cfa_aspace`` *New* 3446 3447 The ``DW_CFA_LLVM_def_cfa_aspace`` instruction takes a single unsigned 3448 LEB128 operand representing an address space identifier AS for those 3449 architectures that support multiple address spaces. The required action is 3450 to define the current CFA rule to be the memory location description L that 3451 is the result of evaluating the DWARF expression ``DW_OP_constu AS; 3452 DW_OP_aspace_bregx R, D`` where R and D are the old CFA register number and 3453 byte displacement respectively. 3454 3455 If AS is not one of the values defined by the target architecture's 3456 ``DW_ASPACE_*`` values then the DWARF expression is ill-formed. 3457 34587. ``DW_CFA_def_cfa_expression`` 3459 3460 The ``DW_CFA_def_cfa_expression`` instruction takes a single operand encoded 3461 as a ``DW_FORM_exprloc`` value representing a DWARF expression E. The 3462 required action is to define the current CFA rule to be the memory location 3463 description computed by evaluating E. 3464 3465 *See :ref:`amdgpu-dwarf-call-frame-instructions` regarding restrictions on 3466 the DWARF expression operators that can be used in E.* 3467 3468 If the result of evaluating E is not a memory location description with bit 3469 offset that is a multiple of 8 (the byte size), then the DWARF is 3470 ill-formed. 3471 3472.. _amdgpu-dwarf-register-rule-instructions: 3473 3474Register Rule Instructions 3475########################## 3476 3477.. note:: 3478 3479 For AMDGPU: 3480 3481 * The register number follows the numbering defined in 3482 :ref:`amdgpu-dwarf-register-mapping`. 3483 34841. ``DW_CFA_undefined`` 3485 3486 The ``DW_CFA_undefined`` instruction takes a single unsigned LEB128 operand 3487 that represents a register number R. The required action is to set the rule 3488 for the register specified by R to ``undefined``. 3489 34902. ``DW_CFA_same_value`` 3491 3492 The ``DW_CFA_same_value`` instruction takes a single unsigned LEB128 operand 3493 that represents a register number R. The required action is to set the rule 3494 for the register specified by R to ``same value``. 3495 34963. ``DW_CFA_offset`` 3497 3498 The ``DW_CFA_offset`` instruction takes two operands: a register number R 3499 (encoded with the opcode) and an unsigned LEB128 constant representing a 3500 factored displacement D. The required action is to change the rule for the 3501 register specified by R to be an *offset(D*data_alignment_factor)* rule. 3502 3503 .. note:: 3504 3505 Seems this should be named ``DW_CFA_offset_uf`` since the offset is 3506 unsigned factored. 3507 35084. ``DW_CFA_offset_extended`` 3509 3510 The ``DW_CFA_offset_extended`` instruction takes two unsigned LEB128 3511 operands representing a register number R and a factored displacement D. 3512 This instruction is identical to ``DW_CFA_offset`` except for the encoding 3513 and size of the register operand. 3514 3515 .. note:: 3516 3517 Seems this should be named ``DW_CFA_offset_extended_uf`` since the 3518 displacement is unsigned factored. 3519 35205. ``DW_CFA_offset_extended_sf`` 3521 3522 The ``DW_CFA_offset_extended_sf`` instruction takes two operands: an 3523 unsigned LEB128 value representing a register number R and a signed LEB128 3524 factored displacement D. This instruction is identical to 3525 ``DW_CFA_offset_extended`` except that D is signed. 3526 35276. ``DW_CFA_val_offset`` 3528 3529 The ``DW_CFA_val_offset`` instruction takes two unsigned LEB128 operands 3530 representing a register number R and a factored displacement D. The required 3531 action is to change the rule for the register indicated by R to be a 3532 *val_offset(D*data_alignment_factor)* rule. 3533 3534 .. note:: 3535 3536 Seems this should be named ``DW_CFA_val_offset_uf`` since the displacement 3537 is unsigned factored. 3538 35397. ``DW_CFA_val_offset_sf`` 3540 3541 The ``DW_CFA_val_offset_sf`` instruction takes two operands: an unsigned 3542 LEB128 value representing a register number R and a signed LEB128 factored 3543 displacement D. This instruction is identical to ``DW_CFA_val_offset`` 3544 except that D is signed. 3545 35468. ``DW_CFA_register`` 3547 3548 The ``DW_CFA_register`` instruction takes two unsigned LEB128 operands 3549 representing register numbers R1 and R2 respectively. The required action is 3550 to set the rule for the register specified by R1 to be *register(R)* where R 3551 is R2. 3552 35539. ``DW_CFA_expression`` 3554 3555 The ``DW_CFA_expression`` instruction takes two operands: an unsigned LEB128 3556 value representing a register number R, and a ``DW_FORM_block`` value 3557 representing a DWARF expression E. The required action is to change the rule 3558 for the register specified by R to be an *expression(E)* rule. The memory 3559 location description of the current CFA is pushed on the DWARF stack prior 3560 to execution of E. 3561 3562 *That is, the DWARF expression computes the location description where the 3563 register value can be retrieved.* 3564 3565 *See :ref:`amdgpu-dwarf-call-frame-instructions` regarding restrictions on 3566 the DWARF expression operators that can be used in E.* 3567 356810. ``DW_CFA_val_expression`` 3569 3570 The ``DW_CFA_val_expression`` instruction takes two operands: an unsigned 3571 LEB128 value representing a register number R, and a ``DW_FORM_block`` value 3572 representing a DWARF expression E. The required action is to change the rule 3573 for the register specified by R to be a *val_expression(E)* rule. The memory 3574 location description of the current CFA is pushed on the DWARF evaluation 3575 stack prior to execution of E. 3576 3577 *That is, E computes the value of register R.* 3578 3579 *See :ref:`amdgpu-dwarf-call-frame-instructions` regarding restrictions on 3580 the DWARF expression operators that can be used in E.* 3581 3582 If the result of evaluating E is not a value with a base type size that 3583 matches the register size, then the DWARF is ill-formed. 3584 358511. ``DW_CFA_restore`` 3586 3587 The ``DW_CFA_restore`` instruction takes a single operand (encoded with the 3588 opcode) that represents a register number R. The required action is to 3589 change the rule for the register specified by R to the rule assigned it by 3590 the initial_instructions in the CIE. 3591 359212. ``DW_CFA_restore_extended`` 3593 3594 The ``DW_CFA_restore_extended`` instruction takes a single unsigned LEB128 3595 operand that represents a register number R. This instruction is identical 3596 to ``DW_CFA_restore`` except for the encoding and size of the register 3597 operand. 3598 3599Row State Instructions 3600###################### 3601 3602These instructions are the same as in DWARF 5. 3603 3604Call Frame Calling Address 3605++++++++++++++++++++++++++ 3606 3607*When virtually unwinding frames, consumers frequently wish to obtain the 3608address of the instruction which called a subroutine. This information is not 3609always provided. Typically, however, one of the registers in the virtual unwind 3610table is the Return Address.* 3611 3612If a Return Address register is defined in the virtual unwind table, and its 3613rule is undefined (for example, by ``DW_CFA_undefined``), then there is no 3614return address and no call address, and the virtual unwind of stack activations 3615is complete. 3616 3617*In most cases the return address is in the same context as the calling address, 3618but that need not be the case, especially if the producer knows in some way the 3619call never will return. The context of the ’return address’ might be on a 3620different line, in a different lexical block, or past the end of the calling 3621subroutine. If a consumer were to assume that it was in the same context as the 3622calling address, the virtual unwind might fail.* 3623 3624*For architectures with constant-length instructions where the return address 3625immediately follows the call instruction, a simple solution is to subtract the 3626length of an instruction from the return address to obtain the calling 3627instruction. For architectures with variable-length instructions (for example, 3628x86), this is not possible. However, subtracting 1 from the return address, 3629although not guaranteed to provide the exact calling address, generally will 3630produce an address within the same context as the calling address, and that 3631usually is sufficient.* 3632 3633.. note:: 3634 3635 For AMDGPU the instructions are variable size and a consumer can subtract 1 3636 from the return address to get the address of a byte within the call site 3637 instructions. 3638 3639Call Frame Information Instruction Encodings 3640++++++++++++++++++++++++++++++++++++++++++++ 3641 3642The following table gives the encoding of the DWARF call frame information 3643instructions added for AMDGPU. 3644 3645.. table:: AMDGPU DWARF Call Frame Information Instruction Encodings 3646 :name: amdgpu-dwarf-call-frame-information-instruction-encodings-table 3647 3648 =================================== ==== ==== ============== ================ 3649 Instruction High Low Operand 1 Operand 1 3650 2 6 3651 Bits Bits 3652 =================================== ==== ==== ============== ================ 3653 DW_CFA_LLVM_def_cfa_aspace 0 0Xxx ULEB128 3654 =================================== ==== ==== ============== ================ 3655 3656Line Table 3657~~~~~~~~~~ 3658 3659.. note:: 3660 3661 AMDGPU does not use the ``isa`` state machine registers and always sets it to 3662 0. 3663 3664.. TODO:: 3665 3666 Should the ``isa`` state machine register be used to indicate if the code is 3667 in wave32 or wave64 mode? Or used to specify the architecture ISA? 3668 3669Accelerated Access 3670~~~~~~~~~~~~~~~~~~ 3671 3672Lookup By Name 3673++++++++++++++ 3674 3675.. note:: 3676 3677 For AMDGPU: 3678 3679 * The rule for debugger information entries included in the name 3680 index in the optional ``.debug_names`` section is extended to also include 3681 named ``DW_TAG_variable`` debugging information entries with a 3682 ``DW_AT_location`` attribute that includes a 3683 ``DW_OP_LLVM_form_aspace_address`` operation. 3684 3685 * The lookup by name section header ``augmentation_string`` string field contains: 3686 3687 :: 3688 3689 [amd:v0.0] 3690 3691 The "vX.Y" specifies the major X and minor Y version number of the AMDGPU 3692 extensions used in the DWARF of the compilation unit. The version number 3693 conforms to [SEMVER]_. 3694 3695Lookup By Address 3696+++++++++++++++++ 3697 3698.. note:: 3699 3700 For AMDGPU: 3701 3702 * The lookup by address section header table: 3703 3704 ``address_size`` (ubyte) 3705 Match the address size for the ``Global`` address space defined in 3706 :ref:`amdgpu-dwarf-address-space-mapping-table`. 3707 3708 ``segment_selector_size`` (ubyte) 3709 AMDGPU does not use a segment selector so this is 0. The entries in the 3710 ``.debug_aranges`` do not have a segment selector. 3711 3712Data Representation 3713~~~~~~~~~~~~~~~~~~~ 3714 3715.. _amdgpu-dwarf-32-bit-and-64-bit-dwarf-formats: 3716 371732-Bit and 64-Bit DWARF Formats 3718+++++++++++++++++++++++++++++++ 3719 3720.. note:: 3721 3722 For AMDGPU: 3723 3724 * For the ``amdgcn`` target only 64-bit process address space is supported 3725 * The producer can generate either 32-bit or 64-bit DWARF format. 3726 37271. Within the body of the ``.debug_info`` section, certain forms of attribute 3728 value depend on the choice of DWARF format as follows. For the 32-bit DWARF 3729 format, the value is a 4-byte unsigned integer; for the 64-bit DWARF format, 3730 the value is an 8-byte unsigned integer. 3731 3732 .. table:: AMDGPU DWARF ``.debug_info`` section attribute sizes 3733 :name: amdgpu-dwarf-debug-info-section-attribute-sizes 3734 3735 =================================== ===================================== 3736 Form Role 3737 =================================== ===================================== 3738 DW_FORM_line_strp offset in ``.debug_line_str`` 3739 DW_FORM_ref_addr offset in ``.debug_info`` 3740 DW_FORM_sec_offset offset in a section other than 3741 ``.debug_info`` or ``.debug_str`` 3742 DW_FORM_strp offset in ``.debug_str`` 3743 DW_FORM_strp_sup offset in ``.debug_str`` section of 3744 supplementary object file 3745 DW_OP_call_ref offset in ``.debug_info`` 3746 DW_OP_implicit_pointer offset in ``.debug_info`` 3747 DW_OP_LLVM_aspace_implicit_pointer offset in ``.debug_info`` 3748 =================================== ===================================== 3749 3750Unit Headers 3751++++++++++++ 3752 3753.. note:: 3754 3755 For AMDGPU: 3756 3757 * For AMDGPU the ``address_size`` field of the DWARF unit headers matches the 3758 address size for the ``Global`` address space defined in 3759 :ref:`amdgpu-dwarf-address-space-mapping-table`. 3760 3761.. _amdgpu-dwarf-amdgpu-dw-at-llvm-lane-pc: 3762 3763AMDGPU DW_AT_LLVM_lane_pc 3764~~~~~~~~~~~~~~~~~~~~~~~~~ 3765 3766The ``DW_AT_LLVM_lane_pc`` attribute can be used to specify the program location 3767of the separate lanes of a SIMT thread. See 3768:ref:`amdgpu-dwarf-debugging-information-entry-attributes`. 3769 3770If the lane is an active lane then this will be the same as the current program 3771location. 3772 3773If the lane is inactive, but was active on entry to the subprogram, then this is 3774the program location in the subprogram at which execution of the lane is 3775conceptual positioned. 3776 3777If the lane was not active on entry to the subprogram, then this will be the 3778undefined location. A client debugger can check if the lane is part of a valid 3779work-group by checking that the lane is in the range of the associated 3780work-group within the grid, accounting for partial work-groups. If it is not 3781then the debugger can omit any information for the lane. Otherwise, the debugger 3782may repeatedly unwind the stack and inspect the ``DW_AT_LLVM_lane_pc`` of the 3783calling subprogram until it finds a non-undefined location. Conceptually the 3784lane only has the call frames that it has a non-undefined 3785``DW_AT_LLVM_lane_pc``. 3786 3787The following example illustrates how the AMDGPU backend can generate a location 3788list for the nested ``IF/THEN/ELSE`` structures of the following subprogram 3789pseudo code for a target with 64 lanes per wave. 3790 3791.. code:: 3792 :number-lines: 3793 3794 SUBPROGRAM X 3795 BEGIN 3796 a; 3797 IF (c1) THEN 3798 b; 3799 IF (c2) THEN 3800 c; 3801 ELSE 3802 d; 3803 ENDIF 3804 e; 3805 ELSE 3806 f; 3807 ENDIF 3808 g; 3809 END 3810 3811The AMDGPU backend may generate the following pseudo LLVM MIR to manipulate the 3812execution mask (``EXEC``) to linearized the control flow. The condition is 3813evaluated to make a mask of the lanes for which the condition evaluates to true. 3814First the ``THEN`` region is executed by setting the ``EXEC`` mask to the 3815logical ``AND`` of the current ``EXEC`` mask with the condition mask. Then the 3816``ELSE`` region is executed by negating the ``EXEC`` mask and logical ``AND`` of 3817the saved ``EXEC`` mask at the start of the region. After the ``IF/THEN/ELSE`` 3818region the ``EXEC`` mask is restored to the value it had at the beginning of the 3819region. This is shown below. Other approaches are possible, but the basic 3820concept is the same. 3821 3822.. code:: 3823 :number-lines: 3824 3825 $lex_start: 3826 a; 3827 %1 = EXEC 3828 %2 = c1 3829 $lex_1_start: 3830 EXEC = %1 & %2 3831 $if_1_then: 3832 b; 3833 %3 = EXEC 3834 %4 = c2 3835 $lex_1_1_start: 3836 EXEC = %3 & %4 3837 $lex_1_1_then: 3838 c; 3839 EXEC = ~EXEC & %3 3840 $lex_1_1_else: 3841 d; 3842 EXEC = %3 3843 $lex_1_1_end: 3844 e; 3845 EXEC = ~EXEC & %1 3846 $lex_1_else: 3847 f; 3848 EXEC = %1 3849 $lex_1_end: 3850 g; 3851 $lex_end: 3852 3853To create the location list that defines the location description of a vector of 3854lane program locations, the LLVM MIR ``DBG_VALUE`` pseudo instruction can be 3855used to annotate the linearized control flow. This can be done by defining an 3856artificial variable for the lane PC. The location list created for it is used to 3857define the value of the ``DW_AT_LLVM_lane_pc`` attribute. 3858 3859A DWARF procedure is defined for each well nested structured control flow region 3860which provides the conceptual lane program location for a lane if it is not 3861active (namely it is divergent). The expression for each region inherits the 3862value of the immediately enclosing region and modifies it according to the 3863semantics of the region. 3864 3865For an ``IF/THEN/ELSE`` region the divergent program location is at the start of 3866the region for the ``THEN`` region since it is executed first. For the ``ELSE`` 3867region the divergent program location is at the end of the ``IF/THEN/ELSE`` 3868region since the ``THEN`` region has completed. 3869 3870The lane PC artificial variable is assigned at each region transition. It uses 3871the immediately enclosing region's DWARF procedure to compute the program 3872location for each lane assuming they are divergent, and then modifies the result 3873by inserting the current program location for each lane that the ``EXEC`` mask 3874indicates is active. 3875 3876By having separate DWARF procedures for each region, they can be reused to 3877define the value for any nested region. This reduces the amount of DWARF 3878required. 3879 3880The following provides an example using pseudo LLVM MIR. 3881 3882.. code:: 3883 :number-lines: 3884 3885 $lex_start: 3886 DEFINE_DWARF %__uint_64 = DW_TAG_base_type[ 3887 DW_AT_name = "__uint64"; 3888 DW_AT_byte_size = 8; 3889 DW_AT_encoding = DW_ATE_unsigned; 3890 ]; 3891 DEFINE_DWARF %__active_lane_pc = DW_TAG_dwarf_procedure[ 3892 DW_AT_name = "__active_lane_pc"; 3893 DW_AT_location = [ 3894 DW_OP_regx PC; 3895 DW_OP_LLVM_extend 64, 64; 3896 DW_OP_regval_type EXEC, %uint_64; 3897 DW_OP_LLVM_select_bit_piece 64, 64; 3898 ]; 3899 ]; 3900 DEFINE_DWARF %__divergent_lane_pc = DW_TAG_dwarf_procedure[ 3901 DW_AT_name = "__divergent_lane_pc"; 3902 DW_AT_location = [ 3903 DW_OP_LLVM_undefined; 3904 DW_OP_LLVM_extend 64, 64; 3905 ]; 3906 ]; 3907 DBG_VALUE $noreg, $noreg, %DW_AT_LLVM_lane_pc, DIExpression[ 3908 DW_OP_call_ref %__divergent_lane_pc; 3909 DW_OP_call_ref %__active_lane_pc; 3910 ]; 3911 a; 3912 %1 = EXEC; 3913 DBG_VALUE %1, $noreg, %__lex_1_save_exec; 3914 %2 = c1; 3915 $lex_1_start: 3916 EXEC = %1 & %2; 3917 $lex_1_then: 3918 DEFINE_DWARF %__divergent_lane_pc_1_then = DW_TAG_dwarf_procedure[ 3919 DW_AT_name = "__divergent_lane_pc_1_then"; 3920 DW_AT_location = DIExpression[ 3921 DW_OP_call_ref %__divergent_lane_pc; 3922 DW_OP_xaddr &lex_1_start; 3923 DW_OP_stack_value; 3924 DW_OP_LLVM_extend 64, 64; 3925 DW_OP_call_ref %__lex_1_save_exec; 3926 DW_OP_deref_type 64, %__uint_64; 3927 DW_OP_LLVM_select_bit_piece 64, 64; 3928 ]; 3929 ]; 3930 DBG_VALUE $noreg, $noreg, %DW_AT_LLVM_lane_pc, DIExpression[ 3931 DW_OP_call_ref %__divergent_lane_pc_1_then; 3932 DW_OP_call_ref %__active_lane_pc; 3933 ]; 3934 b; 3935 %3 = EXEC; 3936 DBG_VALUE %3, %__lex_1_1_save_exec; 3937 %4 = c2; 3938 $lex_1_1_start: 3939 EXEC = %3 & %4; 3940 $lex_1_1_then: 3941 DEFINE_DWARF %__divergent_lane_pc_1_1_then = DW_TAG_dwarf_procedure[ 3942 DW_AT_name = "__divergent_lane_pc_1_1_then"; 3943 DW_AT_location = DIExpression[ 3944 DW_OP_call_ref %__divergent_lane_pc_1_then; 3945 DW_OP_xaddr &lex_1_1_start; 3946 DW_OP_stack_value; 3947 DW_OP_LLVM_extend 64, 64; 3948 DW_OP_call_ref %__lex_1_1_save_exec; 3949 DW_OP_deref_type 64, %__uint_64; 3950 DW_OP_LLVM_select_bit_piece 64, 64; 3951 ]; 3952 ]; 3953 DBG_VALUE $noreg, $noreg, %DW_AT_LLVM_lane_pc, DIExpression[ 3954 DW_OP_call_ref %__divergent_lane_pc_1_1_then; 3955 DW_OP_call_ref %__active_lane_pc; 3956 ]; 3957 c; 3958 EXEC = ~EXEC & %3; 3959 $lex_1_1_else: 3960 DEFINE_DWARF %__divergent_lane_pc_1_1_else = DW_TAG_dwarf_procedure[ 3961 DW_AT_name = "__divergent_lane_pc_1_1_else"; 3962 DW_AT_location = DIExpression[ 3963 DW_OP_call_ref %__divergent_lane_pc_1_then; 3964 DW_OP_xaddr &lex_1_1_end; 3965 DW_OP_stack_value; 3966 DW_OP_LLVM_extend 64, 64; 3967 DW_OP_call_ref %__lex_1_1_save_exec; 3968 DW_OP_deref_type 64, %__uint_64; 3969 DW_OP_LLVM_select_bit_piece 64, 64; 3970 ]; 3971 ]; 3972 DBG_VALUE $noreg, $noreg, %DW_AT_LLVM_lane_pc, DIExpression[ 3973 DW_OP_call_ref %__divergent_lane_pc_1_1_else; 3974 DW_OP_call_ref %__active_lane_pc; 3975 ]; 3976 d; 3977 EXEC = %3; 3978 $lex_1_1_end: 3979 DBG_VALUE $noreg, $noreg, %DW_AT_LLVM_lane_pc, DIExpression[ 3980 DW_OP_call_ref %__divergent_lane_pc; 3981 DW_OP_call_ref %__active_lane_pc; 3982 ]; 3983 e; 3984 EXEC = ~EXEC & %1; 3985 $lex_1_else: 3986 DEFINE_DWARF %__divergent_lane_pc_1_else = DW_TAG_dwarf_procedure[ 3987 DW_AT_name = "__divergent_lane_pc_1_else"; 3988 DW_AT_location = DIExpression[ 3989 DW_OP_call_ref %__divergent_lane_pc; 3990 DW_OP_xaddr &lex_1_end; 3991 DW_OP_stack_value; 3992 DW_OP_LLVM_extend 64, 64; 3993 DW_OP_call_ref %__lex_1_save_exec; 3994 DW_OP_deref_type 64, %__uint_64; 3995 DW_OP_LLVM_select_bit_piece 64, 64; 3996 ]; 3997 ]; 3998 DBG_VALUE $noreg, $noreg, %DW_AT_LLVM_lane_pc, DIExpression[ 3999 DW_OP_call_ref %__divergent_lane_pc_1_else; 4000 DW_OP_call_ref %__active_lane_pc; 4001 ]; 4002 f; 4003 EXEC = %1; 4004 $lex_1_end: 4005 DBG_VALUE $noreg, $noreg, %DW_AT_LLVM_lane_pc DIExpression[ 4006 DW_OP_call_ref %__divergent_lane_pc; 4007 DW_OP_call_ref %__active_lane_pc; 4008 ]; 4009 g; 4010 $lex_end: 4011 4012The DWARF procedure ``%__active_lane_pc`` is used to update the lane pc elements 4013that are active with the current program location. 4014 4015Artificial variables %__lex_1_save_exec and %__lex_1_1_save_exec are created for 4016the execution masks saved on entry to a region. Using the ``DBG_VALUE`` pseudo 4017instruction, location lists that describes where they are allocated at any given 4018program location will be created. The compiler may allocate them to registers, 4019or spill them to memory. 4020 4021The DWARF procedures for each region use saved execution mask value to only 4022update the lanes that are active on entry to the region. All other lanes retain 4023the value of the enclosing region where they were last active. If they were not 4024active on entry to the subprogram, then will have the undefined location 4025description. 4026 4027Other structured control flow regions can be handled similarly. For example, 4028loops would set the divergent program location for the region at the end of the 4029loop. Any lanes active will be in the loop, and any lanes not active must have 4030exited the loop. 4031 4032An ``IF/THEN/ELSEIF/ELSEIF/...`` region can be treated as a nest of 4033``IF/THEN/ELSE`` regions. 4034 4035The DWARF procedures can use the active lane artificial variable described in 4036:ref:`amdgpu-dwarf-amdgpu-dw-at-llvm-active-lane` rather than the actual 4037``EXEC`` mask in order to support whole or quad wave mode. 4038 4039.. _amdgpu-dwarf-amdgpu-dw-at-llvm-active-lane: 4040 4041AMDGPU DW_AT_LLVM_active_lane 4042~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 4043 4044The ``DW_AT_LLVM_active_lane`` attribute can be used to specify the lanes that 4045are conceptually active for a SIMT thread. See 4046:ref:`amdgpu-dwarf-debugging-information-entry-attributes`. 4047 4048The execution mask may be modified to implement whole or quad wave mode 4049operations. For example, all lanes may need to temporarily be made active to 4050execute a whole wave operation. Such regions would save the ``EXEC`` mask, 4051update it to enable the necessary lanes, perform the operations, and then 4052restore the ``EXEC`` mask from the saved value. While executing the whole wave 4053region, the conceptual execution mask is the saved value, not the ``EXEC`` 4054value. 4055 4056This is handled by defining an artificial variable for the active lane mask. The 4057active lane mask artificial variable would be the actual ``EXEC`` mask for 4058normal regions, and the saved execution mask for regions where the mask is 4059temporarily updated. The location list created for this artificial variable is 4060used to define the value of the ``DW_AT_LLVM_active_lane`` attribute. 4061 4062Source Text 4063~~~~~~~~~~~ 4064 4065Source text for online-compiled programs (e.g. those compiled by the OpenCL 4066runtime) may be embedded into the DWARF v5 line table using the ``clang 4067-gembed-source`` option, described in table :ref:`amdgpu-debug-options`. 4068 4069For example: 4070 4071``-gembed-source`` 4072 Enable the embedded source DWARF v5 extension. 4073``-gno-embed-source`` 4074 Disable the embedded source DWARF v5 extension. 4075 4076 .. table:: AMDGPU Debug Options 4077 :name: amdgpu-debug-options 4078 4079 ==================== ================================================== 4080 Debug Flag Description 4081 ==================== ================================================== 4082 -g[no-]embed-source Enable/disable embedding source text in DWARF 4083 debug sections. Useful for environments where 4084 source cannot be written to disk, such as 4085 when performing online compilation. 4086 ==================== ================================================== 4087 4088This option enables one extended content types in the DWARF v5 Line Number 4089Program Header, which is used to encode embedded source. 4090 4091 .. table:: AMDGPU DWARF Line Number Program Header Extended Content Types 4092 :name: amdgpu-dwarf-extended-content-types 4093 4094 ============================ ====================== 4095 Content Type Form 4096 ============================ ====================== 4097 ``DW_LNCT_LLVM_source`` ``DW_FORM_line_strp`` 4098 ============================ ====================== 4099 4100The source field will contain the UTF-8 encoded, null-terminated source text 4101with ``'\n'`` line endings. When the source field is present, consumers can use 4102the embedded source instead of attempting to discover the source on disk. When 4103the source field is absent, consumers can access the file to get the source 4104text. 4105 4106The above content type appears in the ``file_name_entry_format`` field of the 4107line table prologue, and its corresponding value appear in the ``file_names`` 4108field. The current encoding of the content type is documented in table 4109:ref:`amdgpu-dwarf-extended-content-types-encoding` 4110 4111 .. table:: AMDGPU DWARF Line Number Program Header Extended Content Types Encoding 4112 :name: amdgpu-dwarf-extended-content-types-encoding 4113 4114 ============================ ==================== 4115 Content Type Value 4116 ============================ ==================== 4117 ``DW_LNCT_LLVM_source`` 0x2001 4118 ============================ ==================== 4119 4120.. _amdgpu-code-conventions: 4121 4122Code Conventions 4123================ 4124 4125This section provides code conventions used for each supported target triple OS 4126(see :ref:`amdgpu-target-triples`). 4127 4128AMDHSA 4129------ 4130 4131This section provides code conventions used when the target triple OS is 4132``amdhsa`` (see :ref:`amdgpu-target-triples`). 4133 4134.. _amdgpu-amdhsa-code-object-target-identification: 4135 4136Code Object Target Identification 4137~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 4138 4139The AMDHSA OS uses the following syntax to specify the code object 4140target as a single string: 4141 4142 ``<Architecture>-<Vendor>-<OS>-<Environment>-<Processor><Target Features>`` 4143 4144Where: 4145 4146 - ``<Architecture>``, ``<Vendor>``, ``<OS>`` and ``<Environment>`` 4147 are the same as the *Target Triple* (see 4148 :ref:`amdgpu-target-triples`). 4149 4150 - ``<Processor>`` is the same as the *Processor* (see 4151 :ref:`amdgpu-processors`). 4152 4153 - ``<Target Features>`` is a list of the enabled *Target Features* 4154 (see :ref:`amdgpu-target-features`), each prefixed by a plus, that 4155 apply to *Processor*. The list must be in the same order as listed 4156 in the table :ref:`amdgpu-target-feature-table`. Note that *Target 4157 Features* must be included in the list if they are enabled even if 4158 that is the default for *Processor*. 4159 4160For example: 4161 4162 ``"amdgcn-amd-amdhsa--gfx902+xnack"`` 4163 4164.. _amdgpu-amdhsa-code-object-metadata: 4165 4166Code Object Metadata 4167~~~~~~~~~~~~~~~~~~~~ 4168 4169The code object metadata specifies extensible metadata associated with the code 4170objects executed on HSA [HSA]_ compatible runtimes such as AMD's ROCm 4171[AMD-ROCm]_. The encoding and semantics of this metadata depends on the code 4172object version; see :ref:`amdgpu-amdhsa-code-object-metadata-v2` and 4173:ref:`amdgpu-amdhsa-code-object-metadata-v3`. 4174 4175Code object metadata is specified in a note record (see 4176:ref:`amdgpu-note-records`) and is required when the target triple OS is 4177``amdhsa`` (see :ref:`amdgpu-target-triples`). It must contain the minimum 4178information necessary to support the ROCM kernel queries. For example, the 4179segment sizes needed in a dispatch packet. In addition, a high level language 4180runtime may require other information to be included. For example, the AMD 4181OpenCL runtime records kernel argument information. 4182 4183.. _amdgpu-amdhsa-code-object-metadata-v2: 4184 4185Code Object V2 Metadata (-mattr=-code-object-v3) 4186++++++++++++++++++++++++++++++++++++++++++++++++ 4187 4188.. warning:: Code Object V2 is not the default code object version emitted by 4189 this version of LLVM. For a description of the metadata generated with the 4190 default configuration (Code Object V3) see 4191 :ref:`amdgpu-amdhsa-code-object-metadata-v3`. 4192 4193Code object V2 metadata is specified by the ``NT_AMD_AMDGPU_METADATA`` note 4194record (see :ref:`amdgpu-note-records-v2`). 4195 4196The metadata is specified as a YAML formatted string (see [YAML]_ and 4197:doc:`YamlIO`). 4198 4199.. TODO:: 4200 4201 Is the string null terminated? It probably should not if YAML allows it to 4202 contain null characters, otherwise it should be. 4203 4204The metadata is represented as a single YAML document comprised of the mapping 4205defined in table :ref:`amdgpu-amdhsa-code-object-metadata-map-table-v2` and 4206referenced tables. 4207 4208For boolean values, the string values of ``false`` and ``true`` are used for 4209false and true respectively. 4210 4211Additional information can be added to the mappings. To avoid conflicts, any 4212non-AMD key names should be prefixed by "*vendor-name*.". 4213 4214 .. table:: AMDHSA Code Object V2 Metadata Map 4215 :name: amdgpu-amdhsa-code-object-metadata-map-table-v2 4216 4217 ========== ============== ========= ======================================= 4218 String Key Value Type Required? Description 4219 ========== ============== ========= ======================================= 4220 "Version" sequence of Required - The first integer is the major 4221 2 integers version. Currently 1. 4222 - The second integer is the minor 4223 version. Currently 0. 4224 "Printf" sequence of Each string is encoded information 4225 strings about a printf function call. The 4226 encoded information is organized as 4227 fields separated by colon (':'): 4228 4229 ``ID:N:S[0]:S[1]:...:S[N-1]:FormatString`` 4230 4231 where: 4232 4233 ``ID`` 4234 A 32-bit integer as a unique id for 4235 each printf function call 4236 4237 ``N`` 4238 A 32-bit integer equal to the number 4239 of arguments of printf function call 4240 minus 1 4241 4242 ``S[i]`` (where i = 0, 1, ... , N-1) 4243 32-bit integers for the size in bytes 4244 of the i-th FormatString argument of 4245 the printf function call 4246 4247 FormatString 4248 The format string passed to the 4249 printf function call. 4250 "Kernels" sequence of Required Sequence of the mappings for each 4251 mapping kernel in the code object. See 4252 :ref:`amdgpu-amdhsa-code-object-kernel-metadata-map-table-v2` 4253 for the definition of the mapping. 4254 ========== ============== ========= ======================================= 4255 4256.. 4257 4258 .. table:: AMDHSA Code Object V2 Kernel Metadata Map 4259 :name: amdgpu-amdhsa-code-object-kernel-metadata-map-table-v2 4260 4261 ================= ============== ========= ================================ 4262 String Key Value Type Required? Description 4263 ================= ============== ========= ================================ 4264 "Name" string Required Source name of the kernel. 4265 "SymbolName" string Required Name of the kernel 4266 descriptor ELF symbol. 4267 "Language" string Source language of the kernel. 4268 Values include: 4269 4270 - "OpenCL C" 4271 - "OpenCL C++" 4272 - "HCC" 4273 - "OpenMP" 4274 4275 "LanguageVersion" sequence of - The first integer is the major 4276 2 integers version. 4277 - The second integer is the 4278 minor version. 4279 "Attrs" mapping Mapping of kernel attributes. 4280 See 4281 :ref:`amdgpu-amdhsa-code-object-kernel-attribute-metadata-map-table-v2` 4282 for the mapping definition. 4283 "Args" sequence of Sequence of mappings of the 4284 mapping kernel arguments. See 4285 :ref:`amdgpu-amdhsa-code-object-kernel-argument-metadata-map-table-v2` 4286 for the definition of the mapping. 4287 "CodeProps" mapping Mapping of properties related to 4288 the kernel code. See 4289 :ref:`amdgpu-amdhsa-code-object-kernel-code-properties-metadata-map-table-v2` 4290 for the mapping definition. 4291 ================= ============== ========= ================================ 4292 4293.. 4294 4295 .. table:: AMDHSA Code Object V2 Kernel Attribute Metadata Map 4296 :name: amdgpu-amdhsa-code-object-kernel-attribute-metadata-map-table-v2 4297 4298 =================== ============== ========= ============================== 4299 String Key Value Type Required? Description 4300 =================== ============== ========= ============================== 4301 "ReqdWorkGroupSize" sequence of If not 0, 0, 0 then all values 4302 3 integers must be >=1 and the dispatch 4303 work-group size X, Y, Z must 4304 correspond to the specified 4305 values. Defaults to 0, 0, 0. 4306 4307 Corresponds to the OpenCL 4308 ``reqd_work_group_size`` 4309 attribute. 4310 "WorkGroupSizeHint" sequence of The dispatch work-group size 4311 3 integers X, Y, Z is likely to be the 4312 specified values. 4313 4314 Corresponds to the OpenCL 4315 ``work_group_size_hint`` 4316 attribute. 4317 "VecTypeHint" string The name of a scalar or vector 4318 type. 4319 4320 Corresponds to the OpenCL 4321 ``vec_type_hint`` attribute. 4322 4323 "RuntimeHandle" string The external symbol name 4324 associated with a kernel. 4325 OpenCL runtime allocates a 4326 global buffer for the symbol 4327 and saves the kernel's address 4328 to it, which is used for 4329 device side enqueueing. Only 4330 available for device side 4331 enqueued kernels. 4332 =================== ============== ========= ============================== 4333 4334.. 4335 4336 .. table:: AMDHSA Code Object V2 Kernel Argument Metadata Map 4337 :name: amdgpu-amdhsa-code-object-kernel-argument-metadata-map-table-v2 4338 4339 ================= ============== ========= ================================ 4340 String Key Value Type Required? Description 4341 ================= ============== ========= ================================ 4342 "Name" string Kernel argument name. 4343 "TypeName" string Kernel argument type name. 4344 "Size" integer Required Kernel argument size in bytes. 4345 "Align" integer Required Kernel argument alignment in 4346 bytes. Must be a power of two. 4347 "ValueKind" string Required Kernel argument kind that 4348 specifies how to set up the 4349 corresponding argument. 4350 Values include: 4351 4352 "ByValue" 4353 The argument is copied 4354 directly into the kernarg. 4355 4356 "GlobalBuffer" 4357 A global address space pointer 4358 to the buffer data is passed 4359 in the kernarg. 4360 4361 "DynamicSharedPointer" 4362 A group address space pointer 4363 to dynamically allocated LDS 4364 is passed in the kernarg. 4365 4366 "Sampler" 4367 A global address space 4368 pointer to a S# is passed in 4369 the kernarg. 4370 4371 "Image" 4372 A global address space 4373 pointer to a T# is passed in 4374 the kernarg. 4375 4376 "Pipe" 4377 A global address space pointer 4378 to an OpenCL pipe is passed in 4379 the kernarg. 4380 4381 "Queue" 4382 A global address space pointer 4383 to an OpenCL device enqueue 4384 queue is passed in the 4385 kernarg. 4386 4387 "HiddenGlobalOffsetX" 4388 The OpenCL grid dispatch 4389 global offset for the X 4390 dimension is passed in the 4391 kernarg. 4392 4393 "HiddenGlobalOffsetY" 4394 The OpenCL grid dispatch 4395 global offset for the Y 4396 dimension is passed in the 4397 kernarg. 4398 4399 "HiddenGlobalOffsetZ" 4400 The OpenCL grid dispatch 4401 global offset for the Z 4402 dimension is passed in the 4403 kernarg. 4404 4405 "HiddenNone" 4406 An argument that is not used 4407 by the kernel. Space needs to 4408 be left for it, but it does 4409 not need to be set up. 4410 4411 "HiddenPrintfBuffer" 4412 A global address space pointer 4413 to the runtime printf buffer 4414 is passed in kernarg. 4415 4416 "HiddenHostcallBuffer" 4417 A global address space pointer 4418 to the runtime hostcall buffer 4419 is passed in kernarg. 4420 4421 "HiddenDefaultQueue" 4422 A global address space pointer 4423 to the OpenCL device enqueue 4424 queue that should be used by 4425 the kernel by default is 4426 passed in the kernarg. 4427 4428 "HiddenCompletionAction" 4429 A global address space pointer 4430 to help link enqueued kernels into 4431 the ancestor tree for determining 4432 when the parent kernel has finished. 4433 4434 "HiddenMultiGridSyncArg" 4435 A global address space pointer for 4436 multi-grid synchronization is 4437 passed in the kernarg. 4438 4439 "ValueType" string Required Kernel argument value type. Only 4440 present if "ValueKind" is 4441 "ByValue". For vector data 4442 types, the value is for the 4443 element type. Values include: 4444 4445 - "Struct" 4446 - "I8" 4447 - "U8" 4448 - "I16" 4449 - "U16" 4450 - "F16" 4451 - "I32" 4452 - "U32" 4453 - "F32" 4454 - "I64" 4455 - "U64" 4456 - "F64" 4457 4458 .. TODO:: 4459 How can it be determined if a 4460 vector type, and what size 4461 vector? 4462 "PointeeAlign" integer Alignment in bytes of pointee 4463 type for pointer type kernel 4464 argument. Must be a power 4465 of 2. Only present if 4466 "ValueKind" is 4467 "DynamicSharedPointer". 4468 "AddrSpaceQual" string Kernel argument address space 4469 qualifier. Only present if 4470 "ValueKind" is "GlobalBuffer" or 4471 "DynamicSharedPointer". Values 4472 are: 4473 4474 - "Private" 4475 - "Global" 4476 - "Constant" 4477 - "Local" 4478 - "Generic" 4479 - "Region" 4480 4481 .. TODO:: 4482 Is GlobalBuffer only Global 4483 or Constant? Is 4484 DynamicSharedPointer always 4485 Local? Can HCC allow Generic? 4486 How can Private or Region 4487 ever happen? 4488 "AccQual" string Kernel argument access 4489 qualifier. Only present if 4490 "ValueKind" is "Image" or 4491 "Pipe". Values 4492 are: 4493 4494 - "ReadOnly" 4495 - "WriteOnly" 4496 - "ReadWrite" 4497 4498 .. TODO:: 4499 Does this apply to 4500 GlobalBuffer? 4501 "ActualAccQual" string The actual memory accesses 4502 performed by the kernel on the 4503 kernel argument. Only present if 4504 "ValueKind" is "GlobalBuffer", 4505 "Image", or "Pipe". This may be 4506 more restrictive than indicated 4507 by "AccQual" to reflect what the 4508 kernel actual does. If not 4509 present then the runtime must 4510 assume what is implied by 4511 "AccQual" and "IsConst". Values 4512 are: 4513 4514 - "ReadOnly" 4515 - "WriteOnly" 4516 - "ReadWrite" 4517 4518 "IsConst" boolean Indicates if the kernel argument 4519 is const qualified. Only present 4520 if "ValueKind" is 4521 "GlobalBuffer". 4522 4523 "IsRestrict" boolean Indicates if the kernel argument 4524 is restrict qualified. Only 4525 present if "ValueKind" is 4526 "GlobalBuffer". 4527 4528 "IsVolatile" boolean Indicates if the kernel argument 4529 is volatile qualified. Only 4530 present if "ValueKind" is 4531 "GlobalBuffer". 4532 4533 "IsPipe" boolean Indicates if the kernel argument 4534 is pipe qualified. Only present 4535 if "ValueKind" is "Pipe". 4536 4537 .. TODO:: 4538 Can GlobalBuffer be pipe 4539 qualified? 4540 ================= ============== ========= ================================ 4541 4542.. 4543 4544 .. table:: AMDHSA Code Object V2 Kernel Code Properties Metadata Map 4545 :name: amdgpu-amdhsa-code-object-kernel-code-properties-metadata-map-table-v2 4546 4547 ============================ ============== ========= ===================== 4548 String Key Value Type Required? Description 4549 ============================ ============== ========= ===================== 4550 "KernargSegmentSize" integer Required The size in bytes of 4551 the kernarg segment 4552 that holds the values 4553 of the arguments to 4554 the kernel. 4555 "GroupSegmentFixedSize" integer Required The amount of group 4556 segment memory 4557 required by a 4558 work-group in 4559 bytes. This does not 4560 include any 4561 dynamically allocated 4562 group segment memory 4563 that may be added 4564 when the kernel is 4565 dispatched. 4566 "PrivateSegmentFixedSize" integer Required The amount of fixed 4567 private address space 4568 memory required for a 4569 work-item in 4570 bytes. If the kernel 4571 uses a dynamic call 4572 stack then additional 4573 space must be added 4574 to this value for the 4575 call stack. 4576 "KernargSegmentAlign" integer Required The maximum byte 4577 alignment of 4578 arguments in the 4579 kernarg segment. Must 4580 be a power of 2. 4581 "WavefrontSize" integer Required Wavefront size. Must 4582 be a power of 2. 4583 "NumSGPRs" integer Required Number of scalar 4584 registers used by a 4585 wavefront for 4586 GFX6-GFX10. This 4587 includes the special 4588 SGPRs for VCC, Flat 4589 Scratch (GFX7-GFX10) 4590 and XNACK (for 4591 GFX8-GFX10). It does 4592 not include the 16 4593 SGPR added if a trap 4594 handler is 4595 enabled. It is not 4596 rounded up to the 4597 allocation 4598 granularity. 4599 "NumVGPRs" integer Required Number of vector 4600 registers used by 4601 each work-item for 4602 GFX6-GFX10 4603 "MaxFlatWorkGroupSize" integer Required Maximum flat 4604 work-group size 4605 supported by the 4606 kernel in work-items. 4607 Must be >=1 and 4608 consistent with 4609 ReqdWorkGroupSize if 4610 not 0, 0, 0. 4611 "NumSpilledSGPRs" integer Number of stores from 4612 a scalar register to 4613 a register allocator 4614 created spill 4615 location. 4616 "NumSpilledVGPRs" integer Number of stores from 4617 a vector register to 4618 a register allocator 4619 created spill 4620 location. 4621 ============================ ============== ========= ===================== 4622 4623.. _amdgpu-amdhsa-code-object-metadata-v3: 4624 4625Code Object V3 Metadata (-mattr=+code-object-v3) 4626++++++++++++++++++++++++++++++++++++++++++++++++ 4627 4628Code object V3 metadata is specified by the ``NT_AMDGPU_METADATA`` note record 4629(see :ref:`amdgpu-note-records-v3`). 4630 4631The metadata is represented as Message Pack formatted binary data (see 4632[MsgPack]_). The top level is a Message Pack map that includes the 4633keys defined in table 4634:ref:`amdgpu-amdhsa-code-object-metadata-map-table-v3` and referenced 4635tables. 4636 4637Additional information can be added to the maps. To avoid conflicts, 4638any key names should be prefixed by "*vendor-name*." where 4639``vendor-name`` can be the name of the vendor and specific vendor 4640tool that generates the information. The prefix is abbreviated to 4641simply "." when it appears within a map that has been added by the 4642same *vendor-name*. 4643 4644 .. table:: AMDHSA Code Object V3 Metadata Map 4645 :name: amdgpu-amdhsa-code-object-metadata-map-table-v3 4646 4647 ================= ============== ========= ======================================= 4648 String Key Value Type Required? Description 4649 ================= ============== ========= ======================================= 4650 "amdhsa.version" sequence of Required - The first integer is the major 4651 2 integers version. Currently 1. 4652 - The second integer is the minor 4653 version. Currently 0. 4654 "amdhsa.printf" sequence of Each string is encoded information 4655 strings about a printf function call. The 4656 encoded information is organized as 4657 fields separated by colon (':'): 4658 4659 ``ID:N:S[0]:S[1]:...:S[N-1]:FormatString`` 4660 4661 where: 4662 4663 ``ID`` 4664 A 32-bit integer as a unique id for 4665 each printf function call 4666 4667 ``N`` 4668 A 32-bit integer equal to the number 4669 of arguments of printf function call 4670 minus 1 4671 4672 ``S[i]`` (where i = 0, 1, ... , N-1) 4673 32-bit integers for the size in bytes 4674 of the i-th FormatString argument of 4675 the printf function call 4676 4677 FormatString 4678 The format string passed to the 4679 printf function call. 4680 "amdhsa.kernels" sequence of Required Sequence of the maps for each 4681 map kernel in the code object. See 4682 :ref:`amdgpu-amdhsa-code-object-kernel-metadata-map-table-v3` 4683 for the definition of the keys included 4684 in that map. 4685 ================= ============== ========= ======================================= 4686 4687.. 4688 4689 .. table:: AMDHSA Code Object V3 Kernel Metadata Map 4690 :name: amdgpu-amdhsa-code-object-kernel-metadata-map-table-v3 4691 4692 =================================== ============== ========= ================================ 4693 String Key Value Type Required? Description 4694 =================================== ============== ========= ================================ 4695 ".name" string Required Source name of the kernel. 4696 ".symbol" string Required Name of the kernel 4697 descriptor ELF symbol. 4698 ".language" string Source language of the kernel. 4699 Values include: 4700 4701 - "OpenCL C" 4702 - "OpenCL C++" 4703 - "HCC" 4704 - "HIP" 4705 - "OpenMP" 4706 - "Assembler" 4707 4708 ".language_version" sequence of - The first integer is the major 4709 2 integers version. 4710 - The second integer is the 4711 minor version. 4712 ".args" sequence of Sequence of maps of the 4713 map kernel arguments. See 4714 :ref:`amdgpu-amdhsa-code-object-kernel-argument-metadata-map-table-v3` 4715 for the definition of the keys 4716 included in that map. 4717 ".reqd_workgroup_size" sequence of If not 0, 0, 0 then all values 4718 3 integers must be >=1 and the dispatch 4719 work-group size X, Y, Z must 4720 correspond to the specified 4721 values. Defaults to 0, 0, 0. 4722 4723 Corresponds to the OpenCL 4724 ``reqd_work_group_size`` 4725 attribute. 4726 ".workgroup_size_hint" sequence of The dispatch work-group size 4727 3 integers X, Y, Z is likely to be the 4728 specified values. 4729 4730 Corresponds to the OpenCL 4731 ``work_group_size_hint`` 4732 attribute. 4733 ".vec_type_hint" string The name of a scalar or vector 4734 type. 4735 4736 Corresponds to the OpenCL 4737 ``vec_type_hint`` attribute. 4738 4739 ".device_enqueue_symbol" string The external symbol name 4740 associated with a kernel. 4741 OpenCL runtime allocates a 4742 global buffer for the symbol 4743 and saves the kernel's address 4744 to it, which is used for 4745 device side enqueueing. Only 4746 available for device side 4747 enqueued kernels. 4748 ".kernarg_segment_size" integer Required The size in bytes of 4749 the kernarg segment 4750 that holds the values 4751 of the arguments to 4752 the kernel. 4753 ".group_segment_fixed_size" integer Required The amount of group 4754 segment memory 4755 required by a 4756 work-group in 4757 bytes. This does not 4758 include any 4759 dynamically allocated 4760 group segment memory 4761 that may be added 4762 when the kernel is 4763 dispatched. 4764 ".private_segment_fixed_size" integer Required The amount of fixed 4765 private address space 4766 memory required for a 4767 work-item in 4768 bytes. If the kernel 4769 uses a dynamic call 4770 stack then additional 4771 space must be added 4772 to this value for the 4773 call stack. 4774 ".kernarg_segment_align" integer Required The maximum byte 4775 alignment of 4776 arguments in the 4777 kernarg segment. Must 4778 be a power of 2. 4779 ".wavefront_size" integer Required Wavefront size. Must 4780 be a power of 2. 4781 ".sgpr_count" integer Required Number of scalar 4782 registers required by a 4783 wavefront for 4784 GFX6-GFX9. A register 4785 is required if it is 4786 used explicitly, or 4787 if a higher numbered 4788 register is used 4789 explicitly. This 4790 includes the special 4791 SGPRs for VCC, Flat 4792 Scratch (GFX7-GFX9) 4793 and XNACK (for 4794 GFX8-GFX9). It does 4795 not include the 16 4796 SGPR added if a trap 4797 handler is 4798 enabled. It is not 4799 rounded up to the 4800 allocation 4801 granularity. 4802 ".vgpr_count" integer Required Number of vector 4803 registers required by 4804 each work-item for 4805 GFX6-GFX9. A register 4806 is required if it is 4807 used explicitly, or 4808 if a higher numbered 4809 register is used 4810 explicitly. 4811 ".max_flat_workgroup_size" integer Required Maximum flat 4812 work-group size 4813 supported by the 4814 kernel in work-items. 4815 Must be >=1 and 4816 consistent with 4817 ReqdWorkGroupSize if 4818 not 0, 0, 0. 4819 ".sgpr_spill_count" integer Number of stores from 4820 a scalar register to 4821 a register allocator 4822 created spill 4823 location. 4824 ".vgpr_spill_count" integer Number of stores from 4825 a vector register to 4826 a register allocator 4827 created spill 4828 location. 4829 =================================== ============== ========= ================================ 4830 4831.. 4832 4833 .. table:: AMDHSA Code Object V3 Kernel Argument Metadata Map 4834 :name: amdgpu-amdhsa-code-object-kernel-argument-metadata-map-table-v3 4835 4836 ====================== ============== ========= ================================ 4837 String Key Value Type Required? Description 4838 ====================== ============== ========= ================================ 4839 ".name" string Kernel argument name. 4840 ".type_name" string Kernel argument type name. 4841 ".size" integer Required Kernel argument size in bytes. 4842 ".offset" integer Required Kernel argument offset in 4843 bytes. The offset must be a 4844 multiple of the alignment 4845 required by the argument. 4846 ".value_kind" string Required Kernel argument kind that 4847 specifies how to set up the 4848 corresponding argument. 4849 Values include: 4850 4851 "by_value" 4852 The argument is copied 4853 directly into the kernarg. 4854 4855 "global_buffer" 4856 A global address space pointer 4857 to the buffer data is passed 4858 in the kernarg. 4859 4860 "dynamic_shared_pointer" 4861 A group address space pointer 4862 to dynamically allocated LDS 4863 is passed in the kernarg. 4864 4865 "sampler" 4866 A global address space 4867 pointer to a S# is passed in 4868 the kernarg. 4869 4870 "image" 4871 A global address space 4872 pointer to a T# is passed in 4873 the kernarg. 4874 4875 "pipe" 4876 A global address space pointer 4877 to an OpenCL pipe is passed in 4878 the kernarg. 4879 4880 "queue" 4881 A global address space pointer 4882 to an OpenCL device enqueue 4883 queue is passed in the 4884 kernarg. 4885 4886 "hidden_global_offset_x" 4887 The OpenCL grid dispatch 4888 global offset for the X 4889 dimension is passed in the 4890 kernarg. 4891 4892 "hidden_global_offset_y" 4893 The OpenCL grid dispatch 4894 global offset for the Y 4895 dimension is passed in the 4896 kernarg. 4897 4898 "hidden_global_offset_z" 4899 The OpenCL grid dispatch 4900 global offset for the Z 4901 dimension is passed in the 4902 kernarg. 4903 4904 "hidden_none" 4905 An argument that is not used 4906 by the kernel. Space needs to 4907 be left for it, but it does 4908 not need to be set up. 4909 4910 "hidden_printf_buffer" 4911 A global address space pointer 4912 to the runtime printf buffer 4913 is passed in kernarg. 4914 4915 "hidden_hostcall_buffer" 4916 A global address space pointer 4917 to the runtime hostcall buffer 4918 is passed in kernarg. 4919 4920 "hidden_default_queue" 4921 A global address space pointer 4922 to the OpenCL device enqueue 4923 queue that should be used by 4924 the kernel by default is 4925 passed in the kernarg. 4926 4927 "hidden_completion_action" 4928 A global address space pointer 4929 to help link enqueued kernels into 4930 the ancestor tree for determining 4931 when the parent kernel has finished. 4932 4933 "hidden_multigrid_sync_arg" 4934 A global address space pointer for 4935 multi-grid synchronization is 4936 passed in the kernarg. 4937 4938 ".value_type" string Required Kernel argument value type. Only 4939 present if ".value_kind" is 4940 "by_value". For vector data 4941 types, the value is for the 4942 element type. Values include: 4943 4944 - "struct" 4945 - "i8" 4946 - "u8" 4947 - "i16" 4948 - "u16" 4949 - "f16" 4950 - "i32" 4951 - "u32" 4952 - "f32" 4953 - "i64" 4954 - "u64" 4955 - "f64" 4956 4957 .. TODO:: 4958 How can it be determined if a 4959 vector type, and what size 4960 vector? 4961 ".pointee_align" integer Alignment in bytes of pointee 4962 type for pointer type kernel 4963 argument. Must be a power 4964 of 2. Only present if 4965 ".value_kind" is 4966 "dynamic_shared_pointer". 4967 ".address_space" string Kernel argument address space 4968 qualifier. Only present if 4969 ".value_kind" is "global_buffer" or 4970 "dynamic_shared_pointer". Values 4971 are: 4972 4973 - "private" 4974 - "global" 4975 - "constant" 4976 - "local" 4977 - "generic" 4978 - "region" 4979 4980 .. TODO:: 4981 Is "global_buffer" only "global" 4982 or "constant"? Is 4983 "dynamic_shared_pointer" always 4984 "local"? Can HCC allow "generic"? 4985 How can "private" or "region" 4986 ever happen? 4987 ".access" string Kernel argument access 4988 qualifier. Only present if 4989 ".value_kind" is "image" or 4990 "pipe". Values 4991 are: 4992 4993 - "read_only" 4994 - "write_only" 4995 - "read_write" 4996 4997 .. TODO:: 4998 Does this apply to 4999 "global_buffer"? 5000 ".actual_access" string The actual memory accesses 5001 performed by the kernel on the 5002 kernel argument. Only present if 5003 ".value_kind" is "global_buffer", 5004 "image", or "pipe". This may be 5005 more restrictive than indicated 5006 by ".access" to reflect what the 5007 kernel actual does. If not 5008 present then the runtime must 5009 assume what is implied by 5010 ".access" and ".is_const" . Values 5011 are: 5012 5013 - "read_only" 5014 - "write_only" 5015 - "read_write" 5016 5017 ".is_const" boolean Indicates if the kernel argument 5018 is const qualified. Only present 5019 if ".value_kind" is 5020 "global_buffer". 5021 5022 ".is_restrict" boolean Indicates if the kernel argument 5023 is restrict qualified. Only 5024 present if ".value_kind" is 5025 "global_buffer". 5026 5027 ".is_volatile" boolean Indicates if the kernel argument 5028 is volatile qualified. Only 5029 present if ".value_kind" is 5030 "global_buffer". 5031 5032 ".is_pipe" boolean Indicates if the kernel argument 5033 is pipe qualified. Only present 5034 if ".value_kind" is "pipe". 5035 5036 .. TODO:: 5037 Can "global_buffer" be pipe 5038 qualified? 5039 ====================== ============== ========= ================================ 5040 5041.. 5042 5043Kernel Dispatch 5044~~~~~~~~~~~~~~~ 5045 5046The HSA architected queuing language (AQL) defines a user space memory 5047interface that can be used to control the dispatch of kernels, in an agent 5048independent way. An agent can have zero or more AQL queues created for it using 5049the ROCm runtime, in which AQL packets (all of which are 64 bytes) can be 5050placed. See the *HSA Platform System Architecture Specification* [HSA]_ for the 5051AQL queue mechanics and packet layouts. 5052 5053The packet processor of a kernel agent is responsible for detecting and 5054dispatching HSA kernels from the AQL queues associated with it. For AMD GPUs the 5055packet processor is implemented by the hardware command processor (CP), 5056asynchronous dispatch controller (ADC) and shader processor input controller 5057(SPI). 5058 5059The ROCm runtime can be used to allocate an AQL queue object. It uses the kernel 5060mode driver to initialize and register the AQL queue with CP. 5061 5062To dispatch a kernel the following actions are performed. This can occur in the 5063CPU host program, or from an HSA kernel executing on a GPU. 5064 50651. A pointer to an AQL queue for the kernel agent on which the kernel is to be 5066 executed is obtained. 50672. A pointer to the kernel descriptor (see 5068 :ref:`amdgpu-amdhsa-kernel-descriptor`) of the kernel to execute is obtained. 5069 It must be for a kernel that is contained in a code object that that was 5070 loaded by the ROCm runtime on the kernel agent with which the AQL queue is 5071 associated. 50723. Space is allocated for the kernel arguments using the ROCm runtime allocator 5073 for a memory region with the kernarg property for the kernel agent that will 5074 execute the kernel. It must be at least 16 byte aligned. 50754. Kernel argument values are assigned to the kernel argument memory 5076 allocation. The layout is defined in the *HSA Programmer's Language 5077 Reference* [HSA]_. For AMDGPU the kernel execution directly accesses the 5078 kernel argument memory in the same way constant memory is accessed. (Note 5079 that the HSA specification allows an implementation to copy the kernel 5080 argument contents to another location that is accessed by the kernel.) 50815. An AQL kernel dispatch packet is created on the AQL queue. The ROCm runtime 5082 api uses 64-bit atomic operations to reserve space in the AQL queue for the 5083 packet. The packet must be set up, and the final write must use an atomic 5084 store release to set the packet kind to ensure the packet contents are 5085 visible to the kernel agent. AQL defines a doorbell signal mechanism to 5086 notify the kernel agent that the AQL queue has been updated. These rules, and 5087 the layout of the AQL queue and kernel dispatch packet is defined in the *HSA 5088 System Architecture Specification* [HSA]_. 50896. A kernel dispatch packet includes information about the actual dispatch, 5090 such as grid and work-group size, together with information from the code 5091 object about the kernel, such as segment sizes. The ROCm runtime queries on 5092 the kernel symbol can be used to obtain the code object values which are 5093 recorded in the :ref:`amdgpu-amdhsa-code-object-metadata`. 50947. CP executes micro-code and is responsible for detecting and setting up the 5095 GPU to execute the wavefronts of a kernel dispatch. 50968. CP ensures that when the a wavefront starts executing the kernel machine 5097 code, the scalar general purpose registers (SGPR) and vector general purpose 5098 registers (VGPR) are set up as required by the machine code. The required 5099 setup is defined in the :ref:`amdgpu-amdhsa-kernel-descriptor`. The initial 5100 register state is defined in 5101 :ref:`amdgpu-amdhsa-initial-kernel-execution-state`. 51029. The prolog of the kernel machine code (see 5103 :ref:`amdgpu-amdhsa-kernel-prolog`) sets up the machine state as necessary 5104 before continuing executing the machine code that corresponds to the kernel. 510510. When the kernel dispatch has completed execution, CP signals the completion 5106 signal specified in the kernel dispatch packet if not 0. 5107 5108Image and Samplers 5109~~~~~~~~~~~~~~~~~~ 5110 5111Image and sample handles created by the ROCm runtime are 64-bit addresses of a 5112hardware 32 byte V# and 48 byte S# object respectively. In order to support the 5113HSA ``query_sampler`` operations two extra dwords are used to store the HSA BRIG 5114enumeration values for the queries that are not trivially deducible from the S# 5115representation. 5116 5117HSA Signals 5118~~~~~~~~~~~ 5119 5120HSA signal handles created by the ROCm runtime are 64-bit addresses of a 5121structure allocated in memory accessible from both the CPU and GPU. The 5122structure is defined by the ROCm runtime and subject to change between releases 5123(see [AMD-ROCm-github]_). 5124 5125.. _amdgpu-amdhsa-hsa-aql-queue: 5126 5127HSA AQL Queue 5128~~~~~~~~~~~~~ 5129 5130The HSA AQL queue structure is defined by the ROCm runtime and subject to change 5131between releases (see [AMD-ROCm-github]_). For some processors it contains 5132fields needed to implement certain language features such as the flat address 5133aperture bases. It also contains fields used by CP such as managing the 5134allocation of scratch memory. 5135 5136.. _amdgpu-amdhsa-kernel-descriptor: 5137 5138Kernel Descriptor 5139~~~~~~~~~~~~~~~~~ 5140 5141A kernel descriptor consists of the information needed by CP to initiate the 5142execution of a kernel, including the entry point address of the machine code 5143that implements the kernel. 5144 5145Kernel Descriptor for GFX6-GFX10 5146++++++++++++++++++++++++++++++++ 5147 5148CP microcode requires the Kernel descriptor to be allocated on 64 byte 5149alignment. 5150 5151 .. table:: Kernel Descriptor for GFX6-GFX10 5152 :name: amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table 5153 5154 ======= ======= =============================== ============================ 5155 Bits Size Field Name Description 5156 ======= ======= =============================== ============================ 5157 31:0 4 bytes GROUP_SEGMENT_FIXED_SIZE The amount of fixed local 5158 address space memory 5159 required for a work-group 5160 in bytes. This does not 5161 include any dynamically 5162 allocated local address 5163 space memory that may be 5164 added when the kernel is 5165 dispatched. 5166 63:32 4 bytes PRIVATE_SEGMENT_FIXED_SIZE The amount of fixed 5167 private address space 5168 memory required for a 5169 work-item in bytes. If 5170 is_dynamic_callstack is 1 5171 then additional space must 5172 be added to this value for 5173 the call stack. 5174 127:64 8 bytes Reserved, must be 0. 5175 191:128 8 bytes KERNEL_CODE_ENTRY_BYTE_OFFSET Byte offset (possibly 5176 negative) from base 5177 address of kernel 5178 descriptor to kernel's 5179 entry point instruction 5180 which must be 256 byte 5181 aligned. 5182 351:272 20 Reserved, must be 0. 5183 bytes 5184 383:352 4 bytes COMPUTE_PGM_RSRC3 GFX6-9 5185 Reserved, must be 0. 5186 GFX10 5187 Compute Shader (CS) 5188 program settings used by 5189 CP to set up 5190 ``COMPUTE_PGM_RSRC3`` 5191 configuration 5192 register. See 5193 :ref:`amdgpu-amdhsa-compute_pgm_rsrc3-gfx10-table`. 5194 415:384 4 bytes COMPUTE_PGM_RSRC1 Compute Shader (CS) 5195 program settings used by 5196 CP to set up 5197 ``COMPUTE_PGM_RSRC1`` 5198 configuration 5199 register. See 5200 :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`. 5201 447:416 4 bytes COMPUTE_PGM_RSRC2 Compute Shader (CS) 5202 program settings used by 5203 CP to set up 5204 ``COMPUTE_PGM_RSRC2`` 5205 configuration 5206 register. See 5207 :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`. 5208 448 1 bit ENABLE_SGPR_PRIVATE_SEGMENT Enable the setup of the 5209 _BUFFER SGPR user data registers 5210 (see 5211 :ref:`amdgpu-amdhsa-initial-kernel-execution-state`). 5212 5213 The total number of SGPR 5214 user data registers 5215 requested must not exceed 5216 16 and match value in 5217 ``compute_pgm_rsrc2.user_sgpr.user_sgpr_count``. 5218 Any requests beyond 16 5219 will be ignored. 5220 449 1 bit ENABLE_SGPR_DISPATCH_PTR *see above* 5221 450 1 bit ENABLE_SGPR_QUEUE_PTR *see above* 5222 451 1 bit ENABLE_SGPR_KERNARG_SEGMENT_PTR *see above* 5223 452 1 bit ENABLE_SGPR_DISPATCH_ID *see above* 5224 453 1 bit ENABLE_SGPR_FLAT_SCRATCH_INIT *see above* 5225 454 1 bit ENABLE_SGPR_PRIVATE_SEGMENT *see above* 5226 _SIZE 5227 457:455 3 bits Reserved, must be 0. 5228 458 1 bit ENABLE_WAVEFRONT_SIZE32 GFX6-9 5229 Reserved, must be 0. 5230 GFX10 5231 - If 0 execute in 5232 wavefront size 64 mode. 5233 - If 1 execute in 5234 native wavefront size 5235 32 mode. 5236 463:459 5 bits Reserved, must be 0. 5237 511:464 6 bytes Reserved, must be 0. 5238 512 **Total size 64 bytes.** 5239 ======= ==================================================================== 5240 5241.. 5242 5243 .. table:: compute_pgm_rsrc1 for GFX6-GFX10 5244 :name: amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table 5245 5246 ======= ======= =============================== =========================================================================== 5247 Bits Size Field Name Description 5248 ======= ======= =============================== =========================================================================== 5249 5:0 6 bits GRANULATED_WORKITEM_VGPR_COUNT Number of vector register 5250 blocks used by each work-item; 5251 granularity is device 5252 specific: 5253 5254 GFX6-GFX9 5255 - vgprs_used 0..256 5256 - max(0, ceil(vgprs_used / 4) - 1) 5257 GFX10 (wavefront size 64) 5258 - max_vgpr 1..256 5259 - max(0, ceil(vgprs_used / 4) - 1) 5260 GFX10 (wavefront size 32) 5261 - max_vgpr 1..256 5262 - max(0, ceil(vgprs_used / 8) - 1) 5263 5264 Where vgprs_used is defined 5265 as the highest VGPR number 5266 explicitly referenced plus 5267 one. 5268 5269 Used by CP to set up 5270 ``COMPUTE_PGM_RSRC1.VGPRS``. 5271 5272 The 5273 :ref:`amdgpu-assembler` 5274 calculates this 5275 automatically for the 5276 selected processor from 5277 values provided to the 5278 `.amdhsa_kernel` directive 5279 by the 5280 `.amdhsa_next_free_vgpr` 5281 nested directive (see 5282 :ref:`amdhsa-kernel-directives-table`). 5283 9:6 4 bits GRANULATED_WAVEFRONT_SGPR_COUNT Number of scalar register 5284 blocks used by a wavefront; 5285 granularity is device 5286 specific: 5287 5288 GFX6-GFX8 5289 - sgprs_used 0..112 5290 - max(0, ceil(sgprs_used / 8) - 1) 5291 GFX9 5292 - sgprs_used 0..112 5293 - 2 * max(0, ceil(sgprs_used / 16) - 1) 5294 GFX10 5295 Reserved, must be 0. 5296 (128 SGPRs always 5297 allocated.) 5298 5299 Where sgprs_used is 5300 defined as the highest 5301 SGPR number explicitly 5302 referenced plus one, plus 5303 a target-specific number 5304 of additional special 5305 SGPRs for VCC, 5306 FLAT_SCRATCH (GFX7+) and 5307 XNACK_MASK (GFX8+), and 5308 any additional 5309 target-specific 5310 limitations. It does not 5311 include the 16 SGPRs added 5312 if a trap handler is 5313 enabled. 5314 5315 The target-specific 5316 limitations and special 5317 SGPR layout are defined in 5318 the hardware 5319 documentation, which can 5320 be found in the 5321 :ref:`amdgpu-processors` 5322 table. 5323 5324 Used by CP to set up 5325 ``COMPUTE_PGM_RSRC1.SGPRS``. 5326 5327 The 5328 :ref:`amdgpu-assembler` 5329 calculates this 5330 automatically for the 5331 selected processor from 5332 values provided to the 5333 `.amdhsa_kernel` directive 5334 by the 5335 `.amdhsa_next_free_sgpr` 5336 and `.amdhsa_reserve_*` 5337 nested directives (see 5338 :ref:`amdhsa-kernel-directives-table`). 5339 11:10 2 bits PRIORITY Must be 0. 5340 5341 Start executing wavefront 5342 at the specified priority. 5343 5344 CP is responsible for 5345 filling in 5346 ``COMPUTE_PGM_RSRC1.PRIORITY``. 5347 13:12 2 bits FLOAT_ROUND_MODE_32 Wavefront starts execution 5348 with specified rounding 5349 mode for single (32 5350 bit) floating point 5351 precision floating point 5352 operations. 5353 5354 Floating point rounding 5355 mode values are defined in 5356 :ref:`amdgpu-amdhsa-floating-point-rounding-mode-enumeration-values-table`. 5357 5358 Used by CP to set up 5359 ``COMPUTE_PGM_RSRC1.FLOAT_MODE``. 5360 15:14 2 bits FLOAT_ROUND_MODE_16_64 Wavefront starts execution 5361 with specified rounding 5362 denorm mode for half/double (16 5363 and 64-bit) floating point 5364 precision floating point 5365 operations. 5366 5367 Floating point rounding 5368 mode values are defined in 5369 :ref:`amdgpu-amdhsa-floating-point-rounding-mode-enumeration-values-table`. 5370 5371 Used by CP to set up 5372 ``COMPUTE_PGM_RSRC1.FLOAT_MODE``. 5373 17:16 2 bits FLOAT_DENORM_MODE_32 Wavefront starts execution 5374 with specified denorm mode 5375 for single (32 5376 bit) floating point 5377 precision floating point 5378 operations. 5379 5380 Floating point denorm mode 5381 values are defined in 5382 :ref:`amdgpu-amdhsa-floating-point-denorm-mode-enumeration-values-table`. 5383 5384 Used by CP to set up 5385 ``COMPUTE_PGM_RSRC1.FLOAT_MODE``. 5386 19:18 2 bits FLOAT_DENORM_MODE_16_64 Wavefront starts execution 5387 with specified denorm mode 5388 for half/double (16 5389 and 64-bit) floating point 5390 precision floating point 5391 operations. 5392 5393 Floating point denorm mode 5394 values are defined in 5395 :ref:`amdgpu-amdhsa-floating-point-denorm-mode-enumeration-values-table`. 5396 5397 Used by CP to set up 5398 ``COMPUTE_PGM_RSRC1.FLOAT_MODE``. 5399 20 1 bit PRIV Must be 0. 5400 5401 Start executing wavefront 5402 in privilege trap handler 5403 mode. 5404 5405 CP is responsible for 5406 filling in 5407 ``COMPUTE_PGM_RSRC1.PRIV``. 5408 21 1 bit ENABLE_DX10_CLAMP Wavefront starts execution 5409 with DX10 clamp mode 5410 enabled. Used by the vector 5411 ALU to force DX10 style 5412 treatment of NaN's (when 5413 set, clamp NaN to zero, 5414 otherwise pass NaN 5415 through). 5416 5417 Used by CP to set up 5418 ``COMPUTE_PGM_RSRC1.DX10_CLAMP``. 5419 22 1 bit DEBUG_MODE Must be 0. 5420 5421 Start executing wavefront 5422 in single step mode. 5423 5424 CP is responsible for 5425 filling in 5426 ``COMPUTE_PGM_RSRC1.DEBUG_MODE``. 5427 23 1 bit ENABLE_IEEE_MODE Wavefront starts execution 5428 with IEEE mode 5429 enabled. Floating point 5430 opcodes that support 5431 exception flag gathering 5432 will quiet and propagate 5433 signaling-NaN inputs per 5434 IEEE 754-2008. Min_dx10 and 5435 max_dx10 become IEEE 5436 754-2008 compliant due to 5437 signaling-NaN propagation 5438 and quieting. 5439 5440 Used by CP to set up 5441 ``COMPUTE_PGM_RSRC1.IEEE_MODE``. 5442 24 1 bit BULKY Must be 0. 5443 5444 Only one work-group allowed 5445 to execute on a compute 5446 unit. 5447 5448 CP is responsible for 5449 filling in 5450 ``COMPUTE_PGM_RSRC1.BULKY``. 5451 25 1 bit CDBG_USER Must be 0. 5452 5453 Flag that can be used to 5454 control debugging code. 5455 5456 CP is responsible for 5457 filling in 5458 ``COMPUTE_PGM_RSRC1.CDBG_USER``. 5459 26 1 bit FP16_OVFL GFX6-GFX8 5460 Reserved, must be 0. 5461 GFX9-GFX10 5462 Wavefront starts execution 5463 with specified fp16 overflow 5464 mode. 5465 5466 - If 0, fp16 overflow generates 5467 +/-INF values. 5468 - If 1, fp16 overflow that is the 5469 result of an +/-INF input value 5470 or divide by 0 produces a +/-INF, 5471 otherwise clamps computed 5472 overflow to +/-MAX_FP16 as 5473 appropriate. 5474 5475 Used by CP to set up 5476 ``COMPUTE_PGM_RSRC1.FP16_OVFL``. 5477 28:27 2 bits Reserved, must be 0. 5478 29 1 bit WGP_MODE GFX6-GFX9 5479 Reserved, must be 0. 5480 GFX10 5481 - If 0 execute work-groups in 5482 CU wavefront execution mode. 5483 - If 1 execute work-groups on 5484 in WGP wavefront execution mode. 5485 5486 See :ref:`amdgpu-amdhsa-memory-model`. 5487 5488 Used by CP to set up 5489 ``COMPUTE_PGM_RSRC1.WGP_MODE``. 5490 30 1 bit MEM_ORDERED GFX6-9 5491 Reserved, must be 0. 5492 GFX10 5493 Controls the behavior of the 5494 waitcnt's vmcnt and vscnt 5495 counters. 5496 5497 - If 0 vmcnt reports completion 5498 of load and atomic with return 5499 out of order with sample 5500 instructions, and the vscnt 5501 reports the completion of 5502 store and atomic without 5503 return in order. 5504 - If 1 vmcnt reports completion 5505 of load, atomic with return 5506 and sample instructions in 5507 order, and the vscnt reports 5508 the completion of store and 5509 atomic without return in order. 5510 5511 Used by CP to set up 5512 ``COMPUTE_PGM_RSRC1.MEM_ORDERED``. 5513 31 1 bit FWD_PROGRESS GFX6-9 5514 Reserved, must be 0. 5515 GFX10 5516 - If 0 execute SIMD wavefronts 5517 using oldest first policy. 5518 - If 1 execute SIMD wavefronts to 5519 ensure wavefronts will make some 5520 forward progress. 5521 5522 Used by CP to set up 5523 ``COMPUTE_PGM_RSRC1.FWD_PROGRESS``. 5524 32 **Total size 4 bytes** 5525 ======= =================================================================================================================== 5526 5527.. 5528 5529 .. table:: compute_pgm_rsrc2 for GFX6-GFX10 5530 :name: amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table 5531 5532 ======= ======= =============================== =========================================================================== 5533 Bits Size Field Name Description 5534 ======= ======= =============================== =========================================================================== 5535 0 1 bit ENABLE_SGPR_PRIVATE_SEGMENT Enable the setup of the 5536 _WAVEFRONT_OFFSET SGPR wavefront scratch offset 5537 system register (see 5538 :ref:`amdgpu-amdhsa-initial-kernel-execution-state`). 5539 5540 Used by CP to set up 5541 ``COMPUTE_PGM_RSRC2.SCRATCH_EN``. 5542 5:1 5 bits USER_SGPR_COUNT The total number of SGPR 5543 user data registers 5544 requested. This number must 5545 match the number of user 5546 data registers enabled. 5547 5548 Used by CP to set up 5549 ``COMPUTE_PGM_RSRC2.USER_SGPR``. 5550 6 1 bit ENABLE_TRAP_HANDLER Must be 0. 5551 5552 This bit represents 5553 ``COMPUTE_PGM_RSRC2.TRAP_PRESENT``, 5554 which is set by the CP if 5555 the runtime has installed a 5556 trap handler. 5557 7 1 bit ENABLE_SGPR_WORKGROUP_ID_X Enable the setup of the 5558 system SGPR register for 5559 the work-group id in the X 5560 dimension (see 5561 :ref:`amdgpu-amdhsa-initial-kernel-execution-state`). 5562 5563 Used by CP to set up 5564 ``COMPUTE_PGM_RSRC2.TGID_X_EN``. 5565 8 1 bit ENABLE_SGPR_WORKGROUP_ID_Y Enable the setup of the 5566 system SGPR register for 5567 the work-group id in the Y 5568 dimension (see 5569 :ref:`amdgpu-amdhsa-initial-kernel-execution-state`). 5570 5571 Used by CP to set up 5572 ``COMPUTE_PGM_RSRC2.TGID_Y_EN``. 5573 9 1 bit ENABLE_SGPR_WORKGROUP_ID_Z Enable the setup of the 5574 system SGPR register for 5575 the work-group id in the Z 5576 dimension (see 5577 :ref:`amdgpu-amdhsa-initial-kernel-execution-state`). 5578 5579 Used by CP to set up 5580 ``COMPUTE_PGM_RSRC2.TGID_Z_EN``. 5581 10 1 bit ENABLE_SGPR_WORKGROUP_INFO Enable the setup of the 5582 system SGPR register for 5583 work-group information (see 5584 :ref:`amdgpu-amdhsa-initial-kernel-execution-state`). 5585 5586 Used by CP to set up 5587 ``COMPUTE_PGM_RSRC2.TGID_SIZE_EN``. 5588 12:11 2 bits ENABLE_VGPR_WORKITEM_ID Enable the setup of the 5589 VGPR system registers used 5590 for the work-item ID. 5591 :ref:`amdgpu-amdhsa-system-vgpr-work-item-id-enumeration-values-table` 5592 defines the values. 5593 5594 Used by CP to set up 5595 ``COMPUTE_PGM_RSRC2.TIDIG_CMP_CNT``. 5596 13 1 bit ENABLE_EXCEPTION_ADDRESS_WATCH Must be 0. 5597 5598 Wavefront starts execution 5599 with address watch 5600 exceptions enabled which 5601 are generated when L1 has 5602 witnessed a thread access 5603 an *address of 5604 interest*. 5605 5606 CP is responsible for 5607 filling in the address 5608 watch bit in 5609 ``COMPUTE_PGM_RSRC2.EXCP_EN_MSB`` 5610 according to what the 5611 runtime requests. 5612 14 1 bit ENABLE_EXCEPTION_MEMORY Must be 0. 5613 5614 Wavefront starts execution 5615 with memory violation 5616 exceptions exceptions 5617 enabled which are generated 5618 when a memory violation has 5619 occurred for this wavefront from 5620 L1 or LDS 5621 (write-to-read-only-memory, 5622 mis-aligned atomic, LDS 5623 address out of range, 5624 illegal address, etc.). 5625 5626 CP sets the memory 5627 violation bit in 5628 ``COMPUTE_PGM_RSRC2.EXCP_EN_MSB`` 5629 according to what the 5630 runtime requests. 5631 23:15 9 bits GRANULATED_LDS_SIZE Must be 0. 5632 5633 CP uses the rounded value 5634 from the dispatch packet, 5635 not this value, as the 5636 dispatch may contain 5637 dynamically allocated group 5638 segment memory. CP writes 5639 directly to 5640 ``COMPUTE_PGM_RSRC2.LDS_SIZE``. 5641 5642 Amount of group segment 5643 (LDS) to allocate for each 5644 work-group. Granularity is 5645 device specific: 5646 5647 GFX6: 5648 roundup(lds-size / (64 * 4)) 5649 GFX7-GFX10: 5650 roundup(lds-size / (128 * 4)) 5651 5652 24 1 bit ENABLE_EXCEPTION_IEEE_754_FP Wavefront starts execution 5653 _INVALID_OPERATION with specified exceptions 5654 enabled. 5655 5656 Used by CP to set up 5657 ``COMPUTE_PGM_RSRC2.EXCP_EN`` 5658 (set from bits 0..6). 5659 5660 IEEE 754 FP Invalid 5661 Operation 5662 25 1 bit ENABLE_EXCEPTION_FP_DENORMAL FP Denormal one or more 5663 _SOURCE input operands is a 5664 denormal number 5665 26 1 bit ENABLE_EXCEPTION_IEEE_754_FP IEEE 754 FP Division by 5666 _DIVISION_BY_ZERO Zero 5667 27 1 bit ENABLE_EXCEPTION_IEEE_754_FP IEEE 754 FP FP Overflow 5668 _OVERFLOW 5669 28 1 bit ENABLE_EXCEPTION_IEEE_754_FP IEEE 754 FP Underflow 5670 _UNDERFLOW 5671 29 1 bit ENABLE_EXCEPTION_IEEE_754_FP IEEE 754 FP Inexact 5672 _INEXACT 5673 30 1 bit ENABLE_EXCEPTION_INT_DIVIDE_BY Integer Division by Zero 5674 _ZERO (rcp_iflag_f32 instruction 5675 only) 5676 31 1 bit Reserved, must be 0. 5677 32 **Total size 4 bytes.** 5678 ======= =================================================================================================================== 5679 5680.. 5681 5682 .. table:: compute_pgm_rsrc3 for GFX10 5683 :name: amdgpu-amdhsa-compute_pgm_rsrc3-gfx10-table 5684 5685 ======= ======= =============================== =========================================================================== 5686 Bits Size Field Name Description 5687 ======= ======= =============================== =========================================================================== 5688 3:0 4 bits SHARED_VGPR_COUNT Number of shared VGPRs for wavefront size 64. Granularity 8. Value 0-120. 5689 compute_pgm_rsrc1.vgprs + shared_vgpr_cnt cannot exceed 64. 5690 31:4 28 Reserved, must be 0. 5691 bits 5692 32 **Total size 4 bytes.** 5693 ======= =================================================================================================================== 5694 5695.. 5696 5697 .. table:: Floating Point Rounding Mode Enumeration Values 5698 :name: amdgpu-amdhsa-floating-point-rounding-mode-enumeration-values-table 5699 5700 ====================================== ===== ============================== 5701 Enumeration Name Value Description 5702 ====================================== ===== ============================== 5703 FLOAT_ROUND_MODE_NEAR_EVEN 0 Round Ties To Even 5704 FLOAT_ROUND_MODE_PLUS_INFINITY 1 Round Toward +infinity 5705 FLOAT_ROUND_MODE_MINUS_INFINITY 2 Round Toward -infinity 5706 FLOAT_ROUND_MODE_ZERO 3 Round Toward 0 5707 ====================================== ===== ============================== 5708 5709.. 5710 5711 .. table:: Floating Point Denorm Mode Enumeration Values 5712 :name: amdgpu-amdhsa-floating-point-denorm-mode-enumeration-values-table 5713 5714 ====================================== ===== ============================== 5715 Enumeration Name Value Description 5716 ====================================== ===== ============================== 5717 FLOAT_DENORM_MODE_FLUSH_SRC_DST 0 Flush Source and Destination 5718 Denorms 5719 FLOAT_DENORM_MODE_FLUSH_DST 1 Flush Output Denorms 5720 FLOAT_DENORM_MODE_FLUSH_SRC 2 Flush Source Denorms 5721 FLOAT_DENORM_MODE_FLUSH_NONE 3 No Flush 5722 ====================================== ===== ============================== 5723 5724.. 5725 5726 .. table:: System VGPR Work-Item ID Enumeration Values 5727 :name: amdgpu-amdhsa-system-vgpr-work-item-id-enumeration-values-table 5728 5729 ======================================== ===== ============================ 5730 Enumeration Name Value Description 5731 ======================================== ===== ============================ 5732 SYSTEM_VGPR_WORKITEM_ID_X 0 Set work-item X dimension 5733 ID. 5734 SYSTEM_VGPR_WORKITEM_ID_X_Y 1 Set work-item X and Y 5735 dimensions ID. 5736 SYSTEM_VGPR_WORKITEM_ID_X_Y_Z 2 Set work-item X, Y and Z 5737 dimensions ID. 5738 SYSTEM_VGPR_WORKITEM_ID_UNDEFINED 3 Undefined. 5739 ======================================== ===== ============================ 5740 5741.. _amdgpu-amdhsa-initial-kernel-execution-state: 5742 5743Initial Kernel Execution State 5744~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 5745 5746This section defines the register state that will be set up by the packet 5747processor prior to the start of execution of every wavefront. This is limited by 5748the constraints of the hardware controllers of CP/ADC/SPI. 5749 5750The order of the SGPR registers is defined, but the compiler can specify which 5751ones are actually setup in the kernel descriptor using the ``enable_sgpr_*`` bit 5752fields (see :ref:`amdgpu-amdhsa-kernel-descriptor`). The register numbers used 5753for enabled registers are dense starting at SGPR0: the first enabled register is 5754SGPR0, the next enabled register is SGPR1 etc.; disabled registers do not have 5755an SGPR number. 5756 5757The initial SGPRs comprise up to 16 User SRGPs that are set by CP and apply to 5758all wavefronts of the grid. It is possible to specify more than 16 User SGPRs 5759using the ``enable_sgpr_*`` bit fields, in which case only the first 16 are 5760actually initialized. These are then immediately followed by the System SGPRs 5761that are set up by ADC/SPI and can have different values for each wavefront of 5762the grid dispatch. 5763 5764SGPR register initial state is defined in 5765:ref:`amdgpu-amdhsa-sgpr-register-set-up-order-table`. 5766 5767 .. table:: SGPR Register Set Up Order 5768 :name: amdgpu-amdhsa-sgpr-register-set-up-order-table 5769 5770 ========== ========================== ====== ============================== 5771 SGPR Order Name Number Description 5772 (kernel descriptor enable of 5773 field) SGPRs 5774 ========== ========================== ====== ============================== 5775 First Private Segment Buffer 4 V# that can be used, together 5776 (enable_sgpr_private with Scratch Wavefront Offset 5777 _segment_buffer) as an offset, to access the 5778 private address space using a 5779 segment address. 5780 5781 CP uses the value provided by 5782 the runtime. 5783 then Dispatch Ptr 2 64-bit address of AQL dispatch 5784 (enable_sgpr_dispatch_ptr) packet for kernel dispatch 5785 actually executing. 5786 then Queue Ptr 2 64-bit address of amd_queue_t 5787 (enable_sgpr_queue_ptr) object for AQL queue on which 5788 the dispatch packet was 5789 queued. 5790 then Kernarg Segment Ptr 2 64-bit address of Kernarg 5791 (enable_sgpr_kernarg segment. This is directly 5792 _segment_ptr) copied from the 5793 kernarg_address in the kernel 5794 dispatch packet. 5795 5796 Having CP load it once avoids 5797 loading it at the beginning of 5798 every wavefront. 5799 then Dispatch Id 2 64-bit Dispatch ID of the 5800 (enable_sgpr_dispatch_id) dispatch packet being 5801 executed. 5802 then Flat Scratch Init 2 This is 2 SGPRs: 5803 (enable_sgpr_flat_scratch 5804 _init) GFX6 5805 Not supported. 5806 GFX7-GFX8 5807 The first SGPR is a 32-bit 5808 byte offset from 5809 ``SH_HIDDEN_PRIVATE_BASE_VIMID`` 5810 to per SPI base of memory 5811 for scratch for the queue 5812 executing the kernel 5813 dispatch. CP obtains this 5814 from the runtime. (The 5815 Scratch Segment Buffer base 5816 address is 5817 ``SH_HIDDEN_PRIVATE_BASE_VIMID`` 5818 plus this offset.) The value 5819 of Scratch Wavefront Offset must 5820 be added to this offset by 5821 the kernel machine code, 5822 right shifted by 8, and 5823 moved to the FLAT_SCRATCH_HI 5824 SGPR register. 5825 FLAT_SCRATCH_HI corresponds 5826 to SGPRn-4 on GFX7, and 5827 SGPRn-6 on GFX8 (where SGPRn 5828 is the highest numbered SGPR 5829 allocated to the wavefront). 5830 FLAT_SCRATCH_HI is 5831 multiplied by 256 (as it is 5832 in units of 256 bytes) and 5833 added to 5834 ``SH_HIDDEN_PRIVATE_BASE_VIMID`` 5835 to calculate the per wavefront 5836 FLAT SCRATCH BASE in flat 5837 memory instructions that 5838 access the scratch 5839 aperture. 5840 5841 The second SGPR is 32-bit 5842 byte size of a single 5843 work-item's scratch memory 5844 usage. CP obtains this from 5845 the runtime, and it is 5846 always a multiple of DWORD. 5847 CP checks that the value in 5848 the kernel dispatch packet 5849 Private Segment Byte Size is 5850 not larger, and requests the 5851 runtime to increase the 5852 queue's scratch size if 5853 necessary. The kernel code 5854 must move it to 5855 FLAT_SCRATCH_LO which is 5856 SGPRn-3 on GFX7 and SGPRn-5 5857 on GFX8. FLAT_SCRATCH_LO is 5858 used as the FLAT SCRATCH 5859 SIZE in flat memory 5860 instructions. Having CP load 5861 it once avoids loading it at 5862 the beginning of every 5863 wavefront. 5864 GFX9-GFX10 5865 This is the 5866 64-bit base address of the 5867 per SPI scratch backing 5868 memory managed by SPI for 5869 the queue executing the 5870 kernel dispatch. CP obtains 5871 this from the runtime (and 5872 divides it if there are 5873 multiple Shader Arrays each 5874 with its own SPI). The value 5875 of Scratch Wavefront Offset must 5876 be added by the kernel 5877 machine code and the result 5878 moved to the FLAT_SCRATCH 5879 SGPR which is SGPRn-6 and 5880 SGPRn-5. It is used as the 5881 FLAT SCRATCH BASE in flat 5882 memory instructions. 5883 then Private Segment Size 1 The 32-bit byte size of a 5884 (enable_sgpr_private single 5885 work-item's 5886 scratch_segment_size) memory 5887 allocation. This is the 5888 value from the kernel 5889 dispatch packet Private 5890 Segment Byte Size rounded up 5891 by CP to a multiple of 5892 DWORD. 5893 5894 Having CP load it once avoids 5895 loading it at the beginning of 5896 every wavefront. 5897 5898 This is not used for 5899 GFX7-GFX8 since it is the same 5900 value as the second SGPR of 5901 Flat Scratch Init. However, it 5902 may be needed for GFX9-GFX10 which 5903 changes the meaning of the 5904 Flat Scratch Init value. 5905 then Grid Work-Group Count X 1 32-bit count of the number of 5906 (enable_sgpr_grid work-groups in the X dimension 5907 _workgroup_count_X) for the grid being 5908 executed. Computed from the 5909 fields in the kernel dispatch 5910 packet as ((grid_size.x + 5911 workgroup_size.x - 1) / 5912 workgroup_size.x). 5913 then Grid Work-Group Count Y 1 32-bit count of the number of 5914 (enable_sgpr_grid work-groups in the Y dimension 5915 _workgroup_count_Y && for the grid being 5916 less than 16 previous executed. Computed from the 5917 SGPRs) fields in the kernel dispatch 5918 packet as ((grid_size.y + 5919 workgroup_size.y - 1) / 5920 workgroupSize.y). 5921 5922 Only initialized if <16 5923 previous SGPRs initialized. 5924 then Grid Work-Group Count Z 1 32-bit count of the number of 5925 (enable_sgpr_grid work-groups in the Z dimension 5926 _workgroup_count_Z && for the grid being 5927 less than 16 previous executed. Computed from the 5928 SGPRs) fields in the kernel dispatch 5929 packet as ((grid_size.z + 5930 workgroup_size.z - 1) / 5931 workgroupSize.z). 5932 5933 Only initialized if <16 5934 previous SGPRs initialized. 5935 then Work-Group Id X 1 32-bit work-group id in X 5936 (enable_sgpr_workgroup_id dimension of grid for 5937 _X) wavefront. 5938 then Work-Group Id Y 1 32-bit work-group id in Y 5939 (enable_sgpr_workgroup_id dimension of grid for 5940 _Y) wavefront. 5941 then Work-Group Id Z 1 32-bit work-group id in Z 5942 (enable_sgpr_workgroup_id dimension of grid for 5943 _Z) wavefront. 5944 then Work-Group Info 1 {first_wavefront, 14'b0000, 5945 (enable_sgpr_workgroup ordered_append_term[10:0], 5946 _info) threadgroup_size_in_wavefronts[5:0]} 5947 then Scratch Wavefront Offset 1 32-bit byte offset from base 5948 (enable_sgpr_private of scratch base of queue 5949 _segment_wavefront_offset) executing the kernel 5950 dispatch. Must be used as an 5951 offset with Private 5952 segment address when using 5953 Scratch Segment Buffer. It 5954 must be used to set up FLAT 5955 SCRATCH for flat addressing 5956 (see 5957 :ref:`amdgpu-amdhsa-kernel-prolog-flat-scratch`). 5958 ========== ========================== ====== ============================== 5959 5960The order of the VGPR registers is defined, but the compiler can specify which 5961ones are actually setup in the kernel descriptor using the ``enable_vgpr*`` bit 5962fields (see :ref:`amdgpu-amdhsa-kernel-descriptor`). The register numbers used 5963for enabled registers are dense starting at VGPR0: the first enabled register is 5964VGPR0, the next enabled register is VGPR1 etc.; disabled registers do not have a 5965VGPR number. 5966 5967VGPR register initial state is defined in 5968:ref:`amdgpu-amdhsa-vgpr-register-set-up-order-table`. 5969 5970 .. table:: VGPR Register Set Up Order 5971 :name: amdgpu-amdhsa-vgpr-register-set-up-order-table 5972 5973 ========== ========================== ====== ============================== 5974 VGPR Order Name Number Description 5975 (kernel descriptor enable of 5976 field) VGPRs 5977 ========== ========================== ====== ============================== 5978 First Work-Item Id X 1 32-bit work item id in X 5979 (Always initialized) dimension of work-group for 5980 wavefront lane. 5981 then Work-Item Id Y 1 32-bit work item id in Y 5982 (enable_vgpr_workitem_id dimension of work-group for 5983 > 0) wavefront lane. 5984 then Work-Item Id Z 1 32-bit work item id in Z 5985 (enable_vgpr_workitem_id dimension of work-group for 5986 > 1) wavefront lane. 5987 ========== ========================== ====== ============================== 5988 5989The setting of registers is done by GPU CP/ADC/SPI hardware as follows: 5990 59911. SGPRs before the Work-Group Ids are set by CP using the 16 User Data 5992 registers. 59932. Work-group Id registers X, Y, Z are set by ADC which supports any 5994 combination including none. 59953. Scratch Wavefront Offset is set by SPI in a per wavefront basis which is why 5996 its value cannot included with the flat scratch init value which is per 5997 queue. 59984. The VGPRs are set by SPI which only supports specifying either (X), (X, Y) 5999 or (X, Y, Z). 6000 6001Flat Scratch register pair are adjacent SGRRs so they can be moved as a 64-bit 6002value to the hardware required SGPRn-3 and SGPRn-4 respectively. 6003 6004The global segment can be accessed either using buffer instructions (GFX6 which 6005has V# 64-bit address support), flat instructions (GFX7-GFX10), or global 6006instructions (GFX9-GFX10). 6007 6008If buffer operations are used then the compiler can generate a V# with the 6009following properties: 6010 6011* base address of 0 6012* no swizzle 6013* ATC: 1 if IOMMU present (such as APU) 6014* ptr64: 1 6015* MTYPE set to support memory coherence that matches the runtime (such as CC for 6016 APU and NC for dGPU). 6017 6018.. _amdgpu-amdhsa-kernel-prolog: 6019 6020Kernel Prolog 6021~~~~~~~~~~~~~ 6022 6023The compiler performs initialization in the kernel prologue depending on the 6024target and information about things like stack usage in the kernel and called 6025functions. Some of this initialization requires the compiler to request certain 6026User and System SGPRs be present in the 6027:ref:`amdgpu-amdhsa-initial-kernel-execution-state` via the 6028:ref:`amdgpu-amdhsa-kernel-descriptor`. 6029 6030.. _amdgpu-amdhsa-kernel-prolog-cfi: 6031 6032CFI 6033+++ 6034 60351. The CFI return address is undefined. 60362. The CFI CFA is defined using an expression which evaluates to a memory 6037 location description for the private segment address ``0``. 6038 6039.. _amdgpu-amdhsa-kernel-prolog-m0: 6040 6041M0 6042++ 6043 6044GFX6-GFX8 6045 The M0 register must be initialized with a value at least the total LDS size 6046 if the kernel may access LDS via DS or flat operations. Total LDS size is 6047 available in dispatch packet. For M0, it is also possible to use maximum 6048 possible value of LDS for given target (0x7FFF for GFX6 and 0xFFFF for 6049 GFX7-GFX8). 6050GFX9-GFX10 6051 The M0 register is not used for range checking LDS accesses and so does not 6052 need to be initialized in the prolog. 6053 6054.. _amdgpu-amdhsa-kernel-prolog-stack-pointer: 6055 6056Stack Pointer 6057+++++++++++++ 6058 6059If the kernel has function calls it must set up the ABI stack pointer described 6060in :ref:`amdgpu-amdhsa-function-call-convention-non-kernel-functions` by 6061setting SGPR32 to the the unswizzled scratch offset of the address past the 6062last local allocation. 6063 6064.. _amdgpu-amdhsa-kernel-prolog-frame-pointer: 6065 6066Frame Pointer 6067+++++++++++++ 6068 6069If the kernel needs a frame pointer for the reasons defined in 6070``SIFrameLowering`` then SGPR33 is used and is always set to ``0`` in the 6071kernel prolog. If a frame pointer is not required then all uses of the frame 6072pointer are replaced with immediate ``0`` offsets. 6073 6074.. _amdgpu-amdhsa-kernel-prolog-flat-scratch: 6075 6076Flat Scratch 6077++++++++++++ 6078 6079If the kernel or any function it calls may use flat operations to access 6080scratch memory, the prolog code must set up the FLAT_SCRATCH register pair 6081(FLAT_SCRATCH_LO/FLAT_SCRATCH_HI which are in SGPRn-4/SGPRn-3). Initialization 6082uses Flat Scratch Init and Scratch Wavefront Offset SGPR registers (see 6083:ref:`amdgpu-amdhsa-initial-kernel-execution-state`): 6084 6085GFX6 6086 Flat scratch is not supported. 6087 6088GFX7-GFX8 6089 6090 1. The low word of Flat Scratch Init is 32-bit byte offset from 6091 ``SH_HIDDEN_PRIVATE_BASE_VIMID`` to the base of scratch backing memory 6092 being managed by SPI for the queue executing the kernel dispatch. This is 6093 the same value used in the Scratch Segment Buffer V# base address. The 6094 prolog must add the value of Scratch Wavefront Offset to get the 6095 wavefront's byte scratch backing memory offset from 6096 ``SH_HIDDEN_PRIVATE_BASE_VIMID``. Since FLAT_SCRATCH_LO is in units of 256 6097 bytes, the offset must be right shifted by 8 before moving into 6098 FLAT_SCRATCH_LO. 6099 2. The second word of Flat Scratch Init is 32-bit byte size of a single 6100 work-items scratch memory usage. This is directly loaded from the kernel 6101 dispatch packet Private Segment Byte Size and rounded up to a multiple of 6102 DWORD. Having CP load it once avoids loading it at the beginning of every 6103 wavefront. The prolog must move it to FLAT_SCRATCH_LO for use as FLAT 6104 SCRATCH SIZE. 6105 6106GFX9-GFX10 6107 The Flat Scratch Init is the 64-bit address of the base of scratch backing 6108 memory being managed by SPI for the queue executing the kernel dispatch. The 6109 prolog must add the value of Scratch Wavefront Offset and moved to the 6110 FLAT_SCRATCH pair for use as the flat scratch base in flat memory 6111 instructions. 6112 6113.. _amdgpu-amdhsa-kernel-prolog-private-segment-buffer: 6114 6115Private Segment Buffer 6116++++++++++++++++++++++ 6117 6118A set of four SGPRs beginning at a four-aligned SGPR index are always selected 6119to serve as the scratch V# for the kernel as follows: 6120 6121 - If it is know during instruction selection that there is stack usage, 6122 SGPR0-3 is reserved for use as the scratch V#. Stack usage is assumed if 6123 optimisations are disabled (``-O0``), if stack objects already exist (for 6124 locals, etc.), or if there are any function calls. 6125 6126 - Otherwise, four high numbered SGPRs beginning at a four-aligned SGPR index 6127 are reserved for the tentative scratch V#. These will be used if it is 6128 determined that spilling is needed. 6129 6130 - If no use is made of the tentative scratch V#, then it is unreserved 6131 and the register count is determined ignoring it. 6132 - If use is made of the tenatative scratch V#, then its register numbers 6133 are shifted to the first four-aligned SGPR index after the highest one 6134 allocated by the register allocator, and all uses are updated. The 6135 register count includes them in the shifted location. 6136 - In either case, if the processor has the SGPR allocation bug, the 6137 tentative allocation is not shifted or unreserved in order to ensure 6138 the register count is higher to workaround the bug. 6139 6140 .. note:: 6141 6142 This approach of using a tentative scratch V# and shifting the register 6143 numbers if used avoids having to perform register allocation a second 6144 time if the tentative V# is eliminated. This is more efficient and 6145 avoids the problem that the second register allocation may perform 6146 spilling which will fail as there is no longer a scratch V#. 6147 6148When the kernel prolog code is being emitted it is known whether the scratch V# 6149described above is actually used. If it is, the prolog code must set it up by 6150copying the Private Segment Buffer to the scratch V# registers and then adding 6151the Private Segment Wavefront Offset to the queue base address in the V#. The 6152result is a V# with a base address pointing to the beginning of the wavefront 6153scratch backing memory. 6154 6155The Private Segment Buffer is always requested, but the Private Segment 6156Wavefront Offset is only requested if it is used (see 6157:ref:`amdgpu-amdhsa-initial-kernel-execution-state`). 6158 6159.. _amdgpu-amdhsa-memory-model: 6160 6161Memory Model 6162~~~~~~~~~~~~ 6163 6164This section describes the mapping of LLVM memory model onto AMDGPU machine code 6165(see :ref:`memmodel`). 6166 6167The AMDGPU backend supports the memory synchronization scopes specified in 6168:ref:`amdgpu-memory-scopes`. 6169 6170The code sequences used to implement the memory model are defined in table 6171:ref:`amdgpu-amdhsa-memory-model-code-sequences-gfx6-gfx10-table`. 6172 6173The sequences specify the order of instructions that a single thread must 6174execute. The ``s_waitcnt`` and ``buffer_wbinvl1_vol`` are defined with respect 6175to other memory instructions executed by the same thread. This allows them to be 6176moved earlier or later which can allow them to be combined with other instances 6177of the same instruction, or hoisted/sunk out of loops to improve 6178performance. Only the instructions related to the memory model are given; 6179additional ``s_waitcnt`` instructions are required to ensure registers are 6180defined before being used. These may be able to be combined with the memory 6181model ``s_waitcnt`` instructions as described above. 6182 6183The AMDGPU backend supports the following memory models: 6184 6185 HSA Memory Model [HSA]_ 6186 The HSA memory model uses a single happens-before relation for all address 6187 spaces (see :ref:`amdgpu-address-spaces`). 6188 OpenCL Memory Model [OpenCL]_ 6189 The OpenCL memory model which has separate happens-before relations for the 6190 global and local address spaces. Only a fence specifying both global and 6191 local address space, and seq_cst instructions join the relationships. Since 6192 the LLVM ``memfence`` instruction does not allow an address space to be 6193 specified the OpenCL fence has to conservatively assume both local and 6194 global address space was specified. However, optimizations can often be 6195 done to eliminate the additional ``s_waitcnt`` instructions when there are 6196 no intervening memory instructions which access the corresponding address 6197 space. The code sequences in the table indicate what can be omitted for the 6198 OpenCL memory. The target triple environment is used to determine if the 6199 source language is OpenCL (see :ref:`amdgpu-opencl`). 6200 6201``ds/flat_load/store/atomic`` instructions to local memory are termed LDS 6202operations. 6203 6204``buffer/global/flat_load/store/atomic`` instructions to global memory are 6205termed vector memory operations. 6206 6207For GFX6-GFX9: 6208 6209* Each agent has multiple shader arrays (SA). 6210* Each SA has multiple compute units (CU). 6211* Each CU has multiple SIMDs that execute wavefronts. 6212* The wavefronts for a single work-group are executed in the same CU but may be 6213 executed by different SIMDs. 6214* Each CU has a single LDS memory shared by the wavefronts of the work-groups 6215 executing on it. 6216* All LDS operations of a CU are performed as wavefront wide operations in a 6217 global order and involve no caching. Completion is reported to a wavefront in 6218 execution order. 6219* The LDS memory has multiple request queues shared by the SIMDs of a 6220 CU. Therefore, the LDS operations performed by different wavefronts of a 6221 work-group can be reordered relative to each other, which can result in 6222 reordering the visibility of vector memory operations with respect to LDS 6223 operations of other wavefronts in the same work-group. A ``s_waitcnt 6224 lgkmcnt(0)`` is required to ensure synchronization between LDS operations and 6225 vector memory operations between wavefronts of a work-group, but not between 6226 operations performed by the same wavefront. 6227* The vector memory operations are performed as wavefront wide operations and 6228 completion is reported to a wavefront in execution order. The exception is 6229 that for GFX7-GFX9 ``flat_load/store/atomic`` instructions can report out of 6230 vector memory order if they access LDS memory, and out of LDS operation order 6231 if they access global memory. 6232* The vector memory operations access a single vector L1 cache shared by all 6233 SIMDs a CU. Therefore, no special action is required for coherence between the 6234 lanes of a single wavefront, or for coherence between wavefronts in the same 6235 work-group. A ``buffer_wbinvl1_vol`` is required for coherence between 6236 wavefronts executing in different work-groups as they may be executing on 6237 different CUs. 6238* The scalar memory operations access a scalar L1 cache shared by all wavefronts 6239 on a group of CUs. The scalar and vector L1 caches are not coherent. However, 6240 scalar operations are used in a restricted way so do not impact the memory 6241 model. See :ref:`amdgpu-address-spaces`. 6242* The vector and scalar memory operations use an L2 cache shared by all CUs on 6243 the same agent. 6244* The L2 cache has independent channels to service disjoint ranges of virtual 6245 addresses. 6246* Each CU has a separate request queue per channel. Therefore, the vector and 6247 scalar memory operations performed by wavefronts executing in different 6248 work-groups (which may be executing on different CUs) of an agent can be 6249 reordered relative to each other. A ``s_waitcnt vmcnt(0)`` is required to 6250 ensure synchronization between vector memory operations of different CUs. It 6251 ensures a previous vector memory operation has completed before executing a 6252 subsequent vector memory or LDS operation and so can be used to meet the 6253 requirements of acquire and release. 6254* The L2 cache can be kept coherent with other agents on some targets, or ranges 6255 of virtual addresses can be set up to bypass it to ensure system coherence. 6256 6257For GFX10: 6258 6259* Each agent has multiple shader arrays (SA). 6260* Each SA has multiple work-group processors (WGP). 6261* Each WGP has multiple compute units (CU). 6262* Each CU has multiple SIMDs that execute wavefronts. 6263* The wavefronts for a single work-group are executed in the same 6264 WGP. In CU wavefront execution mode the wavefronts may be executed by 6265 different SIMDs in the same CU. In WGP wavefront execution mode the 6266 wavefronts may be executed by different SIMDs in different CUs in the same 6267 WGP. 6268* Each WGP has a single LDS memory shared by the wavefronts of the work-groups 6269 executing on it. 6270* All LDS operations of a WGP are performed as wavefront wide operations in a 6271 global order and involve no caching. Completion is reported to a wavefront in 6272 execution order. 6273* The LDS memory has multiple request queues shared by the SIMDs of a 6274 WGP. Therefore, the LDS operations performed by different wavefronts of a 6275 work-group can be reordered relative to each other, which can result in 6276 reordering the visibility of vector memory operations with respect to LDS 6277 operations of other wavefronts in the same work-group. A ``s_waitcnt 6278 lgkmcnt(0)`` is required to ensure synchronization between LDS operations and 6279 vector memory operations between wavefronts of a work-group, but not between 6280 operations performed by the same wavefront. 6281* The vector memory operations are performed as wavefront wide operations. 6282 Completion of load/store/sample operations are reported to a wavefront in 6283 execution order of other load/store/sample operations performed by that 6284 wavefront. 6285* The vector memory operations access a vector L0 cache. There is a single L0 6286 cache per CU. Each SIMD of a CU accesses the same L0 cache. Therefore, no 6287 special action is required for coherence between the lanes of a single 6288 wavefront. However, a ``BUFFER_GL0_INV`` is required for coherence between 6289 wavefronts executing in the same work-group as they may be executing on SIMDs 6290 of different CUs that access different L0s. A ``BUFFER_GL0_INV`` is also 6291 required for coherence between wavefronts executing in different work-groups 6292 as they may be executing on different WGPs. 6293* The scalar memory operations access a scalar L0 cache shared by all wavefronts 6294 on a WGP. The scalar and vector L0 caches are not coherent. However, scalar 6295 operations are used in a restricted way so do not impact the memory model. See 6296 :ref:`amdgpu-address-spaces`. 6297* The vector and scalar memory L0 caches use an L1 cache shared by all WGPs on 6298 the same SA. Therefore, no special action is required for coherence between 6299 the wavefronts of a single work-group. However, a ``BUFFER_GL1_INV`` is 6300 required for coherence between wavefronts executing in different work-groups 6301 as they may be executing on different SAs that access different L1s. 6302* The L1 caches have independent quadrants to service disjoint ranges of virtual 6303 addresses. 6304* Each L0 cache has a separate request queue per L1 quadrant. Therefore, the 6305 vector and scalar memory operations performed by different wavefronts, whether 6306 executing in the same or different work-groups (which may be executing on 6307 different CUs accessing different L0s), can be reordered relative to each 6308 other. A ``s_waitcnt vmcnt(0) & vscnt(0)`` is required to ensure 6309 synchronization between vector memory operations of different wavefronts. It 6310 ensures a previous vector memory operation has completed before executing a 6311 subsequent vector memory or LDS operation and so can be used to meet the 6312 requirements of acquire, release and sequential consistency. 6313* The L1 caches use an L2 cache shared by all SAs on the same agent. 6314* The L2 cache has independent channels to service disjoint ranges of virtual 6315 addresses. 6316* Each L1 quadrant of a single SA accesses a different L2 channel. Each L1 6317 quadrant has a separate request queue per L2 channel. Therefore, the vector 6318 and scalar memory operations performed by wavefronts executing in different 6319 work-groups (which may be executing on different SAs) of an agent can be 6320 reordered relative to each other. A ``s_waitcnt vmcnt(0) & vscnt(0)`` is 6321 required to ensure synchronization between vector memory operations of 6322 different SAs. It ensures a previous vector memory operation has completed 6323 before executing a subsequent vector memory and so can be used to meet the 6324 requirements of acquire, release and sequential consistency. 6325* The L2 cache can be kept coherent with other agents on some targets, or ranges 6326 of virtual addresses can be set up to bypass it to ensure system coherence. 6327 6328Private address space uses ``buffer_load/store`` using the scratch V# 6329(GFX6-GFX8), or ``scratch_load/store`` (GFX9-GFX10). Since only a single thread 6330is accessing the memory, atomic memory orderings are not meaningful and all 6331accesses are treated as non-atomic. 6332 6333Constant address space uses ``buffer/global_load`` instructions (or equivalent 6334scalar memory instructions). Since the constant address space contents do not 6335change during the execution of a kernel dispatch it is not legal to perform 6336stores, and atomic memory orderings are not meaningful and all access are 6337treated as non-atomic. 6338 6339A memory synchronization scope wider than work-group is not meaningful for the 6340group (LDS) address space and is treated as work-group. 6341 6342The memory model does not support the region address space which is treated as 6343non-atomic. 6344 6345Acquire memory ordering is not meaningful on store atomic instructions and is 6346treated as non-atomic. 6347 6348Release memory ordering is not meaningful on load atomic instructions and is 6349treated a non-atomic. 6350 6351Acquire-release memory ordering is not meaningful on load or store atomic 6352instructions and is treated as acquire and release respectively. 6353 6354AMDGPU backend only uses scalar memory operations to access memory that is 6355proven to not change during the execution of the kernel dispatch. This includes 6356constant address space and global address space for program scope const 6357variables. Therefore the kernel machine code does not have to maintain the 6358scalar L1 cache to ensure it is coherent with the vector L1 cache. The scalar 6359and vector L1 caches are invalidated between kernel dispatches by CP since 6360constant address space data may change between kernel dispatch executions. See 6361:ref:`amdgpu-address-spaces`. 6362 6363The one exception is if scalar writes are used to spill SGPR registers. In this 6364case the AMDGPU backend ensures the memory location used to spill is never 6365accessed by vector memory operations at the same time. If scalar writes are used 6366then a ``s_dcache_wb`` is inserted before the ``s_endpgm`` and before a function 6367return since the locations may be used for vector memory instructions by a 6368future wavefront that uses the same scratch area, or a function call that 6369creates a frame at the same address, respectively. There is no need for a 6370``s_dcache_inv`` as all scalar writes are write-before-read in the same thread. 6371 6372For GFX6-GFX9, scratch backing memory (which is used for the private address 6373space) is accessed with MTYPE NC_NV (non-coherent non-volatile). Since the 6374private address space is only accessed by a single thread, and is always 6375write-before-read, there is never a need to invalidate these entries from the L1 6376cache. Hence all cache invalidates are done as ``*_vol`` to only invalidate the 6377volatile cache lines. 6378 6379For GFX10, scratch backing memory (which is used for the private address space) 6380is accessed with MTYPE NC (non-coherent). Since the private address space is 6381only accessed by a single thread, and is always write-before-read, there is 6382never a need to invalidate these entries from the L0 or L1 caches. 6383 6384For GFX10, wavefronts are executed in native mode with in-order reporting of 6385loads and sample instructions. In this mode vmcnt reports completion of load, 6386atomic with return and sample instructions in order, and the vscnt reports the 6387completion of store and atomic without return in order. See ``MEM_ORDERED`` 6388field in :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`. 6389 6390In GFX10, wavefronts can be executed in WGP or CU wavefront execution mode: 6391 6392* In WGP wavefront execution mode the wavefronts of a work-group are executed 6393 on the SIMDs of both CUs of the WGP. Therefore, explicit management of the per 6394 CU L0 caches is required for work-group synchronization. Also accesses to L1 6395 at work-group scope need to be explicitly ordered as the accesses from 6396 different CUs are not ordered. 6397* In CU wavefront execution mode the wavefronts of a work-group are executed on 6398 the SIMDs of a single CU of the WGP. Therefore, all global memory access by 6399 the work-group access the same L0 which in turn ensures L1 accesses are 6400 ordered and so do not require explicit management of the caches for 6401 work-group synchronization. 6402 6403See ``WGP_MODE`` field in 6404:ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table` and 6405:ref:`amdgpu-target-features`. 6406 6407On dGPU the kernarg backing memory is accessed as UC (uncached) to avoid needing 6408to invalidate the L2 cache. For GFX6-GFX9, this also causes it to be treated as 6409non-volatile and so is not invalidated by ``*_vol``. On APU it is accessed as CC 6410(cache coherent) and so the L2 cache will be coherent with the CPU and other 6411agents. 6412 6413 .. table:: AMDHSA Memory Model Code Sequences GFX6-GFX10 6414 :name: amdgpu-amdhsa-memory-model-code-sequences-gfx6-gfx10-table 6415 6416 ============ ============ ============== ========== =============================== ================================== 6417 LLVM Instr LLVM Memory LLVM Memory AMDGPU AMDGPU Machine Code AMDGPU Machine Code 6418 Ordering Sync Scope Address GFX6-9 GFX10 6419 Space 6420 ============ ============ ============== ========== =============================== ================================== 6421 **Non-Atomic** 6422 ---------------------------------------------------------------------------------------------------------------------- 6423 load *none* *none* - global - !volatile & !nontemporal - !volatile & !nontemporal 6424 - generic 6425 - private 1. buffer/global/flat_load 1. buffer/global/flat_load 6426 - constant 6427 - volatile & !nontemporal - volatile & !nontemporal 6428 6429 1. buffer/global/flat_load 1. buffer/global/flat_load 6430 glc=1 glc=1 dlc=1 6431 6432 - nontemporal - nontemporal 6433 6434 1. buffer/global/flat_load 1. buffer/global/flat_load 6435 glc=1 slc=1 slc=1 6436 6437 load *none* *none* - local 1. ds_load 1. ds_load 6438 store *none* *none* - global - !nontemporal - !nontemporal 6439 - generic 6440 - private 1. buffer/global/flat_store 1. buffer/global/flat_store 6441 - constant 6442 - nontemporal - nontemporal 6443 6444 1. buffer/global/flat_store 1. buffer/global/flat_store 6445 glc=1 slc=1 slc=1 6446 6447 store *none* *none* - local 1. ds_store 1. ds_store 6448 **Unordered Atomic** 6449 ---------------------------------------------------------------------------------------------------------------------- 6450 load atomic unordered *any* *any* *Same as non-atomic*. *Same as non-atomic*. 6451 store atomic unordered *any* *any* *Same as non-atomic*. *Same as non-atomic*. 6452 atomicrmw unordered *any* *any* *Same as monotonic *Same as monotonic 6453 atomic*. atomic*. 6454 **Monotonic Atomic** 6455 ---------------------------------------------------------------------------------------------------------------------- 6456 load atomic monotonic - singlethread - global 1. buffer/global/flat_load 1. buffer/global/flat_load 6457 - wavefront - generic 6458 load atomic monotonic - workgroup - global 1. buffer/global/flat_load 1. buffer/global/flat_load 6459 - generic glc=1 6460 6461 - If CU wavefront execution mode, omit glc=1. 6462 6463 load atomic monotonic - singlethread - local 1. ds_load 1. ds_load 6464 - wavefront 6465 - workgroup 6466 load atomic monotonic - agent - global 1. buffer/global/flat_load 1. buffer/global/flat_load 6467 - system - generic glc=1 glc=1 dlc=1 6468 store atomic monotonic - singlethread - global 1. buffer/global/flat_store 1. buffer/global/flat_store 6469 - wavefront - generic 6470 - workgroup 6471 - agent 6472 - system 6473 store atomic monotonic - singlethread - local 1. ds_store 1. ds_store 6474 - wavefront 6475 - workgroup 6476 atomicrmw monotonic - singlethread - global 1. buffer/global/flat_atomic 1. buffer/global/flat_atomic 6477 - wavefront - generic 6478 - workgroup 6479 - agent 6480 - system 6481 atomicrmw monotonic - singlethread - local 1. ds_atomic 1. ds_atomic 6482 - wavefront 6483 - workgroup 6484 **Acquire Atomic** 6485 ---------------------------------------------------------------------------------------------------------------------- 6486 load atomic acquire - singlethread - global 1. buffer/global/ds/flat_load 1. buffer/global/ds/flat_load 6487 - wavefront - local 6488 - generic 6489 load atomic acquire - workgroup - global 1. buffer/global/flat_load 1. buffer/global_load glc=1 6490 6491 - If CU wavefront execution mode, omit glc=1. 6492 6493 2. s_waitcnt vmcnt(0) 6494 6495 - If CU wavefront execution mode, omit. 6496 - Must happen before 6497 the following buffer_gl0_inv 6498 and before any following 6499 global/generic 6500 load/load 6501 atomic/store/store 6502 atomic/atomicrmw. 6503 6504 3. buffer_gl0_inv 6505 6506 - If CU wavefront execution mode, omit. 6507 - Ensures that 6508 following 6509 loads will not see 6510 stale data. 6511 6512 load atomic acquire - workgroup - local 1. ds_load 1. ds_load 6513 2. s_waitcnt lgkmcnt(0) 2. s_waitcnt lgkmcnt(0) 6514 6515 - If OpenCL, omit. - If OpenCL, omit. 6516 - Must happen before - Must happen before 6517 any following the following buffer_gl0_inv 6518 global/generic and before any following 6519 load/load global/generic load/load 6520 atomic/store/store atomic/store/store 6521 atomic/atomicrmw. atomic/atomicrmw. 6522 - Ensures any - Ensures any 6523 following global following global 6524 data read is no data read is no 6525 older than the load older than the load 6526 atomic value being atomic value being 6527 acquired. acquired. 6528 6529 3. buffer_gl0_inv 6530 6531 - If CU wavefront execution mode, omit. 6532 - If OpenCL, omit. 6533 - Ensures that 6534 following 6535 loads will not see 6536 stale data. 6537 6538 load atomic acquire - workgroup - generic 1. flat_load 1. flat_load glc=1 6539 6540 - If CU wavefront execution mode, omit glc=1. 6541 6542 2. s_waitcnt lgkmcnt(0) 2. s_waitcnt lgkmcnt(0) & 6543 vmcnt(0) 6544 6545 - If CU wavefront execution mode, omit vmcnt. 6546 - If OpenCL, omit. - If OpenCL, omit 6547 lgkmcnt(0). 6548 - Must happen before - Must happen before 6549 any following the following 6550 global/generic buffer_gl0_inv and any 6551 load/load following global/generic 6552 atomic/store/store load/load 6553 atomic/atomicrmw. atomic/store/store 6554 atomic/atomicrmw. 6555 - Ensures any - Ensures any 6556 following global following global 6557 data read is no data read is no 6558 older than the load older than the load 6559 atomic value being atomic value being 6560 acquired. acquired. 6561 6562 3. buffer_gl0_inv 6563 6564 - If CU wavefront execution mode, omit. 6565 - Ensures that 6566 following 6567 loads will not see 6568 stale data. 6569 6570 load atomic acquire - agent - global 1. buffer/global/flat_load 1. buffer/global_load 6571 - system glc=1 glc=1 dlc=1 6572 2. s_waitcnt vmcnt(0) 2. s_waitcnt vmcnt(0) 6573 6574 - Must happen before - Must happen before 6575 following following 6576 buffer_wbinvl1_vol. buffer_gl*_inv. 6577 - Ensures the load - Ensures the load 6578 has completed has completed 6579 before invalidating before invalidating 6580 the cache. the caches. 6581 6582 3. buffer_wbinvl1_vol 3. buffer_gl0_inv; 6583 buffer_gl1_inv 6584 6585 - Must happen before - Must happen before 6586 any following any following 6587 global/generic global/generic 6588 load/load load/load 6589 atomic/atomicrmw. atomic/atomicrmw. 6590 - Ensures that - Ensures that 6591 following following 6592 loads will not see loads will not see 6593 stale global data. stale global data. 6594 6595 load atomic acquire - agent - generic 1. flat_load glc=1 1. flat_load glc=1 dlc=1 6596 - system 2. s_waitcnt vmcnt(0) & 2. s_waitcnt vmcnt(0) & 6597 lgkmcnt(0) lgkmcnt(0) 6598 6599 - If OpenCL omit - If OpenCL omit 6600 lgkmcnt(0). lgkmcnt(0). 6601 - Must happen before - Must happen before 6602 following following 6603 buffer_wbinvl1_vol. buffer_gl*_invl. 6604 - Ensures the flat_load - Ensures the flat_load 6605 has completed has completed 6606 before invalidating before invalidating 6607 the cache. the caches. 6608 6609 3. buffer_wbinvl1_vol 3. buffer_gl0_inv; 6610 buffer_gl1_inv 6611 6612 - Must happen before - Must happen before 6613 any following any following 6614 global/generic global/generic 6615 load/load load/load 6616 atomic/atomicrmw. atomic/atomicrmw. 6617 - Ensures that - Ensures that 6618 following loads following loads 6619 will not see stale will not see stale 6620 global data. global data. 6621 6622 atomicrmw acquire - singlethread - global 1. buffer/global/ds/flat_atomic 1. buffer/global/ds/flat_atomic 6623 - wavefront - local 6624 - generic 6625 atomicrmw acquire - workgroup - global 1. buffer/global/flat_atomic 1. buffer/global_atomic 6626 2. s_waitcnt vm/vscnt(0) 6627 6628 - If CU wavefront execution mode, omit. 6629 - Use vmcnt if atomic with 6630 return and vscnt if atomic 6631 with no-return. 6632 - Must happen before 6633 the following buffer_gl0_inv 6634 and before any following 6635 global/generic 6636 load/load 6637 atomic/store/store 6638 atomic/atomicrmw. 6639 6640 3. buffer_gl0_inv 6641 6642 - If CU wavefront execution mode, omit. 6643 - Ensures that 6644 following 6645 loads will not see 6646 stale data. 6647 6648 atomicrmw acquire - workgroup - local 1. ds_atomic 1. ds_atomic 6649 2. waitcnt lgkmcnt(0) 2. waitcnt lgkmcnt(0) 6650 6651 - If OpenCL, omit. - If OpenCL, omit. 6652 - Must happen before - Must happen before 6653 any following the following 6654 global/generic buffer_gl0_inv. 6655 load/load 6656 atomic/store/store 6657 atomic/atomicrmw. 6658 - Ensures any - Ensures any 6659 following global following global 6660 data read is no data read is no 6661 older than the older than the 6662 atomicrmw value atomicrmw value 6663 being acquired. being acquired. 6664 6665 3. buffer_gl0_inv 6666 6667 - If OpenCL omit. 6668 - Ensures that 6669 following 6670 loads will not see 6671 stale data. 6672 6673 atomicrmw acquire - workgroup - generic 1. flat_atomic 1. flat_atomic 6674 2. waitcnt lgkmcnt(0) 2. waitcnt lgkmcnt(0) & 6675 vm/vscnt(0) 6676 6677 - If CU wavefront execution mode, omit vm/vscnt. 6678 - If OpenCL, omit. - If OpenCL, omit 6679 waitcnt lgkmcnt(0).. 6680 - Use vmcnt if atomic with 6681 return and vscnt if atomic 6682 with no-return. 6683 waitcnt lgkmcnt(0). 6684 - Must happen before - Must happen before 6685 any following the following 6686 global/generic buffer_gl0_inv. 6687 load/load 6688 atomic/store/store 6689 atomic/atomicrmw. 6690 - Ensures any - Ensures any 6691 following global following global 6692 data read is no data read is no 6693 older than the older than the 6694 atomicrmw value atomicrmw value 6695 being acquired. being acquired. 6696 6697 3. buffer_gl0_inv 6698 6699 - If CU wavefront execution mode, omit. 6700 - Ensures that 6701 following 6702 loads will not see 6703 stale data. 6704 6705 atomicrmw acquire - agent - global 1. buffer/global/flat_atomic 1. buffer/global_atomic 6706 - system 2. s_waitcnt vmcnt(0) 2. s_waitcnt vm/vscnt(0) 6707 6708 - Use vmcnt if atomic with 6709 return and vscnt if atomic 6710 with no-return. 6711 waitcnt lgkmcnt(0). 6712 - Must happen before - Must happen before 6713 following following 6714 buffer_wbinvl1_vol. buffer_gl*_inv. 6715 - Ensures the - Ensures the 6716 atomicrmw has atomicrmw has 6717 completed before completed before 6718 invalidating the invalidating the 6719 cache. caches. 6720 6721 3. buffer_wbinvl1_vol 3. buffer_gl0_inv; 6722 buffer_gl1_inv 6723 6724 - Must happen before - Must happen before 6725 any following any following 6726 global/generic global/generic 6727 load/load load/load 6728 atomic/atomicrmw. atomic/atomicrmw. 6729 - Ensures that - Ensures that 6730 following loads following loads 6731 will not see stale will not see stale 6732 global data. global data. 6733 6734 atomicrmw acquire - agent - generic 1. flat_atomic 1. flat_atomic 6735 - system 2. s_waitcnt vmcnt(0) & 2. s_waitcnt vm/vscnt(0) & 6736 lgkmcnt(0) lgkmcnt(0) 6737 6738 - If OpenCL, omit - If OpenCL, omit 6739 lgkmcnt(0). lgkmcnt(0). 6740 - Use vmcnt if atomic with 6741 return and vscnt if atomic 6742 with no-return. 6743 - Must happen before - Must happen before 6744 following following 6745 buffer_wbinvl1_vol. buffer_gl*_inv. 6746 - Ensures the - Ensures the 6747 atomicrmw has atomicrmw has 6748 completed before completed before 6749 invalidating the invalidating the 6750 cache. caches. 6751 6752 3. buffer_wbinvl1_vol 3. buffer_gl0_inv; 6753 buffer_gl1_inv 6754 6755 - Must happen before - Must happen before 6756 any following any following 6757 global/generic global/generic 6758 load/load load/load 6759 atomic/atomicrmw. atomic/atomicrmw. 6760 - Ensures that - Ensures that 6761 following loads following loads 6762 will not see stale will not see stale 6763 global data. global data. 6764 6765 fence acquire - singlethread *none* *none* *none* 6766 - wavefront 6767 fence acquire - workgroup *none* 1. s_waitcnt lgkmcnt(0) 1. s_waitcnt lgkmcnt(0) & 6768 vmcnt(0) & vscnt(0) 6769 6770 - If CU wavefront execution mode, omit vmcnt and 6771 vscnt. 6772 - If OpenCL and - If OpenCL and 6773 address space is address space is 6774 not generic, omit. not generic, omit 6775 lgkmcnt(0). 6776 - If OpenCL and 6777 address space is 6778 local, omit 6779 vmcnt(0) and vscnt(0). 6780 - However, since LLVM - However, since LLVM 6781 currently has no currently has no 6782 address space on address space on 6783 the fence need to the fence need to 6784 conservatively conservatively 6785 always generate. If always generate. If 6786 fence had an fence had an 6787 address space then address space then 6788 set to address set to address 6789 space of OpenCL space of OpenCL 6790 fence flag, or to fence flag, or to 6791 generic if both generic if both 6792 local and global local and global 6793 flags are flags are 6794 specified. specified. 6795 - Must happen after 6796 any preceding 6797 local/generic load 6798 atomic/atomicrmw 6799 with an equal or 6800 wider sync scope 6801 and memory ordering 6802 stronger than 6803 unordered (this is 6804 termed the 6805 fence-paired-atomic). 6806 - Must happen before 6807 any following 6808 global/generic 6809 load/load 6810 atomic/store/store 6811 atomic/atomicrmw. 6812 - Ensures any 6813 following global 6814 data read is no 6815 older than the 6816 value read by the 6817 fence-paired-atomic. 6818 - Could be split into 6819 separate s_waitcnt 6820 vmcnt(0), s_waitcnt 6821 vscnt(0) and s_waitcnt 6822 lgkmcnt(0) to allow 6823 them to be 6824 independently moved 6825 according to the 6826 following rules. 6827 - s_waitcnt vmcnt(0) 6828 must happen after 6829 any preceding 6830 global/generic load 6831 atomic/ 6832 atomicrmw-with-return-value 6833 with an equal or 6834 wider sync scope 6835 and memory ordering 6836 stronger than 6837 unordered (this is 6838 termed the 6839 fence-paired-atomic). 6840 - s_waitcnt vscnt(0) 6841 must happen after 6842 any preceding 6843 global/generic 6844 atomicrmw-no-return-value 6845 with an equal or 6846 wider sync scope 6847 and memory ordering 6848 stronger than 6849 unordered (this is 6850 termed the 6851 fence-paired-atomic). 6852 - s_waitcnt lgkmcnt(0) 6853 must happen after 6854 any preceding 6855 local/generic load 6856 atomic/atomicrmw 6857 with an equal or 6858 wider sync scope 6859 and memory ordering 6860 stronger than 6861 unordered (this is 6862 termed the 6863 fence-paired-atomic). 6864 - Must happen before 6865 the following 6866 buffer_gl0_inv. 6867 - Ensures that the 6868 fence-paired atomic 6869 has completed 6870 before invalidating 6871 the 6872 cache. Therefore 6873 any following 6874 locations read must 6875 be no older than 6876 the value read by 6877 the 6878 fence-paired-atomic. 6879 6880 3. buffer_gl0_inv 6881 6882 - If CU wavefront execution mode, omit. 6883 - Ensures that 6884 following 6885 loads will not see 6886 stale data. 6887 6888 fence acquire - agent *none* 1. s_waitcnt lgkmcnt(0) & 1. s_waitcnt lgkmcnt(0) & 6889 - system vmcnt(0) vmcnt(0) & vscnt(0) 6890 6891 - If OpenCL and - If OpenCL and 6892 address space is address space is 6893 not generic, omit not generic, omit 6894 lgkmcnt(0). lgkmcnt(0). 6895 - If OpenCL and 6896 address space is 6897 local, omit 6898 vmcnt(0) and vscnt(0). 6899 - However, since LLVM - However, since LLVM 6900 currently has no currently has no 6901 address space on address space on 6902 the fence need to the fence need to 6903 conservatively conservatively 6904 always generate always generate 6905 (see comment for (see comment for 6906 previous fence). previous fence). 6907 - Could be split into 6908 separate s_waitcnt 6909 vmcnt(0) and 6910 s_waitcnt 6911 lgkmcnt(0) to allow 6912 them to be 6913 independently moved 6914 according to the 6915 following rules. 6916 - s_waitcnt vmcnt(0) 6917 must happen after 6918 any preceding 6919 global/generic load 6920 atomic/atomicrmw 6921 with an equal or 6922 wider sync scope 6923 and memory ordering 6924 stronger than 6925 unordered (this is 6926 termed the 6927 fence-paired-atomic). 6928 - s_waitcnt lgkmcnt(0) 6929 must happen after 6930 any preceding 6931 local/generic load 6932 atomic/atomicrmw 6933 with an equal or 6934 wider sync scope 6935 and memory ordering 6936 stronger than 6937 unordered (this is 6938 termed the 6939 fence-paired-atomic). 6940 - Must happen before 6941 the following 6942 buffer_wbinvl1_vol. 6943 - Ensures that the 6944 fence-paired atomic 6945 has completed 6946 before invalidating 6947 the 6948 cache. Therefore 6949 any following 6950 locations read must 6951 be no older than 6952 the value read by 6953 the 6954 fence-paired-atomic. 6955 - Could be split into 6956 separate s_waitcnt 6957 vmcnt(0), s_waitcnt 6958 vscnt(0) and s_waitcnt 6959 lgkmcnt(0) to allow 6960 them to be 6961 independently moved 6962 according to the 6963 following rules. 6964 - s_waitcnt vmcnt(0) 6965 must happen after 6966 any preceding 6967 global/generic load 6968 atomic/ 6969 atomicrmw-with-return-value 6970 with an equal or 6971 wider sync scope 6972 and memory ordering 6973 stronger than 6974 unordered (this is 6975 termed the 6976 fence-paired-atomic). 6977 - s_waitcnt vscnt(0) 6978 must happen after 6979 any preceding 6980 global/generic 6981 atomicrmw-no-return-value 6982 with an equal or 6983 wider sync scope 6984 and memory ordering 6985 stronger than 6986 unordered (this is 6987 termed the 6988 fence-paired-atomic). 6989 - s_waitcnt lgkmcnt(0) 6990 must happen after 6991 any preceding 6992 local/generic load 6993 atomic/atomicrmw 6994 with an equal or 6995 wider sync scope 6996 and memory ordering 6997 stronger than 6998 unordered (this is 6999 termed the 7000 fence-paired-atomic). 7001 - Must happen before 7002 the following 7003 buffer_gl*_inv. 7004 - Ensures that the 7005 fence-paired atomic 7006 has completed 7007 before invalidating 7008 the 7009 caches. Therefore 7010 any following 7011 locations read must 7012 be no older than 7013 the value read by 7014 the 7015 fence-paired-atomic. 7016 7017 2. buffer_wbinvl1_vol 2. buffer_gl0_inv; 7018 buffer_gl1_inv 7019 7020 - Must happen before any - Must happen before any 7021 following global/generic following global/generic 7022 load/load load/load 7023 atomic/store/store atomic/store/store 7024 atomic/atomicrmw. atomic/atomicrmw. 7025 - Ensures that - Ensures that 7026 following loads following loads 7027 will not see stale will not see stale 7028 global data. global data. 7029 7030 **Release Atomic** 7031 ---------------------------------------------------------------------------------------------------------------------- 7032 store atomic release - singlethread - global 1. buffer/global/ds/flat_store 1. buffer/global/ds/flat_store 7033 - wavefront - local 7034 - generic 7035 store atomic release - workgroup - global 1. s_waitcnt lgkmcnt(0) 1. s_waitcnt lgkmcnt(0) & 7036 vmcnt(0) & vscnt(0) 7037 7038 - If CU wavefront execution mode, omit vmcnt and 7039 vscnt. 7040 - If OpenCL, omit. - If OpenCL, omit 7041 lgkmcnt(0). 7042 - Must happen after 7043 any preceding 7044 local/generic 7045 load/store/load 7046 atomic/store 7047 atomic/atomicrmw. 7048 - Could be split into 7049 separate s_waitcnt 7050 vmcnt(0), s_waitcnt 7051 vscnt(0) and s_waitcnt 7052 lgkmcnt(0) to allow 7053 them to be 7054 independently moved 7055 according to the 7056 following rules. 7057 - s_waitcnt vmcnt(0) 7058 must happen after 7059 any preceding 7060 global/generic load/load 7061 atomic/ 7062 atomicrmw-with-return-value. 7063 - s_waitcnt vscnt(0) 7064 must happen after 7065 any preceding 7066 global/generic 7067 store/store 7068 atomic/ 7069 atomicrmw-no-return-value. 7070 - s_waitcnt lgkmcnt(0) 7071 must happen after 7072 any preceding 7073 local/generic 7074 load/store/load 7075 atomic/store 7076 atomic/atomicrmw. 7077 - Must happen before - Must happen before 7078 the following the following 7079 store. store. 7080 - Ensures that all - Ensures that all 7081 memory operations memory operations 7082 to local have have 7083 completed before completed before 7084 performing the performing the 7085 store that is being store that is being 7086 released. released. 7087 7088 2. buffer/global/flat_store 2. buffer/global_store 7089 store atomic release - workgroup - local 1. waitcnt vmcnt(0) & vscnt(0) 7090 7091 - If CU wavefront execution mode, omit. 7092 - If OpenCL, omit. 7093 - Could be split into 7094 separate s_waitcnt 7095 vmcnt(0) and s_waitcnt 7096 vscnt(0) to allow 7097 them to be 7098 independently moved 7099 according to the 7100 following rules. 7101 - s_waitcnt vmcnt(0) 7102 must happen after 7103 any preceding 7104 global/generic load/load 7105 atomic/ 7106 atomicrmw-with-return-value. 7107 - s_waitcnt vscnt(0) 7108 must happen after 7109 any preceding 7110 global/generic 7111 store/store atomic/ 7112 atomicrmw-no-return-value. 7113 - Must happen before 7114 the following 7115 store. 7116 - Ensures that all 7117 global memory 7118 operations have 7119 completed before 7120 performing the 7121 store that is being 7122 released. 7123 7124 1. ds_store 2. ds_store 7125 store atomic release - workgroup - generic 1. s_waitcnt lgkmcnt(0) 1. s_waitcnt lgkmcnt(0) & 7126 vmcnt(0) & vscnt(0) 7127 7128 - If CU wavefront execution mode, omit vmcnt and 7129 vscnt. 7130 - If OpenCL, omit. - If OpenCL, omit 7131 lgkmcnt(0). 7132 - Must happen after 7133 any preceding 7134 local/generic 7135 load/store/load 7136 atomic/store 7137 atomic/atomicrmw. 7138 - Could be split into 7139 separate s_waitcnt 7140 vmcnt(0), s_waitcnt 7141 vscnt(0) and s_waitcnt 7142 lgkmcnt(0) to allow 7143 them to be 7144 independently moved 7145 according to the 7146 following rules. 7147 - s_waitcnt vmcnt(0) 7148 must happen after 7149 any preceding 7150 global/generic load/load 7151 atomic/ 7152 atomicrmw-with-return-value. 7153 - s_waitcnt vscnt(0) 7154 must happen after 7155 any preceding 7156 global/generic 7157 store/store 7158 atomic/ 7159 atomicrmw-no-return-value. 7160 - s_waitcnt lgkmcnt(0) 7161 must happen after 7162 any preceding 7163 local/generic load/store/load 7164 atomic/store atomic/atomicrmw. 7165 - Must happen before - Must happen before 7166 the following the following 7167 store. store. 7168 - Ensures that all - Ensures that all 7169 memory operations memory operations 7170 to local have have 7171 completed before completed before 7172 performing the performing the 7173 store that is being store that is being 7174 released. released. 7175 7176 2. flat_store 2. flat_store 7177 store atomic release - agent - global 1. s_waitcnt lgkmcnt(0) & 1. s_waitcnt lgkmcnt(0) & 7178 - system - generic vmcnt(0) vmcnt(0) & vscnt(0) 7179 7180 - If OpenCL, omit - If OpenCL, omit 7181 lgkmcnt(0). lgkmcnt(0). 7182 - Could be split into - Could be split into 7183 separate s_waitcnt separate s_waitcnt 7184 vmcnt(0) and vmcnt(0), s_waitcnt vscnt(0) 7185 s_waitcnt and s_waitcnt 7186 lgkmcnt(0) to allow lgkmcnt(0) to allow 7187 them to be them to be 7188 independently moved independently moved 7189 according to the according to the 7190 following rules. following rules. 7191 - s_waitcnt vmcnt(0) - s_waitcnt vmcnt(0) 7192 must happen after must happen after 7193 any preceding any preceding 7194 global/generic global/generic 7195 load/store/load load/load 7196 atomic/store atomic/ 7197 atomic/atomicrmw. atomicrmw-with-return-value. 7198 - s_waitcnt vscnt(0) 7199 must happen after 7200 any preceding 7201 global/generic 7202 store/store atomic/ 7203 atomicrmw-no-return-value. 7204 - s_waitcnt lgkmcnt(0) - s_waitcnt lgkmcnt(0) 7205 must happen after must happen after 7206 any preceding any preceding 7207 local/generic local/generic 7208 load/store/load load/store/load 7209 atomic/store atomic/store 7210 atomic/atomicrmw. atomic/atomicrmw. 7211 - Must happen before - Must happen before 7212 the following the following 7213 store. store. 7214 - Ensures that all - Ensures that all 7215 memory operations memory operations 7216 to memory have to memory have 7217 completed before completed before 7218 performing the performing the 7219 store that is being store that is being 7220 released. released. 7221 7222 2. buffer/global/ds/flat_store 2. buffer/global/ds/flat_store 7223 atomicrmw release - singlethread - global 1. buffer/global/ds/flat_atomic 1. buffer/global/ds/flat_atomic 7224 - wavefront - local 7225 - generic 7226 atomicrmw release - workgroup - global 1. s_waitcnt lgkmcnt(0) 1. s_waitcnt lgkmcnt(0) & 7227 vmcnt(0) & vscnt(0) 7228 7229 - If CU wavefront execution mode, omit vmcnt and 7230 vscnt. 7231 - If OpenCL, omit. 7232 7233 - Must happen after 7234 any preceding 7235 local/generic 7236 load/store/load 7237 atomic/store 7238 atomic/atomicrmw. 7239 - Could be split into 7240 separate s_waitcnt 7241 vmcnt(0), s_waitcnt 7242 vscnt(0) and s_waitcnt 7243 lgkmcnt(0) to allow 7244 them to be 7245 independently moved 7246 according to the 7247 following rules. 7248 - s_waitcnt vmcnt(0) 7249 must happen after 7250 any preceding 7251 global/generic load/load 7252 atomic/ 7253 atomicrmw-with-return-value. 7254 - s_waitcnt vscnt(0) 7255 must happen after 7256 any preceding 7257 global/generic 7258 store/store 7259 atomic/ 7260 atomicrmw-no-return-value. 7261 - s_waitcnt lgkmcnt(0) 7262 must happen after 7263 any preceding 7264 local/generic 7265 load/store/load 7266 atomic/store 7267 atomic/atomicrmw. 7268 - Must happen before - Must happen before 7269 the following the following 7270 atomicrmw. atomicrmw. 7271 - Ensures that all - Ensures that all 7272 memory operations memory operations 7273 to local have have 7274 completed before completed before 7275 performing the performing the 7276 atomicrmw that is atomicrmw that is 7277 being released. being released. 7278 7279 2. buffer/global/flat_atomic 2. buffer/global_atomic 7280 atomicrmw release - workgroup - local 1. waitcnt vmcnt(0) & vscnt(0) 7281 7282 - If CU wavefront execution mode, omit. 7283 - If OpenCL, omit. 7284 - Could be split into 7285 separate s_waitcnt 7286 vmcnt(0) and s_waitcnt 7287 vscnt(0) to allow 7288 them to be 7289 independently moved 7290 according to the 7291 following rules. 7292 - s_waitcnt vmcnt(0) 7293 must happen after 7294 any preceding 7295 global/generic load/load 7296 atomic/ 7297 atomicrmw-with-return-value. 7298 - s_waitcnt vscnt(0) 7299 must happen after 7300 any preceding 7301 global/generic 7302 store/store atomic/ 7303 atomicrmw-no-return-value. 7304 - Must happen before 7305 the following 7306 store. 7307 - Ensures that all 7308 global memory 7309 operations have 7310 completed before 7311 performing the 7312 store that is being 7313 released. 7314 7315 1. ds_atomic 2. ds_atomic 7316 atomicrmw release - workgroup - generic 1. s_waitcnt lgkmcnt(0) 1. s_waitcnt lgkmcnt(0) & 7317 vmcnt(0) & vscnt(0) 7318 7319 - If CU wavefront execution mode, omit vmcnt and 7320 vscnt. 7321 - If OpenCL, omit. - If OpenCL, omit 7322 waitcnt lgkmcnt(0). 7323 - Must happen after 7324 any preceding 7325 local/generic 7326 load/store/load 7327 atomic/store 7328 atomic/atomicrmw. 7329 - Could be split into 7330 separate s_waitcnt 7331 vmcnt(0), s_waitcnt 7332 vscnt(0) and s_waitcnt 7333 lgkmcnt(0) to allow 7334 them to be 7335 independently moved 7336 according to the 7337 following rules. 7338 - s_waitcnt vmcnt(0) 7339 must happen after 7340 any preceding 7341 global/generic load/load 7342 atomic/ 7343 atomicrmw-with-return-value. 7344 - s_waitcnt vscnt(0) 7345 must happen after 7346 any preceding 7347 global/generic 7348 store/store 7349 atomic/ 7350 atomicrmw-no-return-value. 7351 - s_waitcnt lgkmcnt(0) 7352 must happen after 7353 any preceding 7354 local/generic load/store/load 7355 atomic/store atomic/atomicrmw. 7356 - Must happen before - Must happen before 7357 the following the following 7358 atomicrmw. atomicrmw. 7359 - Ensures that all - Ensures that all 7360 memory operations memory operations 7361 to local have have 7362 completed before completed before 7363 performing the performing the 7364 atomicrmw that is atomicrmw that is 7365 being released. being released. 7366 7367 2. flat_atomic 2. flat_atomic 7368 atomicrmw release - agent - global 1. s_waitcnt lgkmcnt(0) & 1. s_waitcnt lkkmcnt(0) & 7369 - system - generic vmcnt(0) vmcnt(0) & vscnt(0) 7370 7371 - If OpenCL, omit - If OpenCL, omit 7372 lgkmcnt(0). lgkmcnt(0). 7373 - Could be split into - Could be split into 7374 separate s_waitcnt separate s_waitcnt 7375 vmcnt(0) and vmcnt(0), s_waitcnt 7376 s_waitcnt vscnt(0) and s_waitcnt 7377 lgkmcnt(0) to allow lgkmcnt(0) to allow 7378 them to be them to be 7379 independently moved independently moved 7380 according to the according to the 7381 following rules. following rules. 7382 - s_waitcnt vmcnt(0) - s_waitcnt vmcnt(0) 7383 must happen after must happen after 7384 any preceding any preceding 7385 global/generic global/generic 7386 load/store/load load/load atomic/ 7387 atomic/store atomicrmw-with-return-value. 7388 atomic/atomicrmw. 7389 - s_waitcnt vscnt(0) 7390 must happen after 7391 any preceding 7392 global/generic 7393 store/store atomic/ 7394 atomicrmw-no-return-value. 7395 - s_waitcnt lgkmcnt(0) - s_waitcnt lgkmcnt(0) 7396 must happen after must happen after 7397 any preceding any preceding 7398 local/generic local/generic 7399 load/store/load load/store/load 7400 atomic/store atomic/store 7401 atomic/atomicrmw. atomic/atomicrmw. 7402 - Must happen before - Must happen before 7403 the following the following 7404 atomicrmw. atomicrmw. 7405 - Ensures that all - Ensures that all 7406 memory operations memory operations 7407 to global and local to global and local 7408 have completed have completed 7409 before performing before performing 7410 the atomicrmw that the atomicrmw that 7411 is being released. is being released. 7412 7413 2. buffer/global/ds/flat_atomic 2. buffer/global/ds/flat_atomic 7414 fence release - singlethread *none* *none* *none* 7415 - wavefront 7416 fence release - workgroup *none* 1. s_waitcnt lgkmcnt(0) 1. s_waitcnt lgkmcnt(0) & 7417 vmcnt(0) & vscnt(0) 7418 7419 - If CU wavefront execution mode, omit vmcnt and 7420 vscnt. 7421 - If OpenCL and - If OpenCL and 7422 address space is address space is 7423 not generic, omit. not generic, omit 7424 lgkmcnt(0). 7425 - If OpenCL and 7426 address space is 7427 local, omit 7428 vmcnt(0) and vscnt(0). 7429 - However, since LLVM - However, since LLVM 7430 currently has no currently has no 7431 address space on address space on 7432 the fence need to the fence need to 7433 conservatively conservatively 7434 always generate. If always generate. If 7435 fence had an fence had an 7436 address space then address space then 7437 set to address set to address 7438 space of OpenCL space of OpenCL 7439 fence flag, or to fence flag, or to 7440 generic if both generic if both 7441 local and global local and global 7442 flags are flags are 7443 specified. specified. 7444 - Must happen after 7445 any preceding 7446 local/generic 7447 load/load 7448 atomic/store/store 7449 atomic/atomicrmw. 7450 - Could be split into 7451 separate s_waitcnt 7452 vmcnt(0), s_waitcnt 7453 vscnt(0) and s_waitcnt 7454 lgkmcnt(0) to allow 7455 them to be 7456 independently moved 7457 according to the 7458 following rules. 7459 - s_waitcnt vmcnt(0) 7460 must happen after 7461 any preceding 7462 global/generic 7463 load/load 7464 atomic/ 7465 atomicrmw-with-return-value. 7466 - s_waitcnt vscnt(0) 7467 must happen after 7468 any preceding 7469 global/generic 7470 store/store atomic/ 7471 atomicrmw-no-return-value. 7472 - s_waitcnt lgkmcnt(0) 7473 must happen after 7474 any preceding 7475 local/generic 7476 load/store/load 7477 atomic/store atomic/ 7478 atomicrmw. 7479 - Must happen before - Must happen before 7480 any following store any following store 7481 atomic/atomicrmw atomic/atomicrmw 7482 with an equal or with an equal or 7483 wider sync scope wider sync scope 7484 and memory ordering and memory ordering 7485 stronger than stronger than 7486 unordered (this is unordered (this is 7487 termed the termed the 7488 fence-paired-atomic). fence-paired-atomic). 7489 - Ensures that all - Ensures that all 7490 memory operations memory operations 7491 to local have have 7492 completed before completed before 7493 performing the performing the 7494 following following 7495 fence-paired-atomic. fence-paired-atomic. 7496 7497 fence release - agent *none* 1. s_waitcnt lgkmcnt(0) & 1. s_waitcnt lgkmcnt(0) & 7498 - system vmcnt(0) vmcnt(0) & vscnt(0) 7499 7500 - If OpenCL and - If OpenCL and 7501 address space is address space is 7502 not generic, omit not generic, omit 7503 lgkmcnt(0). lgkmcnt(0). 7504 - If OpenCL and - If OpenCL and 7505 address space is address space is 7506 local, omit local, omit 7507 vmcnt(0). vmcnt(0) and vscnt(0). 7508 - However, since LLVM - However, since LLVM 7509 currently has no currently has no 7510 address space on address space on 7511 the fence need to the fence need to 7512 conservatively conservatively 7513 always generate. If always generate. If 7514 fence had an fence had an 7515 address space then address space then 7516 set to address set to address 7517 space of OpenCL space of OpenCL 7518 fence flag, or to fence flag, or to 7519 generic if both generic if both 7520 local and global local and global 7521 flags are flags are 7522 specified. specified. 7523 - Could be split into - Could be split into 7524 separate s_waitcnt separate s_waitcnt 7525 vmcnt(0) and vmcnt(0), s_waitcnt 7526 s_waitcnt vscnt(0) and s_waitcnt 7527 lgkmcnt(0) to allow lgkmcnt(0) to allow 7528 them to be them to be 7529 independently moved independently moved 7530 according to the according to the 7531 following rules. following rules. 7532 - s_waitcnt vmcnt(0) - s_waitcnt vmcnt(0) 7533 must happen after must happen after 7534 any preceding any preceding 7535 global/generic global/generic 7536 load/store/load load/load atomic/ 7537 atomic/store atomicrmw-with-return-value. 7538 atomic/atomicrmw. 7539 - s_waitcnt vscnt(0) 7540 must happen after 7541 any preceding 7542 global/generic 7543 store/store atomic/ 7544 atomicrmw-no-return-value. 7545 - s_waitcnt lgkmcnt(0) - s_waitcnt lgkmcnt(0) 7546 must happen after must happen after 7547 any preceding any preceding 7548 local/generic local/generic 7549 load/store/load load/store/load 7550 atomic/store atomic/store 7551 atomic/atomicrmw. atomic/atomicrmw. 7552 - Must happen before - Must happen before 7553 any following store any following store 7554 atomic/atomicrmw atomic/atomicrmw 7555 with an equal or with an equal or 7556 wider sync scope wider sync scope 7557 and memory ordering and memory ordering 7558 stronger than stronger than 7559 unordered (this is unordered (this is 7560 termed the termed the 7561 fence-paired-atomic). fence-paired-atomic). 7562 - Ensures that all - Ensures that all 7563 memory operations memory operations 7564 have have 7565 completed before completed before 7566 performing the performing the 7567 following following 7568 fence-paired-atomic. fence-paired-atomic. 7569 7570 **Acquire-Release Atomic** 7571 ---------------------------------------------------------------------------------------------------------------------- 7572 atomicrmw acq_rel - singlethread - global 1. buffer/global/ds/flat_atomic 1. buffer/global/ds/flat_atomic 7573 - wavefront - local 7574 - generic 7575 atomicrmw acq_rel - workgroup - global 1. s_waitcnt lgkmcnt(0) 1. s_waitcnt lgkmcnt(0) & 7576 vmcnt(0) & vscnt(0) 7577 7578 - If CU wavefront execution mode, omit vmcnt and 7579 vscnt. 7580 - If OpenCL, omit. - If OpenCL, omit 7581 s_waitcnt lgkmcnt(0). 7582 - Must happen after - Must happen after 7583 any preceding any preceding 7584 local/generic local/generic 7585 load/store/load load/store/load 7586 atomic/store atomic/store 7587 atomic/atomicrmw. atomic/atomicrmw. 7588 - Could be split into 7589 separate s_waitcnt 7590 vmcnt(0), s_waitcnt 7591 vscnt(0) and s_waitcnt 7592 lgkmcnt(0) to allow 7593 them to be 7594 independently moved 7595 according to the 7596 following rules. 7597 - s_waitcnt vmcnt(0) 7598 must happen after 7599 any preceding 7600 global/generic load/load 7601 atomic/ 7602 atomicrmw-with-return-value. 7603 - s_waitcnt vscnt(0) 7604 must happen after 7605 any preceding 7606 global/generic 7607 store/store 7608 atomic/ 7609 atomicrmw-no-return-value. 7610 - s_waitcnt lgkmcnt(0) 7611 must happen after 7612 any preceding 7613 local/generic load/store/load 7614 atomic/store atomic/atomicrmw. 7615 - Must happen before - Must happen before 7616 the following the following 7617 atomicrmw. atomicrmw. 7618 - Ensures that all - Ensures that all 7619 memory operations memory operations 7620 to local have have 7621 completed before completed before 7622 performing the performing the 7623 atomicrmw that is atomicrmw that is 7624 being released. being released. 7625 7626 2. buffer/global/flat_atomic 2. buffer/global_atomic 7627 3. s_waitcnt vm/vscnt(0) 7628 7629 - If CU wavefront execution mode, omit vm/vscnt. 7630 - Use vmcnt if atomic with 7631 return and vscnt if atomic 7632 with no-return. 7633 waitcnt lgkmcnt(0). 7634 - Must happen before 7635 the following 7636 buffer_gl0_inv. 7637 - Ensures any 7638 following global 7639 data read is no 7640 older than the 7641 atomicrmw value 7642 being acquired. 7643 7644 4. buffer_gl0_inv 7645 7646 - If CU wavefront execution mode, omit. 7647 - Ensures that 7648 following 7649 loads will not see 7650 stale data. 7651 7652 atomicrmw acq_rel - workgroup - local 1. waitcnt vmcnt(0) & vscnt(0) 7653 7654 - If CU wavefront execution mode, omit. 7655 - If OpenCL, omit. 7656 - Could be split into 7657 separate s_waitcnt 7658 vmcnt(0) and s_waitcnt 7659 vscnt(0) to allow 7660 them to be 7661 independently moved 7662 according to the 7663 following rules. 7664 - s_waitcnt vmcnt(0) 7665 must happen after 7666 any preceding 7667 global/generic load/load 7668 atomic/ 7669 atomicrmw-with-return-value. 7670 - s_waitcnt vscnt(0) 7671 must happen after 7672 any preceding 7673 global/generic 7674 store/store atomic/ 7675 atomicrmw-no-return-value. 7676 - Must happen before 7677 the following 7678 store. 7679 - Ensures that all 7680 global memory 7681 operations have 7682 completed before 7683 performing the 7684 store that is being 7685 released. 7686 7687 1. ds_atomic 2. ds_atomic 7688 2. s_waitcnt lgkmcnt(0) 3. s_waitcnt lgkmcnt(0) 7689 7690 - If OpenCL, omit. - If OpenCL, omit. 7691 - Must happen before - Must happen before 7692 any following the following 7693 global/generic buffer_gl0_inv. 7694 load/load 7695 atomic/store/store 7696 atomic/atomicrmw. 7697 - Ensures any - Ensures any 7698 following global following global 7699 data read is no data read is no 7700 older than the load older than the load 7701 atomic value being atomic value being 7702 acquired. acquired. 7703 7704 4. buffer_gl0_inv 7705 7706 - If CU wavefront execution mode, omit. 7707 - If OpenCL omit. 7708 - Ensures that 7709 following 7710 loads will not see 7711 stale data. 7712 7713 atomicrmw acq_rel - workgroup - generic 1. s_waitcnt lgkmcnt(0) 1. s_waitcnt lgkmcnt(0) & 7714 vmcnt(0) & vscnt(0) 7715 7716 - If CU wavefront execution mode, omit vmcnt and 7717 vscnt. 7718 - If OpenCL, omit. - If OpenCL, omit 7719 waitcnt lgkmcnt(0). 7720 - Must happen after 7721 any preceding 7722 local/generic 7723 load/store/load 7724 atomic/store 7725 atomic/atomicrmw. 7726 - Could be split into 7727 separate s_waitcnt 7728 vmcnt(0), s_waitcnt 7729 vscnt(0) and s_waitcnt 7730 lgkmcnt(0) to allow 7731 them to be 7732 independently moved 7733 according to the 7734 following rules. 7735 - s_waitcnt vmcnt(0) 7736 must happen after 7737 any preceding 7738 global/generic load/load 7739 atomic/ 7740 atomicrmw-with-return-value. 7741 - s_waitcnt vscnt(0) 7742 must happen after 7743 any preceding 7744 global/generic 7745 store/store 7746 atomic/ 7747 atomicrmw-no-return-value. 7748 - s_waitcnt lgkmcnt(0) 7749 must happen after 7750 any preceding 7751 local/generic load/store/load 7752 atomic/store atomic/atomicrmw. 7753 - Must happen before - Must happen before 7754 the following the following 7755 atomicrmw. atomicrmw. 7756 - Ensures that all - Ensures that all 7757 memory operations memory operations 7758 to local have have 7759 completed before completed before 7760 performing the performing the 7761 atomicrmw that is atomicrmw that is 7762 being released. being released. 7763 7764 2. flat_atomic 2. flat_atomic 7765 3. s_waitcnt lgkmcnt(0) 3. s_waitcnt lgkmcnt(0) & 7766 vm/vscnt(0) 7767 7768 - If CU wavefront execution mode, omit vm/vscnt. 7769 - If OpenCL, omit. - If OpenCL, omit 7770 waitcnt lgkmcnt(0). 7771 - Must happen before - Must happen before 7772 any following the following 7773 global/generic buffer_gl0_inv. 7774 load/load 7775 atomic/store/store 7776 atomic/atomicrmw. 7777 - Ensures any - Ensures any 7778 following global following global 7779 data read is no data read is no 7780 older than the load older than the load 7781 atomic value being atomic value being 7782 acquired. acquired. 7783 7784 3. buffer_gl0_inv 7785 7786 - If CU wavefront execution mode, omit. 7787 - Ensures that 7788 following 7789 loads will not see 7790 stale data. 7791 7792 atomicrmw acq_rel - agent - global 1. s_waitcnt lgkmcnt(0) & 1. s_waitcnt lgkmcnt(0) & 7793 - system vmcnt(0) vmcnt(0) & vscnt(0) 7794 7795 - If OpenCL, omit - If OpenCL, omit 7796 lgkmcnt(0). lgkmcnt(0). 7797 - Could be split into - Could be split into 7798 separate s_waitcnt separate s_waitcnt 7799 vmcnt(0) and vmcnt(0), s_waitcnt 7800 s_waitcnt vscnt(0) and s_waitcnt 7801 lgkmcnt(0) to allow lgkmcnt(0) to allow 7802 them to be them to be 7803 independently moved independently moved 7804 according to the according to the 7805 following rules. following rules. 7806 - s_waitcnt vmcnt(0) - s_waitcnt vmcnt(0) 7807 must happen after must happen after 7808 any preceding any preceding 7809 global/generic global/generic 7810 load/store/load load/load atomic/ 7811 atomic/store atomicrmw-with-return-value. 7812 atomic/atomicrmw. 7813 - s_waitcnt vscnt(0) 7814 must happen after 7815 any preceding 7816 global/generic 7817 store/store atomic/ 7818 atomicrmw-no-return-value. 7819 - s_waitcnt lgkmcnt(0) - s_waitcnt lgkmcnt(0) 7820 must happen after must happen after 7821 any preceding any preceding 7822 local/generic local/generic 7823 load/store/load load/store/load 7824 atomic/store atomic/store 7825 atomic/atomicrmw. atomic/atomicrmw. 7826 - Must happen before - Must happen before 7827 the following the following 7828 atomicrmw. atomicrmw. 7829 - Ensures that all - Ensures that all 7830 memory operations memory operations 7831 to global have to global have 7832 completed before completed before 7833 performing the performing the 7834 atomicrmw that is atomicrmw that is 7835 being released. being released. 7836 7837 2. buffer/global/flat_atomic 2. buffer/global_atomic 7838 3. s_waitcnt vmcnt(0) 3. s_waitcnt vm/vscnt(0) 7839 7840 - Use vmcnt if atomic with 7841 return and vscnt if atomic 7842 with no-return. 7843 waitcnt lgkmcnt(0). 7844 - Must happen before - Must happen before 7845 following following 7846 buffer_wbinvl1_vol. buffer_gl*_inv. 7847 - Ensures the - Ensures the 7848 atomicrmw has atomicrmw has 7849 completed before completed before 7850 invalidating the invalidating the 7851 cache. caches. 7852 7853 4. buffer_wbinvl1_vol 4. buffer_gl0_inv; 7854 buffer_gl1_inv 7855 7856 - Must happen before - Must happen before 7857 any following any following 7858 global/generic global/generic 7859 load/load load/load 7860 atomic/atomicrmw. atomic/atomicrmw. 7861 - Ensures that - Ensures that 7862 following loads following loads 7863 will not see stale will not see stale 7864 global data. global data. 7865 7866 atomicrmw acq_rel - agent - generic 1. s_waitcnt lgkmcnt(0) & 1. s_waitcnt lgkmcnt(0) & 7867 - system vmcnt(0) vmcnt(0) & vscnt(0) 7868 7869 - If OpenCL, omit - If OpenCL, omit 7870 lgkmcnt(0). lgkmcnt(0). 7871 - Could be split into - Could be split into 7872 separate s_waitcnt separate s_waitcnt 7873 vmcnt(0) and vmcnt(0), s_waitcnt 7874 s_waitcnt vscnt(0) and s_waitcnt 7875 lgkmcnt(0) to allow lgkmcnt(0) to allow 7876 them to be them to be 7877 independently moved independently moved 7878 according to the according to the 7879 following rules. following rules. 7880 - s_waitcnt vmcnt(0) - s_waitcnt vmcnt(0) 7881 must happen after must happen after 7882 any preceding any preceding 7883 global/generic global/generic 7884 load/store/load load/load atomic 7885 atomic/store atomicrmw-with-return-value. 7886 atomic/atomicrmw. 7887 - s_waitcnt vscnt(0) 7888 must happen after 7889 any preceding 7890 global/generic 7891 store/store atomic/ 7892 atomicrmw-no-return-value. 7893 - s_waitcnt lgkmcnt(0) - s_waitcnt lgkmcnt(0) 7894 must happen after must happen after 7895 any preceding any preceding 7896 local/generic local/generic 7897 load/store/load load/store/load 7898 atomic/store atomic/store 7899 atomic/atomicrmw. atomic/atomicrmw. 7900 - Must happen before - Must happen before 7901 the following the following 7902 atomicrmw. atomicrmw. 7903 - Ensures that all - Ensures that all 7904 memory operations memory operations 7905 to global have have 7906 completed before completed before 7907 performing the performing the 7908 atomicrmw that is atomicrmw that is 7909 being released. being released. 7910 7911 2. flat_atomic 2. flat_atomic 7912 3. s_waitcnt vmcnt(0) & 3. s_waitcnt vm/vscnt(0) & 7913 lgkmcnt(0) lgkmcnt(0) 7914 7915 - If OpenCL, omit - If OpenCL, omit 7916 lgkmcnt(0). lgkmcnt(0). 7917 - Use vmcnt if atomic with 7918 return and vscnt if atomic 7919 with no-return. 7920 - Must happen before - Must happen before 7921 following following 7922 buffer_wbinvl1_vol. buffer_gl*_inv. 7923 - Ensures the - Ensures the 7924 atomicrmw has atomicrmw has 7925 completed before completed before 7926 invalidating the invalidating the 7927 cache. caches. 7928 7929 4. buffer_wbinvl1_vol 4. buffer_gl0_inv; 7930 buffer_gl1_inv 7931 7932 - Must happen before - Must happen before 7933 any following any following 7934 global/generic global/generic 7935 load/load load/load 7936 atomic/atomicrmw. atomic/atomicrmw. 7937 - Ensures that - Ensures that 7938 following loads following loads 7939 will not see stale will not see stale 7940 global data. global data. 7941 7942 fence acq_rel - singlethread *none* *none* *none* 7943 - wavefront 7944 fence acq_rel - workgroup *none* 1. s_waitcnt lgkmcnt(0) 1. s_waitcnt lgkmcnt(0) & 7945 vmcnt(0) & vscnt(0) 7946 7947 - If CU wavefront execution mode, omit vmcnt and 7948 vscnt. 7949 - If OpenCL and - If OpenCL and 7950 address space is address space is 7951 not generic, omit. not generic, omit 7952 lgkmcnt(0). 7953 - If OpenCL and 7954 address space is 7955 local, omit 7956 vmcnt(0) and vscnt(0). 7957 - However, - However, 7958 since LLVM since LLVM 7959 currently has no currently has no 7960 address space on address space on 7961 the fence need to the fence need to 7962 conservatively conservatively 7963 always generate always generate 7964 (see comment for (see comment for 7965 previous fence). previous fence). 7966 - Must happen after 7967 any preceding 7968 local/generic 7969 load/load 7970 atomic/store/store 7971 atomic/atomicrmw. 7972 - Could be split into 7973 separate s_waitcnt 7974 vmcnt(0), s_waitcnt 7975 vscnt(0) and s_waitcnt 7976 lgkmcnt(0) to allow 7977 them to be 7978 independently moved 7979 according to the 7980 following rules. 7981 - s_waitcnt vmcnt(0) 7982 must happen after 7983 any preceding 7984 global/generic 7985 load/load 7986 atomic/ 7987 atomicrmw-with-return-value. 7988 - s_waitcnt vscnt(0) 7989 must happen after 7990 any preceding 7991 global/generic 7992 store/store atomic/ 7993 atomicrmw-no-return-value. 7994 - s_waitcnt lgkmcnt(0) 7995 must happen after 7996 any preceding 7997 local/generic 7998 load/store/load 7999 atomic/store atomic/ 8000 atomicrmw. 8001 - Must happen before - Must happen before 8002 any following any following 8003 global/generic global/generic 8004 load/load load/load 8005 atomic/store/store atomic/store/store 8006 atomic/atomicrmw. atomic/atomicrmw. 8007 - Ensures that all - Ensures that all 8008 memory operations memory operations 8009 to local have have 8010 completed before completed before 8011 performing any performing any 8012 following global following global 8013 memory operations. memory operations. 8014 - Ensures that the - Ensures that the 8015 preceding preceding 8016 local/generic load local/generic load 8017 atomic/atomicrmw atomic/atomicrmw 8018 with an equal or with an equal or 8019 wider sync scope wider sync scope 8020 and memory ordering and memory ordering 8021 stronger than stronger than 8022 unordered (this is unordered (this is 8023 termed the termed the 8024 acquire-fence-paired-atomic acquire-fence-paired-atomic 8025 ) has completed ) has completed 8026 before following before following 8027 global memory global memory 8028 operations. This operations. This 8029 satisfies the satisfies the 8030 requirements of requirements of 8031 acquire. acquire. 8032 - Ensures that all - Ensures that all 8033 previous memory previous memory 8034 operations have operations have 8035 completed before a completed before a 8036 following following 8037 local/generic store local/generic store 8038 atomic/atomicrmw atomic/atomicrmw 8039 with an equal or with an equal or 8040 wider sync scope wider sync scope 8041 and memory ordering and memory ordering 8042 stronger than stronger than 8043 unordered (this is unordered (this is 8044 termed the termed the 8045 release-fence-paired-atomic release-fence-paired-atomic 8046 ). This satisfies the ). This satisfies the 8047 requirements of requirements of 8048 release. release. 8049 - Must happen before 8050 the following 8051 buffer_gl0_inv. 8052 - Ensures that the 8053 acquire-fence-paired 8054 atomic has completed 8055 before invalidating 8056 the 8057 cache. Therefore 8058 any following 8059 locations read must 8060 be no older than 8061 the value read by 8062 the 8063 acquire-fence-paired-atomic. 8064 8065 3. buffer_gl0_inv 8066 8067 - If CU wavefront execution mode, omit. 8068 - Ensures that 8069 following 8070 loads will not see 8071 stale data. 8072 8073 fence acq_rel - agent *none* 1. s_waitcnt lgkmcnt(0) & 1. s_waitcnt lgkmcnt(0) & 8074 - system vmcnt(0) vmcnt(0) & vscnt(0) 8075 8076 - If OpenCL and - If OpenCL and 8077 address space is address space is 8078 not generic, omit not generic, omit 8079 lgkmcnt(0). lgkmcnt(0). 8080 - If OpenCL and 8081 address space is 8082 local, omit 8083 vmcnt(0) and vscnt(0). 8084 - However, since LLVM - However, since LLVM 8085 currently has no currently has no 8086 address space on address space on 8087 the fence need to the fence need to 8088 conservatively conservatively 8089 always generate always generate 8090 (see comment for (see comment for 8091 previous fence). previous fence). 8092 - Could be split into - Could be split into 8093 separate s_waitcnt separate s_waitcnt 8094 vmcnt(0) and vmcnt(0), s_waitcnt 8095 s_waitcnt vscnt(0) and s_waitcnt 8096 lgkmcnt(0) to allow lgkmcnt(0) to allow 8097 them to be them to be 8098 independently moved independently moved 8099 according to the according to the 8100 following rules. following rules. 8101 - s_waitcnt vmcnt(0) - s_waitcnt vmcnt(0) 8102 must happen after must happen after 8103 any preceding any preceding 8104 global/generic global/generic 8105 load/store/load load/load 8106 atomic/store atomic/ 8107 atomic/atomicrmw. atomicrmw-with-return-value. 8108 - s_waitcnt vscnt(0) 8109 must happen after 8110 any preceding 8111 global/generic 8112 store/store atomic/ 8113 atomicrmw-no-return-value. 8114 - s_waitcnt lgkmcnt(0) - s_waitcnt lgkmcnt(0) 8115 must happen after must happen after 8116 any preceding any preceding 8117 local/generic local/generic 8118 load/store/load load/store/load 8119 atomic/store atomic/store 8120 atomic/atomicrmw. atomic/atomicrmw. 8121 - Must happen before - Must happen before 8122 the following the following 8123 buffer_wbinvl1_vol. buffer_gl*_inv. 8124 - Ensures that the - Ensures that the 8125 preceding preceding 8126 global/local/generic global/local/generic 8127 load load 8128 atomic/atomicrmw atomic/atomicrmw 8129 with an equal or with an equal or 8130 wider sync scope wider sync scope 8131 and memory ordering and memory ordering 8132 stronger than stronger than 8133 unordered (this is unordered (this is 8134 termed the termed the 8135 acquire-fence-paired-atomic acquire-fence-paired-atomic 8136 ) has completed ) has completed 8137 before invalidating before invalidating 8138 the cache. This the caches. This 8139 satisfies the satisfies the 8140 requirements of requirements of 8141 acquire. acquire. 8142 - Ensures that all - Ensures that all 8143 previous memory previous memory 8144 operations have operations have 8145 completed before a completed before a 8146 following following 8147 global/local/generic global/local/generic 8148 store store 8149 atomic/atomicrmw atomic/atomicrmw 8150 with an equal or with an equal or 8151 wider sync scope wider sync scope 8152 and memory ordering and memory ordering 8153 stronger than stronger than 8154 unordered (this is unordered (this is 8155 termed the termed the 8156 release-fence-paired-atomic release-fence-paired-atomic 8157 ). This satisfies the ). This satisfies the 8158 requirements of requirements of 8159 release. release. 8160 8161 2. buffer_wbinvl1_vol 2. buffer_gl0_inv; 8162 buffer_gl1_inv 8163 8164 - Must happen before - Must happen before 8165 any following any following 8166 global/generic global/generic 8167 load/load load/load 8168 atomic/store/store atomic/store/store 8169 atomic/atomicrmw. atomic/atomicrmw. 8170 - Ensures that - Ensures that 8171 following loads following loads 8172 will not see stale will not see stale 8173 global data. This global data. This 8174 satisfies the satisfies the 8175 requirements of requirements of 8176 acquire. acquire. 8177 8178 **Sequential Consistent Atomic** 8179 ---------------------------------------------------------------------------------------------------------------------- 8180 load atomic seq_cst - singlethread - global *Same as corresponding *Same as corresponding 8181 - wavefront - local load atomic acquire, load atomic acquire, 8182 - generic except must generated except must generated 8183 all instructions even all instructions even 8184 for OpenCL.* for OpenCL.* 8185 load atomic seq_cst - workgroup - global 1. s_waitcnt lgkmcnt(0) 1. s_waitcnt lgkmcnt(0) & 8186 - generic vmcnt(0) & vscnt(0) 8187 8188 - If CU wavefront execution mode, omit vmcnt and 8189 vscnt. 8190 - Could be split into 8191 separate s_waitcnt 8192 vmcnt(0), s_waitcnt 8193 vscnt(0) and s_waitcnt 8194 lgkmcnt(0) to allow 8195 them to be 8196 independently moved 8197 according to the 8198 following rules. 8199 - Must - waitcnt lgkmcnt(0) must 8200 happen after happen after 8201 preceding preceding 8202 global/generic load local load 8203 atomic/store atomic/store 8204 atomic/atomicrmw atomic/atomicrmw 8205 with memory with memory 8206 ordering of seq_cst ordering of seq_cst 8207 and with equal or and with equal or 8208 wider sync scope. wider sync scope. 8209 (Note that seq_cst (Note that seq_cst 8210 fences have their fences have their 8211 own s_waitcnt own s_waitcnt 8212 lgkmcnt(0) and so do lgkmcnt(0) and so do 8213 not need to be not need to be 8214 considered.) considered.) 8215 - waitcnt vmcnt(0) 8216 Must happen after 8217 preceding 8218 global/generic load 8219 atomic/ 8220 atomicrmw-with-return-value 8221 with memory 8222 ordering of seq_cst 8223 and with equal or 8224 wider sync scope. 8225 (Note that seq_cst 8226 fences have their 8227 own s_waitcnt 8228 vmcnt(0) and so do 8229 not need to be 8230 considered.) 8231 - waitcnt vscnt(0) 8232 Must happen after 8233 preceding 8234 global/generic store 8235 atomic/ 8236 atomicrmw-no-return-value 8237 with memory 8238 ordering of seq_cst 8239 and with equal or 8240 wider sync scope. 8241 (Note that seq_cst 8242 fences have their 8243 own s_waitcnt 8244 vscnt(0) and so do 8245 not need to be 8246 considered.) 8247 - Ensures any - Ensures any 8248 preceding preceding 8249 sequential sequential 8250 consistent local consistent global/local 8251 memory instructions memory instructions 8252 have completed have completed 8253 before executing before executing 8254 this sequentially this sequentially 8255 consistent consistent 8256 instruction. This instruction. This 8257 prevents reordering prevents reordering 8258 a seq_cst store a seq_cst store 8259 followed by a followed by a 8260 seq_cst load. (Note seq_cst load. (Note 8261 that seq_cst is that seq_cst is 8262 stronger than stronger than 8263 acquire/release as acquire/release as 8264 the reordering of the reordering of 8265 load acquire load acquire 8266 followed by a store followed by a store 8267 release is release is 8268 prevented by the prevented by the 8269 waitcnt of waitcnt of 8270 the release, but the release, but 8271 there is nothing there is nothing 8272 preventing a store preventing a store 8273 release followed by release followed by 8274 load acquire from load acquire from 8275 competing out of competing out of 8276 order.) order.) 8277 8278 2. *Following 2. *Following 8279 instructions same as instructions same as 8280 corresponding load corresponding load 8281 atomic acquire, atomic acquire, 8282 except must generated except must generated 8283 all instructions even all instructions even 8284 for OpenCL.* for OpenCL.* 8285 load atomic seq_cst - workgroup - local *Same as corresponding 8286 load atomic acquire, 8287 except must generated 8288 all instructions even 8289 for OpenCL.* 8290 8291 1. s_waitcnt vmcnt(0) & vscnt(0) 8292 8293 - If CU wavefront execution mode, omit. 8294 - Could be split into 8295 separate s_waitcnt 8296 vmcnt(0) and s_waitcnt 8297 vscnt(0) to allow 8298 them to be 8299 independently moved 8300 according to the 8301 following rules. 8302 - waitcnt vmcnt(0) 8303 Must happen after 8304 preceding 8305 global/generic load 8306 atomic/ 8307 atomicrmw-with-return-value 8308 with memory 8309 ordering of seq_cst 8310 and with equal or 8311 wider sync scope. 8312 (Note that seq_cst 8313 fences have their 8314 own s_waitcnt 8315 vmcnt(0) and so do 8316 not need to be 8317 considered.) 8318 - waitcnt vscnt(0) 8319 Must happen after 8320 preceding 8321 global/generic store 8322 atomic/ 8323 atomicrmw-no-return-value 8324 with memory 8325 ordering of seq_cst 8326 and with equal or 8327 wider sync scope. 8328 (Note that seq_cst 8329 fences have their 8330 own s_waitcnt 8331 vscnt(0) and so do 8332 not need to be 8333 considered.) 8334 - Ensures any 8335 preceding 8336 sequential 8337 consistent global 8338 memory instructions 8339 have completed 8340 before executing 8341 this sequentially 8342 consistent 8343 instruction. This 8344 prevents reordering 8345 a seq_cst store 8346 followed by a 8347 seq_cst load. (Note 8348 that seq_cst is 8349 stronger than 8350 acquire/release as 8351 the reordering of 8352 load acquire 8353 followed by a store 8354 release is 8355 prevented by the 8356 waitcnt of 8357 the release, but 8358 there is nothing 8359 preventing a store 8360 release followed by 8361 load acquire from 8362 competing out of 8363 order.) 8364 8365 2. *Following 8366 instructions same as 8367 corresponding load 8368 atomic acquire, 8369 except must generated 8370 all instructions even 8371 for OpenCL.* 8372 8373 load atomic seq_cst - agent - global 1. s_waitcnt lgkmcnt(0) & 1. s_waitcnt lgkmcnt(0) & 8374 - system - generic vmcnt(0) vmcnt(0) & vscnt(0) 8375 8376 - Could be split into - Could be split into 8377 separate s_waitcnt separate s_waitcnt 8378 vmcnt(0) vmcnt(0), s_waitcnt 8379 and s_waitcnt vscnt(0) and s_waitcnt 8380 lgkmcnt(0) to allow lgkmcnt(0) to allow 8381 them to be them to be 8382 independently moved independently moved 8383 according to the according to the 8384 following rules. following rules. 8385 - waitcnt lgkmcnt(0) - waitcnt lgkmcnt(0) 8386 must happen after must happen after 8387 preceding preceding 8388 global/generic load local load 8389 atomic/store atomic/store 8390 atomic/atomicrmw atomic/atomicrmw 8391 with memory with memory 8392 ordering of seq_cst ordering of seq_cst 8393 and with equal or and with equal or 8394 wider sync scope. wider sync scope. 8395 (Note that seq_cst (Note that seq_cst 8396 fences have their fences have their 8397 own s_waitcnt own s_waitcnt 8398 lgkmcnt(0) and so do lgkmcnt(0) and so do 8399 not need to be not need to be 8400 considered.) considered.) 8401 - waitcnt vmcnt(0) - waitcnt vmcnt(0) 8402 must happen after must happen after 8403 preceding preceding 8404 global/generic load global/generic load 8405 atomic/store atomic/ 8406 atomic/atomicrmw atomicrmw-with-return-value 8407 with memory with memory 8408 ordering of seq_cst ordering of seq_cst 8409 and with equal or and with equal or 8410 wider sync scope. wider sync scope. 8411 (Note that seq_cst (Note that seq_cst 8412 fences have their fences have their 8413 own s_waitcnt own s_waitcnt 8414 vmcnt(0) and so do vmcnt(0) and so do 8415 not need to be not need to be 8416 considered.) considered.) 8417 - waitcnt vscnt(0) 8418 Must happen after 8419 preceding 8420 global/generic store 8421 atomic/ 8422 atomicrmw-no-return-value 8423 with memory 8424 ordering of seq_cst 8425 and with equal or 8426 wider sync scope. 8427 (Note that seq_cst 8428 fences have their 8429 own s_waitcnt 8430 vscnt(0) and so do 8431 not need to be 8432 considered.) 8433 - Ensures any - Ensures any 8434 preceding preceding 8435 sequential sequential 8436 consistent global consistent global 8437 memory instructions memory instructions 8438 have completed have completed 8439 before executing before executing 8440 this sequentially this sequentially 8441 consistent consistent 8442 instruction. This instruction. This 8443 prevents reordering prevents reordering 8444 a seq_cst store a seq_cst store 8445 followed by a followed by a 8446 seq_cst load. (Note seq_cst load. (Note 8447 that seq_cst is that seq_cst is 8448 stronger than stronger than 8449 acquire/release as acquire/release as 8450 the reordering of the reordering of 8451 load acquire load acquire 8452 followed by a store followed by a store 8453 release is release is 8454 prevented by the prevented by the 8455 waitcnt of waitcnt of 8456 the release, but the release, but 8457 there is nothing there is nothing 8458 preventing a store preventing a store 8459 release followed by release followed by 8460 load acquire from load acquire from 8461 competing out of competing out of 8462 order.) order.) 8463 8464 2. *Following 2. *Following 8465 instructions same as instructions same as 8466 corresponding load corresponding load 8467 atomic acquire, atomic acquire, 8468 except must generated except must generated 8469 all instructions even all instructions even 8470 for OpenCL.* for OpenCL.* 8471 store atomic seq_cst - singlethread - global *Same as corresponding *Same as corresponding 8472 - wavefront - local store atomic release, store atomic release, 8473 - workgroup - generic except must generated except must generated 8474 all instructions even all instructions even 8475 for OpenCL.* for OpenCL.* 8476 store atomic seq_cst - agent - global *Same as corresponding *Same as corresponding 8477 - system - generic store atomic release, store atomic release, 8478 except must generated except must generated 8479 all instructions even all instructions even 8480 for OpenCL.* for OpenCL.* 8481 atomicrmw seq_cst - singlethread - global *Same as corresponding *Same as corresponding 8482 - wavefront - local atomicrmw acq_rel, atomicrmw acq_rel, 8483 - workgroup - generic except must generated except must generated 8484 all instructions even all instructions even 8485 for OpenCL.* for OpenCL.* 8486 atomicrmw seq_cst - agent - global *Same as corresponding *Same as corresponding 8487 - system - generic atomicrmw acq_rel, atomicrmw acq_rel, 8488 except must generated except must generated 8489 all instructions even all instructions even 8490 for OpenCL.* for OpenCL.* 8491 fence seq_cst - singlethread *none* *Same as corresponding *Same as corresponding 8492 - wavefront fence acq_rel, fence acq_rel, 8493 - workgroup except must generated except must generated 8494 - agent all instructions even all instructions even 8495 - system for OpenCL.* for OpenCL.* 8496 ============ ============ ============== ========== =============================== ================================== 8497 8498The memory order also adds the single thread optimization constrains defined in 8499table 8500:ref:`amdgpu-amdhsa-memory-model-single-thread-optimization-constraints-gfx6-gfx10-table`. 8501 8502 .. table:: AMDHSA Memory Model Single Thread Optimization Constraints GFX6-GFX10 8503 :name: amdgpu-amdhsa-memory-model-single-thread-optimization-constraints-gfx6-gfx10-table 8504 8505 ============ ============================================================== 8506 LLVM Memory Optimization Constraints 8507 Ordering 8508 ============ ============================================================== 8509 unordered *none* 8510 monotonic *none* 8511 acquire - If a load atomic/atomicrmw then no following load/load 8512 atomic/store/ store atomic/atomicrmw/fence instruction can 8513 be moved before the acquire. 8514 - If a fence then same as load atomic, plus no preceding 8515 associated fence-paired-atomic can be moved after the fence. 8516 release - If a store atomic/atomicrmw then no preceding load/load 8517 atomic/store/ store atomic/atomicrmw/fence instruction can 8518 be moved after the release. 8519 - If a fence then same as store atomic, plus no following 8520 associated fence-paired-atomic can be moved before the 8521 fence. 8522 acq_rel Same constraints as both acquire and release. 8523 seq_cst - If a load atomic then same constraints as acquire, plus no 8524 preceding sequentially consistent load atomic/store 8525 atomic/atomicrmw/fence instruction can be moved after the 8526 seq_cst. 8527 - If a store atomic then the same constraints as release, plus 8528 no following sequentially consistent load atomic/store 8529 atomic/atomicrmw/fence instruction can be moved before the 8530 seq_cst. 8531 - If an atomicrmw/fence then same constraints as acq_rel. 8532 ============ ============================================================== 8533 8534Trap Handler ABI 8535~~~~~~~~~~~~~~~~ 8536 8537For code objects generated by AMDGPU backend for HSA [HSA]_ compatible runtimes 8538(such as ROCm [AMD-ROCm]_), the runtime installs a trap handler that supports 8539the ``s_trap`` instruction with the following usage: 8540 8541 .. table:: AMDGPU Trap Handler for AMDHSA OS 8542 :name: amdgpu-trap-handler-for-amdhsa-os-table 8543 8544 =================== =============== =============== ======================= 8545 Usage Code Sequence Trap Handler Description 8546 Inputs 8547 =================== =============== =============== ======================= 8548 reserved ``s_trap 0x00`` Reserved by hardware. 8549 ``debugtrap(arg)`` ``s_trap 0x01`` ``SGPR0-1``: Reserved for HSA 8550 ``queue_ptr`` ``debugtrap`` 8551 ``VGPR0``: intrinsic (not 8552 ``arg`` implemented). 8553 ``llvm.trap`` ``s_trap 0x02`` ``SGPR0-1``: Causes dispatch to be 8554 ``queue_ptr`` terminated and its 8555 associated queue put 8556 into the error state. 8557 ``llvm.debugtrap`` ``s_trap 0x03`` - If debugger not 8558 installed then 8559 behaves as a 8560 no-operation. The 8561 trap handler is 8562 entered and 8563 immediately returns 8564 to continue 8565 execution of the 8566 wavefront. 8567 - If the debugger is 8568 installed, causes 8569 the debug trap to be 8570 reported by the 8571 debugger and the 8572 wavefront is put in 8573 the halt state until 8574 resumed by the 8575 debugger. 8576 reserved ``s_trap 0x04`` Reserved. 8577 reserved ``s_trap 0x05`` Reserved. 8578 reserved ``s_trap 0x06`` Reserved. 8579 debugger breakpoint ``s_trap 0x07`` Reserved for debugger 8580 breakpoints. 8581 reserved ``s_trap 0x08`` Reserved. 8582 reserved ``s_trap 0xfe`` Reserved. 8583 reserved ``s_trap 0xff`` Reserved. 8584 =================== =============== =============== ======================= 8585 8586.. _amdgpu-amdhsa-function-call-convention: 8587 8588Call Convention 8589~~~~~~~~~~~~~~~ 8590 8591.. note:: 8592 8593 This section is currently incomplete and has inakkuracies. It is WIP that will 8594 be updated as information is determined. 8595 8596See :ref:`amdgpu-dwarf-address-space-mapping` for information on swizzled 8597addresses. Unswizzled addresses are normal linear addresses. 8598 8599.. _amdgpu-amdhsa-function-call-convention-kernel-functions: 8600 8601Kernel Functions 8602++++++++++++++++ 8603 8604This section describes the call convention ABI for the outer kernel function. 8605 8606See :ref:`amdgpu-amdhsa-initial-kernel-execution-state` for the kernel call 8607convention. 8608 8609The following is not part of the AMDGPU kernel calling convention but describes 8610how the AMDGPU implements function calls: 8611 86121. Clang decides the kernarg layout to match the *HSA Programmer's Language 8613 Reference* [HSA]_. 8614 8615 - All structs are passed directly. 8616 - Lambda values are passed *TBA*. 8617 8618 .. TODO:: 8619 8620 - Does this really follow HSA rules? Or are structs >16 bytes passed 8621 by-value struct? 8622 - What is ABI for lambda values? 8623 86244. The kernel performs certain setup in its prolog, as described in 8625 :ref:`amdgpu-amdhsa-kernel-prolog`. 8626 8627.. _amdgpu-amdhsa-function-call-convention-non-kernel-functions: 8628 8629Non-Kernel Functions 8630++++++++++++++++++++ 8631 8632This section describes the call convention ABI for functions other than the 8633outer kernel function. 8634 8635If a kernel has function calls then scratch is always allocated and used for 8636the call stack which grows from low address to high address using the swizzled 8637scratch address space. 8638 8639On entry to a function: 8640 86411. SGPR0-3 contain a V# with the following properties (see 8642 :ref:`amdgpu-amdhsa-kernel-prolog-private-segment-buffer`): 8643 8644 * Base address pointing to the beginning of the wavefront scratch backing 8645 memory. 8646 * Swizzled with dword element size and stride of wavefront size elements. 8647 86482. The FLAT_SCRATCH register pair is setup. See 8649 :ref:`amdgpu-amdhsa-kernel-prolog-flat-scratch`. 86503. GFX6-8: M0 register set to the size of LDS in bytes. See 8651 :ref:`amdgpu-amdhsa-kernel-prolog-m0`. 86524. The EXEC register is set to the lanes active on entry to the function. 86535. MODE register: *TBD* 86546. VGPR0-31 and SGPR4-29 are used to pass function input arguments as described 8655 below. 86567. SGPR30-31 return address (RA). The code address that the function must 8657 return to when it completes. The value is undefined if the function is *no 8658 return*. 86598. SGPR32 is used for the stack pointer (SP). It is an unswizzled scratch 8660 offset relative to the beginning of the wavefront scratch backing memory. 8661 8662 The unswizzled SP can be used with buffer instructions as an unswizzled SGPR 8663 offset with the scratch V# in SGPR0-3 to access the stack in a swizzled 8664 manner. 8665 8666 The unswizzled SP value can be converted into the swizzled SP value by: 8667 8668 | swizzled SP = unswizzled SP / wavefront size 8669 8670 This may be used to obtain the private address space address of stack 8671 objects and to convert this address to a flat address by adding the flat 8672 scratch aperture base address. 8673 8674 The swizzled SP value is always 4 bytes aligned for the ``r600`` 8675 architecture and 16 byte aligned for the ``amdgcn`` architecture. 8676 8677 .. note:: 8678 8679 The ``amdgcn`` value is selected to avoid dynamic stack alignment for the 8680 OpenCL language which has the largest base type defined as 16 bytes. 8681 8682 On entry, the swizzled SP value is the address of the first function 8683 argument passed on the stack. Other stack passed arguments are positive 8684 offsets from the entry swizzled SP value. 8685 8686 The function may use positive offsets beyond the last stack passed argument 8687 for stack allocated local variables and register spill slots. If necessary 8688 the function may align these to greater alignment than 16 bytes. After these 8689 the function may dynamically allocate space for such things as runtime sized 8690 ``alloca`` local allocations. 8691 8692 If the function calls another function, it will place any stack allocated 8693 arguments after the last local allocation and adjust SGPR32 to the address 8694 after the last local allocation. 8695 86969. All other registers are unspecified. 869710. Any necessary ``waitcnt`` has been performed to ensure memory is available 8698 to the function. 8699 8700On exit from a function: 8701 87021. VGPR0-31 and SGPR4-29 are used to pass function result arguments as 8703 described below. Any registers used are considered clobbered registers. 87042. The following registers are preserved and have the same value as on entry: 8705 8706 * FLAT_SCRATCH 8707 * EXEC 8708 * GFX6-8: M0 8709 * All SGPR and VGPR registers except the clobbered registers of SGPR4-31 and 8710 VGPR0-31. 8711 8712 For the AMDGPU backend, an inter-procedural register allocation (IPRA) 8713 optimization may mark some of clobbered SGPR4-31 and VGPR0-31 registers as 8714 preserved if it can be determined that the called function does not change 8715 their value. 8716 87172. The PC is set to the RA provided on entry. 87183. MODE register: *TBD*. 87194. All other registers are clobbered. 87205. Any necessary ``waitcnt`` has been performed to ensure memory accessed by 8721 function is available to the caller. 8722 8723.. TODO:: 8724 8725 - On gfx908 are all ACC registers clobbered? 8726 8727 - How are function results returned? The address of structured types is passed 8728 by reference, but what about other types? 8729 8730The function input arguments are made up of the formal arguments explicitly 8731declared by the source language function plus the implicit input arguments used 8732by the implementation. 8733 8734The source language input arguments are: 8735 87361. Any source language implicit ``this`` or ``self`` argument comes first as a 8737 pointer type. 87382. Followed by the function formal arguments in left to right source order. 8739 8740The source language result arguments are: 8741 87421. The function result argument. 8743 8744The source language input or result struct type arguments that are less than or 8745equal to 16 bytes, are decomposed recursively into their base type fields, and 8746each field is passed as if a separate argument. For input arguments, if the 8747called function requires the struct to be in memory, for example because its 8748address is taken, then the function body is responsible for allocating a stack 8749location and copying the field arguments into it. Clang terms this *direct 8750struct*. 8751 8752The source language input struct type arguments that are greater than 16 bytes, 8753are passed by reference. The caller is responsible for allocating a stack 8754location to make a copy of the struct value and pass the address as the input 8755argument. The called function is responsible to perform the dereference when 8756accessing the input argument. Clang terms this *by-value struct*. 8757 8758A source language result struct type argument that is greater than 16 bytes, is 8759returned by reference. The caller is responsible for allocating a stack location 8760to hold the result value and passes the address as the last input argument 8761(before the implicit input arguments). In this case there are no result 8762arguments. The called function is responsible to perform the dereference when 8763storing the result value. Clang terms this *structured return (sret)*. 8764 8765*TODO: correct the sret definition.* 8766 8767.. TODO:: 8768 8769 Is this definition correct? Or is sret only used if passing in registers, and 8770 pass as non-decomposed struct as stack argument? Or something else? Is the 8771 memory location in the caller stack frame, or a stack memory argument and so 8772 no address is passed as the caller can directly write to the argument stack 8773 location. But then the stack location is still live after return. If an 8774 argument stack location is it the first stack argument or the last one? 8775 8776Lambda argument types are treated as struct types with an implementation defined 8777set of fields. 8778 8779.. TODO:: 8780 8781 Need to specify the ABI for lambda types for AMDGPU. 8782 8783For AMDGPU backend all source language arguments (including the decomposed 8784struct type arguments) are passed in VGPRs unless marked ``inreg`` in which case 8785they are passed in SGPRs. 8786 8787The AMDGPU backend walks the function call graph from the leaves to determine 8788which implicit input arguments are used, propagating to each caller of the 8789function. The used implicit arguments are appended to the function arguments 8790after the source language arguments in the following order: 8791 8792.. TODO:: 8793 8794 Is recursion or external functions supported? 8795 87961. Work-Item ID (1 VGPR) 8797 8798 The X, Y and Z work-item ID are packed into a single VGRP with the following 8799 layout. Only fields actually used by the function are set. The other bits 8800 are undefined. 8801 8802 The values come from the initial kernel execution state. See 8803 :ref:`amdgpu-amdhsa-vgpr-register-set-up-order-table`. 8804 8805 .. table:: Work-item implict argument layout 8806 :name: amdgpu-amdhsa-workitem-implict-argument-layout-table 8807 8808 ======= ======= ============== 8809 Bits Size Field Name 8810 ======= ======= ============== 8811 9:0 10 bits X Work-Item ID 8812 19:10 10 bits Y Work-Item ID 8813 29:20 10 bits Z Work-Item ID 8814 31:30 2 bits Unused 8815 ======= ======= ============== 8816 88172. Dispatch Ptr (2 SGPRs) 8818 8819 The value comes from the initial kernel execution state. See 8820 :ref:`amdgpu-amdhsa-sgpr-register-set-up-order-table`. 8821 88223. Queue Ptr (2 SGPRs) 8823 8824 The value comes from the initial kernel execution state. See 8825 :ref:`amdgpu-amdhsa-sgpr-register-set-up-order-table`. 8826 88274. Kernarg Segment Ptr (2 SGPRs) 8828 8829 The value comes from the initial kernel execution state. See 8830 :ref:`amdgpu-amdhsa-sgpr-register-set-up-order-table`. 8831 88325. Dispatch id (2 SGPRs) 8833 8834 The value comes from the initial kernel execution state. See 8835 :ref:`amdgpu-amdhsa-sgpr-register-set-up-order-table`. 8836 88376. Work-Group ID X (1 SGPR) 8838 8839 The value comes from the initial kernel execution state. See 8840 :ref:`amdgpu-amdhsa-sgpr-register-set-up-order-table`. 8841 88427. Work-Group ID Y (1 SGPR) 8843 8844 The value comes from the initial kernel execution state. See 8845 :ref:`amdgpu-amdhsa-sgpr-register-set-up-order-table`. 8846 88478. Work-Group ID Z (1 SGPR) 8848 8849 The value comes from the initial kernel execution state. See 8850 :ref:`amdgpu-amdhsa-sgpr-register-set-up-order-table`. 8851 88529. Implicit Argument Ptr (2 SGPRs) 8853 8854 The value is computed by adding an offset to Kernarg Segment Ptr to get the 8855 global address space pointer to the first kernarg implicit argument. 8856 8857The input and result arguments are assigned in order in the following manner: 8858 8859..note:: 8860 8861 There are likely some errors and ommissions in the following description that 8862 need correction. 8863 8864 ..TODO:: 8865 8866 Check the clang source code to decipher how funtion arguments and return 8867 results are handled. Also see the AMDGPU specific values used. 8868 8869* VGPR arguments are assigned to consecutive VGPRs starting at VGPR0 up to 8870 VGPR31. 8871 8872 If there are more arguments than will fit in these registers, the remaining 8873 arguments are allocated on the stack in order on naturally aligned 8874 addresses. 8875 8876 .. TODO:: 8877 8878 How are overly aligned structures allocated on the stack? 8879 8880* SGPR arguments are assigned to consecutive SGPRs starting at SGPR0 up to 8881 SGPR29. 8882 8883 If there are more arguments than will fit in these registers, the remaining 8884 arguments are allocated on the stack in order on naturally aligned 8885 addresses. 8886 8887Note that decomposed struct type arguments may have some fields passed in 8888registers and some in memory. 8889 8890..TODO:: 8891 8892 So a struct which can pass some fields as decomposed register arguments, will 8893 pass the rest as decomposed stack elements? But an arguent that will not start 8894 in registers will not be decomposed and will be passed as a non-decomposed 8895 stack value? 8896 8897The following is not part of the AMDGPU function calling convention but 8898describes how the AMDGPU implements function calls: 8899 89001. SGPR33 is used as a frame pointer (FP) if necessary. Like the SP it is an 8901 unswizzled scratch address. It is only needed if runtime sized ``alloca`` 8902 are used, or for the reasons defined in ``SIFrameLowering``. 89032. Runtime stack alignment is not currently supported. 8904 8905 .. TODO:: 8906 8907 - If runtime stack alignment is supported then will an extra argument 8908 pointer register be used? 8909 89102. Allocating SGPR arguments on the stack are not supported. 8911 89123. No CFI is currently generated. See :ref:`amdgpu-call-frame-information`. 8913 8914 ..note:: 8915 8916 CFI will be generated that defines the CFA as the unswizzled address 8917 relative to the wave scratch base in the unswizzled private address space 8918 of the lowest address stack allocated local variable. 8919 8920 ``DW_AT_frame_base`` will be defined as the swizzled address in the 8921 swizzled private address space by dividing the CFA by the wavefront size 8922 (since CFA is always at least dword aligned which matches the scratch 8923 swizzle element size). 8924 8925 If no dynamic stack alignment was performed, the stack allocated arguments 8926 are accessed as negative offsets relative to ``DW_AT_frame_base``, and the 8927 local variables and register spill slots are accessed as positive offsets 8928 relative to ``DW_AT_frame_base``. 8929 89304. Function argument passing is implemented by copying the input physical 8931 registers to virtual registers on entry. The register allocator can spill if 8932 necessary. These are copied back to physical registers at call sites. The 8933 net effect is that each function call can have these values in entirely 8934 distinct locations. The IPRA can help avoid shuffling argument registers. 89355. Call sites are implemented by setting up the arguments at positive offsets 8936 from SP. Then SP is incremented to account for the known frame size before 8937 the call and decremented after the call. 8938 8939 ..note:: 8940 8941 The CFI will reflect the changed calculation needed to compute the CFA 8942 from SP. 8943 89446. 4 byte spill slots are used in the stack frame. One slot is allocated for an 8945 emergency spill slot. Buffer instructions are used for stack accesses and 8946 not the ``flat_scratch`` instruction. 8947 8948 ..TODO:: 8949 8950 Explain when the emergency spill slot is used. 8951 8952.. TODO:: 8953 8954 Possible broken issues: 8955 8956 - Stack arguments must be aligned to required alignment. 8957 - Stack is aligned to max(16, max formal argument alignment) 8958 - Direct argument < 64 bits should check register budget. 8959 - Register budget calculation should respect ``inreg`` for SGPR. 8960 - SGPR overflow is not handled. 8961 - struct with 1 member unpeeling is not checking size of member. 8962 - ``sret`` is after ``this`` pointer. 8963 - Caller is not implementing stack realignment: need an extra pointer. 8964 - Should say AMDGPU passes FP rather than SP. 8965 - Should CFI define CFA as address of locals or arguments. Difference is 8966 apparent when have implemented dynamic alignment. 8967 - If ``SCRATCH`` instruction could allow negative offsets then can make FP be 8968 highest address of stack frame and use negative offset for locals. Would 8969 allow SP to be the same as FP and could support signal-handler-like as now 8970 have a real SP for the top of the stack. 8971 - How is ``sret`` passed on the stack? In argument stack area? Can it overlay 8972 arguments? 8973 8974AMDPAL 8975------ 8976 8977This section provides code conventions used when the target triple OS is 8978``amdpal`` (see :ref:`amdgpu-target-triples`) for passing runtime parameters 8979from the application/runtime to each invocation of a hardware shader. These 8980parameters include both generic, application-controlled parameters called 8981*user data* as well as system-generated parameters that are a product of the 8982draw or dispatch execution. 8983 8984User Data 8985~~~~~~~~~ 8986 8987Each hardware stage has a set of 32-bit *user data registers* which can be 8988written from a command buffer and then loaded into SGPRs when waves are launched 8989via a subsequent dispatch or draw operation. This is the way most arguments are 8990passed from the application/runtime to a hardware shader. 8991 8992Compute User Data 8993~~~~~~~~~~~~~~~~~ 8994 8995Compute shader user data mappings are simpler than graphics shaders, and have a 8996fixed mapping. 8997 8998Note that there are always 10 available *user data entries* in registers - 8999entries beyond that limit must be fetched from memory (via the spill table 9000pointer) by the shader. 9001 9002 .. table:: PAL Compute Shader User Data Registers 9003 :name: pal-compute-user-data-registers 9004 9005 ============= ================================ 9006 User Register Description 9007 ============= ================================ 9008 0 Global Internal Table (32-bit pointer) 9009 1 Per-Shader Internal Table (32-bit pointer) 9010 2 - 11 Application-Controlled User Data (10 32-bit values) 9011 12 Spill Table (32-bit pointer) 9012 13 - 14 Thread Group Count (64-bit pointer) 9013 15 GDS Range 9014 ============= ================================ 9015 9016Graphics User Data 9017~~~~~~~~~~~~~~~~~~ 9018 9019Graphics pipelines support a much more flexible user data mapping: 9020 9021 .. table:: PAL Graphics Shader User Data Registers 9022 :name: pal-graphics-user-data-registers 9023 9024 ============= ================================ 9025 User Register Description 9026 ============= ================================ 9027 0 Global Internal Table (32-bit pointer) 9028 + Per-Shader Internal Table (32-bit pointer) 9029 + 1-15 Application Controlled User Data 9030 (1-15 Contiguous 32-bit Values in Registers) 9031 + Spill Table (32-bit pointer) 9032 + Draw Index (First Stage Only) 9033 + Vertex Offset (First Stage Only) 9034 + Instance Offset (First Stage Only) 9035 ============= ================================ 9036 9037 The placement of the global internal table remains fixed in the first *user 9038 data SGPR register*. Otherwise all parameters are optional, and can be mapped 9039 to any desired *user data SGPR register*, with the following restrictions: 9040 9041 * Draw Index, Vertex Offset, and Instance Offset can only be used by the first 9042 active hardware stage in a graphics pipeline (i.e. where the API vertex 9043 shader runs). 9044 9045 * Application-controlled user data must be mapped into a contiguous range of 9046 user data registers. 9047 9048 * The application-controlled user data range supports compaction remapping, so 9049 only *entries* that are actually consumed by the shader must be assigned to 9050 corresponding *registers*. Note that in order to support an efficient runtime 9051 implementation, the remapping must pack *registers* in the same order as 9052 *entries*, with unused *entries* removed. 9053 9054.. _pal_global_internal_table: 9055 9056Global Internal Table 9057~~~~~~~~~~~~~~~~~~~~~ 9058 9059The global internal table is a table of *shader resource descriptors* (SRDs) 9060that define how certain engine-wide, runtime-managed resources should be 9061accessed from a shader. The majority of these resources have HW-defined formats, 9062and it is up to the compiler to write/read data as required by the target 9063hardware. 9064 9065The following table illustrates the required format: 9066 9067 .. table:: PAL Global Internal Table 9068 :name: pal-git-table 9069 9070 ============= ================================ 9071 Offset Description 9072 ============= ================================ 9073 0-3 Graphics Scratch SRD 9074 4-7 Compute Scratch SRD 9075 8-11 ES/GS Ring Output SRD 9076 12-15 ES/GS Ring Input SRD 9077 16-19 GS/VS Ring Output #0 9078 20-23 GS/VS Ring Output #1 9079 24-27 GS/VS Ring Output #2 9080 28-31 GS/VS Ring Output #3 9081 32-35 GS/VS Ring Input SRD 9082 36-39 Tessellation Factor Buffer SRD 9083 40-43 Off-Chip LDS Buffer SRD 9084 44-47 Off-Chip Param Cache Buffer SRD 9085 48-51 Sample Position Buffer SRD 9086 52 vaRange::ShadowDescriptorTable High Bits 9087 ============= ================================ 9088 9089 The pointer to the global internal table passed to the shader as user data 9090 is a 32-bit pointer. The top 32 bits should be assumed to be the same as 9091 the top 32 bits of the pipeline, so the shader may use the program 9092 counter's top 32 bits. 9093 9094Unspecified OS 9095-------------- 9096 9097This section provides code conventions used when the target triple OS is 9098empty (see :ref:`amdgpu-target-triples`). 9099 9100Trap Handler ABI 9101~~~~~~~~~~~~~~~~ 9102 9103For code objects generated by AMDGPU backend for non-amdhsa OS, the runtime does 9104not install a trap handler. The ``llvm.trap`` and ``llvm.debugtrap`` 9105instructions are handled as follows: 9106 9107 .. table:: AMDGPU Trap Handler for Non-AMDHSA OS 9108 :name: amdgpu-trap-handler-for-non-amdhsa-os-table 9109 9110 =============== =============== =========================================== 9111 Usage Code Sequence Description 9112 =============== =============== =========================================== 9113 llvm.trap s_endpgm Causes wavefront to be terminated. 9114 llvm.debugtrap *none* Compiler warning given that there is no 9115 trap handler installed. 9116 =============== =============== =========================================== 9117 9118Source Languages 9119================ 9120 9121.. _amdgpu-opencl: 9122 9123OpenCL 9124------ 9125 9126When the language is OpenCL the following differences occur: 9127 91281. The OpenCL memory model is used (see :ref:`amdgpu-amdhsa-memory-model`). 91292. The AMDGPU backend appends additional arguments to the kernel's explicit 9130 arguments for the AMDHSA OS (see 9131 :ref:`opencl-kernel-implicit-arguments-appended-for-amdhsa-os-table`). 91323. Additional metadata is generated 9133 (see :ref:`amdgpu-amdhsa-code-object-metadata`). 9134 9135 .. table:: OpenCL kernel implicit arguments appended for AMDHSA OS 9136 :name: opencl-kernel-implicit-arguments-appended-for-amdhsa-os-table 9137 9138 ======== ==== ========= =========================================== 9139 Position Byte Byte Description 9140 Size Alignment 9141 ======== ==== ========= =========================================== 9142 1 8 8 OpenCL Global Offset X 9143 2 8 8 OpenCL Global Offset Y 9144 3 8 8 OpenCL Global Offset Z 9145 4 8 8 OpenCL address of printf buffer 9146 5 8 8 OpenCL address of virtual queue used by 9147 enqueue_kernel. 9148 6 8 8 OpenCL address of AqlWrap struct used by 9149 enqueue_kernel. 9150 7 8 8 Pointer argument used for Multi-gird 9151 synchronization. 9152 ======== ==== ========= =========================================== 9153 9154.. _amdgpu-hcc: 9155 9156HCC 9157--- 9158 9159When the language is HCC the following differences occur: 9160 91611. The HSA memory model is used (see :ref:`amdgpu-amdhsa-memory-model`). 9162 9163.. _amdgpu-assembler: 9164 9165Assembler 9166--------- 9167 9168AMDGPU backend has LLVM-MC based assembler which is currently in development. 9169It supports AMDGCN GFX6-GFX10. 9170 9171This section describes general syntax for instructions and operands. 9172 9173Instructions 9174~~~~~~~~~~~~ 9175 9176.. toctree:: 9177 :hidden: 9178 9179 AMDGPU/AMDGPUAsmGFX7 9180 AMDGPU/AMDGPUAsmGFX8 9181 AMDGPU/AMDGPUAsmGFX9 9182 AMDGPU/AMDGPUAsmGFX900 9183 AMDGPU/AMDGPUAsmGFX904 9184 AMDGPU/AMDGPUAsmGFX906 9185 AMDGPU/AMDGPUAsmGFX908 9186 AMDGPU/AMDGPUAsmGFX10 9187 AMDGPU/AMDGPUAsmGFX1011 9188 AMDGPUModifierSyntax 9189 AMDGPUOperandSyntax 9190 AMDGPUInstructionSyntax 9191 AMDGPUInstructionNotation 9192 9193An instruction has the following :doc:`syntax<AMDGPUInstructionSyntax>`: 9194 9195 | ``<``\ *opcode*\ ``> <``\ *operand0*\ ``>, <``\ *operand1*\ ``>,... 9196 <``\ *modifier0*\ ``> <``\ *modifier1*\ ``>...`` 9197 9198:doc:`Operands<AMDGPUOperandSyntax>` are comma-separated while 9199:doc:`modifiers<AMDGPUModifierSyntax>` are space-separated. 9200 9201The order of operands and modifiers is fixed. 9202Most modifiers are optional and may be omitted. 9203 9204Links to detailed instruction syntax description may be found in the following 9205table. Note that features under development are not included 9206in this description. 9207 9208 =================================== ======================================= 9209 Core ISA ISA Extensions 9210 =================================== ======================================= 9211 :doc:`GFX7<AMDGPU/AMDGPUAsmGFX7>` \- 9212 :doc:`GFX8<AMDGPU/AMDGPUAsmGFX8>` \- 9213 :doc:`GFX9<AMDGPU/AMDGPUAsmGFX9>` :doc:`gfx900<AMDGPU/AMDGPUAsmGFX900>` 9214 9215 :doc:`gfx902<AMDGPU/AMDGPUAsmGFX900>` 9216 9217 :doc:`gfx904<AMDGPU/AMDGPUAsmGFX904>` 9218 9219 :doc:`gfx906<AMDGPU/AMDGPUAsmGFX906>` 9220 9221 :doc:`gfx908<AMDGPU/AMDGPUAsmGFX908>` 9222 9223 :doc:`gfx909<AMDGPU/AMDGPUAsmGFX900>` 9224 9225 :doc:`GFX10<AMDGPU/AMDGPUAsmGFX10>` :doc:`gfx1011<AMDGPU/AMDGPUAsmGFX1011>` 9226 9227 :doc:`gfx1012<AMDGPU/AMDGPUAsmGFX1011>` 9228 =================================== ======================================= 9229 9230For more information about instructions, their semantics and supported 9231combinations of operands, refer to one of instruction set architecture manuals 9232[AMD-GCN-GFX6]_, [AMD-GCN-GFX7]_, [AMD-GCN-GFX8]_, [AMD-GCN-GFX9]_ and 9233[AMD-GCN-GFX10]_. 9234 9235Operands 9236~~~~~~~~ 9237 9238Detailed description of operands may be found :doc:`here<AMDGPUOperandSyntax>`. 9239 9240Modifiers 9241~~~~~~~~~ 9242 9243Detailed description of modifiers may be found 9244:doc:`here<AMDGPUModifierSyntax>`. 9245 9246Instruction Examples 9247~~~~~~~~~~~~~~~~~~~~ 9248 9249DS 9250++ 9251 9252.. code-block:: nasm 9253 9254 ds_add_u32 v2, v4 offset:16 9255 ds_write_src2_b64 v2 offset0:4 offset1:8 9256 ds_cmpst_f32 v2, v4, v6 9257 ds_min_rtn_f64 v[8:9], v2, v[4:5] 9258 9259For full list of supported instructions, refer to "LDS/GDS instructions" in ISA 9260Manual. 9261 9262FLAT 9263++++ 9264 9265.. code-block:: nasm 9266 9267 flat_load_dword v1, v[3:4] 9268 flat_store_dwordx3 v[3:4], v[5:7] 9269 flat_atomic_swap v1, v[3:4], v5 glc 9270 flat_atomic_cmpswap v1, v[3:4], v[5:6] glc slc 9271 flat_atomic_fmax_x2 v[1:2], v[3:4], v[5:6] glc 9272 9273For full list of supported instructions, refer to "FLAT instructions" in ISA 9274Manual. 9275 9276MUBUF 9277+++++ 9278 9279.. code-block:: nasm 9280 9281 buffer_load_dword v1, off, s[4:7], s1 9282 buffer_store_dwordx4 v[1:4], v2, ttmp[4:7], s1 offen offset:4 glc tfe 9283 buffer_store_format_xy v[1:2], off, s[4:7], s1 9284 buffer_wbinvl1 9285 buffer_atomic_inc v1, v2, s[8:11], s4 idxen offset:4 slc 9286 9287For full list of supported instructions, refer to "MUBUF Instructions" in ISA 9288Manual. 9289 9290SMRD/SMEM 9291+++++++++ 9292 9293.. code-block:: nasm 9294 9295 s_load_dword s1, s[2:3], 0xfc 9296 s_load_dwordx8 s[8:15], s[2:3], s4 9297 s_load_dwordx16 s[88:103], s[2:3], s4 9298 s_dcache_inv_vol 9299 s_memtime s[4:5] 9300 9301For full list of supported instructions, refer to "Scalar Memory Operations" in 9302ISA Manual. 9303 9304SOP1 9305++++ 9306 9307.. code-block:: nasm 9308 9309 s_mov_b32 s1, s2 9310 s_mov_b64 s[0:1], 0x80000000 9311 s_cmov_b32 s1, 200 9312 s_wqm_b64 s[2:3], s[4:5] 9313 s_bcnt0_i32_b64 s1, s[2:3] 9314 s_swappc_b64 s[2:3], s[4:5] 9315 s_cbranch_join s[4:5] 9316 9317For full list of supported instructions, refer to "SOP1 Instructions" in ISA 9318Manual. 9319 9320SOP2 9321++++ 9322 9323.. code-block:: nasm 9324 9325 s_add_u32 s1, s2, s3 9326 s_and_b64 s[2:3], s[4:5], s[6:7] 9327 s_cselect_b32 s1, s2, s3 9328 s_andn2_b32 s2, s4, s6 9329 s_lshr_b64 s[2:3], s[4:5], s6 9330 s_ashr_i32 s2, s4, s6 9331 s_bfm_b64 s[2:3], s4, s6 9332 s_bfe_i64 s[2:3], s[4:5], s6 9333 s_cbranch_g_fork s[4:5], s[6:7] 9334 9335For full list of supported instructions, refer to "SOP2 Instructions" in ISA 9336Manual. 9337 9338SOPC 9339++++ 9340 9341.. code-block:: nasm 9342 9343 s_cmp_eq_i32 s1, s2 9344 s_bitcmp1_b32 s1, s2 9345 s_bitcmp0_b64 s[2:3], s4 9346 s_setvskip s3, s5 9347 9348For full list of supported instructions, refer to "SOPC Instructions" in ISA 9349Manual. 9350 9351SOPP 9352++++ 9353 9354.. code-block:: nasm 9355 9356 s_barrier 9357 s_nop 2 9358 s_endpgm 9359 s_waitcnt 0 ; Wait for all counters to be 0 9360 s_waitcnt vmcnt(0) & expcnt(0) & lgkmcnt(0) ; Equivalent to above 9361 s_waitcnt vmcnt(1) ; Wait for vmcnt counter to be 1. 9362 s_sethalt 9 9363 s_sleep 10 9364 s_sendmsg 0x1 9365 s_sendmsg sendmsg(MSG_INTERRUPT) 9366 s_trap 1 9367 9368For full list of supported instructions, refer to "SOPP Instructions" in ISA 9369Manual. 9370 9371Unless otherwise mentioned, little verification is performed on the operands 9372of SOPP Instructions, so it is up to the programmer to be familiar with the 9373range or acceptable values. 9374 9375VALU 9376++++ 9377 9378For vector ALU instruction opcodes (VOP1, VOP2, VOP3, VOPC, VOP_DPP, VOP_SDWA), 9379the assembler will automatically use optimal encoding based on its operands. To 9380force specific encoding, one can add a suffix to the opcode of the instruction: 9381 9382* _e32 for 32-bit VOP1/VOP2/VOPC 9383* _e64 for 64-bit VOP3 9384* _dpp for VOP_DPP 9385* _sdwa for VOP_SDWA 9386 9387VOP1/VOP2/VOP3/VOPC examples: 9388 9389.. code-block:: nasm 9390 9391 v_mov_b32 v1, v2 9392 v_mov_b32_e32 v1, v2 9393 v_nop 9394 v_cvt_f64_i32_e32 v[1:2], v2 9395 v_floor_f32_e32 v1, v2 9396 v_bfrev_b32_e32 v1, v2 9397 v_add_f32_e32 v1, v2, v3 9398 v_mul_i32_i24_e64 v1, v2, 3 9399 v_mul_i32_i24_e32 v1, -3, v3 9400 v_mul_i32_i24_e32 v1, -100, v3 9401 v_addc_u32 v1, s[0:1], v2, v3, s[2:3] 9402 v_max_f16_e32 v1, v2, v3 9403 9404VOP_DPP examples: 9405 9406.. code-block:: nasm 9407 9408 v_mov_b32 v0, v0 quad_perm:[0,2,1,1] 9409 v_sin_f32 v0, v0 row_shl:1 row_mask:0xa bank_mask:0x1 bound_ctrl:0 9410 v_mov_b32 v0, v0 wave_shl:1 9411 v_mov_b32 v0, v0 row_mirror 9412 v_mov_b32 v0, v0 row_bcast:31 9413 v_mov_b32 v0, v0 quad_perm:[1,3,0,1] row_mask:0xa bank_mask:0x1 bound_ctrl:0 9414 v_add_f32 v0, v0, |v0| row_shl:1 row_mask:0xa bank_mask:0x1 bound_ctrl:0 9415 v_max_f16 v1, v2, v3 row_shl:1 row_mask:0xa bank_mask:0x1 bound_ctrl:0 9416 9417VOP_SDWA examples: 9418 9419.. code-block:: nasm 9420 9421 v_mov_b32 v1, v2 dst_sel:BYTE_0 dst_unused:UNUSED_PRESERVE src0_sel:DWORD 9422 v_min_u32 v200, v200, v1 dst_sel:WORD_1 dst_unused:UNUSED_PAD src0_sel:BYTE_1 src1_sel:DWORD 9423 v_sin_f32 v0, v0 dst_unused:UNUSED_PAD src0_sel:WORD_1 9424 v_fract_f32 v0, |v0| dst_sel:DWORD dst_unused:UNUSED_PAD src0_sel:WORD_1 9425 v_cmpx_le_u32 vcc, v1, v2 src0_sel:BYTE_2 src1_sel:WORD_0 9426 9427For full list of supported instructions, refer to "Vector ALU instructions". 9428 9429.. TODO:: 9430 9431 Remove once we switch to code object v3 by default. 9432 9433.. _amdgpu-amdhsa-assembler-predefined-symbols-v2: 9434 9435Code Object V2 Predefined Symbols (-mattr=-code-object-v3) 9436~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 9437 9438.. warning:: Code Object V2 is not the default code object version emitted by 9439 this version of LLVM. For a description of the predefined symbols available 9440 with the default configuration (Code Object V3) see 9441 :ref:`amdgpu-amdhsa-assembler-predefined-symbols-v3`. 9442 9443The AMDGPU assembler defines and updates some symbols automatically. These 9444symbols do not affect code generation. 9445 9446.option.machine_version_major 9447+++++++++++++++++++++++++++++ 9448 9449Set to the GFX major generation number of the target being assembled for. For 9450example, when assembling for a "GFX9" target this will be set to the integer 9451value "9". The possible GFX major generation numbers are presented in 9452:ref:`amdgpu-processors`. 9453 9454.option.machine_version_minor 9455+++++++++++++++++++++++++++++ 9456 9457Set to the GFX minor generation number of the target being assembled for. For 9458example, when assembling for a "GFX810" target this will be set to the integer 9459value "1". The possible GFX minor generation numbers are presented in 9460:ref:`amdgpu-processors`. 9461 9462.option.machine_version_stepping 9463++++++++++++++++++++++++++++++++ 9464 9465Set to the GFX stepping generation number of the target being assembled for. 9466For example, when assembling for a "GFX704" target this will be set to the 9467integer value "4". The possible GFX stepping generation numbers are presented 9468in :ref:`amdgpu-processors`. 9469 9470.kernel.vgpr_count 9471++++++++++++++++++ 9472 9473Set to zero each time a 9474:ref:`amdgpu-amdhsa-assembler-directive-amdgpu_hsa_kernel` directive is 9475encountered. At each instruction, if the current value of this symbol is less 9476than or equal to the maximum VPGR number explicitly referenced within that 9477instruction then the symbol value is updated to equal that VGPR number plus 9478one. 9479 9480.kernel.sgpr_count 9481++++++++++++++++++ 9482 9483Set to zero each time a 9484:ref:`amdgpu-amdhsa-assembler-directive-amdgpu_hsa_kernel` directive is 9485encountered. At each instruction, if the current value of this symbol is less 9486than or equal to the maximum VPGR number explicitly referenced within that 9487instruction then the symbol value is updated to equal that SGPR number plus 9488one. 9489 9490.. _amdgpu-amdhsa-assembler-directives-v2: 9491 9492Code Object V2 Directives (-mattr=-code-object-v3) 9493~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 9494 9495.. warning:: Code Object V2 is not the default code object version emitted by 9496 this version of LLVM. For a description of the directives supported with 9497 the default configuration (Code Object V3) see 9498 :ref:`amdgpu-amdhsa-assembler-directives-v3`. 9499 9500AMDGPU ABI defines auxiliary data in output code object. In assembly source, 9501one can specify them with assembler directives. 9502 9503.hsa_code_object_version major, minor 9504+++++++++++++++++++++++++++++++++++++ 9505 9506*major* and *minor* are integers that specify the version of the HSA code 9507object that will be generated by the assembler. 9508 9509.hsa_code_object_isa [major, minor, stepping, vendor, arch] 9510+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++ 9511 9512 9513*major*, *minor*, and *stepping* are all integers that describe the instruction 9514set architecture (ISA) version of the assembly program. 9515 9516*vendor* and *arch* are quoted strings. *vendor* should always be equal to 9517"AMD" and *arch* should always be equal to "AMDGPU". 9518 9519By default, the assembler will derive the ISA version, *vendor*, and *arch* 9520from the value of the -mcpu option that is passed to the assembler. 9521 9522.. _amdgpu-amdhsa-assembler-directive-amdgpu_hsa_kernel: 9523 9524.amdgpu_hsa_kernel (name) 9525+++++++++++++++++++++++++ 9526 9527This directives specifies that the symbol with given name is a kernel entry 9528point (label) and the object should contain corresponding symbol of type 9529STT_AMDGPU_HSA_KERNEL. 9530 9531.amd_kernel_code_t 9532++++++++++++++++++ 9533 9534This directive marks the beginning of a list of key / value pairs that are used 9535to specify the amd_kernel_code_t object that will be emitted by the assembler. 9536The list must be terminated by the *.end_amd_kernel_code_t* directive. For any 9537amd_kernel_code_t values that are unspecified a default value will be used. The 9538default value for all keys is 0, with the following exceptions: 9539 9540- *amd_code_version_major* defaults to 1. 9541- *amd_kernel_code_version_minor* defaults to 2. 9542- *amd_machine_kind* defaults to 1. 9543- *amd_machine_version_major*, *machine_version_minor*, and 9544 *amd_machine_version_stepping* are derived from the value of the -mcpu option 9545 that is passed to the assembler. 9546- *kernel_code_entry_byte_offset* defaults to 256. 9547- *wavefront_size* defaults 6 for all targets before GFX10. For GFX10 onwards 9548 defaults to 6 if target feature ``wavefrontsize64`` is enabled, otherwise 5. 9549 Note that wavefront size is specified as a power of two, so a value of **n** 9550 means a size of 2^ **n**. 9551- *call_convention* defaults to -1. 9552- *kernarg_segment_alignment*, *group_segment_alignment*, and 9553 *private_segment_alignment* default to 4. Note that alignments are specified 9554 as a power of 2, so a value of **n** means an alignment of 2^ **n**. 9555- *enable_wgp_mode* defaults to 1 if target feature ``cumode`` is disabled for 9556 GFX10 onwards. 9557- *enable_mem_ordered* defaults to 1 for GFX10 onwards. 9558 9559The *.amd_kernel_code_t* directive must be placed immediately after the 9560function label and before any instructions. 9561 9562For a full list of amd_kernel_code_t keys, refer to AMDGPU ABI document, 9563comments in lib/Target/AMDGPU/AmdKernelCodeT.h and test/CodeGen/AMDGPU/hsa.s. 9564 9565.. _amdgpu-amdhsa-assembler-example-v2: 9566 9567Code Object V2 Example Source Code (-mattr=-code-object-v3) 9568~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 9569 9570.. warning:: Code Object V2 is not the default code object version emitted by 9571 this version of LLVM. For a description of the directives supported with 9572 the default configuration (Code Object V3) see 9573 :ref:`amdgpu-amdhsa-assembler-example-v3`. 9574 9575Here is an example of a minimal assembly source file, defining one HSA kernel: 9576 9577.. code:: 9578 :number-lines: 9579 9580 .hsa_code_object_version 1,0 9581 .hsa_code_object_isa 9582 9583 .hsatext 9584 .globl hello_world 9585 .p2align 8 9586 .amdgpu_hsa_kernel hello_world 9587 9588 hello_world: 9589 9590 .amd_kernel_code_t 9591 enable_sgpr_kernarg_segment_ptr = 1 9592 is_ptr64 = 1 9593 compute_pgm_rsrc1_vgprs = 0 9594 compute_pgm_rsrc1_sgprs = 0 9595 compute_pgm_rsrc2_user_sgpr = 2 9596 compute_pgm_rsrc1_wgp_mode = 0 9597 compute_pgm_rsrc1_mem_ordered = 0 9598 compute_pgm_rsrc1_fwd_progress = 1 9599 .end_amd_kernel_code_t 9600 9601 s_load_dwordx2 s[0:1], s[0:1] 0x0 9602 v_mov_b32 v0, 3.14159 9603 s_waitcnt lgkmcnt(0) 9604 v_mov_b32 v1, s0 9605 v_mov_b32 v2, s1 9606 flat_store_dword v[1:2], v0 9607 s_endpgm 9608 .Lfunc_end0: 9609 .size hello_world, .Lfunc_end0-hello_world 9610 9611.. _amdgpu-amdhsa-assembler-predefined-symbols-v3: 9612 9613Code Object V3 Predefined Symbols (-mattr=+code-object-v3) 9614~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 9615 9616The AMDGPU assembler defines and updates some symbols automatically. These 9617symbols do not affect code generation. 9618 9619.amdgcn.gfx_generation_number 9620+++++++++++++++++++++++++++++ 9621 9622Set to the GFX major generation number of the target being assembled for. For 9623example, when assembling for a "GFX9" target this will be set to the integer 9624value "9". The possible GFX major generation numbers are presented in 9625:ref:`amdgpu-processors`. 9626 9627.amdgcn.gfx_generation_minor 9628++++++++++++++++++++++++++++ 9629 9630Set to the GFX minor generation number of the target being assembled for. For 9631example, when assembling for a "GFX810" target this will be set to the integer 9632value "1". The possible GFX minor generation numbers are presented in 9633:ref:`amdgpu-processors`. 9634 9635.amdgcn.gfx_generation_stepping 9636+++++++++++++++++++++++++++++++ 9637 9638Set to the GFX stepping generation number of the target being assembled for. 9639For example, when assembling for a "GFX704" target this will be set to the 9640integer value "4". The possible GFX stepping generation numbers are presented 9641in :ref:`amdgpu-processors`. 9642 9643.. _amdgpu-amdhsa-assembler-symbol-next_free_vgpr: 9644 9645.amdgcn.next_free_vgpr 9646++++++++++++++++++++++ 9647 9648Set to zero before assembly begins. At each instruction, if the current value 9649of this symbol is less than or equal to the maximum VGPR number explicitly 9650referenced within that instruction then the symbol value is updated to equal 9651that VGPR number plus one. 9652 9653May be used to set the `.amdhsa_next_free_vpgr` directive in 9654:ref:`amdhsa-kernel-directives-table`. 9655 9656May be set at any time, e.g. manually set to zero at the start of each kernel. 9657 9658.. _amdgpu-amdhsa-assembler-symbol-next_free_sgpr: 9659 9660.amdgcn.next_free_sgpr 9661++++++++++++++++++++++ 9662 9663Set to zero before assembly begins. At each instruction, if the current value 9664of this symbol is less than or equal the maximum SGPR number explicitly 9665referenced within that instruction then the symbol value is updated to equal 9666that SGPR number plus one. 9667 9668May be used to set the `.amdhsa_next_free_spgr` directive in 9669:ref:`amdhsa-kernel-directives-table`. 9670 9671May be set at any time, e.g. manually set to zero at the start of each kernel. 9672 9673.. _amdgpu-amdhsa-assembler-directives-v3: 9674 9675Code Object V3 Directives (-mattr=+code-object-v3) 9676~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 9677 9678Directives which begin with ``.amdgcn`` are valid for all ``amdgcn`` 9679architecture processors, and are not OS-specific. Directives which begin with 9680``.amdhsa`` are specific to ``amdgcn`` architecture processors when the 9681``amdhsa`` OS is specified. See :ref:`amdgpu-target-triples` and 9682:ref:`amdgpu-processors`. 9683 9684.amdgcn_target <target> 9685+++++++++++++++++++++++ 9686 9687Optional directive which declares the target supported by the containing 9688assembler source file. Valid values are described in 9689:ref:`amdgpu-amdhsa-code-object-target-identification`. Used by the assembler 9690to validate command-line options such as ``-triple``, ``-mcpu``, and those 9691which specify target features. 9692 9693.amdhsa_kernel <name> 9694+++++++++++++++++++++ 9695 9696Creates a correctly aligned AMDHSA kernel descriptor and a symbol, 9697``<name>.kd``, in the current location of the current section. Only valid when 9698the OS is ``amdhsa``. ``<name>`` must be a symbol that labels the first 9699instruction to execute, and does not need to be previously defined. 9700 9701Marks the beginning of a list of directives used to generate the bytes of a 9702kernel descriptor, as described in :ref:`amdgpu-amdhsa-kernel-descriptor`. 9703Directives which may appear in this list are described in 9704:ref:`amdhsa-kernel-directives-table`. Directives may appear in any order, must 9705be valid for the target being assembled for, and cannot be repeated. Directives 9706support the range of values specified by the field they reference in 9707:ref:`amdgpu-amdhsa-kernel-descriptor`. If a directive is not specified, it is 9708assumed to have its default value, unless it is marked as "Required", in which 9709case it is an error to omit the directive. This list of directives is 9710terminated by an ``.end_amdhsa_kernel`` directive. 9711 9712 .. table:: AMDHSA Kernel Assembler Directives 9713 :name: amdhsa-kernel-directives-table 9714 9715 ======================================================== =================== ============ =================== 9716 Directive Default Supported On Description 9717 ======================================================== =================== ============ =================== 9718 ``.amdhsa_group_segment_fixed_size`` 0 GFX6-GFX10 Controls GROUP_SEGMENT_FIXED_SIZE in 9719 :ref:`amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table`. 9720 ``.amdhsa_private_segment_fixed_size`` 0 GFX6-GFX10 Controls PRIVATE_SEGMENT_FIXED_SIZE in 9721 :ref:`amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table`. 9722 ``.amdhsa_user_sgpr_private_segment_buffer`` 0 GFX6-GFX10 Controls ENABLE_SGPR_PRIVATE_SEGMENT_BUFFER in 9723 :ref:`amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table`. 9724 ``.amdhsa_user_sgpr_dispatch_ptr`` 0 GFX6-GFX10 Controls ENABLE_SGPR_DISPATCH_PTR in 9725 :ref:`amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table`. 9726 ``.amdhsa_user_sgpr_queue_ptr`` 0 GFX6-GFX10 Controls ENABLE_SGPR_QUEUE_PTR in 9727 :ref:`amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table`. 9728 ``.amdhsa_user_sgpr_kernarg_segment_ptr`` 0 GFX6-GFX10 Controls ENABLE_SGPR_KERNARG_SEGMENT_PTR in 9729 :ref:`amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table`. 9730 ``.amdhsa_user_sgpr_dispatch_id`` 0 GFX6-GFX10 Controls ENABLE_SGPR_DISPATCH_ID in 9731 :ref:`amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table`. 9732 ``.amdhsa_user_sgpr_flat_scratch_init`` 0 GFX6-GFX10 Controls ENABLE_SGPR_FLAT_SCRATCH_INIT in 9733 :ref:`amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table`. 9734 ``.amdhsa_user_sgpr_private_segment_size`` 0 GFX6-GFX10 Controls ENABLE_SGPR_PRIVATE_SEGMENT_SIZE in 9735 :ref:`amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table`. 9736 ``.amdhsa_wavefront_size32`` Target GFX10 Controls ENABLE_WAVEFRONT_SIZE32 in 9737 Feature :ref:`amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table`. 9738 Specific 9739 (-wavefrontsize64) 9740 ``.amdhsa_system_sgpr_private_segment_wavefront_offset`` 0 GFX6-GFX10 Controls ENABLE_SGPR_PRIVATE_SEGMENT_WAVEFRONT_OFFSET in 9741 :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`. 9742 ``.amdhsa_system_sgpr_workgroup_id_x`` 1 GFX6-GFX10 Controls ENABLE_SGPR_WORKGROUP_ID_X in 9743 :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`. 9744 ``.amdhsa_system_sgpr_workgroup_id_y`` 0 GFX6-GFX10 Controls ENABLE_SGPR_WORKGROUP_ID_Y in 9745 :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`. 9746 ``.amdhsa_system_sgpr_workgroup_id_z`` 0 GFX6-GFX10 Controls ENABLE_SGPR_WORKGROUP_ID_Z in 9747 :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`. 9748 ``.amdhsa_system_sgpr_workgroup_info`` 0 GFX6-GFX10 Controls ENABLE_SGPR_WORKGROUP_INFO in 9749 :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`. 9750 ``.amdhsa_system_vgpr_workitem_id`` 0 GFX6-GFX10 Controls ENABLE_VGPR_WORKITEM_ID in 9751 :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`. 9752 Possible values are defined in 9753 :ref:`amdgpu-amdhsa-system-vgpr-work-item-id-enumeration-values-table`. 9754 ``.amdhsa_next_free_vgpr`` Required GFX6-GFX10 Maximum VGPR number explicitly referenced, plus one. 9755 Used to calculate GRANULATED_WORKITEM_VGPR_COUNT in 9756 :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`. 9757 ``.amdhsa_next_free_sgpr`` Required GFX6-GFX10 Maximum SGPR number explicitly referenced, plus one. 9758 Used to calculate GRANULATED_WAVEFRONT_SGPR_COUNT in 9759 :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`. 9760 ``.amdhsa_reserve_vcc`` 1 GFX6-GFX10 Whether the kernel may use the special VCC SGPR. 9761 Used to calculate GRANULATED_WAVEFRONT_SGPR_COUNT in 9762 :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`. 9763 ``.amdhsa_reserve_flat_scratch`` 1 GFX7-GFX10 Whether the kernel may use flat instructions to access 9764 scratch memory. Used to calculate 9765 GRANULATED_WAVEFRONT_SGPR_COUNT in 9766 :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`. 9767 ``.amdhsa_reserve_xnack_mask`` Target GFX8-GFX10 Whether the kernel may trigger XNACK replay. 9768 Feature Used to calculate GRANULATED_WAVEFRONT_SGPR_COUNT in 9769 Specific :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`. 9770 (+xnack) 9771 ``.amdhsa_float_round_mode_32`` 0 GFX6-GFX10 Controls FLOAT_ROUND_MODE_32 in 9772 :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`. 9773 Possible values are defined in 9774 :ref:`amdgpu-amdhsa-floating-point-rounding-mode-enumeration-values-table`. 9775 ``.amdhsa_float_round_mode_16_64`` 0 GFX6-GFX10 Controls FLOAT_ROUND_MODE_16_64 in 9776 :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`. 9777 Possible values are defined in 9778 :ref:`amdgpu-amdhsa-floating-point-rounding-mode-enumeration-values-table`. 9779 ``.amdhsa_float_denorm_mode_32`` 0 GFX6-GFX10 Controls FLOAT_DENORM_MODE_32 in 9780 :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`. 9781 Possible values are defined in 9782 :ref:`amdgpu-amdhsa-floating-point-denorm-mode-enumeration-values-table`. 9783 ``.amdhsa_float_denorm_mode_16_64`` 3 GFX6-GFX10 Controls FLOAT_DENORM_MODE_16_64 in 9784 :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`. 9785 Possible values are defined in 9786 :ref:`amdgpu-amdhsa-floating-point-denorm-mode-enumeration-values-table`. 9787 ``.amdhsa_dx10_clamp`` 1 GFX6-GFX10 Controls ENABLE_DX10_CLAMP in 9788 :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`. 9789 ``.amdhsa_ieee_mode`` 1 GFX6-GFX10 Controls ENABLE_IEEE_MODE in 9790 :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`. 9791 ``.amdhsa_fp16_overflow`` 0 GFX9-GFX10 Controls FP16_OVFL in 9792 :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`. 9793 ``.amdhsa_workgroup_processor_mode`` Target GFX10 Controls ENABLE_WGP_MODE in 9794 Feature :ref:`amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table`. 9795 Specific 9796 (-cumode) 9797 ``.amdhsa_memory_ordered`` 1 GFX10 Controls MEM_ORDERED in 9798 :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`. 9799 ``.amdhsa_forward_progress`` 0 GFX10 Controls FWD_PROGRESS in 9800 :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`. 9801 ``.amdhsa_exception_fp_ieee_invalid_op`` 0 GFX6-GFX10 Controls ENABLE_EXCEPTION_IEEE_754_FP_INVALID_OPERATION in 9802 :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`. 9803 ``.amdhsa_exception_fp_denorm_src`` 0 GFX6-GFX10 Controls ENABLE_EXCEPTION_FP_DENORMAL_SOURCE in 9804 :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`. 9805 ``.amdhsa_exception_fp_ieee_div_zero`` 0 GFX6-GFX10 Controls ENABLE_EXCEPTION_IEEE_754_FP_DIVISION_BY_ZERO in 9806 :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`. 9807 ``.amdhsa_exception_fp_ieee_overflow`` 0 GFX6-GFX10 Controls ENABLE_EXCEPTION_IEEE_754_FP_OVERFLOW in 9808 :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`. 9809 ``.amdhsa_exception_fp_ieee_underflow`` 0 GFX6-GFX10 Controls ENABLE_EXCEPTION_IEEE_754_FP_UNDERFLOW in 9810 :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`. 9811 ``.amdhsa_exception_fp_ieee_inexact`` 0 GFX6-GFX10 Controls ENABLE_EXCEPTION_IEEE_754_FP_INEXACT in 9812 :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`. 9813 ``.amdhsa_exception_int_div_zero`` 0 GFX6-GFX10 Controls ENABLE_EXCEPTION_INT_DIVIDE_BY_ZERO in 9814 :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`. 9815 ======================================================== =================== ============ =================== 9816 9817.amdgpu_metadata 9818++++++++++++++++ 9819 9820Optional directive which declares the contents of the ``NT_AMDGPU_METADATA`` 9821note record (see :ref:`amdgpu-elf-note-records-table-v3`). 9822 9823The contents must be in the [YAML]_ markup format, with the same structure and 9824semantics described in :ref:`amdgpu-amdhsa-code-object-metadata-v3`. 9825 9826This directive is terminated by an ``.end_amdgpu_metadata`` directive. 9827 9828.. _amdgpu-amdhsa-assembler-example-v3: 9829 9830Code Object V3 Example Source Code (-mattr=+code-object-v3) 9831~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 9832 9833Here is an example of a minimal assembly source file, defining one HSA kernel: 9834 9835.. code:: 9836 :number-lines: 9837 9838 .amdgcn_target "amdgcn-amd-amdhsa--gfx900+xnack" // optional 9839 9840 .text 9841 .globl hello_world 9842 .p2align 8 9843 .type hello_world,@function 9844 hello_world: 9845 s_load_dwordx2 s[0:1], s[0:1] 0x0 9846 v_mov_b32 v0, 3.14159 9847 s_waitcnt lgkmcnt(0) 9848 v_mov_b32 v1, s0 9849 v_mov_b32 v2, s1 9850 flat_store_dword v[1:2], v0 9851 s_endpgm 9852 .Lfunc_end0: 9853 .size hello_world, .Lfunc_end0-hello_world 9854 9855 .rodata 9856 .p2align 6 9857 .amdhsa_kernel hello_world 9858 .amdhsa_user_sgpr_kernarg_segment_ptr 1 9859 .amdhsa_next_free_vgpr .amdgcn.next_free_vgpr 9860 .amdhsa_next_free_sgpr .amdgcn.next_free_sgpr 9861 .end_amdhsa_kernel 9862 9863 .amdgpu_metadata 9864 --- 9865 amdhsa.version: 9866 - 1 9867 - 0 9868 amdhsa.kernels: 9869 - .name: hello_world 9870 .symbol: hello_world.kd 9871 .kernarg_segment_size: 48 9872 .group_segment_fixed_size: 0 9873 .private_segment_fixed_size: 0 9874 .kernarg_segment_align: 4 9875 .wavefront_size: 64 9876 .sgpr_count: 2 9877 .vgpr_count: 3 9878 .max_flat_workgroup_size: 256 9879 ... 9880 .end_amdgpu_metadata 9881 9882If an assembly source file contains multiple kernels and/or functions, the 9883:ref:`amdgpu-amdhsa-assembler-symbol-next_free_vgpr` and 9884:ref:`amdgpu-amdhsa-assembler-symbol-next_free_sgpr` symbols may be reset using 9885the ``.set <symbol>, <expression>`` directive. For example, in the case of two 9886kernels, where ``function1`` is only called from ``kernel1`` it is sufficient 9887to group the function with the kernel that calls it and reset the symbols 9888between the two connected components: 9889 9890.. code:: 9891 :number-lines: 9892 9893 .amdgcn_target "amdgcn-amd-amdhsa--gfx900+xnack" // optional 9894 9895 // gpr tracking symbols are implicitly set to zero 9896 9897 .text 9898 .globl kern0 9899 .p2align 8 9900 .type kern0,@function 9901 kern0: 9902 // ... 9903 s_endpgm 9904 .Lkern0_end: 9905 .size kern0, .Lkern0_end-kern0 9906 9907 .rodata 9908 .p2align 6 9909 .amdhsa_kernel kern0 9910 // ... 9911 .amdhsa_next_free_vgpr .amdgcn.next_free_vgpr 9912 .amdhsa_next_free_sgpr .amdgcn.next_free_sgpr 9913 .end_amdhsa_kernel 9914 9915 // reset symbols to begin tracking usage in func1 and kern1 9916 .set .amdgcn.next_free_vgpr, 0 9917 .set .amdgcn.next_free_sgpr, 0 9918 9919 .text 9920 .hidden func1 9921 .global func1 9922 .p2align 2 9923 .type func1,@function 9924 func1: 9925 // ... 9926 s_setpc_b64 s[30:31] 9927 .Lfunc1_end: 9928 .size func1, .Lfunc1_end-func1 9929 9930 .globl kern1 9931 .p2align 8 9932 .type kern1,@function 9933 kern1: 9934 // ... 9935 s_getpc_b64 s[4:5] 9936 s_add_u32 s4, s4, func1@rel32@lo+4 9937 s_addc_u32 s5, s5, func1@rel32@lo+4 9938 s_swappc_b64 s[30:31], s[4:5] 9939 // ... 9940 s_endpgm 9941 .Lkern1_end: 9942 .size kern1, .Lkern1_end-kern1 9943 9944 .rodata 9945 .p2align 6 9946 .amdhsa_kernel kern1 9947 // ... 9948 .amdhsa_next_free_vgpr .amdgcn.next_free_vgpr 9949 .amdhsa_next_free_sgpr .amdgcn.next_free_sgpr 9950 .end_amdhsa_kernel 9951 9952These symbols cannot identify connected components in order to automatically 9953track the usage for each kernel. However, in some cases careful organization of 9954the kernels and functions in the source file means there is minimal additional 9955effort required to accurately calculate GPR usage. 9956 9957Additional Documentation 9958======================== 9959 9960.. [AMD-RADEON-HD-2000-3000] `AMD R6xx shader ISA <http://developer.amd.com/wordpress/media/2012/10/R600_Instruction_Set_Architecture.pdf>`__ 9961.. [AMD-RADEON-HD-4000] `AMD R7xx shader ISA <http://developer.amd.com/wordpress/media/2012/10/R700-Family_Instruction_Set_Architecture.pdf>`__ 9962.. [AMD-RADEON-HD-5000] `AMD Evergreen shader ISA <http://developer.amd.com/wordpress/media/2012/10/AMD_Evergreen-Family_Instruction_Set_Architecture.pdf>`__ 9963.. [AMD-RADEON-HD-6000] `AMD Cayman/Trinity shader ISA <http://developer.amd.com/wordpress/media/2012/10/AMD_HD_6900_Series_Instruction_Set_Architecture.pdf>`__ 9964.. [AMD-GCN-GFX6] `AMD Southern Islands Series ISA <http://developer.amd.com/wordpress/media/2012/12/AMD_Southern_Islands_Instruction_Set_Architecture.pdf>`__ 9965.. [AMD-GCN-GFX7] `AMD Sea Islands Series ISA <http://developer.amd.com/wordpress/media/2013/07/AMD_Sea_Islands_Instruction_Set_Architecture.pdf>`_ 9966.. [AMD-GCN-GFX8] `AMD GCN3 Instruction Set Architecture <http://amd-dev.wpengine.netdna-cdn.com/wordpress/media/2013/12/AMD_GCN3_Instruction_Set_Architecture_rev1.1.pdf>`__ 9967.. [AMD-GCN-GFX9] `AMD "Vega" Instruction Set Architecture <http://developer.amd.com/wordpress/media/2013/12/Vega_Shader_ISA_28July2017.pdf>`__ 9968.. [AMD-GCN-GFX10] `AMD "RDNA 1.0" Instruction Set Architecture <https://gpuopen.com/wp-content/uploads/2019/08/RDNA_Shader_ISA_5August2019.pdf>`__ 9969.. [AMD-ROCm] `ROCm: Open Platform for Development, Discovery and Education Around GPU Computing <http://gpuopen.com/compute-product/rocm/>`__ 9970.. [AMD-ROCm-github] `ROCm github <http://github.com/RadeonOpenCompute>`__ 9971.. [HSA] `Heterogeneous System Architecture (HSA) Foundation <http://www.hsafoundation.com/>`__ 9972.. [HIP] `HIP Programming Guide <https://rocm-documentation.readthedocs.io/en/latest/Programming_Guides/Programming-Guides.html#hip-programing-guide>`__ 9973.. [ELF] `Executable and Linkable Format (ELF) <http://www.sco.com/developers/gabi/>`__ 9974.. [DWARF] `DWARF Debugging Information Format <http://dwarfstd.org/>`__ 9975.. [YAML] `YAML Ain't Markup Language (YAML™) Version 1.2 <http://www.yaml.org/spec/1.2/spec.html>`__ 9976.. [MsgPack] `Message Pack <http://www.msgpack.org/>`__ 9977.. [SEMVER] `Semantic Versioning <https://semver.org/>`__ 9978.. [OpenCL] `The OpenCL Specification Version 2.0 <http://www.khronos.org/registry/cl/specs/opencl-2.0.pdf>`__ 9979.. [HRF] `Heterogeneous-race-free Memory Models <http://benedictgaster.org/wp-content/uploads/2014/01/asplos269-FINAL.pdf>`__ 9980.. [CLANG-ATTR] `Attributes in Clang <https://clang.llvm.org/docs/AttributeReference.html>`__ 9981