1=============================
2User Guide for AMDGPU Backend
3=============================
4
5.. contents::
6   :local:
7
8Introduction
9============
10
11The AMDGPU backend provides ISA code generation for AMD GPUs, starting with the
12R600 family up until the current GCN families. It lives in the
13``llvm/lib/Target/AMDGPU`` directory.
14
15LLVM
16====
17
18.. _amdgpu-target-triples:
19
20Target Triples
21--------------
22
23Use the ``clang -target <Architecture>-<Vendor>-<OS>-<Environment>`` option to
24specify the target triple:
25
26  .. table:: AMDGPU Architectures
27     :name: amdgpu-architecture-table
28
29     ============ ==============================================================
30     Architecture Description
31     ============ ==============================================================
32     ``r600``     AMD GPUs HD2XXX-HD6XXX for graphics and compute shaders.
33     ``amdgcn``   AMD GPUs GCN GFX6 onwards for graphics and compute shaders.
34     ============ ==============================================================
35
36  .. table:: AMDGPU Vendors
37     :name: amdgpu-vendor-table
38
39     ============ ==============================================================
40     Vendor       Description
41     ============ ==============================================================
42     ``amd``      Can be used for all AMD GPU usage.
43     ``mesa3d``   Can be used if the OS is ``mesa3d``.
44     ============ ==============================================================
45
46  .. table:: AMDGPU Operating Systems
47     :name: amdgpu-os-table
48
49     ============== ============================================================
50     OS             Description
51     ============== ============================================================
52     *<empty>*      Defaults to the *unknown* OS.
53     ``amdhsa``     Compute kernels executed on HSA [HSA]_ compatible runtimes
54                    such as AMD's ROCm [AMD-ROCm]_.
55     ``amdpal``     Graphic shaders and compute kernels executed on AMD PAL
56                    runtime.
57     ``mesa3d``     Graphic shaders and compute kernels executed on Mesa 3D
58                    runtime.
59     ============== ============================================================
60
61  .. table:: AMDGPU Environments
62     :name: amdgpu-environment-table
63
64     ============ ==============================================================
65     Environment  Description
66     ============ ==============================================================
67     *<empty>*    Default.
68     ============ ==============================================================
69
70.. _amdgpu-processors:
71
72Processors
73----------
74
75Use the ``clang -mcpu <Processor>`` option to specify the AMDGPU processor. The
76names from both the *Processor* and *Alternative Processor* can be used.
77
78  .. table:: AMDGPU Processors
79     :name: amdgpu-processor-table
80
81     =========== =============== ============ ===== ================= ======= ======================
82     Processor   Alternative     Target       dGPU/ Target            ROCm    Example
83                 Processor       Triple       APU   Features          Support Products
84                                 Architecture       Supported
85                                                    [Default]
86     =========== =============== ============ ===== ================= ======= ======================
87     **Radeon HD 2000/3000 Series (R600)** [AMD-RADEON-HD-2000-3000]_
88     -----------------------------------------------------------------------------------------------
89     ``r600``                    ``r600``     dGPU
90     ``r630``                    ``r600``     dGPU
91     ``rs880``                   ``r600``     dGPU
92     ``rv670``                   ``r600``     dGPU
93     **Radeon HD 4000 Series (R700)** [AMD-RADEON-HD-4000]_
94     -----------------------------------------------------------------------------------------------
95     ``rv710``                   ``r600``     dGPU
96     ``rv730``                   ``r600``     dGPU
97     ``rv770``                   ``r600``     dGPU
98     **Radeon HD 5000 Series (Evergreen)** [AMD-RADEON-HD-5000]_
99     -----------------------------------------------------------------------------------------------
100     ``cedar``                   ``r600``     dGPU
101     ``cypress``                 ``r600``     dGPU
102     ``juniper``                 ``r600``     dGPU
103     ``redwood``                 ``r600``     dGPU
104     ``sumo``                    ``r600``     dGPU
105     **Radeon HD 6000 Series (Northern Islands)** [AMD-RADEON-HD-6000]_
106     -----------------------------------------------------------------------------------------------
107     ``barts``                   ``r600``     dGPU
108     ``caicos``                  ``r600``     dGPU
109     ``cayman``                  ``r600``     dGPU
110     ``turks``                   ``r600``     dGPU
111     **GCN GFX6 (Southern Islands (SI))** [AMD-GCN-GFX6]_
112     -----------------------------------------------------------------------------------------------
113     ``gfx600``  - ``tahiti``    ``amdgcn``   dGPU
114     ``gfx601``  - ``hainan``    ``amdgcn``   dGPU
115                 - ``oland``
116                 - ``pitcairn``
117                 - ``verde``
118     **GCN GFX7 (Sea Islands (CI))** [AMD-GCN-GFX7]_
119     -----------------------------------------------------------------------------------------------
120     ``gfx700``  - ``kaveri``    ``amdgcn``   APU                             - A6-7000
121                                                                              - A6 Pro-7050B
122                                                                              - A8-7100
123                                                                              - A8 Pro-7150B
124                                                                              - A10-7300
125                                                                              - A10 Pro-7350B
126                                                                              - FX-7500
127                                                                              - A8-7200P
128                                                                              - A10-7400P
129                                                                              - FX-7600P
130     ``gfx701``  - ``hawaii``    ``amdgcn``   dGPU                    ROCm    - FirePro W8100
131                                                                              - FirePro W9100
132                                                                              - FirePro S9150
133                                                                              - FirePro S9170
134     ``gfx702``                  ``amdgcn``   dGPU                    ROCm    - Radeon R9 290
135                                                                              - Radeon R9 290x
136                                                                              - Radeon R390
137                                                                              - Radeon R390x
138     ``gfx703``  - ``kabini``    ``amdgcn``   APU                             - E1-2100
139                 - ``mullins``                                                - E1-2200
140                                                                              - E1-2500
141                                                                              - E2-3000
142                                                                              - E2-3800
143                                                                              - A4-5000
144                                                                              - A4-5100
145                                                                              - A6-5200
146                                                                              - A4 Pro-3340B
147     ``gfx704``  - ``bonaire``   ``amdgcn``   dGPU                            - Radeon HD 7790
148                                                                              - Radeon HD 8770
149                                                                              - R7 260
150                                                                              - R7 260X
151     **GCN GFX8 (Volcanic Islands (VI))** [AMD-GCN-GFX8]_
152     -----------------------------------------------------------------------------------------------
153     ``gfx801``  - ``carrizo``   ``amdgcn``   APU   - xnack                   - A6-8500P
154                                                      [on]                    - Pro A6-8500B
155                                                                              - A8-8600P
156                                                                              - Pro A8-8600B
157                                                                              - FX-8800P
158                                                                              - Pro A12-8800B
159     \                           ``amdgcn``   APU   - xnack           ROCm    - A10-8700P
160                                                      [on]                    - Pro A10-8700B
161                                                                              - A10-8780P
162     \                           ``amdgcn``   APU   - xnack                   - A10-9600P
163                                                      [on]                    - A10-9630P
164                                                                              - A12-9700P
165                                                                              - A12-9730P
166                                                                              - FX-9800P
167                                                                              - FX-9830P
168     \                           ``amdgcn``   APU   - xnack                   - E2-9010
169                                                      [on]                    - A6-9210
170                                                                              - A9-9410
171     ``gfx802``  - ``iceland``   ``amdgcn``   dGPU  - xnack           ROCm    - FirePro S7150
172                 - ``tonga``                          [off]                   - FirePro S7100
173                                                                              - FirePro W7100
174                                                                              - Radeon R285
175                                                                              - Radeon R9 380
176                                                                              - Radeon R9 385
177                                                                              - Mobile FirePro
178                                                                                M7170
179     ``gfx803``  - ``fiji``      ``amdgcn``   dGPU  - xnack           ROCm    - Radeon R9 Nano
180                                                      [off]                   - Radeon R9 Fury
181                                                                              - Radeon R9 FuryX
182                                                                              - Radeon Pro Duo
183                                                                              - FirePro S9300x2
184                                                                              - Radeon Instinct MI8
185     \           - ``polaris10`` ``amdgcn``   dGPU  - xnack           ROCm    - Radeon RX 470
186                                                      [off]                   - Radeon RX 480
187                                                                              - Radeon Instinct MI6
188     \           - ``polaris11`` ``amdgcn``   dGPU  - xnack           ROCm    - Radeon RX 460
189                                                      [off]
190     ``gfx810``  - ``stoney``    ``amdgcn``   APU   - xnack
191                                                      [on]
192     **GCN GFX9** [AMD-GCN-GFX9]_
193     -----------------------------------------------------------------------------------------------
194     ``gfx900``                  ``amdgcn``   dGPU  - xnack           ROCm    - Radeon Vega
195                                                      [off]                     Frontier Edition
196                                                                              - Radeon RX Vega 56
197                                                                              - Radeon RX Vega 64
198                                                                              - Radeon RX Vega 64
199                                                                                Liquid
200                                                                              - Radeon Instinct MI25
201     ``gfx902``                  ``amdgcn``   APU   - xnack                   - Ryzen 3 2200G
202                                                      [on]                    - Ryzen 5 2400G
203     ``gfx904``                  ``amdgcn``   dGPU  - xnack                   *TBA*
204                                                      [off]
205                                                                              .. TODO::
206                                                                                 Add product
207                                                                                 names.
208     ``gfx906``                  ``amdgcn``   dGPU  - xnack                   - Radeon Instinct MI50
209                                                      [off]                   - Radeon Instinct MI60
210     ``gfx908``                  ``amdgcn``   dGPU  - xnack                   *TBA*
211                                                      [off]
212                                                      sram-ecc
213                                                      [on]
214     ``gfx909``                  ``amdgcn``   APU   - xnack                   *TBA* (Raven Ridge 2)
215                                                      [on]
216                                                                              .. TODO::
217                                                                                 Add product
218                                                                                 names.
219     **GCN GFX10** [AMD-GCN-GFX10]_
220     -----------------------------------------------------------------------------------------------
221     ``gfx1010``                 ``amdgcn``   dGPU  - xnack                   *TBA*
222                                                      [off]
223                                                    - wavefrontsize64
224                                                      [off]
225                                                    - cumode
226                                                      [off]
227                                                                              .. TODO::
228                                                                                 Add product
229                                                                                 names.
230     ``gfx1011``                 ``amdgcn``   dGPU  - xnack                   *TBA*
231                                                      [off]
232                                                    - wavefrontsize64
233                                                      [off]
234                                                    - cumode
235                                                      [off]
236                                                                              .. TODO::
237                                                                                 Add product
238                                                                                 names.
239     ``gfx1012``                 ``amdgcn``   dGPU  - xnack                   *TBA*
240                                                      [off]
241                                                    - wavefrontsize64
242                                                      [off]
243                                                    - cumode
244                                                      [off]
245                                                                              .. TODO::
246                                                                                 Add product
247                                                                                 names.
248     =========== =============== ============ ===== ================= ======= ======================
249
250.. _amdgpu-target-features:
251
252Target Features
253---------------
254
255Target features control how code is generated to support certain
256processor specific features. Not all target features are supported by
257all processors. The runtime must ensure that the features supported by
258the device used to execute the code match the features enabled when
259generating the code. A mismatch of features may result in incorrect
260execution, or a reduction in performance.
261
262The target features supported by each processor, and the default value
263used if not specified explicitly, is listed in
264:ref:`amdgpu-processor-table`.
265
266Use the ``clang -m[no-]<TargetFeature>`` option to specify the AMDGPU
267target features.
268
269For example:
270
271``-mxnack``
272  Enable the ``xnack`` feature.
273``-mno-xnack``
274  Disable the ``xnack`` feature.
275
276  .. table:: AMDGPU Target Features
277     :name: amdgpu-target-feature-table
278
279     ====================== ==================================================
280     Target Feature         Description
281     ====================== ==================================================
282     -m[no-]xnack           Enable/disable generating code that has
283                            memory clauses that are compatible with
284                            having XNACK replay enabled.
285
286                            This is used for demand paging and page
287                            migration. If XNACK replay is enabled in
288                            the device, then if a page fault occurs
289                            the code may execute incorrectly if the
290                            ``xnack`` feature is not enabled. Executing
291                            code that has the feature enabled on a
292                            device that does not have XNACK replay
293                            enabled will execute correctly, but may
294                            be less performant than code with the
295                            feature disabled.
296
297     -m[no-]sram-ecc        Enable/disable generating code that assumes SRAM
298                            ECC is enabled/disabled.
299
300     -m[no-]wavefrontsize64 Control the default wavefront size used when
301                            generating code for kernels. When disabled
302                            native wavefront size 32 is used, when enabled
303                            wavefront size 64 is used.
304
305     -m[no-]cumode          Control the default wavefront execution mode used
306                            when generating code for kernels. When disabled
307                            native WGP wavefront execution mode is used,
308                            when enabled CU wavefront execution mode is used
309                            (see :ref:`amdgpu-amdhsa-memory-model`).
310     ====================== ==================================================
311
312.. _amdgpu-address-spaces:
313
314Address Spaces
315--------------
316
317The AMDGPU architecture supports a number of memory address spaces. The address
318space names use the OpenCL standard names, with some additions.
319
320The AMDGPU address spaces correspond to architecture-specific LLVM address
321space numbers used in LLVM IR.
322
323The AMDGPU address spaces are described in
324:ref:`amdgpu-address-spaces-table`. Only 64-bit process address spaces are
325supported for the ``amdgcn`` target.
326
327  .. table:: AMDGPU Address Spaces
328     :name: amdgpu-address-spaces-table
329
330     ================================= =============== =========== ================ ======= ============================
331     ..                                                                                     64-Bit Process Address Space
332     --------------------------------- --------------- ----------- ---------------- ------------------------------------
333     Address Space Name                LLVM IR Address HSA Segment Hardware         Address NULL Value
334                                       Space Number    Name        Name             Size
335     ================================= =============== =========== ================ ======= ============================
336     Generic                           0               flat        flat             64      0x0000000000000000
337     Global                            1               global      global           64      0x0000000000000000
338     Region                            2               N/A         GDS              32      *not implemented for AMDHSA*
339     Local                             3               group       LDS              32      0xFFFFFFFF
340     Constant                          4               constant    *same as global* 64      0x0000000000000000
341     Private                           5               private     scratch          32      0x00000000
342     Constant 32-bit                   6               *TODO*
343     Buffer Fat Pointer (experimental) 7               *TODO*
344     ================================= =============== =========== ================ ======= ============================
345
346**Generic**
347  The generic address space uses the hardware flat address support available in
348  GFX7-GFX10. This uses two fixed ranges of virtual addresses (the private and
349  local apertures), that are outside the range of addressable global memory, to
350  map from a flat address to a private or local address.
351
352  FLAT instructions can take a flat address and access global, private
353  (scratch), and group (LDS) memory depending on if the address is within one
354  of the aperture ranges. Flat access to scratch requires hardware aperture
355  setup and setup in the kernel prologue (see
356  :ref:`amdgpu-amdhsa-kernel-prolog-flat-scratch`). Flat access to LDS requires
357  hardware aperture setup and M0 (GFX7-GFX8) register setup (see
358  :ref:`amdgpu-amdhsa-kernel-prolog-m0`).
359
360  To convert between a private or group address space address (termed a segment
361  address) and a flat address the base address of the corresponding aperture
362  can be used. For GFX7-GFX8 these are available in the
363  :ref:`amdgpu-amdhsa-hsa-aql-queue` the address of which can be obtained with
364  Queue Ptr SGPR (see :ref:`amdgpu-amdhsa-initial-kernel-execution-state`). For
365  GFX9-GFX10 the aperture base addresses are directly available as inline
366  constant registers ``SRC_SHARED_BASE/LIMIT`` and ``SRC_PRIVATE_BASE/LIMIT``.
367  In 64-bit address mode the aperture sizes are 2^32 bytes and the base is
368  aligned to 2^32 which makes it easier to convert from flat to segment or
369  segment to flat.
370
371  A global address space address has the same value when used as a flat address
372  so no conversion is needed.
373
374**Global and Constant**
375  The global and constant address spaces both use global virtual addresses,
376  which are the same virtual address space used by the CPU. However, some
377  virtual addresses may only be accessible to the CPU, some only accessible
378  by the GPU, and some by both.
379
380  Using the constant address space indicates that the data will not change
381  during the execution of the kernel. This allows scalar read instructions to
382  be used. The vector and scalar L1 caches are invalidated of volatile data
383  before each kernel dispatch execution to allow constant memory to change
384  values between kernel dispatches.
385
386**Region**
387  The region address space uses the hardware Global Data Store (GDS). All
388  wavefronts executing on the same device will access the same memory for any
389  given region address. However, the same region address accessed by wavefronts
390  executing on different devices will access different memory. It is higher
391  performance than global memory. It is allocated by the runtime. The data
392  store (DS) instructions can be used to access it.
393
394**Local**
395  The local address space uses the hardware Local Data Store (LDS) which is
396  automatically allocated when the hardware creates the wavefronts of a
397  work-group, and freed when all the wavefronts of a work-group have
398  terminated. All wavefronts belonging to the same work-group will access the
399  same memory for any given local address. However, the same local address
400  accessed by wavefronts belonging to different work-groups will access
401  different memory. It is higher performance than global memory. The data store
402  (DS) instructions can be used to access it.
403
404**Private**
405  The private address space uses the hardware scratch memory support which
406  automatically allocates memory when it creates a wavefront, and frees it when
407  a wavefronts terminates. The memory accessed by a lane of a wavefront for any
408  given private address will be different to the memory accessed by another lane
409  of the same or different wavefront for the same private address.
410
411  If a kernel dispatch uses scratch, then the hardware allocates memory from a
412  pool of backing memory allocated by the runtime for each wavefront. The lanes
413  of the wavefront access this using dword (4 byte) interleaving. The mapping
414  used from private address to backing memory address is:
415
416    ``wavefront-scratch-base +
417    ((private-address / 4) * wavefront-size * 4) +
418    (wavefront-lane-id * 4) + (private-address % 4)``
419
420  If each lane of a wavefront accesses the same private address, the
421  interleaving results in adjacent dwords being accessed and hence requires
422  fewer cache lines to be fetched.
423
424  There are different ways that the wavefront scratch base address is
425  determined by a wavefront (see
426  :ref:`amdgpu-amdhsa-initial-kernel-execution-state`).
427
428  Scratch memory can be accessed in an interleaved manner using buffer
429  instructions with the scratch buffer descriptor and per wavefront scratch
430  offset, by the scratch instructions, or by flat instructions. Multi-dword
431  access is not supported except by flat and scratch instructions in
432  GFX9-GFX10.
433
434**Constant 32-bit**
435  *TODO*
436
437**Buffer Fat Pointer**
438  The buffer fat pointer is an experimental address space that is currently
439  unsupported in the backend. It exposes a non-integral pointer that is in
440  the future intended to support the modelling of 128-bit buffer descriptors
441  plus a 32-bit offset into the buffer (in total encapsulating a 160-bit
442  *pointer*), allowing normal LLVM load/store/atomic operations to be used to
443  model the buffer descriptors used heavily in graphics workloads targeting
444  the backend.
445
446.. _amdgpu-memory-scopes:
447
448Memory Scopes
449-------------
450
451This section provides LLVM memory synchronization scopes supported by the AMDGPU
452backend memory model when the target triple OS is ``amdhsa`` (see
453:ref:`amdgpu-amdhsa-memory-model` and :ref:`amdgpu-target-triples`).
454
455The memory model supported is based on the HSA memory model [HSA]_ which is
456based in turn on HRF-indirect with scope inclusion [HRF]_. The happens-before
457relation is transitive over the synchronizes-with relation independent of scope,
458and synchronizes-with allows the memory scope instances to be inclusive (see
459table :ref:`amdgpu-amdhsa-llvm-sync-scopes-table`).
460
461This is different to the OpenCL [OpenCL]_ memory model which does not have scope
462inclusion and requires the memory scopes to exactly match. However, this
463is conservatively correct for OpenCL.
464
465  .. table:: AMDHSA LLVM Sync Scopes
466     :name: amdgpu-amdhsa-llvm-sync-scopes-table
467
468     ======================= ===================================================
469     LLVM Sync Scope         Description
470     ======================= ===================================================
471     *none*                  The default: ``system``.
472
473                             Synchronizes with, and participates in modification
474                             and seq_cst total orderings with, other operations
475                             (except image operations) for all address spaces
476                             (except private, or generic that accesses private)
477                             provided the other operation's sync scope is:
478
479                             - ``system``.
480                             - ``agent`` and executed by a thread on the same
481                               agent.
482                             - ``workgroup`` and executed by a thread in the
483                               same work-group.
484                             - ``wavefront`` and executed by a thread in the
485                               same wavefront.
486
487     ``agent``               Synchronizes with, and participates in modification
488                             and seq_cst total orderings with, other operations
489                             (except image operations) for all address spaces
490                             (except private, or generic that accesses private)
491                             provided the other operation's sync scope is:
492
493                             - ``system`` or ``agent`` and executed by a thread
494                               on the same agent.
495                             - ``workgroup`` and executed by a thread in the
496                               same work-group.
497                             - ``wavefront`` and executed by a thread in the
498                               same wavefront.
499
500     ``workgroup``           Synchronizes with, and participates in modification
501                             and seq_cst total orderings with, other operations
502                             (except image operations) for all address spaces
503                             (except private, or generic that accesses private)
504                             provided the other operation's sync scope is:
505
506                             - ``system``, ``agent`` or ``workgroup`` and
507                               executed by a thread in the same work-group.
508                             - ``wavefront`` and executed by a thread in the
509                               same wavefront.
510
511     ``wavefront``           Synchronizes with, and participates in modification
512                             and seq_cst total orderings with, other operations
513                             (except image operations) for all address spaces
514                             (except private, or generic that accesses private)
515                             provided the other operation's sync scope is:
516
517                             - ``system``, ``agent``, ``workgroup`` or
518                               ``wavefront`` and executed by a thread in the
519                               same wavefront.
520
521     ``singlethread``        Only synchronizes with, and participates in
522                             modification and seq_cst total orderings with,
523                             other operations (except image operations) running
524                             in the same thread for all address spaces (for
525                             example, in signal handlers).
526
527     ``one-as``              Same as ``system`` but only synchronizes with other
528                             operations within the same address space.
529
530     ``agent-one-as``        Same as ``agent`` but only synchronizes with other
531                             operations within the same address space.
532
533     ``workgroup-one-as``    Same as ``workgroup`` but only synchronizes with
534                             other operations within the same address space.
535
536     ``wavefront-one-as``    Same as ``wavefront`` but only synchronizes with
537                             other operations within the same address space.
538
539     ``singlethread-one-as`` Same as ``singlethread`` but only synchronizes with
540                             other operations within the same address space.
541     ======================= ===================================================
542
543AMDGPU Intrinsics
544-----------------
545
546The AMDGPU backend implements the following LLVM IR intrinsics.
547
548*This section is WIP.*
549
550.. TODO::
551
552   List AMDGPU intrinsics.
553
554AMDGPU Attributes
555-----------------
556
557The AMDGPU backend supports the following LLVM IR attributes.
558
559  .. table:: AMDGPU LLVM IR Attributes
560     :name: amdgpu-llvm-ir-attributes-table
561
562     ======================================= ==========================================================
563     LLVM Attribute                          Description
564     ======================================= ==========================================================
565     "amdgpu-flat-work-group-size"="min,max" Specify the minimum and maximum flat work group sizes that
566                                             will be specified when the kernel is dispatched. Generated
567                                             by the ``amdgpu_flat_work_group_size`` CLANG attribute [CLANG-ATTR]_.
568     "amdgpu-implicitarg-num-bytes"="n"      Number of kernel argument bytes to add to the kernel
569                                             argument block size for the implicit arguments. This
570                                             varies by OS and language (for OpenCL see
571                                             :ref:`opencl-kernel-implicit-arguments-appended-for-amdhsa-os-table`).
572     "amdgpu-num-sgpr"="n"                   Specifies the number of SGPRs to use. Generated by
573                                             the ``amdgpu_num_sgpr`` CLANG attribute [CLANG-ATTR]_.
574     "amdgpu-num-vgpr"="n"                   Specifies the number of VGPRs to use. Generated by the
575                                             ``amdgpu_num_vgpr`` CLANG attribute [CLANG-ATTR]_.
576     "amdgpu-waves-per-eu"="m,n"             Specify the minimum and maximum number of waves per
577                                             execution unit. Generated by the ``amdgpu_waves_per_eu``
578                                             CLANG attribute [CLANG-ATTR]_.
579     "amdgpu-ieee" true/false.               Specify whether the function expects the IEEE field of the
580                                             mode register to be set on entry. Overrides the default for
581                                             the calling convention.
582     "amdgpu-dx10-clamp" true/false.         Specify whether the function expects the DX10_CLAMP field of
583                                             the mode register to be set on entry. Overrides the default
584                                             for the calling convention.
585     ======================================= ==========================================================
586
587Code Object
588===========
589
590The AMDGPU backend generates a standard ELF [ELF]_ relocatable code object that
591can be linked by ``lld`` to produce a standard ELF shared code object which can
592be loaded and executed on an AMDGPU target.
593
594Header
595------
596
597The AMDGPU backend uses the following ELF header:
598
599  .. table:: AMDGPU ELF Header
600     :name: amdgpu-elf-header-table
601
602     ========================== ===============================
603     Field                      Value
604     ========================== ===============================
605     ``e_ident[EI_CLASS]``      ``ELFCLASS64``
606     ``e_ident[EI_DATA]``       ``ELFDATA2LSB``
607     ``e_ident[EI_OSABI]``      - ``ELFOSABI_NONE``
608                                - ``ELFOSABI_AMDGPU_HSA``
609                                - ``ELFOSABI_AMDGPU_PAL``
610                                - ``ELFOSABI_AMDGPU_MESA3D``
611     ``e_ident[EI_ABIVERSION]`` - ``ELFABIVERSION_AMDGPU_HSA``
612                                - ``ELFABIVERSION_AMDGPU_PAL``
613                                - ``ELFABIVERSION_AMDGPU_MESA3D``
614     ``e_type``                 - ``ET_REL``
615                                - ``ET_DYN``
616     ``e_machine``              ``EM_AMDGPU``
617     ``e_entry``                0
618     ``e_flags``                See :ref:`amdgpu-elf-header-e_flags-table`
619     ========================== ===============================
620
621..
622
623  .. table:: AMDGPU ELF Header Enumeration Values
624     :name: amdgpu-elf-header-enumeration-values-table
625
626     =============================== =====
627     Name                            Value
628     =============================== =====
629     ``EM_AMDGPU``                   224
630     ``ELFOSABI_NONE``               0
631     ``ELFOSABI_AMDGPU_HSA``         64
632     ``ELFOSABI_AMDGPU_PAL``         65
633     ``ELFOSABI_AMDGPU_MESA3D``      66
634     ``ELFABIVERSION_AMDGPU_HSA``    1
635     ``ELFABIVERSION_AMDGPU_PAL``    0
636     ``ELFABIVERSION_AMDGPU_MESA3D`` 0
637     =============================== =====
638
639``e_ident[EI_CLASS]``
640  The ELF class is:
641
642  * ``ELFCLASS32`` for ``r600`` architecture.
643
644  * ``ELFCLASS64`` for ``amdgcn`` architecture which only supports 64-bit
645    process address space applications.
646
647``e_ident[EI_DATA]``
648  All AMDGPU targets use ``ELFDATA2LSB`` for little-endian byte ordering.
649
650``e_ident[EI_OSABI]``
651  One of the following AMDGPU architecture specific OS ABIs
652  (see :ref:`amdgpu-os-table`):
653
654  * ``ELFOSABI_NONE`` for *unknown* OS.
655
656  * ``ELFOSABI_AMDGPU_HSA`` for ``amdhsa`` OS.
657
658  * ``ELFOSABI_AMDGPU_PAL`` for ``amdpal`` OS.
659
660  * ``ELFOSABI_AMDGPU_MESA3D`` for ``mesa3D`` OS.
661
662``e_ident[EI_ABIVERSION]``
663  The ABI version of the AMDGPU architecture specific OS ABI to which the code
664  object conforms:
665
666  * ``ELFABIVERSION_AMDGPU_HSA`` is used to specify the version of AMD HSA
667    runtime ABI.
668
669  * ``ELFABIVERSION_AMDGPU_PAL`` is used to specify the version of AMD PAL
670    runtime ABI.
671
672  * ``ELFABIVERSION_AMDGPU_MESA3D`` is used to specify the version of AMD MESA
673    3D runtime ABI.
674
675``e_type``
676  Can be one of the following values:
677
678
679  ``ET_REL``
680    The type produced by the AMDGPU backend compiler as it is relocatable code
681    object.
682
683  ``ET_DYN``
684    The type produced by the linker as it is a shared code object.
685
686  The AMD HSA runtime loader requires a ``ET_DYN`` code object.
687
688``e_machine``
689  The value ``EM_AMDGPU`` is used for the machine for all processors supported
690  by the ``r600`` and ``amdgcn`` architectures (see
691  :ref:`amdgpu-processor-table`). The specific processor is specified in the
692  ``EF_AMDGPU_MACH`` bit field of the ``e_flags`` (see
693  :ref:`amdgpu-elf-header-e_flags-table`).
694
695``e_entry``
696  The entry point is 0 as the entry points for individual kernels must be
697  selected in order to invoke them through AQL packets.
698
699``e_flags``
700  The AMDGPU backend uses the following ELF header flags:
701
702  .. table:: AMDGPU ELF Header ``e_flags``
703     :name: amdgpu-elf-header-e_flags-table
704
705     ================================= ========== =============================
706     Name                              Value      Description
707     ================================= ========== =============================
708     **AMDGPU Processor Flag**                    See :ref:`amdgpu-processor-table`.
709     -------------------------------------------- -----------------------------
710     ``EF_AMDGPU_MACH``                0x000000ff AMDGPU processor selection
711                                                  mask for
712                                                  ``EF_AMDGPU_MACH_xxx`` values
713                                                  defined in
714                                                  :ref:`amdgpu-ef-amdgpu-mach-table`.
715     ``EF_AMDGPU_XNACK``               0x00000100 Indicates if the ``xnack``
716                                                  target feature is
717                                                  enabled for all code
718                                                  contained in the code object.
719                                                  If the processor
720                                                  does not support the
721                                                  ``xnack`` target
722                                                  feature then must
723                                                  be 0.
724                                                  See
725                                                  :ref:`amdgpu-target-features`.
726     ``EF_AMDGPU_SRAM_ECC``            0x00000200 Indicates if the ``sram-ecc``
727                                                  target feature is
728                                                  enabled for all code
729                                                  contained in the code object.
730                                                  If the processor
731                                                  does not support the
732                                                  ``sram-ecc`` target
733                                                  feature then must
734                                                  be 0.
735                                                  See
736                                                  :ref:`amdgpu-target-features`.
737     ================================= ========== =============================
738
739  .. table:: AMDGPU ``EF_AMDGPU_MACH`` Values
740     :name: amdgpu-ef-amdgpu-mach-table
741
742     ================================= ========== =============================
743     Name                              Value      Description (see
744                                                  :ref:`amdgpu-processor-table`)
745     ================================= ========== =============================
746     ``EF_AMDGPU_MACH_NONE``           0x000      *not specified*
747     ``EF_AMDGPU_MACH_R600_R600``      0x001      ``r600``
748     ``EF_AMDGPU_MACH_R600_R630``      0x002      ``r630``
749     ``EF_AMDGPU_MACH_R600_RS880``     0x003      ``rs880``
750     ``EF_AMDGPU_MACH_R600_RV670``     0x004      ``rv670``
751     ``EF_AMDGPU_MACH_R600_RV710``     0x005      ``rv710``
752     ``EF_AMDGPU_MACH_R600_RV730``     0x006      ``rv730``
753     ``EF_AMDGPU_MACH_R600_RV770``     0x007      ``rv770``
754     ``EF_AMDGPU_MACH_R600_CEDAR``     0x008      ``cedar``
755     ``EF_AMDGPU_MACH_R600_CYPRESS``   0x009      ``cypress``
756     ``EF_AMDGPU_MACH_R600_JUNIPER``   0x00a      ``juniper``
757     ``EF_AMDGPU_MACH_R600_REDWOOD``   0x00b      ``redwood``
758     ``EF_AMDGPU_MACH_R600_SUMO``      0x00c      ``sumo``
759     ``EF_AMDGPU_MACH_R600_BARTS``     0x00d      ``barts``
760     ``EF_AMDGPU_MACH_R600_CAICOS``    0x00e      ``caicos``
761     ``EF_AMDGPU_MACH_R600_CAYMAN``    0x00f      ``cayman``
762     ``EF_AMDGPU_MACH_R600_TURKS``     0x010      ``turks``
763     *reserved*                        0x011 -    Reserved for ``r600``
764                                       0x01f      architecture processors.
765     ``EF_AMDGPU_MACH_AMDGCN_GFX600``  0x020      ``gfx600``
766     ``EF_AMDGPU_MACH_AMDGCN_GFX601``  0x021      ``gfx601``
767     ``EF_AMDGPU_MACH_AMDGCN_GFX700``  0x022      ``gfx700``
768     ``EF_AMDGPU_MACH_AMDGCN_GFX701``  0x023      ``gfx701``
769     ``EF_AMDGPU_MACH_AMDGCN_GFX702``  0x024      ``gfx702``
770     ``EF_AMDGPU_MACH_AMDGCN_GFX703``  0x025      ``gfx703``
771     ``EF_AMDGPU_MACH_AMDGCN_GFX704``  0x026      ``gfx704``
772     *reserved*                        0x027      Reserved.
773     ``EF_AMDGPU_MACH_AMDGCN_GFX801``  0x028      ``gfx801``
774     ``EF_AMDGPU_MACH_AMDGCN_GFX802``  0x029      ``gfx802``
775     ``EF_AMDGPU_MACH_AMDGCN_GFX803``  0x02a      ``gfx803``
776     ``EF_AMDGPU_MACH_AMDGCN_GFX810``  0x02b      ``gfx810``
777     ``EF_AMDGPU_MACH_AMDGCN_GFX900``  0x02c      ``gfx900``
778     ``EF_AMDGPU_MACH_AMDGCN_GFX902``  0x02d      ``gfx902``
779     ``EF_AMDGPU_MACH_AMDGCN_GFX904``  0x02e      ``gfx904``
780     ``EF_AMDGPU_MACH_AMDGCN_GFX906``  0x02f      ``gfx906``
781     ``EF_AMDGPU_MACH_AMDGCN_GFX908``  0x030      ``gfx908``
782     ``EF_AMDGPU_MACH_AMDGCN_GFX909``  0x031      ``gfx909``
783     *reserved*                        0x032      Reserved.
784     ``EF_AMDGPU_MACH_AMDGCN_GFX1010`` 0x033      ``gfx1010``
785     ``EF_AMDGPU_MACH_AMDGCN_GFX1011`` 0x034      ``gfx1011``
786     ``EF_AMDGPU_MACH_AMDGCN_GFX1012`` 0x035      ``gfx1012``
787     ================================= ========== =============================
788
789Sections
790--------
791
792An AMDGPU target ELF code object has the standard ELF sections which include:
793
794  .. table:: AMDGPU ELF Sections
795     :name: amdgpu-elf-sections-table
796
797     ================== ================ =================================
798     Name               Type             Attributes
799     ================== ================ =================================
800     ``.bss``           ``SHT_NOBITS``   ``SHF_ALLOC`` + ``SHF_WRITE``
801     ``.data``          ``SHT_PROGBITS`` ``SHF_ALLOC`` + ``SHF_WRITE``
802     ``.debug_``\ *\**  ``SHT_PROGBITS`` *none*
803     ``.dynamic``       ``SHT_DYNAMIC``  ``SHF_ALLOC``
804     ``.dynstr``        ``SHT_PROGBITS`` ``SHF_ALLOC``
805     ``.dynsym``        ``SHT_PROGBITS`` ``SHF_ALLOC``
806     ``.got``           ``SHT_PROGBITS`` ``SHF_ALLOC`` + ``SHF_WRITE``
807     ``.hash``          ``SHT_HASH``     ``SHF_ALLOC``
808     ``.note``          ``SHT_NOTE``     *none*
809     ``.rela``\ *name*  ``SHT_RELA``     *none*
810     ``.rela.dyn``      ``SHT_RELA``     *none*
811     ``.rodata``        ``SHT_PROGBITS`` ``SHF_ALLOC``
812     ``.shstrtab``      ``SHT_STRTAB``   *none*
813     ``.strtab``        ``SHT_STRTAB``   *none*
814     ``.symtab``        ``SHT_SYMTAB``   *none*
815     ``.text``          ``SHT_PROGBITS`` ``SHF_ALLOC`` + ``SHF_EXECINSTR``
816     ================== ================ =================================
817
818These sections have their standard meanings (see [ELF]_) and are only generated
819if needed.
820
821``.debug``\ *\**
822  The standard DWARF sections. See :ref:`amdgpu-dwarf` for information on the
823  DWARF produced by the AMDGPU backend.
824
825``.dynamic``, ``.dynstr``, ``.dynsym``, ``.hash``
826  The standard sections used by a dynamic loader.
827
828``.note``
829  See :ref:`amdgpu-note-records` for the note records supported by the AMDGPU
830  backend.
831
832``.rela``\ *name*, ``.rela.dyn``
833  For relocatable code objects, *name* is the name of the section that the
834  relocation records apply. For example, ``.rela.text`` is the section name for
835  relocation records associated with the ``.text`` section.
836
837  For linked shared code objects, ``.rela.dyn`` contains all the relocation
838  records from each of the relocatable code object's ``.rela``\ *name* sections.
839
840  See :ref:`amdgpu-relocation-records` for the relocation records supported by
841  the AMDGPU backend.
842
843``.text``
844  The executable machine code for the kernels and functions they call. Generated
845  as position independent code. See :ref:`amdgpu-code-conventions` for
846  information on conventions used in the isa generation.
847
848.. _amdgpu-note-records:
849
850Note Records
851------------
852
853The AMDGPU backend code object contains ELF note records in the ``.note``
854section. The set of generated notes and their semantics depend on the code
855object version; see :ref:`amdgpu-note-records-v2` and
856:ref:`amdgpu-note-records-v3`.
857
858As required by ``ELFCLASS32`` and ``ELFCLASS64``, minimal zero byte padding
859must be generated after the ``name`` field to ensure the ``desc`` field is 4
860byte aligned. In addition, minimal zero byte padding must be generated to
861ensure the ``desc`` field size is a multiple of 4 bytes. The ``sh_addralign``
862field of the ``.note`` section must be at least 4 to indicate at least 8 byte
863alignment.
864
865.. _amdgpu-note-records-v2:
866
867Code Object V2 Note Records (-mattr=-code-object-v3)
868~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
869
870.. warning:: Code Object V2 is not the default code object version emitted by
871  this version of LLVM. For a description of the notes generated with the
872  default configuration (Code Object V3) see :ref:`amdgpu-note-records-v3`.
873
874The AMDGPU backend code object uses the following ELF note record in the
875``.note`` section when compiling for Code Object V2 (-mattr=-code-object-v3).
876
877Additional note records may be present, but any which are not documented here
878are deprecated and should not be used.
879
880  .. table:: AMDGPU Code Object V2 ELF Note Records
881     :name: amdgpu-elf-note-records-table-v2
882
883     ===== ============================== ======================================
884     Name  Type                           Description
885     ===== ============================== ======================================
886     "AMD" ``NT_AMD_AMDGPU_HSA_METADATA`` <metadata null terminated string>
887     ===== ============================== ======================================
888
889..
890
891  .. table:: AMDGPU Code Object V2 ELF Note Record Enumeration Values
892     :name: amdgpu-elf-note-record-enumeration-values-table-v2
893
894     ============================== =====
895     Name                           Value
896     ============================== =====
897     *reserved*                       0-9
898     ``NT_AMD_AMDGPU_HSA_METADATA``    10
899     *reserved*                        11
900     ============================== =====
901
902``NT_AMD_AMDGPU_HSA_METADATA``
903  Specifies extensible metadata associated with the code objects executed on HSA
904  [HSA]_ compatible runtimes such as AMD's ROCm [AMD-ROCm]_. It is required when
905  the target triple OS is ``amdhsa`` (see :ref:`amdgpu-target-triples`). See
906  :ref:`amdgpu-amdhsa-code-object-metadata-v2` for the syntax of the code
907  object metadata string.
908
909.. _amdgpu-note-records-v3:
910
911Code Object V3 Note Records (-mattr=+code-object-v3)
912~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
913
914The AMDGPU backend code object uses the following ELF note record in the
915``.note`` section when compiling for Code Object V3 (-mattr=+code-object-v3).
916
917Additional note records may be present, but any which are not documented here
918are deprecated and should not be used.
919
920  .. table:: AMDGPU Code Object V3 ELF Note Records
921     :name: amdgpu-elf-note-records-table-v3
922
923     ======== ============================== ======================================
924     Name     Type                           Description
925     ======== ============================== ======================================
926     "AMDGPU" ``NT_AMDGPU_METADATA``         Metadata in Message Pack [MsgPack]_
927                                             binary format.
928     ======== ============================== ======================================
929
930..
931
932  .. table:: AMDGPU Code Object V3 ELF Note Record Enumeration Values
933     :name: amdgpu-elf-note-record-enumeration-values-table-v3
934
935     ============================== =====
936     Name                           Value
937     ============================== =====
938     *reserved*                     0-31
939     ``NT_AMDGPU_METADATA``         32
940     ============================== =====
941
942``NT_AMDGPU_METADATA``
943  Specifies extensible metadata associated with an AMDGPU code
944  object. It is encoded as a map in the Message Pack [MsgPack]_ binary
945  data format. See :ref:`amdgpu-amdhsa-code-object-metadata-v3` for the
946  map keys defined for the ``amdhsa`` OS.
947
948.. _amdgpu-symbols:
949
950Symbols
951-------
952
953Symbols include the following:
954
955  .. table:: AMDGPU ELF Symbols
956     :name: amdgpu-elf-symbols-table
957
958     ===================== ================== ================ ==================
959     Name                  Type               Section          Description
960     ===================== ================== ================ ==================
961     *link-name*           ``STT_OBJECT``     - ``.data``      Global variable
962                                              - ``.rodata``
963                                              - ``.bss``
964     *link-name*\ ``.kd``  ``STT_OBJECT``     - ``.rodata``    Kernel descriptor
965     *link-name*           ``STT_FUNC``       - ``.text``      Kernel entry point
966     *link-name*           ``STT_OBJECT``     - SHN_AMDGPU_LDS Global variable in LDS
967     ===================== ================== ================ ==================
968
969Global variable
970  Global variables both used and defined by the compilation unit.
971
972  If the symbol is defined in the compilation unit then it is allocated in the
973  appropriate section according to if it has initialized data or is readonly.
974
975  If the symbol is external then its section is ``STN_UNDEF`` and the loader
976  will resolve relocations using the definition provided by another code object
977  or explicitly defined by the runtime.
978
979  If the symbol resides in local/group memory (LDS) then its section is the
980  special processor-specific section name ``SHN_AMDGPU_LDS``, and the
981  ``st_value`` field describes alignment requirements as it does for common
982  symbols.
983
984  .. TODO::
985
986     Add description of linked shared object symbols. Seems undefined symbols
987     are marked as STT_NOTYPE.
988
989Kernel descriptor
990  Every HSA kernel has an associated kernel descriptor. It is the address of the
991  kernel descriptor that is used in the AQL dispatch packet used to invoke the
992  kernel, not the kernel entry point. The layout of the HSA kernel descriptor is
993  defined in :ref:`amdgpu-amdhsa-kernel-descriptor`.
994
995Kernel entry point
996  Every HSA kernel also has a symbol for its machine code entry point.
997
998.. _amdgpu-relocation-records:
999
1000Relocation Records
1001------------------
1002
1003AMDGPU backend generates ``Elf64_Rela`` relocation records. Supported
1004relocatable fields are:
1005
1006``word32``
1007  This specifies a 32-bit field occupying 4 bytes with arbitrary byte
1008  alignment. These values use the same byte order as other word values in the
1009  AMDGPU architecture.
1010
1011``word64``
1012  This specifies a 64-bit field occupying 8 bytes with arbitrary byte
1013  alignment. These values use the same byte order as other word values in the
1014  AMDGPU architecture.
1015
1016Following notations are used for specifying relocation calculations:
1017
1018**A**
1019  Represents the addend used to compute the value of the relocatable field.
1020
1021**G**
1022  Represents the offset into the global offset table at which the relocation
1023  entry's symbol will reside during execution.
1024
1025**GOT**
1026  Represents the address of the global offset table.
1027
1028**P**
1029  Represents the place (section offset for ``et_rel`` or address for ``et_dyn``)
1030  of the storage unit being relocated (computed using ``r_offset``).
1031
1032**S**
1033  Represents the value of the symbol whose index resides in the relocation
1034  entry. Relocations not using this must specify a symbol index of
1035  ``STN_UNDEF``.
1036
1037**B**
1038  Represents the base address of a loaded executable or shared object which is
1039  the difference between the ELF address and the actual load address.
1040  Relocations using this are only valid in executable or shared objects.
1041
1042The following relocation types are supported:
1043
1044  .. table:: AMDGPU ELF Relocation Records
1045     :name: amdgpu-elf-relocation-records-table
1046
1047     ========================== ======= =====  ==========  ==============================
1048     Relocation Type            Kind    Value  Field       Calculation
1049     ========================== ======= =====  ==========  ==============================
1050     ``R_AMDGPU_NONE``                  0      *none*      *none*
1051     ``R_AMDGPU_ABS32_LO``      Static, 1      ``word32``  (S + A) & 0xFFFFFFFF
1052                                Dynamic
1053     ``R_AMDGPU_ABS32_HI``      Static, 2      ``word32``  (S + A) >> 32
1054                                Dynamic
1055     ``R_AMDGPU_ABS64``         Static, 3      ``word64``  S + A
1056                                Dynamic
1057     ``R_AMDGPU_REL32``         Static  4      ``word32``  S + A - P
1058     ``R_AMDGPU_REL64``         Static  5      ``word64``  S + A - P
1059     ``R_AMDGPU_ABS32``         Static, 6      ``word32``  S + A
1060                                Dynamic
1061     ``R_AMDGPU_GOTPCREL``      Static  7      ``word32``  G + GOT + A - P
1062     ``R_AMDGPU_GOTPCREL32_LO`` Static  8      ``word32``  (G + GOT + A - P) & 0xFFFFFFFF
1063     ``R_AMDGPU_GOTPCREL32_HI`` Static  9      ``word32``  (G + GOT + A - P) >> 32
1064     ``R_AMDGPU_REL32_LO``      Static  10     ``word32``  (S + A - P) & 0xFFFFFFFF
1065     ``R_AMDGPU_REL32_HI``      Static  11     ``word32``  (S + A - P) >> 32
1066     *reserved*                         12
1067     ``R_AMDGPU_RELATIVE64``    Dynamic 13     ``word64``  B + A
1068     ========================== ======= =====  ==========  ==============================
1069
1070``R_AMDGPU_ABS32_LO`` and ``R_AMDGPU_ABS32_HI`` are only supported by
1071the ``mesa3d`` OS, which does not support ``R_AMDGPU_ABS64``.
1072
1073There is no current OS loader support for 32-bit programs and so
1074``R_AMDGPU_ABS32`` is not used.
1075
1076.. _amdgpu-dwarf:
1077
1078DWARF
1079-----
1080
1081.. warning::
1082   This section describes a **provisional proposal** that is not currently
1083   fully implemented and is subject to change.
1084
1085Standard DWARF [DWARF]_ sections can be generated. These contain information
1086that maps the code object executable code and data to the source language
1087constructs. It can be used by tools such as debuggers and profilers.
1088
1089This section defines the AMDGPU target specific DWARF. It applies to DWARF
1090Version 4 and 5.
1091
1092.. _amdgpu-dwarf-overview:
1093
1094Overview
1095~~~~~~~~
1096
1097The AMDGPU has several features that require additional DWARF functionality in
1098order to support optimized code.
1099
1100A single code object can contain code for kernels that have different wave
1101sizes. The vector registers and some scalar registers are based on the wave
1102size. AMDGPU defines distinct DWARF registers for each wave size. This
1103simplifies the consumer of the DWARF so that each register has a fixed size,
1104rather than being dynamic according to the wave mode. Similarly, distinct DWARF
1105registers are defined for those registers that vary in size according to the
1106process address size. This allows a consumer to treat a specific AMDGPU target
1107as a single architecture regardless of how it is configured. The compiler
1108explicitly specifies the registers that match the mode of the code it is
1109generating.
1110
1111AMDGPU optimized code may spill vector registers to non-global address space
1112memory, and this spilling may be done only for lanes that are active on entry to
1113the subprogram. To support this, a location description that can be created as a
1114masked select is required.
1115
1116Since the active lane mask may be held in a register, a way to get the value of
1117a register on entry to a subprogram is required. To support this an operation
1118that returns the caller value of a register as specified by the Call Frame
1119Information (see :ref:`amdgpu-call-frame-information`) is required.
1120
1121Current DWARF uses an empty expression to indicate an undefined location
1122description. Since the masked select composite location description operation
1123takes more than one location description, it is necessary to have an explicit
1124way to specify an undefined location description. Otherwise it is not possible
1125to specify that a particular one of the input location descriptions is
1126undefined.
1127
1128CFI describes restoring callee saved registers that are spilled. Currently CFI
1129only allows a location description that is a register, memory address, or
1130implicit location description. AMDGPU optimized code may spill scalar registers
1131into portions of vector registers. This requires extending CFI to allow any
1132location description.
1133
1134The vector registers of the AMDGPU are represented as their full wave size,
1135meaning the wave size times the dword size. This reflects the actual hardware,
1136and allows the compiler to generate DWARF for languages that map a thread to the
1137complete wave. It also allows more efficient DWARF to be generated to describe
1138the CFI as only a single expression is required for the whole vector register,
1139rather than a separate expression for each lane's dword of the vector register.
1140It also allows the compiler to produce DWARF that indexes the vector register if
1141it spills scalar registers into portions of a vector registers.
1142
1143Since DWARF stack value entries have a base type and AMDGPU registers are a
1144vector of dwords, the ability to specify that a base type is a vector is
1145required.
1146
1147If the source language is mapped onto the AMDGPU wavefronts in a SIMT manner,
1148then the variable DWARF location expressions must compute the location for a
1149single lane of the wavefront. Therefore, a DWARF operator is required to denote
1150the current lane, much like ``DW_OP_push_object_address`` denotes the current
1151object. The ``DW_OP_*piece`` operators only allow literal indices. Therefore, a
1152composite location description is required that can take a computed index of a
1153location description (such as a vector register).
1154
1155If the source language is mapped onto the AMDGPU wavefronts in a SIMT manner the
1156compiler can use the AMDGPU execution mask register to control which lanes are
1157active. To describe the conceptual location of non-active lanes a DWARF
1158expression is needed that can compute a per lane PC. For efficiency, this is
1159done for the wave as a whole. This expression benefits by having a masked select
1160composite location description operation. This requires an attribute for source
1161location of each lane. The AMDGPU may update the execution mask for whole wave
1162operations and so needs an attribute that computes the current active lane mask.
1163
1164AMDGPU needs to be able to describe addresses that are in different kinds of
1165memory. Optimized code may need to describe a variable that resides in pieces
1166that are in different kinds of storage which may include parts of registers,
1167memory that is in a mixture of memory kinds, implicit values, or be undefined.
1168DWARF has the concept of segment addresses. However, the segment cannot be
1169specified within a DWARF expression, which is only able to specify the offset
1170portion of a segment address. The segment index is only provided by the entity
1171that species the DWARF expression. Therefore, the segment index is a property
1172that can only be put on complete objects, such as a variable. That makes it only
1173suitable for describing an entity (such as variable or subprogram code) that is
1174in a single kind of memory. Therefore, AMDGPU uses the DWARF concept of address
1175spaces. For example, a variable may be allocated in a register that is partially
1176spilled to the call stack which is in the private address space, and partially
1177spilled to the local address space.
1178
1179DWARF uses the concept of an address in many expression operators but does not
1180define how it relates to address spaces. For example,
1181``DW_OP_push_object_address`` pushes the address of an object. Other contexts
1182implicitly push an address on the stack before evaluating an expression. For
1183example, the ``DW_AT_use_location`` attribute of the
1184``DW_TAG_ptr_to_member_type``. The expression that uses the address needs to do
1185so in a general way and not need to be dependent on the address space of the
1186address. For example, a pointer to member value may want to be applied to an
1187object that may reside in any address space.
1188
1189The number of registers and the cost of memory operations is much higher for
1190AMDGPU than a typical CPU. The compiler attempts to optimize whole variables and
1191arrays into registers. Currently DWARF only allows ``DW_OP_push_object_address``
1192and related operations to work with a global memory location. To support AMDGPU
1193optimized code it is required to generalize DWARF to allow any location
1194description to be used. This allows registers, or composite location
1195descriptions that may be a mixture of memory, registers, or even implicit
1196values.
1197
1198Allowing a location description to be an entry on the DWARF stack allows them to
1199compose naturally. It allows objects to be located in any kind of memory address
1200space, in registers, be implicit values, be undefined, or a composite of any of
1201these.
1202
1203By extending DWARF carefully, all existing DWARF expressions can retain their
1204current semantic meaning. DWARF has implicit conversions that convert from a
1205value that is treated as an address in the default address space to a memory
1206location description. This can be extended to allow a default address space
1207memory location description to be implicitly converted back to its address
1208value. To allow composition of composite location descriptions, an explicit
1209operator that indicates the end is required. This can be implied if the end of a
1210DWARF expression is reached, allowing current DWARF expressions to remain legal.
1211
1212The ``DW_OP_plus`` and ``DW_OP_minus`` can be defined to operate on a memory
1213location description in the default target architecture address space and a
1214generic type, and produce a memory location description. This allows them to
1215continue to be used to offset an address. To generalize offsetting to any
1216location description, including location descriptions that describe when bytes
1217are in registers, are implicit, or a composite of these, the
1218``DW_OP_LLVM_offset`` and ``DW_OP_LLVM_bit_offset`` operations are added. These
1219do not perform wrapping which would be hard to define for location descriptions
1220of non-memory kinds. This allows ``DW_OP_push_object_address`` to push a
1221location description that may be in a register, or be an implicit value, and the
1222DWARF expression of ``DW_TAG_ptr_to_member_type`` can contain
1223``DW_OP_LLVM_offset`` to offset within it. ``DW_OP_LLVM_bit_offset`` generalizes
1224DWARF to work with bit fields.
1225
1226The DWARF ``DW_OP_xderef*`` operation allows a value to be converted into an
1227address of a specified address space which is then read. But provides no way to
1228create a memory location description for an address in the non-default address
1229space. For example, AMDGPU variables can be allocated in the local address space
1230at a fixed address. It is required to have an operation to create an address in
1231a specific address space that can be used to define the location description of
1232the variable. Defining this operation to produce a location description allows
1233the size of addresses in an address space to be larger than the generic type.
1234
1235If an operation had to produce a value that can be implicitly converted to a
1236memory location description, then it would be limited to the size of the generic
1237type which matches the size of the default address space. Its value would be
1238unspecified and likely not match any value in the actual program. By making the
1239result a location description, it allows a consumer great freedom in how it
1240implements it. The implicit conversion back to a value can be limited only to
1241the default address space to maintain compatibility.
1242
1243Similarly ``DW_OP_breg*`` treats the register as containing an address in the
1244default address space. It is required to be able to specify the address space of
1245the register value.
1246
1247Almost all uses of addresses in DWARF are limited to defining location
1248descriptions, or to be dereferenced to read memory. The exception is
1249``DW_CFA_val_offset`` which uses the address to set the value of a register. By
1250defining the CFA DWARF expression as being a memory location description, it can
1251maintain what address space it is, and that can be used to convert the offset
1252address back to an address in that address space. (An alternative is to defined
1253``DW_CFA_val_offset`` to implicitly use the default address space, and add
1254another operation that specifies the address space.)
1255
1256This approach allows all existing DWARF to have the identical semantics. It
1257allows the compiler to explicitly specify the address space it is using. For
1258example, a compiler could choose to access private memory in a swizzled manner
1259when mapping a source language to a wave in a SIMT manner, or to access it in an
1260unswizzled manner if mapping the same language with the wave being the thread.
1261It also allows the compiler to mix the address space it uses to access private
1262memory. For example, for SIMT it can still spill entire vector registers in an
1263unswizzled manner, while using swizzled for SIMT variable access. This approach
1264allows memory location descriptions for different address spaces to be combined
1265using the regular ``DW_OP_*piece`` operators.
1266
1267Location descriptions are an abstraction of storage, they give freedom to the
1268consumer on how to implement them. They allow the address space to encode lane
1269information so they can be used to read memory with only the memory description
1270and no extra arguments. The same set of operations can operate on locations
1271independent of their kind of storage. The ``DW_OP_deref*`` therefore can be used
1272on any storage kind. ``DW_OP_xderef*`` is unnecessary except to become a more
1273compact way to convert a segment address followed by dereferencing it.
1274
1275Several approaches were considered, and the one proposed appears to be the
1276cleanest and offers the greatest improvement of DWARF's ability to support
1277optimized code. Examining the gdb debugger and LLVM compiler, it appears only to
1278require modest changes as they both already have to support general use of
1279location descriptions. It is anticipated that will be the case for other
1280debuggers and compilers.
1281
1282The following provides the definitions for the additional operators, as well as
1283clarifying how existing expression operators, CFI operators, and attributes
1284behave with respect to generalized location descriptions that support address
1285spaces. It has been defined such that it is backwards compatible with DWARF 5.
1286The definitions are intended to fully define well-formed DWARF in a consistent
1287style. Some sections are organized to mirror the DWARF 5 specification
1288structure, with non-normative text shown in *italics*.
1289
1290.. _amdgpu-dwarf-language-names:
1291
1292Language Names
1293~~~~~~~~~~~~~~
1294
1295Language codes defined for use with the ``DW_AT_language`` attribute are
1296defined in :ref:`amdgpu-dwarf-language-names-table`.
1297
1298.. table:: AMDGPU DWARF Language Names
1299   :name: amdgpu-dwarf-language-names-table
1300
1301   ==================== ====== =================== =============================
1302   Language Name        Code   Default Lower Bound Description
1303   ==================== ====== =================== =============================
1304   ``DW_LANG_LLVM_HIP`` 0x8100 0                   AMD HIP Language. See [HIP]_.
1305   ==================== ====== =================== =============================
1306
1307The ``DW_LANG_LLVM_HIP`` language can be supported by extending the C++
1308language.
1309
1310.. _amdgpu-dwarf-register-mapping:
1311
1312Register Mapping
1313~~~~~~~~~~~~~~~~
1314
1315DWARF registers are encoded as numbers, which are mapped to architecture
1316registers. The mapping for AMDGPU is defined in
1317:ref:`amdgpu-dwarf-register-mapping-table`.
1318
1319.. table:: AMDGPU DWARF Register Mapping
1320   :name: amdgpu-dwarf-register-mapping-table
1321
1322   ============== ================= ======== ==================================
1323   DWARF Register AMDGPU Register   Bit Size Description
1324   ============== ================= ======== ==================================
1325   0              PC_32             32       Program Counter (PC) when
1326                                             executing in a 32-bit process
1327                                             address space. Used in the CFI to
1328                                             describe the PC of the calling
1329                                             frame.
1330   1              EXEC_MASK_32      32       Execution Mask Register when
1331                                             executing in wave 32 mode.
1332   2-15           *Reserved*
1333   16             PC_64             64       Program Counter (PC) when
1334                                             executing in a 64-bit process
1335                                             address space. Used in the CFI to
1336                                             describe the PC of the calling
1337                                             frame.
1338   17             EXEC_MASK_64      64       Execution Mask Register when
1339                                             executing in wave 64 mode.
1340   18-31          *Reserved*
1341   32-95          SGPR0-SGPR63      32       Scalar General Purpose
1342                                             Registers.
1343   96-127         *Reserved*
1344   128-511        *Reserved*
1345   512-1023       *Reserved*
1346   1024-1087      *Reserved*
1347   1088-1129      SGPR64-SGPR105    32       Scalar General Purpose Registers
1348   1130-1535      *Reserved*
1349   1536-1791      VGPR0-VGPR255     32*32    Vector General Purpose Registers
1350                                             when executing in wave 32 mode.
1351   1792-2047      *Reserved*
1352   2048-2303      AGPR0-AGPR255     32*32    Vector Accumulation Registers
1353                                             when executing in wave 32 mode.
1354   2304-2559      *Reserved*
1355   2560-2815      VGPR0-VGPR255     64*32    Vector General Purpose Registers
1356                                             when executing in wave 64 mode.
1357   2816-3071      *Reserved*
1358   3072-3327      AGPR0-AGPR255     64*32    Vector Accumulation Registers
1359                                             when executing in wave 64 mode.
1360   3328-3583      *Reserved*
1361   ============== ================= ======== ==================================
1362
1363The vector registers are represented as the full size for the wavefront. They
1364are organized as consecutive dwords (32-bits), one per lane, with the dword at
1365the least significant bit position corresponding to lane 0 and so forth. DWARF
1366location expressions involving the ``DW_OP_LLVM_offset`` and
1367``DW_OP_LLVM_push_lane`` operations are used to select the part of the vector
1368register corresponding to the lane that is executing the current thread of
1369execution in languages that are implemented using a SIMD or SIMT execution
1370model.
1371
1372If the wavefront size is 32 lanes then the wave 32 mode register definitions
1373are used. If the wavefront size is 64 lanes then the wave 64 mode register
1374definitions are used. Some AMDGPU targets support executing in both wave 32
1375and wave 64 mode. The register definitions corresponding to the wave mode
1376of the generated code will be used.
1377
1378If code is generated to execute in a 32-bit process address space then the
137932-bit process address space register definitions are used. If code is
1380generated to execute in a 64-bit process address space then the 64-bit process
1381address space register definitions are used. The ``amdgcn`` target only
1382supports the 64-bit process address space.
1383
1384Address Class Mapping
1385~~~~~~~~~~~~~~~~~~~~~
1386
1387DWARF address classes are used for languages with the concept of memory address
1388spaces. They are used in the ``DW_AT_address_class`` attribute for pointer type,
1389reference type, subroutine, and subroutine type debugger information entries
1390(DIEs).
1391
1392The address class mapping for AMDGPU is defined in
1393:ref:`amdgpu-dwarf-address-class-mapping-table`.
1394
1395.. table:: AMDGPU DWARF Address Class Mapping
1396   :name: amdgpu-dwarf-address-class-mapping-table
1397
1398   =========================== ===== =================
1399   DWARF                             AMDGPU
1400   --------------------------------- -----------------
1401   Address Class Name          Value Address Space
1402   =========================== ===== =================
1403   ``DW_ADDR_none``            0x00  Generic (Flat)
1404   ``DW_ADDR_AMDGPU_global``   0x01  Global
1405   ``DW_ADDR_AMDGPU_region``   0x02  Region (GDS)
1406   ``DW_ADDR_AMDGPU_local``    0x03  Local (group/LDS)
1407   ``DW_ADDR_AMDGPU_constant`` 0x04  Global
1408   ``DW_ADDR_AMDGPU_private``  0x05  Private (Scratch)
1409   =========================== ===== =================
1410
1411See :ref:`amdgpu-address-spaces` for information on the AMDGPU address spaces
1412including address size and NULL value.
1413
1414For AMDGPU the address class encodes the address class as declared in the
1415source language type.
1416
1417For AMDGPU if no ``DW_AT_address_class`` attribute is present, then the
1418``DW_ADDR_none`` address class is used.
1419
1420.. note::
1421
1422  The ``DW_ADDR_none`` default was defined as ``Generic`` and not ``Global``
1423  to match the LLVM address space ordering. This ordering was chosen to better
1424  support CUDA-like languages such as HIP that do not have address spaces in
1425  the language type system, but do allow variables to be allocated in
1426  different address spaces. So effectively all CUDA and HIP source language
1427  addresses are generic.
1428
1429.. note::
1430
1431  Currently DWARF defines address class values as architecture specific. It
1432  is unclear how language specific address spaces are intended to be
1433  represented in DWARF.
1434
1435  For example, OpenCL defines address spaces for ``global``, ``local``,
1436  ``constant``, and ``private``. These are part of the type system and are
1437  modifies to pointer types. In addition, OpenCL defines ``generic`` pointers
1438  that can reference either the ``global``, ``local``, or ``private`` address
1439  spaces. To support the OpenCL language the debugger would want to support
1440  casting pointers between the ``generic`` and other address spaces, and
1441  possibly using pointer casting to form an address for a specific address
1442  space out of an integral value.
1443
1444  The method to use to dereference a pointer type or reference type value is
1445  defined in DWARF expressions using ``DW_OP_xderef*`` which uses an
1446  architecture specific address space.
1447
1448  DWARF defines the ``DW_AT_address_class`` attribute on pointer types and
1449  reference types. It specifies the method to use to dereference them. Why
1450  is the value of this not the same as the address space value used in
1451  ``DW_OP_xderef*`` since in both cases it is architecture specific and the
1452  architecture presumably will use the same set of methods to dereference
1453  pointers in both cases?
1454
1455  Since ``DW_AT_address_class`` uses an architecture specific value it cannot
1456  in general capture the source language address space type modifier concept.
1457  On some architectures all source language address space modifies may
1458  actually use the same method for dereferencing pointers.
1459
1460  One possibility is for DWARF to add an ``DW_TAG_LLVM_address_class_type``
1461  type modifier that can be applied to a pointer type and reference type. The
1462  ``DW_AT_address_class`` attribute could be re-defined to not be architecture
1463  specific and instead define generalized language values that will support
1464  OpenCL and other languages using address spaces. The ``DW_AT_address_class``
1465  could be defined to not be applied to pointer or reference types, but
1466  instead only to the ``DW_TAG_LLVM_address_class_type`` type modifier entry.
1467
1468  If a pointer type or reference type is not modified by
1469  ``DW_TAG_LLVM_address_class_type`` or if ``DW_TAG_LLVM_address_class_type``
1470  has no ``DW_AT_address_class`` attribute, then the pointer type or reference
1471  type would be defined to use the ``DW_ADDR_none`` address class as
1472  currently. Since modifiers can be chained, it would need to be defined if
1473  multiple ``DW_TAG_LLVM_address_class_type`` modifies was legal, and if so if
1474  the outermost one is the one that takes precedence.
1475
1476  A target implementation that supports multiple address spaces would need to
1477  map ``DW_ADDR_none`` appropriately to support CUDA-like languages
1478  that have no address classes in the type system, but do support variable
1479  allocation in address spaces. See the above note that describes why AMDGPU
1480  choose to make ``DW_ADDR_none`` map to the ``Generic`` AMDGPU address space
1481  and not the ``Global`` address space.
1482
1483  An alternative would be to define ``DW_ADDR_none`` as being the global
1484  address class and then change ``DW_ADDR_global`` to ``DW_ADDR_generic``.
1485  Compilers generating DWARF for CUDA-like languages would then have to define
1486  every CUDA-like language pointer type or reference type using
1487  ``DW_TAG_LLVM_address_class_type`` with a ``DW_AT_address_class`` attribute
1488  of ``DW_ADDR_generic`` to match the language semantics. The AMDGPU
1489  alternative avoids needing to do this and seems to fit better into how CLANG
1490  and LLVM have added support for the CUDA-like languages on top of existing
1491  C++ language support.
1492
1493  A new ``DW_AT_address_space`` attribute could be defined that can be applied
1494  to pointer type, reference type, subroutine, and subroutine type to describe
1495  how objects having the given type are dereferenced or called (the role that
1496  ``DW_AT_address_class`` currently provides). The values of
1497  ``DW_AT_address_space`` would be architecture specific and the same as used
1498  in ``DW_OP_xderef*``.
1499
1500.. _amdgpu-dwarf-address-space-mapping:
1501
1502Address Space Mapping
1503~~~~~~~~~~~~~~~~~~~~~
1504
1505DWARF address spaces are used in location expressions to describe the memory
1506space where data resides. Address spaces correspond to a target specific memory
1507space and are not tied to any source language concept.
1508
1509The AMDGPU address space mapping is defined in
1510:ref:`amdgpu-dwarf-address-space-mapping-table`.
1511
1512.. table:: AMDGPU DWARF Address Space Mapping
1513   :name: amdgpu-dwarf-address-space-mapping-table
1514
1515   ======================================= ===== ======= ======== ================= =======================
1516   DWARF                                                          AMDGPU            Notes
1517   --------------------------------------- ----- ---------------- ----------------- -----------------------
1518   Address Space Name                      Value Address Bit Size Address Space
1519   --------------------------------------- ----- ------- -------- ----------------- -----------------------
1520   ..                                            64-bit  32-bit
1521                                                 process process
1522                                                 address address
1523                                                 space   space
1524   ======================================= ===== ======= ======== ================= =======================
1525   ``DW_ASPACE_none``                      0x00  8       4        Global            *default address space*
1526   ``DW_ASPACE_AMDGPU_generic``            0x01  8       4        Generic (Flat)
1527   ``DW_ASPACE_AMDGPU_region``             0x02  4       4        Region (GDS)
1528   ``DW_ASPACE_AMDGPU_local``              0x03  4       4        Local (group/LDS)
1529   *Reserved*                              0x04
1530   ``DW_ASPACE_AMDGPU_private_lane``       0x05  4       4        Private (Scratch) *focused lane*
1531   ``DW_ASPACE_AMDGPU_private_wave``       0x06  4       4        Private (Scratch) *unswizzled wave*
1532   *Reserved*                              0x07-
1533                                           0x1F
1534   ``DW_ASPACE_AMDGPU_private_lane<0-63>`` 0x20- 4       4        Private (Scratch) *specific lane*
1535                                           0x5F
1536   ======================================= ===== ======= ======== ================= =======================
1537
1538See :ref:`amdgpu-address-spaces` for information on the AMDGPU address spaces
1539including address size and NULL value.
1540
1541The ``DW_ASPACE_none`` address space is the default address space used in DWARF
1542operations that do not specify an address space. It therefore has to map to the
1543global address space so that the ``DW_OP_addr*`` and related operations can
1544refer to addresses in the program code.
1545
1546The ``DW_ASPACE_AMDGPU_generic`` address space allows location expressions to
1547specify the flat address space. If the address corresponds to an address in the
1548local address space then it corresponds to the wave that is executing the
1549focused thread of execution. If the address corresponds to an address in the
1550private address space then it corresponds to the lane that is executing the
1551focused thread of execution for languages that are implemented using a SIMD or
1552SIMT execution model.
1553
1554.. note::
1555
1556  CUDA-like languages such as HIP that do not have address spaces in the
1557  language type system, but do allow variables to be allocated in different
1558  address spaces, will need to explicitly specify the
1559  ``DW_ASPACE_AMDGPU_generic`` address space in the DWARF operations as the
1560  default address space is the global address space.
1561
1562The ``DW_ASPACE_AMDGPU_local`` address space allows location expressions to
1563specify the local address space corresponding to the wave that is executing the
1564focused thread of execution.
1565
1566The ``DW_ASPACE_AMDGPU_private_lane`` address space allows location expressions
1567to specify the private address space corresponding to the lane that is
1568executing the focused thread of execution for languages that are implemented
1569using a SIMD or SIMT execution model.
1570
1571The ``DW_ASPACE_AMDGPU_private_wave`` address space allows location expressions
1572to specify the unswizzled private address space corresponding to the wave that
1573is executing the focused thread of execution. The wave view of private memory
1574is the per wave unswizzled backing memory layout defined in
1575:ref:`amdgpu-address-spaces`, such that address 0 corresponds to the first
1576location for the backing memory of the wave (namely the address is not offset
1577by ``wavefront-scratch-base``). So to convert from a
1578``DW_ASPACE_AMDGPU_private_lane`` to a ``DW_ASPACE_AMDGPU_private_wave``
1579segment address perform the following:
1580
1581::
1582
1583  private-address-wave =
1584    ((private-address-lane / 4) * wavefront-size * 4) +
1585    (wavefront-lane-id * 4) + (private-address-lane % 4)
1586
1587If the ``DW_ASPACE_AMDGPU_private_lane`` segment address is dword aligned and
1588the start of the dwords for each lane starting with lane 0 is required, then
1589this simplifies to:
1590
1591::
1592
1593  private-address-wave =
1594    private-address-lane * wavefront-size
1595
1596A compiler can use this address space to read a complete spilled vector
1597register back into a complete vector register in the CFI. The frame pointer can
1598be a private lane segment address which is dword aligned, which can be shifted
1599to multiply by the wave size, and then used to form a private wave segment
1600address that gives a location for a contiguous set of dwords, one per lane,
1601where the vector register dwords are spilled. The compiler knows the wave size
1602since it generates the code. Note that the type of the address may have to be
1603converted as the size of a private lane segment address may be smaller than the
1604size of a private wave segment address.
1605
1606The ``DW_ASPACE_AMDGPU_private_lane<n>`` address space allows location
1607expressions to specify the private address space corresponding to a specific
1608lane. For example, this can be used when the compiler spills scalar registers
1609to scratch memory, with each scalar register being saved to a different lane's
1610scratch memory.
1611
1612.. _amdgpu-dwarf-expressions:
1613
1614Expressions
1615~~~~~~~~~~~
1616
1617The following sections define the new DWARF expression operator used by AMDGPU,
1618as well as clarifying the extensions to already existing DWARF 5 operations.
1619
1620DWARF expressions describe how to compute a value or specify a location
1621description. An expression is encoded as a stream of operations, each consisting
1622of an opcode followed by zero or more literal operands. The number of operands
1623is implied by the opcode.
1624
1625Operations represent a postfix operation on a simple stack machine. They can act
1626on entries on the stack, including adding entries and removing entries. If the
1627kind of a stack entry does not match the kind required by the operation, and is
1628not implicitly convertible to the required kind, then the DWARF expression is
1629ill-formed.
1630
1631Each stack entry can be one of two kinds: a value or a location description.
1632Value stack entries are described in :ref:`amdgpu-value-operations` and
1633location description stack entries are described in
1634:ref:`amdgpu-location-description-operations`.
1635
1636*The evaluation of a DWARF expression can provide the location description of an
1637object, the value of an array bound, the length of a dynamic string, the desired
1638value itself, and so on.*
1639
1640The result of the evaluation of a DWARF expression is defined as:
1641
1642* If evaluation of the DWARF expression is on behalf of a ``DW_OP_call*``
1643  operation for a ``DW_AT_location`` attribute that belongs to a
1644  ``DW_TAG_dwarf_procedure`` debugging information entry, then all the entries
1645  on the stack are left, and execution of the DWARF expression containing the
1646  ``DW_OP_call*`` operation continues.
1647
1648* If evaluation of the DWARF expression requires a location description, then:
1649
1650  * If the stack is empty, an undefined location description is returned.
1651
1652  * If the top stack entry is a location description, or can be converted to
1653    one, then the, possibly converted, location description is returned. Any
1654    other entries on the stack are discarded.
1655
1656  * Otherwise the DWARF expression is ill-formed.
1657
1658    .. note::
1659
1660      Could define this case as returning an implicit location description as
1661      if the ``DW_OP_implicit`` operation is performed.
1662
1663* If evaluation of the DWARF expression requires a value, then:
1664
1665  * If the top stack entry is a value, or can be converted to one, then the,
1666    possibly converted, value is returned. Any other entries on the stack are
1667    discarded.
1668
1669  * Otherwise the DWARF expression is ill-formed.
1670
1671.. _amdgpu-stack-operations:
1672
1673Stack Operations
1674++++++++++++++++
1675
1676The following operations manipulate the DWARF stack. Operations that index
1677the stack assume that the top of the stack (most recently added entry) has index
16780. They allow the stack entries to be either a value or location description.
1679
1680If any stack entry accessed by a stack operation is an incomplete composite
1681location description, then the DWARF expression is ill-formed.
1682
1683.. note::
1684
1685  These operations now support stack entries that are values and location
1686  descriptions.
1687
1688.. note::
1689
1690  If it is desired to also make them work with incomplete composite location
1691  descriptions then would need to define that the composite location storage
1692  specified by the incomplete composite location description is also replicated
1693  when a copy is pushed. This ensures that each copy of the incomplete composite
1694  location description can updated the composite location storage they specify
1695  independently.
1696
16971.  ``DW_OP_dup``
1698
1699    ``DW_OP_dup`` duplicates the stack entry at the top of the stack.
1700
17012.  ``DW_OP_drop``
1702
1703    ``DW_OP_drop`` pops the stack entry at the top of the stack and discards it.
1704
17053.  ``DW_OP_pick``
1706
1707    ``DW_OP_pick`` has a single unsigned 1-byte operand that is treated as an
1708    index I. A copy of the stack entry with index I is pushed onto the stack.
1709
17104.  ``DW_OP_over``
1711
1712    ``DW_OP_over`` pushes a copy of the entry entry with index 1.
1713
1714    *This is equivalent to a ``DW_OP_pick 1`` operation.*
1715
17165.  ``DW_OP_swap``
1717
1718    ``DW_OP_swap`` swaps the top two stack entries. The entry at the top of the
1719    stack becomes the second stack entry, and the second stack entry becomes the
1720    top of the stack.
1721
17226.  ``DW_OP_rot``
1723
1724    ``DW_OP_rot`` rotates the first three stack entries. The entry at the top of
1725    the stack becomes the third stack entry, the second entry becomes the top of
1726    the stack, and the third entry becomes the second entry.
1727
1728.. _amdgpu-value-operations:
1729
1730Value Operations
1731++++++++++++++++
1732
1733Each value stack entry has a type and a value, and can represent a value of
1734any supported base type of the target machine. The base type specifies the size
1735and encoding of the value.
1736
1737.. note::
1738
1739  It may be better to add an implicit pointer value kind that is produced when
1740  ``DW_OP_deref*`` retrieves the full contents of an implicit pointer location
1741  storage created by the ``DW_OP_implicit_pointer`` or
1742  ``DW_OP_LLVM_aspace_implicit_pointer`` operations.
1743
1744Instead of a base type, value stack entries can have a distinguished generic
1745type, which is an integral type that has the size of an address in the target
1746architecture default address space on the target machine and unspecified
1747signedness.
1748
1749*The generic type is the same as the unspecified type used for stack operations
1750defined in DWARF Version 4 and before.*
1751
1752An integral type is a base type that has an encoding of ``DW_ATE_signed``,
1753``DW_ATE_signed_char``, ``DW_ATE_unsigned``, ``DW_ATE_unsigned_char``,
1754``DW_ATE_boolean``, or any target architecture defined integral encoding in the
1755inclusive range ``DW_ATE_lo_user`` to ``DW_ATE_hi_user``.
1756
1757.. note::
1758
1759  Unclear if ``DW_ATE_address`` is an integral type. gdb does not seem to
1760  consider as integral.
1761
17621.  ``DW_OP_LLVM_push_lane`` *New*
1763
1764    ``DW_OP_LLVM_push_lane`` pushes a value with the generic type that is the
1765    target architecture lane identifier of the thread of execution for which a
1766    user presented expression is currently being evaluated. For languages that
1767    are implemented using a SIMD or SIMT execution model this is the lane number
1768    that corresponds to the source language thread of execution upon which the
1769    user is focused. Otherwise this is the value 0.
1770
1771    For AMDGPU, the lane identifier returned by ``DW_OP_LLVM_push_lane``
1772    corresponds to the the hardware lane number which is numbered from 0 to the
1773    wavefront size minus 1.
1774
17752.  ``DW_OP_entry_value``
1776
1777    ``DW_OP_entry_value`` pushes the value that the described location held upon
1778    entering the current subprogram.
1779
1780    It has two operands. The first is an unsigned LEB128 integer. The second is
1781    a block of bytes, with a length equal to the first operand, treated as a
1782    DWARF expression E.
1783
1784    E is evaluated as if it had been evaluated upon entering the current
1785    subprogram. E assumes no values are present on the DWARF stack initially and
1786    results in exactly one value being pushed on the DWARF stack when completed.
1787
1788    ``DW_OP_push_object_address`` is not meaningful inside of this DWARF
1789    operation.
1790
1791    If the result of E is a register location description (see
1792    :ref:`amdgpu-register-location-descriptions`), ``DW_OP_entry_value`` pushes
1793    the value that register had upon entering the current subprogram. The value
1794    entry type is the target machine register base type. If the register value
1795    is undefined or the register location description bit offset is not 0, then
1796    the DWARF expression is ill-formed.
1797
1798    *The register location description provides a more compact form for the case
1799    where the value was in a register on entry to the subprogram.*
1800
1801    Otherwise, the expression result is required to be a value, and
1802    ``DW_OP_entry_value`` pushes that value.
1803
1804    *The values needed to evaluate* ``DW_OP_entry_value`` *could be obtained in
1805    several ways. The consumer could suspend execution on entry to the
1806    subprogram, record values needed by* ``DW_OP_entry_value`` *expressions
1807    within the subprogram, and then continue; when evaluating*
1808    ``DW_OP_entry_value``\ *, the consumer would use these recorded values
1809    rather than the current values. Or, when evaluating* ``DW_OP_entry_value``\
1810    *, the consumer could virtually unwind using the Call Frame Information
1811    (see* :ref:`amdgpu-call-frame-information`\ *) to recover register values
1812    that might have been clobbered since the subprogram entry point.*
1813
1814    .. note::
1815
1816      Unclear why this operation is defined this way. If the expression is
1817      simply using existing variables then it is just a regular expression. It
1818      is unclear how the compiler instructs the consumer how to create the saved
1819      copies of the variables on entry. Seems only the compiler knows how to do
1820      this. If the main purpose is only to read the entry value of a register
1821      using CFI then would be better to have an operation that explicitly does
1822      just that such as ``DW_OP_LLVM_call_frame_entry_reg``.
1823
1824.. _amdgpu-location-description-operations:
1825
1826Location Description Operations
1827+++++++++++++++++++++++++++++++
1828
1829Information about the location of program objects is provided by location
1830descriptions. Location descriptions specify the storage that holds the program
1831objects, and a position within the storage.
1832
1833A location storage is a linear stream of bits that can hold values. Each
1834location storage has a size in bits and can be accessed using a zero-based bit
1835offset. The ordering of bits within location storage uses the bit numbering and
1836direction conventions that are appropriate to the current language on the target
1837architecture.
1838
1839.. note::
1840
1841  For AMDGPU bytes are ordered with least significant bytes first, and bits are
1842  ordered within bytes with least significant bits first.
1843
1844There are five kinds of location storage: undefined, memory, register, implicit,
1845and composite. Memory and register location storage corresponds to the target
1846architecture memory address spaces and registers. Implicit location storage
1847corresponds to fixed values that can only be read. Undefined location storage
1848indicates no value is available and therefore cannot be read or written.
1849Composite location storage allows a mixture of these where some bits come from
1850one kind of location storage and some from another kind of location storage.
1851
1852.. note::
1853
1854  It may be better to add an implicit pointer location storage kind for
1855  ``DW_OP_implicit_pointer`` or ``DW_OP_LLVM_aspace_implicit_pointer``.
1856
1857Location description stack entries specify a location storage to which they
1858refer, and a bit offset relative to the start of the location storage.
1859
1860General Operations
1861##################
1862
18631.  ``DW_OP_LLVM_offset`` *New*
1864
1865    ``DW_OP_LLVM_offset`` pops two stack entries. The first must be an integral
1866    type value that is treated as a byte displacement D. The second must be a
1867    location description L.
1868
1869    It adds the value of D scaled by 8 (the byte size) to the bit offset of L,
1870    and pushes the updated L.
1871
1872    If the updated bit offset of L is less than 0 or greater than or equal to
1873    the size of the location storage specified by L, then the DWARF expression
1874    is ill-formed.
1875
18762.  ``DW_OP_LLVM_offset_uconst`` *New*
1877
1878    ``DW_OP_LLVM_offset_uconst`` has a single unsigned LEB128 integer operand
1879    that is treated as a displacement D.
1880
1881    It pops one stack entry that must be a location description L. It adds the
1882    value of D scaled by 8 (the byte size) to the bit offset of L, and pushes
1883    the updated L.
1884
1885    If the updated bit offset of L is less than 0 or greater than or equal to
1886    the size of the location storage specified by L, then the DWARF expression
1887    is ill-formed.
1888
1889    *This operation is supplied specifically to be able to encode more field
1890    displacements in two bytes than can be done with* ``DW_OP_lit<n>
1891    DW_OP_LLVM_offset``\ *.*
1892
18933.  ``DW_OP_LLVM_bit_offset`` *New*
1894
1895    ``DW_OP_LLVM_bit_offset`` pops two stack entries. The first must be an
1896    integral type value that is treated as a bit displacement D. The second must
1897    be a location description L.
1898
1899    It adds the value of D to the bit offset of L, and pushes the updated L.
1900
1901    If the updated bit offset of L is less than 0 or greater than or equal to
1902    the size of the location storage specified by L, then the DWARF expression
1903    is ill-formed.
1904
19054.  ``DW_OP_deref``
1906
1907    The ``DW_OP_deref`` operation pops one stack entry that must be a location
1908    description L.
1909
1910    A value of the bit size of the generic type is retrieved from the location
1911    storage specified by L starting at the bit offset specified by L. The
1912    retrieved generic type value V is pushed on the stack.
1913
1914    If any bit of the value is retrieved from the undefined location storage, or
1915    the offset of any bit exceeds the size of the location storage specified by
1916    L, then the DWARF expression is ill-formed.
1917
1918    See :ref:`amdgpu-implicit-location-descriptions` for special rules
1919    concerning implicit location descriptions created by the
1920    ``DW_OP_implicit_pointer`` and ``DW_OP_LLVM_implicit_aspace_pointer``
1921    operations.
1922
19235.  ``DW_OP_deref_size``
1924
1925    ``DW_OP_deref_size`` has a single 1-byte unsigned integral constant treated
1926    as a byte result size S.
1927
1928    It pops one stack entry that must be a location description L.
1929
1930    A value of S scaled by 8 (the byte size) bits is retrieved from the location
1931    storage specified by L starting at the bit offset specified by L. The value
1932    V retrieved is zero-extended to the bit size of the generic type before
1933    being pushed onto the stack with the generic type.
1934
1935    If S is larger than the byte size of the generic type, if any bit of the
1936    value is retrieved from the undefined location storage, or if the offset of
1937    any bit exceeds the size of the location storage specified by L, then the
1938    DWARF expression is ill-formed.
1939
1940    See :ref:`amdgpu-implicit-location-descriptions` for special rules
1941    concerning implicit location descriptions created by the
1942    ``DW_OP_implicit_pointer`` and ``DW_OP_LLVM_implicit_aspace_pointer``
1943    operations.
1944
19456.  ``DW_OP_deref_type``
1946
1947    ``DW_OP_deref_type`` has two operands. The first is a 1-byte unsigned
1948    integral constant whose value S is the same as the size of the base type
1949    referenced by the second operand. The second operand is an unsigned LEB128
1950    integer that represents the offset of a debugging information entry E in the
1951    current compilation unit, which must be a ``DW_TAG_base_type`` entry that
1952    provides the type of the result value.
1953
1954    It pops one stack entry that must be a location description L. A value of
1955    the bit size S is retrieved from the location storage specified by L
1956    starting at the bit offset specified by the L. The retrieved result type
1957    value V is pushed on the stack.
1958
1959    If any bit of the value is retrieved from the undefined location storage, or
1960    if the offset of any bit exceeds the size of the specified location storage,
1961    then the DWARF expression is ill-formed.
1962
1963    See :ref:`amdgpu-implicit-location-descriptions` for special rules
1964    concerning implicit location descriptions created by the
1965    ``DW_OP_implicit_pointer`` and ``DW_OP_LLVM_implicit_aspace_pointer``
1966    operations.
1967
1968    *While the size of the pushed value could be inferred from the base type
1969    definition, it is encoded explicitly into the operation so that the
1970    operation can be parsed easily without reference to the* ``.debug_info``
1971    *section.*
1972
19737.  ``DW_OP_xderef`` *Deprecated*
1974
1975    ``DW_OP_xderef`` pops two stack entries. The first must be an integral type
1976    value that is treated as an address A. The second must be an integral type
1977    value that is treated as an address space identifier AS for those
1978    architectures that support multiple address spaces.
1979
1980    The operation is equivalent to performing ``DW_OP_swap;
1981    DW_OP_LLVM_form_aspace_address; DW_OP_deref``. The retrieved generic type
1982    value V is left on the stack.
1983
19848.  ``DW_OP_xderef_size`` *Deprecated*
1985
1986    ``DW_OP_xderef_size`` has a single 1-byte unsigned integral constant treated
1987    as a byte result size S.
1988
1989    It pops two stack entries. The first must be an integral type value that is
1990    treated as an address A. The second must be an integral type value that is
1991    treated as an address space identifier AS for those architectures that
1992    support multiple address spaces.
1993
1994    The operation is equivalent to performing ``DW_OP_swap;
1995    DW_OP_LLVM_form_aspace_address; DW_OP_deref_size S``. The zero-extended
1996    retrieved generic type value V is left on the stack.
1997
19989.  ``DW_OP_xderef_type`` *Deprecated*
1999
2000    ``DW_OP_xderef_type`` has two operands. The first is a 1-byte unsigned
2001    integral constant S whose value is the same as the size of the base type
2002    referenced by the second operand. The second operand is an unsigned LEB128
2003    integer R that represents the offset of a debugging information entry E in
2004    the current compilation unit, which must be a ``DW_TAG_base_type`` entry
2005    that provides the type of the result value.
2006
2007    It pops two stack entries. The first must be an integral type value that is
2008    treated as an address A. The second must be an integral type value that is
2009    treated as an address space identifier AS for those architectures that
2010    support multiple address spaces.
2011
2012    The operation is equivalent to performing ``DW_OP_swap;
2013    DW_OP_LLVM_form_aspace_address; DW_OP_deref_type S R``. The retrieved result
2014    type value V is left on the stack.
2015
201610. ``DW_OP_push_object_address``
2017
2018    ``DW_OP_push_object_address`` pushes the location description L of the
2019    object currently being evaluated as part of evaluation of a user presented
2020    expression.
2021
2022    This object may correspond to an independent variable described by its own
2023    debugging information entry or it may be a component of an array, structure,
2024    or class whose address has been dynamically determined by an earlier step
2025    during user expression evaluation.
2026
2027    *This operator provides explicit functionality (especially for arrays
2028    involving descriptions) that is analogous to the implicit push of the base
2029    address of a structure prior to evaluation of a
2030    ``DW_AT_data_member_location`` to access a data member of a structure.*
2031
203211. ``DW_OP_call2, DW_OP_call4, DW_OP_call_ref``
2033
2034    ``DW_OP_call2``, ``DW_OP_call4``, and ``DW_OP_call_ref`` perform DWARF
2035    procedure calls during evaluation of a DWARF expression or location
2036    description.
2037
2038    ``DW_OP_call2`` and ``DW_OP_call4``, have one operand that is a 2- or 4-byte
2039    unsigned offset, respectively, of a debugging information entry D in the
2040    current compilation unit.
2041
2042    ``DW_OP_LLVM_call_ref`` has one operand that is a 4-byte unsigned value in
2043    the 32-bit DWARF format, or an 8-byte unsigned value in the 64-bit DWARF
2044    format, that is treated as an offset of a debugging information entry D in a
2045    ``.debug_info`` section, which may be contained in an executable or shared
2046    object file other than that containing the operator. For references from one
2047    executable or shared object file to another, the relocation must be
2048    performed by the consumer.
2049
2050    *Operand interpretation of* ``DW_OP_call2``\ *,* ``DW_OP_call4``\ *, and*
2051    ``DW_OP_call_ref`` *is exactly like that for* ``DW_FORM_ref2``\ *,
2052    ``DW_FORM_ref4``\ *, and* ``DW_FORM_ref_addr``\ *, respectively.*
2053
2054    If D has a ``DW_AT_location`` attribute, then the DWARF expression E
2055    corresponding to the current program location is selected.
2056
2057    .. note::
2058
2059      To allow ``DW_OP_call*`` to compute the location description for any
2060      variable or formal parameter regardless of whether the producer has
2061      optimized it to a constant, the following rule could be added:
2062
2063      .. note::
2064
2065        If D has a ``DW_AT_const_value`` attribute, then a DWARF expression E
2066        consisting a ``DW_OP_implicit_value`` operation with the value of the
2067        ``DW_AT_const_value`` attribute is selected.
2068
2069      This would be consistent with ``DW_OP_implicit_pointer``.
2070
2071      Alternatively, could deprecate using ``DW_AT_const_value`` for
2072      ``DW_TAG_variable`` and ``DW_TAG_formal_parameter`` debugger information
2073      entries that are constants and instead use ``DW_AT_location`` with an
2074      implicit location description instead, then this rule would not be
2075      required.
2076
2077    Otherwise, an empty expression E is selected.
2078
2079    If D is a ``DW_TAG_dwarf_procedure`` debugging information entry, then E is
2080    evaluated using the same DWARF expression stack. Any existing stack entries
2081    may be accessed and/or removed in the evaluation of E, and the evaluation of
2082    E may add any new stack entries.
2083
2084    *Values on the stack at the time of the call may be used as parameters by
2085    the called expression and values left on the stack by the called expression
2086    may be used as return values by prior agreement between the calling and
2087    called expressions.*
2088
2089    Otherwise, E is evaluated on a separate DWARF stack and the resulting
2090    location description L is pushed on the ``DW_OP_call*`` operation's stack.
2091
2092    .. note:
2093
2094      In DWARF 5, if D does not have a ``DW_AT_location`` then ``DW_OP_call*``
2095      is defined to have no effect. It is unclear that this is the right
2096      definition as a producer should be able to rely on using ``DW_OP_call*``
2097      to get a location description for any non-\ ``DW_TAG_dwarf_procedure``
2098      debugging information entries, and should not be creating DWARF with
2099      ``DW_OP_call*`` to a ``DW_TAG_dwarf_procedure`` that does not have a
2100      ``DW_AT_location`` attribute.
2101
210212. ``DW_OP_LLVM_call_frame_entry_reg`` *New*
2103
2104    ``DW_OP_LLVM_call_frame_entry_reg`` has a single unsigned LEB128 integer
2105    operand that is treated as a target architecture register number R.
2106
2107    It pushes a location description L that holds the value of register R on
2108    entry to the current subprogram as defined by the Call Frame Information
2109    (see :ref:`amdgpu-call-frame-information`).
2110
2111    *If there is no Call Frame Information defined, then the default rules for
2112    the target architecture are used. If the register rule is* undefined\ *,
2113    then the undefined location description is pushed. If the register rule is*
2114    same value\ *, then a register location description for R is pushed.*
2115
2116Undefined Location Descriptions
2117###############################
2118
2119The undefined location storage represents a piece or all of an object that is
2120present in the source but not in the object code (perhaps due to optimization).
2121Neither reading or writing to the undefined location storage is meaningful.
2122
2123An undefined location description specifies the undefined location storage.
2124There is no concept of the size of the undefined location storage, nor of a bit
2125offset for an undefined location description. The ``DW_OP_LLVM_*offset``
2126operations leave an undefined location description unchanged. The
2127``DW_OP_*piece`` operations can explicitly or implicitly specify an undefined
2128location description, allowing any size and offset to be specified, and results
2129in a part with all undefined bits.
2130
21311.  ``DW_OP_LLVM_undefined`` *New*
2132
2133    ``DW_OP_LLVM_undefined`` pushes an undefined location description L.
2134
2135Memory Location Descriptions
2136############################
2137
2138There is a memory location storage that corresponds to each of the target
2139architecture linear memory address spaces. The size of each memory location
2140storage corresponds to the range of the addresses in the address space.
2141
2142*It is target architecture defined how address space location storage maps to
2143target architecture physical memory. For example, they may be independent memory
2144or more than one location storage may alias the same physical memory possibly at
2145different offsets and with different interleaving. The mapping may also be
2146dictated by the source language address classes.*
2147
2148A memory location description specifies a memory location storage. The bit
2149offset corresponds to an address in the address space scaled by 8 (the byte
2150size). Bits accessed using a memory location description, access the
2151corresponding target architecture memory starting at the bit offset.
2152
2153``DW_ASPACE_none`` is defined as the target architecture default address space.
2154
2155*The target architecture default address space for AMDGPU is the global address
2156space.*
2157
2158If a stack entry is required to be a location description, but it is a value
2159with the generic type, then it is implicitly convert to a memory location
2160description that specifies memory in the target architecture default address
2161space with a bit offset equal to the value scaled by 8 (the byte size).
2162
2163  .. note::
2164
2165    If want to allow any integral type value to be implicitly converted to a
2166    memory location description in the target architecture default address
2167    space:
2168
2169    .. note::
2170
2171      If a stack entry is required to be a location description, but it is a
2172      value with an integral type, then it is implicitly convert to a memory
2173      location description. The stack entry value is zero extended to the size
2174      of the generic type and the least significant generic type size bits are
2175      treated as a twos-complement unsigned value to be used as an address. The
2176      converted memory location description specifies memory location storage
2177      corresponding to the target architecture default address space with a bit
2178      offset equal to the address scaled by 8 (the byte size).
2179
2180    The implicit conversion could also be defined as target specific. For
2181    example, gdb checks if the value is an integral type. If it is not it gives
2182    an error. Otherwise, gdb zero-extends the value to 64 bits. If the gdb
2183    target defines a hook function then it is called and it can modify the 64
2184    bit value, possibly sign extending the original value. Finally, gdb treats
2185    the 64 bit value as a memory location address.
2186
2187If a stack entry is required to be a location description, but it is an implicit
2188pointer value IPV with the target architecture default address space, then it is
2189implicitly convert to the location description specified by IPV. See
2190:ref:`amdgpu-implicit-location-descriptions`.
2191
2192If a stack entry is required to be a value with a generic type, but it is a
2193memory location description in the target architecture default address space
2194with a bit offset that is a multiple of 8, then it is implicitly converted to a
2195value with a generic type that is equal to the bit offset divided by 8 (the byte
2196size).
2197
21981.  ``DW_OP_addr``
2199
2200    ``DW_OP_addr`` has a single byte constant value operand, which has the size
2201    of the generic type, treated as an address A.
2202
2203    It pushes a memory location description L on the stack that specifies the
2204    memory location storage for the target architecture default address space
2205    with a bit offset equal to A scaled by 8 (the byte size).
2206
2207    *If the DWARF is part of a code object, then A may need to be relocated. For
2208    example, in the ELF code object format, A must be adjusted by the difference
2209    between the ELF segment virtual address and the virtual address at which the
2210    segment is loaded.*
2211
22122.  ``DW_OP_addrx``
2213
2214    ``DW_OP_addrx`` has a single unsigned LEB128 integer operand that is treated
2215    as a zero-based index into the ``.debug_addr`` section relative to the value
2216    of the ``DW_AT_addr_base`` attribute of the associated compilation unit. The
2217    address value A in the ``.debug_addr`` section has the size of generic type.
2218
2219    It pushes a memory location description L on the stack that specifies the
2220    memory location storage for the target architecture default address space
2221    with a bit offset equal to A scaled by 8 (the byte size).
2222
2223    *If the DWARF is part of a code object, then A may need to be relocated. For
2224    example, in the ELF code object format, A must be adjusted by the difference
2225    between the ELF segment virtual address and the virtual address at which the
2226    segment is loaded.*
2227
22283.  ``DW_OP_LLVM_form_aspace_address`` *New*
2229
2230    ``DW_OP_LLVM_form_aspace_address`` pops top two stack entries. The first
2231    must be an integral type value that is treated as an address space
2232    identifier AS for those architectures that support multiple address spaces.
2233    The second must be an integral type value that is treated as an address A.
2234
2235    The address size S is defined as the address bit size of the target
2236    architecture's address space that corresponds to AS.
2237
2238    A is adjusted by zero extending it to S bits and the least significant S
2239    bits are treated as a twos-complement unsigned value.
2240
2241    ``DW_OP_LLVM_form_aspace_address`` pushes a memory location description L
2242    that specifies the memory location storage that corresponds to AS, with a
2243    bit offset equal to the adjusted A scaled by 8 (the byte size).
2244
2245    If AS is not one of the values defined by the target architecture's
2246    ``DW_ASPACE_*`` values, then the DWARF expression is ill-formed.
2247
2248    See :ref:`amdgpu-implicit-location-descriptions` for special rules
2249    concerning implicit pointer values produced by dereferencing implicit
2250    location descriptions created by the ``DW_OP_implicit_pointer`` and
2251    ``DW_OP_LLVM_implicit_aspace_pointer`` operations.
2252
2253    The AMDGPU address spaces are defined in
2254    :ref:`amdgpu-dwarf-address-space-mapping-table`.
2255
22564.  ``DW_OP_form_tls_address``
2257
2258    ``DW_OP_form_tls_address`` pops one stack entry that must be an integral
2259    type value, and treats it as a thread-local storage address.
2260
2261    ``DW_OP_form_tls_address`` pushes a memory location description L for the
2262    target architecture default address space that corresponds to the
2263    thread-local storage address.
2264
2265    The meaning of the thread-local storage address is defined by the run-time
2266    environment. If the run-time environment supports multiple thread-local
2267    storage blocks for a single thread, then the block corresponding to the
2268    executable or shared library containing this DWARF expression is used.
2269
2270    *Some implementations of C, C++, Fortran, and other languages, support a
2271    thread-local storage class. Variables with this storage class have distinct
2272    values and addresses in distinct threads, much as automatic variables have
2273    distinct values and addresses in each function invocation. Typically, there
2274    is a single block of storage containing all thread-local variables declared
2275    in the main executable, and a separate block for the variables declared in
2276    each shared library. Each thread-local variable can then be accessed in its
2277    block using an identifier. This identifier is typically an offset into the
2278    block and pushed onto the DWARF stack by one of the* ``DW_OP_const<n><x>``
2279    *operations prior to the* ``DW_OP_form_tls_address`` *operation. Computing
2280    the address of the appropriate block can be complex (in some cases, the
2281    compiler emits a function call to do it), and difficult to describe using
2282    ordinary DWARF location descriptions. Instead of forcing complex
2283    thread-local storage calculations into the DWARF expressions, the*
2284    ``DW_OP_form_tls_address`` *allows the consumer to perform the computation
2285    based on the run-time environment.*
2286
22875.  ``DW_OP_call_frame_cfa``
2288
2289    ``DW_OP_call_frame_cfa`` pushes the memory location description L of the
2290    Canonical Frame Address (CFA) of the current function, obtained from the
2291    Call Frame Information (see :ref:`amdgpu-call-frame-information`).
2292
2293    *Although the value of* ``DW_AT_frame_base`` *can be computed using other
2294    DWARF expression operators, in some cases this would require an extensive
2295    location list because the values of the registers used in computing the CFA
2296    change during a subroutine. If the Call Frame Information is present, then
2297    it already encodes such changes, and it is space efficient to reference
2298    that.*
2299
23006.  ``DW_OP_fbreg``
2301
2302    ``DW_OP_fbreg`` has a single signed LEB128 integer operand that is treated
2303    as a byte displacement D.
2304
2305    The DWARF expression E corresponding to the current program location is
2306    selected from the ``DW_AT_frame_base`` attribute of the current function and
2307    evaluated. The resulting memory location description L's bit offset is
2308    updated as if the ``DW_OP_LLVM_offset D`` operation were applied. The
2309    updated L is pushed.
2310
2311    *This is typically a stack pointer register plus or minus some offset.*
2312
23137.  ``DW_OP_breg0, DW_OP_breg1, ..., DW_OP_breg31``
2314
2315    The ``DW_OP_breg<n>`` operations encode the numbers of up to 32 registers,
2316    numbered from 0 through 31, inclusive. The register number R corresponds to
2317    the ``n`` in the operation name.
2318
2319    They have a single signed LEB128 integer operand that is treated as a byte
2320    displacement D.
2321
2322    The address space identifier AS is defined as the one corresponding to the
2323    target architecture's default address space.
2324
2325    The address size S is defined as the address bit size of the target
2326    architecture's address space corresponding to AS.
2327
2328    The contents of the register specified by R is retrieved as a
2329    twos-complement unsigned value and zero extended to S bits. D is added and
2330    the least significant S bits are treated as a twos-complement unsigned value
2331    to be used as an address A.
2332
2333    They push a memory location description L that specifies the memory location
2334    storage that corresponds to AS, with a bit offset equal to A scaled by 8
2335    (the byte size).
2336
23378.  ``DW_OP_bregx``
2338
2339    ``DW_OP_bregx`` has two operands. The first is an unsigned LEB128 integer
2340    that is treated as a register number R. The second is a signed LEB128
2341    integer that is treated as a byte displacement D.
2342
2343    The action is the same as for ``DW_OP_breg<n>`` except that R is used as the
2344    register number and D is used as the byte displacement.
2345
23469.  ``DW_OP_LLVM_aspace_bregx`` *New*
2347
2348    ``DW_OP_LLVM_aspace_bregx`` has two operands. The first is an unsigned
2349    LEB128 integer that is treated as a register number R. The second is a
2350    signed LEB128 integer that is treated as a byte displacement D. It pops one
2351    stack entry that is required to be an integral type value that is treated as
2352    an address space identifier AS for those architectures that support multiple
2353    address spaces.
2354
2355    The action is the same as for ``DW_OP_breg<n>`` except that R is used as the
2356    register number, D is used as the byte displacement, and AS is used as the
2357    address space identifier.
2358
2359    If AS is not one of the values defined by the target architecture's
2360    ``DW_ASPACE_*`` values, then the DWARF expression is ill-formed.
2361
2362    .. note::
2363
2364      Could also consider adding ``DW_OP_aspace_breg0, DW_OP_aspace_breg1, ...,
2365      DW_OP_aspace_bref31`` which would save encoding size.
2366
2367.. _amdgpu-register-location-descriptions:
2368
2369Register Location Descriptions
2370##############################
2371
2372There is a register location storage that corresponds to each of the target
2373architecture registers. The size of each register location storage corresponds
2374to the size of the corresponding target architecture register.
2375
2376A register location description specifies a register location storage. The bit
2377offset corresponds to a bit position within the register. Bits accessed using a
2378register location description, access the corresponding target architecture
2379register starting at the bit offset.
2380
23811.  ``DW_OP_reg0, DW_OP_reg1, ..., DW_OP_reg31``
2382
2383    ``DW_OP_reg<n>`` operations encode the numbers of up to 32 registers,
2384    numbered from 0 through 31, inclusive. The target architecture register
2385    number R corresponds to the ``n`` in the operation name.
2386
2387    ``DW_OP_reg<n>`` pushes a register location description L that specifies the
2388    register location storage that corresponds to R, with a bit offset of 0.
2389
23902.  ``DW_OP_regx``
2391
2392    ``DW_OP_regx`` has a single unsigned LEB128 integer operand that is treated
2393    as a target architecture register number R.
2394
2395    ``DW_OP_regx`` pushes a register location description L that specifies the
2396    register location storage that corresponds to R, with a bit offset of 0.
2397
2398*These operations name a register location. To fetch the contents of a register,
2399it is necessary to use* ``DW_OP_regval_type``\ *, or one of the register based
2400addressing operations such as* ``DW_OP_bregx``\ *, or using* ``DW_OP_deref*``
2401*on a register location description.*
2402
2403.. _amdgpu-implicit-location-descriptions:
2404
2405Implicit Location Descriptions
2406##############################
2407
2408Implicit location storage represents a piece or all of an object which has no
2409actual location in the program but whose contents are nonetheless known, either
2410as a constant or can be computed from other locations and values in the program.
2411
2412An implicit location description specifies an implicit location storage. The bit
2413offset corresponds to a bit position within the implicit location storage. Bits
2414accessed using an implicit location description, access the corresponding
2415implicit storage value starting at the bit offset.
2416
24171.  ``DW_OP_implicit_value``
2418
2419    ``DW_OP_implicit_value`` has two operands. The first is an unsigned LEB128
2420    integer treated as a byte size S. The second is a block of bytes with a
2421    length equal to S treated as a literal value V.
2422
2423    An implicit location storage LS is created with the literal value V and a
2424    size of S. An implicit location description L is pushed that specifies LS
2425    with a bit offset of 0.
2426
24272.  ``DW_OP_stack_value``
2428
2429    ``DW_OP_stack_value`` pops one stack entry that must be a value treated as a
2430    literal value V.
2431
2432    An implicit location storage LS is created with the literal value V and a
2433    size equal to V's base type size. An implicit location description L is
2434    pushed that specifies LS with a bit offset of 0.
2435
2436    The ``DW_OP_stack_value`` operation specifies that the object does not exist
2437    in memory but its value is nonetheless known and is at the top of the DWARF
2438    expression stack. In this form of location description, the DWARF expression
2439    represents the actual value of the object, rather than its location.
2440
2441    See :ref:`amdgpu-implicit-location-descriptions` for special rules
2442    concerning implicit pointer values produced by dereferencing implicit
2443    location descriptions created by the ``DW_OP_implicit_pointer`` and
2444    ``DW_OP_LLVM_implicit_aspace_pointer`` operations.
2445
2446    .. note::
2447
2448      Since location descriptions are allowed on the stack, the
2449      ``DW_OP_stack_value`` operation no longer terminates the DWARF expression.
2450
24513.  ``DW_OP_implicit_pointer``
2452
2453    *An optimizing compiler may eliminate a pointer, while still retaining the
2454    value that the pointer addressed.* ``DW_OP_implicit_pointer`` *allows a
2455    producer to describe this value.*
2456
2457    ``DW_OP_implicit_pointer`` specifies that the object is a pointer to the
2458    target architecture default address space that cannot be represented as a
2459    real pointer, even though the value it would point to can be described. In
2460    this form of location description, the DWARF expression refers to a
2461    debugging information entry that represents the actual location description
2462    of the object to which the pointer would point. Thus, a consumer of the
2463    debug information would be able to access the the dereferenced pointer, even
2464    when it cannot access of the pointer itself.
2465
2466    ``DW_OP_implicit_pointer`` has two operands. The first is a 4-byte unsigned
2467    value in the 32-bit DWARF format, or an 8-byte unsigned value in the 64-bit
2468    DWARF format, that is treated as a debugging information entry reference R.
2469    The second is a signed LEB128 integer that is treated as a byte
2470    displacement D.
2471
2472    R is used as the offset of a debugging information entry E in a
2473    ``.debug_info`` section, which may be contained in an executable or shared
2474    object file other than that containing the operator. For references from one
2475    executable or shared object file to another, the relocation must be
2476    performed by the consumer.
2477
2478    *The first operand interpretation is exactly like that for*
2479    ``DW_FORM_ref_addr``\ *.*
2480
2481    The address space identifier AS is defined as the one corresponding to the
2482    target architecture's default address space.
2483
2484    The address size S is defined as the address bit size of the target
2485    architecture's address space corresponding to AS.
2486
2487    An implicit location storage LS is created that has the bit size of S. An
2488    implicit location description L is pushed that specifies LS and has a bit
2489    offset of 0.
2490
2491    If a ``DW_OP_deref*`` operation pops a location description L' and retrieves
2492    S' bits where some retrieved bits come from LS such that either:
2493
2494    1.  L' is an implicit location description that specifies LS with bit offset
2495        0, and S' equals S.
2496
2497    2.  L' is a complete composite location description that specifies a
2498        canonical form composite location storage LS'. The bits retrieved all
2499        come from a single part P' of LS'. P' has a bit size of S and has
2500        an implicit location description PL'. PL' specifies LS with a bit offset
2501        of 0.
2502
2503    Then the value V pushed by the ``DW_OP_deref*`` operation is an implicit
2504    pointer value IPV with an address space of AS, a debugging information entry
2505    of E, and a base type of T. If AS is the target architecture default address
2506    space, then T is the generic type. Otherwise, T is an architecture specific
2507    integral type with a bit size equal to S.
2508
2509    Otherwise, if a ``DW_OP_deref*`` operation is applied to a location
2510    description such that some retrieved bits come from LS, then the DWARF
2511    expression is ill-formed.
2512
2513    If IPV is either implicitly converted to a location description (only done
2514    if AS is the target architecture default address space) or used by
2515    ``DW_OP_LLVM_form_aspace_address`` (only done if the address space specified
2516    is AS), then the resulting location description is:
2517
2518    * If E has a ``DW_AT_location`` attribute, the DWARF expression
2519      corresponding to the current program location is selected and evaluated
2520      from the ``DW_AT_location`` attribute. The expression result is the
2521      resulting location description RL.
2522
2523    * If E has a ``DW_AT_const_value`` attribute, then an implicit location
2524      storage RLS is created from the ``DW_AT_const_value`` attribute's value,
2525      with a size matching the size of the ``DW_AT_const_value`` attribute's
2526      value. The resulting implicit location description RL specifies RLS with a
2527      bit offset of 0.
2528
2529      .. note::
2530
2531        If deprecate using ``DW_AT_const_value`` for variables and formal
2532        parameters and instead use ``DW_AT_location`` with an implicit location
2533        description instead, then this rule would not be required.
2534
2535    * Otherwise the DWARF expression is ill-formed.
2536
2537    The bit offset of RL is updated as if the ``DW_OP_LLVM_offset D`` operation
2538    were applied.
2539
2540    If a ``DW_OP_stack_value`` operation pops a value that is the same as IPV,
2541    then it pushes a location description that is the same as L.
2542
2543    The DWARF expression is ill-formed if it accesses LS or IPV in any other
2544    manner.
2545
2546    *The restrictions on how an implicit pointer location description created by
2547    ``DW_OP_implicit_pointer`` and ``DW_OP_LLVM_aspace_implicit_pointer``, or an
2548    implicit pointer value created by ``DW_OP_deref*``, can be used are to
2549    simplify the DWARF consumer.*
2550
25514.  ``DW_OP_LLVM_aspace_implicit_pointer`` *New*
2552
2553    ``DW_OP_LLVM_aspace_implicit_pointer`` has two operands that are the same as
2554    for ``DW_OP_implicit_pointer``.
2555
2556    It pops one stack entry that must be an integral type value that is treated
2557    as an address space identifier AS for those architectures that support
2558    multiple address spaces.
2559
2560    The implicit location description L that is pushed is the same as for
2561    ``DW_OP_implicit_pointer`` except that the address space identifier used is
2562    AS.
2563
2564    If AS is not one of the values defined by the target architecture's
2565    ``DW_ASPACE_*`` values, then the DWARF expression is ill-formed.
2566
2567*The debugging information entry referenced by a* ``DW_OP_implicit_pointer`` or
2568``DW_OP_LLVM_aspace_implicit_pointer`` *operation is typically a*
2569``DW_TAG_variable`` *or* ``DW_TAG_formal_parameter`` *entry whose*
2570``DW_AT_location`` *attribute gives a second DWARF expression or a location list
2571that describes the value of the object, but the referenced entry may be any
2572entry that contains a* ``DW_AT_location`` *or* ``DW_AT_const_value`` *attribute
2573(for example,* ``DW_TAG_dwarf_procedure``\ *). By using the second DWARF
2574expression, a consumer can reconstruct the value of the object when asked to
2575dereference the pointer described by the original DWARF expression containing
2576the* ``DW_OP_implicit_pointer`` or ``DW_OP_LLVM_aspace_implicit_pointer``
2577*operation.*
2578
2579Composite Location Descriptions
2580###############################
2581
2582A composite location storage represents an object or value which may be
2583contained in part of another location storage, or contained in parts of more
2584than one location storage.
2585
2586Each part has a part location description L and a part bit size S. The bits of
2587the part comprise S contiguous bits from the location storage specified by L,
2588starting at the bit offset specified by L. All the bits must be within the size
2589of the location storage specified by L or the DWARF expression is ill-formed.
2590
2591A composite location storage can have zero or more parts. The parts are
2592contiguous such that the zero-based location storage bit index will range over
2593each part with no gaps between them. Therefore, the size of a composite location
2594storage is the size of its parts. The DWARF expression is ill-formed if the size
2595of the contiguous location storage is larger than the size of the memory
2596location storage corresponding to the target architecture's largest address
2597space.
2598
2599The canonical form of a composite location storage is computed by applying the
2600following steps to a composite location storage:
2601
26021.  If any part P has a composite location description L, it is replaced by a
2603    copy of the parts of the composite location storage specified by L that are
2604    selected by the bit size of P starting at the bit offset of L. The location
2605    description of the first copied part has its bit offset updated as
2606    necessary, and the last copied part has its bit size updated as necessary,
2607    to reflect the bits selected by P. This rule is applied repeatedly until no
2608    part has a composite location description.
2609
26102.  If the size on any part is zero, it is removed.
2611
26123.  If any adjacent parts P\ :sup:`1` to P\ :sup:`n` have location descriptions
2613    that specify the same location storage LS such that the bits selected form a
2614    contiguous portion of LS, then they are replaced by a single new part P'. P'
2615    has a location description L that specifies LS with the same bit offset as
2616    P\ :sup:`1`\ 's location description, and a bit size equal to the sum of the
2617    bit sizes of P\ :sup:`1` to P\ :sup:`n` inclusive.
2618
2619A composite location description specifies the canonical form of a composite
2620location storage and a bit offset.
2621
2622There are operations that push a composite location description that specifies a
2623composite location storage that is created by the operation.
2624
2625There are other operations that allow a composite location storage and a
2626composite location description that specifies it to be created incrementally.
2627Each part is described by a separate operation. There may be one or more
2628operations to create the final composite location storage and associated
2629description. A series of such operations describes the parts of the composite
2630location storage that are in the order that the associated part operations are
2631executed.
2632
2633To support incremental creation, a composite location description can be in an
2634incomplete state. When an incremental operation operates on an incomplete
2635composite location description, it adds a new part, otherwise it creates a new
2636composite location description. The ``DW_OP_LLVM_piece_end`` operation
2637explicitly makes an incomplete composite location description complete.
2638
2639If the top stack entry is an incomplete composite location description after the
2640execution of a DWARF expression has completed, it is converted to a complete
2641composite location description.
2642
2643If a stack entry is required to be a location description, but it is an
2644incomplete composite location description, then the DWARF expression is
2645ill-formed.
2646
2647*Note that a DWARF expression may arbitrarily compose composite location
2648descriptions from any other location description, including other composite
2649location descriptions.*
2650
2651*The incremental composite location description operations are defined to be
2652compatible with the definitions in DWARF 5 and earlier.*
2653
26541.  ``DW_OP_piece``
2655
2656    ``DW_OP_piece`` has a single unsigned LEB128 integer that is treated as a
2657    byte size S.
2658
2659    The action is based on the context:
2660
2661    * If the stack is empty, then an incomplete composite location description
2662      L is pushed that specifies a new composite location storage LS and has a
2663      bit offset of 0. LS has a single part P that specifies the undefined
2664      location description, and has a bit size of S scaled by 8 (the byte size).
2665
2666    * If the top stack entry is an incomplete composite location description L,
2667      then the composite location storage LS that it specifies is updated to
2668      append a part that specifies an undefined location description, and has a
2669      bit size S scaled by 8 (the byte size).
2670
2671    * If the top stack entry is a location description or can be converted to
2672      one, then it is popped and treated as a part location description PL.
2673      Then:
2674
2675      * If the stack is empty or the top stack entry is not an incomplete
2676        composite location description, then an incomplete composite location
2677        description L is pushed that specifies a new composite location storage
2678        LS. LS has a single part that specifies PL, and has a bit size of S
2679        scaled by 8 (the byte size).
2680
2681      * Otherwise, the composite location storage LS specified by the top stack
2682        incomplete composite location description L is updated to append a part
2683        that specifies PL, and has a bit size S scaled by 8 (the byte size).
2684
2685    * Otherwise, the DWARF expression is ill-formed
2686
2687    If LS is not in canonical form it is updated to be in canonical form.
2688
2689    *Many compilers store a single variable in sets of registers, or store a
2690    variable partially in memory and partially in registers.* ``DW_OP_piece``
2691    *provides a way of describing how large a part of a variable a particular
2692    DWARF location description refers to.*
2693
2694    *If a computed byte displacement is required, the* ``DW_OP_LLVM_offset``
2695    *can be used to update the part location description.*
2696
26972.  ``DW_OP_bit_piece``
2698
2699    ``DW_OP_bit_piece`` has two operands. The first is an unsigned LEB128
2700    integer that is treated as the part bit size S. The second is an unsigned
2701    LEB128 integer that is treated as a bit displacement D.
2702
2703    The action is the same as for ``DW_OP_piece`` except that any part created
2704    has the bit size S, and the location description of any created part has its
2705    bit offset updated as if the ``DW_OP_LLVM_bit_offset D`` operation were
2706    applied.
2707
2708    *If a computed bit displacement is required, the* ``DW_OP_LLVM_bit_offset``
2709    *can be used to update the part location description.*
2710
2711    .. note::
2712
2713      The bit offset operand is not needed as ``DW_OP_LLVM_bit_offset`` can be
2714      used on the part's location description.
2715
27163.  ``DW_OP_LLVM_piece_end`` *New*
2717
2718    If the top stack entry is an incomplete composite location description L,
2719    then it is updated to be a complete composite location description with the
2720    same parts. Otherwise, the DWARF expression is ill-formed.
2721
27224.  ``DW_OP_LLVM_extend`` *New*
2723
2724    ``DW_OP_LLVM_extend`` has two operands. The first is an unsigned LEB128
2725    integer that is treated as the element bit size S. The second is an unsigned
2726    LEB128 integer that is treated as a count C.
2727
2728    It pops one stack entry that must be a location description and is treated
2729    as the part location description PL.
2730
2731    A complete composite location description L is pushed that comprises C parts
2732    that each specify PL and have a bit size of S.
2733
2734    The DWARF expression is ill-formed if the element bit size or count are 0.
2735
27365.  ``DW_OP_LLVM_select_bit_piece`` *New*
2737
2738    ``DW_OP_LLVM_select_bit_piece`` has two operands. The first is an unsigned
2739    LEB128 integer that is treated as the element bit size S. The second is an
2740    unsigned LEB128 integer that is treated as a count C.
2741
2742    It pops three stack entries. The first must be an integral type value that
2743    is treated as a bit mask value M. The second must be a location description
2744    that is treated as the one-location description L1. The third must be a
2745    location description that is treated as the zero-location description L0.
2746
2747    A complete composite location description L is pushed that specifies a new
2748    composite location storage LS. LS comprises C parts that each specify a part
2749    location description PL and have a bit size of S. The PL for part N is
2750    defined as:
2751
2752    1.  If the Nth least significant bit of M is a zero then the PL for part N
2753        is the same as L0, otherwise it is the same as L1.
2754
2755    2.  The PL for part N is updated as if the ``DW_OP_LLVM_bit_offset N*S``
2756        operation was applied.
2757
2758    If LS is not in canonical form it is updated to be in canonical form.
2759
2760    The DWARF expression is ill-formed if S or C are 0, or if the bit size of M
2761    is less than C.
2762
2763``DW_OP_bit_piece`` *is used instead of* ``DW_OP_piece`` *when the piece to be
2764assembled into a value or assigned to is not byte-sized or is not at the start
2765of the part location description.*
2766
2767.. note::
2768
2769  For AMDGPU:
2770
2771  * In CFI expressions ``DW_OP_LLVM_select_bit_piece`` is used to describe
2772    unwinding vector registers that are spilled under the execution mask to
2773    memory: the zero location description is the vector register, and the one
2774    location description is the spilled memory location. The
2775    ``DW_OP_LLVM_form_aspace_address`` is used to specify the address space of
2776    the memory location description.
2777
2778  * ``DW_OP_LLVM_select_bit_piece`` is used by the ``lane_pc`` attribute
2779    expression where divergent control flow is controlled by the execution mask.
2780    An undefined location description together with ``DW_OP_LLVM_extend`` is
2781    used to indicate the lane was not active on entry to the subprogram.
2782
2783Expression Operation Encodings
2784++++++++++++++++++++++++++++++
2785
2786The following table gives the encoding of the DWARF expression operations added
2787for AMDGPU.
2788
2789.. table:: AMDGPU DWARF Expression Operation Encodings
2790   :name: amdgpu-dwarf-expression-operation-encodings-table
2791
2792   ================================== ===== ======== ===============================
2793   Operation                          Code  Number   Notes
2794                                            of
2795                                            Operands
2796   ================================== ===== ======== ===============================
2797   DW_OP_LLVM_form_aspace_address     0xe7     0
2798   DW_OP_LLVM_push_lane               0xea     0
2799   DW_OP_LLVM_offset                  0xe9     0
2800   DW_OP_LLVM_offset_uconst           *TBD*    1     ULEB128 byte displacement
2801   DW_OP_LLVM_bit_offset              *TBD*    0
2802   DW_OP_LLVM_call_frame_entry_reg    *TBD*    1     ULEB128 register number
2803   DW_OP_LLVM_undefined               *TBD*    0
2804   DW_OP_LLVM_aspace_bregx            *TBD*    2     ULEB128 register number,
2805                                                     ULEB128 byte displacement
2806   DW_OP_LLVM_aspace_implicit_pointer *TBD*    2     4- or 8-byte offset of DIE,
2807                                                     SLEB128 byte displacement
2808   DW_OP_LLVM_piece_end               *TBD*    0
2809   DW_OP_LLVM_extend                  *TBD*    2     ULEB128 bit size,
2810                                                     ULEB128 count
2811   DW_OP_LLVM_select_bit_piece        *TBD*    2     ULEB128 bit size,
2812                                                     ULEB128 count
2813   ================================== ===== ======== ===============================
2814
2815.. _amdgpu-dwarf-debugging-information-entry-attributes:
2816
2817Debugging Information Entry Attributes
2818~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
2819
2820This section provides changes to existing debugger information attributes and
2821defines attributes added by the AMDGPU target.
2822
28231.  ``DW_AT_location``
2824
2825    If the result of the ``DW_AT_location`` DWARF expression is required to be a
2826    location description, then it may have any kind of location description (see
2827    :ref:`amdgpu-location-description-operations`).
2828
28292.  ``DW_AT_const_value``
2830
2831    .. note::
2832
2833      Could deprecate using the ``DW_AT_const_value`` attribute for
2834      ``DW_TAG_variable`` or ``DW_TAG_formal_parameter`` debugger information
2835      entries that are constants. Instead, ``DW_AT_location`` could be used with
2836      a DWARF expression that produces an implicit location description now that
2837      any location description can be used within a DWARF expression. This
2838      allows the ``DW_OP_call*`` operations to be used to push the location
2839      description of any variable regardless of how it is optimized.
2840
28413.  ``DW_AT_frame_base``
2842
2843    A ``DW_TAG_subprogram`` or ``DW_TAG_entry_point`` debugger information entry
2844    may have a ``DW_AT_frame_base`` attribute, whose value is a DWARF expression
2845    or location list that describes the *frame base* for the subroutine or entry
2846    point.
2847
2848    If the result of the DWARF expression is a register location description,
2849    then the ``DW_OP_deref`` operation is applied to compute the frame base
2850    memory location description in the target architecture default address
2851    space.
2852
2853    .. note::
2854
2855      This rule could be removed and require the producer to create the
2856      required location descriptor directly using ``DW_OP_call_frame_cfa``,
2857      ``DW_OP_fbreg``, ``DW_OP_breg*``, or ``DW_OP_LLVM-aspace_bregx``. This
2858      would also then allow a target to implement the call frames withing a
2859      large register.
2860
2861    Otherwise, the result of the DWARF expression is required to be a memory
2862    location description in any of the target architecture address spaces which
2863    is the frame base.
2864
28654.  ``DW_AT_data_member_location``
2866
2867    For a ``DW_AT_data_member_location`` attribute there are two cases:
2868
2869    1.  If the value is an integer constant, it is the offset in bytes from the
2870        beginning of the containing entity. If the beginning of the containing
2871        entity has a non-zero bit offset then the beginning of the member entry
2872        has that same bit offset as well.
2873
2874    2.  Otherwise, the value must be a DWARF expression or location list. The
2875        DWARF expression E corresponding to the current program location is
2876        selected. The location description of the beginning of the containing
2877        entity is pushed on the DWARF stack before E is evaluated. The result of
2878        the evaluation is the location description of the base of the member
2879        entry.
2880
2881        .. note::
2882
2883          The beginning of the containing entity can now be any location
2884          description and can be bit aligned.
2885
28865.  ``DW_AT_use_location``
2887
2888    The ``DW_TAG_ptr_to_member_type`` debugging information entry has a
2889    ``DW_AT_use_location`` attribute whose value is a DWARF expression or
2890    location list. The DWARF expression E corresponding to the current program
2891    location is selected. It is used to computes the location description of the
2892    member of the class to which the pointer to member entry points
2893
2894    *The method used to find the location description of a given member of a
2895    class or structure is common to any instance of that class or structure and
2896    to any instance of the pointer or member type. The method is thus associated
2897    with the type entry, rather than with each instance of the type.*
2898
2899    The ``DW_AT_use_location`` description is used in conjunction with the
2900    location descriptions for a particular object of the given pointer to member
2901    type and for a particular structure or class instance.
2902
2903    Two values are pushed onto the DWARF expression stack before E is evaluated.
2904    The first value pushed is the value of the pointer to member object itself.
2905    The second value pushed is the location description of the base of the
2906    entire structure or union instance containing the member whose address is
2907    being calculated.
2908
29096.  ``DW_AT_data_location``
2910
2911    The ``DW_AT_data_location`` attribute may be used with any type that
2912    provides one or more levels of hidden indirection and/or run-time parameters
2913    in its representation. Its value is a DWARF expression E which computes the
2914    location description of the data for an object. When this attribute is
2915    omitted, the location description of the data is the same as the location
2916    description of the object.
2917
2918    *E will typically begin with ``DW_OP_push_object_address`` which loads the
2919    location description of the object which can then serve as a descriptor in
2920    subsequent calculation.*
2921
29227.  ``DW_AT_vtable_elem_location``
2923
2924    An entry for a virtual function also has a ``DW_AT_vtable_elem_location``
2925    attribute whose value is a DWARF expression or location list. The DWARF
2926    expression E corresponding to the current program location is selected. The
2927    location description of the object of the enclosing type is pushed onto the
2928    expression stack before E is evaluated. The resulting location description
2929    is the slot for the function within the virtual function table for the
2930    enclosing class.
2931
29328.  ``DW_AT_static_link``
2933
2934    If a ``DW_TAG_subprogram`` or ``DW_TAG_entry_point`` debugger information
2935    entry is nested, it may have a ``DW_AT_static_link`` attribute, whose value
2936    is a DWARF expression or location list. The DWARF expression E corresponding
2937    to the current program location is selected. The result of evaluating E is
2938    the frame base memory location description of the relevant instance of the
2939    subroutine that immediately encloses the subroutine or entry point.
2940
29419.  ``DW_AT_return_addr``
2942
2943    A ``DW_TAG_subprogram``, ``DW_TAG_inlined_subroutine``, or
2944    ``DW_TAG_entry_point`` debugger information entry may have a
2945    ``DW_AT_return_addr`` attribute, whose value is a DWARF expression or
2946    location list. The DWARF expression E corresponding to the current program
2947    location is selected. The result of evaluating E is the location description
2948    for the place where the return address for the subroutine or entry point is
2949    stored.
2950
2951    .. note::
2952
2953      It is unclear why ``DW_TAG_inlined_subroutine`` has a
2954      ``DW_AT_return_addr`` attribute but not a ``DW_AT_frame_base`` or
2955      ``DW_AT_static_link`` attribute. Seems it would either have all of them or
2956      none. Since inlined subprograms do not have a frame it seems they would
2957      have none of these attributes.
2958
295910. ``DW_AT_LLVM_lanes`` *New*
2960
2961    For languages that are implemented using a SIMD or SIMT execution model, a
2962    ``DW_TAG_subprogram``, ``DW_TAG_inlined_subroutine``, or
2963    ``DW_TAG_entry_point`` debugger information entry may have a
2964    ``DW_AT_LLVM_lanes`` attribute whose value is an integer constant that is
2965    the number of lanes per thread.
2966
2967    If not present, the default value of 1 is used.
2968
2969    The DWARF is ill-formed if the value is 0.
2970
297111. ``DW_AT_LLVM_lane_pc`` *New*
2972
2973    For languages that are implemented using a SIMD or SIMT execution model, a
2974    ``DW_TAG_subprogram``, ``DW_TAG_inlined_subroutine``, or
2975    ``DW_TAG_entry_point`` debugging information entry may have a
2976    ``DW_AT_LLVM_lane_pc`` attribute whose value is a DWARF expression or
2977    location list. The DWARF expression E corresponding to the current program
2978    location is selected. The result of evaluating E is a location description
2979    that references a wave size vector of generic type elements. Each element
2980    holds the conceptual program location of the corresponding lane, where the
2981    least significant element corresponds to the first target architecture lane
2982    identifier and so forth. If the lane was not active when the subprogram was
2983    called, its element is an undefined location description.
2984
2985    *``DW_AT_LLVM_lane_pc`` allows the compiler to indicate conceptually where
2986    each lane of a SIMT thread is positioned even when it is in divergent
2987    control flow that is not active.*
2988
2989    If not present, the thread is not being used in a SIMT manner, and the
2990    thread's program location is used.
2991
2992    *See* :ref:`amdgpu-dwarf-amdgpu-dw-at-llvm-lane-pc` *for AMDGPU
2993    information.*
2994
299512. ``DW_AT_LLVM_active_lane`` *New*
2996
2997    For languages that are implemented using a SIMD or SIMT execution model, a
2998    ``DW_TAG_subprogram``, ``DW_TAG_inlined_subroutine``, or
2999    ``DW_TAG_entry_point`` debugger information entry may have a
3000    ``DW_AT_LLVM_active_lane`` attribute whose value is a DWARF expression or
3001    location list. The DWARF expression E corresponding to the current program
3002    location is selected. The result of evaluating E is a integral value that is
3003    the mask of active lanes for the current program location. The Nth least
3004    significant bit of the mask corresponds to the Nth lane. If the bit is 1 the
3005    lane is active, otherwise it is inactive.
3006
3007    *Some targets may update the target architecture execution mask for regions
3008    of code that must execute with different sets of lanes than the current
3009    active lanes. For example, some code must execute in whole wave mode.
3010    ``DW_AT_LLVM_active_lane` allows the compiler can provide the means to
3011    determine the actual active lanes.*
3012
3013    If not present and ``DW_AT_LLVM_lanes`` is greater than 1, then the target
3014    architecture execution mask is used.
3015
3016    *See* :ref:`amdgpu-dwarf-amdgpu-dw-at-llvm-active-lane` *for AMDGPU
3017    information.*
3018
301913. ``DW_AT_LLVM_vector_size`` *New*
3020
3021    A base type V may have the ``DW_AT_LLVM_vector_size`` attribute whose value
3022    is an integer constant that is the vector size S.
3023
3024    The representation of a vector base type is as S contiguous elements, each
3025    one having the representation of a base type E that is the same as V without
3026    the ``DW_AT_LLVM_vector_size`` attribute.
3027
3028    If not present, the base type is not a vector.
3029
3030    The DWARF is ill-formed if S not greater than 0.
3031
3032    .. note::
3033
3034      LLVM has mention of non-upstreamed debugger information entry that is
3035      intended to support vector types. However, that was not for a base type
3036      so would not be suitable as the type of a stack value entry. But perhaps
3037      that could be replaced by using this attribute.
3038
303914. ``DW_AT_LLVM_augmentation`` *New*
3040
3041    A compilation unit may have a ``DW_AT_LLVM_augmentation`` attribute, whose
3042    value is an augmentation string.
3043
3044    *The augmentation string allows users to indicate that there is additional
3045    target-specific information in the debugging information entries. For
3046    example, this might be information about the version of target-specific
3047    extensions that are being used.*
3048
3049    If not present, or if the string is empty, then the compilation unit has no
3050    augmentation string.
3051
3052    .. note::
3053
3054      For AMDGPU, the augmentation string contains:
3055
3056      ::
3057
3058        [amd:v0.0]
3059
3060      The "vX.Y" specifies the major X and minor Y version number of the AMDGPU
3061      extensions used in the DWARF of the compilation unit. The version number
3062      conforms to [SEMVER]_.
3063
3064Attribute Encodings
3065+++++++++++++++++++
3066
3067The following table gives the encoding of the debugging information entry
3068attributes added for AMDGPU.
3069
3070.. table:: AMDGPU DWARF Attribute Encodings
3071   :name: amdgpu-dwarf-attribute-encodings-table
3072
3073   ================================== ===== ====================================
3074   Attribute Name                     Value Classes
3075   ================================== ===== ====================================
3076   DW_AT_LLVM_lanes                         constant
3077   DW_AT_LLVM_lane_pc                       exprloc, loclist
3078   DW_AT_LLVM_active_lane                   exprloc, loclist
3079   DW_AT_LLVM_vector_size                   constant
3080   DW_AT_LLVM_augmentation                  string
3081   ================================== ===== ====================================
3082
3083.. _amdgpu-call-frame-information:
3084
3085Call Frame Information
3086~~~~~~~~~~~~~~~~~~~~~~
3087
3088DWARF Call Frame Information describes how an agent can virtually *unwind*
3089call frames in a running process or core dump.
3090
3091.. note::
3092
3093  AMDGPU conforms to the DWARF standard with additional support added for
3094  address spaces. Register unwind DWARF expressions are generalized to allow any
3095  location description, including composite and implicit location descriptions.
3096
3097Structure of Call Frame Information
3098+++++++++++++++++++++++++++++++++++
3099
3100The register rules are:
3101
3102*undefined*
3103  A register that has this rule has no recoverable value in the previous frame.
3104  (By convention, it is not preserved by a callee.)
3105
3106*same value*
3107  This register has not been modified from the previous frame. (By convention,
3108  it is preserved by the callee, but the callee has not modified it.)
3109
3110*offset(N)*
3111  The previous value of this register is saved at the location description
3112  computed as if the ``DW_OP_LLVM_offset N`` operation is applied to the current
3113  CFA memory location description where N is a signed byte offset.
3114
3115*val_offset(N)*
3116  The previous value of this register is the address in the address space of the
3117  memory location description computed as if the ``DW_OP_LLVM_offset N``
3118  operation is applied to the current CFA memory location description where N is
3119  a signed byte displacement.
3120
3121  If the register size does not match the size of an address in the address
3122  space of the current CFA memory location description, then the DWARF is
3123  ill-formed .
3124
3125*register(R)*
3126  The previous value of this register is stored in another register numbered R.
3127
3128  If the register sizes do not match, then the DWARF is ill-formed.
3129
3130*expression(E)*
3131  The previous value of this register is located at the location description
3132  produced by executing the DWARF expression E (see
3133  :ref:`amdgpu-dwarf-expressions`).
3134
3135*val_expression(E)*
3136  The previous value of this register is the value produced by executing the
3137  DWARF expression E (see :ref:`amdgpu-dwarf-expressions`).
3138
3139  If value type size does not match the register size, then the DWARF is
3140  ill-formed.
3141
3142*architectural*
3143  The rule is defined externally to this specification by the augmenter.
3144
3145A Common Information Entry holds information that is shared among many Frame
3146Description Entries. There is at least one CIE in every non-empty
3147``.debug_frame`` section. A CIE contains the following fields, in order:
3148
31491.  ``length`` (initial length)
3150
3151    A constant that gives the number of bytes of the CIE structure, not
3152    including the length field itself. The size of the length field plus the
3153    value of length must be an integral multiple of the address size specified
3154    in the ``address_size`` field.
3155
31562.  ``CIE_id`` (4 or 8 bytes, see
3157    :ref:`amdgpu-dwarf-32-bit-and-64-bit-dwarf-formats`)
3158
3159    A constant that is used to distinguish CIEs from FDEs.
3160
3161    In the 32-bit DWARF format, the value of the CIE id in the CIE header is
3162    0xffffffff; in the 64-bit DWARF format, the value is 0xffffffffffffffff.
3163
31643.  ``version`` (ubyte)
3165
3166    A version number. This number is specific to the call frame information and
3167    is independent of the DWARF version number.
3168
3169    The value of the CIE version number is 4.
3170
31714.  ``augmentation`` (sequence of UTF-8 characters)
3172
3173    A null-terminated UTF-8 string that identifies the augmentation to this CIE
3174    or to the FDEs that use it. If a reader encounters an augmentation string
3175    that is unexpected, then only the following fields can be read:
3176
3177    * CIE: length, CIE_id, version, augmentation
3178    * FDE: length, CIE_pointer, initial_location, address_range
3179
3180    If there is no augmentation, this value is a zero byte.
3181
3182    *The augmentation string allows users to indicate that there is additional
3183    target-specific information in the CIE or FDE which is needed to virtually
3184    unwind a stack frame. For example, this might be information about
3185    dynamically allocated data which needs to be freed on exit from the
3186    routine.*
3187
3188    *Because the .debug_frame section is useful independently of any
3189    ``.debug_info`` section, the augmentation string always uses UTF-8
3190    encoding.*
3191
3192    .. note::
3193
3194      For AMDGPU, the augmentation string contains:
3195
3196      ::
3197
3198        [amd:v0.0]
3199
3200      The "vX.Y" specifies the major X and minor Y version number of the AMDGPU
3201      extensions used in the DWARF of the compilation unit. The version number
3202      conforms to [SEMVER]_.
3203
32045.  ``address_size`` (ubyte)
3205
3206    The size of a target address in this CIE and any FDEs that use it, in bytes.
3207    If a compilation unit exists for this frame, its address size must match the
3208    address size here.
3209
3210    .. note::
3211
3212      For AMDGPU:
3213
3214      * The address size for the ``Global`` address space defined in
3215        :ref:`amdgpu-dwarf-address-space-mapping-table`.
3216
32176.  ``segment_selector_size`` (ubyte)
3218
3219    The size of a segment selector in this CIE and any FDEs that use it, in
3220    bytes.
3221
3222    .. note::
3223
3224      For AMDGPU:
3225
3226      * Does not use a segment selector so this is 0.
3227
32287.  ``code_alignment_factor`` (unsigned LEB128)
3229
3230    A constant that is factored out of all advance location instructions (see
3231    :ref:`amdgpu-dwarf-row-creation-instructions`). The resulting value is
3232    ``(operand * code_alignment_factor)``.
3233
3234    .. note::
3235
3236      For AMDGPU:
3237
3238      * 4 bytes.
3239
3240    .. TODO::
3241
3242       Add to :ref:`amdgpu-processor-table` table.
3243
32448.  ``data_alignment_factor`` (signed LEB128)
3245
3246    A constant that is factored out of certain offset instructions (see
3247    :ref:`amdgpu-dwarf-cfa-definition-instructions` and
3248    :ref:`amdgpu-dwarf-register-rule-instructions`). The resulting value is
3249    ``(operand * data_alignment_factor)``.
3250
3251    .. note::
3252
3253      For AMDGPU:
3254
3255      * 4 bytes.
3256
3257    .. TODO::
3258
3259       Add to :ref:`amdgpu-processor-table` table.
3260
32619.  ``return_address_register`` (unsigned LEB128)
3262
3263    An unsigned LEB128 constant that indicates which column in the rule table
3264    represents the return address of the function. Note that this column might
3265    not correspond to an actual machine register.
3266
3267    .. note::
3268
3269      For AMDGPU:
3270
3271      * ``PC_32`` for 32-bit processes and ``PC_64`` for
3272        64-bit processes defined in :ref:`amdgpu-dwarf-register-mapping`.
3273
327410. ``initial_instructions`` (array of ubyte)
3275
3276    A sequence of rules that are interpreted to create the initial setting of
3277    each column in the table.
3278
3279    The default rule for all columns before interpretation of the initial
3280    instructions is the undefined rule. However, an ABI authoring body or a
3281    compilation system authoring body may specify an alternate default value for
3282    any or all columns.
3283
3284    .. note::
3285
3286      For AMDGPU:
3287
3288      * Since a subprogram A with fewer registers can be called from subprogram
3289        B that has more allocated, A will not change any of the extra registers
3290        as it cannot access them. Therefore, The default rule for all columns is
3291        ``same value``.
3292
329311. ``padding`` (array of ubyte)
3294
3295    Enough ``DW_CFA_nop`` instructions to make the size of this entry match the
3296    length value above.
3297
3298An FDE contains the following fields, in order:
3299
33001.  ``length`` (initial length)
3301
3302    A constant that gives the number of bytes of the header and instruction
3303    stream for this function, not including the length field itself. The size of
3304    the length field plus the value of length must be an integral multiple of
3305    the address size.
3306
33072.  ``CIE_pointer`` (4 or 8 bytes, see
3308    :ref:`amdgpu-dwarf-32-bit-and-64-bit-dwarf-formats`)
3309
3310    A constant offset into the ``.debug_frame`` section that denotes the CIE
3311    that is associated with this FDE.
3312
33133.  ``initial_location`` (segment selector and target address)
3314
3315    The address of the first location associated with this table entry. If the
3316    segment_selector_size field of this FDE’s CIE is non-zero, the initial
3317    location is preceded by a segment selector of the given length.
3318
33194.  ``address_range`` (target address)
3320
3321    The number of bytes of program instructions described by this entry.
3322
33235.  ``instructions`` (array of ubyte)
3324
3325    A sequence of table defining instructions that are described in
3326    :ref:`amdgpu-dwarf-call-frame-instructions`.
3327
33286.  ``padding`` (array of ubyte)
3329
3330    Enough ``DW_CFA_nop`` instructions to make the size of this entry match the
3331    length value above.
3332
3333.. _amdgpu-dwarf-call-frame-instructions:
3334
3335Call Frame Instructions
3336+++++++++++++++++++++++
3337
3338Some call frame instructions have operands that are encoded as DWARF expressions
3339E (see :ref:`amdgpu-dwarf-expressions`). The DWARF operators that can be used in
3340E have the following restrictions:
3341
3342* ``DW_OP_addrx``, ``DW_OP_call2``, ``DW_OP_call4``, ``DW_OP_call_ref``,
3343  ``DW_OP_const_type``, ``DW_OP_constx``, ``DW_OP_convert``,
3344  ``DW_OP_deref_type``, ``DW_OP_regval_type``, and ``DW_OP_reinterpret``
3345  operators are not allowed because the call frame information must not depend
3346  on other debug sections.
3347
3348* ``DW_OP_push_object_address`` is not allowed because there is no object
3349  context to provide a value to push.
3350
3351* ``DW_OP_call_frame_cfa`` and ``DW_OP_entry_value`` are not allowed because
3352  their use would be circular.
3353
3354* ``DW_OP_LLVM_call_frame_entry_reg`` is not allowed if evaluating E causes a
3355  circular dependency between ``DW_OP_LLVM_call_frame_entry_reg`` operators.
3356
3357  *For example, if a register R1 has a* ``DW_CFA_def_cfa_expression``
3358  *instruction that evaluates a* ``DW_OP_LLVM_call_frame_entry_reg`` *operator
3359  that specifies register R2, and register R2 has a*
3360  ``DW_CFA_def_cfa_expression`` *instruction that that evaluates a*
3361  ``DW_OP_LLVM_call_frame_entry_reg`` *operator that specifies register R1.*
3362
3363*Call frame instructions to which these restrictions apply include*
3364``DW_CFA_def_cfa_expression``\ *,* ``DW_CFA_expression``\ *, and*
3365``DW_CFA_val_expression``\ *.*
3366
3367.. _amdgpu-dwarf-row-creation-instructions:
3368
3369Row Creation Instructions
3370#########################
3371
3372These instructions are the same as in DWARF 5.
3373
3374.. _amdgpu-dwarf-cfa-definition-instructions:
3375
3376CFA Definition Instructions
3377###########################
3378
33791.  ``DW_CFA_def_cfa``
3380
3381    The ``DW_CFA_def_cfa`` instruction takes two unsigned LEB128 operands
3382    representing a register number R and a (non-factored) byte displacement D.
3383    The required action is to define the current CFA rule to be the memory
3384    location description that is the result of evaluating the DWARF expression
3385    ``DW_OP_bregx R, D``.
3386
3387    .. note::
3388
3389      Could also consider adding ``DW_CFA_def_aspace_cfa`` and
3390      ``DW_CFA_def_aspace_cfa_sf`` which allow a register R, offset D, and
3391      address space AS to be specified. For example, that would save a byte of
3392      encoding over using ``DW_CFA_def_cfa R, D; DW_CFA_LLVM_def_cfa_aspace
3393      AS;``.
3394
33952.  ``DW_CFA_def_cfa_sf``
3396
3397    The ``DW_CFA_def_cfa_sf`` instruction takes two operands: an unsigned LEB128
3398    value representing a register number R and a signed LEB128 factored byte
3399    displacement D. The required action is to define the current CFA rule to be
3400    the memory location description that is the result of evaluating the DWARF
3401    expression ``DW_OP_bregx R, D*data_alignment_factor``.
3402
3403    *The action is the same as ``DW_CFA_def_cfa`` except that the second operand
3404    is signed and factored.*
3405
34063.  ``DW_CFA_def_cfa_register``
3407
3408    The ``DW_CFA_def_cfa_register`` instruction takes a single unsigned LEB128
3409    operand representing a register number R. The required action is to define
3410    the current CFA rule to be the memory location description that is the
3411    result of evaluating the DWARF expression ``DW_OP_constu AS;
3412    DW_OP_aspace_bregx R, D`` where D and AS are the old CFA byte displacement
3413    and address space respectively.
3414
3415    If the subprogram has no current CFA rule, or the rule was defined by a
3416    ``DW_CFA_def_cfa_expression`` instruction, then the DWARF is ill-formed.
3417
34184.  ``DW_CFA_def_cfa_offset``
3419
3420    The ``DW_CFA_def_cfa_offset`` instruction takes a single unsigned LEB128
3421    operand representing a (non-factored) byte displacement D. The required
3422    action is to define the current CFA rule to be the memory location
3423    description that is the result of evaluating the DWARF expression
3424    ``DW_OP_constu AS; DW_OP_aspace_bregx R, D`` where R and AS are the old CFA
3425    register number and address space respectively.
3426
3427    If the subprogram has no current CFA rule, or the rule was defined by a
3428    ``DW_CFA_def_cfa_expression`` instruction, then the DWARF is ill-formed.
3429
34305.  ``DW_CFA_def_cfa_offset_sf``
3431
3432    The ``DW_CFA_def_cfa_offset_sf`` instruction takes a signed LEB128 operand
3433    representing a factored byte displacement D. The required action is to
3434    define the current CFA rule to be the memory location description that is
3435    the result of evaluating the DWARF expression ``DW_OP_constu AS;
3436    DW_OP_aspace_bregx R, D*data_alignment_factor`` where R and AS are the old
3437    CFA register number and address space respectively.
3438
3439    If the subprogram has no current CFA rule, or the rule was defined by a
3440    ``DW_CFA_def_cfa_expression`` instruction, then the DWARF is ill-formed.
3441
3442    *The action is the same as ``DW_CFA_def_cfa_offset`` except that the operand
3443    is signed and factored.*
3444
34456.  ``DW_CFA_LLVM_def_cfa_aspace`` *New*
3446
3447    The ``DW_CFA_LLVM_def_cfa_aspace`` instruction takes a single unsigned
3448    LEB128 operand representing an address space identifier AS for those
3449    architectures that support multiple address spaces. The required action is
3450    to define the current CFA rule to be the memory location description L that
3451    is the result of evaluating the DWARF expression ``DW_OP_constu AS;
3452    DW_OP_aspace_bregx R, D`` where R and D are the old CFA register number and
3453    byte displacement respectively.
3454
3455    If AS is not one of the values defined by the target architecture's
3456    ``DW_ASPACE_*`` values then the DWARF expression is ill-formed.
3457
34587.  ``DW_CFA_def_cfa_expression``
3459
3460    The ``DW_CFA_def_cfa_expression`` instruction takes a single operand encoded
3461    as a ``DW_FORM_exprloc`` value representing a DWARF expression E. The
3462    required action is to define the current CFA rule to be the memory location
3463    description computed by evaluating E.
3464
3465    *See :ref:`amdgpu-dwarf-call-frame-instructions` regarding restrictions on
3466    the DWARF expression operators that can be used in E.*
3467
3468    If the result of evaluating E is not a memory location description with bit
3469    offset that is a multiple of 8 (the byte size), then the DWARF is
3470    ill-formed.
3471
3472.. _amdgpu-dwarf-register-rule-instructions:
3473
3474Register Rule Instructions
3475##########################
3476
3477.. note::
3478
3479  For AMDGPU:
3480
3481  * The register number follows the numbering defined in
3482    :ref:`amdgpu-dwarf-register-mapping`.
3483
34841.  ``DW_CFA_undefined``
3485
3486    The ``DW_CFA_undefined`` instruction takes a single unsigned LEB128 operand
3487    that represents a register number R. The required action is to set the rule
3488    for the register specified by R to ``undefined``.
3489
34902.  ``DW_CFA_same_value``
3491
3492    The ``DW_CFA_same_value`` instruction takes a single unsigned LEB128 operand
3493    that represents a register number R. The required action is to set the rule
3494    for the register specified by R to ``same value``.
3495
34963.  ``DW_CFA_offset``
3497
3498    The ``DW_CFA_offset`` instruction takes two operands: a register number R
3499    (encoded with the opcode) and an unsigned LEB128 constant representing a
3500    factored displacement D. The required action is to change the rule for the
3501    register specified by R to be an *offset(D*data_alignment_factor)* rule.
3502
3503    .. note::
3504
3505      Seems this should be named ``DW_CFA_offset_uf`` since the offset is
3506      unsigned factored.
3507
35084.  ``DW_CFA_offset_extended``
3509
3510    The ``DW_CFA_offset_extended`` instruction takes two unsigned LEB128
3511    operands representing a register number R and a factored displacement D.
3512    This instruction is identical to ``DW_CFA_offset`` except for the encoding
3513    and size of the register operand.
3514
3515    .. note::
3516
3517      Seems this should be named ``DW_CFA_offset_extended_uf`` since the
3518      displacement is unsigned factored.
3519
35205.  ``DW_CFA_offset_extended_sf``
3521
3522    The ``DW_CFA_offset_extended_sf`` instruction takes two operands: an
3523    unsigned LEB128 value representing a register number R and a signed LEB128
3524    factored displacement D. This instruction is identical to
3525    ``DW_CFA_offset_extended`` except that D is signed.
3526
35276.  ``DW_CFA_val_offset``
3528
3529    The ``DW_CFA_val_offset`` instruction takes two unsigned LEB128 operands
3530    representing a register number R and a factored displacement D. The required
3531    action is to change the rule for the register indicated by R to be a
3532    *val_offset(D*data_alignment_factor)* rule.
3533
3534    .. note::
3535
3536      Seems this should be named ``DW_CFA_val_offset_uf`` since the displacement
3537      is unsigned factored.
3538
35397.  ``DW_CFA_val_offset_sf``
3540
3541    The ``DW_CFA_val_offset_sf`` instruction takes two operands: an unsigned
3542    LEB128 value representing a register number R and a signed LEB128 factored
3543    displacement D. This instruction is identical to ``DW_CFA_val_offset``
3544    except that D is signed.
3545
35468.  ``DW_CFA_register``
3547
3548    The ``DW_CFA_register`` instruction takes two unsigned LEB128 operands
3549    representing register numbers R1 and R2 respectively. The required action is
3550    to set the rule for the register specified by R1 to be *register(R)* where R
3551    is R2.
3552
35539.  ``DW_CFA_expression``
3554
3555    The ``DW_CFA_expression`` instruction takes two operands: an unsigned LEB128
3556    value representing a register number R, and a ``DW_FORM_block`` value
3557    representing a DWARF expression E. The required action is to change the rule
3558    for the register specified by R to be an *expression(E)* rule. The memory
3559    location description of the current CFA is pushed on the DWARF stack prior
3560    to execution of E.
3561
3562    *That is, the DWARF expression computes the location description where the
3563    register value can be retrieved.*
3564
3565    *See :ref:`amdgpu-dwarf-call-frame-instructions` regarding restrictions on
3566    the DWARF expression operators that can be used in E.*
3567
356810. ``DW_CFA_val_expression``
3569
3570    The ``DW_CFA_val_expression`` instruction takes two operands: an unsigned
3571    LEB128 value representing a register number R, and a ``DW_FORM_block`` value
3572    representing a DWARF expression E. The required action is to change the rule
3573    for the register specified by R to be a *val_expression(E)* rule. The memory
3574    location description of the current CFA is pushed on the DWARF evaluation
3575    stack prior to execution of E.
3576
3577    *That is, E computes the value of register R.*
3578
3579    *See :ref:`amdgpu-dwarf-call-frame-instructions` regarding restrictions on
3580    the DWARF expression operators that can be used in E.*
3581
3582    If the result of evaluating E is not a value with a base type size that
3583    matches the register size, then the DWARF is ill-formed.
3584
358511. ``DW_CFA_restore``
3586
3587    The ``DW_CFA_restore`` instruction takes a single operand (encoded with the
3588    opcode) that represents a register number R. The required action is to
3589    change the rule for the register specified by R to the rule assigned it by
3590    the initial_instructions in the CIE.
3591
359212. ``DW_CFA_restore_extended``
3593
3594    The ``DW_CFA_restore_extended`` instruction takes a single unsigned LEB128
3595    operand that represents a register number R. This instruction is identical
3596    to ``DW_CFA_restore`` except for the encoding and size of the register
3597    operand.
3598
3599Row State Instructions
3600######################
3601
3602These instructions are the same as in DWARF 5.
3603
3604Call Frame Calling Address
3605++++++++++++++++++++++++++
3606
3607*When virtually unwinding frames, consumers frequently wish to obtain the
3608address of the instruction which called a subroutine. This information is not
3609always provided. Typically, however, one of the registers in the virtual unwind
3610table is the Return Address.*
3611
3612If a Return Address register is defined in the virtual unwind table, and its
3613rule is undefined (for example, by ``DW_CFA_undefined``), then there is no
3614return address and no call address, and the virtual unwind of stack activations
3615is complete.
3616
3617*In most cases the return address is in the same context as the calling address,
3618but that need not be the case, especially if the producer knows in some way the
3619call never will return. The context of the ’return address’ might be on a
3620different line, in a different lexical block, or past the end of the calling
3621subroutine. If a consumer were to assume that it was in the same context as the
3622calling address, the virtual unwind might fail.*
3623
3624*For architectures with constant-length instructions where the return address
3625immediately follows the call instruction, a simple solution is to subtract the
3626length of an instruction from the return address to obtain the calling
3627instruction. For architectures with variable-length instructions (for example,
3628x86), this is not possible. However, subtracting 1 from the return address,
3629although not guaranteed to provide the exact calling address, generally will
3630produce an address within the same context as the calling address, and that
3631usually is sufficient.*
3632
3633.. note::
3634
3635  For AMDGPU the instructions are variable size and a consumer can subtract 1
3636  from the return address to get the address of a byte within the call site
3637  instructions.
3638
3639Call Frame Information Instruction Encodings
3640++++++++++++++++++++++++++++++++++++++++++++
3641
3642The following table gives the encoding of the DWARF call frame information
3643instructions added for AMDGPU.
3644
3645.. table:: AMDGPU DWARF Call Frame Information Instruction Encodings
3646   :name: amdgpu-dwarf-call-frame-information-instruction-encodings-table
3647
3648   =================================== ==== ==== ============== ================
3649   Instruction                         High Low  Operand 1      Operand 1
3650                                       2    6
3651                                       Bits Bits
3652   =================================== ==== ==== ============== ================
3653   DW_CFA_LLVM_def_cfa_aspace          0    0Xxx ULEB128
3654   =================================== ==== ==== ============== ================
3655
3656Line Table
3657~~~~~~~~~~
3658
3659.. note::
3660
3661  AMDGPU does not use the ``isa`` state machine registers and always sets it to
3662  0.
3663
3664.. TODO::
3665
3666  Should the ``isa`` state machine register be used to indicate if the code is
3667  in wave32 or wave64 mode? Or used to specify the architecture ISA?
3668
3669Accelerated Access
3670~~~~~~~~~~~~~~~~~~
3671
3672Lookup By Name
3673++++++++++++++
3674
3675.. note::
3676
3677  For AMDGPU:
3678
3679  * The rule for debugger information entries included in the name
3680    index in the optional ``.debug_names`` section is extended to also include
3681    named ``DW_TAG_variable`` debugging information entries with a
3682    ``DW_AT_location`` attribute that includes a
3683    ``DW_OP_LLVM_form_aspace_address`` operation.
3684
3685  * The lookup by name section header ``augmentation_string`` string field contains:
3686
3687    ::
3688
3689      [amd:v0.0]
3690
3691    The "vX.Y" specifies the major X and minor Y version number of the AMDGPU
3692    extensions used in the DWARF of the compilation unit. The version number
3693    conforms to [SEMVER]_.
3694
3695Lookup By Address
3696+++++++++++++++++
3697
3698.. note::
3699
3700  For AMDGPU:
3701
3702  * The lookup by address section header table:
3703
3704    ``address_size`` (ubyte)
3705      Match the address size for the ``Global`` address space defined in
3706      :ref:`amdgpu-dwarf-address-space-mapping-table`.
3707
3708    ``segment_selector_size`` (ubyte)
3709      AMDGPU does not use a segment selector so this is 0. The entries in the
3710      ``.debug_aranges`` do not have a segment selector.
3711
3712Data Representation
3713~~~~~~~~~~~~~~~~~~~
3714
3715.. _amdgpu-dwarf-32-bit-and-64-bit-dwarf-formats:
3716
371732-Bit and 64-Bit DWARF Formats
3718+++++++++++++++++++++++++++++++
3719
3720.. note::
3721
3722  For AMDGPU:
3723
3724  * For the ``amdgcn`` target only 64-bit process address space is supported
3725  * The producer can generate either 32-bit or 64-bit DWARF format.
3726
37271.  Within the body of the ``.debug_info`` section, certain forms of attribute
3728    value depend on the choice of DWARF format as follows. For the 32-bit DWARF
3729    format, the value is a 4-byte unsigned integer; for the 64-bit DWARF format,
3730    the value is an 8-byte unsigned integer.
3731
3732    .. table:: AMDGPU DWARF ``.debug_info`` section attribute sizes
3733      :name: amdgpu-dwarf-debug-info-section-attribute-sizes
3734
3735      =================================== =====================================
3736      Form                                Role
3737      =================================== =====================================
3738      DW_FORM_line_strp                   offset in ``.debug_line_str``
3739      DW_FORM_ref_addr                    offset in ``.debug_info``
3740      DW_FORM_sec_offset                  offset in a section other than
3741                                          ``.debug_info`` or ``.debug_str``
3742      DW_FORM_strp                        offset in ``.debug_str``
3743      DW_FORM_strp_sup                    offset in ``.debug_str`` section of
3744                                          supplementary object file
3745      DW_OP_call_ref                      offset in ``.debug_info``
3746      DW_OP_implicit_pointer              offset in ``.debug_info``
3747      DW_OP_LLVM_aspace_implicit_pointer  offset in ``.debug_info``
3748      =================================== =====================================
3749
3750Unit Headers
3751++++++++++++
3752
3753.. note::
3754
3755  For AMDGPU:
3756
3757  * For AMDGPU the ``address_size`` field of the DWARF unit headers matches the
3758    address size for the ``Global`` address space defined in
3759    :ref:`amdgpu-dwarf-address-space-mapping-table`.
3760
3761.. _amdgpu-dwarf-amdgpu-dw-at-llvm-lane-pc:
3762
3763AMDGPU DW_AT_LLVM_lane_pc
3764~~~~~~~~~~~~~~~~~~~~~~~~~
3765
3766The ``DW_AT_LLVM_lane_pc`` attribute can be used to specify the program location
3767of the separate lanes of a SIMT thread. See
3768:ref:`amdgpu-dwarf-debugging-information-entry-attributes`.
3769
3770If the lane is an active lane then this will be the same as the current program
3771location.
3772
3773If the lane is inactive, but was active on entry to the subprogram, then this is
3774the program location in the subprogram at which execution of the lane is
3775conceptual positioned.
3776
3777If the lane was not active on entry to the subprogram, then this will be the
3778undefined location. A client debugger can check if the lane is part of a valid
3779work-group by checking that the lane is in the range of the associated
3780work-group within the grid, accounting for partial work-groups. If it is not
3781then the debugger can omit any information for the lane. Otherwise, the debugger
3782may repeatedly unwind the stack and inspect the ``DW_AT_LLVM_lane_pc`` of the
3783calling subprogram until it finds a non-undefined location. Conceptually the
3784lane only has the call frames that it has a non-undefined
3785``DW_AT_LLVM_lane_pc``.
3786
3787The following example illustrates how the AMDGPU backend can generate a location
3788list for the nested ``IF/THEN/ELSE`` structures of the following subprogram
3789pseudo code for a target with 64 lanes per wave.
3790
3791.. code::
3792  :number-lines:
3793
3794  SUBPROGRAM X
3795  BEGIN
3796    a;
3797    IF (c1) THEN
3798      b;
3799      IF (c2) THEN
3800        c;
3801      ELSE
3802        d;
3803      ENDIF
3804      e;
3805    ELSE
3806      f;
3807    ENDIF
3808    g;
3809  END
3810
3811The AMDGPU backend may generate the following pseudo LLVM MIR to manipulate the
3812execution mask (``EXEC``) to linearized the control flow. The condition is
3813evaluated to make a mask of the lanes for which the condition evaluates to true.
3814First the ``THEN`` region is executed by setting the ``EXEC`` mask to the
3815logical ``AND`` of the current ``EXEC`` mask with the condition mask. Then the
3816``ELSE`` region is executed by negating the ``EXEC`` mask and logical ``AND`` of
3817the saved ``EXEC`` mask at the start of the region. After the ``IF/THEN/ELSE``
3818region the ``EXEC`` mask is restored to the value it had at the beginning of the
3819region. This is shown below. Other approaches are possible, but the basic
3820concept is the same.
3821
3822.. code::
3823  :number-lines:
3824
3825  $lex_start:
3826    a;
3827    %1 = EXEC
3828    %2 = c1
3829  $lex_1_start:
3830    EXEC = %1 & %2
3831  $if_1_then:
3832      b;
3833      %3 = EXEC
3834      %4 = c2
3835  $lex_1_1_start:
3836      EXEC = %3 & %4
3837  $lex_1_1_then:
3838        c;
3839      EXEC = ~EXEC & %3
3840  $lex_1_1_else:
3841        d;
3842      EXEC = %3
3843  $lex_1_1_end:
3844      e;
3845    EXEC = ~EXEC & %1
3846  $lex_1_else:
3847      f;
3848    EXEC = %1
3849  $lex_1_end:
3850    g;
3851  $lex_end:
3852
3853To create the location list that defines the location description of a vector of
3854lane program locations, the LLVM MIR ``DBG_VALUE`` pseudo instruction can be
3855used to annotate the linearized control flow. This can be done by defining an
3856artificial variable for the lane PC. The location list created for it is used to
3857define the value of the ``DW_AT_LLVM_lane_pc`` attribute.
3858
3859A DWARF procedure is defined for each well nested structured control flow region
3860which provides the conceptual lane program location for a lane if it is not
3861active (namely it is divergent). The expression for each region inherits the
3862value of the immediately enclosing region and modifies it according to the
3863semantics of the region.
3864
3865For an ``IF/THEN/ELSE`` region the divergent program location is at the start of
3866the region for the ``THEN`` region since it is executed first. For the ``ELSE``
3867region the divergent program location is at the end of the ``IF/THEN/ELSE``
3868region since the ``THEN`` region has completed.
3869
3870The lane PC artificial variable is assigned at each region transition. It uses
3871the immediately enclosing region's DWARF procedure to compute the program
3872location for each lane assuming they are divergent, and then modifies the result
3873by inserting the current program location for each lane that the ``EXEC`` mask
3874indicates is active.
3875
3876By having separate DWARF procedures for each region, they can be reused to
3877define the value for any nested region. This reduces the amount of DWARF
3878required.
3879
3880The following provides an example using pseudo LLVM MIR.
3881
3882.. code::
3883  :number-lines:
3884
3885  $lex_start:
3886    DEFINE_DWARF %__uint_64 = DW_TAG_base_type[
3887      DW_AT_name = "__uint64";
3888      DW_AT_byte_size = 8;
3889      DW_AT_encoding = DW_ATE_unsigned;
3890    ];
3891    DEFINE_DWARF %__active_lane_pc = DW_TAG_dwarf_procedure[
3892      DW_AT_name = "__active_lane_pc";
3893      DW_AT_location = [
3894        DW_OP_regx PC;
3895        DW_OP_LLVM_extend 64, 64;
3896        DW_OP_regval_type EXEC, %uint_64;
3897        DW_OP_LLVM_select_bit_piece 64, 64;
3898      ];
3899    ];
3900    DEFINE_DWARF %__divergent_lane_pc = DW_TAG_dwarf_procedure[
3901      DW_AT_name = "__divergent_lane_pc";
3902      DW_AT_location = [
3903        DW_OP_LLVM_undefined;
3904        DW_OP_LLVM_extend 64, 64;
3905      ];
3906    ];
3907    DBG_VALUE $noreg, $noreg, %DW_AT_LLVM_lane_pc, DIExpression[
3908      DW_OP_call_ref %__divergent_lane_pc;
3909      DW_OP_call_ref %__active_lane_pc;
3910    ];
3911    a;
3912    %1 = EXEC;
3913    DBG_VALUE %1, $noreg, %__lex_1_save_exec;
3914    %2 = c1;
3915  $lex_1_start:
3916    EXEC = %1 & %2;
3917  $lex_1_then:
3918      DEFINE_DWARF %__divergent_lane_pc_1_then = DW_TAG_dwarf_procedure[
3919        DW_AT_name = "__divergent_lane_pc_1_then";
3920        DW_AT_location = DIExpression[
3921          DW_OP_call_ref %__divergent_lane_pc;
3922          DW_OP_xaddr &lex_1_start;
3923          DW_OP_stack_value;
3924          DW_OP_LLVM_extend 64, 64;
3925          DW_OP_call_ref %__lex_1_save_exec;
3926          DW_OP_deref_type 64, %__uint_64;
3927          DW_OP_LLVM_select_bit_piece 64, 64;
3928        ];
3929      ];
3930      DBG_VALUE $noreg, $noreg, %DW_AT_LLVM_lane_pc, DIExpression[
3931        DW_OP_call_ref %__divergent_lane_pc_1_then;
3932        DW_OP_call_ref %__active_lane_pc;
3933      ];
3934      b;
3935      %3 = EXEC;
3936      DBG_VALUE %3, %__lex_1_1_save_exec;
3937      %4 = c2;
3938  $lex_1_1_start:
3939      EXEC = %3 & %4;
3940  $lex_1_1_then:
3941        DEFINE_DWARF %__divergent_lane_pc_1_1_then = DW_TAG_dwarf_procedure[
3942          DW_AT_name = "__divergent_lane_pc_1_1_then";
3943          DW_AT_location = DIExpression[
3944            DW_OP_call_ref %__divergent_lane_pc_1_then;
3945            DW_OP_xaddr &lex_1_1_start;
3946            DW_OP_stack_value;
3947            DW_OP_LLVM_extend 64, 64;
3948            DW_OP_call_ref %__lex_1_1_save_exec;
3949            DW_OP_deref_type 64, %__uint_64;
3950            DW_OP_LLVM_select_bit_piece 64, 64;
3951          ];
3952        ];
3953        DBG_VALUE $noreg, $noreg, %DW_AT_LLVM_lane_pc, DIExpression[
3954          DW_OP_call_ref %__divergent_lane_pc_1_1_then;
3955          DW_OP_call_ref %__active_lane_pc;
3956        ];
3957        c;
3958      EXEC = ~EXEC & %3;
3959  $lex_1_1_else:
3960        DEFINE_DWARF %__divergent_lane_pc_1_1_else = DW_TAG_dwarf_procedure[
3961          DW_AT_name = "__divergent_lane_pc_1_1_else";
3962          DW_AT_location = DIExpression[
3963            DW_OP_call_ref %__divergent_lane_pc_1_then;
3964            DW_OP_xaddr &lex_1_1_end;
3965            DW_OP_stack_value;
3966            DW_OP_LLVM_extend 64, 64;
3967            DW_OP_call_ref %__lex_1_1_save_exec;
3968            DW_OP_deref_type 64, %__uint_64;
3969            DW_OP_LLVM_select_bit_piece 64, 64;
3970          ];
3971        ];
3972        DBG_VALUE $noreg, $noreg, %DW_AT_LLVM_lane_pc, DIExpression[
3973          DW_OP_call_ref %__divergent_lane_pc_1_1_else;
3974          DW_OP_call_ref %__active_lane_pc;
3975        ];
3976        d;
3977      EXEC = %3;
3978  $lex_1_1_end:
3979      DBG_VALUE $noreg, $noreg, %DW_AT_LLVM_lane_pc, DIExpression[
3980        DW_OP_call_ref %__divergent_lane_pc;
3981        DW_OP_call_ref %__active_lane_pc;
3982      ];
3983      e;
3984    EXEC = ~EXEC & %1;
3985  $lex_1_else:
3986      DEFINE_DWARF %__divergent_lane_pc_1_else = DW_TAG_dwarf_procedure[
3987        DW_AT_name = "__divergent_lane_pc_1_else";
3988        DW_AT_location = DIExpression[
3989          DW_OP_call_ref %__divergent_lane_pc;
3990          DW_OP_xaddr &lex_1_end;
3991          DW_OP_stack_value;
3992          DW_OP_LLVM_extend 64, 64;
3993          DW_OP_call_ref %__lex_1_save_exec;
3994          DW_OP_deref_type 64, %__uint_64;
3995          DW_OP_LLVM_select_bit_piece 64, 64;
3996        ];
3997      ];
3998      DBG_VALUE $noreg, $noreg, %DW_AT_LLVM_lane_pc, DIExpression[
3999        DW_OP_call_ref %__divergent_lane_pc_1_else;
4000        DW_OP_call_ref %__active_lane_pc;
4001      ];
4002      f;
4003    EXEC = %1;
4004  $lex_1_end:
4005    DBG_VALUE $noreg, $noreg, %DW_AT_LLVM_lane_pc DIExpression[
4006      DW_OP_call_ref %__divergent_lane_pc;
4007      DW_OP_call_ref %__active_lane_pc;
4008    ];
4009    g;
4010  $lex_end:
4011
4012The DWARF procedure ``%__active_lane_pc`` is used to update the lane pc elements
4013that are active with the current program location.
4014
4015Artificial variables %__lex_1_save_exec and %__lex_1_1_save_exec are created for
4016the execution masks saved on entry to a region. Using the ``DBG_VALUE`` pseudo
4017instruction, location lists that describes where they are allocated at any given
4018program location will be created. The compiler may allocate them to registers,
4019or spill them to memory.
4020
4021The DWARF procedures for each region use saved execution mask value to only
4022update the lanes that are active on entry to the region. All other lanes retain
4023the value of the enclosing region where they were last active. If they were not
4024active on entry to the subprogram, then will have the undefined location
4025description.
4026
4027Other structured control flow regions can be handled similarly. For example,
4028loops would set the divergent program location for the region at the end of the
4029loop. Any lanes active will be in the loop, and any lanes not active must have
4030exited the loop.
4031
4032An ``IF/THEN/ELSEIF/ELSEIF/...`` region can be treated as a nest of
4033``IF/THEN/ELSE`` regions.
4034
4035The DWARF procedures can use the active lane artificial variable described in
4036:ref:`amdgpu-dwarf-amdgpu-dw-at-llvm-active-lane` rather than the actual
4037``EXEC`` mask in order to support whole or quad wave mode.
4038
4039.. _amdgpu-dwarf-amdgpu-dw-at-llvm-active-lane:
4040
4041AMDGPU DW_AT_LLVM_active_lane
4042~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
4043
4044The ``DW_AT_LLVM_active_lane`` attribute can be used to specify the lanes that
4045are conceptually active for a SIMT thread. See
4046:ref:`amdgpu-dwarf-debugging-information-entry-attributes`.
4047
4048The execution mask may be modified to implement whole or quad wave mode
4049operations. For example, all lanes may need to temporarily be made active to
4050execute a whole wave operation. Such regions would save the ``EXEC`` mask,
4051update it to enable the necessary lanes, perform the operations, and then
4052restore the ``EXEC`` mask from the saved value. While executing the whole wave
4053region, the conceptual execution mask is the saved value, not the ``EXEC``
4054value.
4055
4056This is handled by defining an artificial variable for the active lane mask. The
4057active lane mask artificial variable would be the actual ``EXEC`` mask for
4058normal regions, and the saved execution mask for regions where the mask is
4059temporarily updated. The location list created for this artificial variable is
4060used to define the value of the ``DW_AT_LLVM_active_lane`` attribute.
4061
4062Source Text
4063~~~~~~~~~~~
4064
4065Source text for online-compiled programs (e.g. those compiled by the OpenCL
4066runtime) may be embedded into the DWARF v5 line table using the ``clang
4067-gembed-source`` option, described in table :ref:`amdgpu-debug-options`.
4068
4069For example:
4070
4071``-gembed-source``
4072  Enable the embedded source DWARF v5 extension.
4073``-gno-embed-source``
4074  Disable the embedded source DWARF v5 extension.
4075
4076  .. table:: AMDGPU Debug Options
4077     :name: amdgpu-debug-options
4078
4079     ==================== ==================================================
4080     Debug Flag           Description
4081     ==================== ==================================================
4082     -g[no-]embed-source  Enable/disable embedding source text in DWARF
4083                          debug sections. Useful for environments where
4084                          source cannot be written to disk, such as
4085                          when performing online compilation.
4086     ==================== ==================================================
4087
4088This option enables one extended content types in the DWARF v5 Line Number
4089Program Header, which is used to encode embedded source.
4090
4091  .. table:: AMDGPU DWARF Line Number Program Header Extended Content Types
4092     :name: amdgpu-dwarf-extended-content-types
4093
4094     ============================  ======================
4095     Content Type                  Form
4096     ============================  ======================
4097     ``DW_LNCT_LLVM_source``       ``DW_FORM_line_strp``
4098     ============================  ======================
4099
4100The source field will contain the UTF-8 encoded, null-terminated source text
4101with ``'\n'`` line endings. When the source field is present, consumers can use
4102the embedded source instead of attempting to discover the source on disk. When
4103the source field is absent, consumers can access the file to get the source
4104text.
4105
4106The above content type appears in the ``file_name_entry_format`` field of the
4107line table prologue, and its corresponding value appear in the ``file_names``
4108field. The current encoding of the content type is documented in table
4109:ref:`amdgpu-dwarf-extended-content-types-encoding`
4110
4111  .. table:: AMDGPU DWARF Line Number Program Header Extended Content Types Encoding
4112     :name: amdgpu-dwarf-extended-content-types-encoding
4113
4114     ============================  ====================
4115     Content Type                  Value
4116     ============================  ====================
4117     ``DW_LNCT_LLVM_source``       0x2001
4118     ============================  ====================
4119
4120.. _amdgpu-code-conventions:
4121
4122Code Conventions
4123================
4124
4125This section provides code conventions used for each supported target triple OS
4126(see :ref:`amdgpu-target-triples`).
4127
4128AMDHSA
4129------
4130
4131This section provides code conventions used when the target triple OS is
4132``amdhsa`` (see :ref:`amdgpu-target-triples`).
4133
4134.. _amdgpu-amdhsa-code-object-target-identification:
4135
4136Code Object Target Identification
4137~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
4138
4139The AMDHSA OS uses the following syntax to specify the code object
4140target as a single string:
4141
4142  ``<Architecture>-<Vendor>-<OS>-<Environment>-<Processor><Target Features>``
4143
4144Where:
4145
4146  - ``<Architecture>``, ``<Vendor>``, ``<OS>`` and ``<Environment>``
4147    are the same as the *Target Triple* (see
4148    :ref:`amdgpu-target-triples`).
4149
4150  - ``<Processor>`` is the same as the *Processor* (see
4151    :ref:`amdgpu-processors`).
4152
4153  - ``<Target Features>`` is a list of the enabled *Target Features*
4154    (see :ref:`amdgpu-target-features`), each prefixed by a plus, that
4155    apply to *Processor*. The list must be in the same order as listed
4156    in the table :ref:`amdgpu-target-feature-table`. Note that *Target
4157    Features* must be included in the list if they are enabled even if
4158    that is the default for *Processor*.
4159
4160For example:
4161
4162  ``"amdgcn-amd-amdhsa--gfx902+xnack"``
4163
4164.. _amdgpu-amdhsa-code-object-metadata:
4165
4166Code Object Metadata
4167~~~~~~~~~~~~~~~~~~~~
4168
4169The code object metadata specifies extensible metadata associated with the code
4170objects executed on HSA [HSA]_ compatible runtimes such as AMD's ROCm
4171[AMD-ROCm]_. The encoding and semantics of this metadata depends on the code
4172object version; see :ref:`amdgpu-amdhsa-code-object-metadata-v2` and
4173:ref:`amdgpu-amdhsa-code-object-metadata-v3`.
4174
4175Code object metadata is specified in a note record (see
4176:ref:`amdgpu-note-records`) and is required when the target triple OS is
4177``amdhsa`` (see :ref:`amdgpu-target-triples`). It must contain the minimum
4178information necessary to support the ROCM kernel queries. For example, the
4179segment sizes needed in a dispatch packet. In addition, a high level language
4180runtime may require other information to be included. For example, the AMD
4181OpenCL runtime records kernel argument information.
4182
4183.. _amdgpu-amdhsa-code-object-metadata-v2:
4184
4185Code Object V2 Metadata (-mattr=-code-object-v3)
4186++++++++++++++++++++++++++++++++++++++++++++++++
4187
4188.. warning:: Code Object V2 is not the default code object version emitted by
4189  this version of LLVM. For a description of the metadata generated with the
4190  default configuration (Code Object V3) see
4191  :ref:`amdgpu-amdhsa-code-object-metadata-v3`.
4192
4193Code object V2 metadata is specified by the ``NT_AMD_AMDGPU_METADATA`` note
4194record (see :ref:`amdgpu-note-records-v2`).
4195
4196The metadata is specified as a YAML formatted string (see [YAML]_ and
4197:doc:`YamlIO`).
4198
4199.. TODO::
4200
4201  Is the string null terminated? It probably should not if YAML allows it to
4202  contain null characters, otherwise it should be.
4203
4204The metadata is represented as a single YAML document comprised of the mapping
4205defined in table :ref:`amdgpu-amdhsa-code-object-metadata-map-table-v2` and
4206referenced tables.
4207
4208For boolean values, the string values of ``false`` and ``true`` are used for
4209false and true respectively.
4210
4211Additional information can be added to the mappings. To avoid conflicts, any
4212non-AMD key names should be prefixed by "*vendor-name*.".
4213
4214  .. table:: AMDHSA Code Object V2 Metadata Map
4215     :name: amdgpu-amdhsa-code-object-metadata-map-table-v2
4216
4217     ========== ============== ========= =======================================
4218     String Key Value Type     Required? Description
4219     ========== ============== ========= =======================================
4220     "Version"  sequence of    Required  - The first integer is the major
4221                2 integers                 version. Currently 1.
4222                                         - The second integer is the minor
4223                                           version. Currently 0.
4224     "Printf"   sequence of              Each string is encoded information
4225                strings                  about a printf function call. The
4226                                         encoded information is organized as
4227                                         fields separated by colon (':'):
4228
4229                                         ``ID:N:S[0]:S[1]:...:S[N-1]:FormatString``
4230
4231                                         where:
4232
4233                                         ``ID``
4234                                           A 32-bit integer as a unique id for
4235                                           each printf function call
4236
4237                                         ``N``
4238                                           A 32-bit integer equal to the number
4239                                           of arguments of printf function call
4240                                           minus 1
4241
4242                                         ``S[i]`` (where i = 0, 1, ... , N-1)
4243                                           32-bit integers for the size in bytes
4244                                           of the i-th FormatString argument of
4245                                           the printf function call
4246
4247                                         FormatString
4248                                           The format string passed to the
4249                                           printf function call.
4250     "Kernels"  sequence of    Required  Sequence of the mappings for each
4251                mapping                  kernel in the code object. See
4252                                         :ref:`amdgpu-amdhsa-code-object-kernel-metadata-map-table-v2`
4253                                         for the definition of the mapping.
4254     ========== ============== ========= =======================================
4255
4256..
4257
4258  .. table:: AMDHSA Code Object V2 Kernel Metadata Map
4259     :name: amdgpu-amdhsa-code-object-kernel-metadata-map-table-v2
4260
4261     ================= ============== ========= ================================
4262     String Key        Value Type     Required? Description
4263     ================= ============== ========= ================================
4264     "Name"            string         Required  Source name of the kernel.
4265     "SymbolName"      string         Required  Name of the kernel
4266                                                descriptor ELF symbol.
4267     "Language"        string                   Source language of the kernel.
4268                                                Values include:
4269
4270                                                - "OpenCL C"
4271                                                - "OpenCL C++"
4272                                                - "HCC"
4273                                                - "OpenMP"
4274
4275     "LanguageVersion" sequence of              - The first integer is the major
4276                       2 integers                 version.
4277                                                - The second integer is the
4278                                                  minor version.
4279     "Attrs"           mapping                  Mapping of kernel attributes.
4280                                                See
4281                                                :ref:`amdgpu-amdhsa-code-object-kernel-attribute-metadata-map-table-v2`
4282                                                for the mapping definition.
4283     "Args"            sequence of              Sequence of mappings of the
4284                       mapping                  kernel arguments. See
4285                                                :ref:`amdgpu-amdhsa-code-object-kernel-argument-metadata-map-table-v2`
4286                                                for the definition of the mapping.
4287     "CodeProps"       mapping                  Mapping of properties related to
4288                                                the kernel code. See
4289                                                :ref:`amdgpu-amdhsa-code-object-kernel-code-properties-metadata-map-table-v2`
4290                                                for the mapping definition.
4291     ================= ============== ========= ================================
4292
4293..
4294
4295  .. table:: AMDHSA Code Object V2 Kernel Attribute Metadata Map
4296     :name: amdgpu-amdhsa-code-object-kernel-attribute-metadata-map-table-v2
4297
4298     =================== ============== ========= ==============================
4299     String Key          Value Type     Required? Description
4300     =================== ============== ========= ==============================
4301     "ReqdWorkGroupSize" sequence of              If not 0, 0, 0 then all values
4302                         3 integers               must be >=1 and the dispatch
4303                                                  work-group size X, Y, Z must
4304                                                  correspond to the specified
4305                                                  values. Defaults to 0, 0, 0.
4306
4307                                                  Corresponds to the OpenCL
4308                                                  ``reqd_work_group_size``
4309                                                  attribute.
4310     "WorkGroupSizeHint" sequence of              The dispatch work-group size
4311                         3 integers               X, Y, Z is likely to be the
4312                                                  specified values.
4313
4314                                                  Corresponds to the OpenCL
4315                                                  ``work_group_size_hint``
4316                                                  attribute.
4317     "VecTypeHint"       string                   The name of a scalar or vector
4318                                                  type.
4319
4320                                                  Corresponds to the OpenCL
4321                                                  ``vec_type_hint`` attribute.
4322
4323     "RuntimeHandle"     string                   The external symbol name
4324                                                  associated with a kernel.
4325                                                  OpenCL runtime allocates a
4326                                                  global buffer for the symbol
4327                                                  and saves the kernel's address
4328                                                  to it, which is used for
4329                                                  device side enqueueing. Only
4330                                                  available for device side
4331                                                  enqueued kernels.
4332     =================== ============== ========= ==============================
4333
4334..
4335
4336  .. table:: AMDHSA Code Object V2 Kernel Argument Metadata Map
4337     :name: amdgpu-amdhsa-code-object-kernel-argument-metadata-map-table-v2
4338
4339     ================= ============== ========= ================================
4340     String Key        Value Type     Required? Description
4341     ================= ============== ========= ================================
4342     "Name"            string                   Kernel argument name.
4343     "TypeName"        string                   Kernel argument type name.
4344     "Size"            integer        Required  Kernel argument size in bytes.
4345     "Align"           integer        Required  Kernel argument alignment in
4346                                                bytes. Must be a power of two.
4347     "ValueKind"       string         Required  Kernel argument kind that
4348                                                specifies how to set up the
4349                                                corresponding argument.
4350                                                Values include:
4351
4352                                                "ByValue"
4353                                                  The argument is copied
4354                                                  directly into the kernarg.
4355
4356                                                "GlobalBuffer"
4357                                                  A global address space pointer
4358                                                  to the buffer data is passed
4359                                                  in the kernarg.
4360
4361                                                "DynamicSharedPointer"
4362                                                  A group address space pointer
4363                                                  to dynamically allocated LDS
4364                                                  is passed in the kernarg.
4365
4366                                                "Sampler"
4367                                                  A global address space
4368                                                  pointer to a S# is passed in
4369                                                  the kernarg.
4370
4371                                                "Image"
4372                                                  A global address space
4373                                                  pointer to a T# is passed in
4374                                                  the kernarg.
4375
4376                                                "Pipe"
4377                                                  A global address space pointer
4378                                                  to an OpenCL pipe is passed in
4379                                                  the kernarg.
4380
4381                                                "Queue"
4382                                                  A global address space pointer
4383                                                  to an OpenCL device enqueue
4384                                                  queue is passed in the
4385                                                  kernarg.
4386
4387                                                "HiddenGlobalOffsetX"
4388                                                  The OpenCL grid dispatch
4389                                                  global offset for the X
4390                                                  dimension is passed in the
4391                                                  kernarg.
4392
4393                                                "HiddenGlobalOffsetY"
4394                                                  The OpenCL grid dispatch
4395                                                  global offset for the Y
4396                                                  dimension is passed in the
4397                                                  kernarg.
4398
4399                                                "HiddenGlobalOffsetZ"
4400                                                  The OpenCL grid dispatch
4401                                                  global offset for the Z
4402                                                  dimension is passed in the
4403                                                  kernarg.
4404
4405                                                "HiddenNone"
4406                                                  An argument that is not used
4407                                                  by the kernel. Space needs to
4408                                                  be left for it, but it does
4409                                                  not need to be set up.
4410
4411                                                "HiddenPrintfBuffer"
4412                                                  A global address space pointer
4413                                                  to the runtime printf buffer
4414                                                  is passed in kernarg.
4415
4416                                                "HiddenHostcallBuffer"
4417                                                  A global address space pointer
4418                                                  to the runtime hostcall buffer
4419                                                  is passed in kernarg.
4420
4421                                                "HiddenDefaultQueue"
4422                                                  A global address space pointer
4423                                                  to the OpenCL device enqueue
4424                                                  queue that should be used by
4425                                                  the kernel by default is
4426                                                  passed in the kernarg.
4427
4428                                                "HiddenCompletionAction"
4429                                                  A global address space pointer
4430                                                  to help link enqueued kernels into
4431                                                  the ancestor tree for determining
4432                                                  when the parent kernel has finished.
4433
4434                                                "HiddenMultiGridSyncArg"
4435                                                  A global address space pointer for
4436                                                  multi-grid synchronization is
4437                                                  passed in the kernarg.
4438
4439     "ValueType"       string         Required  Kernel argument value type. Only
4440                                                present if "ValueKind" is
4441                                                "ByValue". For vector data
4442                                                types, the value is for the
4443                                                element type. Values include:
4444
4445                                                - "Struct"
4446                                                - "I8"
4447                                                - "U8"
4448                                                - "I16"
4449                                                - "U16"
4450                                                - "F16"
4451                                                - "I32"
4452                                                - "U32"
4453                                                - "F32"
4454                                                - "I64"
4455                                                - "U64"
4456                                                - "F64"
4457
4458                                                .. TODO::
4459                                                   How can it be determined if a
4460                                                   vector type, and what size
4461                                                   vector?
4462     "PointeeAlign"    integer                  Alignment in bytes of pointee
4463                                                type for pointer type kernel
4464                                                argument. Must be a power
4465                                                of 2. Only present if
4466                                                "ValueKind" is
4467                                                "DynamicSharedPointer".
4468     "AddrSpaceQual"   string                   Kernel argument address space
4469                                                qualifier. Only present if
4470                                                "ValueKind" is "GlobalBuffer" or
4471                                                "DynamicSharedPointer". Values
4472                                                are:
4473
4474                                                - "Private"
4475                                                - "Global"
4476                                                - "Constant"
4477                                                - "Local"
4478                                                - "Generic"
4479                                                - "Region"
4480
4481                                                .. TODO::
4482                                                   Is GlobalBuffer only Global
4483                                                   or Constant? Is
4484                                                   DynamicSharedPointer always
4485                                                   Local? Can HCC allow Generic?
4486                                                   How can Private or Region
4487                                                   ever happen?
4488     "AccQual"         string                   Kernel argument access
4489                                                qualifier. Only present if
4490                                                "ValueKind" is "Image" or
4491                                                "Pipe". Values
4492                                                are:
4493
4494                                                - "ReadOnly"
4495                                                - "WriteOnly"
4496                                                - "ReadWrite"
4497
4498                                                .. TODO::
4499                                                   Does this apply to
4500                                                   GlobalBuffer?
4501     "ActualAccQual"   string                   The actual memory accesses
4502                                                performed by the kernel on the
4503                                                kernel argument. Only present if
4504                                                "ValueKind" is "GlobalBuffer",
4505                                                "Image", or "Pipe". This may be
4506                                                more restrictive than indicated
4507                                                by "AccQual" to reflect what the
4508                                                kernel actual does. If not
4509                                                present then the runtime must
4510                                                assume what is implied by
4511                                                "AccQual" and "IsConst". Values
4512                                                are:
4513
4514                                                - "ReadOnly"
4515                                                - "WriteOnly"
4516                                                - "ReadWrite"
4517
4518     "IsConst"         boolean                  Indicates if the kernel argument
4519                                                is const qualified. Only present
4520                                                if "ValueKind" is
4521                                                "GlobalBuffer".
4522
4523     "IsRestrict"      boolean                  Indicates if the kernel argument
4524                                                is restrict qualified. Only
4525                                                present if "ValueKind" is
4526                                                "GlobalBuffer".
4527
4528     "IsVolatile"      boolean                  Indicates if the kernel argument
4529                                                is volatile qualified. Only
4530                                                present if "ValueKind" is
4531                                                "GlobalBuffer".
4532
4533     "IsPipe"          boolean                  Indicates if the kernel argument
4534                                                is pipe qualified. Only present
4535                                                if "ValueKind" is "Pipe".
4536
4537                                                .. TODO::
4538                                                   Can GlobalBuffer be pipe
4539                                                   qualified?
4540     ================= ============== ========= ================================
4541
4542..
4543
4544  .. table:: AMDHSA Code Object V2 Kernel Code Properties Metadata Map
4545     :name: amdgpu-amdhsa-code-object-kernel-code-properties-metadata-map-table-v2
4546
4547     ============================ ============== ========= =====================
4548     String Key                   Value Type     Required? Description
4549     ============================ ============== ========= =====================
4550     "KernargSegmentSize"         integer        Required  The size in bytes of
4551                                                           the kernarg segment
4552                                                           that holds the values
4553                                                           of the arguments to
4554                                                           the kernel.
4555     "GroupSegmentFixedSize"      integer        Required  The amount of group
4556                                                           segment memory
4557                                                           required by a
4558                                                           work-group in
4559                                                           bytes. This does not
4560                                                           include any
4561                                                           dynamically allocated
4562                                                           group segment memory
4563                                                           that may be added
4564                                                           when the kernel is
4565                                                           dispatched.
4566     "PrivateSegmentFixedSize"    integer        Required  The amount of fixed
4567                                                           private address space
4568                                                           memory required for a
4569                                                           work-item in
4570                                                           bytes. If the kernel
4571                                                           uses a dynamic call
4572                                                           stack then additional
4573                                                           space must be added
4574                                                           to this value for the
4575                                                           call stack.
4576     "KernargSegmentAlign"        integer        Required  The maximum byte
4577                                                           alignment of
4578                                                           arguments in the
4579                                                           kernarg segment. Must
4580                                                           be a power of 2.
4581     "WavefrontSize"              integer        Required  Wavefront size. Must
4582                                                           be a power of 2.
4583     "NumSGPRs"                   integer        Required  Number of scalar
4584                                                           registers used by a
4585                                                           wavefront for
4586                                                           GFX6-GFX10. This
4587                                                           includes the special
4588                                                           SGPRs for VCC, Flat
4589                                                           Scratch (GFX7-GFX10)
4590                                                           and XNACK (for
4591                                                           GFX8-GFX10). It does
4592                                                           not include the 16
4593                                                           SGPR added if a trap
4594                                                           handler is
4595                                                           enabled. It is not
4596                                                           rounded up to the
4597                                                           allocation
4598                                                           granularity.
4599     "NumVGPRs"                   integer        Required  Number of vector
4600                                                           registers used by
4601                                                           each work-item for
4602                                                           GFX6-GFX10
4603     "MaxFlatWorkGroupSize"       integer        Required  Maximum flat
4604                                                           work-group size
4605                                                           supported by the
4606                                                           kernel in work-items.
4607                                                           Must be >=1 and
4608                                                           consistent with
4609                                                           ReqdWorkGroupSize if
4610                                                           not 0, 0, 0.
4611     "NumSpilledSGPRs"            integer                  Number of stores from
4612                                                           a scalar register to
4613                                                           a register allocator
4614                                                           created spill
4615                                                           location.
4616     "NumSpilledVGPRs"            integer                  Number of stores from
4617                                                           a vector register to
4618                                                           a register allocator
4619                                                           created spill
4620                                                           location.
4621     ============================ ============== ========= =====================
4622
4623.. _amdgpu-amdhsa-code-object-metadata-v3:
4624
4625Code Object V3 Metadata (-mattr=+code-object-v3)
4626++++++++++++++++++++++++++++++++++++++++++++++++
4627
4628Code object V3 metadata is specified by the ``NT_AMDGPU_METADATA`` note record
4629(see :ref:`amdgpu-note-records-v3`).
4630
4631The metadata is represented as Message Pack formatted binary data (see
4632[MsgPack]_). The top level is a Message Pack map that includes the
4633keys defined in table
4634:ref:`amdgpu-amdhsa-code-object-metadata-map-table-v3` and referenced
4635tables.
4636
4637Additional information can be added to the maps. To avoid conflicts,
4638any key names should be prefixed by "*vendor-name*." where
4639``vendor-name`` can be the name of the vendor and specific vendor
4640tool that generates the information. The prefix is abbreviated to
4641simply "." when it appears within a map that has been added by the
4642same *vendor-name*.
4643
4644  .. table:: AMDHSA Code Object V3 Metadata Map
4645     :name: amdgpu-amdhsa-code-object-metadata-map-table-v3
4646
4647     ================= ============== ========= =======================================
4648     String Key        Value Type     Required? Description
4649     ================= ============== ========= =======================================
4650     "amdhsa.version"  sequence of    Required  - The first integer is the major
4651                       2 integers                 version. Currently 1.
4652                                                - The second integer is the minor
4653                                                  version. Currently 0.
4654     "amdhsa.printf"   sequence of              Each string is encoded information
4655                       strings                  about a printf function call. The
4656                                                encoded information is organized as
4657                                                fields separated by colon (':'):
4658
4659                                                ``ID:N:S[0]:S[1]:...:S[N-1]:FormatString``
4660
4661                                                where:
4662
4663                                                ``ID``
4664                                                  A 32-bit integer as a unique id for
4665                                                  each printf function call
4666
4667                                                ``N``
4668                                                  A 32-bit integer equal to the number
4669                                                  of arguments of printf function call
4670                                                  minus 1
4671
4672                                                ``S[i]`` (where i = 0, 1, ... , N-1)
4673                                                  32-bit integers for the size in bytes
4674                                                  of the i-th FormatString argument of
4675                                                  the printf function call
4676
4677                                                FormatString
4678                                                  The format string passed to the
4679                                                  printf function call.
4680     "amdhsa.kernels"  sequence of    Required  Sequence of the maps for each
4681                       map                      kernel in the code object. See
4682                                                :ref:`amdgpu-amdhsa-code-object-kernel-metadata-map-table-v3`
4683                                                for the definition of the keys included
4684                                                in that map.
4685     ================= ============== ========= =======================================
4686
4687..
4688
4689  .. table:: AMDHSA Code Object V3 Kernel Metadata Map
4690     :name: amdgpu-amdhsa-code-object-kernel-metadata-map-table-v3
4691
4692     =================================== ============== ========= ================================
4693     String Key                          Value Type     Required? Description
4694     =================================== ============== ========= ================================
4695     ".name"                             string         Required  Source name of the kernel.
4696     ".symbol"                           string         Required  Name of the kernel
4697                                                                  descriptor ELF symbol.
4698     ".language"                         string                   Source language of the kernel.
4699                                                                  Values include:
4700
4701                                                                  - "OpenCL C"
4702                                                                  - "OpenCL C++"
4703                                                                  - "HCC"
4704                                                                  - "HIP"
4705                                                                  - "OpenMP"
4706                                                                  - "Assembler"
4707
4708     ".language_version"                 sequence of              - The first integer is the major
4709                                         2 integers                 version.
4710                                                                  - The second integer is the
4711                                                                    minor version.
4712     ".args"                             sequence of              Sequence of maps of the
4713                                         map                      kernel arguments. See
4714                                                                  :ref:`amdgpu-amdhsa-code-object-kernel-argument-metadata-map-table-v3`
4715                                                                  for the definition of the keys
4716                                                                  included in that map.
4717     ".reqd_workgroup_size"              sequence of              If not 0, 0, 0 then all values
4718                                         3 integers               must be >=1 and the dispatch
4719                                                                  work-group size X, Y, Z must
4720                                                                  correspond to the specified
4721                                                                  values. Defaults to 0, 0, 0.
4722
4723                                                                  Corresponds to the OpenCL
4724                                                                  ``reqd_work_group_size``
4725                                                                  attribute.
4726     ".workgroup_size_hint"              sequence of              The dispatch work-group size
4727                                         3 integers               X, Y, Z is likely to be the
4728                                                                  specified values.
4729
4730                                                                  Corresponds to the OpenCL
4731                                                                  ``work_group_size_hint``
4732                                                                  attribute.
4733     ".vec_type_hint"                    string                   The name of a scalar or vector
4734                                                                  type.
4735
4736                                                                  Corresponds to the OpenCL
4737                                                                  ``vec_type_hint`` attribute.
4738
4739     ".device_enqueue_symbol"            string                   The external symbol name
4740                                                                  associated with a kernel.
4741                                                                  OpenCL runtime allocates a
4742                                                                  global buffer for the symbol
4743                                                                  and saves the kernel's address
4744                                                                  to it, which is used for
4745                                                                  device side enqueueing. Only
4746                                                                  available for device side
4747                                                                  enqueued kernels.
4748     ".kernarg_segment_size"             integer        Required  The size in bytes of
4749                                                                  the kernarg segment
4750                                                                  that holds the values
4751                                                                  of the arguments to
4752                                                                  the kernel.
4753     ".group_segment_fixed_size"         integer        Required  The amount of group
4754                                                                  segment memory
4755                                                                  required by a
4756                                                                  work-group in
4757                                                                  bytes. This does not
4758                                                                  include any
4759                                                                  dynamically allocated
4760                                                                  group segment memory
4761                                                                  that may be added
4762                                                                  when the kernel is
4763                                                                  dispatched.
4764     ".private_segment_fixed_size"       integer        Required  The amount of fixed
4765                                                                  private address space
4766                                                                  memory required for a
4767                                                                  work-item in
4768                                                                  bytes. If the kernel
4769                                                                  uses a dynamic call
4770                                                                  stack then additional
4771                                                                  space must be added
4772                                                                  to this value for the
4773                                                                  call stack.
4774     ".kernarg_segment_align"            integer        Required  The maximum byte
4775                                                                  alignment of
4776                                                                  arguments in the
4777                                                                  kernarg segment. Must
4778                                                                  be a power of 2.
4779     ".wavefront_size"                   integer        Required  Wavefront size. Must
4780                                                                  be a power of 2.
4781     ".sgpr_count"                       integer        Required  Number of scalar
4782                                                                  registers required by a
4783                                                                  wavefront for
4784                                                                  GFX6-GFX9. A register
4785                                                                  is required if it is
4786                                                                  used explicitly, or
4787                                                                  if a higher numbered
4788                                                                  register is used
4789                                                                  explicitly. This
4790                                                                  includes the special
4791                                                                  SGPRs for VCC, Flat
4792                                                                  Scratch (GFX7-GFX9)
4793                                                                  and XNACK (for
4794                                                                  GFX8-GFX9). It does
4795                                                                  not include the 16
4796                                                                  SGPR added if a trap
4797                                                                  handler is
4798                                                                  enabled. It is not
4799                                                                  rounded up to the
4800                                                                  allocation
4801                                                                  granularity.
4802     ".vgpr_count"                       integer        Required  Number of vector
4803                                                                  registers required by
4804                                                                  each work-item for
4805                                                                  GFX6-GFX9. A register
4806                                                                  is required if it is
4807                                                                  used explicitly, or
4808                                                                  if a higher numbered
4809                                                                  register is used
4810                                                                  explicitly.
4811     ".max_flat_workgroup_size"          integer        Required  Maximum flat
4812                                                                  work-group size
4813                                                                  supported by the
4814                                                                  kernel in work-items.
4815                                                                  Must be >=1 and
4816                                                                  consistent with
4817                                                                  ReqdWorkGroupSize if
4818                                                                  not 0, 0, 0.
4819     ".sgpr_spill_count"                 integer                  Number of stores from
4820                                                                  a scalar register to
4821                                                                  a register allocator
4822                                                                  created spill
4823                                                                  location.
4824     ".vgpr_spill_count"                 integer                  Number of stores from
4825                                                                  a vector register to
4826                                                                  a register allocator
4827                                                                  created spill
4828                                                                  location.
4829     =================================== ============== ========= ================================
4830
4831..
4832
4833  .. table:: AMDHSA Code Object V3 Kernel Argument Metadata Map
4834     :name: amdgpu-amdhsa-code-object-kernel-argument-metadata-map-table-v3
4835
4836     ====================== ============== ========= ================================
4837     String Key             Value Type     Required? Description
4838     ====================== ============== ========= ================================
4839     ".name"                string                   Kernel argument name.
4840     ".type_name"           string                   Kernel argument type name.
4841     ".size"                integer        Required  Kernel argument size in bytes.
4842     ".offset"              integer        Required  Kernel argument offset in
4843                                                     bytes. The offset must be a
4844                                                     multiple of the alignment
4845                                                     required by the argument.
4846     ".value_kind"          string         Required  Kernel argument kind that
4847                                                     specifies how to set up the
4848                                                     corresponding argument.
4849                                                     Values include:
4850
4851                                                     "by_value"
4852                                                       The argument is copied
4853                                                       directly into the kernarg.
4854
4855                                                     "global_buffer"
4856                                                       A global address space pointer
4857                                                       to the buffer data is passed
4858                                                       in the kernarg.
4859
4860                                                     "dynamic_shared_pointer"
4861                                                       A group address space pointer
4862                                                       to dynamically allocated LDS
4863                                                       is passed in the kernarg.
4864
4865                                                     "sampler"
4866                                                       A global address space
4867                                                       pointer to a S# is passed in
4868                                                       the kernarg.
4869
4870                                                     "image"
4871                                                       A global address space
4872                                                       pointer to a T# is passed in
4873                                                       the kernarg.
4874
4875                                                     "pipe"
4876                                                       A global address space pointer
4877                                                       to an OpenCL pipe is passed in
4878                                                       the kernarg.
4879
4880                                                     "queue"
4881                                                       A global address space pointer
4882                                                       to an OpenCL device enqueue
4883                                                       queue is passed in the
4884                                                       kernarg.
4885
4886                                                     "hidden_global_offset_x"
4887                                                       The OpenCL grid dispatch
4888                                                       global offset for the X
4889                                                       dimension is passed in the
4890                                                       kernarg.
4891
4892                                                     "hidden_global_offset_y"
4893                                                       The OpenCL grid dispatch
4894                                                       global offset for the Y
4895                                                       dimension is passed in the
4896                                                       kernarg.
4897
4898                                                     "hidden_global_offset_z"
4899                                                       The OpenCL grid dispatch
4900                                                       global offset for the Z
4901                                                       dimension is passed in the
4902                                                       kernarg.
4903
4904                                                     "hidden_none"
4905                                                       An argument that is not used
4906                                                       by the kernel. Space needs to
4907                                                       be left for it, but it does
4908                                                       not need to be set up.
4909
4910                                                     "hidden_printf_buffer"
4911                                                       A global address space pointer
4912                                                       to the runtime printf buffer
4913                                                       is passed in kernarg.
4914
4915                                                     "hidden_hostcall_buffer"
4916                                                       A global address space pointer
4917                                                       to the runtime hostcall buffer
4918                                                       is passed in kernarg.
4919
4920                                                     "hidden_default_queue"
4921                                                       A global address space pointer
4922                                                       to the OpenCL device enqueue
4923                                                       queue that should be used by
4924                                                       the kernel by default is
4925                                                       passed in the kernarg.
4926
4927                                                     "hidden_completion_action"
4928                                                       A global address space pointer
4929                                                       to help link enqueued kernels into
4930                                                       the ancestor tree for determining
4931                                                       when the parent kernel has finished.
4932
4933                                                     "hidden_multigrid_sync_arg"
4934                                                       A global address space pointer for
4935                                                       multi-grid synchronization is
4936                                                       passed in the kernarg.
4937
4938     ".value_type"          string         Required  Kernel argument value type. Only
4939                                                     present if ".value_kind" is
4940                                                     "by_value". For vector data
4941                                                     types, the value is for the
4942                                                     element type. Values include:
4943
4944                                                     - "struct"
4945                                                     - "i8"
4946                                                     - "u8"
4947                                                     - "i16"
4948                                                     - "u16"
4949                                                     - "f16"
4950                                                     - "i32"
4951                                                     - "u32"
4952                                                     - "f32"
4953                                                     - "i64"
4954                                                     - "u64"
4955                                                     - "f64"
4956
4957                                                     .. TODO::
4958                                                        How can it be determined if a
4959                                                        vector type, and what size
4960                                                        vector?
4961     ".pointee_align"       integer                  Alignment in bytes of pointee
4962                                                     type for pointer type kernel
4963                                                     argument. Must be a power
4964                                                     of 2. Only present if
4965                                                     ".value_kind" is
4966                                                     "dynamic_shared_pointer".
4967     ".address_space"       string                   Kernel argument address space
4968                                                     qualifier. Only present if
4969                                                     ".value_kind" is "global_buffer" or
4970                                                     "dynamic_shared_pointer". Values
4971                                                     are:
4972
4973                                                     - "private"
4974                                                     - "global"
4975                                                     - "constant"
4976                                                     - "local"
4977                                                     - "generic"
4978                                                     - "region"
4979
4980                                                     .. TODO::
4981                                                        Is "global_buffer" only "global"
4982                                                        or "constant"? Is
4983                                                        "dynamic_shared_pointer" always
4984                                                        "local"? Can HCC allow "generic"?
4985                                                        How can "private" or "region"
4986                                                        ever happen?
4987     ".access"              string                   Kernel argument access
4988                                                     qualifier. Only present if
4989                                                     ".value_kind" is "image" or
4990                                                     "pipe". Values
4991                                                     are:
4992
4993                                                     - "read_only"
4994                                                     - "write_only"
4995                                                     - "read_write"
4996
4997                                                     .. TODO::
4998                                                        Does this apply to
4999                                                        "global_buffer"?
5000     ".actual_access"       string                   The actual memory accesses
5001                                                     performed by the kernel on the
5002                                                     kernel argument. Only present if
5003                                                     ".value_kind" is "global_buffer",
5004                                                     "image", or "pipe". This may be
5005                                                     more restrictive than indicated
5006                                                     by ".access" to reflect what the
5007                                                     kernel actual does. If not
5008                                                     present then the runtime must
5009                                                     assume what is implied by
5010                                                     ".access" and ".is_const"      . Values
5011                                                     are:
5012
5013                                                     - "read_only"
5014                                                     - "write_only"
5015                                                     - "read_write"
5016
5017     ".is_const"            boolean                  Indicates if the kernel argument
5018                                                     is const qualified. Only present
5019                                                     if ".value_kind" is
5020                                                     "global_buffer".
5021
5022     ".is_restrict"         boolean                  Indicates if the kernel argument
5023                                                     is restrict qualified. Only
5024                                                     present if ".value_kind" is
5025                                                     "global_buffer".
5026
5027     ".is_volatile"         boolean                  Indicates if the kernel argument
5028                                                     is volatile qualified. Only
5029                                                     present if ".value_kind" is
5030                                                     "global_buffer".
5031
5032     ".is_pipe"             boolean                  Indicates if the kernel argument
5033                                                     is pipe qualified. Only present
5034                                                     if ".value_kind" is "pipe".
5035
5036                                                     .. TODO::
5037                                                        Can "global_buffer" be pipe
5038                                                        qualified?
5039     ====================== ============== ========= ================================
5040
5041..
5042
5043Kernel Dispatch
5044~~~~~~~~~~~~~~~
5045
5046The HSA architected queuing language (AQL) defines a user space memory
5047interface that can be used to control the dispatch of kernels, in an agent
5048independent way. An agent can have zero or more AQL queues created for it using
5049the ROCm runtime, in which AQL packets (all of which are 64 bytes) can be
5050placed. See the *HSA Platform System Architecture Specification* [HSA]_ for the
5051AQL queue mechanics and packet layouts.
5052
5053The packet processor of a kernel agent is responsible for detecting and
5054dispatching HSA kernels from the AQL queues associated with it. For AMD GPUs the
5055packet processor is implemented by the hardware command processor (CP),
5056asynchronous dispatch controller (ADC) and shader processor input controller
5057(SPI).
5058
5059The ROCm runtime can be used to allocate an AQL queue object. It uses the kernel
5060mode driver to initialize and register the AQL queue with CP.
5061
5062To dispatch a kernel the following actions are performed. This can occur in the
5063CPU host program, or from an HSA kernel executing on a GPU.
5064
50651. A pointer to an AQL queue for the kernel agent on which the kernel is to be
5066   executed is obtained.
50672. A pointer to the kernel descriptor (see
5068   :ref:`amdgpu-amdhsa-kernel-descriptor`) of the kernel to execute is obtained.
5069   It must be for a kernel that is contained in a code object that that was
5070   loaded by the ROCm runtime on the kernel agent with which the AQL queue is
5071   associated.
50723. Space is allocated for the kernel arguments using the ROCm runtime allocator
5073   for a memory region with the kernarg property for the kernel agent that will
5074   execute the kernel. It must be at least 16 byte aligned.
50754. Kernel argument values are assigned to the kernel argument memory
5076   allocation. The layout is defined in the *HSA Programmer's Language
5077   Reference* [HSA]_. For AMDGPU the kernel execution directly accesses the
5078   kernel argument memory in the same way constant memory is accessed. (Note
5079   that the HSA specification allows an implementation to copy the kernel
5080   argument contents to another location that is accessed by the kernel.)
50815. An AQL kernel dispatch packet is created on the AQL queue. The ROCm runtime
5082   api uses 64-bit atomic operations to reserve space in the AQL queue for the
5083   packet. The packet must be set up, and the final write must use an atomic
5084   store release to set the packet kind to ensure the packet contents are
5085   visible to the kernel agent. AQL defines a doorbell signal mechanism to
5086   notify the kernel agent that the AQL queue has been updated. These rules, and
5087   the layout of the AQL queue and kernel dispatch packet is defined in the *HSA
5088   System Architecture Specification* [HSA]_.
50896. A kernel dispatch packet includes information about the actual dispatch,
5090   such as grid and work-group size, together with information from the code
5091   object about the kernel, such as segment sizes. The ROCm runtime queries on
5092   the kernel symbol can be used to obtain the code object values which are
5093   recorded in the :ref:`amdgpu-amdhsa-code-object-metadata`.
50947. CP executes micro-code and is responsible for detecting and setting up the
5095   GPU to execute the wavefronts of a kernel dispatch.
50968. CP ensures that when the a wavefront starts executing the kernel machine
5097   code, the scalar general purpose registers (SGPR) and vector general purpose
5098   registers (VGPR) are set up as required by the machine code. The required
5099   setup is defined in the :ref:`amdgpu-amdhsa-kernel-descriptor`. The initial
5100   register state is defined in
5101   :ref:`amdgpu-amdhsa-initial-kernel-execution-state`.
51029. The prolog of the kernel machine code (see
5103   :ref:`amdgpu-amdhsa-kernel-prolog`) sets up the machine state as necessary
5104   before continuing executing the machine code that corresponds to the kernel.
510510. When the kernel dispatch has completed execution, CP signals the completion
5106    signal specified in the kernel dispatch packet if not 0.
5107
5108Image and Samplers
5109~~~~~~~~~~~~~~~~~~
5110
5111Image and sample handles created by the ROCm runtime are 64-bit addresses of a
5112hardware 32 byte V# and 48 byte S# object respectively. In order to support the
5113HSA ``query_sampler`` operations two extra dwords are used to store the HSA BRIG
5114enumeration values for the queries that are not trivially deducible from the S#
5115representation.
5116
5117HSA Signals
5118~~~~~~~~~~~
5119
5120HSA signal handles created by the ROCm runtime are 64-bit addresses of a
5121structure allocated in memory accessible from both the CPU and GPU. The
5122structure is defined by the ROCm runtime and subject to change between releases
5123(see [AMD-ROCm-github]_).
5124
5125.. _amdgpu-amdhsa-hsa-aql-queue:
5126
5127HSA AQL Queue
5128~~~~~~~~~~~~~
5129
5130The HSA AQL queue structure is defined by the ROCm runtime and subject to change
5131between releases (see [AMD-ROCm-github]_). For some processors it contains
5132fields needed to implement certain language features such as the flat address
5133aperture bases. It also contains fields used by CP such as managing the
5134allocation of scratch memory.
5135
5136.. _amdgpu-amdhsa-kernel-descriptor:
5137
5138Kernel Descriptor
5139~~~~~~~~~~~~~~~~~
5140
5141A kernel descriptor consists of the information needed by CP to initiate the
5142execution of a kernel, including the entry point address of the machine code
5143that implements the kernel.
5144
5145Kernel Descriptor for GFX6-GFX10
5146++++++++++++++++++++++++++++++++
5147
5148CP microcode requires the Kernel descriptor to be allocated on 64 byte
5149alignment.
5150
5151  .. table:: Kernel Descriptor for GFX6-GFX10
5152     :name: amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table
5153
5154     ======= ======= =============================== ============================
5155     Bits    Size    Field Name                      Description
5156     ======= ======= =============================== ============================
5157     31:0    4 bytes GROUP_SEGMENT_FIXED_SIZE        The amount of fixed local
5158                                                     address space memory
5159                                                     required for a work-group
5160                                                     in bytes. This does not
5161                                                     include any dynamically
5162                                                     allocated local address
5163                                                     space memory that may be
5164                                                     added when the kernel is
5165                                                     dispatched.
5166     63:32   4 bytes PRIVATE_SEGMENT_FIXED_SIZE      The amount of fixed
5167                                                     private address space
5168                                                     memory required for a
5169                                                     work-item in bytes. If
5170                                                     is_dynamic_callstack is 1
5171                                                     then additional space must
5172                                                     be added to this value for
5173                                                     the call stack.
5174     127:64  8 bytes                                 Reserved, must be 0.
5175     191:128 8 bytes KERNEL_CODE_ENTRY_BYTE_OFFSET   Byte offset (possibly
5176                                                     negative) from base
5177                                                     address of kernel
5178                                                     descriptor to kernel's
5179                                                     entry point instruction
5180                                                     which must be 256 byte
5181                                                     aligned.
5182     351:272 20                                      Reserved, must be 0.
5183             bytes
5184     383:352 4 bytes COMPUTE_PGM_RSRC3               GFX6-9
5185                                                       Reserved, must be 0.
5186                                                     GFX10
5187                                                       Compute Shader (CS)
5188                                                       program settings used by
5189                                                       CP to set up
5190                                                       ``COMPUTE_PGM_RSRC3``
5191                                                       configuration
5192                                                       register. See
5193                                                       :ref:`amdgpu-amdhsa-compute_pgm_rsrc3-gfx10-table`.
5194     415:384 4 bytes COMPUTE_PGM_RSRC1               Compute Shader (CS)
5195                                                     program settings used by
5196                                                     CP to set up
5197                                                     ``COMPUTE_PGM_RSRC1``
5198                                                     configuration
5199                                                     register. See
5200                                                     :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`.
5201     447:416 4 bytes COMPUTE_PGM_RSRC2               Compute Shader (CS)
5202                                                     program settings used by
5203                                                     CP to set up
5204                                                     ``COMPUTE_PGM_RSRC2``
5205                                                     configuration
5206                                                     register. See
5207                                                     :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`.
5208     448     1 bit   ENABLE_SGPR_PRIVATE_SEGMENT     Enable the setup of the
5209                     _BUFFER                         SGPR user data registers
5210                                                     (see
5211                                                     :ref:`amdgpu-amdhsa-initial-kernel-execution-state`).
5212
5213                                                     The total number of SGPR
5214                                                     user data registers
5215                                                     requested must not exceed
5216                                                     16 and match value in
5217                                                     ``compute_pgm_rsrc2.user_sgpr.user_sgpr_count``.
5218                                                     Any requests beyond 16
5219                                                     will be ignored.
5220     449     1 bit   ENABLE_SGPR_DISPATCH_PTR        *see above*
5221     450     1 bit   ENABLE_SGPR_QUEUE_PTR           *see above*
5222     451     1 bit   ENABLE_SGPR_KERNARG_SEGMENT_PTR *see above*
5223     452     1 bit   ENABLE_SGPR_DISPATCH_ID         *see above*
5224     453     1 bit   ENABLE_SGPR_FLAT_SCRATCH_INIT   *see above*
5225     454     1 bit   ENABLE_SGPR_PRIVATE_SEGMENT     *see above*
5226                     _SIZE
5227     457:455 3 bits                                  Reserved, must be 0.
5228     458     1 bit   ENABLE_WAVEFRONT_SIZE32         GFX6-9
5229                                                       Reserved, must be 0.
5230                                                     GFX10
5231                                                       - If 0 execute in
5232                                                         wavefront size 64 mode.
5233                                                       - If 1 execute in
5234                                                         native wavefront size
5235                                                         32 mode.
5236     463:459 5 bits                                  Reserved, must be 0.
5237     511:464 6 bytes                                 Reserved, must be 0.
5238     512     **Total size 64 bytes.**
5239     ======= ====================================================================
5240
5241..
5242
5243  .. table:: compute_pgm_rsrc1 for GFX6-GFX10
5244     :name: amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table
5245
5246     ======= ======= =============================== ===========================================================================
5247     Bits    Size    Field Name                      Description
5248     ======= ======= =============================== ===========================================================================
5249     5:0     6 bits  GRANULATED_WORKITEM_VGPR_COUNT  Number of vector register
5250                                                     blocks used by each work-item;
5251                                                     granularity is device
5252                                                     specific:
5253
5254                                                     GFX6-GFX9
5255                                                       - vgprs_used 0..256
5256                                                       - max(0, ceil(vgprs_used / 4) - 1)
5257                                                     GFX10 (wavefront size 64)
5258                                                       - max_vgpr 1..256
5259                                                       - max(0, ceil(vgprs_used / 4) - 1)
5260                                                     GFX10 (wavefront size 32)
5261                                                       - max_vgpr 1..256
5262                                                       - max(0, ceil(vgprs_used / 8) - 1)
5263
5264                                                     Where vgprs_used is defined
5265                                                     as the highest VGPR number
5266                                                     explicitly referenced plus
5267                                                     one.
5268
5269                                                     Used by CP to set up
5270                                                     ``COMPUTE_PGM_RSRC1.VGPRS``.
5271
5272                                                     The
5273                                                     :ref:`amdgpu-assembler`
5274                                                     calculates this
5275                                                     automatically for the
5276                                                     selected processor from
5277                                                     values provided to the
5278                                                     `.amdhsa_kernel` directive
5279                                                     by the
5280                                                     `.amdhsa_next_free_vgpr`
5281                                                     nested directive (see
5282                                                     :ref:`amdhsa-kernel-directives-table`).
5283     9:6     4 bits  GRANULATED_WAVEFRONT_SGPR_COUNT Number of scalar register
5284                                                     blocks used by a wavefront;
5285                                                     granularity is device
5286                                                     specific:
5287
5288                                                     GFX6-GFX8
5289                                                       - sgprs_used 0..112
5290                                                       - max(0, ceil(sgprs_used / 8) - 1)
5291                                                     GFX9
5292                                                       - sgprs_used 0..112
5293                                                       - 2 * max(0, ceil(sgprs_used / 16) - 1)
5294                                                     GFX10
5295                                                       Reserved, must be 0.
5296                                                       (128 SGPRs always
5297                                                       allocated.)
5298
5299                                                     Where sgprs_used is
5300                                                     defined as the highest
5301                                                     SGPR number explicitly
5302                                                     referenced plus one, plus
5303                                                     a target-specific number
5304                                                     of additional special
5305                                                     SGPRs for VCC,
5306                                                     FLAT_SCRATCH (GFX7+) and
5307                                                     XNACK_MASK (GFX8+), and
5308                                                     any additional
5309                                                     target-specific
5310                                                     limitations. It does not
5311                                                     include the 16 SGPRs added
5312                                                     if a trap handler is
5313                                                     enabled.
5314
5315                                                     The target-specific
5316                                                     limitations and special
5317                                                     SGPR layout are defined in
5318                                                     the hardware
5319                                                     documentation, which can
5320                                                     be found in the
5321                                                     :ref:`amdgpu-processors`
5322                                                     table.
5323
5324                                                     Used by CP to set up
5325                                                     ``COMPUTE_PGM_RSRC1.SGPRS``.
5326
5327                                                     The
5328                                                     :ref:`amdgpu-assembler`
5329                                                     calculates this
5330                                                     automatically for the
5331                                                     selected processor from
5332                                                     values provided to the
5333                                                     `.amdhsa_kernel` directive
5334                                                     by the
5335                                                     `.amdhsa_next_free_sgpr`
5336                                                     and `.amdhsa_reserve_*`
5337                                                     nested directives (see
5338                                                     :ref:`amdhsa-kernel-directives-table`).
5339     11:10   2 bits  PRIORITY                        Must be 0.
5340
5341                                                     Start executing wavefront
5342                                                     at the specified priority.
5343
5344                                                     CP is responsible for
5345                                                     filling in
5346                                                     ``COMPUTE_PGM_RSRC1.PRIORITY``.
5347     13:12   2 bits  FLOAT_ROUND_MODE_32             Wavefront starts execution
5348                                                     with specified rounding
5349                                                     mode for single (32
5350                                                     bit) floating point
5351                                                     precision floating point
5352                                                     operations.
5353
5354                                                     Floating point rounding
5355                                                     mode values are defined in
5356                                                     :ref:`amdgpu-amdhsa-floating-point-rounding-mode-enumeration-values-table`.
5357
5358                                                     Used by CP to set up
5359                                                     ``COMPUTE_PGM_RSRC1.FLOAT_MODE``.
5360     15:14   2 bits  FLOAT_ROUND_MODE_16_64          Wavefront starts execution
5361                                                     with specified rounding
5362                                                     denorm mode for half/double (16
5363                                                     and 64-bit) floating point
5364                                                     precision floating point
5365                                                     operations.
5366
5367                                                     Floating point rounding
5368                                                     mode values are defined in
5369                                                     :ref:`amdgpu-amdhsa-floating-point-rounding-mode-enumeration-values-table`.
5370
5371                                                     Used by CP to set up
5372                                                     ``COMPUTE_PGM_RSRC1.FLOAT_MODE``.
5373     17:16   2 bits  FLOAT_DENORM_MODE_32            Wavefront starts execution
5374                                                     with specified denorm mode
5375                                                     for single (32
5376                                                     bit)  floating point
5377                                                     precision floating point
5378                                                     operations.
5379
5380                                                     Floating point denorm mode
5381                                                     values are defined in
5382                                                     :ref:`amdgpu-amdhsa-floating-point-denorm-mode-enumeration-values-table`.
5383
5384                                                     Used by CP to set up
5385                                                     ``COMPUTE_PGM_RSRC1.FLOAT_MODE``.
5386     19:18   2 bits  FLOAT_DENORM_MODE_16_64         Wavefront starts execution
5387                                                     with specified denorm mode
5388                                                     for half/double (16
5389                                                     and 64-bit) floating point
5390                                                     precision floating point
5391                                                     operations.
5392
5393                                                     Floating point denorm mode
5394                                                     values are defined in
5395                                                     :ref:`amdgpu-amdhsa-floating-point-denorm-mode-enumeration-values-table`.
5396
5397                                                     Used by CP to set up
5398                                                     ``COMPUTE_PGM_RSRC1.FLOAT_MODE``.
5399     20      1 bit   PRIV                            Must be 0.
5400
5401                                                     Start executing wavefront
5402                                                     in privilege trap handler
5403                                                     mode.
5404
5405                                                     CP is responsible for
5406                                                     filling in
5407                                                     ``COMPUTE_PGM_RSRC1.PRIV``.
5408     21      1 bit   ENABLE_DX10_CLAMP               Wavefront starts execution
5409                                                     with DX10 clamp mode
5410                                                     enabled. Used by the vector
5411                                                     ALU to force DX10 style
5412                                                     treatment of NaN's (when
5413                                                     set, clamp NaN to zero,
5414                                                     otherwise pass NaN
5415                                                     through).
5416
5417                                                     Used by CP to set up
5418                                                     ``COMPUTE_PGM_RSRC1.DX10_CLAMP``.
5419     22      1 bit   DEBUG_MODE                      Must be 0.
5420
5421                                                     Start executing wavefront
5422                                                     in single step mode.
5423
5424                                                     CP is responsible for
5425                                                     filling in
5426                                                     ``COMPUTE_PGM_RSRC1.DEBUG_MODE``.
5427     23      1 bit   ENABLE_IEEE_MODE                Wavefront starts execution
5428                                                     with IEEE mode
5429                                                     enabled. Floating point
5430                                                     opcodes that support
5431                                                     exception flag gathering
5432                                                     will quiet and propagate
5433                                                     signaling-NaN inputs per
5434                                                     IEEE 754-2008. Min_dx10 and
5435                                                     max_dx10 become IEEE
5436                                                     754-2008 compliant due to
5437                                                     signaling-NaN propagation
5438                                                     and quieting.
5439
5440                                                     Used by CP to set up
5441                                                     ``COMPUTE_PGM_RSRC1.IEEE_MODE``.
5442     24      1 bit   BULKY                           Must be 0.
5443
5444                                                     Only one work-group allowed
5445                                                     to execute on a compute
5446                                                     unit.
5447
5448                                                     CP is responsible for
5449                                                     filling in
5450                                                     ``COMPUTE_PGM_RSRC1.BULKY``.
5451     25      1 bit   CDBG_USER                       Must be 0.
5452
5453                                                     Flag that can be used to
5454                                                     control debugging code.
5455
5456                                                     CP is responsible for
5457                                                     filling in
5458                                                     ``COMPUTE_PGM_RSRC1.CDBG_USER``.
5459     26      1 bit   FP16_OVFL                       GFX6-GFX8
5460                                                       Reserved, must be 0.
5461                                                     GFX9-GFX10
5462                                                       Wavefront starts execution
5463                                                       with specified fp16 overflow
5464                                                       mode.
5465
5466                                                       - If 0, fp16 overflow generates
5467                                                         +/-INF values.
5468                                                       - If 1, fp16 overflow that is the
5469                                                         result of an +/-INF input value
5470                                                         or divide by 0 produces a +/-INF,
5471                                                         otherwise clamps computed
5472                                                         overflow to +/-MAX_FP16 as
5473                                                         appropriate.
5474
5475                                                       Used by CP to set up
5476                                                       ``COMPUTE_PGM_RSRC1.FP16_OVFL``.
5477     28:27   2 bits                                  Reserved, must be 0.
5478     29      1 bit    WGP_MODE                       GFX6-GFX9
5479                                                       Reserved, must be 0.
5480                                                     GFX10
5481                                                       - If 0 execute work-groups in
5482                                                         CU wavefront execution mode.
5483                                                       - If 1 execute work-groups on
5484                                                         in WGP wavefront execution mode.
5485
5486                                                       See :ref:`amdgpu-amdhsa-memory-model`.
5487
5488                                                       Used by CP to set up
5489                                                       ``COMPUTE_PGM_RSRC1.WGP_MODE``.
5490     30      1 bit    MEM_ORDERED                    GFX6-9
5491                                                       Reserved, must be 0.
5492                                                     GFX10
5493                                                       Controls the behavior of the
5494                                                       waitcnt's vmcnt and vscnt
5495                                                       counters.
5496
5497                                                       - If 0 vmcnt reports completion
5498                                                         of load and atomic with return
5499                                                         out of order with sample
5500                                                         instructions, and the vscnt
5501                                                         reports the completion of
5502                                                         store and atomic without
5503                                                         return in order.
5504                                                       - If 1 vmcnt reports completion
5505                                                         of load, atomic with return
5506                                                         and sample instructions in
5507                                                         order, and the vscnt reports
5508                                                         the completion of store and
5509                                                         atomic without return in order.
5510
5511                                                       Used by CP to set up
5512                                                       ``COMPUTE_PGM_RSRC1.MEM_ORDERED``.
5513     31      1 bit    FWD_PROGRESS                   GFX6-9
5514                                                       Reserved, must be 0.
5515                                                     GFX10
5516                                                       - If 0 execute SIMD wavefronts
5517                                                         using oldest first policy.
5518                                                       - If 1 execute SIMD wavefronts to
5519                                                         ensure wavefronts will make some
5520                                                         forward progress.
5521
5522                                                       Used by CP to set up
5523                                                       ``COMPUTE_PGM_RSRC1.FWD_PROGRESS``.
5524     32      **Total size 4 bytes**
5525     ======= ===================================================================================================================
5526
5527..
5528
5529  .. table:: compute_pgm_rsrc2 for GFX6-GFX10
5530     :name: amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table
5531
5532     ======= ======= =============================== ===========================================================================
5533     Bits    Size    Field Name                      Description
5534     ======= ======= =============================== ===========================================================================
5535     0       1 bit   ENABLE_SGPR_PRIVATE_SEGMENT     Enable the setup of the
5536                     _WAVEFRONT_OFFSET               SGPR wavefront scratch offset
5537                                                     system register (see
5538                                                     :ref:`amdgpu-amdhsa-initial-kernel-execution-state`).
5539
5540                                                     Used by CP to set up
5541                                                     ``COMPUTE_PGM_RSRC2.SCRATCH_EN``.
5542     5:1     5 bits  USER_SGPR_COUNT                 The total number of SGPR
5543                                                     user data registers
5544                                                     requested. This number must
5545                                                     match the number of user
5546                                                     data registers enabled.
5547
5548                                                     Used by CP to set up
5549                                                     ``COMPUTE_PGM_RSRC2.USER_SGPR``.
5550     6       1 bit   ENABLE_TRAP_HANDLER             Must be 0.
5551
5552                                                     This bit represents
5553                                                     ``COMPUTE_PGM_RSRC2.TRAP_PRESENT``,
5554                                                     which is set by the CP if
5555                                                     the runtime has installed a
5556                                                     trap handler.
5557     7       1 bit   ENABLE_SGPR_WORKGROUP_ID_X      Enable the setup of the
5558                                                     system SGPR register for
5559                                                     the work-group id in the X
5560                                                     dimension (see
5561                                                     :ref:`amdgpu-amdhsa-initial-kernel-execution-state`).
5562
5563                                                     Used by CP to set up
5564                                                     ``COMPUTE_PGM_RSRC2.TGID_X_EN``.
5565     8       1 bit   ENABLE_SGPR_WORKGROUP_ID_Y      Enable the setup of the
5566                                                     system SGPR register for
5567                                                     the work-group id in the Y
5568                                                     dimension (see
5569                                                     :ref:`amdgpu-amdhsa-initial-kernel-execution-state`).
5570
5571                                                     Used by CP to set up
5572                                                     ``COMPUTE_PGM_RSRC2.TGID_Y_EN``.
5573     9       1 bit   ENABLE_SGPR_WORKGROUP_ID_Z      Enable the setup of the
5574                                                     system SGPR register for
5575                                                     the work-group id in the Z
5576                                                     dimension (see
5577                                                     :ref:`amdgpu-amdhsa-initial-kernel-execution-state`).
5578
5579                                                     Used by CP to set up
5580                                                     ``COMPUTE_PGM_RSRC2.TGID_Z_EN``.
5581     10      1 bit   ENABLE_SGPR_WORKGROUP_INFO      Enable the setup of the
5582                                                     system SGPR register for
5583                                                     work-group information (see
5584                                                     :ref:`amdgpu-amdhsa-initial-kernel-execution-state`).
5585
5586                                                     Used by CP to set up
5587                                                     ``COMPUTE_PGM_RSRC2.TGID_SIZE_EN``.
5588     12:11   2 bits  ENABLE_VGPR_WORKITEM_ID         Enable the setup of the
5589                                                     VGPR system registers used
5590                                                     for the work-item ID.
5591                                                     :ref:`amdgpu-amdhsa-system-vgpr-work-item-id-enumeration-values-table`
5592                                                     defines the values.
5593
5594                                                     Used by CP to set up
5595                                                     ``COMPUTE_PGM_RSRC2.TIDIG_CMP_CNT``.
5596     13      1 bit   ENABLE_EXCEPTION_ADDRESS_WATCH  Must be 0.
5597
5598                                                     Wavefront starts execution
5599                                                     with address watch
5600                                                     exceptions enabled which
5601                                                     are generated when L1 has
5602                                                     witnessed a thread access
5603                                                     an *address of
5604                                                     interest*.
5605
5606                                                     CP is responsible for
5607                                                     filling in the address
5608                                                     watch bit in
5609                                                     ``COMPUTE_PGM_RSRC2.EXCP_EN_MSB``
5610                                                     according to what the
5611                                                     runtime requests.
5612     14      1 bit   ENABLE_EXCEPTION_MEMORY         Must be 0.
5613
5614                                                     Wavefront starts execution
5615                                                     with memory violation
5616                                                     exceptions exceptions
5617                                                     enabled which are generated
5618                                                     when a memory violation has
5619                                                     occurred for this wavefront from
5620                                                     L1 or LDS
5621                                                     (write-to-read-only-memory,
5622                                                     mis-aligned atomic, LDS
5623                                                     address out of range,
5624                                                     illegal address, etc.).
5625
5626                                                     CP sets the memory
5627                                                     violation bit in
5628                                                     ``COMPUTE_PGM_RSRC2.EXCP_EN_MSB``
5629                                                     according to what the
5630                                                     runtime requests.
5631     23:15   9 bits  GRANULATED_LDS_SIZE             Must be 0.
5632
5633                                                     CP uses the rounded value
5634                                                     from the dispatch packet,
5635                                                     not this value, as the
5636                                                     dispatch may contain
5637                                                     dynamically allocated group
5638                                                     segment memory. CP writes
5639                                                     directly to
5640                                                     ``COMPUTE_PGM_RSRC2.LDS_SIZE``.
5641
5642                                                     Amount of group segment
5643                                                     (LDS) to allocate for each
5644                                                     work-group. Granularity is
5645                                                     device specific:
5646
5647                                                     GFX6:
5648                                                       roundup(lds-size / (64 * 4))
5649                                                     GFX7-GFX10:
5650                                                       roundup(lds-size / (128 * 4))
5651
5652     24      1 bit   ENABLE_EXCEPTION_IEEE_754_FP    Wavefront starts execution
5653                     _INVALID_OPERATION              with specified exceptions
5654                                                     enabled.
5655
5656                                                     Used by CP to set up
5657                                                     ``COMPUTE_PGM_RSRC2.EXCP_EN``
5658                                                     (set from bits 0..6).
5659
5660                                                     IEEE 754 FP Invalid
5661                                                     Operation
5662     25      1 bit   ENABLE_EXCEPTION_FP_DENORMAL    FP Denormal one or more
5663                     _SOURCE                         input operands is a
5664                                                     denormal number
5665     26      1 bit   ENABLE_EXCEPTION_IEEE_754_FP    IEEE 754 FP Division by
5666                     _DIVISION_BY_ZERO               Zero
5667     27      1 bit   ENABLE_EXCEPTION_IEEE_754_FP    IEEE 754 FP FP Overflow
5668                     _OVERFLOW
5669     28      1 bit   ENABLE_EXCEPTION_IEEE_754_FP    IEEE 754 FP Underflow
5670                     _UNDERFLOW
5671     29      1 bit   ENABLE_EXCEPTION_IEEE_754_FP    IEEE 754 FP Inexact
5672                     _INEXACT
5673     30      1 bit   ENABLE_EXCEPTION_INT_DIVIDE_BY  Integer Division by Zero
5674                     _ZERO                           (rcp_iflag_f32 instruction
5675                                                     only)
5676     31      1 bit                                   Reserved, must be 0.
5677     32      **Total size 4 bytes.**
5678     ======= ===================================================================================================================
5679
5680..
5681
5682  .. table:: compute_pgm_rsrc3 for GFX10
5683     :name: amdgpu-amdhsa-compute_pgm_rsrc3-gfx10-table
5684
5685     ======= ======= =============================== ===========================================================================
5686     Bits    Size    Field Name                      Description
5687     ======= ======= =============================== ===========================================================================
5688     3:0     4 bits  SHARED_VGPR_COUNT               Number of shared VGPRs for wavefront size 64. Granularity 8. Value 0-120.
5689                                                     compute_pgm_rsrc1.vgprs + shared_vgpr_cnt cannot exceed 64.
5690     31:4    28                                      Reserved, must be 0.
5691             bits
5692     32      **Total size 4 bytes.**
5693     ======= ===================================================================================================================
5694
5695..
5696
5697  .. table:: Floating Point Rounding Mode Enumeration Values
5698     :name: amdgpu-amdhsa-floating-point-rounding-mode-enumeration-values-table
5699
5700     ====================================== ===== ==============================
5701     Enumeration Name                       Value Description
5702     ====================================== ===== ==============================
5703     FLOAT_ROUND_MODE_NEAR_EVEN             0     Round Ties To Even
5704     FLOAT_ROUND_MODE_PLUS_INFINITY         1     Round Toward +infinity
5705     FLOAT_ROUND_MODE_MINUS_INFINITY        2     Round Toward -infinity
5706     FLOAT_ROUND_MODE_ZERO                  3     Round Toward 0
5707     ====================================== ===== ==============================
5708
5709..
5710
5711  .. table:: Floating Point Denorm Mode Enumeration Values
5712     :name: amdgpu-amdhsa-floating-point-denorm-mode-enumeration-values-table
5713
5714     ====================================== ===== ==============================
5715     Enumeration Name                       Value Description
5716     ====================================== ===== ==============================
5717     FLOAT_DENORM_MODE_FLUSH_SRC_DST        0     Flush Source and Destination
5718                                                  Denorms
5719     FLOAT_DENORM_MODE_FLUSH_DST            1     Flush Output Denorms
5720     FLOAT_DENORM_MODE_FLUSH_SRC            2     Flush Source Denorms
5721     FLOAT_DENORM_MODE_FLUSH_NONE           3     No Flush
5722     ====================================== ===== ==============================
5723
5724..
5725
5726  .. table:: System VGPR Work-Item ID Enumeration Values
5727     :name: amdgpu-amdhsa-system-vgpr-work-item-id-enumeration-values-table
5728
5729     ======================================== ===== ============================
5730     Enumeration Name                         Value Description
5731     ======================================== ===== ============================
5732     SYSTEM_VGPR_WORKITEM_ID_X                0     Set work-item X dimension
5733                                                    ID.
5734     SYSTEM_VGPR_WORKITEM_ID_X_Y              1     Set work-item X and Y
5735                                                    dimensions ID.
5736     SYSTEM_VGPR_WORKITEM_ID_X_Y_Z            2     Set work-item X, Y and Z
5737                                                    dimensions ID.
5738     SYSTEM_VGPR_WORKITEM_ID_UNDEFINED        3     Undefined.
5739     ======================================== ===== ============================
5740
5741.. _amdgpu-amdhsa-initial-kernel-execution-state:
5742
5743Initial Kernel Execution State
5744~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
5745
5746This section defines the register state that will be set up by the packet
5747processor prior to the start of execution of every wavefront. This is limited by
5748the constraints of the hardware controllers of CP/ADC/SPI.
5749
5750The order of the SGPR registers is defined, but the compiler can specify which
5751ones are actually setup in the kernel descriptor using the ``enable_sgpr_*`` bit
5752fields (see :ref:`amdgpu-amdhsa-kernel-descriptor`). The register numbers used
5753for enabled registers are dense starting at SGPR0: the first enabled register is
5754SGPR0, the next enabled register is SGPR1 etc.; disabled registers do not have
5755an SGPR number.
5756
5757The initial SGPRs comprise up to 16 User SRGPs that are set by CP and apply to
5758all wavefronts of the grid. It is possible to specify more than 16 User SGPRs
5759using the ``enable_sgpr_*`` bit fields, in which case only the first 16 are
5760actually initialized. These are then immediately followed by the System SGPRs
5761that are set up by ADC/SPI and can have different values for each wavefront of
5762the grid dispatch.
5763
5764SGPR register initial state is defined in
5765:ref:`amdgpu-amdhsa-sgpr-register-set-up-order-table`.
5766
5767  .. table:: SGPR Register Set Up Order
5768     :name: amdgpu-amdhsa-sgpr-register-set-up-order-table
5769
5770     ========== ========================== ====== ==============================
5771     SGPR Order Name                       Number Description
5772                (kernel descriptor enable  of
5773                field)                     SGPRs
5774     ========== ========================== ====== ==============================
5775     First      Private Segment Buffer     4      V# that can be used, together
5776                (enable_sgpr_private              with Scratch Wavefront Offset
5777                _segment_buffer)                  as an offset, to access the
5778                                                  private address space using a
5779                                                  segment address.
5780
5781                                                  CP uses the value provided by
5782                                                  the runtime.
5783     then       Dispatch Ptr               2      64-bit address of AQL dispatch
5784                (enable_sgpr_dispatch_ptr)        packet for kernel dispatch
5785                                                  actually executing.
5786     then       Queue Ptr                  2      64-bit address of amd_queue_t
5787                (enable_sgpr_queue_ptr)           object for AQL queue on which
5788                                                  the dispatch packet was
5789                                                  queued.
5790     then       Kernarg Segment Ptr        2      64-bit address of Kernarg
5791                (enable_sgpr_kernarg              segment. This is directly
5792                _segment_ptr)                     copied from the
5793                                                  kernarg_address in the kernel
5794                                                  dispatch packet.
5795
5796                                                  Having CP load it once avoids
5797                                                  loading it at the beginning of
5798                                                  every wavefront.
5799     then       Dispatch Id                2      64-bit Dispatch ID of the
5800                (enable_sgpr_dispatch_id)         dispatch packet being
5801                                                  executed.
5802     then       Flat Scratch Init          2      This is 2 SGPRs:
5803                (enable_sgpr_flat_scratch
5804                _init)                            GFX6
5805                                                    Not supported.
5806                                                  GFX7-GFX8
5807                                                    The first SGPR is a 32-bit
5808                                                    byte offset from
5809                                                    ``SH_HIDDEN_PRIVATE_BASE_VIMID``
5810                                                    to per SPI base of memory
5811                                                    for scratch for the queue
5812                                                    executing the kernel
5813                                                    dispatch. CP obtains this
5814                                                    from the runtime. (The
5815                                                    Scratch Segment Buffer base
5816                                                    address is
5817                                                    ``SH_HIDDEN_PRIVATE_BASE_VIMID``
5818                                                    plus this offset.) The value
5819                                                    of Scratch Wavefront Offset must
5820                                                    be added to this offset by
5821                                                    the kernel machine code,
5822                                                    right shifted by 8, and
5823                                                    moved to the FLAT_SCRATCH_HI
5824                                                    SGPR register.
5825                                                    FLAT_SCRATCH_HI corresponds
5826                                                    to SGPRn-4 on GFX7, and
5827                                                    SGPRn-6 on GFX8 (where SGPRn
5828                                                    is the highest numbered SGPR
5829                                                    allocated to the wavefront).
5830                                                    FLAT_SCRATCH_HI is
5831                                                    multiplied by 256 (as it is
5832                                                    in units of 256 bytes) and
5833                                                    added to
5834                                                    ``SH_HIDDEN_PRIVATE_BASE_VIMID``
5835                                                    to calculate the per wavefront
5836                                                    FLAT SCRATCH BASE in flat
5837                                                    memory instructions that
5838                                                    access the scratch
5839                                                    aperture.
5840
5841                                                    The second SGPR is 32-bit
5842                                                    byte size of a single
5843                                                    work-item's scratch memory
5844                                                    usage. CP obtains this from
5845                                                    the runtime, and it is
5846                                                    always a multiple of DWORD.
5847                                                    CP checks that the value in
5848                                                    the kernel dispatch packet
5849                                                    Private Segment Byte Size is
5850                                                    not larger, and requests the
5851                                                    runtime to increase the
5852                                                    queue's scratch size if
5853                                                    necessary. The kernel code
5854                                                    must move it to
5855                                                    FLAT_SCRATCH_LO which is
5856                                                    SGPRn-3 on GFX7 and SGPRn-5
5857                                                    on GFX8. FLAT_SCRATCH_LO is
5858                                                    used as the FLAT SCRATCH
5859                                                    SIZE in flat memory
5860                                                    instructions. Having CP load
5861                                                    it once avoids loading it at
5862                                                    the beginning of every
5863                                                    wavefront.
5864                                                  GFX9-GFX10
5865                                                    This is the
5866                                                    64-bit base address of the
5867                                                    per SPI scratch backing
5868                                                    memory managed by SPI for
5869                                                    the queue executing the
5870                                                    kernel dispatch. CP obtains
5871                                                    this from the runtime (and
5872                                                    divides it if there are
5873                                                    multiple Shader Arrays each
5874                                                    with its own SPI). The value
5875                                                    of Scratch Wavefront Offset must
5876                                                    be added by the kernel
5877                                                    machine code and the result
5878                                                    moved to the FLAT_SCRATCH
5879                                                    SGPR which is SGPRn-6 and
5880                                                    SGPRn-5. It is used as the
5881                                                    FLAT SCRATCH BASE in flat
5882                                                    memory instructions.
5883     then       Private Segment Size       1      The 32-bit byte size of a
5884                                                  (enable_sgpr_private single
5885                                                  work-item's
5886                                                  scratch_segment_size) memory
5887                                                  allocation. This is the
5888                                                  value from the kernel
5889                                                  dispatch packet Private
5890                                                  Segment Byte Size rounded up
5891                                                  by CP to a multiple of
5892                                                  DWORD.
5893
5894                                                  Having CP load it once avoids
5895                                                  loading it at the beginning of
5896                                                  every wavefront.
5897
5898                                                  This is not used for
5899                                                  GFX7-GFX8 since it is the same
5900                                                  value as the second SGPR of
5901                                                  Flat Scratch Init. However, it
5902                                                  may be needed for GFX9-GFX10 which
5903                                                  changes the meaning of the
5904                                                  Flat Scratch Init value.
5905     then       Grid Work-Group Count X    1      32-bit count of the number of
5906                (enable_sgpr_grid                 work-groups in the X dimension
5907                _workgroup_count_X)               for the grid being
5908                                                  executed. Computed from the
5909                                                  fields in the kernel dispatch
5910                                                  packet as ((grid_size.x +
5911                                                  workgroup_size.x - 1) /
5912                                                  workgroup_size.x).
5913     then       Grid Work-Group Count Y    1      32-bit count of the number of
5914                (enable_sgpr_grid                 work-groups in the Y dimension
5915                _workgroup_count_Y &&             for the grid being
5916                less than 16 previous             executed. Computed from the
5917                SGPRs)                            fields in the kernel dispatch
5918                                                  packet as ((grid_size.y +
5919                                                  workgroup_size.y - 1) /
5920                                                  workgroupSize.y).
5921
5922                                                  Only initialized if <16
5923                                                  previous SGPRs initialized.
5924     then       Grid Work-Group Count Z    1      32-bit count of the number of
5925                (enable_sgpr_grid                 work-groups in the Z dimension
5926                _workgroup_count_Z &&             for the grid being
5927                less than 16 previous             executed. Computed from the
5928                SGPRs)                            fields in the kernel dispatch
5929                                                  packet as ((grid_size.z +
5930                                                  workgroup_size.z - 1) /
5931                                                  workgroupSize.z).
5932
5933                                                  Only initialized if <16
5934                                                  previous SGPRs initialized.
5935     then       Work-Group Id X            1      32-bit work-group id in X
5936                (enable_sgpr_workgroup_id         dimension of grid for
5937                _X)                               wavefront.
5938     then       Work-Group Id Y            1      32-bit work-group id in Y
5939                (enable_sgpr_workgroup_id         dimension of grid for
5940                _Y)                               wavefront.
5941     then       Work-Group Id Z            1      32-bit work-group id in Z
5942                (enable_sgpr_workgroup_id         dimension of grid for
5943                _Z)                               wavefront.
5944     then       Work-Group Info            1      {first_wavefront, 14'b0000,
5945                (enable_sgpr_workgroup            ordered_append_term[10:0],
5946                _info)                            threadgroup_size_in_wavefronts[5:0]}
5947     then       Scratch Wavefront Offset   1      32-bit byte offset from base
5948                (enable_sgpr_private              of scratch base of queue
5949                _segment_wavefront_offset)        executing the kernel
5950                                                  dispatch. Must be used as an
5951                                                  offset with Private
5952                                                  segment address when using
5953                                                  Scratch Segment Buffer. It
5954                                                  must be used to set up FLAT
5955                                                  SCRATCH for flat addressing
5956                                                  (see
5957                                                  :ref:`amdgpu-amdhsa-kernel-prolog-flat-scratch`).
5958     ========== ========================== ====== ==============================
5959
5960The order of the VGPR registers is defined, but the compiler can specify which
5961ones are actually setup in the kernel descriptor using the ``enable_vgpr*`` bit
5962fields (see :ref:`amdgpu-amdhsa-kernel-descriptor`). The register numbers used
5963for enabled registers are dense starting at VGPR0: the first enabled register is
5964VGPR0, the next enabled register is VGPR1 etc.; disabled registers do not have a
5965VGPR number.
5966
5967VGPR register initial state is defined in
5968:ref:`amdgpu-amdhsa-vgpr-register-set-up-order-table`.
5969
5970  .. table:: VGPR Register Set Up Order
5971     :name: amdgpu-amdhsa-vgpr-register-set-up-order-table
5972
5973     ========== ========================== ====== ==============================
5974     VGPR Order Name                       Number Description
5975                (kernel descriptor enable  of
5976                field)                     VGPRs
5977     ========== ========================== ====== ==============================
5978     First      Work-Item Id X             1      32-bit work item id in X
5979                (Always initialized)              dimension of work-group for
5980                                                  wavefront lane.
5981     then       Work-Item Id Y             1      32-bit work item id in Y
5982                (enable_vgpr_workitem_id          dimension of work-group for
5983                > 0)                              wavefront lane.
5984     then       Work-Item Id Z             1      32-bit work item id in Z
5985                (enable_vgpr_workitem_id          dimension of work-group for
5986                > 1)                              wavefront lane.
5987     ========== ========================== ====== ==============================
5988
5989The setting of registers is done by GPU CP/ADC/SPI hardware as follows:
5990
59911. SGPRs before the Work-Group Ids are set by CP using the 16 User Data
5992   registers.
59932. Work-group Id registers X, Y, Z are set by ADC which supports any
5994   combination including none.
59953. Scratch Wavefront Offset is set by SPI in a per wavefront basis which is why
5996   its value cannot included with the flat scratch init value which is per
5997   queue.
59984. The VGPRs are set by SPI which only supports specifying either (X), (X, Y)
5999   or (X, Y, Z).
6000
6001Flat Scratch register pair are adjacent SGRRs so they can be moved as a 64-bit
6002value to the hardware required SGPRn-3 and SGPRn-4 respectively.
6003
6004The global segment can be accessed either using buffer instructions (GFX6 which
6005has V# 64-bit address support), flat instructions (GFX7-GFX10), or global
6006instructions (GFX9-GFX10).
6007
6008If buffer operations are used then the compiler can generate a V# with the
6009following properties:
6010
6011* base address of 0
6012* no swizzle
6013* ATC: 1 if IOMMU present (such as APU)
6014* ptr64: 1
6015* MTYPE set to support memory coherence that matches the runtime (such as CC for
6016  APU and NC for dGPU).
6017
6018.. _amdgpu-amdhsa-kernel-prolog:
6019
6020Kernel Prolog
6021~~~~~~~~~~~~~
6022
6023The compiler performs initialization in the kernel prologue depending on the
6024target and information about things like stack usage in the kernel and called
6025functions. Some of this initialization requires the compiler to request certain
6026User and System SGPRs be present in the
6027:ref:`amdgpu-amdhsa-initial-kernel-execution-state` via the
6028:ref:`amdgpu-amdhsa-kernel-descriptor`.
6029
6030.. _amdgpu-amdhsa-kernel-prolog-cfi:
6031
6032CFI
6033+++
6034
60351. The CFI return address is undefined.
60362. The CFI CFA is defined using an expression which evaluates to a memory
6037   location description for the private segment address ``0``.
6038
6039.. _amdgpu-amdhsa-kernel-prolog-m0:
6040
6041M0
6042++
6043
6044GFX6-GFX8
6045  The M0 register must be initialized with a value at least the total LDS size
6046  if the kernel may access LDS via DS or flat operations. Total LDS size is
6047  available in dispatch packet. For M0, it is also possible to use maximum
6048  possible value of LDS for given target (0x7FFF for GFX6 and 0xFFFF for
6049  GFX7-GFX8).
6050GFX9-GFX10
6051  The M0 register is not used for range checking LDS accesses and so does not
6052  need to be initialized in the prolog.
6053
6054.. _amdgpu-amdhsa-kernel-prolog-stack-pointer:
6055
6056Stack Pointer
6057+++++++++++++
6058
6059If the kernel has function calls it must set up the ABI stack pointer described
6060in :ref:`amdgpu-amdhsa-function-call-convention-non-kernel-functions` by
6061setting SGPR32 to the the unswizzled scratch offset of the address past the
6062last local allocation.
6063
6064.. _amdgpu-amdhsa-kernel-prolog-frame-pointer:
6065
6066Frame Pointer
6067+++++++++++++
6068
6069If the kernel needs a frame pointer for the reasons defined in
6070``SIFrameLowering`` then SGPR33 is used and is always set to ``0`` in the
6071kernel prolog. If a frame pointer is not required then all uses of the frame
6072pointer are replaced with immediate ``0`` offsets.
6073
6074.. _amdgpu-amdhsa-kernel-prolog-flat-scratch:
6075
6076Flat Scratch
6077++++++++++++
6078
6079If the kernel or any function it calls may use flat operations to access
6080scratch memory, the prolog code must set up the FLAT_SCRATCH register pair
6081(FLAT_SCRATCH_LO/FLAT_SCRATCH_HI which are in SGPRn-4/SGPRn-3). Initialization
6082uses Flat Scratch Init and Scratch Wavefront Offset SGPR registers (see
6083:ref:`amdgpu-amdhsa-initial-kernel-execution-state`):
6084
6085GFX6
6086  Flat scratch is not supported.
6087
6088GFX7-GFX8
6089
6090  1. The low word of Flat Scratch Init is 32-bit byte offset from
6091     ``SH_HIDDEN_PRIVATE_BASE_VIMID`` to the base of scratch backing memory
6092     being managed by SPI for the queue executing the kernel dispatch. This is
6093     the same value used in the Scratch Segment Buffer V# base address. The
6094     prolog must add the value of Scratch Wavefront Offset to get the
6095     wavefront's byte scratch backing memory offset from
6096     ``SH_HIDDEN_PRIVATE_BASE_VIMID``. Since FLAT_SCRATCH_LO is in units of 256
6097     bytes, the offset must be right shifted by 8 before moving into
6098     FLAT_SCRATCH_LO.
6099  2. The second word of Flat Scratch Init is 32-bit byte size of a single
6100     work-items scratch memory usage. This is directly loaded from the kernel
6101     dispatch packet Private Segment Byte Size and rounded up to a multiple of
6102     DWORD. Having CP load it once avoids loading it at the beginning of every
6103     wavefront. The prolog must move it to FLAT_SCRATCH_LO for use as FLAT
6104     SCRATCH SIZE.
6105
6106GFX9-GFX10
6107  The Flat Scratch Init is the 64-bit address of the base of scratch backing
6108  memory being managed by SPI for the queue executing the kernel dispatch. The
6109  prolog must add the value of Scratch Wavefront Offset and moved to the
6110  FLAT_SCRATCH pair for use as the flat scratch base in flat memory
6111  instructions.
6112
6113.. _amdgpu-amdhsa-kernel-prolog-private-segment-buffer:
6114
6115Private Segment Buffer
6116++++++++++++++++++++++
6117
6118A set of four SGPRs beginning at a four-aligned SGPR index are always selected
6119to serve as the scratch V# for the kernel as follows:
6120
6121  - If it is know during instruction selection that there is stack usage,
6122    SGPR0-3 is reserved for use as the scratch V#.  Stack usage is assumed if
6123    optimisations are disabled (``-O0``), if stack objects already exist (for
6124    locals, etc.), or if there are any function calls.
6125
6126  - Otherwise, four high numbered SGPRs beginning at a four-aligned SGPR index
6127    are reserved for the tentative scratch V#. These will be used if it is
6128    determined that spilling is needed.
6129
6130    - If no use is made of the tentative scratch V#, then it is unreserved
6131      and the register count is determined ignoring it.
6132    - If use is made of the tenatative scratch V#, then its register numbers
6133      are shifted to the first four-aligned SGPR index after the highest one
6134      allocated by the register allocator, and all uses are updated. The
6135      register count includes them in the shifted location.
6136    - In either case, if the processor has the SGPR allocation bug, the
6137      tentative allocation is not shifted or unreserved in order to ensure
6138      the register count is higher to workaround the bug.
6139
6140    .. note::
6141
6142      This approach of using a tentative scratch V# and shifting the register
6143      numbers if used avoids having to perform register allocation a second
6144      time if the tentative V# is eliminated. This is more efficient and
6145      avoids the problem that the second register allocation may perform
6146      spilling which will fail as there is no longer a scratch V#.
6147
6148When the kernel prolog code is being emitted it is known whether the scratch V#
6149described above is actually used. If it is, the prolog code must set it up by
6150copying the Private Segment Buffer to the scratch V# registers and then adding
6151the Private Segment Wavefront Offset to the queue base address in the V#. The
6152result is a V# with a base address pointing to the beginning of the wavefront
6153scratch backing memory.
6154
6155The Private Segment Buffer is always requested, but the Private Segment
6156Wavefront Offset is only requested if it is used (see
6157:ref:`amdgpu-amdhsa-initial-kernel-execution-state`).
6158
6159.. _amdgpu-amdhsa-memory-model:
6160
6161Memory Model
6162~~~~~~~~~~~~
6163
6164This section describes the mapping of LLVM memory model onto AMDGPU machine code
6165(see :ref:`memmodel`).
6166
6167The AMDGPU backend supports the memory synchronization scopes specified in
6168:ref:`amdgpu-memory-scopes`.
6169
6170The code sequences used to implement the memory model are defined in table
6171:ref:`amdgpu-amdhsa-memory-model-code-sequences-gfx6-gfx10-table`.
6172
6173The sequences specify the order of instructions that a single thread must
6174execute. The ``s_waitcnt`` and ``buffer_wbinvl1_vol`` are defined with respect
6175to other memory instructions executed by the same thread. This allows them to be
6176moved earlier or later which can allow them to be combined with other instances
6177of the same instruction, or hoisted/sunk out of loops to improve
6178performance. Only the instructions related to the memory model are given;
6179additional ``s_waitcnt`` instructions are required to ensure registers are
6180defined before being used. These may be able to be combined with the memory
6181model ``s_waitcnt`` instructions as described above.
6182
6183The AMDGPU backend supports the following memory models:
6184
6185  HSA Memory Model [HSA]_
6186    The HSA memory model uses a single happens-before relation for all address
6187    spaces (see :ref:`amdgpu-address-spaces`).
6188  OpenCL Memory Model [OpenCL]_
6189    The OpenCL memory model which has separate happens-before relations for the
6190    global and local address spaces. Only a fence specifying both global and
6191    local address space, and seq_cst instructions join the relationships. Since
6192    the LLVM ``memfence`` instruction does not allow an address space to be
6193    specified the OpenCL fence has to conservatively assume both local and
6194    global address space was specified. However, optimizations can often be
6195    done to eliminate the additional ``s_waitcnt`` instructions when there are
6196    no intervening memory instructions which access the corresponding address
6197    space. The code sequences in the table indicate what can be omitted for the
6198    OpenCL memory. The target triple environment is used to determine if the
6199    source language is OpenCL (see :ref:`amdgpu-opencl`).
6200
6201``ds/flat_load/store/atomic`` instructions to local memory are termed LDS
6202operations.
6203
6204``buffer/global/flat_load/store/atomic`` instructions to global memory are
6205termed vector memory operations.
6206
6207For GFX6-GFX9:
6208
6209* Each agent has multiple shader arrays (SA).
6210* Each SA has multiple compute units (CU).
6211* Each CU has multiple SIMDs that execute wavefronts.
6212* The wavefronts for a single work-group are executed in the same CU but may be
6213  executed by different SIMDs.
6214* Each CU has a single LDS memory shared by the wavefronts of the work-groups
6215  executing on it.
6216* All LDS operations of a CU are performed as wavefront wide operations in a
6217  global order and involve no caching. Completion is reported to a wavefront in
6218  execution order.
6219* The LDS memory has multiple request queues shared by the SIMDs of a
6220  CU. Therefore, the LDS operations performed by different wavefronts of a
6221  work-group can be reordered relative to each other, which can result in
6222  reordering the visibility of vector memory operations with respect to LDS
6223  operations of other wavefronts in the same work-group. A ``s_waitcnt
6224  lgkmcnt(0)`` is required to ensure synchronization between LDS operations and
6225  vector memory operations between wavefronts of a work-group, but not between
6226  operations performed by the same wavefront.
6227* The vector memory operations are performed as wavefront wide operations and
6228  completion is reported to a wavefront in execution order. The exception is
6229  that for GFX7-GFX9 ``flat_load/store/atomic`` instructions can report out of
6230  vector memory order if they access LDS memory, and out of LDS operation order
6231  if they access global memory.
6232* The vector memory operations access a single vector L1 cache shared by all
6233  SIMDs a CU. Therefore, no special action is required for coherence between the
6234  lanes of a single wavefront, or for coherence between wavefronts in the same
6235  work-group. A ``buffer_wbinvl1_vol`` is required for coherence between
6236  wavefronts executing in different work-groups as they may be executing on
6237  different CUs.
6238* The scalar memory operations access a scalar L1 cache shared by all wavefronts
6239  on a group of CUs. The scalar and vector L1 caches are not coherent. However,
6240  scalar operations are used in a restricted way so do not impact the memory
6241  model. See :ref:`amdgpu-address-spaces`.
6242* The vector and scalar memory operations use an L2 cache shared by all CUs on
6243  the same agent.
6244* The L2 cache has independent channels to service disjoint ranges of virtual
6245  addresses.
6246* Each CU has a separate request queue per channel. Therefore, the vector and
6247  scalar memory operations performed by wavefronts executing in different
6248  work-groups (which may be executing on different CUs) of an agent can be
6249  reordered relative to each other. A ``s_waitcnt vmcnt(0)`` is required to
6250  ensure synchronization between vector memory operations of different CUs. It
6251  ensures a previous vector memory operation has completed before executing a
6252  subsequent vector memory or LDS operation and so can be used to meet the
6253  requirements of acquire and release.
6254* The L2 cache can be kept coherent with other agents on some targets, or ranges
6255  of virtual addresses can be set up to bypass it to ensure system coherence.
6256
6257For GFX10:
6258
6259* Each agent has multiple shader arrays (SA).
6260* Each SA has multiple work-group processors (WGP).
6261* Each WGP has multiple compute units (CU).
6262* Each CU has multiple SIMDs that execute wavefronts.
6263* The wavefronts for a single work-group are executed in the same
6264  WGP. In CU wavefront execution mode the wavefronts may be executed by
6265  different SIMDs in the same CU. In WGP wavefront execution mode the
6266  wavefronts may be executed by different SIMDs in different CUs in the same
6267  WGP.
6268* Each WGP has a single LDS memory shared by the wavefronts of the work-groups
6269  executing on it.
6270* All LDS operations of a WGP are performed as wavefront wide operations in a
6271  global order and involve no caching. Completion is reported to a wavefront in
6272  execution order.
6273* The LDS memory has multiple request queues shared by the SIMDs of a
6274  WGP. Therefore, the LDS operations performed by different wavefronts of a
6275  work-group can be reordered relative to each other, which can result in
6276  reordering the visibility of vector memory operations with respect to LDS
6277  operations of other wavefronts in the same work-group. A ``s_waitcnt
6278  lgkmcnt(0)`` is required to ensure synchronization between LDS operations and
6279  vector memory operations between wavefronts of a work-group, but not between
6280  operations performed by the same wavefront.
6281* The vector memory operations are performed as wavefront wide operations.
6282  Completion of load/store/sample operations are reported to a wavefront in
6283  execution order of other load/store/sample operations performed by that
6284  wavefront.
6285* The vector memory operations access a vector L0 cache. There is a single L0
6286  cache per CU. Each SIMD of a CU accesses the same L0 cache. Therefore, no
6287  special action is required for coherence between the lanes of a single
6288  wavefront. However, a ``BUFFER_GL0_INV`` is required for coherence between
6289  wavefronts executing in the same work-group as they may be executing on SIMDs
6290  of different CUs that access different L0s. A ``BUFFER_GL0_INV`` is also
6291  required for coherence between wavefronts executing in different work-groups
6292  as they may be executing on different WGPs.
6293* The scalar memory operations access a scalar L0 cache shared by all wavefronts
6294  on a WGP. The scalar and vector L0 caches are not coherent. However, scalar
6295  operations are used in a restricted way so do not impact the memory model. See
6296  :ref:`amdgpu-address-spaces`.
6297* The vector and scalar memory L0 caches use an L1 cache shared by all WGPs on
6298  the same SA. Therefore, no special action is required for coherence between
6299  the wavefronts of a single work-group. However, a ``BUFFER_GL1_INV`` is
6300  required for coherence between wavefronts executing in different work-groups
6301  as they may be executing on different SAs that access different L1s.
6302* The L1 caches have independent quadrants to service disjoint ranges of virtual
6303  addresses.
6304* Each L0 cache has a separate request queue per L1 quadrant. Therefore, the
6305  vector and scalar memory operations performed by different wavefronts, whether
6306  executing in the same or different work-groups (which may be executing on
6307  different CUs accessing different L0s), can be reordered relative to each
6308  other. A ``s_waitcnt vmcnt(0) & vscnt(0)`` is required to ensure
6309  synchronization between vector memory operations of different wavefronts. It
6310  ensures a previous vector memory operation has completed before executing a
6311  subsequent vector memory or LDS operation and so can be used to meet the
6312  requirements of acquire, release and sequential consistency.
6313* The L1 caches use an L2 cache shared by all SAs on the same agent.
6314* The L2 cache has independent channels to service disjoint ranges of virtual
6315  addresses.
6316* Each L1 quadrant of a single SA accesses a different L2 channel. Each L1
6317  quadrant has a separate request queue per L2 channel. Therefore, the vector
6318  and scalar memory operations performed by wavefronts executing in different
6319  work-groups (which may be executing on different SAs) of an agent can be
6320  reordered relative to each other. A ``s_waitcnt vmcnt(0) & vscnt(0)`` is
6321  required to ensure synchronization between vector memory operations of
6322  different SAs. It ensures a previous vector memory operation has completed
6323  before executing a subsequent vector memory and so can be used to meet the
6324  requirements of acquire, release and sequential consistency.
6325* The L2 cache can be kept coherent with other agents on some targets, or ranges
6326  of virtual addresses can be set up to bypass it to ensure system coherence.
6327
6328Private address space uses ``buffer_load/store`` using the scratch V#
6329(GFX6-GFX8), or ``scratch_load/store`` (GFX9-GFX10). Since only a single thread
6330is accessing the memory, atomic memory orderings are not meaningful and all
6331accesses are treated as non-atomic.
6332
6333Constant address space uses ``buffer/global_load`` instructions (or equivalent
6334scalar memory instructions). Since the constant address space contents do not
6335change during the execution of a kernel dispatch it is not legal to perform
6336stores, and atomic memory orderings are not meaningful and all access are
6337treated as non-atomic.
6338
6339A memory synchronization scope wider than work-group is not meaningful for the
6340group (LDS) address space and is treated as work-group.
6341
6342The memory model does not support the region address space which is treated as
6343non-atomic.
6344
6345Acquire memory ordering is not meaningful on store atomic instructions and is
6346treated as non-atomic.
6347
6348Release memory ordering is not meaningful on load atomic instructions and is
6349treated a non-atomic.
6350
6351Acquire-release memory ordering is not meaningful on load or store atomic
6352instructions and is treated as acquire and release respectively.
6353
6354AMDGPU backend only uses scalar memory operations to access memory that is
6355proven to not change during the execution of the kernel dispatch. This includes
6356constant address space and global address space for program scope const
6357variables. Therefore the kernel machine code does not have to maintain the
6358scalar L1 cache to ensure it is coherent with the vector L1 cache. The scalar
6359and vector L1 caches are invalidated between kernel dispatches by CP since
6360constant address space data may change between kernel dispatch executions. See
6361:ref:`amdgpu-address-spaces`.
6362
6363The one exception is if scalar writes are used to spill SGPR registers. In this
6364case the AMDGPU backend ensures the memory location used to spill is never
6365accessed by vector memory operations at the same time. If scalar writes are used
6366then a ``s_dcache_wb`` is inserted before the ``s_endpgm`` and before a function
6367return since the locations may be used for vector memory instructions by a
6368future wavefront that uses the same scratch area, or a function call that
6369creates a frame at the same address, respectively. There is no need for a
6370``s_dcache_inv`` as all scalar writes are write-before-read in the same thread.
6371
6372For GFX6-GFX9, scratch backing memory (which is used for the private address
6373space) is accessed with MTYPE NC_NV (non-coherent non-volatile). Since the
6374private address space is only accessed by a single thread, and is always
6375write-before-read, there is never a need to invalidate these entries from the L1
6376cache. Hence all cache invalidates are done as ``*_vol`` to only invalidate the
6377volatile cache lines.
6378
6379For GFX10, scratch backing memory (which is used for the private address space)
6380is accessed with MTYPE NC (non-coherent). Since the private address space is
6381only accessed by a single thread, and is always write-before-read, there is
6382never a need to invalidate these entries from the L0 or L1 caches.
6383
6384For GFX10, wavefronts are executed in native mode with in-order reporting of
6385loads and sample instructions. In this mode vmcnt reports completion of load,
6386atomic with return and sample instructions in order, and the vscnt reports the
6387completion of store and atomic without return in order. See ``MEM_ORDERED``
6388field in :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`.
6389
6390In GFX10, wavefronts can be executed in WGP or CU wavefront execution mode:
6391
6392* In WGP wavefront execution mode the wavefronts of a work-group are executed
6393  on the SIMDs of both CUs of the WGP. Therefore, explicit management of the per
6394  CU L0 caches is required for work-group synchronization. Also accesses to L1
6395  at work-group scope need to be explicitly ordered as the accesses from
6396  different CUs are not ordered.
6397* In CU wavefront execution mode the wavefronts of a work-group are executed on
6398  the SIMDs of a single CU of the WGP. Therefore, all global memory access by
6399  the work-group access the same L0 which in turn ensures L1 accesses are
6400  ordered and so do not require explicit management of the caches for
6401  work-group synchronization.
6402
6403See ``WGP_MODE`` field in
6404:ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table` and
6405:ref:`amdgpu-target-features`.
6406
6407On dGPU the kernarg backing memory is accessed as UC (uncached) to avoid needing
6408to invalidate the L2 cache. For GFX6-GFX9, this also causes it to be treated as
6409non-volatile and so is not invalidated by ``*_vol``. On APU it is accessed as CC
6410(cache coherent) and so the L2 cache will be coherent with the CPU and other
6411agents.
6412
6413  .. table:: AMDHSA Memory Model Code Sequences GFX6-GFX10
6414     :name: amdgpu-amdhsa-memory-model-code-sequences-gfx6-gfx10-table
6415
6416     ============ ============ ============== ========== =============================== ==================================
6417     LLVM Instr   LLVM Memory  LLVM Memory    AMDGPU     AMDGPU Machine Code             AMDGPU Machine Code
6418                  Ordering     Sync Scope     Address    GFX6-9                          GFX10
6419                                              Space
6420     ============ ============ ============== ========== =============================== ==================================
6421     **Non-Atomic**
6422     ----------------------------------------------------------------------------------------------------------------------
6423     load         *none*       *none*         - global   - !volatile & !nontemporal      - !volatile & !nontemporal
6424                                              - generic
6425                                              - private    1. buffer/global/flat_load      1. buffer/global/flat_load
6426                                              - constant
6427                                                         - volatile & !nontemporal       - volatile & !nontemporal
6428
6429                                                           1. buffer/global/flat_load      1. buffer/global/flat_load
6430                                                              glc=1                           glc=1 dlc=1
6431
6432                                                         - nontemporal                   - nontemporal
6433
6434                                                           1. buffer/global/flat_load      1. buffer/global/flat_load
6435                                                              glc=1 slc=1                     slc=1
6436
6437     load         *none*       *none*         - local    1. ds_load                      1. ds_load
6438     store        *none*       *none*         - global   - !nontemporal                  - !nontemporal
6439                                              - generic
6440                                              - private    1. buffer/global/flat_store     1. buffer/global/flat_store
6441                                              - constant
6442                                                         - nontemporal                   - nontemporal
6443
6444                                                           1. buffer/global/flat_store      1. buffer/global/flat_store
6445                                                              glc=1 slc=1                      slc=1
6446
6447     store        *none*       *none*         - local    1. ds_store                     1. ds_store
6448     **Unordered Atomic**
6449     ----------------------------------------------------------------------------------------------------------------------
6450     load atomic  unordered    *any*          *any*      *Same as non-atomic*.           *Same as non-atomic*.
6451     store atomic unordered    *any*          *any*      *Same as non-atomic*.           *Same as non-atomic*.
6452     atomicrmw    unordered    *any*          *any*      *Same as monotonic              *Same as monotonic
6453                                                         atomic*.                        atomic*.
6454     **Monotonic Atomic**
6455     ----------------------------------------------------------------------------------------------------------------------
6456     load atomic  monotonic    - singlethread - global   1. buffer/global/flat_load      1. buffer/global/flat_load
6457                               - wavefront    - generic
6458     load atomic  monotonic    - workgroup    - global   1. buffer/global/flat_load      1. buffer/global/flat_load
6459                                              - generic                                     glc=1
6460
6461                                                                                           - If CU wavefront execution mode, omit glc=1.
6462
6463     load atomic  monotonic    - singlethread - local    1. ds_load                      1. ds_load
6464                               - wavefront
6465                               - workgroup
6466     load atomic  monotonic    - agent        - global   1. buffer/global/flat_load      1. buffer/global/flat_load
6467                               - system       - generic     glc=1                           glc=1 dlc=1
6468     store atomic monotonic    - singlethread - global   1. buffer/global/flat_store     1. buffer/global/flat_store
6469                               - wavefront    - generic
6470                               - workgroup
6471                               - agent
6472                               - system
6473     store atomic monotonic    - singlethread - local    1. ds_store                     1. ds_store
6474                               - wavefront
6475                               - workgroup
6476     atomicrmw    monotonic    - singlethread - global   1. buffer/global/flat_atomic    1. buffer/global/flat_atomic
6477                               - wavefront    - generic
6478                               - workgroup
6479                               - agent
6480                               - system
6481     atomicrmw    monotonic    - singlethread - local    1. ds_atomic                    1. ds_atomic
6482                               - wavefront
6483                               - workgroup
6484     **Acquire Atomic**
6485     ----------------------------------------------------------------------------------------------------------------------
6486     load atomic  acquire      - singlethread - global   1. buffer/global/ds/flat_load   1. buffer/global/ds/flat_load
6487                               - wavefront    - local
6488                                              - generic
6489     load atomic  acquire      - workgroup    - global   1. buffer/global/flat_load      1. buffer/global_load glc=1
6490
6491                                                                                           - If CU wavefront execution mode, omit glc=1.
6492
6493                                                                                         2. s_waitcnt vmcnt(0)
6494
6495                                                                                           - If CU wavefront execution mode, omit.
6496                                                                                           - Must happen before
6497                                                                                             the following buffer_gl0_inv
6498                                                                                             and before any following
6499                                                                                             global/generic
6500                                                                                             load/load
6501                                                                                             atomic/store/store
6502                                                                                             atomic/atomicrmw.
6503
6504                                                                                         3. buffer_gl0_inv
6505
6506                                                                                           - If CU wavefront execution mode, omit.
6507                                                                                           - Ensures that
6508                                                                                             following
6509                                                                                             loads will not see
6510                                                                                             stale data.
6511
6512     load atomic  acquire      - workgroup    - local    1. ds_load                      1. ds_load
6513                                                         2. s_waitcnt lgkmcnt(0)         2. s_waitcnt lgkmcnt(0)
6514
6515                                                           - If OpenCL, omit.              - If OpenCL, omit.
6516                                                           - Must happen before            - Must happen before
6517                                                             any following                   the following buffer_gl0_inv
6518                                                             global/generic                  and before any following
6519                                                             load/load                       global/generic load/load
6520                                                             atomic/store/store              atomic/store/store
6521                                                             atomic/atomicrmw.               atomic/atomicrmw.
6522                                                           - Ensures any                   - Ensures any
6523                                                             following global                following global
6524                                                             data read is no                 data read is no
6525                                                             older than the load             older than the load
6526                                                             atomic value being              atomic value being
6527                                                             acquired.                       acquired.
6528
6529                                                                                         3. buffer_gl0_inv
6530
6531                                                                                           - If CU wavefront execution mode, omit.
6532                                                                                           - If OpenCL, omit.
6533                                                                                           - Ensures that
6534                                                                                             following
6535                                                                                             loads will not see
6536                                                                                             stale data.
6537
6538     load atomic  acquire      - workgroup    - generic  1. flat_load                    1. flat_load glc=1
6539
6540                                                                                           - If CU wavefront execution mode, omit glc=1.
6541
6542                                                         2. s_waitcnt lgkmcnt(0)         2. s_waitcnt lgkmcnt(0) &
6543                                                                                            vmcnt(0)
6544
6545                                                                                           - If CU wavefront execution mode, omit vmcnt.
6546                                                           - If OpenCL, omit.              - If OpenCL, omit
6547                                                                                             lgkmcnt(0).
6548                                                           - Must happen before            - Must happen before
6549                                                             any following                   the following
6550                                                             global/generic                  buffer_gl0_inv and any
6551                                                             load/load                       following global/generic
6552                                                             atomic/store/store              load/load
6553                                                             atomic/atomicrmw.               atomic/store/store
6554                                                                                             atomic/atomicrmw.
6555                                                           - Ensures any                   - Ensures any
6556                                                             following global                following global
6557                                                             data read is no                 data read is no
6558                                                             older than the load             older than the load
6559                                                             atomic value being              atomic value being
6560                                                             acquired.                       acquired.
6561
6562                                                                                         3. buffer_gl0_inv
6563
6564                                                                                           - If CU wavefront execution mode, omit.
6565                                                                                           - Ensures that
6566                                                                                             following
6567                                                                                             loads will not see
6568                                                                                             stale data.
6569
6570     load atomic  acquire      - agent        - global   1. buffer/global/flat_load      1. buffer/global_load
6571                               - system                     glc=1                           glc=1 dlc=1
6572                                                         2. s_waitcnt vmcnt(0)           2. s_waitcnt vmcnt(0)
6573
6574                                                           - Must happen before            - Must happen before
6575                                                             following                       following
6576                                                             buffer_wbinvl1_vol.             buffer_gl*_inv.
6577                                                           - Ensures the load              - Ensures the load
6578                                                             has completed                   has completed
6579                                                             before invalidating             before invalidating
6580                                                             the cache.                      the caches.
6581
6582                                                         3. buffer_wbinvl1_vol           3. buffer_gl0_inv;
6583                                                                                            buffer_gl1_inv
6584
6585                                                           - Must happen before            - Must happen before
6586                                                             any following                   any following
6587                                                             global/generic                  global/generic
6588                                                             load/load                       load/load
6589                                                             atomic/atomicrmw.               atomic/atomicrmw.
6590                                                           - Ensures that                  - Ensures that
6591                                                             following                       following
6592                                                             loads will not see              loads will not see
6593                                                             stale global data.              stale global data.
6594
6595     load atomic  acquire      - agent        - generic  1. flat_load glc=1              1. flat_load glc=1 dlc=1
6596                               - system                  2. s_waitcnt vmcnt(0) &         2. s_waitcnt vmcnt(0) &
6597                                                            lgkmcnt(0)                      lgkmcnt(0)
6598
6599                                                           - If OpenCL omit                - If OpenCL omit
6600                                                             lgkmcnt(0).                     lgkmcnt(0).
6601                                                           - Must happen before            - Must happen before
6602                                                             following                       following
6603                                                             buffer_wbinvl1_vol.             buffer_gl*_invl.
6604                                                           - Ensures the flat_load         - Ensures the flat_load
6605                                                             has completed                   has completed
6606                                                             before invalidating             before invalidating
6607                                                             the cache.                      the caches.
6608
6609                                                         3. buffer_wbinvl1_vol           3. buffer_gl0_inv;
6610                                                                                            buffer_gl1_inv
6611
6612                                                           - Must happen before            - Must happen before
6613                                                             any following                   any following
6614                                                             global/generic                  global/generic
6615                                                             load/load                       load/load
6616                                                             atomic/atomicrmw.               atomic/atomicrmw.
6617                                                           - Ensures that                  - Ensures that
6618                                                             following loads                 following loads
6619                                                             will not see stale              will not see stale
6620                                                             global data.                    global data.
6621
6622     atomicrmw    acquire      - singlethread - global   1. buffer/global/ds/flat_atomic 1. buffer/global/ds/flat_atomic
6623                               - wavefront    - local
6624                                              - generic
6625     atomicrmw    acquire      - workgroup    - global   1. buffer/global/flat_atomic    1. buffer/global_atomic
6626                                                                                         2. s_waitcnt vm/vscnt(0)
6627
6628                                                                                           - If CU wavefront execution mode, omit.
6629                                                                                           - Use vmcnt if atomic with
6630                                                                                             return and vscnt if atomic
6631                                                                                             with no-return.
6632                                                                                           - Must happen before
6633                                                                                             the following buffer_gl0_inv
6634                                                                                             and before any following
6635                                                                                             global/generic
6636                                                                                             load/load
6637                                                                                             atomic/store/store
6638                                                                                             atomic/atomicrmw.
6639
6640                                                                                         3. buffer_gl0_inv
6641
6642                                                                                           - If CU wavefront execution mode, omit.
6643                                                                                           - Ensures that
6644                                                                                             following
6645                                                                                             loads will not see
6646                                                                                             stale data.
6647
6648     atomicrmw    acquire      - workgroup    - local    1. ds_atomic                    1. ds_atomic
6649                                                         2. waitcnt lgkmcnt(0)           2. waitcnt lgkmcnt(0)
6650
6651                                                           - If OpenCL, omit.              - If OpenCL, omit.
6652                                                           - Must happen before            - Must happen before
6653                                                             any following                   the following
6654                                                             global/generic                  buffer_gl0_inv.
6655                                                             load/load
6656                                                             atomic/store/store
6657                                                             atomic/atomicrmw.
6658                                                           - Ensures any                   - Ensures any
6659                                                             following global                following global
6660                                                             data read is no                 data read is no
6661                                                             older than the                  older than the
6662                                                             atomicrmw value                 atomicrmw value
6663                                                             being acquired.                 being acquired.
6664
6665                                                                                         3. buffer_gl0_inv
6666
6667                                                                                           - If OpenCL omit.
6668                                                                                           - Ensures that
6669                                                                                             following
6670                                                                                             loads will not see
6671                                                                                             stale data.
6672
6673     atomicrmw    acquire      - workgroup    - generic  1. flat_atomic                  1. flat_atomic
6674                                                         2. waitcnt lgkmcnt(0)           2. waitcnt lgkmcnt(0) &
6675                                                                                            vm/vscnt(0)
6676
6677                                                                                           - If CU wavefront execution mode, omit vm/vscnt.
6678                                                           - If OpenCL, omit.              - If OpenCL, omit
6679                                                                                             waitcnt lgkmcnt(0)..
6680                                                                                           - Use vmcnt if atomic with
6681                                                                                             return and vscnt if atomic
6682                                                                                             with no-return.
6683                                                                                             waitcnt lgkmcnt(0).
6684                                                           - Must happen before            - Must happen before
6685                                                             any following                   the following
6686                                                             global/generic                  buffer_gl0_inv.
6687                                                             load/load
6688                                                             atomic/store/store
6689                                                             atomic/atomicrmw.
6690                                                           - Ensures any                   - Ensures any
6691                                                             following global                following global
6692                                                             data read is no                 data read is no
6693                                                             older than the                  older than the
6694                                                             atomicrmw value                 atomicrmw value
6695                                                             being acquired.                 being acquired.
6696
6697                                                                                         3. buffer_gl0_inv
6698
6699                                                                                           - If CU wavefront execution mode, omit.
6700                                                                                           - Ensures that
6701                                                                                             following
6702                                                                                             loads will not see
6703                                                                                             stale data.
6704
6705     atomicrmw    acquire      - agent        - global   1. buffer/global/flat_atomic    1. buffer/global_atomic
6706                               - system                  2. s_waitcnt vmcnt(0)           2. s_waitcnt vm/vscnt(0)
6707
6708                                                                                           - Use vmcnt if atomic with
6709                                                                                             return and vscnt if atomic
6710                                                                                             with no-return.
6711                                                                                             waitcnt lgkmcnt(0).
6712                                                           - Must happen before            - Must happen before
6713                                                             following                       following
6714                                                             buffer_wbinvl1_vol.             buffer_gl*_inv.
6715                                                           - Ensures the                   - Ensures the
6716                                                             atomicrmw has                   atomicrmw has
6717                                                             completed before                completed before
6718                                                             invalidating the                invalidating the
6719                                                             cache.                          caches.
6720
6721                                                         3. buffer_wbinvl1_vol           3. buffer_gl0_inv;
6722                                                                                            buffer_gl1_inv
6723
6724                                                           - Must happen before            - Must happen before
6725                                                             any following                   any following
6726                                                             global/generic                  global/generic
6727                                                             load/load                       load/load
6728                                                             atomic/atomicrmw.               atomic/atomicrmw.
6729                                                           - Ensures that                  - Ensures that
6730                                                             following loads                 following loads
6731                                                             will not see stale              will not see stale
6732                                                             global data.                    global data.
6733
6734     atomicrmw    acquire      - agent        - generic  1. flat_atomic                  1. flat_atomic
6735                               - system                  2. s_waitcnt vmcnt(0) &         2. s_waitcnt vm/vscnt(0) &
6736                                                            lgkmcnt(0)                      lgkmcnt(0)
6737
6738                                                           - If OpenCL, omit               - If OpenCL, omit
6739                                                             lgkmcnt(0).                     lgkmcnt(0).
6740                                                                                           - Use vmcnt if atomic with
6741                                                                                             return and vscnt if atomic
6742                                                                                             with no-return.
6743                                                           - Must happen before            - Must happen before
6744                                                             following                       following
6745                                                             buffer_wbinvl1_vol.             buffer_gl*_inv.
6746                                                           - Ensures the                   - Ensures the
6747                                                             atomicrmw has                   atomicrmw has
6748                                                             completed before                completed before
6749                                                             invalidating the                invalidating the
6750                                                             cache.                          caches.
6751
6752                                                         3. buffer_wbinvl1_vol           3. buffer_gl0_inv;
6753                                                                                            buffer_gl1_inv
6754
6755                                                           - Must happen before            - Must happen before
6756                                                             any following                   any following
6757                                                             global/generic                  global/generic
6758                                                             load/load                       load/load
6759                                                             atomic/atomicrmw.               atomic/atomicrmw.
6760                                                           - Ensures that                  - Ensures that
6761                                                             following loads                 following loads
6762                                                             will not see stale              will not see stale
6763                                                             global data.                    global data.
6764
6765     fence        acquire      - singlethread *none*     *none*                          *none*
6766                               - wavefront
6767     fence        acquire      - workgroup    *none*     1. s_waitcnt lgkmcnt(0)         1. s_waitcnt lgkmcnt(0) &
6768                                                                                            vmcnt(0) & vscnt(0)
6769
6770                                                                                           - If CU wavefront execution mode, omit vmcnt and
6771                                                                                             vscnt.
6772                                                           - If OpenCL and                 - If OpenCL and
6773                                                             address space is                address space is
6774                                                             not generic, omit.              not generic, omit
6775                                                                                             lgkmcnt(0).
6776                                                                                           - If OpenCL and
6777                                                                                             address space is
6778                                                                                             local, omit
6779                                                                                             vmcnt(0) and vscnt(0).
6780                                                           - However, since LLVM           - However, since LLVM
6781                                                             currently has no                currently has no
6782                                                             address space on                address space on
6783                                                             the fence need to               the fence need to
6784                                                             conservatively                  conservatively
6785                                                             always generate. If             always generate. If
6786                                                             fence had an                    fence had an
6787                                                             address space then              address space then
6788                                                             set to address                  set to address
6789                                                             space of OpenCL                 space of OpenCL
6790                                                             fence flag, or to               fence flag, or to
6791                                                             generic if both                 generic if both
6792                                                             local and global                local and global
6793                                                             flags are                       flags are
6794                                                             specified.                      specified.
6795                                                           - Must happen after
6796                                                             any preceding
6797                                                             local/generic load
6798                                                             atomic/atomicrmw
6799                                                             with an equal or
6800                                                             wider sync scope
6801                                                             and memory ordering
6802                                                             stronger than
6803                                                             unordered (this is
6804                                                             termed the
6805                                                             fence-paired-atomic).
6806                                                           - Must happen before
6807                                                             any following
6808                                                             global/generic
6809                                                             load/load
6810                                                             atomic/store/store
6811                                                             atomic/atomicrmw.
6812                                                           - Ensures any
6813                                                             following global
6814                                                             data read is no
6815                                                             older than the
6816                                                             value read by the
6817                                                             fence-paired-atomic.
6818                                                                                           - Could be split into
6819                                                                                             separate s_waitcnt
6820                                                                                             vmcnt(0), s_waitcnt
6821                                                                                             vscnt(0) and s_waitcnt
6822                                                                                             lgkmcnt(0) to allow
6823                                                                                             them to be
6824                                                                                             independently moved
6825                                                                                             according to the
6826                                                                                             following rules.
6827                                                                                           - s_waitcnt vmcnt(0)
6828                                                                                             must happen after
6829                                                                                             any preceding
6830                                                                                             global/generic load
6831                                                                                             atomic/
6832                                                                                             atomicrmw-with-return-value
6833                                                                                             with an equal or
6834                                                                                             wider sync scope
6835                                                                                             and memory ordering
6836                                                                                             stronger than
6837                                                                                             unordered (this is
6838                                                                                             termed the
6839                                                                                             fence-paired-atomic).
6840                                                                                           - s_waitcnt vscnt(0)
6841                                                                                             must happen after
6842                                                                                             any preceding
6843                                                                                             global/generic
6844                                                                                             atomicrmw-no-return-value
6845                                                                                             with an equal or
6846                                                                                             wider sync scope
6847                                                                                             and memory ordering
6848                                                                                             stronger than
6849                                                                                             unordered (this is
6850                                                                                             termed the
6851                                                                                             fence-paired-atomic).
6852                                                                                           - s_waitcnt lgkmcnt(0)
6853                                                                                             must happen after
6854                                                                                             any preceding
6855                                                                                             local/generic load
6856                                                                                             atomic/atomicrmw
6857                                                                                             with an equal or
6858                                                                                             wider sync scope
6859                                                                                             and memory ordering
6860                                                                                             stronger than
6861                                                                                             unordered (this is
6862                                                                                             termed the
6863                                                                                             fence-paired-atomic).
6864                                                                                           - Must happen before
6865                                                                                             the following
6866                                                                                             buffer_gl0_inv.
6867                                                                                           - Ensures that the
6868                                                                                             fence-paired atomic
6869                                                                                             has completed
6870                                                                                             before invalidating
6871                                                                                             the
6872                                                                                             cache. Therefore
6873                                                                                             any following
6874                                                                                             locations read must
6875                                                                                             be no older than
6876                                                                                             the value read by
6877                                                                                             the
6878                                                                                             fence-paired-atomic.
6879
6880                                                                                         3. buffer_gl0_inv
6881
6882                                                                                           - If CU wavefront execution mode, omit.
6883                                                                                           - Ensures that
6884                                                                                             following
6885                                                                                             loads will not see
6886                                                                                             stale data.
6887
6888     fence        acquire      - agent        *none*     1. s_waitcnt lgkmcnt(0) &       1. s_waitcnt lgkmcnt(0) &
6889                               - system                     vmcnt(0)                        vmcnt(0) & vscnt(0)
6890
6891                                                           - If OpenCL and                 - If OpenCL and
6892                                                             address space is                address space is
6893                                                             not generic, omit               not generic, omit
6894                                                             lgkmcnt(0).                     lgkmcnt(0).
6895                                                                                           - If OpenCL and
6896                                                                                             address space is
6897                                                                                             local, omit
6898                                                                                             vmcnt(0) and vscnt(0).
6899                                                           - However, since LLVM           - However, since LLVM
6900                                                             currently has no                currently has no
6901                                                             address space on                address space on
6902                                                             the fence need to               the fence need to
6903                                                             conservatively                  conservatively
6904                                                             always generate                 always generate
6905                                                             (see comment for                (see comment for
6906                                                             previous fence).                previous fence).
6907                                                           - Could be split into
6908                                                             separate s_waitcnt
6909                                                             vmcnt(0) and
6910                                                             s_waitcnt
6911                                                             lgkmcnt(0) to allow
6912                                                             them to be
6913                                                             independently moved
6914                                                             according to the
6915                                                             following rules.
6916                                                           - s_waitcnt vmcnt(0)
6917                                                             must happen after
6918                                                             any preceding
6919                                                             global/generic load
6920                                                             atomic/atomicrmw
6921                                                             with an equal or
6922                                                             wider sync scope
6923                                                             and memory ordering
6924                                                             stronger than
6925                                                             unordered (this is
6926                                                             termed the
6927                                                             fence-paired-atomic).
6928                                                           - s_waitcnt lgkmcnt(0)
6929                                                             must happen after
6930                                                             any preceding
6931                                                             local/generic load
6932                                                             atomic/atomicrmw
6933                                                             with an equal or
6934                                                             wider sync scope
6935                                                             and memory ordering
6936                                                             stronger than
6937                                                             unordered (this is
6938                                                             termed the
6939                                                             fence-paired-atomic).
6940                                                           - Must happen before
6941                                                             the following
6942                                                             buffer_wbinvl1_vol.
6943                                                           - Ensures that the
6944                                                             fence-paired atomic
6945                                                             has completed
6946                                                             before invalidating
6947                                                             the
6948                                                             cache. Therefore
6949                                                             any following
6950                                                             locations read must
6951                                                             be no older than
6952                                                             the value read by
6953                                                             the
6954                                                             fence-paired-atomic.
6955                                                                                           - Could be split into
6956                                                                                             separate s_waitcnt
6957                                                                                             vmcnt(0), s_waitcnt
6958                                                                                             vscnt(0) and s_waitcnt
6959                                                                                             lgkmcnt(0) to allow
6960                                                                                             them to be
6961                                                                                             independently moved
6962                                                                                             according to the
6963                                                                                             following rules.
6964                                                                                           - s_waitcnt vmcnt(0)
6965                                                                                             must happen after
6966                                                                                             any preceding
6967                                                                                             global/generic load
6968                                                                                             atomic/
6969                                                                                             atomicrmw-with-return-value
6970                                                                                             with an equal or
6971                                                                                             wider sync scope
6972                                                                                             and memory ordering
6973                                                                                             stronger than
6974                                                                                             unordered (this is
6975                                                                                             termed the
6976                                                                                             fence-paired-atomic).
6977                                                                                           - s_waitcnt vscnt(0)
6978                                                                                             must happen after
6979                                                                                             any preceding
6980                                                                                             global/generic
6981                                                                                             atomicrmw-no-return-value
6982                                                                                             with an equal or
6983                                                                                             wider sync scope
6984                                                                                             and memory ordering
6985                                                                                             stronger than
6986                                                                                             unordered (this is
6987                                                                                             termed the
6988                                                                                             fence-paired-atomic).
6989                                                                                           - s_waitcnt lgkmcnt(0)
6990                                                                                             must happen after
6991                                                                                             any preceding
6992                                                                                             local/generic load
6993                                                                                             atomic/atomicrmw
6994                                                                                             with an equal or
6995                                                                                             wider sync scope
6996                                                                                             and memory ordering
6997                                                                                             stronger than
6998                                                                                             unordered (this is
6999                                                                                             termed the
7000                                                                                             fence-paired-atomic).
7001                                                                                           - Must happen before
7002                                                                                             the following
7003                                                                                             buffer_gl*_inv.
7004                                                                                           - Ensures that the
7005                                                                                             fence-paired atomic
7006                                                                                             has completed
7007                                                                                             before invalidating
7008                                                                                             the
7009                                                                                             caches. Therefore
7010                                                                                             any following
7011                                                                                             locations read must
7012                                                                                             be no older than
7013                                                                                             the value read by
7014                                                                                             the
7015                                                                                             fence-paired-atomic.
7016
7017                                                         2. buffer_wbinvl1_vol           2. buffer_gl0_inv;
7018                                                                                            buffer_gl1_inv
7019
7020                                                           - Must happen before any        - Must happen before any
7021                                                             following global/generic        following global/generic
7022                                                             load/load                       load/load
7023                                                             atomic/store/store              atomic/store/store
7024                                                             atomic/atomicrmw.               atomic/atomicrmw.
7025                                                           - Ensures that                  - Ensures that
7026                                                             following loads                 following loads
7027                                                             will not see stale              will not see stale
7028                                                             global data.                    global data.
7029
7030     **Release Atomic**
7031     ----------------------------------------------------------------------------------------------------------------------
7032     store atomic release      - singlethread - global   1. buffer/global/ds/flat_store  1. buffer/global/ds/flat_store
7033                               - wavefront    - local
7034                                              - generic
7035     store atomic release      - workgroup    - global   1. s_waitcnt lgkmcnt(0)         1. s_waitcnt lgkmcnt(0) &
7036                                                                                            vmcnt(0) & vscnt(0)
7037
7038                                                                                           - If CU wavefront execution mode, omit vmcnt and
7039                                                                                             vscnt.
7040                                                           - If OpenCL, omit.              - If OpenCL, omit
7041                                                                                             lgkmcnt(0).
7042                                                           - Must happen after
7043                                                             any preceding
7044                                                             local/generic
7045                                                             load/store/load
7046                                                             atomic/store
7047                                                             atomic/atomicrmw.
7048                                                                                           - Could be split into
7049                                                                                             separate s_waitcnt
7050                                                                                             vmcnt(0), s_waitcnt
7051                                                                                             vscnt(0) and s_waitcnt
7052                                                                                             lgkmcnt(0) to allow
7053                                                                                             them to be
7054                                                                                             independently moved
7055                                                                                             according to the
7056                                                                                             following rules.
7057                                                                                           - s_waitcnt vmcnt(0)
7058                                                                                             must happen after
7059                                                                                             any preceding
7060                                                                                             global/generic load/load
7061                                                                                             atomic/
7062                                                                                             atomicrmw-with-return-value.
7063                                                                                           - s_waitcnt vscnt(0)
7064                                                                                             must happen after
7065                                                                                             any preceding
7066                                                                                             global/generic
7067                                                                                             store/store
7068                                                                                             atomic/
7069                                                                                             atomicrmw-no-return-value.
7070                                                                                           - s_waitcnt lgkmcnt(0)
7071                                                                                             must happen after
7072                                                                                             any preceding
7073                                                                                             local/generic
7074                                                                                             load/store/load
7075                                                                                             atomic/store
7076                                                                                             atomic/atomicrmw.
7077                                                           - Must happen before            - Must happen before
7078                                                             the following                   the following
7079                                                             store.                          store.
7080                                                           - Ensures that all              - Ensures that all
7081                                                             memory operations               memory operations
7082                                                             to local have                   have
7083                                                             completed before                completed before
7084                                                             performing the                  performing the
7085                                                             store that is being             store that is being
7086                                                             released.                       released.
7087
7088                                                         2. buffer/global/flat_store     2. buffer/global_store
7089     store atomic release      - workgroup    - local                                    1. waitcnt vmcnt(0) & vscnt(0)
7090
7091                                                                                           - If CU wavefront execution mode, omit.
7092                                                                                           - If OpenCL, omit.
7093                                                                                           - Could be split into
7094                                                                                             separate s_waitcnt
7095                                                                                             vmcnt(0) and s_waitcnt
7096                                                                                             vscnt(0) to allow
7097                                                                                             them to be
7098                                                                                             independently moved
7099                                                                                             according to the
7100                                                                                             following rules.
7101                                                                                           - s_waitcnt vmcnt(0)
7102                                                                                             must happen after
7103                                                                                             any preceding
7104                                                                                             global/generic load/load
7105                                                                                             atomic/
7106                                                                                             atomicrmw-with-return-value.
7107                                                                                           - s_waitcnt vscnt(0)
7108                                                                                             must happen after
7109                                                                                             any preceding
7110                                                                                             global/generic
7111                                                                                             store/store atomic/
7112                                                                                             atomicrmw-no-return-value.
7113                                                                                           - Must happen before
7114                                                                                             the following
7115                                                                                             store.
7116                                                                                           - Ensures that all
7117                                                                                             global memory
7118                                                                                             operations have
7119                                                                                             completed before
7120                                                                                             performing the
7121                                                                                             store that is being
7122                                                                                             released.
7123
7124                                                         1. ds_store                     2. ds_store
7125     store atomic release      - workgroup    - generic  1. s_waitcnt lgkmcnt(0)         1. s_waitcnt lgkmcnt(0) &
7126                                                                                            vmcnt(0) & vscnt(0)
7127
7128                                                                                           - If CU wavefront execution mode, omit vmcnt and
7129                                                                                             vscnt.
7130                                                           - If OpenCL, omit.              - If OpenCL, omit
7131                                                                                             lgkmcnt(0).
7132                                                           - Must happen after
7133                                                             any preceding
7134                                                             local/generic
7135                                                             load/store/load
7136                                                             atomic/store
7137                                                             atomic/atomicrmw.
7138                                                                                           - Could be split into
7139                                                                                             separate s_waitcnt
7140                                                                                             vmcnt(0), s_waitcnt
7141                                                                                             vscnt(0) and s_waitcnt
7142                                                                                             lgkmcnt(0) to allow
7143                                                                                             them to be
7144                                                                                             independently moved
7145                                                                                             according to the
7146                                                                                             following rules.
7147                                                                                           - s_waitcnt vmcnt(0)
7148                                                                                             must happen after
7149                                                                                             any preceding
7150                                                                                             global/generic load/load
7151                                                                                             atomic/
7152                                                                                             atomicrmw-with-return-value.
7153                                                                                           - s_waitcnt vscnt(0)
7154                                                                                             must happen after
7155                                                                                             any preceding
7156                                                                                             global/generic
7157                                                                                             store/store
7158                                                                                             atomic/
7159                                                                                             atomicrmw-no-return-value.
7160                                                                                           - s_waitcnt lgkmcnt(0)
7161                                                                                             must happen after
7162                                                                                             any preceding
7163                                                                                             local/generic load/store/load
7164                                                                                             atomic/store atomic/atomicrmw.
7165                                                           - Must happen before            - Must happen before
7166                                                             the following                   the following
7167                                                             store.                          store.
7168                                                           - Ensures that all              - Ensures that all
7169                                                             memory operations               memory operations
7170                                                             to local have                   have
7171                                                             completed before                completed before
7172                                                             performing the                  performing the
7173                                                             store that is being             store that is being
7174                                                             released.                       released.
7175
7176                                                         2. flat_store                   2. flat_store
7177     store atomic release      - agent        - global   1. s_waitcnt lgkmcnt(0) &         1. s_waitcnt lgkmcnt(0) &
7178                               - system       - generic     vmcnt(0)                          vmcnt(0) & vscnt(0)
7179
7180                                                           - If OpenCL, omit               - If OpenCL, omit
7181                                                             lgkmcnt(0).                     lgkmcnt(0).
7182                                                           - Could be split into           - Could be split into
7183                                                             separate s_waitcnt              separate s_waitcnt
7184                                                             vmcnt(0) and                    vmcnt(0), s_waitcnt vscnt(0)
7185                                                             s_waitcnt                       and s_waitcnt
7186                                                             lgkmcnt(0) to allow             lgkmcnt(0) to allow
7187                                                             them to be                      them to be
7188                                                             independently moved             independently moved
7189                                                             according to the                according to the
7190                                                             following rules.                following rules.
7191                                                           - s_waitcnt vmcnt(0)            - s_waitcnt vmcnt(0)
7192                                                             must happen after               must happen after
7193                                                             any preceding                   any preceding
7194                                                             global/generic                  global/generic
7195                                                             load/store/load                 load/load
7196                                                             atomic/store                    atomic/
7197                                                             atomic/atomicrmw.               atomicrmw-with-return-value.
7198                                                                                           - s_waitcnt vscnt(0)
7199                                                                                             must happen after
7200                                                                                             any preceding
7201                                                                                             global/generic
7202                                                                                             store/store atomic/
7203                                                                                             atomicrmw-no-return-value.
7204                                                           - s_waitcnt lgkmcnt(0)          - s_waitcnt lgkmcnt(0)
7205                                                             must happen after               must happen after
7206                                                             any preceding                   any preceding
7207                                                             local/generic                   local/generic
7208                                                             load/store/load                 load/store/load
7209                                                             atomic/store                    atomic/store
7210                                                             atomic/atomicrmw.               atomic/atomicrmw.
7211                                                           - Must happen before            - Must happen before
7212                                                             the following                   the following
7213                                                             store.                          store.
7214                                                           - Ensures that all              - Ensures that all
7215                                                             memory operations               memory operations
7216                                                             to memory have                  to memory have
7217                                                             completed before                completed before
7218                                                             performing the                  performing the
7219                                                             store that is being             store that is being
7220                                                             released.                       released.
7221
7222                                                         2. buffer/global/ds/flat_store  2. buffer/global/ds/flat_store
7223     atomicrmw    release      - singlethread - global   1. buffer/global/ds/flat_atomic 1. buffer/global/ds/flat_atomic
7224                               - wavefront    - local
7225                                              - generic
7226     atomicrmw    release      - workgroup    - global   1. s_waitcnt lgkmcnt(0)         1. s_waitcnt lgkmcnt(0) &
7227                                                                                            vmcnt(0) & vscnt(0)
7228
7229                                                                                           - If CU wavefront execution mode, omit vmcnt and
7230                                                                                             vscnt.
7231                                                           - If OpenCL, omit.
7232
7233                                                           - Must happen after
7234                                                             any preceding
7235                                                             local/generic
7236                                                             load/store/load
7237                                                             atomic/store
7238                                                             atomic/atomicrmw.
7239                                                                                           - Could be split into
7240                                                                                             separate s_waitcnt
7241                                                                                             vmcnt(0), s_waitcnt
7242                                                                                             vscnt(0) and s_waitcnt
7243                                                                                             lgkmcnt(0) to allow
7244                                                                                             them to be
7245                                                                                             independently moved
7246                                                                                             according to the
7247                                                                                             following rules.
7248                                                                                           - s_waitcnt vmcnt(0)
7249                                                                                             must happen after
7250                                                                                             any preceding
7251                                                                                             global/generic load/load
7252                                                                                             atomic/
7253                                                                                             atomicrmw-with-return-value.
7254                                                                                           - s_waitcnt vscnt(0)
7255                                                                                             must happen after
7256                                                                                             any preceding
7257                                                                                             global/generic
7258                                                                                             store/store
7259                                                                                             atomic/
7260                                                                                             atomicrmw-no-return-value.
7261                                                                                           - s_waitcnt lgkmcnt(0)
7262                                                                                             must happen after
7263                                                                                             any preceding
7264                                                                                             local/generic
7265                                                                                             load/store/load
7266                                                                                             atomic/store
7267                                                                                             atomic/atomicrmw.
7268                                                           - Must happen before            - Must happen before
7269                                                             the following                   the following
7270                                                             atomicrmw.                      atomicrmw.
7271                                                           - Ensures that all              - Ensures that all
7272                                                             memory operations               memory operations
7273                                                             to local have                   have
7274                                                             completed before                completed before
7275                                                             performing the                  performing the
7276                                                             atomicrmw that is               atomicrmw that is
7277                                                             being released.                 being released.
7278
7279                                                         2. buffer/global/flat_atomic    2. buffer/global_atomic
7280     atomicrmw    release      - workgroup    - local                                    1. waitcnt vmcnt(0) & vscnt(0)
7281
7282                                                                                           - If CU wavefront execution mode, omit.
7283                                                                                           - If OpenCL, omit.
7284                                                                                           - Could be split into
7285                                                                                             separate s_waitcnt
7286                                                                                             vmcnt(0) and s_waitcnt
7287                                                                                             vscnt(0) to allow
7288                                                                                             them to be
7289                                                                                             independently moved
7290                                                                                             according to the
7291                                                                                             following rules.
7292                                                                                           - s_waitcnt vmcnt(0)
7293                                                                                             must happen after
7294                                                                                             any preceding
7295                                                                                             global/generic load/load
7296                                                                                             atomic/
7297                                                                                             atomicrmw-with-return-value.
7298                                                                                           - s_waitcnt vscnt(0)
7299                                                                                             must happen after
7300                                                                                             any preceding
7301                                                                                             global/generic
7302                                                                                             store/store atomic/
7303                                                                                             atomicrmw-no-return-value.
7304                                                                                           - Must happen before
7305                                                                                             the following
7306                                                                                             store.
7307                                                                                           - Ensures that all
7308                                                                                             global memory
7309                                                                                             operations have
7310                                                                                             completed before
7311                                                                                             performing the
7312                                                                                             store that is being
7313                                                                                             released.
7314
7315                                                         1. ds_atomic                    2. ds_atomic
7316     atomicrmw    release      - workgroup    - generic  1. s_waitcnt lgkmcnt(0)         1. s_waitcnt lgkmcnt(0) &
7317                                                                                            vmcnt(0) & vscnt(0)
7318
7319                                                                                           - If CU wavefront execution mode, omit vmcnt and
7320                                                                                             vscnt.
7321                                                           - If OpenCL, omit.              - If OpenCL, omit
7322                                                                                             waitcnt lgkmcnt(0).
7323                                                           - Must happen after
7324                                                             any preceding
7325                                                             local/generic
7326                                                             load/store/load
7327                                                             atomic/store
7328                                                             atomic/atomicrmw.
7329                                                                                           - Could be split into
7330                                                                                             separate s_waitcnt
7331                                                                                             vmcnt(0), s_waitcnt
7332                                                                                             vscnt(0) and s_waitcnt
7333                                                                                             lgkmcnt(0) to allow
7334                                                                                             them to be
7335                                                                                             independently moved
7336                                                                                             according to the
7337                                                                                             following rules.
7338                                                                                           - s_waitcnt vmcnt(0)
7339                                                                                             must happen after
7340                                                                                             any preceding
7341                                                                                             global/generic load/load
7342                                                                                             atomic/
7343                                                                                             atomicrmw-with-return-value.
7344                                                                                           - s_waitcnt vscnt(0)
7345                                                                                             must happen after
7346                                                                                             any preceding
7347                                                                                             global/generic
7348                                                                                             store/store
7349                                                                                             atomic/
7350                                                                                             atomicrmw-no-return-value.
7351                                                                                           - s_waitcnt lgkmcnt(0)
7352                                                                                             must happen after
7353                                                                                             any preceding
7354                                                                                             local/generic load/store/load
7355                                                                                             atomic/store atomic/atomicrmw.
7356                                                           - Must happen before            - Must happen before
7357                                                             the following                   the following
7358                                                             atomicrmw.                      atomicrmw.
7359                                                           - Ensures that all              - Ensures that all
7360                                                             memory operations               memory operations
7361                                                             to local have                   have
7362                                                             completed before                completed before
7363                                                             performing the                  performing the
7364                                                             atomicrmw that is               atomicrmw that is
7365                                                             being released.                 being released.
7366
7367                                                         2. flat_atomic                  2. flat_atomic
7368     atomicrmw    release      - agent        - global   1. s_waitcnt lgkmcnt(0) &       1. s_waitcnt lkkmcnt(0) &
7369                               - system       - generic     vmcnt(0)                         vmcnt(0) & vscnt(0)
7370
7371                                                           - If OpenCL, omit               - If OpenCL, omit
7372                                                             lgkmcnt(0).                     lgkmcnt(0).
7373                                                           - Could be split into           - Could be split into
7374                                                             separate s_waitcnt              separate s_waitcnt
7375                                                             vmcnt(0) and                    vmcnt(0), s_waitcnt
7376                                                             s_waitcnt                       vscnt(0) and s_waitcnt
7377                                                             lgkmcnt(0) to allow             lgkmcnt(0) to allow
7378                                                             them to be                      them to be
7379                                                             independently moved             independently moved
7380                                                             according to the                according to the
7381                                                             following rules.                following rules.
7382                                                           - s_waitcnt vmcnt(0)            - s_waitcnt vmcnt(0)
7383                                                             must happen after               must happen after
7384                                                             any preceding                   any preceding
7385                                                             global/generic                  global/generic
7386                                                             load/store/load                 load/load atomic/
7387                                                             atomic/store                    atomicrmw-with-return-value.
7388                                                             atomic/atomicrmw.
7389                                                                                           - s_waitcnt vscnt(0)
7390                                                                                             must happen after
7391                                                                                             any preceding
7392                                                                                             global/generic
7393                                                                                             store/store atomic/
7394                                                                                             atomicrmw-no-return-value.
7395                                                           - s_waitcnt lgkmcnt(0)          - s_waitcnt lgkmcnt(0)
7396                                                             must happen after               must happen after
7397                                                             any preceding                   any preceding
7398                                                             local/generic                   local/generic
7399                                                             load/store/load                 load/store/load
7400                                                             atomic/store                    atomic/store
7401                                                             atomic/atomicrmw.               atomic/atomicrmw.
7402                                                           - Must happen before            - Must happen before
7403                                                             the following                   the following
7404                                                             atomicrmw.                      atomicrmw.
7405                                                           - Ensures that all              - Ensures that all
7406                                                             memory operations               memory operations
7407                                                             to global and local             to global and local
7408                                                             have completed                  have completed
7409                                                             before performing               before performing
7410                                                             the atomicrmw that              the atomicrmw that
7411                                                             is being released.              is being released.
7412
7413                                                         2. buffer/global/ds/flat_atomic 2. buffer/global/ds/flat_atomic
7414     fence        release      - singlethread *none*     *none*                          *none*
7415                               - wavefront
7416     fence        release      - workgroup    *none*     1. s_waitcnt lgkmcnt(0)         1. s_waitcnt lgkmcnt(0) &
7417                                                                                            vmcnt(0) & vscnt(0)
7418
7419                                                                                           - If CU wavefront execution mode, omit vmcnt and
7420                                                                                             vscnt.
7421                                                           - If OpenCL and                 - If OpenCL and
7422                                                             address space is                address space is
7423                                                             not generic, omit.              not generic, omit
7424                                                                                             lgkmcnt(0).
7425                                                                                           - If OpenCL and
7426                                                                                             address space is
7427                                                                                             local, omit
7428                                                                                             vmcnt(0) and vscnt(0).
7429                                                           - However, since LLVM           - However, since LLVM
7430                                                             currently has no                currently has no
7431                                                             address space on                address space on
7432                                                             the fence need to               the fence need to
7433                                                             conservatively                  conservatively
7434                                                             always generate. If             always generate. If
7435                                                             fence had an                    fence had an
7436                                                             address space then              address space then
7437                                                             set to address                  set to address
7438                                                             space of OpenCL                 space of OpenCL
7439                                                             fence flag, or to               fence flag, or to
7440                                                             generic if both                 generic if both
7441                                                             local and global                local and global
7442                                                             flags are                       flags are
7443                                                             specified.                      specified.
7444                                                           - Must happen after
7445                                                             any preceding
7446                                                             local/generic
7447                                                             load/load
7448                                                             atomic/store/store
7449                                                             atomic/atomicrmw.
7450                                                                                           - Could be split into
7451                                                                                             separate s_waitcnt
7452                                                                                             vmcnt(0), s_waitcnt
7453                                                                                             vscnt(0) and s_waitcnt
7454                                                                                             lgkmcnt(0) to allow
7455                                                                                             them to be
7456                                                                                             independently moved
7457                                                                                             according to the
7458                                                                                             following rules.
7459                                                                                           - s_waitcnt vmcnt(0)
7460                                                                                             must happen after
7461                                                                                             any preceding
7462                                                                                             global/generic
7463                                                                                             load/load
7464                                                                                             atomic/
7465                                                                                             atomicrmw-with-return-value.
7466                                                                                           - s_waitcnt vscnt(0)
7467                                                                                             must happen after
7468                                                                                             any preceding
7469                                                                                             global/generic
7470                                                                                             store/store atomic/
7471                                                                                             atomicrmw-no-return-value.
7472                                                                                           - s_waitcnt lgkmcnt(0)
7473                                                                                             must happen after
7474                                                                                             any preceding
7475                                                                                             local/generic
7476                                                                                             load/store/load
7477                                                                                             atomic/store atomic/
7478                                                                                             atomicrmw.
7479                                                           - Must happen before            - Must happen before
7480                                                             any following store             any following store
7481                                                             atomic/atomicrmw                atomic/atomicrmw
7482                                                             with an equal or                with an equal or
7483                                                             wider sync scope                wider sync scope
7484                                                             and memory ordering             and memory ordering
7485                                                             stronger than                   stronger than
7486                                                             unordered (this is              unordered (this is
7487                                                             termed the                      termed the
7488                                                             fence-paired-atomic).           fence-paired-atomic).
7489                                                           - Ensures that all              - Ensures that all
7490                                                             memory operations               memory operations
7491                                                             to local have                   have
7492                                                             completed before                completed before
7493                                                             performing the                  performing the
7494                                                             following                       following
7495                                                             fence-paired-atomic.            fence-paired-atomic.
7496
7497     fence        release      - agent        *none*     1. s_waitcnt lgkmcnt(0) &       1. s_waitcnt lgkmcnt(0) &
7498                               - system                     vmcnt(0)                        vmcnt(0) & vscnt(0)
7499
7500                                                           - If OpenCL and                 - If OpenCL and
7501                                                             address space is                address space is
7502                                                             not generic, omit               not generic, omit
7503                                                             lgkmcnt(0).                     lgkmcnt(0).
7504                                                           - If OpenCL and                 - If OpenCL and
7505                                                             address space is                address space is
7506                                                             local, omit                     local, omit
7507                                                             vmcnt(0).                       vmcnt(0) and vscnt(0).
7508                                                           - However, since LLVM           - However, since LLVM
7509                                                             currently has no                currently has no
7510                                                             address space on                address space on
7511                                                             the fence need to               the fence need to
7512                                                             conservatively                  conservatively
7513                                                             always generate. If             always generate. If
7514                                                             fence had an                    fence had an
7515                                                             address space then              address space then
7516                                                             set to address                  set to address
7517                                                             space of OpenCL                 space of OpenCL
7518                                                             fence flag, or to               fence flag, or to
7519                                                             generic if both                 generic if both
7520                                                             local and global                local and global
7521                                                             flags are                       flags are
7522                                                             specified.                      specified.
7523                                                           - Could be split into           - Could be split into
7524                                                             separate s_waitcnt              separate s_waitcnt
7525                                                             vmcnt(0) and                    vmcnt(0), s_waitcnt
7526                                                             s_waitcnt                       vscnt(0) and s_waitcnt
7527                                                             lgkmcnt(0) to allow             lgkmcnt(0) to allow
7528                                                             them to be                      them to be
7529                                                             independently moved             independently moved
7530                                                             according to the                according to the
7531                                                             following rules.                following rules.
7532                                                           - s_waitcnt vmcnt(0)            - s_waitcnt vmcnt(0)
7533                                                             must happen after               must happen after
7534                                                             any preceding                   any preceding
7535                                                             global/generic                  global/generic
7536                                                             load/store/load                 load/load atomic/
7537                                                             atomic/store                    atomicrmw-with-return-value.
7538                                                             atomic/atomicrmw.
7539                                                                                           - s_waitcnt vscnt(0)
7540                                                                                             must happen after
7541                                                                                             any preceding
7542                                                                                             global/generic
7543                                                                                             store/store atomic/
7544                                                                                             atomicrmw-no-return-value.
7545                                                           - s_waitcnt lgkmcnt(0)          - s_waitcnt lgkmcnt(0)
7546                                                             must happen after               must happen after
7547                                                             any preceding                   any preceding
7548                                                             local/generic                   local/generic
7549                                                             load/store/load                 load/store/load
7550                                                             atomic/store                    atomic/store
7551                                                             atomic/atomicrmw.               atomic/atomicrmw.
7552                                                           - Must happen before            - Must happen before
7553                                                             any following store             any following store
7554                                                             atomic/atomicrmw                atomic/atomicrmw
7555                                                             with an equal or                with an equal or
7556                                                             wider sync scope                wider sync scope
7557                                                             and memory ordering             and memory ordering
7558                                                             stronger than                   stronger than
7559                                                             unordered (this is              unordered (this is
7560                                                             termed the                      termed the
7561                                                             fence-paired-atomic).           fence-paired-atomic).
7562                                                           - Ensures that all              - Ensures that all
7563                                                             memory operations               memory operations
7564                                                             have                            have
7565                                                             completed before                completed before
7566                                                             performing the                  performing the
7567                                                             following                       following
7568                                                             fence-paired-atomic.            fence-paired-atomic.
7569
7570     **Acquire-Release Atomic**
7571     ----------------------------------------------------------------------------------------------------------------------
7572     atomicrmw    acq_rel      - singlethread - global   1. buffer/global/ds/flat_atomic 1. buffer/global/ds/flat_atomic
7573                               - wavefront    - local
7574                                              - generic
7575     atomicrmw    acq_rel      - workgroup    - global   1. s_waitcnt lgkmcnt(0)         1. s_waitcnt lgkmcnt(0) &
7576                                                                                            vmcnt(0) & vscnt(0)
7577
7578                                                                                           - If CU wavefront execution mode, omit vmcnt and
7579                                                                                             vscnt.
7580                                                           - If OpenCL, omit.              - If OpenCL, omit
7581                                                                                             s_waitcnt lgkmcnt(0).
7582                                                           - Must happen after             - Must happen after
7583                                                             any preceding                   any preceding
7584                                                             local/generic                   local/generic
7585                                                             load/store/load                 load/store/load
7586                                                             atomic/store                    atomic/store
7587                                                             atomic/atomicrmw.               atomic/atomicrmw.
7588                                                                                           - Could be split into
7589                                                                                             separate s_waitcnt
7590                                                                                             vmcnt(0), s_waitcnt
7591                                                                                             vscnt(0) and s_waitcnt
7592                                                                                             lgkmcnt(0) to allow
7593                                                                                             them to be
7594                                                                                             independently moved
7595                                                                                             according to the
7596                                                                                             following rules.
7597                                                                                           - s_waitcnt vmcnt(0)
7598                                                                                             must happen after
7599                                                                                             any preceding
7600                                                                                             global/generic load/load
7601                                                                                             atomic/
7602                                                                                             atomicrmw-with-return-value.
7603                                                                                           - s_waitcnt vscnt(0)
7604                                                                                             must happen after
7605                                                                                             any preceding
7606                                                                                             global/generic
7607                                                                                             store/store
7608                                                                                             atomic/
7609                                                                                             atomicrmw-no-return-value.
7610                                                                                           - s_waitcnt lgkmcnt(0)
7611                                                                                             must happen after
7612                                                                                             any preceding
7613                                                                                             local/generic load/store/load
7614                                                                                             atomic/store atomic/atomicrmw.
7615                                                           - Must happen before            - Must happen before
7616                                                             the following                   the following
7617                                                             atomicrmw.                      atomicrmw.
7618                                                           - Ensures that all              - Ensures that all
7619                                                             memory operations               memory operations
7620                                                             to local have                   have
7621                                                             completed before                completed before
7622                                                             performing the                  performing the
7623                                                             atomicrmw that is               atomicrmw that is
7624                                                             being released.                 being released.
7625
7626                                                         2. buffer/global/flat_atomic    2. buffer/global_atomic
7627                                                                                         3. s_waitcnt vm/vscnt(0)
7628
7629                                                                                           - If CU wavefront execution mode, omit vm/vscnt.
7630                                                                                           - Use vmcnt if atomic with
7631                                                                                             return and vscnt if atomic
7632                                                                                             with no-return.
7633                                                                                             waitcnt lgkmcnt(0).
7634                                                                                           - Must happen before
7635                                                                                             the following
7636                                                                                             buffer_gl0_inv.
7637                                                                                           - Ensures any
7638                                                                                             following global
7639                                                                                             data read is no
7640                                                                                             older than the
7641                                                                                             atomicrmw value
7642                                                                                             being acquired.
7643
7644                                                                                         4. buffer_gl0_inv
7645
7646                                                                                           - If CU wavefront execution mode, omit.
7647                                                                                           - Ensures that
7648                                                                                             following
7649                                                                                             loads will not see
7650                                                                                             stale data.
7651
7652     atomicrmw    acq_rel      - workgroup    - local                                    1. waitcnt vmcnt(0) & vscnt(0)
7653
7654                                                                                           - If CU wavefront execution mode, omit.
7655                                                                                           - If OpenCL, omit.
7656                                                                                           - Could be split into
7657                                                                                             separate s_waitcnt
7658                                                                                             vmcnt(0) and s_waitcnt
7659                                                                                             vscnt(0) to allow
7660                                                                                             them to be
7661                                                                                             independently moved
7662                                                                                             according to the
7663                                                                                             following rules.
7664                                                                                           - s_waitcnt vmcnt(0)
7665                                                                                             must happen after
7666                                                                                             any preceding
7667                                                                                             global/generic load/load
7668                                                                                             atomic/
7669                                                                                             atomicrmw-with-return-value.
7670                                                                                           - s_waitcnt vscnt(0)
7671                                                                                             must happen after
7672                                                                                             any preceding
7673                                                                                             global/generic
7674                                                                                             store/store atomic/
7675                                                                                             atomicrmw-no-return-value.
7676                                                                                           - Must happen before
7677                                                                                             the following
7678                                                                                             store.
7679                                                                                           - Ensures that all
7680                                                                                             global memory
7681                                                                                             operations have
7682                                                                                             completed before
7683                                                                                             performing the
7684                                                                                             store that is being
7685                                                                                             released.
7686
7687                                                         1. ds_atomic                    2. ds_atomic
7688                                                         2. s_waitcnt lgkmcnt(0)         3. s_waitcnt lgkmcnt(0)
7689
7690                                                           - If OpenCL, omit.              - If OpenCL, omit.
7691                                                           - Must happen before            - Must happen before
7692                                                             any following                   the following
7693                                                             global/generic                  buffer_gl0_inv.
7694                                                             load/load
7695                                                             atomic/store/store
7696                                                             atomic/atomicrmw.
7697                                                           - Ensures any                   - Ensures any
7698                                                             following global                following global
7699                                                             data read is no                 data read is no
7700                                                             older than the load             older than the load
7701                                                             atomic value being              atomic value being
7702                                                             acquired.                       acquired.
7703
7704                                                                                         4. buffer_gl0_inv
7705
7706                                                                                           - If CU wavefront execution mode, omit.
7707                                                                                           - If OpenCL omit.
7708                                                                                           - Ensures that
7709                                                                                             following
7710                                                                                             loads will not see
7711                                                                                             stale data.
7712
7713     atomicrmw    acq_rel      - workgroup    - generic  1. s_waitcnt lgkmcnt(0)         1. s_waitcnt lgkmcnt(0) &
7714                                                                                            vmcnt(0) & vscnt(0)
7715
7716                                                                                           - If CU wavefront execution mode, omit vmcnt and
7717                                                                                             vscnt.
7718                                                           - If OpenCL, omit.              - If OpenCL, omit
7719                                                                                             waitcnt lgkmcnt(0).
7720                                                           - Must happen after
7721                                                             any preceding
7722                                                             local/generic
7723                                                             load/store/load
7724                                                             atomic/store
7725                                                             atomic/atomicrmw.
7726                                                                                           - Could be split into
7727                                                                                             separate s_waitcnt
7728                                                                                             vmcnt(0), s_waitcnt
7729                                                                                             vscnt(0) and s_waitcnt
7730                                                                                             lgkmcnt(0) to allow
7731                                                                                             them to be
7732                                                                                             independently moved
7733                                                                                             according to the
7734                                                                                             following rules.
7735                                                                                           - s_waitcnt vmcnt(0)
7736                                                                                             must happen after
7737                                                                                             any preceding
7738                                                                                             global/generic load/load
7739                                                                                             atomic/
7740                                                                                             atomicrmw-with-return-value.
7741                                                                                           - s_waitcnt vscnt(0)
7742                                                                                             must happen after
7743                                                                                             any preceding
7744                                                                                             global/generic
7745                                                                                             store/store
7746                                                                                             atomic/
7747                                                                                             atomicrmw-no-return-value.
7748                                                                                           - s_waitcnt lgkmcnt(0)
7749                                                                                             must happen after
7750                                                                                             any preceding
7751                                                                                             local/generic load/store/load
7752                                                                                             atomic/store atomic/atomicrmw.
7753                                                           - Must happen before            - Must happen before
7754                                                             the following                   the following
7755                                                             atomicrmw.                      atomicrmw.
7756                                                           - Ensures that all              - Ensures that all
7757                                                             memory operations               memory operations
7758                                                             to local have                   have
7759                                                             completed before                completed before
7760                                                             performing the                  performing the
7761                                                             atomicrmw that is               atomicrmw that is
7762                                                             being released.                 being released.
7763
7764                                                         2. flat_atomic                  2. flat_atomic
7765                                                         3. s_waitcnt lgkmcnt(0)         3. s_waitcnt lgkmcnt(0) &
7766                                                                                            vm/vscnt(0)
7767
7768                                                                                           - If CU wavefront execution mode, omit vm/vscnt.
7769                                                           - If OpenCL, omit.              - If OpenCL, omit
7770                                                                                             waitcnt lgkmcnt(0).
7771                                                           - Must happen before            - Must happen before
7772                                                             any following                   the following
7773                                                             global/generic                  buffer_gl0_inv.
7774                                                             load/load
7775                                                             atomic/store/store
7776                                                             atomic/atomicrmw.
7777                                                           - Ensures any                   - Ensures any
7778                                                             following global                following global
7779                                                             data read is no                 data read is no
7780                                                             older than the load             older than the load
7781                                                             atomic value being              atomic value being
7782                                                             acquired.                       acquired.
7783
7784                                                                                         3. buffer_gl0_inv
7785
7786                                                                                           - If CU wavefront execution mode, omit.
7787                                                                                           - Ensures that
7788                                                                                             following
7789                                                                                             loads will not see
7790                                                                                             stale data.
7791
7792     atomicrmw    acq_rel      - agent        - global   1. s_waitcnt lgkmcnt(0) &       1. s_waitcnt lgkmcnt(0) &
7793                               - system                     vmcnt(0)                        vmcnt(0) & vscnt(0)
7794
7795                                                           - If OpenCL, omit               - If OpenCL, omit
7796                                                             lgkmcnt(0).                     lgkmcnt(0).
7797                                                           - Could be split into           - Could be split into
7798                                                             separate s_waitcnt              separate s_waitcnt
7799                                                             vmcnt(0) and                    vmcnt(0), s_waitcnt
7800                                                             s_waitcnt                       vscnt(0) and s_waitcnt
7801                                                             lgkmcnt(0) to allow             lgkmcnt(0) to allow
7802                                                             them to be                      them to be
7803                                                             independently moved             independently moved
7804                                                             according to the                according to the
7805                                                             following rules.                following rules.
7806                                                           - s_waitcnt vmcnt(0)            - s_waitcnt vmcnt(0)
7807                                                             must happen after               must happen after
7808                                                             any preceding                   any preceding
7809                                                             global/generic                  global/generic
7810                                                             load/store/load                 load/load atomic/
7811                                                             atomic/store                    atomicrmw-with-return-value.
7812                                                             atomic/atomicrmw.
7813                                                                                           - s_waitcnt vscnt(0)
7814                                                                                             must happen after
7815                                                                                             any preceding
7816                                                                                             global/generic
7817                                                                                             store/store atomic/
7818                                                                                             atomicrmw-no-return-value.
7819                                                           - s_waitcnt lgkmcnt(0)          - s_waitcnt lgkmcnt(0)
7820                                                             must happen after               must happen after
7821                                                             any preceding                   any preceding
7822                                                             local/generic                   local/generic
7823                                                             load/store/load                 load/store/load
7824                                                             atomic/store                    atomic/store
7825                                                             atomic/atomicrmw.               atomic/atomicrmw.
7826                                                           - Must happen before            - Must happen before
7827                                                             the following                   the following
7828                                                             atomicrmw.                      atomicrmw.
7829                                                           - Ensures that all              - Ensures that all
7830                                                             memory operations               memory operations
7831                                                             to global have                  to global have
7832                                                             completed before                completed before
7833                                                             performing the                  performing the
7834                                                             atomicrmw that is               atomicrmw that is
7835                                                             being released.                 being released.
7836
7837                                                         2. buffer/global/flat_atomic    2. buffer/global_atomic
7838                                                         3. s_waitcnt vmcnt(0)           3. s_waitcnt vm/vscnt(0)
7839
7840                                                                                           - Use vmcnt if atomic with
7841                                                                                             return and vscnt if atomic
7842                                                                                             with no-return.
7843                                                                                             waitcnt lgkmcnt(0).
7844                                                           - Must happen before            - Must happen before
7845                                                             following                       following
7846                                                             buffer_wbinvl1_vol.             buffer_gl*_inv.
7847                                                           - Ensures the                   - Ensures the
7848                                                             atomicrmw has                   atomicrmw has
7849                                                             completed before                completed before
7850                                                             invalidating the                invalidating the
7851                                                             cache.                          caches.
7852
7853                                                         4. buffer_wbinvl1_vol           4. buffer_gl0_inv;
7854                                                                                            buffer_gl1_inv
7855
7856                                                           - Must happen before            - Must happen before
7857                                                             any following                   any following
7858                                                             global/generic                  global/generic
7859                                                             load/load                       load/load
7860                                                             atomic/atomicrmw.               atomic/atomicrmw.
7861                                                           - Ensures that                  - Ensures that
7862                                                             following loads                 following loads
7863                                                             will not see stale              will not see stale
7864                                                             global data.                    global data.
7865
7866     atomicrmw    acq_rel      - agent        - generic  1. s_waitcnt lgkmcnt(0) &       1. s_waitcnt lgkmcnt(0) &
7867                               - system                     vmcnt(0)                        vmcnt(0) & vscnt(0)
7868
7869                                                           - If OpenCL, omit               - If OpenCL, omit
7870                                                             lgkmcnt(0).                     lgkmcnt(0).
7871                                                           - Could be split into           - Could be split into
7872                                                             separate s_waitcnt              separate s_waitcnt
7873                                                             vmcnt(0) and                    vmcnt(0), s_waitcnt
7874                                                             s_waitcnt                       vscnt(0) and s_waitcnt
7875                                                             lgkmcnt(0) to allow             lgkmcnt(0) to allow
7876                                                             them to be                      them to be
7877                                                             independently moved             independently moved
7878                                                             according to the                according to the
7879                                                             following rules.                following rules.
7880                                                           - s_waitcnt vmcnt(0)            - s_waitcnt vmcnt(0)
7881                                                             must happen after               must happen after
7882                                                             any preceding                   any preceding
7883                                                             global/generic                  global/generic
7884                                                             load/store/load                 load/load atomic
7885                                                             atomic/store                    atomicrmw-with-return-value.
7886                                                             atomic/atomicrmw.
7887                                                                                           - s_waitcnt vscnt(0)
7888                                                                                             must happen after
7889                                                                                             any preceding
7890                                                                                             global/generic
7891                                                                                             store/store atomic/
7892                                                                                             atomicrmw-no-return-value.
7893                                                           - s_waitcnt lgkmcnt(0)          - s_waitcnt lgkmcnt(0)
7894                                                             must happen after               must happen after
7895                                                             any preceding                   any preceding
7896                                                             local/generic                   local/generic
7897                                                             load/store/load                 load/store/load
7898                                                             atomic/store                    atomic/store
7899                                                             atomic/atomicrmw.               atomic/atomicrmw.
7900                                                           - Must happen before            - Must happen before
7901                                                             the following                   the following
7902                                                             atomicrmw.                      atomicrmw.
7903                                                           - Ensures that all              - Ensures that all
7904                                                             memory operations               memory operations
7905                                                             to global have                  have
7906                                                             completed before                completed before
7907                                                             performing the                  performing the
7908                                                             atomicrmw that is               atomicrmw that is
7909                                                             being released.                 being released.
7910
7911                                                         2. flat_atomic                  2. flat_atomic
7912                                                         3. s_waitcnt vmcnt(0) &         3. s_waitcnt vm/vscnt(0) &
7913                                                            lgkmcnt(0)                      lgkmcnt(0)
7914
7915                                                           - If OpenCL, omit               - If OpenCL, omit
7916                                                             lgkmcnt(0).                     lgkmcnt(0).
7917                                                                                           - Use vmcnt if atomic with
7918                                                                                             return and vscnt if atomic
7919                                                                                             with no-return.
7920                                                           - Must happen before            - Must happen before
7921                                                             following                       following
7922                                                             buffer_wbinvl1_vol.             buffer_gl*_inv.
7923                                                           - Ensures the                   - Ensures the
7924                                                             atomicrmw has                   atomicrmw has
7925                                                             completed before                completed before
7926                                                             invalidating the                invalidating the
7927                                                             cache.                          caches.
7928
7929                                                         4. buffer_wbinvl1_vol           4. buffer_gl0_inv;
7930                                                                                            buffer_gl1_inv
7931
7932                                                           - Must happen before            - Must happen before
7933                                                             any following                   any following
7934                                                             global/generic                  global/generic
7935                                                             load/load                       load/load
7936                                                             atomic/atomicrmw.               atomic/atomicrmw.
7937                                                           - Ensures that                  - Ensures that
7938                                                             following loads                 following loads
7939                                                             will not see stale              will not see stale
7940                                                             global data.                    global data.
7941
7942     fence        acq_rel      - singlethread *none*     *none*                          *none*
7943                               - wavefront
7944     fence        acq_rel      - workgroup    *none*     1. s_waitcnt lgkmcnt(0)         1. s_waitcnt lgkmcnt(0) &
7945                                                                                            vmcnt(0) & vscnt(0)
7946
7947                                                                                           - If CU wavefront execution mode, omit vmcnt and
7948                                                                                             vscnt.
7949                                                           - If OpenCL and                 - If OpenCL and
7950                                                             address space is                address space is
7951                                                             not generic, omit.              not generic, omit
7952                                                                                             lgkmcnt(0).
7953                                                                                           - If OpenCL and
7954                                                                                             address space is
7955                                                                                             local, omit
7956                                                                                             vmcnt(0) and vscnt(0).
7957                                                           - However,                      - However,
7958                                                             since LLVM                      since LLVM
7959                                                             currently has no                currently has no
7960                                                             address space on                address space on
7961                                                             the fence need to               the fence need to
7962                                                             conservatively                  conservatively
7963                                                             always generate                 always generate
7964                                                             (see comment for                (see comment for
7965                                                             previous fence).                previous fence).
7966                                                           - Must happen after
7967                                                             any preceding
7968                                                             local/generic
7969                                                             load/load
7970                                                             atomic/store/store
7971                                                             atomic/atomicrmw.
7972                                                                                           - Could be split into
7973                                                                                             separate s_waitcnt
7974                                                                                             vmcnt(0), s_waitcnt
7975                                                                                             vscnt(0) and s_waitcnt
7976                                                                                             lgkmcnt(0) to allow
7977                                                                                             them to be
7978                                                                                             independently moved
7979                                                                                             according to the
7980                                                                                             following rules.
7981                                                                                           - s_waitcnt vmcnt(0)
7982                                                                                             must happen after
7983                                                                                             any preceding
7984                                                                                             global/generic
7985                                                                                             load/load
7986                                                                                             atomic/
7987                                                                                             atomicrmw-with-return-value.
7988                                                                                           - s_waitcnt vscnt(0)
7989                                                                                             must happen after
7990                                                                                             any preceding
7991                                                                                             global/generic
7992                                                                                             store/store atomic/
7993                                                                                             atomicrmw-no-return-value.
7994                                                                                           - s_waitcnt lgkmcnt(0)
7995                                                                                             must happen after
7996                                                                                             any preceding
7997                                                                                             local/generic
7998                                                                                             load/store/load
7999                                                                                             atomic/store atomic/
8000                                                                                             atomicrmw.
8001                                                           - Must happen before            - Must happen before
8002                                                             any following                   any following
8003                                                             global/generic                  global/generic
8004                                                             load/load                       load/load
8005                                                             atomic/store/store              atomic/store/store
8006                                                             atomic/atomicrmw.               atomic/atomicrmw.
8007                                                           - Ensures that all              - Ensures that all
8008                                                             memory operations               memory operations
8009                                                             to local have                   have
8010                                                             completed before                completed before
8011                                                             performing any                  performing any
8012                                                             following global                following global
8013                                                             memory operations.              memory operations.
8014                                                           - Ensures that the              - Ensures that the
8015                                                             preceding                       preceding
8016                                                             local/generic load              local/generic load
8017                                                             atomic/atomicrmw                atomic/atomicrmw
8018                                                             with an equal or                with an equal or
8019                                                             wider sync scope                wider sync scope
8020                                                             and memory ordering             and memory ordering
8021                                                             stronger than                   stronger than
8022                                                             unordered (this is              unordered (this is
8023                                                             termed the                      termed the
8024                                                             acquire-fence-paired-atomic     acquire-fence-paired-atomic
8025                                                             ) has completed                 ) has completed
8026                                                             before following                before following
8027                                                             global memory                   global memory
8028                                                             operations. This                operations. This
8029                                                             satisfies the                   satisfies the
8030                                                             requirements of                 requirements of
8031                                                             acquire.                        acquire.
8032                                                           - Ensures that all              - Ensures that all
8033                                                             previous memory                 previous memory
8034                                                             operations have                 operations have
8035                                                             completed before a              completed before a
8036                                                             following                       following
8037                                                             local/generic store             local/generic store
8038                                                             atomic/atomicrmw                atomic/atomicrmw
8039                                                             with an equal or                with an equal or
8040                                                             wider sync scope                wider sync scope
8041                                                             and memory ordering             and memory ordering
8042                                                             stronger than                   stronger than
8043                                                             unordered (this is              unordered (this is
8044                                                             termed the                      termed the
8045                                                             release-fence-paired-atomic     release-fence-paired-atomic
8046                                                             ). This satisfies the           ). This satisfies the
8047                                                             requirements of                 requirements of
8048                                                             release.                        release.
8049                                                                                           - Must happen before
8050                                                                                             the following
8051                                                                                             buffer_gl0_inv.
8052                                                                                           - Ensures that the
8053                                                                                             acquire-fence-paired
8054                                                                                             atomic has completed
8055                                                                                             before invalidating
8056                                                                                             the
8057                                                                                             cache. Therefore
8058                                                                                             any following
8059                                                                                             locations read must
8060                                                                                             be no older than
8061                                                                                             the value read by
8062                                                                                             the
8063                                                                                             acquire-fence-paired-atomic.
8064
8065                                                                                         3. buffer_gl0_inv
8066
8067                                                                                           - If CU wavefront execution mode, omit.
8068                                                                                           - Ensures that
8069                                                                                             following
8070                                                                                             loads will not see
8071                                                                                             stale data.
8072
8073     fence        acq_rel      - agent        *none*     1. s_waitcnt lgkmcnt(0) &       1. s_waitcnt lgkmcnt(0) &
8074                               - system                     vmcnt(0)                        vmcnt(0) & vscnt(0)
8075
8076                                                           - If OpenCL and                 - If OpenCL and
8077                                                             address space is                address space is
8078                                                             not generic, omit               not generic, omit
8079                                                             lgkmcnt(0).                     lgkmcnt(0).
8080                                                                                           - If OpenCL and
8081                                                                                             address space is
8082                                                                                             local, omit
8083                                                                                             vmcnt(0) and vscnt(0).
8084                                                           - However, since LLVM           - However, since LLVM
8085                                                             currently has no                currently has no
8086                                                             address space on                address space on
8087                                                             the fence need to               the fence need to
8088                                                             conservatively                  conservatively
8089                                                             always generate                 always generate
8090                                                             (see comment for                (see comment for
8091                                                             previous fence).                previous fence).
8092                                                           - Could be split into           - Could be split into
8093                                                             separate s_waitcnt              separate s_waitcnt
8094                                                             vmcnt(0) and                    vmcnt(0), s_waitcnt
8095                                                             s_waitcnt                       vscnt(0) and s_waitcnt
8096                                                             lgkmcnt(0) to allow             lgkmcnt(0) to allow
8097                                                             them to be                      them to be
8098                                                             independently moved             independently moved
8099                                                             according to the                according to the
8100                                                             following rules.                following rules.
8101                                                           - s_waitcnt vmcnt(0)            - s_waitcnt vmcnt(0)
8102                                                             must happen after               must happen after
8103                                                             any preceding                   any preceding
8104                                                             global/generic                  global/generic
8105                                                             load/store/load                 load/load
8106                                                             atomic/store                    atomic/
8107                                                             atomic/atomicrmw.               atomicrmw-with-return-value.
8108                                                                                           - s_waitcnt vscnt(0)
8109                                                                                             must happen after
8110                                                                                             any preceding
8111                                                                                             global/generic
8112                                                                                             store/store atomic/
8113                                                                                             atomicrmw-no-return-value.
8114                                                           - s_waitcnt lgkmcnt(0)          - s_waitcnt lgkmcnt(0)
8115                                                             must happen after               must happen after
8116                                                             any preceding                   any preceding
8117                                                             local/generic                   local/generic
8118                                                             load/store/load                 load/store/load
8119                                                             atomic/store                    atomic/store
8120                                                             atomic/atomicrmw.               atomic/atomicrmw.
8121                                                           - Must happen before            - Must happen before
8122                                                             the following                   the following
8123                                                             buffer_wbinvl1_vol.             buffer_gl*_inv.
8124                                                           - Ensures that the              - Ensures that the
8125                                                             preceding                       preceding
8126                                                             global/local/generic            global/local/generic
8127                                                             load                            load
8128                                                             atomic/atomicrmw                atomic/atomicrmw
8129                                                             with an equal or                with an equal or
8130                                                             wider sync scope                wider sync scope
8131                                                             and memory ordering             and memory ordering
8132                                                             stronger than                   stronger than
8133                                                             unordered (this is              unordered (this is
8134                                                             termed the                      termed the
8135                                                             acquire-fence-paired-atomic     acquire-fence-paired-atomic
8136                                                             ) has completed                 ) has completed
8137                                                             before invalidating             before invalidating
8138                                                             the cache. This                 the caches. This
8139                                                             satisfies the                   satisfies the
8140                                                             requirements of                 requirements of
8141                                                             acquire.                        acquire.
8142                                                           - Ensures that all              - Ensures that all
8143                                                             previous memory                 previous memory
8144                                                             operations have                 operations have
8145                                                             completed before a              completed before a
8146                                                             following                       following
8147                                                             global/local/generic            global/local/generic
8148                                                             store                           store
8149                                                             atomic/atomicrmw                atomic/atomicrmw
8150                                                             with an equal or                with an equal or
8151                                                             wider sync scope                wider sync scope
8152                                                             and memory ordering             and memory ordering
8153                                                             stronger than                   stronger than
8154                                                             unordered (this is              unordered (this is
8155                                                             termed the                      termed the
8156                                                             release-fence-paired-atomic     release-fence-paired-atomic
8157                                                             ). This satisfies the           ). This satisfies the
8158                                                             requirements of                 requirements of
8159                                                             release.                        release.
8160
8161                                                         2. buffer_wbinvl1_vol           2. buffer_gl0_inv;
8162                                                                                            buffer_gl1_inv
8163
8164                                                           - Must happen before            - Must happen before
8165                                                             any following                   any following
8166                                                             global/generic                  global/generic
8167                                                             load/load                       load/load
8168                                                             atomic/store/store              atomic/store/store
8169                                                             atomic/atomicrmw.               atomic/atomicrmw.
8170                                                           - Ensures that                  - Ensures that
8171                                                             following loads                 following loads
8172                                                             will not see stale              will not see stale
8173                                                             global data. This               global data. This
8174                                                             satisfies the                   satisfies the
8175                                                             requirements of                 requirements of
8176                                                             acquire.                        acquire.
8177
8178     **Sequential Consistent Atomic**
8179     ----------------------------------------------------------------------------------------------------------------------
8180     load atomic  seq_cst      - singlethread - global   *Same as corresponding          *Same as corresponding
8181                               - wavefront    - local    load atomic acquire,            load atomic acquire,
8182                                              - generic  except must generated           except must generated
8183                                                         all instructions even           all instructions even
8184                                                         for OpenCL.*                    for OpenCL.*
8185     load atomic  seq_cst      - workgroup    - global   1. s_waitcnt lgkmcnt(0)         1. s_waitcnt lgkmcnt(0) &
8186                                              - generic                                     vmcnt(0) & vscnt(0)
8187
8188                                                                                           - If CU wavefront execution mode, omit vmcnt and
8189                                                                                             vscnt.
8190                                                                                           - Could be split into
8191                                                                                             separate s_waitcnt
8192                                                                                             vmcnt(0), s_waitcnt
8193                                                                                             vscnt(0) and s_waitcnt
8194                                                                                             lgkmcnt(0) to allow
8195                                                                                             them to be
8196                                                                                             independently moved
8197                                                                                             according to the
8198                                                                                             following rules.
8199                                                           - Must                          - waitcnt lgkmcnt(0) must
8200                                                             happen after                    happen after
8201                                                             preceding                       preceding
8202                                                             global/generic load             local load
8203                                                             atomic/store                    atomic/store
8204                                                             atomic/atomicrmw                atomic/atomicrmw
8205                                                             with memory                     with memory
8206                                                             ordering of seq_cst             ordering of seq_cst
8207                                                             and with equal or               and with equal or
8208                                                             wider sync scope.               wider sync scope.
8209                                                             (Note that seq_cst              (Note that seq_cst
8210                                                             fences have their               fences have their
8211                                                             own s_waitcnt                   own s_waitcnt
8212                                                             lgkmcnt(0) and so do            lgkmcnt(0) and so do
8213                                                             not need to be                  not need to be
8214                                                             considered.)                    considered.)
8215                                                                                           - waitcnt vmcnt(0)
8216                                                                                             Must happen after
8217                                                                                             preceding
8218                                                                                             global/generic load
8219                                                                                             atomic/
8220                                                                                             atomicrmw-with-return-value
8221                                                                                             with memory
8222                                                                                             ordering of seq_cst
8223                                                                                             and with equal or
8224                                                                                             wider sync scope.
8225                                                                                             (Note that seq_cst
8226                                                                                             fences have their
8227                                                                                             own s_waitcnt
8228                                                                                             vmcnt(0) and so do
8229                                                                                             not need to be
8230                                                                                             considered.)
8231                                                                                           - waitcnt vscnt(0)
8232                                                                                             Must happen after
8233                                                                                             preceding
8234                                                                                             global/generic store
8235                                                                                             atomic/
8236                                                                                             atomicrmw-no-return-value
8237                                                                                             with memory
8238                                                                                             ordering of seq_cst
8239                                                                                             and with equal or
8240                                                                                             wider sync scope.
8241                                                                                             (Note that seq_cst
8242                                                                                             fences have their
8243                                                                                             own s_waitcnt
8244                                                                                             vscnt(0) and so do
8245                                                                                             not need to be
8246                                                                                             considered.)
8247                                                           - Ensures any                   - Ensures any
8248                                                             preceding                       preceding
8249                                                             sequential                      sequential
8250                                                             consistent local                consistent global/local
8251                                                             memory instructions             memory instructions
8252                                                             have completed                  have completed
8253                                                             before executing                before executing
8254                                                             this sequentially               this sequentially
8255                                                             consistent                      consistent
8256                                                             instruction. This               instruction. This
8257                                                             prevents reordering             prevents reordering
8258                                                             a seq_cst store                 a seq_cst store
8259                                                             followed by a                   followed by a
8260                                                             seq_cst load. (Note             seq_cst load. (Note
8261                                                             that seq_cst is                 that seq_cst is
8262                                                             stronger than                   stronger than
8263                                                             acquire/release as              acquire/release as
8264                                                             the reordering of               the reordering of
8265                                                             load acquire                    load acquire
8266                                                             followed by a store             followed by a store
8267                                                             release is                      release is
8268                                                             prevented by the                prevented by the
8269                                                             waitcnt of                      waitcnt of
8270                                                             the release, but                the release, but
8271                                                             there is nothing                there is nothing
8272                                                             preventing a store              preventing a store
8273                                                             release followed by             release followed by
8274                                                             load acquire from               load acquire from
8275                                                             competing out of                competing out of
8276                                                             order.)                         order.)
8277
8278                                                         2. *Following                   2. *Following
8279                                                            instructions same as            instructions same as
8280                                                            corresponding load              corresponding load
8281                                                            atomic acquire,                 atomic acquire,
8282                                                            except must generated           except must generated
8283                                                            all instructions even           all instructions even
8284                                                            for OpenCL.*                    for OpenCL.*
8285     load atomic  seq_cst      - workgroup    - local    *Same as corresponding
8286                                                         load atomic acquire,
8287                                                         except must generated
8288                                                         all instructions even
8289                                                         for OpenCL.*
8290
8291                                                                                         1. s_waitcnt vmcnt(0) & vscnt(0)
8292
8293                                                                                           - If CU wavefront execution mode, omit.
8294                                                                                           - Could be split into
8295                                                                                             separate s_waitcnt
8296                                                                                             vmcnt(0) and s_waitcnt
8297                                                                                             vscnt(0) to allow
8298                                                                                             them to be
8299                                                                                             independently moved
8300                                                                                             according to the
8301                                                                                             following rules.
8302                                                                                           - waitcnt vmcnt(0)
8303                                                                                             Must happen after
8304                                                                                             preceding
8305                                                                                             global/generic load
8306                                                                                             atomic/
8307                                                                                             atomicrmw-with-return-value
8308                                                                                             with memory
8309                                                                                             ordering of seq_cst
8310                                                                                             and with equal or
8311                                                                                             wider sync scope.
8312                                                                                             (Note that seq_cst
8313                                                                                             fences have their
8314                                                                                             own s_waitcnt
8315                                                                                             vmcnt(0) and so do
8316                                                                                             not need to be
8317                                                                                             considered.)
8318                                                                                           - waitcnt vscnt(0)
8319                                                                                             Must happen after
8320                                                                                             preceding
8321                                                                                             global/generic store
8322                                                                                             atomic/
8323                                                                                             atomicrmw-no-return-value
8324                                                                                             with memory
8325                                                                                             ordering of seq_cst
8326                                                                                             and with equal or
8327                                                                                             wider sync scope.
8328                                                                                             (Note that seq_cst
8329                                                                                             fences have their
8330                                                                                             own s_waitcnt
8331                                                                                             vscnt(0) and so do
8332                                                                                             not need to be
8333                                                                                             considered.)
8334                                                                                           - Ensures any
8335                                                                                             preceding
8336                                                                                             sequential
8337                                                                                             consistent global
8338                                                                                             memory instructions
8339                                                                                             have completed
8340                                                                                             before executing
8341                                                                                             this sequentially
8342                                                                                             consistent
8343                                                                                             instruction. This
8344                                                                                             prevents reordering
8345                                                                                             a seq_cst store
8346                                                                                             followed by a
8347                                                                                             seq_cst load. (Note
8348                                                                                             that seq_cst is
8349                                                                                             stronger than
8350                                                                                             acquire/release as
8351                                                                                             the reordering of
8352                                                                                             load acquire
8353                                                                                             followed by a store
8354                                                                                             release is
8355                                                                                             prevented by the
8356                                                                                             waitcnt of
8357                                                                                             the release, but
8358                                                                                             there is nothing
8359                                                                                             preventing a store
8360                                                                                             release followed by
8361                                                                                             load acquire from
8362                                                                                             competing out of
8363                                                                                             order.)
8364
8365                                                                                         2. *Following
8366                                                                                            instructions same as
8367                                                                                            corresponding load
8368                                                                                            atomic acquire,
8369                                                                                            except must generated
8370                                                                                            all instructions even
8371                                                                                            for OpenCL.*
8372
8373     load atomic  seq_cst      - agent        - global   1. s_waitcnt lgkmcnt(0) &       1. s_waitcnt lgkmcnt(0) &
8374                               - system       - generic     vmcnt(0)                        vmcnt(0) & vscnt(0)
8375
8376                                                           - Could be split into           - Could be split into
8377                                                             separate s_waitcnt              separate s_waitcnt
8378                                                             vmcnt(0)                        vmcnt(0), s_waitcnt
8379                                                             and s_waitcnt                   vscnt(0) and s_waitcnt
8380                                                             lgkmcnt(0) to allow             lgkmcnt(0) to allow
8381                                                             them to be                      them to be
8382                                                             independently moved             independently moved
8383                                                             according to the                according to the
8384                                                             following rules.                following rules.
8385                                                           - waitcnt lgkmcnt(0)            - waitcnt lgkmcnt(0)
8386                                                             must happen after               must happen after
8387                                                             preceding                       preceding
8388                                                             global/generic load             local load
8389                                                             atomic/store                    atomic/store
8390                                                             atomic/atomicrmw                atomic/atomicrmw
8391                                                             with memory                     with memory
8392                                                             ordering of seq_cst             ordering of seq_cst
8393                                                             and with equal or               and with equal or
8394                                                             wider sync scope.               wider sync scope.
8395                                                             (Note that seq_cst              (Note that seq_cst
8396                                                             fences have their               fences have their
8397                                                             own s_waitcnt                   own s_waitcnt
8398                                                             lgkmcnt(0) and so do            lgkmcnt(0) and so do
8399                                                             not need to be                  not need to be
8400                                                             considered.)                    considered.)
8401                                                           - waitcnt vmcnt(0)              - waitcnt vmcnt(0)
8402                                                             must happen after               must happen after
8403                                                             preceding                       preceding
8404                                                             global/generic load             global/generic load
8405                                                             atomic/store                    atomic/
8406                                                             atomic/atomicrmw                atomicrmw-with-return-value
8407                                                             with memory                     with memory
8408                                                             ordering of seq_cst             ordering of seq_cst
8409                                                             and with equal or               and with equal or
8410                                                             wider sync scope.               wider sync scope.
8411                                                             (Note that seq_cst              (Note that seq_cst
8412                                                             fences have their               fences have their
8413                                                             own s_waitcnt                   own s_waitcnt
8414                                                             vmcnt(0) and so do              vmcnt(0) and so do
8415                                                             not need to be                  not need to be
8416                                                             considered.)                    considered.)
8417                                                                                           - waitcnt vscnt(0)
8418                                                                                             Must happen after
8419                                                                                             preceding
8420                                                                                             global/generic store
8421                                                                                             atomic/
8422                                                                                             atomicrmw-no-return-value
8423                                                                                             with memory
8424                                                                                             ordering of seq_cst
8425                                                                                             and with equal or
8426                                                                                             wider sync scope.
8427                                                                                             (Note that seq_cst
8428                                                                                             fences have their
8429                                                                                             own s_waitcnt
8430                                                                                             vscnt(0) and so do
8431                                                                                             not need to be
8432                                                                                             considered.)
8433                                                           - Ensures any                   - Ensures any
8434                                                             preceding                       preceding
8435                                                             sequential                      sequential
8436                                                             consistent global               consistent global
8437                                                             memory instructions             memory instructions
8438                                                             have completed                  have completed
8439                                                             before executing                before executing
8440                                                             this sequentially               this sequentially
8441                                                             consistent                      consistent
8442                                                             instruction. This               instruction. This
8443                                                             prevents reordering             prevents reordering
8444                                                             a seq_cst store                 a seq_cst store
8445                                                             followed by a                   followed by a
8446                                                             seq_cst load. (Note             seq_cst load. (Note
8447                                                             that seq_cst is                 that seq_cst is
8448                                                             stronger than                   stronger than
8449                                                             acquire/release as              acquire/release as
8450                                                             the reordering of               the reordering of
8451                                                             load acquire                    load acquire
8452                                                             followed by a store             followed by a store
8453                                                             release is                      release is
8454                                                             prevented by the                prevented by the
8455                                                             waitcnt of                      waitcnt of
8456                                                             the release, but                the release, but
8457                                                             there is nothing                there is nothing
8458                                                             preventing a store              preventing a store
8459                                                             release followed by             release followed by
8460                                                             load acquire from               load acquire from
8461                                                             competing out of                competing out of
8462                                                             order.)                         order.)
8463
8464                                                         2. *Following                   2. *Following
8465                                                            instructions same as            instructions same as
8466                                                            corresponding load              corresponding load
8467                                                            atomic acquire,                 atomic acquire,
8468                                                            except must generated           except must generated
8469                                                            all instructions even           all instructions even
8470                                                            for OpenCL.*                    for OpenCL.*
8471     store atomic seq_cst      - singlethread - global   *Same as corresponding          *Same as corresponding
8472                               - wavefront    - local    store atomic release,           store atomic release,
8473                               - workgroup    - generic  except must generated           except must generated
8474                                                         all instructions even           all instructions even
8475                                                         for OpenCL.*                    for OpenCL.*
8476     store atomic seq_cst      - agent        - global   *Same as corresponding          *Same as corresponding
8477                               - system       - generic  store atomic release,           store atomic release,
8478                                                         except must generated           except must generated
8479                                                         all instructions even           all instructions even
8480                                                         for OpenCL.*                    for OpenCL.*
8481     atomicrmw    seq_cst      - singlethread - global   *Same as corresponding          *Same as corresponding
8482                               - wavefront    - local    atomicrmw acq_rel,              atomicrmw acq_rel,
8483                               - workgroup    - generic  except must generated           except must generated
8484                                                         all instructions even           all instructions even
8485                                                         for OpenCL.*                    for OpenCL.*
8486     atomicrmw    seq_cst      - agent        - global   *Same as corresponding          *Same as corresponding
8487                               - system       - generic  atomicrmw acq_rel,              atomicrmw acq_rel,
8488                                                         except must generated           except must generated
8489                                                         all instructions even           all instructions even
8490                                                         for OpenCL.*                    for OpenCL.*
8491     fence        seq_cst      - singlethread *none*     *Same as corresponding          *Same as corresponding
8492                               - wavefront               fence acq_rel,                  fence acq_rel,
8493                               - workgroup               except must generated           except must generated
8494                               - agent                   all instructions even           all instructions even
8495                               - system                  for OpenCL.*                    for OpenCL.*
8496     ============ ============ ============== ========== =============================== ==================================
8497
8498The memory order also adds the single thread optimization constrains defined in
8499table
8500:ref:`amdgpu-amdhsa-memory-model-single-thread-optimization-constraints-gfx6-gfx10-table`.
8501
8502  .. table:: AMDHSA Memory Model Single Thread Optimization Constraints GFX6-GFX10
8503     :name: amdgpu-amdhsa-memory-model-single-thread-optimization-constraints-gfx6-gfx10-table
8504
8505     ============ ==============================================================
8506     LLVM Memory  Optimization Constraints
8507     Ordering
8508     ============ ==============================================================
8509     unordered    *none*
8510     monotonic    *none*
8511     acquire      - If a load atomic/atomicrmw then no following load/load
8512                    atomic/store/ store atomic/atomicrmw/fence instruction can
8513                    be moved before the acquire.
8514                  - If a fence then same as load atomic, plus no preceding
8515                    associated fence-paired-atomic can be moved after the fence.
8516     release      - If a store atomic/atomicrmw then no preceding load/load
8517                    atomic/store/ store atomic/atomicrmw/fence instruction can
8518                    be moved after the release.
8519                  - If a fence then same as store atomic, plus no following
8520                    associated fence-paired-atomic can be moved before the
8521                    fence.
8522     acq_rel      Same constraints as both acquire and release.
8523     seq_cst      - If a load atomic then same constraints as acquire, plus no
8524                    preceding sequentially consistent load atomic/store
8525                    atomic/atomicrmw/fence instruction can be moved after the
8526                    seq_cst.
8527                  - If a store atomic then the same constraints as release, plus
8528                    no following sequentially consistent load atomic/store
8529                    atomic/atomicrmw/fence instruction can be moved before the
8530                    seq_cst.
8531                  - If an atomicrmw/fence then same constraints as acq_rel.
8532     ============ ==============================================================
8533
8534Trap Handler ABI
8535~~~~~~~~~~~~~~~~
8536
8537For code objects generated by AMDGPU backend for HSA [HSA]_ compatible runtimes
8538(such as ROCm [AMD-ROCm]_), the runtime installs a trap handler that supports
8539the ``s_trap`` instruction with the following usage:
8540
8541  .. table:: AMDGPU Trap Handler for AMDHSA OS
8542     :name: amdgpu-trap-handler-for-amdhsa-os-table
8543
8544     =================== =============== =============== =======================
8545     Usage               Code Sequence   Trap Handler    Description
8546                                         Inputs
8547     =================== =============== =============== =======================
8548     reserved            ``s_trap 0x00``                 Reserved by hardware.
8549     ``debugtrap(arg)``  ``s_trap 0x01`` ``SGPR0-1``:    Reserved for HSA
8550                                           ``queue_ptr`` ``debugtrap``
8551                                         ``VGPR0``:      intrinsic (not
8552                                           ``arg``       implemented).
8553     ``llvm.trap``       ``s_trap 0x02`` ``SGPR0-1``:    Causes dispatch to be
8554                                           ``queue_ptr`` terminated and its
8555                                                         associated queue put
8556                                                         into the error state.
8557     ``llvm.debugtrap``  ``s_trap 0x03``                 - If debugger not
8558                                                           installed then
8559                                                           behaves as a
8560                                                           no-operation. The
8561                                                           trap handler is
8562                                                           entered and
8563                                                           immediately returns
8564                                                           to continue
8565                                                           execution of the
8566                                                           wavefront.
8567                                                         - If the debugger is
8568                                                           installed, causes
8569                                                           the debug trap to be
8570                                                           reported by the
8571                                                           debugger and the
8572                                                           wavefront is put in
8573                                                           the halt state until
8574                                                           resumed by the
8575                                                           debugger.
8576     reserved            ``s_trap 0x04``                 Reserved.
8577     reserved            ``s_trap 0x05``                 Reserved.
8578     reserved            ``s_trap 0x06``                 Reserved.
8579     debugger breakpoint ``s_trap 0x07``                 Reserved for debugger
8580                                                         breakpoints.
8581     reserved            ``s_trap 0x08``                 Reserved.
8582     reserved            ``s_trap 0xfe``                 Reserved.
8583     reserved            ``s_trap 0xff``                 Reserved.
8584     =================== =============== =============== =======================
8585
8586.. _amdgpu-amdhsa-function-call-convention:
8587
8588Call Convention
8589~~~~~~~~~~~~~~~
8590
8591.. note::
8592
8593  This section is currently incomplete and has inakkuracies. It is WIP that will
8594  be updated as information is determined.
8595
8596See :ref:`amdgpu-dwarf-address-space-mapping` for information on swizzled
8597addresses. Unswizzled addresses are normal linear addresses.
8598
8599.. _amdgpu-amdhsa-function-call-convention-kernel-functions:
8600
8601Kernel Functions
8602++++++++++++++++
8603
8604This section describes the call convention ABI for the outer kernel function.
8605
8606See :ref:`amdgpu-amdhsa-initial-kernel-execution-state` for the kernel call
8607convention.
8608
8609The following is not part of the AMDGPU kernel calling convention but describes
8610how the AMDGPU implements function calls:
8611
86121.  Clang decides the kernarg layout to match the *HSA Programmer's Language
8613    Reference* [HSA]_.
8614
8615    - All structs are passed directly.
8616    - Lambda values are passed *TBA*.
8617
8618    .. TODO::
8619
8620      - Does this really follow HSA rules? Or are structs >16 bytes passed
8621        by-value struct?
8622      - What is ABI for lambda values?
8623
86244.  The kernel performs certain setup in its prolog, as described in
8625    :ref:`amdgpu-amdhsa-kernel-prolog`.
8626
8627.. _amdgpu-amdhsa-function-call-convention-non-kernel-functions:
8628
8629Non-Kernel Functions
8630++++++++++++++++++++
8631
8632This section describes the call convention ABI for functions other than the
8633outer kernel function.
8634
8635If a kernel has function calls then scratch is always allocated and used for
8636the call stack which grows from low address to high address using the swizzled
8637scratch address space.
8638
8639On entry to a function:
8640
86411.  SGPR0-3 contain a V# with the following properties (see
8642    :ref:`amdgpu-amdhsa-kernel-prolog-private-segment-buffer`):
8643
8644    * Base address pointing to the beginning of the wavefront scratch backing
8645      memory.
8646    * Swizzled with dword element size and stride of wavefront size elements.
8647
86482.  The FLAT_SCRATCH register pair is setup. See
8649    :ref:`amdgpu-amdhsa-kernel-prolog-flat-scratch`.
86503.  GFX6-8: M0 register set to the size of LDS in bytes. See
8651    :ref:`amdgpu-amdhsa-kernel-prolog-m0`.
86524.  The EXEC register is set to the lanes active on entry to the function.
86535.  MODE register: *TBD*
86546.  VGPR0-31 and SGPR4-29 are used to pass function input arguments as described
8655    below.
86567.  SGPR30-31 return address (RA). The code address that the function must
8657    return to when it completes. The value is undefined if the function is *no
8658    return*.
86598.  SGPR32 is used for the stack pointer (SP). It is an unswizzled scratch
8660    offset relative to the beginning of the wavefront scratch backing memory.
8661
8662    The unswizzled SP can be used with buffer instructions as an unswizzled SGPR
8663    offset with the scratch V# in SGPR0-3 to access the stack in a swizzled
8664    manner.
8665
8666    The unswizzled SP value can be converted into the swizzled SP value by:
8667
8668      | swizzled SP = unswizzled SP / wavefront size
8669
8670    This may be used to obtain the private address space address of stack
8671    objects and to convert this address to a flat address by adding the flat
8672    scratch aperture base address.
8673
8674    The swizzled SP value is always 4 bytes aligned for the ``r600``
8675    architecture and 16 byte aligned for the ``amdgcn`` architecture.
8676
8677    .. note::
8678
8679      The ``amdgcn`` value is selected to avoid dynamic stack alignment for the
8680      OpenCL language which has the largest base type defined as 16 bytes.
8681
8682    On entry, the swizzled SP value is the address of the first function
8683    argument passed on the stack. Other stack passed arguments are positive
8684    offsets from the entry swizzled SP value.
8685
8686    The function may use positive offsets beyond the last stack passed argument
8687    for stack allocated local variables and register spill slots. If necessary
8688    the function may align these to greater alignment than 16 bytes. After these
8689    the function may dynamically allocate space for such things as runtime sized
8690    ``alloca`` local allocations.
8691
8692    If the function calls another function, it will place any stack allocated
8693    arguments after the last local allocation and adjust SGPR32 to the address
8694    after the last local allocation.
8695
86969.  All other registers are unspecified.
869710. Any necessary ``waitcnt`` has been performed to ensure memory is available
8698    to the function.
8699
8700On exit from a function:
8701
87021.  VGPR0-31 and SGPR4-29 are used to pass function result arguments as
8703    described below. Any registers used are considered clobbered registers.
87042.  The following registers are preserved and have the same value as on entry:
8705
8706    * FLAT_SCRATCH
8707    * EXEC
8708    * GFX6-8: M0
8709    * All SGPR and VGPR registers except the clobbered registers of SGPR4-31 and
8710      VGPR0-31.
8711
8712      For the AMDGPU backend, an inter-procedural register allocation (IPRA)
8713      optimization may mark some of clobbered SGPR4-31 and VGPR0-31 registers as
8714      preserved if it can be determined that the called function does not change
8715      their value.
8716
87172.  The PC is set to the RA provided on entry.
87183.  MODE register: *TBD*.
87194.  All other registers are clobbered.
87205.  Any necessary ``waitcnt`` has been performed to ensure memory accessed by
8721    function is available to the caller.
8722
8723.. TODO::
8724
8725  - On gfx908 are all ACC registers clobbered?
8726
8727  - How are function results returned? The address of structured types is passed
8728    by reference, but what about other types?
8729
8730The function input arguments are made up of the formal arguments explicitly
8731declared by the source language function plus the implicit input arguments used
8732by the implementation.
8733
8734The source language input arguments are:
8735
87361. Any source language implicit ``this`` or ``self`` argument comes first as a
8737   pointer type.
87382. Followed by the function formal arguments in left to right source order.
8739
8740The source language result arguments are:
8741
87421. The function result argument.
8743
8744The source language input or result struct type arguments that are less than or
8745equal to 16 bytes, are decomposed recursively into their base type fields, and
8746each field is passed as if a separate argument. For input arguments, if the
8747called function requires the struct to be in memory, for example because its
8748address is taken, then the function body is responsible for allocating a stack
8749location and copying the field arguments into it. Clang terms this *direct
8750struct*.
8751
8752The source language input struct type arguments that are greater than 16 bytes,
8753are passed by reference. The caller is responsible for allocating a stack
8754location to make a copy of the struct value and pass the address as the input
8755argument. The called function is responsible to perform the dereference when
8756accessing the input argument. Clang terms this *by-value struct*.
8757
8758A source language result struct type argument that is greater than 16 bytes, is
8759returned by reference. The caller is responsible for allocating a stack location
8760to hold the result value and passes the address as the last input argument
8761(before the implicit input arguments). In this case there are no result
8762arguments. The called function is responsible to perform the dereference when
8763storing the result value. Clang terms this *structured return (sret)*.
8764
8765*TODO: correct the sret definition.*
8766
8767.. TODO::
8768
8769  Is this definition correct? Or is sret only used if passing in registers, and
8770  pass as non-decomposed struct as stack argument? Or something else? Is the
8771  memory location in the caller stack frame, or a stack memory argument and so
8772  no address is passed as the caller can directly write to the argument stack
8773  location. But then the stack location is still live after return. If an
8774  argument stack location is it the first stack argument or the last one?
8775
8776Lambda argument types are treated as struct types with an implementation defined
8777set of fields.
8778
8779.. TODO::
8780
8781  Need to specify the ABI for lambda types for AMDGPU.
8782
8783For AMDGPU backend all source language arguments (including the decomposed
8784struct type arguments) are passed in VGPRs unless marked ``inreg`` in which case
8785they are passed in SGPRs.
8786
8787The AMDGPU backend walks the function call graph from the leaves to determine
8788which implicit input arguments are used, propagating to each caller of the
8789function. The used implicit arguments are appended to the function arguments
8790after the source language arguments in the following order:
8791
8792.. TODO::
8793
8794  Is recursion or external functions supported?
8795
87961.  Work-Item ID (1 VGPR)
8797
8798    The X, Y and Z work-item ID are packed into a single VGRP with the following
8799    layout. Only fields actually used by the function are set. The other bits
8800    are undefined.
8801
8802    The values come from the initial kernel execution state. See
8803    :ref:`amdgpu-amdhsa-vgpr-register-set-up-order-table`.
8804
8805    .. table:: Work-item implict argument layout
8806      :name: amdgpu-amdhsa-workitem-implict-argument-layout-table
8807
8808      ======= ======= ==============
8809      Bits    Size    Field Name
8810      ======= ======= ==============
8811      9:0     10 bits X Work-Item ID
8812      19:10   10 bits Y Work-Item ID
8813      29:20   10 bits Z Work-Item ID
8814      31:30   2 bits  Unused
8815      ======= ======= ==============
8816
88172.  Dispatch Ptr (2 SGPRs)
8818
8819    The value comes from the initial kernel execution state. See
8820    :ref:`amdgpu-amdhsa-sgpr-register-set-up-order-table`.
8821
88223.  Queue Ptr (2 SGPRs)
8823
8824    The value comes from the initial kernel execution state. See
8825    :ref:`amdgpu-amdhsa-sgpr-register-set-up-order-table`.
8826
88274.  Kernarg Segment Ptr (2 SGPRs)
8828
8829    The value comes from the initial kernel execution state. See
8830    :ref:`amdgpu-amdhsa-sgpr-register-set-up-order-table`.
8831
88325.  Dispatch id (2 SGPRs)
8833
8834    The value comes from the initial kernel execution state. See
8835    :ref:`amdgpu-amdhsa-sgpr-register-set-up-order-table`.
8836
88376.  Work-Group ID X (1 SGPR)
8838
8839    The value comes from the initial kernel execution state. See
8840    :ref:`amdgpu-amdhsa-sgpr-register-set-up-order-table`.
8841
88427.  Work-Group ID Y (1 SGPR)
8843
8844    The value comes from the initial kernel execution state. See
8845    :ref:`amdgpu-amdhsa-sgpr-register-set-up-order-table`.
8846
88478.  Work-Group ID Z (1 SGPR)
8848
8849    The value comes from the initial kernel execution state. See
8850    :ref:`amdgpu-amdhsa-sgpr-register-set-up-order-table`.
8851
88529.  Implicit Argument Ptr (2 SGPRs)
8853
8854    The value is computed by adding an offset to Kernarg Segment Ptr to get the
8855    global address space pointer to the first kernarg implicit argument.
8856
8857The input and result arguments are assigned in order in the following manner:
8858
8859..note::
8860
8861  There are likely some errors and ommissions in the following description that
8862  need correction.
8863
8864  ..TODO::
8865
8866    Check the clang source code to decipher how funtion arguments and return
8867    results are handled. Also see the AMDGPU specific values used.
8868
8869* VGPR arguments are assigned to consecutive VGPRs starting at VGPR0 up to
8870  VGPR31.
8871
8872  If there are more arguments than will fit in these registers, the remaining
8873  arguments are allocated on the stack in order on naturally aligned
8874  addresses.
8875
8876  .. TODO::
8877
8878    How are overly aligned structures allocated on the stack?
8879
8880* SGPR arguments are assigned to consecutive SGPRs starting at SGPR0 up to
8881  SGPR29.
8882
8883  If there are more arguments than will fit in these registers, the remaining
8884  arguments are allocated on the stack in order on naturally aligned
8885  addresses.
8886
8887Note that decomposed struct type arguments may have some fields passed in
8888registers and some in memory.
8889
8890..TODO::
8891
8892  So a struct which can pass some fields as decomposed register arguments, will
8893  pass the rest as decomposed stack elements? But an arguent that will not start
8894  in registers will not be decomposed and will be passed as a non-decomposed
8895  stack value?
8896
8897The following is not part of the AMDGPU function calling convention but
8898describes how the AMDGPU implements function calls:
8899
89001.  SGPR33 is used as a frame pointer (FP) if necessary. Like the SP it is an
8901    unswizzled scratch address. It is only needed if runtime sized ``alloca``
8902    are used, or for the reasons defined in ``SIFrameLowering``.
89032.  Runtime stack alignment is not currently supported.
8904
8905    .. TODO::
8906
8907      - If runtime stack alignment is supported then will an extra argument
8908        pointer register be used?
8909
89102.  Allocating SGPR arguments on the stack are not supported.
8911
89123.  No CFI is currently generated. See :ref:`amdgpu-call-frame-information`.
8913
8914    ..note::
8915
8916      CFI will be generated that defines the CFA as the unswizzled address
8917      relative to the wave scratch base in the unswizzled private address space
8918      of the lowest address stack allocated local variable.
8919
8920      ``DW_AT_frame_base`` will be defined as the swizzled address in the
8921      swizzled private address space by dividing the CFA by the wavefront size
8922      (since CFA is always at least dword aligned which matches the scratch
8923      swizzle element size).
8924
8925      If no dynamic stack alignment was performed, the stack allocated arguments
8926      are accessed as negative offsets relative to ``DW_AT_frame_base``, and the
8927      local variables and register spill slots are accessed as positive offsets
8928      relative to ``DW_AT_frame_base``.
8929
89304.  Function argument passing is implemented by copying the input physical
8931    registers to virtual registers on entry. The register allocator can spill if
8932    necessary. These are copied back to physical registers at call sites. The
8933    net effect is that each function call can have these values in entirely
8934    distinct locations. The IPRA can help avoid shuffling argument registers.
89355.  Call sites are implemented by setting up the arguments at positive offsets
8936    from SP. Then SP is incremented to account for the known frame size before
8937    the call and decremented after the call.
8938
8939    ..note::
8940
8941      The CFI will reflect the changed calculation needed to compute the CFA
8942      from SP.
8943
89446.  4 byte spill slots are used in the stack frame. One slot is allocated for an
8945    emergency spill slot. Buffer instructions are used for stack accesses and
8946    not the ``flat_scratch`` instruction.
8947
8948    ..TODO::
8949
8950      Explain when the emergency spill slot is used.
8951
8952.. TODO::
8953
8954  Possible broken issues:
8955
8956  - Stack arguments must be aligned to required alignment.
8957  - Stack is aligned to max(16, max formal argument alignment)
8958  - Direct argument < 64 bits should check register budget.
8959  - Register budget calculation should respect ``inreg`` for SGPR.
8960  - SGPR overflow is not handled.
8961  - struct with 1 member unpeeling is not checking size of member.
8962  - ``sret`` is after ``this`` pointer.
8963  - Caller is not implementing stack realignment: need an extra pointer.
8964  - Should say AMDGPU passes FP rather than SP.
8965  - Should CFI define CFA as address of locals or arguments. Difference is
8966    apparent when have implemented dynamic alignment.
8967  - If ``SCRATCH`` instruction could allow negative offsets then can make FP be
8968    highest address of stack frame and use negative offset for locals. Would
8969    allow SP to be the same as FP and could support signal-handler-like as now
8970    have a real SP for the top of the stack.
8971  - How is ``sret`` passed on the stack? In argument stack area? Can it overlay
8972    arguments?
8973
8974AMDPAL
8975------
8976
8977This section provides code conventions used when the target triple OS is
8978``amdpal`` (see :ref:`amdgpu-target-triples`) for passing runtime parameters
8979from the application/runtime to each invocation of a hardware shader. These
8980parameters include both generic, application-controlled parameters called
8981*user data* as well as system-generated parameters that are a product of the
8982draw or dispatch execution.
8983
8984User Data
8985~~~~~~~~~
8986
8987Each hardware stage has a set of 32-bit *user data registers* which can be
8988written from a command buffer and then loaded into SGPRs when waves are launched
8989via a subsequent dispatch or draw operation. This is the way most arguments are
8990passed from the application/runtime to a hardware shader.
8991
8992Compute User Data
8993~~~~~~~~~~~~~~~~~
8994
8995Compute shader user data mappings are simpler than graphics shaders, and have a
8996fixed mapping.
8997
8998Note that there are always 10 available *user data entries* in registers -
8999entries beyond that limit must be fetched from memory (via the spill table
9000pointer) by the shader.
9001
9002  .. table:: PAL Compute Shader User Data Registers
9003     :name: pal-compute-user-data-registers
9004
9005     ============= ================================
9006     User Register Description
9007     ============= ================================
9008     0             Global Internal Table (32-bit pointer)
9009     1             Per-Shader Internal Table (32-bit pointer)
9010     2 - 11        Application-Controlled User Data (10 32-bit values)
9011     12            Spill Table (32-bit pointer)
9012     13 - 14       Thread Group Count (64-bit pointer)
9013     15            GDS Range
9014     ============= ================================
9015
9016Graphics User Data
9017~~~~~~~~~~~~~~~~~~
9018
9019Graphics pipelines support a much more flexible user data mapping:
9020
9021  .. table:: PAL Graphics Shader User Data Registers
9022     :name: pal-graphics-user-data-registers
9023
9024     ============= ================================
9025     User Register Description
9026     ============= ================================
9027     0             Global Internal Table (32-bit pointer)
9028     +             Per-Shader Internal Table (32-bit pointer)
9029     + 1-15        Application Controlled User Data
9030                   (1-15 Contiguous 32-bit Values in Registers)
9031     +             Spill Table (32-bit pointer)
9032     +             Draw Index (First Stage Only)
9033     +             Vertex Offset (First Stage Only)
9034     +             Instance Offset (First Stage Only)
9035     ============= ================================
9036
9037  The placement of the global internal table remains fixed in the first *user
9038  data SGPR register*. Otherwise all parameters are optional, and can be mapped
9039  to any desired *user data SGPR register*, with the following restrictions:
9040
9041  * Draw Index, Vertex Offset, and Instance Offset can only be used by the first
9042    active hardware stage in a graphics pipeline (i.e. where the API vertex
9043    shader runs).
9044
9045  * Application-controlled user data must be mapped into a contiguous range of
9046    user data registers.
9047
9048  * The application-controlled user data range supports compaction remapping, so
9049    only *entries* that are actually consumed by the shader must be assigned to
9050    corresponding *registers*. Note that in order to support an efficient runtime
9051    implementation, the remapping must pack *registers* in the same order as
9052    *entries*, with unused *entries* removed.
9053
9054.. _pal_global_internal_table:
9055
9056Global Internal Table
9057~~~~~~~~~~~~~~~~~~~~~
9058
9059The global internal table is a table of *shader resource descriptors* (SRDs)
9060that define how certain engine-wide, runtime-managed resources should be
9061accessed from a shader. The majority of these resources have HW-defined formats,
9062and it is up to the compiler to write/read data as required by the target
9063hardware.
9064
9065The following table illustrates the required format:
9066
9067  .. table:: PAL Global Internal Table
9068     :name: pal-git-table
9069
9070     ============= ================================
9071     Offset        Description
9072     ============= ================================
9073     0-3           Graphics Scratch SRD
9074     4-7           Compute Scratch SRD
9075     8-11          ES/GS Ring Output SRD
9076     12-15         ES/GS Ring Input SRD
9077     16-19         GS/VS Ring Output #0
9078     20-23         GS/VS Ring Output #1
9079     24-27         GS/VS Ring Output #2
9080     28-31         GS/VS Ring Output #3
9081     32-35         GS/VS Ring Input SRD
9082     36-39         Tessellation Factor Buffer SRD
9083     40-43         Off-Chip LDS Buffer SRD
9084     44-47         Off-Chip Param Cache Buffer SRD
9085     48-51         Sample Position Buffer SRD
9086     52            vaRange::ShadowDescriptorTable High Bits
9087     ============= ================================
9088
9089  The pointer to the global internal table passed to the shader as user data
9090  is a 32-bit pointer. The top 32 bits should be assumed to be the same as
9091  the top 32 bits of the pipeline, so the shader may use the program
9092  counter's top 32 bits.
9093
9094Unspecified OS
9095--------------
9096
9097This section provides code conventions used when the target triple OS is
9098empty (see :ref:`amdgpu-target-triples`).
9099
9100Trap Handler ABI
9101~~~~~~~~~~~~~~~~
9102
9103For code objects generated by AMDGPU backend for non-amdhsa OS, the runtime does
9104not install a trap handler. The ``llvm.trap`` and ``llvm.debugtrap``
9105instructions are handled as follows:
9106
9107  .. table:: AMDGPU Trap Handler for Non-AMDHSA OS
9108     :name: amdgpu-trap-handler-for-non-amdhsa-os-table
9109
9110     =============== =============== ===========================================
9111     Usage           Code Sequence   Description
9112     =============== =============== ===========================================
9113     llvm.trap       s_endpgm        Causes wavefront to be terminated.
9114     llvm.debugtrap  *none*          Compiler warning given that there is no
9115                                     trap handler installed.
9116     =============== =============== ===========================================
9117
9118Source Languages
9119================
9120
9121.. _amdgpu-opencl:
9122
9123OpenCL
9124------
9125
9126When the language is OpenCL the following differences occur:
9127
91281. The OpenCL memory model is used (see :ref:`amdgpu-amdhsa-memory-model`).
91292. The AMDGPU backend appends additional arguments to the kernel's explicit
9130   arguments for the AMDHSA OS (see
9131   :ref:`opencl-kernel-implicit-arguments-appended-for-amdhsa-os-table`).
91323. Additional metadata is generated
9133   (see :ref:`amdgpu-amdhsa-code-object-metadata`).
9134
9135  .. table:: OpenCL kernel implicit arguments appended for AMDHSA OS
9136     :name: opencl-kernel-implicit-arguments-appended-for-amdhsa-os-table
9137
9138     ======== ==== ========= ===========================================
9139     Position Byte Byte      Description
9140              Size Alignment
9141     ======== ==== ========= ===========================================
9142     1        8    8         OpenCL Global Offset X
9143     2        8    8         OpenCL Global Offset Y
9144     3        8    8         OpenCL Global Offset Z
9145     4        8    8         OpenCL address of printf buffer
9146     5        8    8         OpenCL address of virtual queue used by
9147                             enqueue_kernel.
9148     6        8    8         OpenCL address of AqlWrap struct used by
9149                             enqueue_kernel.
9150     7        8    8         Pointer argument used for Multi-gird
9151                             synchronization.
9152     ======== ==== ========= ===========================================
9153
9154.. _amdgpu-hcc:
9155
9156HCC
9157---
9158
9159When the language is HCC the following differences occur:
9160
91611. The HSA memory model is used (see :ref:`amdgpu-amdhsa-memory-model`).
9162
9163.. _amdgpu-assembler:
9164
9165Assembler
9166---------
9167
9168AMDGPU backend has LLVM-MC based assembler which is currently in development.
9169It supports AMDGCN GFX6-GFX10.
9170
9171This section describes general syntax for instructions and operands.
9172
9173Instructions
9174~~~~~~~~~~~~
9175
9176.. toctree::
9177   :hidden:
9178
9179   AMDGPU/AMDGPUAsmGFX7
9180   AMDGPU/AMDGPUAsmGFX8
9181   AMDGPU/AMDGPUAsmGFX9
9182   AMDGPU/AMDGPUAsmGFX900
9183   AMDGPU/AMDGPUAsmGFX904
9184   AMDGPU/AMDGPUAsmGFX906
9185   AMDGPU/AMDGPUAsmGFX908
9186   AMDGPU/AMDGPUAsmGFX10
9187   AMDGPU/AMDGPUAsmGFX1011
9188   AMDGPUModifierSyntax
9189   AMDGPUOperandSyntax
9190   AMDGPUInstructionSyntax
9191   AMDGPUInstructionNotation
9192
9193An instruction has the following :doc:`syntax<AMDGPUInstructionSyntax>`:
9194
9195  | ``<``\ *opcode*\ ``> <``\ *operand0*\ ``>, <``\ *operand1*\ ``>,...
9196    <``\ *modifier0*\ ``> <``\ *modifier1*\ ``>...``
9197
9198:doc:`Operands<AMDGPUOperandSyntax>` are comma-separated while
9199:doc:`modifiers<AMDGPUModifierSyntax>` are space-separated.
9200
9201The order of operands and modifiers is fixed.
9202Most modifiers are optional and may be omitted.
9203
9204Links to detailed instruction syntax description may be found in the following
9205table. Note that features under development are not included
9206in this description.
9207
9208    =================================== =======================================
9209    Core ISA                            ISA Extensions
9210    =================================== =======================================
9211    :doc:`GFX7<AMDGPU/AMDGPUAsmGFX7>`   \-
9212    :doc:`GFX8<AMDGPU/AMDGPUAsmGFX8>`   \-
9213    :doc:`GFX9<AMDGPU/AMDGPUAsmGFX9>`   :doc:`gfx900<AMDGPU/AMDGPUAsmGFX900>`
9214
9215                                        :doc:`gfx902<AMDGPU/AMDGPUAsmGFX900>`
9216
9217                                        :doc:`gfx904<AMDGPU/AMDGPUAsmGFX904>`
9218
9219                                        :doc:`gfx906<AMDGPU/AMDGPUAsmGFX906>`
9220
9221                                        :doc:`gfx908<AMDGPU/AMDGPUAsmGFX908>`
9222
9223                                        :doc:`gfx909<AMDGPU/AMDGPUAsmGFX900>`
9224
9225    :doc:`GFX10<AMDGPU/AMDGPUAsmGFX10>` :doc:`gfx1011<AMDGPU/AMDGPUAsmGFX1011>`
9226
9227                                        :doc:`gfx1012<AMDGPU/AMDGPUAsmGFX1011>`
9228    =================================== =======================================
9229
9230For more information about instructions, their semantics and supported
9231combinations of operands, refer to one of instruction set architecture manuals
9232[AMD-GCN-GFX6]_, [AMD-GCN-GFX7]_, [AMD-GCN-GFX8]_, [AMD-GCN-GFX9]_ and
9233[AMD-GCN-GFX10]_.
9234
9235Operands
9236~~~~~~~~
9237
9238Detailed description of operands may be found :doc:`here<AMDGPUOperandSyntax>`.
9239
9240Modifiers
9241~~~~~~~~~
9242
9243Detailed description of modifiers may be found
9244:doc:`here<AMDGPUModifierSyntax>`.
9245
9246Instruction Examples
9247~~~~~~~~~~~~~~~~~~~~
9248
9249DS
9250++
9251
9252.. code-block:: nasm
9253
9254  ds_add_u32 v2, v4 offset:16
9255  ds_write_src2_b64 v2 offset0:4 offset1:8
9256  ds_cmpst_f32 v2, v4, v6
9257  ds_min_rtn_f64 v[8:9], v2, v[4:5]
9258
9259For full list of supported instructions, refer to "LDS/GDS instructions" in ISA
9260Manual.
9261
9262FLAT
9263++++
9264
9265.. code-block:: nasm
9266
9267  flat_load_dword v1, v[3:4]
9268  flat_store_dwordx3 v[3:4], v[5:7]
9269  flat_atomic_swap v1, v[3:4], v5 glc
9270  flat_atomic_cmpswap v1, v[3:4], v[5:6] glc slc
9271  flat_atomic_fmax_x2 v[1:2], v[3:4], v[5:6] glc
9272
9273For full list of supported instructions, refer to "FLAT instructions" in ISA
9274Manual.
9275
9276MUBUF
9277+++++
9278
9279.. code-block:: nasm
9280
9281  buffer_load_dword v1, off, s[4:7], s1
9282  buffer_store_dwordx4 v[1:4], v2, ttmp[4:7], s1 offen offset:4 glc tfe
9283  buffer_store_format_xy v[1:2], off, s[4:7], s1
9284  buffer_wbinvl1
9285  buffer_atomic_inc v1, v2, s[8:11], s4 idxen offset:4 slc
9286
9287For full list of supported instructions, refer to "MUBUF Instructions" in ISA
9288Manual.
9289
9290SMRD/SMEM
9291+++++++++
9292
9293.. code-block:: nasm
9294
9295  s_load_dword s1, s[2:3], 0xfc
9296  s_load_dwordx8 s[8:15], s[2:3], s4
9297  s_load_dwordx16 s[88:103], s[2:3], s4
9298  s_dcache_inv_vol
9299  s_memtime s[4:5]
9300
9301For full list of supported instructions, refer to "Scalar Memory Operations" in
9302ISA Manual.
9303
9304SOP1
9305++++
9306
9307.. code-block:: nasm
9308
9309  s_mov_b32 s1, s2
9310  s_mov_b64 s[0:1], 0x80000000
9311  s_cmov_b32 s1, 200
9312  s_wqm_b64 s[2:3], s[4:5]
9313  s_bcnt0_i32_b64 s1, s[2:3]
9314  s_swappc_b64 s[2:3], s[4:5]
9315  s_cbranch_join s[4:5]
9316
9317For full list of supported instructions, refer to "SOP1 Instructions" in ISA
9318Manual.
9319
9320SOP2
9321++++
9322
9323.. code-block:: nasm
9324
9325  s_add_u32 s1, s2, s3
9326  s_and_b64 s[2:3], s[4:5], s[6:7]
9327  s_cselect_b32 s1, s2, s3
9328  s_andn2_b32 s2, s4, s6
9329  s_lshr_b64 s[2:3], s[4:5], s6
9330  s_ashr_i32 s2, s4, s6
9331  s_bfm_b64 s[2:3], s4, s6
9332  s_bfe_i64 s[2:3], s[4:5], s6
9333  s_cbranch_g_fork s[4:5], s[6:7]
9334
9335For full list of supported instructions, refer to "SOP2 Instructions" in ISA
9336Manual.
9337
9338SOPC
9339++++
9340
9341.. code-block:: nasm
9342
9343  s_cmp_eq_i32 s1, s2
9344  s_bitcmp1_b32 s1, s2
9345  s_bitcmp0_b64 s[2:3], s4
9346  s_setvskip s3, s5
9347
9348For full list of supported instructions, refer to "SOPC Instructions" in ISA
9349Manual.
9350
9351SOPP
9352++++
9353
9354.. code-block:: nasm
9355
9356  s_barrier
9357  s_nop 2
9358  s_endpgm
9359  s_waitcnt 0 ; Wait for all counters to be 0
9360  s_waitcnt vmcnt(0) & expcnt(0) & lgkmcnt(0) ; Equivalent to above
9361  s_waitcnt vmcnt(1) ; Wait for vmcnt counter to be 1.
9362  s_sethalt 9
9363  s_sleep 10
9364  s_sendmsg 0x1
9365  s_sendmsg sendmsg(MSG_INTERRUPT)
9366  s_trap 1
9367
9368For full list of supported instructions, refer to "SOPP Instructions" in ISA
9369Manual.
9370
9371Unless otherwise mentioned, little verification is performed on the operands
9372of SOPP Instructions, so it is up to the programmer to be familiar with the
9373range or acceptable values.
9374
9375VALU
9376++++
9377
9378For vector ALU instruction opcodes (VOP1, VOP2, VOP3, VOPC, VOP_DPP, VOP_SDWA),
9379the assembler will automatically use optimal encoding based on its operands. To
9380force specific encoding, one can add a suffix to the opcode of the instruction:
9381
9382* _e32 for 32-bit VOP1/VOP2/VOPC
9383* _e64 for 64-bit VOP3
9384* _dpp for VOP_DPP
9385* _sdwa for VOP_SDWA
9386
9387VOP1/VOP2/VOP3/VOPC examples:
9388
9389.. code-block:: nasm
9390
9391  v_mov_b32 v1, v2
9392  v_mov_b32_e32 v1, v2
9393  v_nop
9394  v_cvt_f64_i32_e32 v[1:2], v2
9395  v_floor_f32_e32 v1, v2
9396  v_bfrev_b32_e32 v1, v2
9397  v_add_f32_e32 v1, v2, v3
9398  v_mul_i32_i24_e64 v1, v2, 3
9399  v_mul_i32_i24_e32 v1, -3, v3
9400  v_mul_i32_i24_e32 v1, -100, v3
9401  v_addc_u32 v1, s[0:1], v2, v3, s[2:3]
9402  v_max_f16_e32 v1, v2, v3
9403
9404VOP_DPP examples:
9405
9406.. code-block:: nasm
9407
9408  v_mov_b32 v0, v0 quad_perm:[0,2,1,1]
9409  v_sin_f32 v0, v0 row_shl:1 row_mask:0xa bank_mask:0x1 bound_ctrl:0
9410  v_mov_b32 v0, v0 wave_shl:1
9411  v_mov_b32 v0, v0 row_mirror
9412  v_mov_b32 v0, v0 row_bcast:31
9413  v_mov_b32 v0, v0 quad_perm:[1,3,0,1] row_mask:0xa bank_mask:0x1 bound_ctrl:0
9414  v_add_f32 v0, v0, |v0| row_shl:1 row_mask:0xa bank_mask:0x1 bound_ctrl:0
9415  v_max_f16 v1, v2, v3 row_shl:1 row_mask:0xa bank_mask:0x1 bound_ctrl:0
9416
9417VOP_SDWA examples:
9418
9419.. code-block:: nasm
9420
9421  v_mov_b32 v1, v2 dst_sel:BYTE_0 dst_unused:UNUSED_PRESERVE src0_sel:DWORD
9422  v_min_u32 v200, v200, v1 dst_sel:WORD_1 dst_unused:UNUSED_PAD src0_sel:BYTE_1 src1_sel:DWORD
9423  v_sin_f32 v0, v0 dst_unused:UNUSED_PAD src0_sel:WORD_1
9424  v_fract_f32 v0, |v0| dst_sel:DWORD dst_unused:UNUSED_PAD src0_sel:WORD_1
9425  v_cmpx_le_u32 vcc, v1, v2 src0_sel:BYTE_2 src1_sel:WORD_0
9426
9427For full list of supported instructions, refer to "Vector ALU instructions".
9428
9429.. TODO::
9430
9431  Remove once we switch to code object v3 by default.
9432
9433.. _amdgpu-amdhsa-assembler-predefined-symbols-v2:
9434
9435Code Object V2 Predefined Symbols (-mattr=-code-object-v3)
9436~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
9437
9438.. warning:: Code Object V2 is not the default code object version emitted by
9439  this version of LLVM. For a description of the predefined symbols available
9440  with the default configuration (Code Object V3) see
9441  :ref:`amdgpu-amdhsa-assembler-predefined-symbols-v3`.
9442
9443The AMDGPU assembler defines and updates some symbols automatically. These
9444symbols do not affect code generation.
9445
9446.option.machine_version_major
9447+++++++++++++++++++++++++++++
9448
9449Set to the GFX major generation number of the target being assembled for. For
9450example, when assembling for a "GFX9" target this will be set to the integer
9451value "9". The possible GFX major generation numbers are presented in
9452:ref:`amdgpu-processors`.
9453
9454.option.machine_version_minor
9455+++++++++++++++++++++++++++++
9456
9457Set to the GFX minor generation number of the target being assembled for. For
9458example, when assembling for a "GFX810" target this will be set to the integer
9459value "1". The possible GFX minor generation numbers are presented in
9460:ref:`amdgpu-processors`.
9461
9462.option.machine_version_stepping
9463++++++++++++++++++++++++++++++++
9464
9465Set to the GFX stepping generation number of the target being assembled for.
9466For example, when assembling for a "GFX704" target this will be set to the
9467integer value "4". The possible GFX stepping generation numbers are presented
9468in :ref:`amdgpu-processors`.
9469
9470.kernel.vgpr_count
9471++++++++++++++++++
9472
9473Set to zero each time a
9474:ref:`amdgpu-amdhsa-assembler-directive-amdgpu_hsa_kernel` directive is
9475encountered. At each instruction, if the current value of this symbol is less
9476than or equal to the maximum VPGR number explicitly referenced within that
9477instruction then the symbol value is updated to equal that VGPR number plus
9478one.
9479
9480.kernel.sgpr_count
9481++++++++++++++++++
9482
9483Set to zero each time a
9484:ref:`amdgpu-amdhsa-assembler-directive-amdgpu_hsa_kernel` directive is
9485encountered. At each instruction, if the current value of this symbol is less
9486than or equal to the maximum VPGR number explicitly referenced within that
9487instruction then the symbol value is updated to equal that SGPR number plus
9488one.
9489
9490.. _amdgpu-amdhsa-assembler-directives-v2:
9491
9492Code Object V2 Directives (-mattr=-code-object-v3)
9493~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
9494
9495.. warning:: Code Object V2 is not the default code object version emitted by
9496  this version of LLVM. For a description of the directives supported with
9497  the default configuration (Code Object V3) see
9498  :ref:`amdgpu-amdhsa-assembler-directives-v3`.
9499
9500AMDGPU ABI defines auxiliary data in output code object. In assembly source,
9501one can specify them with assembler directives.
9502
9503.hsa_code_object_version major, minor
9504+++++++++++++++++++++++++++++++++++++
9505
9506*major* and *minor* are integers that specify the version of the HSA code
9507object that will be generated by the assembler.
9508
9509.hsa_code_object_isa [major, minor, stepping, vendor, arch]
9510+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
9511
9512
9513*major*, *minor*, and *stepping* are all integers that describe the instruction
9514set architecture (ISA) version of the assembly program.
9515
9516*vendor* and *arch* are quoted strings. *vendor* should always be equal to
9517"AMD" and *arch* should always be equal to "AMDGPU".
9518
9519By default, the assembler will derive the ISA version, *vendor*, and *arch*
9520from the value of the -mcpu option that is passed to the assembler.
9521
9522.. _amdgpu-amdhsa-assembler-directive-amdgpu_hsa_kernel:
9523
9524.amdgpu_hsa_kernel (name)
9525+++++++++++++++++++++++++
9526
9527This directives specifies that the symbol with given name is a kernel entry
9528point (label) and the object should contain corresponding symbol of type
9529STT_AMDGPU_HSA_KERNEL.
9530
9531.amd_kernel_code_t
9532++++++++++++++++++
9533
9534This directive marks the beginning of a list of key / value pairs that are used
9535to specify the amd_kernel_code_t object that will be emitted by the assembler.
9536The list must be terminated by the *.end_amd_kernel_code_t* directive. For any
9537amd_kernel_code_t values that are unspecified a default value will be used. The
9538default value for all keys is 0, with the following exceptions:
9539
9540- *amd_code_version_major* defaults to 1.
9541- *amd_kernel_code_version_minor* defaults to 2.
9542- *amd_machine_kind* defaults to 1.
9543- *amd_machine_version_major*, *machine_version_minor*, and
9544  *amd_machine_version_stepping* are derived from the value of the -mcpu option
9545  that is passed to the assembler.
9546- *kernel_code_entry_byte_offset* defaults to 256.
9547- *wavefront_size* defaults 6 for all targets before GFX10. For GFX10 onwards
9548  defaults to 6 if target feature ``wavefrontsize64`` is enabled, otherwise 5.
9549  Note that wavefront size is specified as a power of two, so a value of **n**
9550  means a size of 2^ **n**.
9551- *call_convention* defaults to -1.
9552- *kernarg_segment_alignment*, *group_segment_alignment*, and
9553  *private_segment_alignment* default to 4. Note that alignments are specified
9554  as a power of 2, so a value of **n** means an alignment of 2^ **n**.
9555- *enable_wgp_mode* defaults to 1 if target feature ``cumode`` is disabled for
9556  GFX10 onwards.
9557- *enable_mem_ordered* defaults to 1 for GFX10 onwards.
9558
9559The *.amd_kernel_code_t* directive must be placed immediately after the
9560function label and before any instructions.
9561
9562For a full list of amd_kernel_code_t keys, refer to AMDGPU ABI document,
9563comments in lib/Target/AMDGPU/AmdKernelCodeT.h and test/CodeGen/AMDGPU/hsa.s.
9564
9565.. _amdgpu-amdhsa-assembler-example-v2:
9566
9567Code Object V2 Example Source Code (-mattr=-code-object-v3)
9568~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
9569
9570.. warning:: Code Object V2 is not the default code object version emitted by
9571  this version of LLVM. For a description of the directives supported with
9572  the default configuration (Code Object V3) see
9573  :ref:`amdgpu-amdhsa-assembler-example-v3`.
9574
9575Here is an example of a minimal assembly source file, defining one HSA kernel:
9576
9577.. code::
9578   :number-lines:
9579
9580   .hsa_code_object_version 1,0
9581   .hsa_code_object_isa
9582
9583   .hsatext
9584   .globl  hello_world
9585   .p2align 8
9586   .amdgpu_hsa_kernel hello_world
9587
9588   hello_world:
9589
9590      .amd_kernel_code_t
9591         enable_sgpr_kernarg_segment_ptr = 1
9592         is_ptr64 = 1
9593         compute_pgm_rsrc1_vgprs = 0
9594         compute_pgm_rsrc1_sgprs = 0
9595         compute_pgm_rsrc2_user_sgpr = 2
9596         compute_pgm_rsrc1_wgp_mode = 0
9597         compute_pgm_rsrc1_mem_ordered = 0
9598         compute_pgm_rsrc1_fwd_progress = 1
9599     .end_amd_kernel_code_t
9600
9601     s_load_dwordx2 s[0:1], s[0:1] 0x0
9602     v_mov_b32 v0, 3.14159
9603     s_waitcnt lgkmcnt(0)
9604     v_mov_b32 v1, s0
9605     v_mov_b32 v2, s1
9606     flat_store_dword v[1:2], v0
9607     s_endpgm
9608   .Lfunc_end0:
9609        .size   hello_world, .Lfunc_end0-hello_world
9610
9611.. _amdgpu-amdhsa-assembler-predefined-symbols-v3:
9612
9613Code Object V3 Predefined Symbols (-mattr=+code-object-v3)
9614~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
9615
9616The AMDGPU assembler defines and updates some symbols automatically. These
9617symbols do not affect code generation.
9618
9619.amdgcn.gfx_generation_number
9620+++++++++++++++++++++++++++++
9621
9622Set to the GFX major generation number of the target being assembled for. For
9623example, when assembling for a "GFX9" target this will be set to the integer
9624value "9". The possible GFX major generation numbers are presented in
9625:ref:`amdgpu-processors`.
9626
9627.amdgcn.gfx_generation_minor
9628++++++++++++++++++++++++++++
9629
9630Set to the GFX minor generation number of the target being assembled for. For
9631example, when assembling for a "GFX810" target this will be set to the integer
9632value "1". The possible GFX minor generation numbers are presented in
9633:ref:`amdgpu-processors`.
9634
9635.amdgcn.gfx_generation_stepping
9636+++++++++++++++++++++++++++++++
9637
9638Set to the GFX stepping generation number of the target being assembled for.
9639For example, when assembling for a "GFX704" target this will be set to the
9640integer value "4". The possible GFX stepping generation numbers are presented
9641in :ref:`amdgpu-processors`.
9642
9643.. _amdgpu-amdhsa-assembler-symbol-next_free_vgpr:
9644
9645.amdgcn.next_free_vgpr
9646++++++++++++++++++++++
9647
9648Set to zero before assembly begins. At each instruction, if the current value
9649of this symbol is less than or equal to the maximum VGPR number explicitly
9650referenced within that instruction then the symbol value is updated to equal
9651that VGPR number plus one.
9652
9653May be used to set the `.amdhsa_next_free_vpgr` directive in
9654:ref:`amdhsa-kernel-directives-table`.
9655
9656May be set at any time, e.g. manually set to zero at the start of each kernel.
9657
9658.. _amdgpu-amdhsa-assembler-symbol-next_free_sgpr:
9659
9660.amdgcn.next_free_sgpr
9661++++++++++++++++++++++
9662
9663Set to zero before assembly begins. At each instruction, if the current value
9664of this symbol is less than or equal the maximum SGPR number explicitly
9665referenced within that instruction then the symbol value is updated to equal
9666that SGPR number plus one.
9667
9668May be used to set the `.amdhsa_next_free_spgr` directive in
9669:ref:`amdhsa-kernel-directives-table`.
9670
9671May be set at any time, e.g. manually set to zero at the start of each kernel.
9672
9673.. _amdgpu-amdhsa-assembler-directives-v3:
9674
9675Code Object V3 Directives (-mattr=+code-object-v3)
9676~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
9677
9678Directives which begin with ``.amdgcn`` are valid for all ``amdgcn``
9679architecture processors, and are not OS-specific. Directives which begin with
9680``.amdhsa`` are specific to ``amdgcn`` architecture processors when the
9681``amdhsa`` OS is specified. See :ref:`amdgpu-target-triples` and
9682:ref:`amdgpu-processors`.
9683
9684.amdgcn_target <target>
9685+++++++++++++++++++++++
9686
9687Optional directive which declares the target supported by the containing
9688assembler source file. Valid values are described in
9689:ref:`amdgpu-amdhsa-code-object-target-identification`. Used by the assembler
9690to validate command-line options such as ``-triple``, ``-mcpu``, and those
9691which specify target features.
9692
9693.amdhsa_kernel <name>
9694+++++++++++++++++++++
9695
9696Creates a correctly aligned AMDHSA kernel descriptor and a symbol,
9697``<name>.kd``, in the current location of the current section. Only valid when
9698the OS is ``amdhsa``. ``<name>`` must be a symbol that labels the first
9699instruction to execute, and does not need to be previously defined.
9700
9701Marks the beginning of a list of directives used to generate the bytes of a
9702kernel descriptor, as described in :ref:`amdgpu-amdhsa-kernel-descriptor`.
9703Directives which may appear in this list are described in
9704:ref:`amdhsa-kernel-directives-table`. Directives may appear in any order, must
9705be valid for the target being assembled for, and cannot be repeated. Directives
9706support the range of values specified by the field they reference in
9707:ref:`amdgpu-amdhsa-kernel-descriptor`. If a directive is not specified, it is
9708assumed to have its default value, unless it is marked as "Required", in which
9709case it is an error to omit the directive. This list of directives is
9710terminated by an ``.end_amdhsa_kernel`` directive.
9711
9712  .. table:: AMDHSA Kernel Assembler Directives
9713     :name: amdhsa-kernel-directives-table
9714
9715     ======================================================== =================== ============ ===================
9716     Directive                                                Default             Supported On Description
9717     ======================================================== =================== ============ ===================
9718     ``.amdhsa_group_segment_fixed_size``                     0                   GFX6-GFX10   Controls GROUP_SEGMENT_FIXED_SIZE in
9719                                                                                               :ref:`amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table`.
9720     ``.amdhsa_private_segment_fixed_size``                   0                   GFX6-GFX10   Controls PRIVATE_SEGMENT_FIXED_SIZE in
9721                                                                                               :ref:`amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table`.
9722     ``.amdhsa_user_sgpr_private_segment_buffer``             0                   GFX6-GFX10   Controls ENABLE_SGPR_PRIVATE_SEGMENT_BUFFER in
9723                                                                                               :ref:`amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table`.
9724     ``.amdhsa_user_sgpr_dispatch_ptr``                       0                   GFX6-GFX10   Controls ENABLE_SGPR_DISPATCH_PTR in
9725                                                                                               :ref:`amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table`.
9726     ``.amdhsa_user_sgpr_queue_ptr``                          0                   GFX6-GFX10   Controls ENABLE_SGPR_QUEUE_PTR in
9727                                                                                               :ref:`amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table`.
9728     ``.amdhsa_user_sgpr_kernarg_segment_ptr``                0                   GFX6-GFX10   Controls ENABLE_SGPR_KERNARG_SEGMENT_PTR in
9729                                                                                               :ref:`amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table`.
9730     ``.amdhsa_user_sgpr_dispatch_id``                        0                   GFX6-GFX10   Controls ENABLE_SGPR_DISPATCH_ID in
9731                                                                                               :ref:`amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table`.
9732     ``.amdhsa_user_sgpr_flat_scratch_init``                  0                   GFX6-GFX10   Controls ENABLE_SGPR_FLAT_SCRATCH_INIT in
9733                                                                                               :ref:`amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table`.
9734     ``.amdhsa_user_sgpr_private_segment_size``               0                   GFX6-GFX10   Controls ENABLE_SGPR_PRIVATE_SEGMENT_SIZE in
9735                                                                                               :ref:`amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table`.
9736     ``.amdhsa_wavefront_size32``                             Target              GFX10        Controls ENABLE_WAVEFRONT_SIZE32 in
9737                                                              Feature                          :ref:`amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table`.
9738                                                              Specific
9739                                                              (-wavefrontsize64)
9740     ``.amdhsa_system_sgpr_private_segment_wavefront_offset`` 0                   GFX6-GFX10   Controls ENABLE_SGPR_PRIVATE_SEGMENT_WAVEFRONT_OFFSET in
9741                                                                                               :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`.
9742     ``.amdhsa_system_sgpr_workgroup_id_x``                   1                   GFX6-GFX10   Controls ENABLE_SGPR_WORKGROUP_ID_X in
9743                                                                                               :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`.
9744     ``.amdhsa_system_sgpr_workgroup_id_y``                   0                   GFX6-GFX10   Controls ENABLE_SGPR_WORKGROUP_ID_Y in
9745                                                                                               :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`.
9746     ``.amdhsa_system_sgpr_workgroup_id_z``                   0                   GFX6-GFX10   Controls ENABLE_SGPR_WORKGROUP_ID_Z in
9747                                                                                               :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`.
9748     ``.amdhsa_system_sgpr_workgroup_info``                   0                   GFX6-GFX10   Controls ENABLE_SGPR_WORKGROUP_INFO in
9749                                                                                               :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`.
9750     ``.amdhsa_system_vgpr_workitem_id``                      0                   GFX6-GFX10   Controls ENABLE_VGPR_WORKITEM_ID in
9751                                                                                               :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`.
9752                                                                                               Possible values are defined in
9753                                                                                               :ref:`amdgpu-amdhsa-system-vgpr-work-item-id-enumeration-values-table`.
9754     ``.amdhsa_next_free_vgpr``                               Required            GFX6-GFX10   Maximum VGPR number explicitly referenced, plus one.
9755                                                                                               Used to calculate GRANULATED_WORKITEM_VGPR_COUNT in
9756                                                                                               :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`.
9757     ``.amdhsa_next_free_sgpr``                               Required            GFX6-GFX10   Maximum SGPR number explicitly referenced, plus one.
9758                                                                                               Used to calculate GRANULATED_WAVEFRONT_SGPR_COUNT in
9759                                                                                               :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`.
9760     ``.amdhsa_reserve_vcc``                                  1                   GFX6-GFX10   Whether the kernel may use the special VCC SGPR.
9761                                                                                               Used to calculate GRANULATED_WAVEFRONT_SGPR_COUNT in
9762                                                                                               :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`.
9763     ``.amdhsa_reserve_flat_scratch``                         1                   GFX7-GFX10   Whether the kernel may use flat instructions to access
9764                                                                                               scratch memory. Used to calculate
9765                                                                                               GRANULATED_WAVEFRONT_SGPR_COUNT in
9766                                                                                               :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`.
9767     ``.amdhsa_reserve_xnack_mask``                           Target              GFX8-GFX10   Whether the kernel may trigger XNACK replay.
9768                                                              Feature                          Used to calculate GRANULATED_WAVEFRONT_SGPR_COUNT in
9769                                                              Specific                         :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`.
9770                                                              (+xnack)
9771     ``.amdhsa_float_round_mode_32``                          0                   GFX6-GFX10   Controls FLOAT_ROUND_MODE_32 in
9772                                                                                               :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`.
9773                                                                                               Possible values are defined in
9774                                                                                               :ref:`amdgpu-amdhsa-floating-point-rounding-mode-enumeration-values-table`.
9775     ``.amdhsa_float_round_mode_16_64``                       0                   GFX6-GFX10   Controls FLOAT_ROUND_MODE_16_64 in
9776                                                                                               :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`.
9777                                                                                               Possible values are defined in
9778                                                                                               :ref:`amdgpu-amdhsa-floating-point-rounding-mode-enumeration-values-table`.
9779     ``.amdhsa_float_denorm_mode_32``                         0                   GFX6-GFX10   Controls FLOAT_DENORM_MODE_32 in
9780                                                                                               :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`.
9781                                                                                               Possible values are defined in
9782                                                                                               :ref:`amdgpu-amdhsa-floating-point-denorm-mode-enumeration-values-table`.
9783     ``.amdhsa_float_denorm_mode_16_64``                      3                   GFX6-GFX10   Controls FLOAT_DENORM_MODE_16_64 in
9784                                                                                               :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`.
9785                                                                                               Possible values are defined in
9786                                                                                               :ref:`amdgpu-amdhsa-floating-point-denorm-mode-enumeration-values-table`.
9787     ``.amdhsa_dx10_clamp``                                   1                   GFX6-GFX10   Controls ENABLE_DX10_CLAMP in
9788                                                                                               :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`.
9789     ``.amdhsa_ieee_mode``                                    1                   GFX6-GFX10   Controls ENABLE_IEEE_MODE in
9790                                                                                               :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`.
9791     ``.amdhsa_fp16_overflow``                                0                   GFX9-GFX10   Controls FP16_OVFL in
9792                                                                                               :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`.
9793     ``.amdhsa_workgroup_processor_mode``                     Target              GFX10        Controls ENABLE_WGP_MODE in
9794                                                              Feature                          :ref:`amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table`.
9795                                                              Specific
9796                                                              (-cumode)
9797     ``.amdhsa_memory_ordered``                               1                   GFX10        Controls MEM_ORDERED in
9798                                                                                               :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`.
9799     ``.amdhsa_forward_progress``                             0                   GFX10        Controls FWD_PROGRESS in
9800                                                                                               :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`.
9801     ``.amdhsa_exception_fp_ieee_invalid_op``                 0                   GFX6-GFX10   Controls ENABLE_EXCEPTION_IEEE_754_FP_INVALID_OPERATION in
9802                                                                                               :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`.
9803     ``.amdhsa_exception_fp_denorm_src``                      0                   GFX6-GFX10   Controls ENABLE_EXCEPTION_FP_DENORMAL_SOURCE in
9804                                                                                               :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`.
9805     ``.amdhsa_exception_fp_ieee_div_zero``                   0                   GFX6-GFX10   Controls ENABLE_EXCEPTION_IEEE_754_FP_DIVISION_BY_ZERO in
9806                                                                                               :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`.
9807     ``.amdhsa_exception_fp_ieee_overflow``                   0                   GFX6-GFX10   Controls ENABLE_EXCEPTION_IEEE_754_FP_OVERFLOW in
9808                                                                                               :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`.
9809     ``.amdhsa_exception_fp_ieee_underflow``                  0                   GFX6-GFX10   Controls ENABLE_EXCEPTION_IEEE_754_FP_UNDERFLOW in
9810                                                                                               :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`.
9811     ``.amdhsa_exception_fp_ieee_inexact``                    0                   GFX6-GFX10   Controls ENABLE_EXCEPTION_IEEE_754_FP_INEXACT in
9812                                                                                               :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`.
9813     ``.amdhsa_exception_int_div_zero``                       0                   GFX6-GFX10   Controls ENABLE_EXCEPTION_INT_DIVIDE_BY_ZERO in
9814                                                                                               :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`.
9815     ======================================================== =================== ============ ===================
9816
9817.amdgpu_metadata
9818++++++++++++++++
9819
9820Optional directive which declares the contents of the ``NT_AMDGPU_METADATA``
9821note record (see :ref:`amdgpu-elf-note-records-table-v3`).
9822
9823The contents must be in the [YAML]_ markup format, with the same structure and
9824semantics described in :ref:`amdgpu-amdhsa-code-object-metadata-v3`.
9825
9826This directive is terminated by an ``.end_amdgpu_metadata`` directive.
9827
9828.. _amdgpu-amdhsa-assembler-example-v3:
9829
9830Code Object V3 Example Source Code (-mattr=+code-object-v3)
9831~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
9832
9833Here is an example of a minimal assembly source file, defining one HSA kernel:
9834
9835.. code::
9836   :number-lines:
9837
9838   .amdgcn_target "amdgcn-amd-amdhsa--gfx900+xnack" // optional
9839
9840   .text
9841   .globl hello_world
9842   .p2align 8
9843   .type hello_world,@function
9844   hello_world:
9845     s_load_dwordx2 s[0:1], s[0:1] 0x0
9846     v_mov_b32 v0, 3.14159
9847     s_waitcnt lgkmcnt(0)
9848     v_mov_b32 v1, s0
9849     v_mov_b32 v2, s1
9850     flat_store_dword v[1:2], v0
9851     s_endpgm
9852   .Lfunc_end0:
9853     .size   hello_world, .Lfunc_end0-hello_world
9854
9855   .rodata
9856   .p2align 6
9857   .amdhsa_kernel hello_world
9858     .amdhsa_user_sgpr_kernarg_segment_ptr 1
9859     .amdhsa_next_free_vgpr .amdgcn.next_free_vgpr
9860     .amdhsa_next_free_sgpr .amdgcn.next_free_sgpr
9861   .end_amdhsa_kernel
9862
9863   .amdgpu_metadata
9864   ---
9865   amdhsa.version:
9866     - 1
9867     - 0
9868   amdhsa.kernels:
9869     - .name: hello_world
9870       .symbol: hello_world.kd
9871       .kernarg_segment_size: 48
9872       .group_segment_fixed_size: 0
9873       .private_segment_fixed_size: 0
9874       .kernarg_segment_align: 4
9875       .wavefront_size: 64
9876       .sgpr_count: 2
9877       .vgpr_count: 3
9878       .max_flat_workgroup_size: 256
9879   ...
9880   .end_amdgpu_metadata
9881
9882If an assembly source file contains multiple kernels and/or functions, the
9883:ref:`amdgpu-amdhsa-assembler-symbol-next_free_vgpr` and
9884:ref:`amdgpu-amdhsa-assembler-symbol-next_free_sgpr` symbols may be reset using
9885the ``.set <symbol>, <expression>`` directive. For example, in the case of two
9886kernels, where ``function1`` is only called from ``kernel1`` it is sufficient
9887to group the function with the kernel that calls it and reset the symbols
9888between the two connected components:
9889
9890.. code::
9891   :number-lines:
9892
9893   .amdgcn_target "amdgcn-amd-amdhsa--gfx900+xnack" // optional
9894
9895   // gpr tracking symbols are implicitly set to zero
9896
9897   .text
9898   .globl kern0
9899   .p2align 8
9900   .type kern0,@function
9901   kern0:
9902     // ...
9903     s_endpgm
9904   .Lkern0_end:
9905     .size   kern0, .Lkern0_end-kern0
9906
9907   .rodata
9908   .p2align 6
9909   .amdhsa_kernel kern0
9910     // ...
9911     .amdhsa_next_free_vgpr .amdgcn.next_free_vgpr
9912     .amdhsa_next_free_sgpr .amdgcn.next_free_sgpr
9913   .end_amdhsa_kernel
9914
9915   // reset symbols to begin tracking usage in func1 and kern1
9916   .set .amdgcn.next_free_vgpr, 0
9917   .set .amdgcn.next_free_sgpr, 0
9918
9919   .text
9920   .hidden func1
9921   .global func1
9922   .p2align 2
9923   .type func1,@function
9924   func1:
9925     // ...
9926     s_setpc_b64 s[30:31]
9927   .Lfunc1_end:
9928   .size func1, .Lfunc1_end-func1
9929
9930   .globl kern1
9931   .p2align 8
9932   .type kern1,@function
9933   kern1:
9934     // ...
9935     s_getpc_b64 s[4:5]
9936     s_add_u32 s4, s4, func1@rel32@lo+4
9937     s_addc_u32 s5, s5, func1@rel32@lo+4
9938     s_swappc_b64 s[30:31], s[4:5]
9939     // ...
9940     s_endpgm
9941   .Lkern1_end:
9942     .size   kern1, .Lkern1_end-kern1
9943
9944   .rodata
9945   .p2align 6
9946   .amdhsa_kernel kern1
9947     // ...
9948     .amdhsa_next_free_vgpr .amdgcn.next_free_vgpr
9949     .amdhsa_next_free_sgpr .amdgcn.next_free_sgpr
9950   .end_amdhsa_kernel
9951
9952These symbols cannot identify connected components in order to automatically
9953track the usage for each kernel. However, in some cases careful organization of
9954the kernels and functions in the source file means there is minimal additional
9955effort required to accurately calculate GPR usage.
9956
9957Additional Documentation
9958========================
9959
9960.. [AMD-RADEON-HD-2000-3000] `AMD R6xx shader ISA <http://developer.amd.com/wordpress/media/2012/10/R600_Instruction_Set_Architecture.pdf>`__
9961.. [AMD-RADEON-HD-4000] `AMD R7xx shader ISA <http://developer.amd.com/wordpress/media/2012/10/R700-Family_Instruction_Set_Architecture.pdf>`__
9962.. [AMD-RADEON-HD-5000] `AMD Evergreen shader ISA <http://developer.amd.com/wordpress/media/2012/10/AMD_Evergreen-Family_Instruction_Set_Architecture.pdf>`__
9963.. [AMD-RADEON-HD-6000] `AMD Cayman/Trinity shader ISA <http://developer.amd.com/wordpress/media/2012/10/AMD_HD_6900_Series_Instruction_Set_Architecture.pdf>`__
9964.. [AMD-GCN-GFX6] `AMD Southern Islands Series ISA <http://developer.amd.com/wordpress/media/2012/12/AMD_Southern_Islands_Instruction_Set_Architecture.pdf>`__
9965.. [AMD-GCN-GFX7] `AMD Sea Islands Series ISA <http://developer.amd.com/wordpress/media/2013/07/AMD_Sea_Islands_Instruction_Set_Architecture.pdf>`_
9966.. [AMD-GCN-GFX8] `AMD GCN3 Instruction Set Architecture <http://amd-dev.wpengine.netdna-cdn.com/wordpress/media/2013/12/AMD_GCN3_Instruction_Set_Architecture_rev1.1.pdf>`__
9967.. [AMD-GCN-GFX9] `AMD "Vega" Instruction Set Architecture <http://developer.amd.com/wordpress/media/2013/12/Vega_Shader_ISA_28July2017.pdf>`__
9968.. [AMD-GCN-GFX10] `AMD "RDNA 1.0" Instruction Set Architecture <https://gpuopen.com/wp-content/uploads/2019/08/RDNA_Shader_ISA_5August2019.pdf>`__
9969.. [AMD-ROCm] `ROCm: Open Platform for Development, Discovery and Education Around GPU Computing <http://gpuopen.com/compute-product/rocm/>`__
9970.. [AMD-ROCm-github] `ROCm github <http://github.com/RadeonOpenCompute>`__
9971.. [HSA] `Heterogeneous System Architecture (HSA) Foundation <http://www.hsafoundation.com/>`__
9972.. [HIP] `HIP Programming Guide <https://rocm-documentation.readthedocs.io/en/latest/Programming_Guides/Programming-Guides.html#hip-programing-guide>`__
9973.. [ELF] `Executable and Linkable Format (ELF) <http://www.sco.com/developers/gabi/>`__
9974.. [DWARF] `DWARF Debugging Information Format <http://dwarfstd.org/>`__
9975.. [YAML] `YAML Ain't Markup Language (YAML™) Version 1.2 <http://www.yaml.org/spec/1.2/spec.html>`__
9976.. [MsgPack] `Message Pack <http://www.msgpack.org/>`__
9977.. [SEMVER] `Semantic Versioning <https://semver.org/>`__
9978.. [OpenCL] `The OpenCL Specification Version 2.0 <http://www.khronos.org/registry/cl/specs/opencl-2.0.pdf>`__
9979.. [HRF] `Heterogeneous-race-free Memory Models <http://benedictgaster.org/wp-content/uploads/2014/01/asplos269-FINAL.pdf>`__
9980.. [CLANG-ATTR] `Attributes in Clang <https://clang.llvm.org/docs/AttributeReference.html>`__
9981