1=============================
2User Guide for AMDGPU Backend
3=============================
4
5.. contents::
6   :local:
7
8.. toctree::
9   :hidden:
10
11   AMDGPU/AMDGPUAsmGFX7
12   AMDGPU/AMDGPUAsmGFX8
13   AMDGPU/AMDGPUAsmGFX9
14   AMDGPU/AMDGPUAsmGFX900
15   AMDGPU/AMDGPUAsmGFX904
16   AMDGPU/AMDGPUAsmGFX906
17   AMDGPU/AMDGPUAsmGFX908
18   AMDGPU/AMDGPUAsmGFX10
19   AMDGPU/AMDGPUAsmGFX1011
20   AMDGPUModifierSyntax
21   AMDGPUOperandSyntax
22   AMDGPUInstructionSyntax
23   AMDGPUInstructionNotation
24   AMDGPUDwarfExtensionsForHeterogeneousDebugging
25
26Introduction
27============
28
29The AMDGPU backend provides ISA code generation for AMD GPUs, starting with the
30R600 family up until the current GCN families. It lives in the
31``llvm/lib/Target/AMDGPU`` directory.
32
33LLVM
34====
35
36.. _amdgpu-target-triples:
37
38Target Triples
39--------------
40
41Use the ``clang -target <Architecture>-<Vendor>-<OS>-<Environment>`` option to
42specify the target triple:
43
44  .. table:: AMDGPU Architectures
45     :name: amdgpu-architecture-table
46
47     ============ ==============================================================
48     Architecture Description
49     ============ ==============================================================
50     ``r600``     AMD GPUs HD2XXX-HD6XXX for graphics and compute shaders.
51     ``amdgcn``   AMD GPUs GCN GFX6 onwards for graphics and compute shaders.
52     ============ ==============================================================
53
54  .. table:: AMDGPU Vendors
55     :name: amdgpu-vendor-table
56
57     ============ ==============================================================
58     Vendor       Description
59     ============ ==============================================================
60     ``amd``      Can be used for all AMD GPU usage.
61     ``mesa3d``   Can be used if the OS is ``mesa3d``.
62     ============ ==============================================================
63
64  .. table:: AMDGPU Operating Systems
65     :name: amdgpu-os-table
66
67     ============== ============================================================
68     OS             Description
69     ============== ============================================================
70     *<empty>*      Defaults to the *unknown* OS.
71     ``amdhsa``     Compute kernels executed on HSA [HSA]_ compatible runtimes
72                    such as AMD's ROCm [AMD-ROCm]_.
73     ``amdpal``     Graphic shaders and compute kernels executed on AMD PAL
74                    runtime.
75     ``mesa3d``     Graphic shaders and compute kernels executed on Mesa 3D
76                    runtime.
77     ============== ============================================================
78
79  .. table:: AMDGPU Environments
80     :name: amdgpu-environment-table
81
82     ============ ==============================================================
83     Environment  Description
84     ============ ==============================================================
85     *<empty>*    Default.
86     ============ ==============================================================
87
88.. _amdgpu-processors:
89
90Processors
91----------
92
93Use the ``clang -mcpu <Processor>`` option to specify the AMDGPU processor. The
94names from both the *Processor* and *Alternative Processor* can be used.
95
96  .. table:: AMDGPU Processors
97     :name: amdgpu-processor-table
98
99     =========== =============== ============ ===== ================= ======= ======================
100     Processor   Alternative     Target       dGPU/ Target            ROCm    Example
101                 Processor       Triple       APU   Features          Support Products
102                                 Architecture       Supported
103                                                    [Default]
104     =========== =============== ============ ===== ================= ======= ======================
105     **Radeon HD 2000/3000 Series (R600)** [AMD-RADEON-HD-2000-3000]_
106     -----------------------------------------------------------------------------------------------
107     ``r600``                    ``r600``     dGPU
108     ``r630``                    ``r600``     dGPU
109     ``rs880``                   ``r600``     dGPU
110     ``rv670``                   ``r600``     dGPU
111     **Radeon HD 4000 Series (R700)** [AMD-RADEON-HD-4000]_
112     -----------------------------------------------------------------------------------------------
113     ``rv710``                   ``r600``     dGPU
114     ``rv730``                   ``r600``     dGPU
115     ``rv770``                   ``r600``     dGPU
116     **Radeon HD 5000 Series (Evergreen)** [AMD-RADEON-HD-5000]_
117     -----------------------------------------------------------------------------------------------
118     ``cedar``                   ``r600``     dGPU
119     ``cypress``                 ``r600``     dGPU
120     ``juniper``                 ``r600``     dGPU
121     ``redwood``                 ``r600``     dGPU
122     ``sumo``                    ``r600``     dGPU
123     **Radeon HD 6000 Series (Northern Islands)** [AMD-RADEON-HD-6000]_
124     -----------------------------------------------------------------------------------------------
125     ``barts``                   ``r600``     dGPU
126     ``caicos``                  ``r600``     dGPU
127     ``cayman``                  ``r600``     dGPU
128     ``turks``                   ``r600``     dGPU
129     **GCN GFX6 (Southern Islands (SI))** [AMD-GCN-GFX6]_
130     -----------------------------------------------------------------------------------------------
131     ``gfx600``  - ``tahiti``    ``amdgcn``   dGPU
132     ``gfx601``  - ``hainan``    ``amdgcn``   dGPU
133                 - ``oland``
134                 - ``pitcairn``
135                 - ``verde``
136     **GCN GFX7 (Sea Islands (CI))** [AMD-GCN-GFX7]_
137     -----------------------------------------------------------------------------------------------
138     ``gfx700``  - ``kaveri``    ``amdgcn``   APU                             - A6-7000
139                                                                              - A6 Pro-7050B
140                                                                              - A8-7100
141                                                                              - A8 Pro-7150B
142                                                                              - A10-7300
143                                                                              - A10 Pro-7350B
144                                                                              - FX-7500
145                                                                              - A8-7200P
146                                                                              - A10-7400P
147                                                                              - FX-7600P
148     ``gfx701``  - ``hawaii``    ``amdgcn``   dGPU                    ROCm    - FirePro W8100
149                                                                              - FirePro W9100
150                                                                              - FirePro S9150
151                                                                              - FirePro S9170
152     ``gfx702``                  ``amdgcn``   dGPU                    ROCm    - Radeon R9 290
153                                                                              - Radeon R9 290x
154                                                                              - Radeon R390
155                                                                              - Radeon R390x
156     ``gfx703``  - ``kabini``    ``amdgcn``   APU                             - E1-2100
157                 - ``mullins``                                                - E1-2200
158                                                                              - E1-2500
159                                                                              - E2-3000
160                                                                              - E2-3800
161                                                                              - A4-5000
162                                                                              - A4-5100
163                                                                              - A6-5200
164                                                                              - A4 Pro-3340B
165     ``gfx704``  - ``bonaire``   ``amdgcn``   dGPU                            - Radeon HD 7790
166                                                                              - Radeon HD 8770
167                                                                              - R7 260
168                                                                              - R7 260X
169     **GCN GFX8 (Volcanic Islands (VI))** [AMD-GCN-GFX8]_
170     -----------------------------------------------------------------------------------------------
171     ``gfx801``  - ``carrizo``   ``amdgcn``   APU   - xnack                   - A6-8500P
172                                                      [on]                    - Pro A6-8500B
173                                                                              - A8-8600P
174                                                                              - Pro A8-8600B
175                                                                              - FX-8800P
176                                                                              - Pro A12-8800B
177     \                           ``amdgcn``   APU   - xnack           ROCm    - A10-8700P
178                                                      [on]                    - Pro A10-8700B
179                                                                              - A10-8780P
180     \                           ``amdgcn``   APU   - xnack                   - A10-9600P
181                                                      [on]                    - A10-9630P
182                                                                              - A12-9700P
183                                                                              - A12-9730P
184                                                                              - FX-9800P
185                                                                              - FX-9830P
186     \                           ``amdgcn``   APU   - xnack                   - E2-9010
187                                                      [on]                    - A6-9210
188                                                                              - A9-9410
189     ``gfx802``  - ``iceland``   ``amdgcn``   dGPU  - xnack           ROCm    - FirePro S7150
190                 - ``tonga``                          [off]                   - FirePro S7100
191                                                                              - FirePro W7100
192                                                                              - Radeon R285
193                                                                              - Radeon R9 380
194                                                                              - Radeon R9 385
195                                                                              - Mobile FirePro
196                                                                                M7170
197     ``gfx803``  - ``fiji``      ``amdgcn``   dGPU  - xnack           ROCm    - Radeon R9 Nano
198                                                      [off]                   - Radeon R9 Fury
199                                                                              - Radeon R9 FuryX
200                                                                              - Radeon Pro Duo
201                                                                              - FirePro S9300x2
202                                                                              - Radeon Instinct MI8
203     \           - ``polaris10`` ``amdgcn``   dGPU  - xnack           ROCm    - Radeon RX 470
204                                                      [off]                   - Radeon RX 480
205                                                                              - Radeon Instinct MI6
206     \           - ``polaris11`` ``amdgcn``   dGPU  - xnack           ROCm    - Radeon RX 460
207                                                      [off]
208     ``gfx810``  - ``stoney``    ``amdgcn``   APU   - xnack
209                                                      [on]
210     **GCN GFX9** [AMD-GCN-GFX9]_
211     -----------------------------------------------------------------------------------------------
212     ``gfx900``                  ``amdgcn``   dGPU  - xnack           ROCm    - Radeon Vega
213                                                      [off]                     Frontier Edition
214                                                                              - Radeon RX Vega 56
215                                                                              - Radeon RX Vega 64
216                                                                              - Radeon RX Vega 64
217                                                                                Liquid
218                                                                              - Radeon Instinct MI25
219     ``gfx902``                  ``amdgcn``   APU   - xnack                   - Ryzen 3 2200G
220                                                      [on]                    - Ryzen 5 2400G
221     ``gfx904``                  ``amdgcn``   dGPU  - xnack                   *TBA*
222                                                      [off]
223                                                                              .. TODO::
224                                                                                 Add product
225                                                                                 names.
226     ``gfx906``                  ``amdgcn``   dGPU  - xnack                   - Radeon Instinct MI50
227                                                      [off]                   - Radeon Instinct MI60
228                                                    - sram-ecc                - Radeon VII
229                                                      [off]                   - Radeon Pro VII
230     ``gfx908``                  ``amdgcn``   dGPU  - xnack                   *TBA*
231                                                      [off]
232                                                    - sram-ecc
233                                                      [on]
234                                                                              .. TODO::
235                                                                                 Add product
236                                                                                 names.
237     ``gfx909``                  ``amdgcn``   APU   - xnack                   *TBA*
238                                                      [on]
239                                                                              .. TODO::
240                                                                                 Add product
241                                                                                 names.
242     **GCN GFX10** [AMD-GCN-GFX10]_
243     -----------------------------------------------------------------------------------------------
244     ``gfx1010``                 ``amdgcn``   dGPU  - xnack                   - Radeon RX 5700
245                                                      [off]                   - Radeon RX 5700 XT
246                                                    - wavefrontsize64         - Radeon Pro 5600 XT
247                                                      [off]
248                                                    - cumode
249                                                      [off]
250     ``gfx1011``                 ``amdgcn``   dGPU  - xnack                   - Radeon Pro 5600M
251                                                      [off]
252                                                    - wavefrontsize64
253                                                      [off]
254                                                    - cumode
255                                                      [off]
256     ``gfx1012``                 ``amdgcn``   dGPU  - xnack                   - Radeon RX 5500
257                                                      [off]                   - Radeon RX 5500 XT
258                                                    - wavefrontsize64
259                                                      [off]
260                                                    - cumode
261                                                      [off]
262     ``gfx1030``                 ``amdgcn``   dGPU  - wavefrontsize64         *TBA*
263                                                      [off]
264                                                    - cumode
265                                                      [off]
266                                                                              .. TODO
267                                                                                 Add product
268                                                                                 names.
269     ``gfx1031``                 ``amdgcn``   dGPU  - xnack                   *TBA*
270                                                      [off]
271                                                    - wavefrontsize64
272                                                      [off]
273                                                    - cumode
274                                                      [off]
275                                                                              .. TODO
276                                                                                 Add product
277                                                                                 names.
278     =========== =============== ============ ===== ================= ======= ======================
279
280.. _amdgpu-target-features:
281
282Target Features
283---------------
284
285Target features control how code is generated to support certain
286processor specific features. Not all target features are supported by
287all processors. The runtime must ensure that the features supported by
288the device used to execute the code match the features enabled when
289generating the code. A mismatch of features may result in incorrect
290execution, or a reduction in performance.
291
292The target features supported by each processor, and the default value
293used if not specified explicitly, is listed in
294:ref:`amdgpu-processor-table`.
295
296Use the ``clang -m[no-]<TargetFeature>`` option to specify the AMDGPU
297target features.
298
299For example:
300
301``-mxnack``
302  Enable the ``xnack`` feature.
303``-mno-xnack``
304  Disable the ``xnack`` feature.
305
306  .. table:: AMDGPU Target Features
307     :name: amdgpu-target-feature-table
308
309     ====================== ==================================================
310     Target Feature         Description
311     ====================== ==================================================
312     -m[no-]xnack           Enable/disable generating code that has
313                            memory clauses that are compatible with
314                            having XNACK replay enabled.
315
316                            This is used for demand paging and page
317                            migration. If XNACK replay is enabled in
318                            the device, then if a page fault occurs
319                            the code may execute incorrectly if the
320                            ``xnack`` feature is not enabled. Executing
321                            code that has the feature enabled on a
322                            device that does not have XNACK replay
323                            enabled will execute correctly but may
324                            be less performant than code with the
325                            feature disabled.
326
327     -m[no-]sram-ecc        Enable/disable generating code that assumes SRAM
328                            ECC is enabled/disabled.
329
330     -m[no-]wavefrontsize64 Control the default wavefront size used when
331                            generating code for kernels. When disabled
332                            native wavefront size 32 is used, when enabled
333                            wavefront size 64 is used.
334
335     -m[no-]cumode          Control the default wavefront execution mode used
336                            when generating code for kernels. When disabled
337                            native WGP wavefront execution mode is used,
338                            when enabled CU wavefront execution mode is used
339                            (see :ref:`amdgpu-amdhsa-memory-model`).
340     ====================== ==================================================
341
342.. _amdgpu-address-spaces:
343
344Address Spaces
345--------------
346
347The AMDGPU architecture supports a number of memory address spaces. The address
348space names use the OpenCL standard names, with some additions.
349
350The AMDGPU address spaces correspond to target architecture specific LLVM
351address space numbers used in LLVM IR.
352
353The AMDGPU address spaces are described in
354:ref:`amdgpu-address-spaces-table`. Only 64-bit process address spaces are
355supported for the ``amdgcn`` target.
356
357  .. table:: AMDGPU Address Spaces
358     :name: amdgpu-address-spaces-table
359
360     ================================= =============== =========== ================ ======= ============================
361     ..                                                                                     64-Bit Process Address Space
362     --------------------------------- --------------- ----------- ---------------- ------------------------------------
363     Address Space Name                LLVM IR Address HSA Segment Hardware         Address NULL Value
364                                       Space Number    Name        Name             Size
365     ================================= =============== =========== ================ ======= ============================
366     Generic                           0               flat        flat             64      0x0000000000000000
367     Global                            1               global      global           64      0x0000000000000000
368     Region                            2               N/A         GDS              32      *not implemented for AMDHSA*
369     Local                             3               group       LDS              32      0xFFFFFFFF
370     Constant                          4               constant    *same as global* 64      0x0000000000000000
371     Private                           5               private     scratch          32      0xFFFFFFFF
372     Constant 32-bit                   6               *TODO*                               0x00000000
373     Buffer Fat Pointer (experimental) 7               *TODO*
374     ================================= =============== =========== ================ ======= ============================
375
376**Generic**
377  The generic address space uses the hardware flat address support available in
378  GFX7-GFX10. This uses two fixed ranges of virtual addresses (the private and
379  local apertures), that are outside the range of addressable global memory, to
380  map from a flat address to a private or local address.
381
382  FLAT instructions can take a flat address and access global, private
383  (scratch), and group (LDS) memory depending on if the address is within one
384  of the aperture ranges. Flat access to scratch requires hardware aperture
385  setup and setup in the kernel prologue (see
386  :ref:`amdgpu-amdhsa-kernel-prolog-flat-scratch`). Flat access to LDS requires
387  hardware aperture setup and M0 (GFX7-GFX8) register setup (see
388  :ref:`amdgpu-amdhsa-kernel-prolog-m0`).
389
390  To convert between a private or group address space address (termed a segment
391  address) and a flat address the base address of the corresponding aperture
392  can be used. For GFX7-GFX8 these are available in the
393  :ref:`amdgpu-amdhsa-hsa-aql-queue` the address of which can be obtained with
394  Queue Ptr SGPR (see :ref:`amdgpu-amdhsa-initial-kernel-execution-state`). For
395  GFX9-GFX10 the aperture base addresses are directly available as inline
396  constant registers ``SRC_SHARED_BASE/LIMIT`` and ``SRC_PRIVATE_BASE/LIMIT``.
397  In 64-bit address mode the aperture sizes are 2^32 bytes and the base is
398  aligned to 2^32 which makes it easier to convert from flat to segment or
399  segment to flat.
400
401  A global address space address has the same value when used as a flat address
402  so no conversion is needed.
403
404**Global and Constant**
405  The global and constant address spaces both use global virtual addresses,
406  which are the same virtual address space used by the CPU. However, some
407  virtual addresses may only be accessible to the CPU, some only accessible
408  by the GPU, and some by both.
409
410  Using the constant address space indicates that the data will not change
411  during the execution of the kernel. This allows scalar read instructions to
412  be used. The vector and scalar L1 caches are invalidated of volatile data
413  before each kernel dispatch execution to allow constant memory to change
414  values between kernel dispatches.
415
416**Region**
417  The region address space uses the hardware Global Data Store (GDS). All
418  wavefronts executing on the same device will access the same memory for any
419  given region address. However, the same region address accessed by wavefronts
420  executing on different devices will access different memory. It is higher
421  performance than global memory. It is allocated by the runtime. The data
422  store (DS) instructions can be used to access it.
423
424**Local**
425  The local address space uses the hardware Local Data Store (LDS) which is
426  automatically allocated when the hardware creates the wavefronts of a
427  work-group, and freed when all the wavefronts of a work-group have
428  terminated. All wavefronts belonging to the same work-group will access the
429  same memory for any given local address. However, the same local address
430  accessed by wavefronts belonging to different work-groups will access
431  different memory. It is higher performance than global memory. The data store
432  (DS) instructions can be used to access it.
433
434**Private**
435  The private address space uses the hardware scratch memory support which
436  automatically allocates memory when it creates a wavefront and frees it when
437  a wavefronts terminates. The memory accessed by a lane of a wavefront for any
438  given private address will be different to the memory accessed by another lane
439  of the same or different wavefront for the same private address.
440
441  If a kernel dispatch uses scratch, then the hardware allocates memory from a
442  pool of backing memory allocated by the runtime for each wavefront. The lanes
443  of the wavefront access this using dword (4 byte) interleaving. The mapping
444  used from private address to backing memory address is:
445
446    ``wavefront-scratch-base +
447    ((private-address / 4) * wavefront-size * 4) +
448    (wavefront-lane-id * 4) + (private-address % 4)``
449
450  If each lane of a wavefront accesses the same private address, the
451  interleaving results in adjacent dwords being accessed and hence requires
452  fewer cache lines to be fetched.
453
454  There are different ways that the wavefront scratch base address is
455  determined by a wavefront (see
456  :ref:`amdgpu-amdhsa-initial-kernel-execution-state`).
457
458  Scratch memory can be accessed in an interleaved manner using buffer
459  instructions with the scratch buffer descriptor and per wavefront scratch
460  offset, by the scratch instructions, or by flat instructions. Multi-dword
461  access is not supported except by flat and scratch instructions in
462  GFX9-GFX10.
463
464**Constant 32-bit**
465  *TODO*
466
467**Buffer Fat Pointer**
468  The buffer fat pointer is an experimental address space that is currently
469  unsupported in the backend. It exposes a non-integral pointer that is in
470  the future intended to support the modelling of 128-bit buffer descriptors
471  plus a 32-bit offset into the buffer (in total encapsulating a 160-bit
472  *pointer*), allowing normal LLVM load/store/atomic operations to be used to
473  model the buffer descriptors used heavily in graphics workloads targeting
474  the backend.
475
476.. _amdgpu-memory-scopes:
477
478Memory Scopes
479-------------
480
481This section provides LLVM memory synchronization scopes supported by the AMDGPU
482backend memory model when the target triple OS is ``amdhsa`` (see
483:ref:`amdgpu-amdhsa-memory-model` and :ref:`amdgpu-target-triples`).
484
485The memory model supported is based on the HSA memory model [HSA]_ which is
486based in turn on HRF-indirect with scope inclusion [HRF]_. The happens-before
487relation is transitive over the synchronizes-with relation independent of scope
488and synchronizes-with allows the memory scope instances to be inclusive (see
489table :ref:`amdgpu-amdhsa-llvm-sync-scopes-table`).
490
491This is different to the OpenCL [OpenCL]_ memory model which does not have scope
492inclusion and requires the memory scopes to exactly match. However, this
493is conservatively correct for OpenCL.
494
495  .. table:: AMDHSA LLVM Sync Scopes
496     :name: amdgpu-amdhsa-llvm-sync-scopes-table
497
498     ======================= ===================================================
499     LLVM Sync Scope         Description
500     ======================= ===================================================
501     *none*                  The default: ``system``.
502
503                             Synchronizes with, and participates in modification
504                             and seq_cst total orderings with, other operations
505                             (except image operations) for all address spaces
506                             (except private, or generic that accesses private)
507                             provided the other operation's sync scope is:
508
509                             - ``system``.
510                             - ``agent`` and executed by a thread on the same
511                               agent.
512                             - ``workgroup`` and executed by a thread in the
513                               same work-group.
514                             - ``wavefront`` and executed by a thread in the
515                               same wavefront.
516
517     ``agent``               Synchronizes with, and participates in modification
518                             and seq_cst total orderings with, other operations
519                             (except image operations) for all address spaces
520                             (except private, or generic that accesses private)
521                             provided the other operation's sync scope is:
522
523                             - ``system`` or ``agent`` and executed by a thread
524                               on the same agent.
525                             - ``workgroup`` and executed by a thread in the
526                               same work-group.
527                             - ``wavefront`` and executed by a thread in the
528                               same wavefront.
529
530     ``workgroup``           Synchronizes with, and participates in modification
531                             and seq_cst total orderings with, other operations
532                             (except image operations) for all address spaces
533                             (except private, or generic that accesses private)
534                             provided the other operation's sync scope is:
535
536                             - ``system``, ``agent`` or ``workgroup`` and
537                               executed by a thread in the same work-group.
538                             - ``wavefront`` and executed by a thread in the
539                               same wavefront.
540
541     ``wavefront``           Synchronizes with, and participates in modification
542                             and seq_cst total orderings with, other operations
543                             (except image operations) for all address spaces
544                             (except private, or generic that accesses private)
545                             provided the other operation's sync scope is:
546
547                             - ``system``, ``agent``, ``workgroup`` or
548                               ``wavefront`` and executed by a thread in the
549                               same wavefront.
550
551     ``singlethread``        Only synchronizes with and participates in
552                             modification and seq_cst total orderings with,
553                             other operations (except image operations) running
554                             in the same thread for all address spaces (for
555                             example, in signal handlers).
556
557     ``one-as``              Same as ``system`` but only synchronizes with other
558                             operations within the same address space.
559
560     ``agent-one-as``        Same as ``agent`` but only synchronizes with other
561                             operations within the same address space.
562
563     ``workgroup-one-as``    Same as ``workgroup`` but only synchronizes with
564                             other operations within the same address space.
565
566     ``wavefront-one-as``    Same as ``wavefront`` but only synchronizes with
567                             other operations within the same address space.
568
569     ``singlethread-one-as`` Same as ``singlethread`` but only synchronizes with
570                             other operations within the same address space.
571     ======================= ===================================================
572
573LLVM IR Intrinsics
574------------------
575
576The AMDGPU backend implements the following LLVM IR intrinsics.
577
578*This section is WIP.*
579
580.. TODO::
581
582   List AMDGPU intrinsics.
583
584LLVM IR Attributes
585------------------
586
587The AMDGPU backend supports the following LLVM IR attributes.
588
589  .. table:: AMDGPU LLVM IR Attributes
590     :name: amdgpu-llvm-ir-attributes-table
591
592     ======================================= ==========================================================
593     LLVM Attribute                          Description
594     ======================================= ==========================================================
595     "amdgpu-flat-work-group-size"="min,max" Specify the minimum and maximum flat work group sizes that
596                                             will be specified when the kernel is dispatched. Generated
597                                             by the ``amdgpu_flat_work_group_size`` CLANG attribute [CLANG-ATTR]_.
598     "amdgpu-implicitarg-num-bytes"="n"      Number of kernel argument bytes to add to the kernel
599                                             argument block size for the implicit arguments. This
600                                             varies by OS and language (for OpenCL see
601                                             :ref:`opencl-kernel-implicit-arguments-appended-for-amdhsa-os-table`).
602     "amdgpu-num-sgpr"="n"                   Specifies the number of SGPRs to use. Generated by
603                                             the ``amdgpu_num_sgpr`` CLANG attribute [CLANG-ATTR]_.
604     "amdgpu-num-vgpr"="n"                   Specifies the number of VGPRs to use. Generated by the
605                                             ``amdgpu_num_vgpr`` CLANG attribute [CLANG-ATTR]_.
606     "amdgpu-waves-per-eu"="m,n"             Specify the minimum and maximum number of waves per
607                                             execution unit. Generated by the ``amdgpu_waves_per_eu``
608                                             CLANG attribute [CLANG-ATTR]_.
609     "amdgpu-ieee" true/false.               Specify whether the function expects the IEEE field of the
610                                             mode register to be set on entry. Overrides the default for
611                                             the calling convention.
612     "amdgpu-dx10-clamp" true/false.         Specify whether the function expects the DX10_CLAMP field of
613                                             the mode register to be set on entry. Overrides the default
614                                             for the calling convention.
615     ======================================= ==========================================================
616
617.. _amdgpu-elf-code-object:
618
619ELF Code Object
620===============
621
622The AMDGPU backend generates a standard ELF [ELF]_ relocatable code object that
623can be linked by ``lld`` to produce a standard ELF shared code object which can
624be loaded and executed on an AMDGPU target.
625
626.. _amdgpu-elf-header:
627
628Header
629------
630
631The AMDGPU backend uses the following ELF header:
632
633  .. table:: AMDGPU ELF Header
634     :name: amdgpu-elf-header-table
635
636     ========================== ===============================
637     Field                      Value
638     ========================== ===============================
639     ``e_ident[EI_CLASS]``      ``ELFCLASS64``
640     ``e_ident[EI_DATA]``       ``ELFDATA2LSB``
641     ``e_ident[EI_OSABI]``      - ``ELFOSABI_NONE``
642                                - ``ELFOSABI_AMDGPU_HSA``
643                                - ``ELFOSABI_AMDGPU_PAL``
644                                - ``ELFOSABI_AMDGPU_MESA3D``
645     ``e_ident[EI_ABIVERSION]`` - ``ELFABIVERSION_AMDGPU_HSA``
646                                - ``ELFABIVERSION_AMDGPU_PAL``
647                                - ``ELFABIVERSION_AMDGPU_MESA3D``
648     ``e_type``                 - ``ET_REL``
649                                - ``ET_DYN``
650     ``e_machine``              ``EM_AMDGPU``
651     ``e_entry``                0
652     ``e_flags``                See :ref:`amdgpu-elf-header-e_flags-table`
653     ========================== ===============================
654
655..
656
657  .. table:: AMDGPU ELF Header Enumeration Values
658     :name: amdgpu-elf-header-enumeration-values-table
659
660     =============================== =====
661     Name                            Value
662     =============================== =====
663     ``EM_AMDGPU``                   224
664     ``ELFOSABI_NONE``               0
665     ``ELFOSABI_AMDGPU_HSA``         64
666     ``ELFOSABI_AMDGPU_PAL``         65
667     ``ELFOSABI_AMDGPU_MESA3D``      66
668     ``ELFABIVERSION_AMDGPU_HSA``    1
669     ``ELFABIVERSION_AMDGPU_PAL``    0
670     ``ELFABIVERSION_AMDGPU_MESA3D`` 0
671     =============================== =====
672
673``e_ident[EI_CLASS]``
674  The ELF class is:
675
676  * ``ELFCLASS32`` for ``r600`` architecture.
677
678  * ``ELFCLASS64`` for ``amdgcn`` architecture which only supports 64-bit
679    process address space applications.
680
681``e_ident[EI_DATA]``
682  All AMDGPU targets use ``ELFDATA2LSB`` for little-endian byte ordering.
683
684``e_ident[EI_OSABI]``
685  One of the following AMDGPU target architecture specific OS ABIs
686  (see :ref:`amdgpu-os-table`):
687
688  * ``ELFOSABI_NONE`` for *unknown* OS.
689
690  * ``ELFOSABI_AMDGPU_HSA`` for ``amdhsa`` OS.
691
692  * ``ELFOSABI_AMDGPU_PAL`` for ``amdpal`` OS.
693
694  * ``ELFOSABI_AMDGPU_MESA3D`` for ``mesa3D`` OS.
695
696``e_ident[EI_ABIVERSION]``
697  The ABI version of the AMDGPU target architecture specific OS ABI to which the code
698  object conforms:
699
700  * ``ELFABIVERSION_AMDGPU_HSA`` is used to specify the version of AMD HSA
701    runtime ABI.
702
703  * ``ELFABIVERSION_AMDGPU_PAL`` is used to specify the version of AMD PAL
704    runtime ABI.
705
706  * ``ELFABIVERSION_AMDGPU_MESA3D`` is used to specify the version of AMD MESA
707    3D runtime ABI.
708
709``e_type``
710  Can be one of the following values:
711
712
713  ``ET_REL``
714    The type produced by the AMDGPU backend compiler as it is relocatable code
715    object.
716
717  ``ET_DYN``
718    The type produced by the linker as it is a shared code object.
719
720  The AMD HSA runtime loader requires a ``ET_DYN`` code object.
721
722``e_machine``
723  The value ``EM_AMDGPU`` is used for the machine for all processors supported
724  by the ``r600`` and ``amdgcn`` architectures (see
725  :ref:`amdgpu-processor-table`). The specific processor is specified in the
726  ``EF_AMDGPU_MACH`` bit field of the ``e_flags`` (see
727  :ref:`amdgpu-elf-header-e_flags-table`).
728
729``e_entry``
730  The entry point is 0 as the entry points for individual kernels must be
731  selected in order to invoke them through AQL packets.
732
733``e_flags``
734  The AMDGPU backend uses the following ELF header flags:
735
736  .. table:: AMDGPU ELF Header ``e_flags``
737     :name: amdgpu-elf-header-e_flags-table
738
739     ================================= ========== =============================
740     Name                              Value      Description
741     ================================= ========== =============================
742     **AMDGPU Processor Flag**                    See :ref:`amdgpu-processor-table`.
743     -------------------------------------------- -----------------------------
744     ``EF_AMDGPU_MACH``                0x000000ff AMDGPU processor selection
745                                                  mask for
746                                                  ``EF_AMDGPU_MACH_xxx`` values
747                                                  defined in
748                                                  :ref:`amdgpu-ef-amdgpu-mach-table`.
749     ``EF_AMDGPU_XNACK``               0x00000100 Indicates if the ``xnack``
750                                                  target feature is
751                                                  enabled for all code
752                                                  contained in the code object.
753                                                  If the processor
754                                                  does not support the
755                                                  ``xnack`` target
756                                                  feature then must
757                                                  be 0.
758                                                  See
759                                                  :ref:`amdgpu-target-features`.
760     ``EF_AMDGPU_SRAM_ECC``            0x00000200 Indicates if the ``sram-ecc``
761                                                  target feature is
762                                                  enabled for all code
763                                                  contained in the code object.
764                                                  If the processor
765                                                  does not support the
766                                                  ``sram-ecc`` target
767                                                  feature then must
768                                                  be 0.
769                                                  See
770                                                  :ref:`amdgpu-target-features`.
771     ================================= ========== =============================
772
773  .. table:: AMDGPU ``EF_AMDGPU_MACH`` Values
774     :name: amdgpu-ef-amdgpu-mach-table
775
776     ================================= ========== =============================
777     Name                              Value      Description (see
778                                                  :ref:`amdgpu-processor-table`)
779     ================================= ========== =============================
780     ``EF_AMDGPU_MACH_NONE``           0x000      *not specified*
781     ``EF_AMDGPU_MACH_R600_R600``      0x001      ``r600``
782     ``EF_AMDGPU_MACH_R600_R630``      0x002      ``r630``
783     ``EF_AMDGPU_MACH_R600_RS880``     0x003      ``rs880``
784     ``EF_AMDGPU_MACH_R600_RV670``     0x004      ``rv670``
785     ``EF_AMDGPU_MACH_R600_RV710``     0x005      ``rv710``
786     ``EF_AMDGPU_MACH_R600_RV730``     0x006      ``rv730``
787     ``EF_AMDGPU_MACH_R600_RV770``     0x007      ``rv770``
788     ``EF_AMDGPU_MACH_R600_CEDAR``     0x008      ``cedar``
789     ``EF_AMDGPU_MACH_R600_CYPRESS``   0x009      ``cypress``
790     ``EF_AMDGPU_MACH_R600_JUNIPER``   0x00a      ``juniper``
791     ``EF_AMDGPU_MACH_R600_REDWOOD``   0x00b      ``redwood``
792     ``EF_AMDGPU_MACH_R600_SUMO``      0x00c      ``sumo``
793     ``EF_AMDGPU_MACH_R600_BARTS``     0x00d      ``barts``
794     ``EF_AMDGPU_MACH_R600_CAICOS``    0x00e      ``caicos``
795     ``EF_AMDGPU_MACH_R600_CAYMAN``    0x00f      ``cayman``
796     ``EF_AMDGPU_MACH_R600_TURKS``     0x010      ``turks``
797     *reserved*                        0x011 -    Reserved for ``r600``
798                                       0x01f      architecture processors.
799     ``EF_AMDGPU_MACH_AMDGCN_GFX600``  0x020      ``gfx600``
800     ``EF_AMDGPU_MACH_AMDGCN_GFX601``  0x021      ``gfx601``
801     ``EF_AMDGPU_MACH_AMDGCN_GFX700``  0x022      ``gfx700``
802     ``EF_AMDGPU_MACH_AMDGCN_GFX701``  0x023      ``gfx701``
803     ``EF_AMDGPU_MACH_AMDGCN_GFX702``  0x024      ``gfx702``
804     ``EF_AMDGPU_MACH_AMDGCN_GFX703``  0x025      ``gfx703``
805     ``EF_AMDGPU_MACH_AMDGCN_GFX704``  0x026      ``gfx704``
806     *reserved*                        0x027      Reserved.
807     ``EF_AMDGPU_MACH_AMDGCN_GFX801``  0x028      ``gfx801``
808     ``EF_AMDGPU_MACH_AMDGCN_GFX802``  0x029      ``gfx802``
809     ``EF_AMDGPU_MACH_AMDGCN_GFX803``  0x02a      ``gfx803``
810     ``EF_AMDGPU_MACH_AMDGCN_GFX810``  0x02b      ``gfx810``
811     ``EF_AMDGPU_MACH_AMDGCN_GFX900``  0x02c      ``gfx900``
812     ``EF_AMDGPU_MACH_AMDGCN_GFX902``  0x02d      ``gfx902``
813     ``EF_AMDGPU_MACH_AMDGCN_GFX904``  0x02e      ``gfx904``
814     ``EF_AMDGPU_MACH_AMDGCN_GFX906``  0x02f      ``gfx906``
815     ``EF_AMDGPU_MACH_AMDGCN_GFX908``  0x030      ``gfx908``
816     ``EF_AMDGPU_MACH_AMDGCN_GFX909``  0x031      ``gfx909``
817     *reserved*                        0x032      Reserved.
818     ``EF_AMDGPU_MACH_AMDGCN_GFX1010`` 0x033      ``gfx1010``
819     ``EF_AMDGPU_MACH_AMDGCN_GFX1011`` 0x034      ``gfx1011``
820     ``EF_AMDGPU_MACH_AMDGCN_GFX1012`` 0x035      ``gfx1012``
821     ``EF_AMDGPU_MACH_AMDGCN_GFX1030`` 0x036      ``gfx1030``
822     ``EF_AMDGPU_MACH_AMDGCN_GFX1031`` 0x037      ``gfx1031``
823     ================================= ========== =============================
824
825Sections
826--------
827
828An AMDGPU target ELF code object has the standard ELF sections which include:
829
830  .. table:: AMDGPU ELF Sections
831     :name: amdgpu-elf-sections-table
832
833     ================== ================ =================================
834     Name               Type             Attributes
835     ================== ================ =================================
836     ``.bss``           ``SHT_NOBITS``   ``SHF_ALLOC`` + ``SHF_WRITE``
837     ``.data``          ``SHT_PROGBITS`` ``SHF_ALLOC`` + ``SHF_WRITE``
838     ``.debug_``\ *\**  ``SHT_PROGBITS`` *none*
839     ``.dynamic``       ``SHT_DYNAMIC``  ``SHF_ALLOC``
840     ``.dynstr``        ``SHT_PROGBITS`` ``SHF_ALLOC``
841     ``.dynsym``        ``SHT_PROGBITS`` ``SHF_ALLOC``
842     ``.got``           ``SHT_PROGBITS`` ``SHF_ALLOC`` + ``SHF_WRITE``
843     ``.hash``          ``SHT_HASH``     ``SHF_ALLOC``
844     ``.note``          ``SHT_NOTE``     *none*
845     ``.rela``\ *name*  ``SHT_RELA``     *none*
846     ``.rela.dyn``      ``SHT_RELA``     *none*
847     ``.rodata``        ``SHT_PROGBITS`` ``SHF_ALLOC``
848     ``.shstrtab``      ``SHT_STRTAB``   *none*
849     ``.strtab``        ``SHT_STRTAB``   *none*
850     ``.symtab``        ``SHT_SYMTAB``   *none*
851     ``.text``          ``SHT_PROGBITS`` ``SHF_ALLOC`` + ``SHF_EXECINSTR``
852     ================== ================ =================================
853
854These sections have their standard meanings (see [ELF]_) and are only generated
855if needed.
856
857``.debug``\ *\**
858  The standard DWARF sections. See :ref:`amdgpu-dwarf-debug-information` for
859  information on the DWARF produced by the AMDGPU backend.
860
861``.dynamic``, ``.dynstr``, ``.dynsym``, ``.hash``
862  The standard sections used by a dynamic loader.
863
864``.note``
865  See :ref:`amdgpu-note-records` for the note records supported by the AMDGPU
866  backend.
867
868``.rela``\ *name*, ``.rela.dyn``
869  For relocatable code objects, *name* is the name of the section that the
870  relocation records apply. For example, ``.rela.text`` is the section name for
871  relocation records associated with the ``.text`` section.
872
873  For linked shared code objects, ``.rela.dyn`` contains all the relocation
874  records from each of the relocatable code object's ``.rela``\ *name* sections.
875
876  See :ref:`amdgpu-relocation-records` for the relocation records supported by
877  the AMDGPU backend.
878
879``.text``
880  The executable machine code for the kernels and functions they call. Generated
881  as position independent code. See :ref:`amdgpu-code-conventions` for
882  information on conventions used in the isa generation.
883
884.. _amdgpu-note-records:
885
886Note Records
887------------
888
889The AMDGPU backend code object contains ELF note records in the ``.note``
890section. The set of generated notes and their semantics depend on the code
891object version; see :ref:`amdgpu-note-records-v2` and
892:ref:`amdgpu-note-records-v3`.
893
894As required by ``ELFCLASS32`` and ``ELFCLASS64``, minimal zero-byte padding
895must be generated after the ``name`` field to ensure the ``desc`` field is 4
896byte aligned. In addition, minimal zero-byte padding must be generated to
897ensure the ``desc`` field size is a multiple of 4 bytes. The ``sh_addralign``
898field of the ``.note`` section must be at least 4 to indicate at least 8 byte
899alignment.
900
901.. _amdgpu-note-records-v2:
902
903Code Object V2 Note Records (-mattr=-code-object-v3)
904~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
905
906.. warning:: Code Object V2 is not the default code object version emitted by
907  this version of LLVM. For a description of the notes generated with the
908  default configuration (Code Object V3) see :ref:`amdgpu-note-records-v3`.
909
910The AMDGPU backend code object uses the following ELF note record in the
911``.note`` section when compiling for Code Object V2 (-mattr=-code-object-v3).
912
913Additional note records may be present, but any which are not documented here
914are deprecated and should not be used.
915
916  .. table:: AMDGPU Code Object V2 ELF Note Records
917     :name: amdgpu-elf-note-records-table-v2
918
919     ===== ============================== ======================================
920     Name  Type                           Description
921     ===== ============================== ======================================
922     "AMD" ``NT_AMD_AMDGPU_HSA_METADATA`` <metadata null terminated string>
923     ===== ============================== ======================================
924
925..
926
927  .. table:: AMDGPU Code Object V2 ELF Note Record Enumeration Values
928     :name: amdgpu-elf-note-record-enumeration-values-table-v2
929
930     ============================== =====
931     Name                           Value
932     ============================== =====
933     *reserved*                       0-9
934     ``NT_AMD_AMDGPU_HSA_METADATA``    10
935     *reserved*                        11
936     ============================== =====
937
938``NT_AMD_AMDGPU_HSA_METADATA``
939  Specifies extensible metadata associated with the code objects executed on HSA
940  [HSA]_ compatible runtimes such as AMD's ROCm [AMD-ROCm]_. It is required when
941  the target triple OS is ``amdhsa`` (see :ref:`amdgpu-target-triples`). See
942  :ref:`amdgpu-amdhsa-code-object-metadata-v2` for the syntax of the code
943  object metadata string.
944
945.. _amdgpu-note-records-v3:
946
947Code Object V3 Note Records (-mattr=+code-object-v3)
948~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
949
950The AMDGPU backend code object uses the following ELF note record in the
951``.note`` section when compiling for Code Object V3 (-mattr=+code-object-v3).
952
953Additional note records may be present, but any which are not documented here
954are deprecated and should not be used.
955
956  .. table:: AMDGPU Code Object V3 ELF Note Records
957     :name: amdgpu-elf-note-records-table-v3
958
959     ======== ============================== ======================================
960     Name     Type                           Description
961     ======== ============================== ======================================
962     "AMDGPU" ``NT_AMDGPU_METADATA``         Metadata in Message Pack [MsgPack]_
963                                             binary format.
964     ======== ============================== ======================================
965
966..
967
968  .. table:: AMDGPU Code Object V3 ELF Note Record Enumeration Values
969     :name: amdgpu-elf-note-record-enumeration-values-table-v3
970
971     ============================== =====
972     Name                           Value
973     ============================== =====
974     *reserved*                     0-31
975     ``NT_AMDGPU_METADATA``         32
976     ============================== =====
977
978``NT_AMDGPU_METADATA``
979  Specifies extensible metadata associated with an AMDGPU code
980  object. It is encoded as a map in the Message Pack [MsgPack]_ binary
981  data format. See :ref:`amdgpu-amdhsa-code-object-metadata-v3` for the
982  map keys defined for the ``amdhsa`` OS.
983
984.. _amdgpu-symbols:
985
986Symbols
987-------
988
989Symbols include the following:
990
991  .. table:: AMDGPU ELF Symbols
992     :name: amdgpu-elf-symbols-table
993
994     ===================== ================== ================ ==================
995     Name                  Type               Section          Description
996     ===================== ================== ================ ==================
997     *link-name*           ``STT_OBJECT``     - ``.data``      Global variable
998                                              - ``.rodata``
999                                              - ``.bss``
1000     *link-name*\ ``.kd``  ``STT_OBJECT``     - ``.rodata``    Kernel descriptor
1001     *link-name*           ``STT_FUNC``       - ``.text``      Kernel entry point
1002     *link-name*           ``STT_OBJECT``     - SHN_AMDGPU_LDS Global variable in LDS
1003     ===================== ================== ================ ==================
1004
1005Global variable
1006  Global variables both used and defined by the compilation unit.
1007
1008  If the symbol is defined in the compilation unit then it is allocated in the
1009  appropriate section according to if it has initialized data or is readonly.
1010
1011  If the symbol is external then its section is ``STN_UNDEF`` and the loader
1012  will resolve relocations using the definition provided by another code object
1013  or explicitly defined by the runtime.
1014
1015  If the symbol resides in local/group memory (LDS) then its section is the
1016  special processor specific section name ``SHN_AMDGPU_LDS``, and the
1017  ``st_value`` field describes alignment requirements as it does for common
1018  symbols.
1019
1020  .. TODO::
1021
1022     Add description of linked shared object symbols. Seems undefined symbols
1023     are marked as STT_NOTYPE.
1024
1025Kernel descriptor
1026  Every HSA kernel has an associated kernel descriptor. It is the address of the
1027  kernel descriptor that is used in the AQL dispatch packet used to invoke the
1028  kernel, not the kernel entry point. The layout of the HSA kernel descriptor is
1029  defined in :ref:`amdgpu-amdhsa-kernel-descriptor`.
1030
1031Kernel entry point
1032  Every HSA kernel also has a symbol for its machine code entry point.
1033
1034.. _amdgpu-relocation-records:
1035
1036Relocation Records
1037------------------
1038
1039AMDGPU backend generates ``Elf64_Rela`` relocation records. Supported
1040relocatable fields are:
1041
1042``word32``
1043  This specifies a 32-bit field occupying 4 bytes with arbitrary byte
1044  alignment. These values use the same byte order as other word values in the
1045  AMDGPU architecture.
1046
1047``word64``
1048  This specifies a 64-bit field occupying 8 bytes with arbitrary byte
1049  alignment. These values use the same byte order as other word values in the
1050  AMDGPU architecture.
1051
1052Following notations are used for specifying relocation calculations:
1053
1054**A**
1055  Represents the addend used to compute the value of the relocatable field.
1056
1057**G**
1058  Represents the offset into the global offset table at which the relocation
1059  entry's symbol will reside during execution.
1060
1061**GOT**
1062  Represents the address of the global offset table.
1063
1064**P**
1065  Represents the place (section offset for ``et_rel`` or address for ``et_dyn``)
1066  of the storage unit being relocated (computed using ``r_offset``).
1067
1068**S**
1069  Represents the value of the symbol whose index resides in the relocation
1070  entry. Relocations not using this must specify a symbol index of
1071  ``STN_UNDEF``.
1072
1073**B**
1074  Represents the base address of a loaded executable or shared object which is
1075  the difference between the ELF address and the actual load address.
1076  Relocations using this are only valid in executable or shared objects.
1077
1078The following relocation types are supported:
1079
1080  .. table:: AMDGPU ELF Relocation Records
1081     :name: amdgpu-elf-relocation-records-table
1082
1083     ========================== ======= =====  ==========  ==============================
1084     Relocation Type            Kind    Value  Field       Calculation
1085     ========================== ======= =====  ==========  ==============================
1086     ``R_AMDGPU_NONE``                  0      *none*      *none*
1087     ``R_AMDGPU_ABS32_LO``      Static, 1      ``word32``  (S + A) & 0xFFFFFFFF
1088                                Dynamic
1089     ``R_AMDGPU_ABS32_HI``      Static, 2      ``word32``  (S + A) >> 32
1090                                Dynamic
1091     ``R_AMDGPU_ABS64``         Static, 3      ``word64``  S + A
1092                                Dynamic
1093     ``R_AMDGPU_REL32``         Static  4      ``word32``  S + A - P
1094     ``R_AMDGPU_REL64``         Static  5      ``word64``  S + A - P
1095     ``R_AMDGPU_ABS32``         Static, 6      ``word32``  S + A
1096                                Dynamic
1097     ``R_AMDGPU_GOTPCREL``      Static  7      ``word32``  G + GOT + A - P
1098     ``R_AMDGPU_GOTPCREL32_LO`` Static  8      ``word32``  (G + GOT + A - P) & 0xFFFFFFFF
1099     ``R_AMDGPU_GOTPCREL32_HI`` Static  9      ``word32``  (G + GOT + A - P) >> 32
1100     ``R_AMDGPU_REL32_LO``      Static  10     ``word32``  (S + A - P) & 0xFFFFFFFF
1101     ``R_AMDGPU_REL32_HI``      Static  11     ``word32``  (S + A - P) >> 32
1102     *reserved*                         12
1103     ``R_AMDGPU_RELATIVE64``    Dynamic 13     ``word64``  B + A
1104     ========================== ======= =====  ==========  ==============================
1105
1106``R_AMDGPU_ABS32_LO`` and ``R_AMDGPU_ABS32_HI`` are only supported by
1107the ``mesa3d`` OS, which does not support ``R_AMDGPU_ABS64``.
1108
1109There is no current OS loader support for 32-bit programs and so
1110``R_AMDGPU_ABS32`` is not used.
1111
1112.. _amdgpu-loaded-code-object-path-uniform-resource-identifier:
1113
1114Loaded Code Object Path Uniform Resource Identifier (URI)
1115---------------------------------------------------------
1116
1117The AMD GPU code object loader represents the path of the ELF shared object from
1118which the code object was loaded as a textual Unifom Resource Identifier (URI).
1119Note that the code object is the in memory loaded relocated form of the ELF
1120shared object.  Multiple code objects may be loaded at different memory
1121addresses in the same process from the same ELF shared object.
1122
1123The loaded code object path URI syntax is defined by the following BNF syntax:
1124
1125.. code::
1126
1127  code_object_uri ::== file_uri | memory_uri
1128  file_uri        ::== "file://" file_path [ range_specifier ]
1129  memory_uri      ::== "memory://" process_id range_specifier
1130  range_specifier ::== [ "#" | "?" ] "offset=" number "&" "size=" number
1131  file_path       ::== URI_ENCODED_OS_FILE_PATH
1132  process_id      ::== DECIMAL_NUMBER
1133  number          ::== HEX_NUMBER | DECIMAL_NUMBER | OCTAL_NUMBER
1134
1135**number**
1136  Is a C integral literal where hexadecimal values are prefixed by "0x" or "0X",
1137  and octal values by "0".
1138
1139**file_path**
1140  Is the file's path specified as a URI encoded UTF-8 string. In URI encoding,
1141  every character that is not in the regular expression ``[a-zA-Z0-9/_.~-]`` is
1142  encoded as two uppercase hexidecimal digits proceeded by "%".  Directories in
1143  the path are separated by "/".
1144
1145**offset**
1146  Is a 0-based byte offset to the start of the code object.  For a file URI, it
1147  is from the start of the file specified by the ``file_path``, and if omitted
1148  defaults to 0. For a memory URI, it is the memory address and is required.
1149
1150**size**
1151  Is the number of bytes in the code object.  For a file URI, if omitted it
1152  defaults to the size of the file.  It is required for a memory URI.
1153
1154**process_id**
1155  Is the identity of the process owning the memory.  For Linux it is the C
1156  unsigned integral decimal literal for the process ID (PID).
1157
1158For example:
1159
1160.. code::
1161
1162  file:///dir1/dir2/file1
1163  file:///dir3/dir4/file2#offset=0x2000&size=3000
1164  memory://1234#offset=0x20000&size=3000
1165
1166.. _amdgpu-dwarf-debug-information:
1167
1168DWARF Debug Information
1169=======================
1170
1171.. warning::
1172
1173   This section describes **provisional support** for AMDGPU DWARF [DWARF]_ that
1174   is not currently fully implemented and is subject to change.
1175
1176AMDGPU generates DWARF [DWARF]_ debugging information ELF sections (see
1177:ref:`amdgpu-elf-code-object`) which contain information that maps the code
1178object executable code and data to the source language constructs. It can be
1179used by tools such as debuggers and profilers. It uses features defined in
1180:doc:`AMDGPUDwarfExtensionsForHeterogeneousDebugging` that are made available in
1181DWARF Version 4 and DWARF Version 5 as an LLVM vendor extension.
1182
1183This section defines the AMDGPU target architecture specific DWARF mappings.
1184
1185.. _amdgpu-dwarf-register-identifier:
1186
1187Register Identifier
1188-------------------
1189
1190This section defines the AMDGPU target architecture register numbers used in
1191DWARF operation expressions (see DWARF Version 5 section 2.5 and
1192:ref:`amdgpu-dwarf-operation-expressions`) and Call Frame Information
1193instructions (see DWARF Version 5 section 6.4 and
1194:ref:`amdgpu-dwarf-call-frame-information`).
1195
1196A single code object can contain code for kernels that have different wavefront
1197sizes. The vector registers and some scalar registers are based on the wavefront
1198size. AMDGPU defines distinct DWARF registers for each wavefront size. This
1199simplifies the consumer of the DWARF so that each register has a fixed size,
1200rather than being dynamic according to the wavefront size mode. Similarly,
1201distinct DWARF registers are defined for those registers that vary in size
1202according to the process address size. This allows a consumer to treat a
1203specific AMDGPU processor as a single architecture regardless of how it is
1204configured at run time. The compiler explicitly specifies the DWARF registers
1205that match the mode in which the code it is generating will be executed.
1206
1207DWARF registers are encoded as numbers, which are mapped to architecture
1208registers. The mapping for AMDGPU is defined in
1209:ref:`amdgpu-dwarf-register-mapping-table`. All AMDGPU targets use the same
1210mapping.
1211
1212.. table:: AMDGPU DWARF Register Mapping
1213   :name: amdgpu-dwarf-register-mapping-table
1214
1215   ============== ================= ======== ==================================
1216   DWARF Register AMDGPU Register   Bit Size Description
1217   ============== ================= ======== ==================================
1218   0              PC_32             32       Program Counter (PC) when
1219                                             executing in a 32-bit process
1220                                             address space. Used in the CFI to
1221                                             describe the PC of the calling
1222                                             frame.
1223   1              EXEC_MASK_32      32       Execution Mask Register when
1224                                             executing in wavefront 32 mode.
1225   2-15           *Reserved*                 *Reserved for highly accessed
1226                                             registers using DWARF shortcut.*
1227   16             PC_64             64       Program Counter (PC) when
1228                                             executing in a 64-bit process
1229                                             address space. Used in the CFI to
1230                                             describe the PC of the calling
1231                                             frame.
1232   17             EXEC_MASK_64      64       Execution Mask Register when
1233                                             executing in wavefront 64 mode.
1234   18-31          *Reserved*                 *Reserved for highly accessed
1235                                             registers using DWARF shortcut.*
1236   32-95          SGPR0-SGPR63      32       Scalar General Purpose
1237                                             Registers.
1238   96-127         *Reserved*                 *Reserved for frequently accessed
1239                                             registers using DWARF 1-byte ULEB.*
1240   128            SCC               32       Scalar Condition Code Register.
1241   129-511        *Reserved*                 *Reserved for future Scalar
1242                                             Architectural Registers.*
1243   512            VCC_32            32       Vector Condition Code Register
1244                                             when executing in wavefront 32
1245                                             mode.
1246   513-1023       *Reserved*                 *Reserved for future Vector
1247                                             Architectural Registers when
1248                                             executing in wavefront 32 mode.*
1249   768            VCC_64            32       Vector Condition Code Register
1250                                             when executing in wavefront 64
1251                                             mode.
1252   769-1023       *Reserved*                 *Reserved for future Vector
1253                                             Architectural Registers when
1254                                             executing in wavefront 64 mode.*
1255   1024-1087      *Reserved*                 *Reserved for padding.*
1256   1088-1129      SGPR64-SGPR105    32       Scalar General Purpose Registers.
1257   1130-1535      *Reserved*                 *Reserved for future Scalar
1258                                             General Purpose Registers.*
1259   1536-1791      VGPR0-VGPR255     32*32    Vector General Purpose Registers
1260                                             when executing in wavefront 32
1261                                             mode.
1262   1792-2047      *Reserved*                 *Reserved for future Vector
1263                                             General Purpose Registers when
1264                                             executing in wavefront 32 mode.*
1265   2048-2303      AGPR0-AGPR255     32*32    Vector Accumulation Registers
1266                                             when executing in wavefront 32
1267                                             mode.
1268   2304-2559      *Reserved*                 *Reserved for future Vector
1269                                             Accumulation Registers when
1270                                             executing in wavefront 32 mode.*
1271   2560-2815      VGPR0-VGPR255     64*32    Vector General Purpose Registers
1272                                             when executing in wavefront 64
1273                                             mode.
1274   2816-3071      *Reserved*                 *Reserved for future Vector
1275                                             General Purpose Registers when
1276                                             executing in wavefront 64 mode.*
1277   3072-3327      AGPR0-AGPR255     64*32    Vector Accumulation Registers
1278                                             when executing in wavefront 64
1279                                             mode.
1280   3328-3583      *Reserved*                 *Reserved for future Vector
1281                                             Accumulation Registers when
1282                                             executing in wavefront 64 mode.*
1283   ============== ================= ======== ==================================
1284
1285The vector registers are represented as the full size for the wavefront. They
1286are organized as consecutive dwords (32-bits), one per lane, with the dword at
1287the least significant bit position corresponding to lane 0 and so forth. DWARF
1288location expressions involving the ``DW_OP_LLVM_offset`` and
1289``DW_OP_LLVM_push_lane`` operations are used to select the part of the vector
1290register corresponding to the lane that is executing the current thread of
1291execution in languages that are implemented using a SIMD or SIMT execution
1292model.
1293
1294If the wavefront size is 32 lanes then the wavefront 32 mode register
1295definitions are used. If the wavefront size is 64 lanes then the wavefront 64
1296mode register definitions are used. Some AMDGPU targets support executing in
1297both wavefront 32 and wavefront 64 mode. The register definitions corresponding
1298to the wavefront mode of the generated code will be used.
1299
1300If code is generated to execute in a 32-bit process address space, then the
130132-bit process address space register definitions are used. If code is generated
1302to execute in a 64-bit process address space, then the 64-bit process address
1303space register definitions are used. The ``amdgcn`` target only supports the
130464-bit process address space.
1305
1306.. _amdgpu-dwarf-address-class-identifier:
1307
1308Address Class Identifier
1309------------------------
1310
1311The DWARF address class represents the source language memory space. See DWARF
1312Version 5 section 2.12 which is updated by the *DWARF Extensions For
1313Heterogeneous Debugging* section :ref:`amdgpu-dwarf-segment_addresses`.
1314
1315The DWARF address class mapping used for AMDGPU is defined in
1316:ref:`amdgpu-dwarf-address-class-mapping-table`.
1317
1318.. table:: AMDGPU DWARF Address Class Mapping
1319   :name: amdgpu-dwarf-address-class-mapping-table
1320
1321   ========================= ====== =================
1322   DWARF                            AMDGPU
1323   -------------------------------- -----------------
1324   Address Class Name        Value  Address Space
1325   ========================= ====== =================
1326   ``DW_ADDR_none``          0x0000 Generic (Flat)
1327   ``DW_ADDR_LLVM_global``   0x0001 Global
1328   ``DW_ADDR_LLVM_constant`` 0x0002 Global
1329   ``DW_ADDR_LLVM_group``    0x0003 Local (group/LDS)
1330   ``DW_ADDR_LLVM_private``  0x0004 Private (Scratch)
1331   ``DW_ADDR_AMDGPU_region`` 0x8000 Region (GDS)
1332   ========================= ====== =================
1333
1334The DWARF address class values defined in the *DWARF Extensions For
1335Heterogeneous Debugging* section :ref:`amdgpu-dwarf-segment_addresses` are used.
1336
1337In addition, ``DW_ADDR_AMDGPU_region`` is encoded as a vendor extension. This is
1338available for use for the AMD extension for access to the hardware GDS memory
1339which is scratchpad memory allocated per device.
1340
1341For AMDGPU if no ``DW_AT_address_class`` attribute is present, then the default
1342address class of ``DW_ADDR_none`` is used.
1343
1344See :ref:`amdgpu-dwarf-address-space-identifier` for information on the AMDGPU
1345mapping of DWARF address classes to DWARF address spaces, including address size
1346and NULL value.
1347
1348.. _amdgpu-dwarf-address-space-identifier:
1349
1350Address Space Identifier
1351------------------------
1352
1353DWARF address spaces correspond to target architecture specific linear
1354addressable memory areas. See DWARF Version 5 section 2.12 and *DWARF Extensions
1355For Heterogeneous Debugging* section :ref:`amdgpu-dwarf-segment_addresses`.
1356
1357The DWARF address space mapping used for AMDGPU is defined in
1358:ref:`amdgpu-dwarf-address-space-mapping-table`.
1359
1360.. table:: AMDGPU DWARF Address Space Mapping
1361   :name: amdgpu-dwarf-address-space-mapping-table
1362
1363   ======================================= ===== ======= ======== ================= =======================
1364   DWARF                                                          AMDGPU            Notes
1365   --------------------------------------- ----- ---------------- ----------------- -----------------------
1366   Address Space Name                      Value Address Bit Size Address Space
1367   --------------------------------------- ----- ------- -------- ----------------- -----------------------
1368   ..                                            64-bit  32-bit
1369                                                 process process
1370                                                 address address
1371                                                 space   space
1372   ======================================= ===== ======= ======== ================= =======================
1373   ``DW_ASPACE_none``                      0x00  8       4        Global            *default address space*
1374   ``DW_ASPACE_AMDGPU_generic``            0x01  8       4        Generic (Flat)
1375   ``DW_ASPACE_AMDGPU_region``             0x02  4       4        Region (GDS)
1376   ``DW_ASPACE_AMDGPU_local``              0x03  4       4        Local (group/LDS)
1377   *Reserved*                              0x04
1378   ``DW_ASPACE_AMDGPU_private_lane``       0x05  4       4        Private (Scratch) *focused lane*
1379   ``DW_ASPACE_AMDGPU_private_wave``       0x06  4       4        Private (Scratch) *unswizzled wavefront*
1380   ======================================= ===== ======= ======== ================= =======================
1381
1382See :ref:`amdgpu-address-spaces` for information on the AMDGPU address spaces
1383including address size and NULL value.
1384
1385The ``DW_ASPACE_none`` address space is the default target architecture address
1386space used in DWARF operations that do not specify an address space. It
1387therefore has to map to the global address space so that the ``DW_OP_addr*`` and
1388related operations can refer to addresses in the program code.
1389
1390The ``DW_ASPACE_AMDGPU_generic`` address space allows location expressions to
1391specify the flat address space. If the address corresponds to an address in the
1392local address space, then it corresponds to the wavefront that is executing the
1393focused thread of execution. If the address corresponds to an address in the
1394private address space, then it corresponds to the lane that is executing the
1395focused thread of execution for languages that are implemented using a SIMD or
1396SIMT execution model.
1397
1398.. note::
1399
1400  CUDA-like languages such as HIP that do not have address spaces in the
1401  language type system, but do allow variables to be allocated in different
1402  address spaces, need to explicitly specify the ``DW_ASPACE_AMDGPU_generic``
1403  address space in the DWARF expression operations as the default address space
1404  is the global address space.
1405
1406The ``DW_ASPACE_AMDGPU_local`` address space allows location expressions to
1407specify the local address space corresponding to the wavefront that is executing
1408the focused thread of execution.
1409
1410The ``DW_ASPACE_AMDGPU_private_lane`` address space allows location expressions
1411to specify the private address space corresponding to the lane that is executing
1412the focused thread of execution for languages that are implemented using a SIMD
1413or SIMT execution model.
1414
1415The ``DW_ASPACE_AMDGPU_private_wave`` address space allows location expressions
1416to specify the unswizzled private address space corresponding to the wavefront
1417that is executing the focused thread of execution. The wavefront view of private
1418memory is the per wavefront unswizzled backing memory layout defined in
1419:ref:`amdgpu-address-spaces`, such that address 0 corresponds to the first
1420location for the backing memory of the wavefront (namely the address is not
1421offset by ``wavefront-scratch-base``). The following formula can be used to
1422convert from a ``DW_ASPACE_AMDGPU_private_lane`` address to a
1423``DW_ASPACE_AMDGPU_private_wave`` address:
1424
1425::
1426
1427  private-address-wavefront =
1428    ((private-address-lane / 4) * wavefront-size * 4) +
1429    (wavefront-lane-id * 4) + (private-address-lane % 4)
1430
1431If the ``DW_ASPACE_AMDGPU_private_lane`` address is dword aligned, and the start
1432of the dwords for each lane starting with lane 0 is required, then this
1433simplifies to:
1434
1435::
1436
1437  private-address-wavefront =
1438    private-address-lane * wavefront-size
1439
1440A compiler can use the ``DW_ASPACE_AMDGPU_private_wave`` address space to read a
1441complete spilled vector register back into a complete vector register in the
1442CFI. The frame pointer can be a private lane address which is dword aligned,
1443which can be shifted to multiply by the wavefront size, and then used to form a
1444private wavefront address that gives a location for a contiguous set of dwords,
1445one per lane, where the vector register dwords are spilled. The compiler knows
1446the wavefront size since it generates the code. Note that the type of the
1447address may have to be converted as the size of a
1448``DW_ASPACE_AMDGPU_private_lane`` address may be smaller than the size of a
1449``DW_ASPACE_AMDGPU_private_wave`` address.
1450
1451.. _amdgpu-dwarf-lane-identifier:
1452
1453Lane identifier
1454---------------
1455
1456DWARF lane identifies specify a target architecture lane position for hardware
1457that executes in a SIMD or SIMT manner, and on which a source language maps its
1458threads of execution onto those lanes. The DWARF lane identifier is pushed by
1459the ``DW_OP_LLVM_push_lane`` DWARF expression operation. See DWARF Version 5
1460section 2.5 which is updated by *DWARF Extensions For Heterogeneous Debugging*
1461section :ref:`amdgpu-dwarf-operation-expressions`.
1462
1463For AMDGPU, the lane identifier corresponds to the hardware lane ID of a
1464wavefront. It is numbered from 0 to the wavefront size minus 1.
1465
1466Operation Expressions
1467---------------------
1468
1469DWARF expressions are used to compute program values and the locations of
1470program objects. See DWARF Version 5 section 2.5 and
1471:ref:`amdgpu-dwarf-operation-expressions`.
1472
1473DWARF location descriptions describe how to access storage which includes memory
1474and registers. When accessing storage on AMDGPU, bytes are ordered with least
1475significant bytes first, and bits are ordered within bytes with least
1476significant bits first.
1477
1478For AMDGPU CFI expressions, ``DW_OP_LLVM_select_bit_piece`` is used to describe
1479unwinding vector registers that are spilled under the execution mask to memory:
1480the zero-single location description is the vector register, and the one-single
1481location description is the spilled memory location description. The
1482``DW_OP_LLVM_form_aspace_address`` is used to specify the address space of the
1483memory location description.
1484
1485In AMDGPU expressions, ``DW_OP_LLVM_select_bit_piece`` is used by the
1486``DW_AT_LLVM_lane_pc`` attribute expression where divergent control flow is
1487controlled by the execution mask. An undefined location description together
1488with ``DW_OP_LLVM_extend`` is used to indicate the lane was not active on entry
1489to the subprogram. See :ref:`amdgpu-dwarf-dw-at-llvm-lane-pc` for an example.
1490
1491Debugger Information Entry Attributes
1492-------------------------------------
1493
1494This section describes how certain debugger information entry attributes are
1495used by AMDGPU. See the sections in DWARF Version 5 section 2 which are updated
1496by *DWARF Extensions For Heterogeneous Debugging* section
1497:ref:`amdgpu-dwarf-debugging-information-entry-attributes`.
1498
1499.. _amdgpu-dwarf-dw-at-llvm-lane-pc:
1500
1501``DW_AT_LLVM_lane_pc``
1502~~~~~~~~~~~~~~~~~~~~~~
1503
1504For AMDGPU, the ``DW_AT_LLVM_lane_pc`` attribute is used to specify the program
1505location of the separate lanes of a SIMT thread.
1506
1507If the lane is an active lane then this will be the same as the current program
1508location.
1509
1510If the lane is inactive, but was active on entry to the subprogram, then this is
1511the program location in the subprogram at which execution of the lane is
1512conceptual positioned.
1513
1514If the lane was not active on entry to the subprogram, then this will be the
1515undefined location. A client debugger can check if the lane is part of a valid
1516work-group by checking that the lane is in the range of the associated
1517work-group within the grid, accounting for partial work-groups. If it is not,
1518then the debugger can omit any information for the lane. Otherwise, the debugger
1519may repeatedly unwind the stack and inspect the ``DW_AT_LLVM_lane_pc`` of the
1520calling subprogram until it finds a non-undefined location. Conceptually the
1521lane only has the call frames that it has a non-undefined
1522``DW_AT_LLVM_lane_pc``.
1523
1524The following example illustrates how the AMDGPU backend can generate a DWARF
1525location list expression for the nested ``IF/THEN/ELSE`` structures of the
1526following subprogram pseudo code for a target with 64 lanes per wavefront.
1527
1528.. code::
1529  :number-lines:
1530
1531  SUBPROGRAM X
1532  BEGIN
1533    a;
1534    IF (c1) THEN
1535      b;
1536      IF (c2) THEN
1537        c;
1538      ELSE
1539        d;
1540      ENDIF
1541      e;
1542    ELSE
1543      f;
1544    ENDIF
1545    g;
1546  END
1547
1548The AMDGPU backend may generate the following pseudo LLVM MIR to manipulate the
1549execution mask (``EXEC``) to linearize the control flow. The condition is
1550evaluated to make a mask of the lanes for which the condition evaluates to true.
1551First the ``THEN`` region is executed by setting the ``EXEC`` mask to the
1552logical ``AND`` of the current ``EXEC`` mask with the condition mask. Then the
1553``ELSE`` region is executed by negating the ``EXEC`` mask and logical ``AND`` of
1554the saved ``EXEC`` mask at the start of the region. After the ``IF/THEN/ELSE``
1555region the ``EXEC`` mask is restored to the value it had at the beginning of the
1556region. This is shown below. Other approaches are possible, but the basic
1557concept is the same.
1558
1559.. code::
1560  :number-lines:
1561
1562  $lex_start:
1563    a;
1564    %1 = EXEC
1565    %2 = c1
1566  $lex_1_start:
1567    EXEC = %1 & %2
1568  $if_1_then:
1569      b;
1570      %3 = EXEC
1571      %4 = c2
1572  $lex_1_1_start:
1573      EXEC = %3 & %4
1574  $lex_1_1_then:
1575        c;
1576      EXEC = ~EXEC & %3
1577  $lex_1_1_else:
1578        d;
1579      EXEC = %3
1580  $lex_1_1_end:
1581      e;
1582    EXEC = ~EXEC & %1
1583  $lex_1_else:
1584      f;
1585    EXEC = %1
1586  $lex_1_end:
1587    g;
1588  $lex_end:
1589
1590To create the DWARF location list expression that defines the location
1591description of a vector of lane program locations, the LLVM MIR ``DBG_VALUE``
1592pseudo instruction can be used to annotate the linearized control flow. This can
1593be done by defining an artificial variable for the lane PC. The DWARF location
1594list expression created for it is used as the value of the
1595``DW_AT_LLVM_lane_pc`` attribute on the subprogram's debugger information entry.
1596
1597A DWARF procedure is defined for each well nested structured control flow region
1598which provides the conceptual lane program location for a lane if it is not
1599active (namely it is divergent). The DWARF operation expression for each region
1600conceptually inherits the value of the immediately enclosing region and modifies
1601it according to the semantics of the region.
1602
1603For an ``IF/THEN/ELSE`` region the divergent program location is at the start of
1604the region for the ``THEN`` region since it is executed first. For the ``ELSE``
1605region the divergent program location is at the end of the ``IF/THEN/ELSE``
1606region since the ``THEN`` region has completed.
1607
1608The lane PC artificial variable is assigned at each region transition. It uses
1609the immediately enclosing region's DWARF procedure to compute the program
1610location for each lane assuming they are divergent, and then modifies the result
1611by inserting the current program location for each lane that the ``EXEC`` mask
1612indicates is active.
1613
1614By having separate DWARF procedures for each region, they can be reused to
1615define the value for any nested region. This reduces the total size of the DWARF
1616operation expressions.
1617
1618The following provides an example using pseudo LLVM MIR.
1619
1620.. code::
1621  :number-lines:
1622
1623  $lex_start:
1624    DEFINE_DWARF %__uint_64 = DW_TAG_base_type[
1625      DW_AT_name = "__uint64";
1626      DW_AT_byte_size = 8;
1627      DW_AT_encoding = DW_ATE_unsigned;
1628    ];
1629    DEFINE_DWARF %__active_lane_pc = DW_TAG_dwarf_procedure[
1630      DW_AT_name = "__active_lane_pc";
1631      DW_AT_location = [
1632        DW_OP_regx PC;
1633        DW_OP_LLVM_extend 64, 64;
1634        DW_OP_regval_type EXEC, %uint_64;
1635        DW_OP_LLVM_select_bit_piece 64, 64;
1636      ];
1637    ];
1638    DEFINE_DWARF %__divergent_lane_pc = DW_TAG_dwarf_procedure[
1639      DW_AT_name = "__divergent_lane_pc";
1640      DW_AT_location = [
1641        DW_OP_LLVM_undefined;
1642        DW_OP_LLVM_extend 64, 64;
1643      ];
1644    ];
1645    DBG_VALUE $noreg, $noreg, %DW_AT_LLVM_lane_pc, DIExpression[
1646      DW_OP_call_ref %__divergent_lane_pc;
1647      DW_OP_call_ref %__active_lane_pc;
1648    ];
1649    a;
1650    %1 = EXEC;
1651    DBG_VALUE %1, $noreg, %__lex_1_save_exec;
1652    %2 = c1;
1653  $lex_1_start:
1654    EXEC = %1 & %2;
1655  $lex_1_then:
1656      DEFINE_DWARF %__divergent_lane_pc_1_then = DW_TAG_dwarf_procedure[
1657        DW_AT_name = "__divergent_lane_pc_1_then";
1658        DW_AT_location = DIExpression[
1659          DW_OP_call_ref %__divergent_lane_pc;
1660          DW_OP_addrx &lex_1_start;
1661          DW_OP_stack_value;
1662          DW_OP_LLVM_extend 64, 64;
1663          DW_OP_call_ref %__lex_1_save_exec;
1664          DW_OP_deref_type 64, %__uint_64;
1665          DW_OP_LLVM_select_bit_piece 64, 64;
1666        ];
1667      ];
1668      DBG_VALUE $noreg, $noreg, %DW_AT_LLVM_lane_pc, DIExpression[
1669        DW_OP_call_ref %__divergent_lane_pc_1_then;
1670        DW_OP_call_ref %__active_lane_pc;
1671      ];
1672      b;
1673      %3 = EXEC;
1674      DBG_VALUE %3, %__lex_1_1_save_exec;
1675      %4 = c2;
1676  $lex_1_1_start:
1677      EXEC = %3 & %4;
1678  $lex_1_1_then:
1679        DEFINE_DWARF %__divergent_lane_pc_1_1_then = DW_TAG_dwarf_procedure[
1680          DW_AT_name = "__divergent_lane_pc_1_1_then";
1681          DW_AT_location = DIExpression[
1682            DW_OP_call_ref %__divergent_lane_pc_1_then;
1683            DW_OP_addrx &lex_1_1_start;
1684            DW_OP_stack_value;
1685            DW_OP_LLVM_extend 64, 64;
1686            DW_OP_call_ref %__lex_1_1_save_exec;
1687            DW_OP_deref_type 64, %__uint_64;
1688            DW_OP_LLVM_select_bit_piece 64, 64;
1689          ];
1690        ];
1691        DBG_VALUE $noreg, $noreg, %DW_AT_LLVM_lane_pc, DIExpression[
1692          DW_OP_call_ref %__divergent_lane_pc_1_1_then;
1693          DW_OP_call_ref %__active_lane_pc;
1694        ];
1695        c;
1696      EXEC = ~EXEC & %3;
1697  $lex_1_1_else:
1698        DEFINE_DWARF %__divergent_lane_pc_1_1_else = DW_TAG_dwarf_procedure[
1699          DW_AT_name = "__divergent_lane_pc_1_1_else";
1700          DW_AT_location = DIExpression[
1701            DW_OP_call_ref %__divergent_lane_pc_1_then;
1702            DW_OP_addrx &lex_1_1_end;
1703            DW_OP_stack_value;
1704            DW_OP_LLVM_extend 64, 64;
1705            DW_OP_call_ref %__lex_1_1_save_exec;
1706            DW_OP_deref_type 64, %__uint_64;
1707            DW_OP_LLVM_select_bit_piece 64, 64;
1708          ];
1709        ];
1710        DBG_VALUE $noreg, $noreg, %DW_AT_LLVM_lane_pc, DIExpression[
1711          DW_OP_call_ref %__divergent_lane_pc_1_1_else;
1712          DW_OP_call_ref %__active_lane_pc;
1713        ];
1714        d;
1715      EXEC = %3;
1716  $lex_1_1_end:
1717      DBG_VALUE $noreg, $noreg, %DW_AT_LLVM_lane_pc, DIExpression[
1718        DW_OP_call_ref %__divergent_lane_pc;
1719        DW_OP_call_ref %__active_lane_pc;
1720      ];
1721      e;
1722    EXEC = ~EXEC & %1;
1723  $lex_1_else:
1724      DEFINE_DWARF %__divergent_lane_pc_1_else = DW_TAG_dwarf_procedure[
1725        DW_AT_name = "__divergent_lane_pc_1_else";
1726        DW_AT_location = DIExpression[
1727          DW_OP_call_ref %__divergent_lane_pc;
1728          DW_OP_addrx &lex_1_end;
1729          DW_OP_stack_value;
1730          DW_OP_LLVM_extend 64, 64;
1731          DW_OP_call_ref %__lex_1_save_exec;
1732          DW_OP_deref_type 64, %__uint_64;
1733          DW_OP_LLVM_select_bit_piece 64, 64;
1734        ];
1735      ];
1736      DBG_VALUE $noreg, $noreg, %DW_AT_LLVM_lane_pc, DIExpression[
1737        DW_OP_call_ref %__divergent_lane_pc_1_else;
1738        DW_OP_call_ref %__active_lane_pc;
1739      ];
1740      f;
1741    EXEC = %1;
1742  $lex_1_end:
1743    DBG_VALUE $noreg, $noreg, %DW_AT_LLVM_lane_pc DIExpression[
1744      DW_OP_call_ref %__divergent_lane_pc;
1745      DW_OP_call_ref %__active_lane_pc;
1746    ];
1747    g;
1748  $lex_end:
1749
1750The DWARF procedure ``%__active_lane_pc`` is used to update the lane pc elements
1751that are active, with the current program location.
1752
1753Artificial variables %__lex_1_save_exec and %__lex_1_1_save_exec are created for
1754the execution masks saved on entry to a region. Using the ``DBG_VALUE`` pseudo
1755instruction, location list entries will be created that describe where the
1756artificial variables are allocated at any given program location. The compiler
1757may allocate them to registers or spill them to memory.
1758
1759The DWARF procedures for each region use the values of the saved execution mask
1760artificial variables to only update the lanes that are active on entry to the
1761region. All other lanes retain the value of the enclosing region where they were
1762last active. If they were not active on entry to the subprogram, then will have
1763the undefined location description.
1764
1765Other structured control flow regions can be handled similarly. For example,
1766loops would set the divergent program location for the region at the end of the
1767loop. Any lanes active will be in the loop, and any lanes not active must have
1768exited the loop.
1769
1770An ``IF/THEN/ELSEIF/ELSEIF/...`` region can be treated as a nest of
1771``IF/THEN/ELSE`` regions.
1772
1773The DWARF procedures can use the active lane artificial variable described in
1774:ref:`amdgpu-dwarf-amdgpu-dw-at-llvm-active-lane` rather than the actual
1775``EXEC`` mask in order to support whole or quad wavefront mode.
1776
1777.. _amdgpu-dwarf-amdgpu-dw-at-llvm-active-lane:
1778
1779``DW_AT_LLVM_active_lane``
1780~~~~~~~~~~~~~~~~~~~~~~~~~~
1781
1782The ``DW_AT_LLVM_active_lane`` attribute on a subprogram debugger information
1783entry is used to specify the lanes that are conceptually active for a SIMT
1784thread.
1785
1786The execution mask may be modified to implement whole or quad wavefront mode
1787operations. For example, all lanes may need to temporarily be made active to
1788execute a whole wavefront operation. Such regions would save the ``EXEC`` mask,
1789update it to enable the necessary lanes, perform the operations, and then
1790restore the ``EXEC`` mask from the saved value. While executing the whole
1791wavefront region, the conceptual execution mask is the saved value, not the
1792``EXEC`` value.
1793
1794This is handled by defining an artificial variable for the active lane mask. The
1795active lane mask artificial variable would be the actual ``EXEC`` mask for
1796normal regions, and the saved execution mask for regions where the mask is
1797temporarily updated. The location list expression created for this artificial
1798variable is used to define the value of the ``DW_AT_LLVM_active_lane``
1799attribute.
1800
1801``DW_AT_LLVM_augmentation``
1802~~~~~~~~~~~~~~~~~~~~~~~~~~~
1803
1804For AMDGPU, the ``DW_AT_LLVM_augmentation`` attribute of a compilation unit
1805debugger information entry has the following value for the augmentation string:
1806
1807::
1808
1809  [amdgpu:v0.0]
1810
1811The "vX.Y" specifies the major X and minor Y version number of the AMDGPU
1812extensions used in the DWARF of the compilation unit. The version number
1813conforms to [SEMVER]_.
1814
1815Call Frame Information
1816----------------------
1817
1818DWARF Call Frame Information (CFI) describes how a consumer can virtually
1819*unwind* call frames in a running process or core dump. See DWARF Version 5
1820section 6.4 and :ref:`amdgpu-dwarf-call-frame-information`.
1821
1822For AMDGPU, the Common Information Entry (CIE) fields have the following values:
1823
18241.  ``augmentation`` string contains the following null-terminated UTF-8 string:
1825
1826    ::
1827
1828      [amd:v0.0]
1829
1830    The ``vX.Y`` specifies the major X and minor Y version number of the AMDGPU
1831    extensions used in this CIE or to the FDEs that use it. The version number
1832    conforms to [SEMVER]_.
1833
18342.  ``address_size`` for the ``Global`` address space is defined in
1835    :ref:`amdgpu-dwarf-address-space-identifier`.
1836
18373.  ``segment_selector_size`` is 0 as AMDGPU does not use a segment selector.
1838
18394.  ``code_alignment_factor`` is 4 bytes.
1840
1841    .. TODO::
1842
1843       Add to :ref:`amdgpu-processor-table` table.
1844
18455.  ``data_alignment_factor`` is 4 bytes.
1846
1847    .. TODO::
1848
1849       Add to :ref:`amdgpu-processor-table` table.
1850
18516.  ``return_address_register`` is ``PC_32`` for 32-bit processes and ``PC_64``
1852    for 64-bit processes defined in :ref:`amdgpu-dwarf-register-identifier`.
1853
18547.  ``initial_instructions`` Since a subprogram X with fewer registers can be
1855    called from subprogram Y that has more allocated, X will not change any of
1856    the extra registers as it cannot access them. Therefore, the default rule
1857    for all columns is ``same value``.
1858
1859For AMDGPU the register number follows the numbering defined in
1860:ref:`amdgpu-dwarf-register-identifier`.
1861
1862For AMDGPU the instructions are variable size. A consumer can subtract 1 from
1863the return address to get the address of a byte within the call site
1864instructions. See DWARF Version 5 section 6.4.4.
1865
1866Accelerated Access
1867------------------
1868
1869See DWARF Version 5 section 6.1.
1870
1871Lookup By Name Section Header
1872~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
1873
1874See DWARF Version 5 section 6.1.1.4.1 and :ref:`amdgpu-dwarf-lookup-by-name`.
1875
1876For AMDGPU the lookup by name section header table:
1877
1878``augmentation_string_size`` (uword)
1879
1880  Set to the length of the ``augmentation_string`` value which is always a
1881  multiple of 4.
1882
1883``augmentation_string`` (sequence of UTF-8 characters)
1884
1885  Contains the following UTF-8 string null padded to a multiple of 4 bytes:
1886
1887  ::
1888
1889    [amdgpu:v0.0]
1890
1891  The "vX.Y" specifies the major X and minor Y version number of the AMDGPU
1892  extensions used in the DWARF of this index. The version number conforms to
1893  [SEMVER]_.
1894
1895  .. note::
1896
1897    This is different to the DWARF Version 5 definition that requires the first
1898    4 characters to be the vendor ID. But this is consistent with the other
1899    augmentation strings and does allow multiple vendor contributions. However,
1900    backwards compatibility may be more desirable.
1901
1902Lookup By Address Section Header
1903~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
1904
1905See DWARF Version 5 section 6.1.2.
1906
1907For AMDGPU the lookup by address section header table:
1908
1909``address_size`` (ubyte)
1910
1911  Match the address size for the ``Global`` address space defined in
1912  :ref:`amdgpu-dwarf-address-space-identifier`.
1913
1914``segment_selector_size`` (ubyte)
1915
1916  AMDGPU does not use a segment selector so this is 0. The entries in the
1917  ``.debug_aranges`` do not have a segment selector.
1918
1919Line Number Information
1920-----------------------
1921
1922See DWARF Version 5 section 6.2 and :ref:`amdgpu-dwarf-line-number-information`.
1923
1924AMDGPU does not use the ``isa`` state machine registers and always sets it to 0.
1925The instruction set must be obtained from the ELF file header ``e_flags`` field
1926in the ``EF_AMDGPU_MACH`` bit position (see :ref:`ELF Header
1927<amdgpu-elf-header>`). See DWARF Version 5 section 6.2.2.
1928
1929.. TODO::
1930
1931  Should the ``isa`` state machine register be used to indicate if the code is
1932  in wavefront32 or wavefront64 mode? Or used to specify the architecture ISA?
1933
1934For AMDGPU the line number program header fields have the following values (see
1935DWARF Version 5 section 6.2.4):
1936
1937``address_size`` (ubyte)
1938  Matches the address size for the ``Global`` address space defined in
1939  :ref:`amdgpu-dwarf-address-space-identifier`.
1940
1941``segment_selector_size`` (ubyte)
1942  AMDGPU does not use a segment selector so this is 0.
1943
1944``minimum_instruction_length`` (ubyte)
1945  For GFX9-GFX10 this is 4.
1946
1947``maximum_operations_per_instruction`` (ubyte)
1948  For GFX9-GFX10 this is 1.
1949
1950Source text for online-compiled programs (for example, those compiled by the
1951OpenCL language runtime) may be embedded into the DWARF Version 5 line table.
1952See DWARF Version 5 section 6.2.4.1 which is updated by *DWARF Extensions For
1953Heterogeneous Debugging* section :ref:`DW_LNCT_LLVM_source
1954<amdgpu-dwarf-line-number-information-dw-lnct-llvm-source>`.
1955
1956The Clang option used to control source embedding in AMDGPU is defined in
1957:ref:`amdgpu-clang-debug-options-table`.
1958
1959  .. table:: AMDGPU Clang Debug Options
1960     :name: amdgpu-clang-debug-options-table
1961
1962     ==================== ==================================================
1963     Debug Flag           Description
1964     ==================== ==================================================
1965     -g[no-]embed-source  Enable/disable embedding source text in DWARF
1966                          debug sections. Useful for environments where
1967                          source cannot be written to disk, such as
1968                          when performing online compilation.
1969     ==================== ==================================================
1970
1971For example:
1972
1973``-gembed-source``
1974  Enable the embedded source.
1975
1976``-gno-embed-source``
1977  Disable the embedded source.
1978
197932-Bit and 64-Bit DWARF Formats
1980-------------------------------
1981
1982See DWARF Version 5 section 7.4 and
1983:ref:`amdgpu-dwarf-32-bit-and-64-bit-dwarf-formats`.
1984
1985For AMDGPU:
1986
1987* For the ``amdgcn`` target architecture only the 64-bit process address space
1988  is supported.
1989
1990* The producer can generate either 32-bit or 64-bit DWARF format. LLVM generates
1991  the 32-bit DWARF format.
1992
1993Unit Headers
1994------------
1995
1996For AMDGPU the following values apply for each of the unit headers described in
1997DWARF Version 5 sections 7.5.1.1, 7.5.1.2, and 7.5.1.3:
1998
1999``address_size`` (ubyte)
2000  Matches the address size for the ``Global`` address space defined in
2001  :ref:`amdgpu-dwarf-address-space-identifier`.
2002
2003.. _amdgpu-code-conventions:
2004
2005Code Conventions
2006================
2007
2008This section provides code conventions used for each supported target triple OS
2009(see :ref:`amdgpu-target-triples`).
2010
2011AMDHSA
2012------
2013
2014This section provides code conventions used when the target triple OS is
2015``amdhsa`` (see :ref:`amdgpu-target-triples`).
2016
2017.. _amdgpu-amdhsa-code-object-target-identification:
2018
2019Code Object Target Identification
2020~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
2021
2022The AMDHSA OS uses the following syntax to specify the code object
2023target as a single string:
2024
2025  ``<Architecture>-<Vendor>-<OS>-<Environment>-<Processor><Target Features>``
2026
2027Where:
2028
2029  - ``<Architecture>``, ``<Vendor>``, ``<OS>`` and ``<Environment>``
2030    are the same as the *Target Triple* (see
2031    :ref:`amdgpu-target-triples`).
2032
2033  - ``<Processor>`` is the same as the *Processor* (see
2034    :ref:`amdgpu-processors`).
2035
2036  - ``<Target Features>`` is a list of the enabled *Target Features*
2037    (see :ref:`amdgpu-target-features`), each prefixed by a plus, that
2038    apply to *Processor*. The list must be in the same order as listed
2039    in the table :ref:`amdgpu-target-feature-table`. Note that *Target
2040    Features* must be included in the list if they are enabled even if
2041    that is the default for *Processor*.
2042
2043For example:
2044
2045  ``"amdgcn-amd-amdhsa--gfx902+xnack"``
2046
2047.. _amdgpu-amdhsa-code-object-metadata:
2048
2049Code Object Metadata
2050~~~~~~~~~~~~~~~~~~~~
2051
2052The code object metadata specifies extensible metadata associated with the code
2053objects executed on HSA [HSA]_ compatible runtimes such as AMD's ROCm
2054[AMD-ROCm]_. The encoding and semantics of this metadata depends on the code
2055object version; see :ref:`amdgpu-amdhsa-code-object-metadata-v2` and
2056:ref:`amdgpu-amdhsa-code-object-metadata-v3`.
2057
2058Code object metadata is specified in a note record (see
2059:ref:`amdgpu-note-records`) and is required when the target triple OS is
2060``amdhsa`` (see :ref:`amdgpu-target-triples`). It must contain the minimum
2061information necessary to support the ROCM kernel queries. For example, the
2062segment sizes needed in a dispatch packet. In addition, a high-level language
2063runtime may require other information to be included. For example, the AMD
2064OpenCL runtime records kernel argument information.
2065
2066.. _amdgpu-amdhsa-code-object-metadata-v2:
2067
2068Code Object V2 Metadata (-mattr=-code-object-v3)
2069++++++++++++++++++++++++++++++++++++++++++++++++
2070
2071.. warning:: Code Object V2 is not the default code object version emitted by
2072  this version of LLVM. For a description of the metadata generated with the
2073  default configuration (Code Object V3) see
2074  :ref:`amdgpu-amdhsa-code-object-metadata-v3`.
2075
2076Code object V2 metadata is specified by the ``NT_AMD_AMDGPU_METADATA`` note
2077record (see :ref:`amdgpu-note-records-v2`).
2078
2079The metadata is specified as a YAML formatted string (see [YAML]_ and
2080:doc:`YamlIO`).
2081
2082.. TODO::
2083
2084  Is the string null terminated? It probably should not if YAML allows it to
2085  contain null characters, otherwise it should be.
2086
2087The metadata is represented as a single YAML document comprised of the mapping
2088defined in table :ref:`amdgpu-amdhsa-code-object-metadata-map-table-v2` and
2089referenced tables.
2090
2091For boolean values, the string values of ``false`` and ``true`` are used for
2092false and true respectively.
2093
2094Additional information can be added to the mappings. To avoid conflicts, any
2095non-AMD key names should be prefixed by "*vendor-name*.".
2096
2097  .. table:: AMDHSA Code Object V2 Metadata Map
2098     :name: amdgpu-amdhsa-code-object-metadata-map-table-v2
2099
2100     ========== ============== ========= =======================================
2101     String Key Value Type     Required? Description
2102     ========== ============== ========= =======================================
2103     "Version"  sequence of    Required  - The first integer is the major
2104                2 integers                 version. Currently 1.
2105                                         - The second integer is the minor
2106                                           version. Currently 0.
2107     "Printf"   sequence of              Each string is encoded information
2108                strings                  about a printf function call. The
2109                                         encoded information is organized as
2110                                         fields separated by colon (':'):
2111
2112                                         ``ID:N:S[0]:S[1]:...:S[N-1]:FormatString``
2113
2114                                         where:
2115
2116                                         ``ID``
2117                                           A 32-bit integer as a unique id for
2118                                           each printf function call
2119
2120                                         ``N``
2121                                           A 32-bit integer equal to the number
2122                                           of arguments of printf function call
2123                                           minus 1
2124
2125                                         ``S[i]`` (where i = 0, 1, ... , N-1)
2126                                           32-bit integers for the size in bytes
2127                                           of the i-th FormatString argument of
2128                                           the printf function call
2129
2130                                         FormatString
2131                                           The format string passed to the
2132                                           printf function call.
2133     "Kernels"  sequence of    Required  Sequence of the mappings for each
2134                mapping                  kernel in the code object. See
2135                                         :ref:`amdgpu-amdhsa-code-object-kernel-metadata-map-table-v2`
2136                                         for the definition of the mapping.
2137     ========== ============== ========= =======================================
2138
2139..
2140
2141  .. table:: AMDHSA Code Object V2 Kernel Metadata Map
2142     :name: amdgpu-amdhsa-code-object-kernel-metadata-map-table-v2
2143
2144     ================= ============== ========= ================================
2145     String Key        Value Type     Required? Description
2146     ================= ============== ========= ================================
2147     "Name"            string         Required  Source name of the kernel.
2148     "SymbolName"      string         Required  Name of the kernel
2149                                                descriptor ELF symbol.
2150     "Language"        string                   Source language of the kernel.
2151                                                Values include:
2152
2153                                                - "OpenCL C"
2154                                                - "OpenCL C++"
2155                                                - "HCC"
2156                                                - "OpenMP"
2157
2158     "LanguageVersion" sequence of              - The first integer is the major
2159                       2 integers                 version.
2160                                                - The second integer is the
2161                                                  minor version.
2162     "Attrs"           mapping                  Mapping of kernel attributes.
2163                                                See
2164                                                :ref:`amdgpu-amdhsa-code-object-kernel-attribute-metadata-map-table-v2`
2165                                                for the mapping definition.
2166     "Args"            sequence of              Sequence of mappings of the
2167                       mapping                  kernel arguments. See
2168                                                :ref:`amdgpu-amdhsa-code-object-kernel-argument-metadata-map-table-v2`
2169                                                for the definition of the mapping.
2170     "CodeProps"       mapping                  Mapping of properties related to
2171                                                the kernel code. See
2172                                                :ref:`amdgpu-amdhsa-code-object-kernel-code-properties-metadata-map-table-v2`
2173                                                for the mapping definition.
2174     ================= ============== ========= ================================
2175
2176..
2177
2178  .. table:: AMDHSA Code Object V2 Kernel Attribute Metadata Map
2179     :name: amdgpu-amdhsa-code-object-kernel-attribute-metadata-map-table-v2
2180
2181     =================== ============== ========= ==============================
2182     String Key          Value Type     Required? Description
2183     =================== ============== ========= ==============================
2184     "ReqdWorkGroupSize" sequence of              If not 0, 0, 0 then all values
2185                         3 integers               must be >=1 and the dispatch
2186                                                  work-group size X, Y, Z must
2187                                                  correspond to the specified
2188                                                  values. Defaults to 0, 0, 0.
2189
2190                                                  Corresponds to the OpenCL
2191                                                  ``reqd_work_group_size``
2192                                                  attribute.
2193     "WorkGroupSizeHint" sequence of              The dispatch work-group size
2194                         3 integers               X, Y, Z is likely to be the
2195                                                  specified values.
2196
2197                                                  Corresponds to the OpenCL
2198                                                  ``work_group_size_hint``
2199                                                  attribute.
2200     "VecTypeHint"       string                   The name of a scalar or vector
2201                                                  type.
2202
2203                                                  Corresponds to the OpenCL
2204                                                  ``vec_type_hint`` attribute.
2205
2206     "RuntimeHandle"     string                   The external symbol name
2207                                                  associated with a kernel.
2208                                                  OpenCL runtime allocates a
2209                                                  global buffer for the symbol
2210                                                  and saves the kernel's address
2211                                                  to it, which is used for
2212                                                  device side enqueueing. Only
2213                                                  available for device side
2214                                                  enqueued kernels.
2215     =================== ============== ========= ==============================
2216
2217..
2218
2219  .. table:: AMDHSA Code Object V2 Kernel Argument Metadata Map
2220     :name: amdgpu-amdhsa-code-object-kernel-argument-metadata-map-table-v2
2221
2222     ================= ============== ========= ================================
2223     String Key        Value Type     Required? Description
2224     ================= ============== ========= ================================
2225     "Name"            string                   Kernel argument name.
2226     "TypeName"        string                   Kernel argument type name.
2227     "Size"            integer        Required  Kernel argument size in bytes.
2228     "Align"           integer        Required  Kernel argument alignment in
2229                                                bytes. Must be a power of two.
2230     "ValueKind"       string         Required  Kernel argument kind that
2231                                                specifies how to set up the
2232                                                corresponding argument.
2233                                                Values include:
2234
2235                                                "ByValue"
2236                                                  The argument is copied
2237                                                  directly into the kernarg.
2238
2239                                                "GlobalBuffer"
2240                                                  A global address space pointer
2241                                                  to the buffer data is passed
2242                                                  in the kernarg.
2243
2244                                                "DynamicSharedPointer"
2245                                                  A group address space pointer
2246                                                  to dynamically allocated LDS
2247                                                  is passed in the kernarg.
2248
2249                                                "Sampler"
2250                                                  A global address space
2251                                                  pointer to a S# is passed in
2252                                                  the kernarg.
2253
2254                                                "Image"
2255                                                  A global address space
2256                                                  pointer to a T# is passed in
2257                                                  the kernarg.
2258
2259                                                "Pipe"
2260                                                  A global address space pointer
2261                                                  to an OpenCL pipe is passed in
2262                                                  the kernarg.
2263
2264                                                "Queue"
2265                                                  A global address space pointer
2266                                                  to an OpenCL device enqueue
2267                                                  queue is passed in the
2268                                                  kernarg.
2269
2270                                                "HiddenGlobalOffsetX"
2271                                                  The OpenCL grid dispatch
2272                                                  global offset for the X
2273                                                  dimension is passed in the
2274                                                  kernarg.
2275
2276                                                "HiddenGlobalOffsetY"
2277                                                  The OpenCL grid dispatch
2278                                                  global offset for the Y
2279                                                  dimension is passed in the
2280                                                  kernarg.
2281
2282                                                "HiddenGlobalOffsetZ"
2283                                                  The OpenCL grid dispatch
2284                                                  global offset for the Z
2285                                                  dimension is passed in the
2286                                                  kernarg.
2287
2288                                                "HiddenNone"
2289                                                  An argument that is not used
2290                                                  by the kernel. Space needs to
2291                                                  be left for it, but it does
2292                                                  not need to be set up.
2293
2294                                                "HiddenPrintfBuffer"
2295                                                  A global address space pointer
2296                                                  to the runtime printf buffer
2297                                                  is passed in kernarg.
2298
2299                                                "HiddenHostcallBuffer"
2300                                                  A global address space pointer
2301                                                  to the runtime hostcall buffer
2302                                                  is passed in kernarg.
2303
2304                                                "HiddenDefaultQueue"
2305                                                  A global address space pointer
2306                                                  to the OpenCL device enqueue
2307                                                  queue that should be used by
2308                                                  the kernel by default is
2309                                                  passed in the kernarg.
2310
2311                                                "HiddenCompletionAction"
2312                                                  A global address space pointer
2313                                                  to help link enqueued kernels into
2314                                                  the ancestor tree for determining
2315                                                  when the parent kernel has finished.
2316
2317                                                "HiddenMultiGridSyncArg"
2318                                                  A global address space pointer for
2319                                                  multi-grid synchronization is
2320                                                  passed in the kernarg.
2321
2322     "ValueType"       string                   Unused and deprecated. This should no longer
2323                                                be emitted, but is accepted for compatibility.
2324
2325
2326     "PointeeAlign"    integer                  Alignment in bytes of pointee
2327                                                type for pointer type kernel
2328                                                argument. Must be a power
2329                                                of 2. Only present if
2330                                                "ValueKind" is
2331                                                "DynamicSharedPointer".
2332     "AddrSpaceQual"   string                   Kernel argument address space
2333                                                qualifier. Only present if
2334                                                "ValueKind" is "GlobalBuffer" or
2335                                                "DynamicSharedPointer". Values
2336                                                are:
2337
2338                                                - "Private"
2339                                                - "Global"
2340                                                - "Constant"
2341                                                - "Local"
2342                                                - "Generic"
2343                                                - "Region"
2344
2345                                                .. TODO::
2346                                                   Is GlobalBuffer only Global
2347                                                   or Constant? Is
2348                                                   DynamicSharedPointer always
2349                                                   Local? Can HCC allow Generic?
2350                                                   How can Private or Region
2351                                                   ever happen?
2352     "AccQual"         string                   Kernel argument access
2353                                                qualifier. Only present if
2354                                                "ValueKind" is "Image" or
2355                                                "Pipe". Values
2356                                                are:
2357
2358                                                - "ReadOnly"
2359                                                - "WriteOnly"
2360                                                - "ReadWrite"
2361
2362                                                .. TODO::
2363                                                   Does this apply to
2364                                                   GlobalBuffer?
2365     "ActualAccQual"   string                   The actual memory accesses
2366                                                performed by the kernel on the
2367                                                kernel argument. Only present if
2368                                                "ValueKind" is "GlobalBuffer",
2369                                                "Image", or "Pipe". This may be
2370                                                more restrictive than indicated
2371                                                by "AccQual" to reflect what the
2372                                                kernel actual does. If not
2373                                                present then the runtime must
2374                                                assume what is implied by
2375                                                "AccQual" and "IsConst". Values
2376                                                are:
2377
2378                                                - "ReadOnly"
2379                                                - "WriteOnly"
2380                                                - "ReadWrite"
2381
2382     "IsConst"         boolean                  Indicates if the kernel argument
2383                                                is const qualified. Only present
2384                                                if "ValueKind" is
2385                                                "GlobalBuffer".
2386
2387     "IsRestrict"      boolean                  Indicates if the kernel argument
2388                                                is restrict qualified. Only
2389                                                present if "ValueKind" is
2390                                                "GlobalBuffer".
2391
2392     "IsVolatile"      boolean                  Indicates if the kernel argument
2393                                                is volatile qualified. Only
2394                                                present if "ValueKind" is
2395                                                "GlobalBuffer".
2396
2397     "IsPipe"          boolean                  Indicates if the kernel argument
2398                                                is pipe qualified. Only present
2399                                                if "ValueKind" is "Pipe".
2400
2401                                                .. TODO::
2402                                                   Can GlobalBuffer be pipe
2403                                                   qualified?
2404     ================= ============== ========= ================================
2405
2406..
2407
2408  .. table:: AMDHSA Code Object V2 Kernel Code Properties Metadata Map
2409     :name: amdgpu-amdhsa-code-object-kernel-code-properties-metadata-map-table-v2
2410
2411     ============================ ============== ========= =====================
2412     String Key                   Value Type     Required? Description
2413     ============================ ============== ========= =====================
2414     "KernargSegmentSize"         integer        Required  The size in bytes of
2415                                                           the kernarg segment
2416                                                           that holds the values
2417                                                           of the arguments to
2418                                                           the kernel.
2419     "GroupSegmentFixedSize"      integer        Required  The amount of group
2420                                                           segment memory
2421                                                           required by a
2422                                                           work-group in
2423                                                           bytes. This does not
2424                                                           include any
2425                                                           dynamically allocated
2426                                                           group segment memory
2427                                                           that may be added
2428                                                           when the kernel is
2429                                                           dispatched.
2430     "PrivateSegmentFixedSize"    integer        Required  The amount of fixed
2431                                                           private address space
2432                                                           memory required for a
2433                                                           work-item in
2434                                                           bytes. If the kernel
2435                                                           uses a dynamic call
2436                                                           stack then additional
2437                                                           space must be added
2438                                                           to this value for the
2439                                                           call stack.
2440     "KernargSegmentAlign"        integer        Required  The maximum byte
2441                                                           alignment of
2442                                                           arguments in the
2443                                                           kernarg segment. Must
2444                                                           be a power of 2.
2445     "WavefrontSize"              integer        Required  Wavefront size. Must
2446                                                           be a power of 2.
2447     "NumSGPRs"                   integer        Required  Number of scalar
2448                                                           registers used by a
2449                                                           wavefront for
2450                                                           GFX6-GFX10. This
2451                                                           includes the special
2452                                                           SGPRs for VCC, Flat
2453                                                           Scratch (GFX7-GFX10)
2454                                                           and XNACK (for
2455                                                           GFX8-GFX10). It does
2456                                                           not include the 16
2457                                                           SGPR added if a trap
2458                                                           handler is
2459                                                           enabled. It is not
2460                                                           rounded up to the
2461                                                           allocation
2462                                                           granularity.
2463     "NumVGPRs"                   integer        Required  Number of vector
2464                                                           registers used by
2465                                                           each work-item for
2466                                                           GFX6-GFX10
2467     "MaxFlatWorkGroupSize"       integer        Required  Maximum flat
2468                                                           work-group size
2469                                                           supported by the
2470                                                           kernel in work-items.
2471                                                           Must be >=1 and
2472                                                           consistent with
2473                                                           ReqdWorkGroupSize if
2474                                                           not 0, 0, 0.
2475     "NumSpilledSGPRs"            integer                  Number of stores from
2476                                                           a scalar register to
2477                                                           a register allocator
2478                                                           created spill
2479                                                           location.
2480     "NumSpilledVGPRs"            integer                  Number of stores from
2481                                                           a vector register to
2482                                                           a register allocator
2483                                                           created spill
2484                                                           location.
2485     ============================ ============== ========= =====================
2486
2487.. _amdgpu-amdhsa-code-object-metadata-v3:
2488
2489Code Object V3 Metadata (-mattr=+code-object-v3)
2490++++++++++++++++++++++++++++++++++++++++++++++++
2491
2492Code object V3 metadata is specified by the ``NT_AMDGPU_METADATA`` note record
2493(see :ref:`amdgpu-note-records-v3`).
2494
2495The metadata is represented as Message Pack formatted binary data (see
2496[MsgPack]_). The top level is a Message Pack map that includes the
2497keys defined in table
2498:ref:`amdgpu-amdhsa-code-object-metadata-map-table-v3` and referenced
2499tables.
2500
2501Additional information can be added to the maps. To avoid conflicts,
2502any key names should be prefixed by "*vendor-name*." where
2503``vendor-name`` can be the name of the vendor and specific vendor
2504tool that generates the information. The prefix is abbreviated to
2505simply "." when it appears within a map that has been added by the
2506same *vendor-name*.
2507
2508  .. table:: AMDHSA Code Object V3 Metadata Map
2509     :name: amdgpu-amdhsa-code-object-metadata-map-table-v3
2510
2511     ================= ============== ========= =======================================
2512     String Key        Value Type     Required? Description
2513     ================= ============== ========= =======================================
2514     "amdhsa.version"  sequence of    Required  - The first integer is the major
2515                       2 integers                 version. Currently 1.
2516                                                - The second integer is the minor
2517                                                  version. Currently 0.
2518     "amdhsa.printf"   sequence of              Each string is encoded information
2519                       strings                  about a printf function call. The
2520                                                encoded information is organized as
2521                                                fields separated by colon (':'):
2522
2523                                                ``ID:N:S[0]:S[1]:...:S[N-1]:FormatString``
2524
2525                                                where:
2526
2527                                                ``ID``
2528                                                  A 32-bit integer as a unique id for
2529                                                  each printf function call
2530
2531                                                ``N``
2532                                                  A 32-bit integer equal to the number
2533                                                  of arguments of printf function call
2534                                                  minus 1
2535
2536                                                ``S[i]`` (where i = 0, 1, ... , N-1)
2537                                                  32-bit integers for the size in bytes
2538                                                  of the i-th FormatString argument of
2539                                                  the printf function call
2540
2541                                                FormatString
2542                                                  The format string passed to the
2543                                                  printf function call.
2544     "amdhsa.kernels"  sequence of    Required  Sequence of the maps for each
2545                       map                      kernel in the code object. See
2546                                                :ref:`amdgpu-amdhsa-code-object-kernel-metadata-map-table-v3`
2547                                                for the definition of the keys included
2548                                                in that map.
2549     ================= ============== ========= =======================================
2550
2551..
2552
2553  .. table:: AMDHSA Code Object V3 Kernel Metadata Map
2554     :name: amdgpu-amdhsa-code-object-kernel-metadata-map-table-v3
2555
2556     =================================== ============== ========= ================================
2557     String Key                          Value Type     Required? Description
2558     =================================== ============== ========= ================================
2559     ".name"                             string         Required  Source name of the kernel.
2560     ".symbol"                           string         Required  Name of the kernel
2561                                                                  descriptor ELF symbol.
2562     ".language"                         string                   Source language of the kernel.
2563                                                                  Values include:
2564
2565                                                                  - "OpenCL C"
2566                                                                  - "OpenCL C++"
2567                                                                  - "HCC"
2568                                                                  - "HIP"
2569                                                                  - "OpenMP"
2570                                                                  - "Assembler"
2571
2572     ".language_version"                 sequence of              - The first integer is the major
2573                                         2 integers                 version.
2574                                                                  - The second integer is the
2575                                                                    minor version.
2576     ".args"                             sequence of              Sequence of maps of the
2577                                         map                      kernel arguments. See
2578                                                                  :ref:`amdgpu-amdhsa-code-object-kernel-argument-metadata-map-table-v3`
2579                                                                  for the definition of the keys
2580                                                                  included in that map.
2581     ".reqd_workgroup_size"              sequence of              If not 0, 0, 0 then all values
2582                                         3 integers               must be >=1 and the dispatch
2583                                                                  work-group size X, Y, Z must
2584                                                                  correspond to the specified
2585                                                                  values. Defaults to 0, 0, 0.
2586
2587                                                                  Corresponds to the OpenCL
2588                                                                  ``reqd_work_group_size``
2589                                                                  attribute.
2590     ".workgroup_size_hint"              sequence of              The dispatch work-group size
2591                                         3 integers               X, Y, Z is likely to be the
2592                                                                  specified values.
2593
2594                                                                  Corresponds to the OpenCL
2595                                                                  ``work_group_size_hint``
2596                                                                  attribute.
2597     ".vec_type_hint"                    string                   The name of a scalar or vector
2598                                                                  type.
2599
2600                                                                  Corresponds to the OpenCL
2601                                                                  ``vec_type_hint`` attribute.
2602
2603     ".device_enqueue_symbol"            string                   The external symbol name
2604                                                                  associated with a kernel.
2605                                                                  OpenCL runtime allocates a
2606                                                                  global buffer for the symbol
2607                                                                  and saves the kernel's address
2608                                                                  to it, which is used for
2609                                                                  device side enqueueing. Only
2610                                                                  available for device side
2611                                                                  enqueued kernels.
2612     ".kernarg_segment_size"             integer        Required  The size in bytes of
2613                                                                  the kernarg segment
2614                                                                  that holds the values
2615                                                                  of the arguments to
2616                                                                  the kernel.
2617     ".group_segment_fixed_size"         integer        Required  The amount of group
2618                                                                  segment memory
2619                                                                  required by a
2620                                                                  work-group in
2621                                                                  bytes. This does not
2622                                                                  include any
2623                                                                  dynamically allocated
2624                                                                  group segment memory
2625                                                                  that may be added
2626                                                                  when the kernel is
2627                                                                  dispatched.
2628     ".private_segment_fixed_size"       integer        Required  The amount of fixed
2629                                                                  private address space
2630                                                                  memory required for a
2631                                                                  work-item in
2632                                                                  bytes. If the kernel
2633                                                                  uses a dynamic call
2634                                                                  stack then additional
2635                                                                  space must be added
2636                                                                  to this value for the
2637                                                                  call stack.
2638     ".kernarg_segment_align"            integer        Required  The maximum byte
2639                                                                  alignment of
2640                                                                  arguments in the
2641                                                                  kernarg segment. Must
2642                                                                  be a power of 2.
2643     ".wavefront_size"                   integer        Required  Wavefront size. Must
2644                                                                  be a power of 2.
2645     ".sgpr_count"                       integer        Required  Number of scalar
2646                                                                  registers required by a
2647                                                                  wavefront for
2648                                                                  GFX6-GFX9. A register
2649                                                                  is required if it is
2650                                                                  used explicitly, or
2651                                                                  if a higher numbered
2652                                                                  register is used
2653                                                                  explicitly. This
2654                                                                  includes the special
2655                                                                  SGPRs for VCC, Flat
2656                                                                  Scratch (GFX7-GFX9)
2657                                                                  and XNACK (for
2658                                                                  GFX8-GFX9). It does
2659                                                                  not include the 16
2660                                                                  SGPR added if a trap
2661                                                                  handler is
2662                                                                  enabled. It is not
2663                                                                  rounded up to the
2664                                                                  allocation
2665                                                                  granularity.
2666     ".vgpr_count"                       integer        Required  Number of vector
2667                                                                  registers required by
2668                                                                  each work-item for
2669                                                                  GFX6-GFX9. A register
2670                                                                  is required if it is
2671                                                                  used explicitly, or
2672                                                                  if a higher numbered
2673                                                                  register is used
2674                                                                  explicitly.
2675     ".max_flat_workgroup_size"          integer        Required  Maximum flat
2676                                                                  work-group size
2677                                                                  supported by the
2678                                                                  kernel in work-items.
2679                                                                  Must be >=1 and
2680                                                                  consistent with
2681                                                                  ReqdWorkGroupSize if
2682                                                                  not 0, 0, 0.
2683     ".sgpr_spill_count"                 integer                  Number of stores from
2684                                                                  a scalar register to
2685                                                                  a register allocator
2686                                                                  created spill
2687                                                                  location.
2688     ".vgpr_spill_count"                 integer                  Number of stores from
2689                                                                  a vector register to
2690                                                                  a register allocator
2691                                                                  created spill
2692                                                                  location.
2693     =================================== ============== ========= ================================
2694
2695..
2696
2697  .. table:: AMDHSA Code Object V3 Kernel Argument Metadata Map
2698     :name: amdgpu-amdhsa-code-object-kernel-argument-metadata-map-table-v3
2699
2700     ====================== ============== ========= ================================
2701     String Key             Value Type     Required? Description
2702     ====================== ============== ========= ================================
2703     ".name"                string                   Kernel argument name.
2704     ".type_name"           string                   Kernel argument type name.
2705     ".size"                integer        Required  Kernel argument size in bytes.
2706     ".offset"              integer        Required  Kernel argument offset in
2707                                                     bytes. The offset must be a
2708                                                     multiple of the alignment
2709                                                     required by the argument.
2710     ".value_kind"          string         Required  Kernel argument kind that
2711                                                     specifies how to set up the
2712                                                     corresponding argument.
2713                                                     Values include:
2714
2715                                                     "by_value"
2716                                                       The argument is copied
2717                                                       directly into the kernarg.
2718
2719                                                     "global_buffer"
2720                                                       A global address space pointer
2721                                                       to the buffer data is passed
2722                                                       in the kernarg.
2723
2724                                                     "dynamic_shared_pointer"
2725                                                       A group address space pointer
2726                                                       to dynamically allocated LDS
2727                                                       is passed in the kernarg.
2728
2729                                                     "sampler"
2730                                                       A global address space
2731                                                       pointer to a S# is passed in
2732                                                       the kernarg.
2733
2734                                                     "image"
2735                                                       A global address space
2736                                                       pointer to a T# is passed in
2737                                                       the kernarg.
2738
2739                                                     "pipe"
2740                                                       A global address space pointer
2741                                                       to an OpenCL pipe is passed in
2742                                                       the kernarg.
2743
2744                                                     "queue"
2745                                                       A global address space pointer
2746                                                       to an OpenCL device enqueue
2747                                                       queue is passed in the
2748                                                       kernarg.
2749
2750                                                     "hidden_global_offset_x"
2751                                                       The OpenCL grid dispatch
2752                                                       global offset for the X
2753                                                       dimension is passed in the
2754                                                       kernarg.
2755
2756                                                     "hidden_global_offset_y"
2757                                                       The OpenCL grid dispatch
2758                                                       global offset for the Y
2759                                                       dimension is passed in the
2760                                                       kernarg.
2761
2762                                                     "hidden_global_offset_z"
2763                                                       The OpenCL grid dispatch
2764                                                       global offset for the Z
2765                                                       dimension is passed in the
2766                                                       kernarg.
2767
2768                                                     "hidden_none"
2769                                                       An argument that is not used
2770                                                       by the kernel. Space needs to
2771                                                       be left for it, but it does
2772                                                       not need to be set up.
2773
2774                                                     "hidden_printf_buffer"
2775                                                       A global address space pointer
2776                                                       to the runtime printf buffer
2777                                                       is passed in kernarg.
2778
2779                                                     "hidden_hostcall_buffer"
2780                                                       A global address space pointer
2781                                                       to the runtime hostcall buffer
2782                                                       is passed in kernarg.
2783
2784                                                     "hidden_default_queue"
2785                                                       A global address space pointer
2786                                                       to the OpenCL device enqueue
2787                                                       queue that should be used by
2788                                                       the kernel by default is
2789                                                       passed in the kernarg.
2790
2791                                                     "hidden_completion_action"
2792                                                       A global address space pointer
2793                                                       to help link enqueued kernels into
2794                                                       the ancestor tree for determining
2795                                                       when the parent kernel has finished.
2796
2797                                                     "hidden_multigrid_sync_arg"
2798                                                       A global address space pointer for
2799                                                       multi-grid synchronization is
2800                                                       passed in the kernarg.
2801
2802     ".value_type"          string                    Unused and deprecated. This should no longer
2803                                                      be emitted, but is accepted for compatibility.
2804
2805     ".pointee_align"       integer                  Alignment in bytes of pointee
2806                                                     type for pointer type kernel
2807                                                     argument. Must be a power
2808                                                     of 2. Only present if
2809                                                     ".value_kind" is
2810                                                     "dynamic_shared_pointer".
2811     ".address_space"       string                   Kernel argument address space
2812                                                     qualifier. Only present if
2813                                                     ".value_kind" is "global_buffer" or
2814                                                     "dynamic_shared_pointer". Values
2815                                                     are:
2816
2817                                                     - "private"
2818                                                     - "global"
2819                                                     - "constant"
2820                                                     - "local"
2821                                                     - "generic"
2822                                                     - "region"
2823
2824                                                     .. TODO::
2825                                                        Is "global_buffer" only "global"
2826                                                        or "constant"? Is
2827                                                        "dynamic_shared_pointer" always
2828                                                        "local"? Can HCC allow "generic"?
2829                                                        How can "private" or "region"
2830                                                        ever happen?
2831     ".access"              string                   Kernel argument access
2832                                                     qualifier. Only present if
2833                                                     ".value_kind" is "image" or
2834                                                     "pipe". Values
2835                                                     are:
2836
2837                                                     - "read_only"
2838                                                     - "write_only"
2839                                                     - "read_write"
2840
2841                                                     .. TODO::
2842                                                        Does this apply to
2843                                                        "global_buffer"?
2844     ".actual_access"       string                   The actual memory accesses
2845                                                     performed by the kernel on the
2846                                                     kernel argument. Only present if
2847                                                     ".value_kind" is "global_buffer",
2848                                                     "image", or "pipe". This may be
2849                                                     more restrictive than indicated
2850                                                     by ".access" to reflect what the
2851                                                     kernel actual does. If not
2852                                                     present then the runtime must
2853                                                     assume what is implied by
2854                                                     ".access" and ".is_const"      . Values
2855                                                     are:
2856
2857                                                     - "read_only"
2858                                                     - "write_only"
2859                                                     - "read_write"
2860
2861     ".is_const"            boolean                  Indicates if the kernel argument
2862                                                     is const qualified. Only present
2863                                                     if ".value_kind" is
2864                                                     "global_buffer".
2865
2866     ".is_restrict"         boolean                  Indicates if the kernel argument
2867                                                     is restrict qualified. Only
2868                                                     present if ".value_kind" is
2869                                                     "global_buffer".
2870
2871     ".is_volatile"         boolean                  Indicates if the kernel argument
2872                                                     is volatile qualified. Only
2873                                                     present if ".value_kind" is
2874                                                     "global_buffer".
2875
2876     ".is_pipe"             boolean                  Indicates if the kernel argument
2877                                                     is pipe qualified. Only present
2878                                                     if ".value_kind" is "pipe".
2879
2880                                                     .. TODO::
2881                                                        Can "global_buffer" be pipe
2882                                                        qualified?
2883     ====================== ============== ========= ================================
2884
2885..
2886
2887Kernel Dispatch
2888~~~~~~~~~~~~~~~
2889
2890The HSA architected queuing language (AQL) defines a user space memory
2891interface that can be used to control the dispatch of kernels, in an agent
2892independent way. An agent can have zero or more AQL queues created for it using
2893the ROCm runtime, in which AQL packets (all of which are 64 bytes) can be
2894placed. See the *HSA Platform System Architecture Specification* [HSA]_ for the
2895AQL queue mechanics and packet layouts.
2896
2897The packet processor of a kernel agent is responsible for detecting and
2898dispatching HSA kernels from the AQL queues associated with it. For AMD GPUs the
2899packet processor is implemented by the hardware command processor (CP),
2900asynchronous dispatch controller (ADC) and shader processor input controller
2901(SPI).
2902
2903The ROCm runtime can be used to allocate an AQL queue object. It uses the kernel
2904mode driver to initialize and register the AQL queue with CP.
2905
2906To dispatch a kernel the following actions are performed. This can occur in the
2907CPU host program, or from an HSA kernel executing on a GPU.
2908
29091. A pointer to an AQL queue for the kernel agent on which the kernel is to be
2910   executed is obtained.
29112. A pointer to the kernel descriptor (see
2912   :ref:`amdgpu-amdhsa-kernel-descriptor`) of the kernel to execute is obtained.
2913   It must be for a kernel that is contained in a code object that that was
2914   loaded by the ROCm runtime on the kernel agent with which the AQL queue is
2915   associated.
29163. Space is allocated for the kernel arguments using the ROCm runtime allocator
2917   for a memory region with the kernarg property for the kernel agent that will
2918   execute the kernel. It must be at least 16-byte aligned.
29194. Kernel argument values are assigned to the kernel argument memory
2920   allocation. The layout is defined in the *HSA Programmer's Language
2921   Reference* [HSA]_. For AMDGPU the kernel execution directly accesses the
2922   kernel argument memory in the same way constant memory is accessed. (Note
2923   that the HSA specification allows an implementation to copy the kernel
2924   argument contents to another location that is accessed by the kernel.)
29255. An AQL kernel dispatch packet is created on the AQL queue. The ROCm runtime
2926   api uses 64-bit atomic operations to reserve space in the AQL queue for the
2927   packet. The packet must be set up, and the final write must use an atomic
2928   store release to set the packet kind to ensure the packet contents are
2929   visible to the kernel agent. AQL defines a doorbell signal mechanism to
2930   notify the kernel agent that the AQL queue has been updated. These rules, and
2931   the layout of the AQL queue and kernel dispatch packet is defined in the *HSA
2932   System Architecture Specification* [HSA]_.
29336. A kernel dispatch packet includes information about the actual dispatch,
2934   such as grid and work-group size, together with information from the code
2935   object about the kernel, such as segment sizes. The ROCm runtime queries on
2936   the kernel symbol can be used to obtain the code object values which are
2937   recorded in the :ref:`amdgpu-amdhsa-code-object-metadata`.
29387. CP executes micro-code and is responsible for detecting and setting up the
2939   GPU to execute the wavefronts of a kernel dispatch.
29408. CP ensures that when the a wavefront starts executing the kernel machine
2941   code, the scalar general purpose registers (SGPR) and vector general purpose
2942   registers (VGPR) are set up as required by the machine code. The required
2943   setup is defined in the :ref:`amdgpu-amdhsa-kernel-descriptor`. The initial
2944   register state is defined in
2945   :ref:`amdgpu-amdhsa-initial-kernel-execution-state`.
29469. The prolog of the kernel machine code (see
2947   :ref:`amdgpu-amdhsa-kernel-prolog`) sets up the machine state as necessary
2948   before continuing executing the machine code that corresponds to the kernel.
294910. When the kernel dispatch has completed execution, CP signals the completion
2950    signal specified in the kernel dispatch packet if not 0.
2951
2952Image and Samplers
2953~~~~~~~~~~~~~~~~~~
2954
2955Image and sample handles created by the ROCm runtime are 64-bit addresses of a
2956hardware 32-byte V# and 48 byte S# object respectively. In order to support the
2957HSA ``query_sampler`` operations two extra dwords are used to store the HSA BRIG
2958enumeration values for the queries that are not trivially deducible from the S#
2959representation.
2960
2961HSA Signals
2962~~~~~~~~~~~
2963
2964HSA signal handles created by the ROCm runtime are 64-bit addresses of a
2965structure allocated in memory accessible from both the CPU and GPU. The
2966structure is defined by the ROCm runtime and subject to change between releases
2967(see [AMD-ROCm-github]_).
2968
2969.. _amdgpu-amdhsa-hsa-aql-queue:
2970
2971HSA AQL Queue
2972~~~~~~~~~~~~~
2973
2974The HSA AQL queue structure is defined by the ROCm runtime and subject to change
2975between releases (see [AMD-ROCm-github]_). For some processors it contains
2976fields needed to implement certain language features such as the flat address
2977aperture bases. It also contains fields used by CP such as managing the
2978allocation of scratch memory.
2979
2980.. _amdgpu-amdhsa-kernel-descriptor:
2981
2982Kernel Descriptor
2983~~~~~~~~~~~~~~~~~
2984
2985A kernel descriptor consists of the information needed by CP to initiate the
2986execution of a kernel, including the entry point address of the machine code
2987that implements the kernel.
2988
2989Kernel Descriptor for GFX6-GFX10
2990++++++++++++++++++++++++++++++++
2991
2992CP microcode requires the Kernel descriptor to be allocated on 64-byte
2993alignment.
2994
2995  .. table:: Kernel Descriptor for GFX6-GFX10
2996     :name: amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table
2997
2998     ======= ======= =============================== ============================
2999     Bits    Size    Field Name                      Description
3000     ======= ======= =============================== ============================
3001     31:0    4 bytes GROUP_SEGMENT_FIXED_SIZE        The amount of fixed local
3002                                                     address space memory
3003                                                     required for a work-group
3004                                                     in bytes. This does not
3005                                                     include any dynamically
3006                                                     allocated local address
3007                                                     space memory that may be
3008                                                     added when the kernel is
3009                                                     dispatched.
3010     63:32   4 bytes PRIVATE_SEGMENT_FIXED_SIZE      The amount of fixed
3011                                                     private address space
3012                                                     memory required for a
3013                                                     work-item in bytes. If
3014                                                     is_dynamic_callstack is 1
3015                                                     then additional space must
3016                                                     be added to this value for
3017                                                     the call stack.
3018     127:64  8 bytes                                 Reserved, must be 0.
3019     191:128 8 bytes KERNEL_CODE_ENTRY_BYTE_OFFSET   Byte offset (possibly
3020                                                     negative) from base
3021                                                     address of kernel
3022                                                     descriptor to kernel's
3023                                                     entry point instruction
3024                                                     which must be 256 byte
3025                                                     aligned.
3026     351:272 20                                      Reserved, must be 0.
3027             bytes
3028     383:352 4 bytes COMPUTE_PGM_RSRC3               GFX6-9
3029                                                       Reserved, must be 0.
3030                                                     GFX10
3031                                                       Compute Shader (CS)
3032                                                       program settings used by
3033                                                       CP to set up
3034                                                       ``COMPUTE_PGM_RSRC3``
3035                                                       configuration
3036                                                       register. See
3037                                                       :ref:`amdgpu-amdhsa-compute_pgm_rsrc3-gfx10-table`.
3038     415:384 4 bytes COMPUTE_PGM_RSRC1               Compute Shader (CS)
3039                                                     program settings used by
3040                                                     CP to set up
3041                                                     ``COMPUTE_PGM_RSRC1``
3042                                                     configuration
3043                                                     register. See
3044                                                     :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`.
3045     447:416 4 bytes COMPUTE_PGM_RSRC2               Compute Shader (CS)
3046                                                     program settings used by
3047                                                     CP to set up
3048                                                     ``COMPUTE_PGM_RSRC2``
3049                                                     configuration
3050                                                     register. See
3051                                                     :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`.
3052     448     1 bit   ENABLE_SGPR_PRIVATE_SEGMENT     Enable the setup of the
3053                     _BUFFER                         SGPR user data registers
3054                                                     (see
3055                                                     :ref:`amdgpu-amdhsa-initial-kernel-execution-state`).
3056
3057                                                     The total number of SGPR
3058                                                     user data registers
3059                                                     requested must not exceed
3060                                                     16 and match value in
3061                                                     ``compute_pgm_rsrc2.user_sgpr.user_sgpr_count``.
3062                                                     Any requests beyond 16
3063                                                     will be ignored.
3064     449     1 bit   ENABLE_SGPR_DISPATCH_PTR        *see above*
3065     450     1 bit   ENABLE_SGPR_QUEUE_PTR           *see above*
3066     451     1 bit   ENABLE_SGPR_KERNARG_SEGMENT_PTR *see above*
3067     452     1 bit   ENABLE_SGPR_DISPATCH_ID         *see above*
3068     453     1 bit   ENABLE_SGPR_FLAT_SCRATCH_INIT   *see above*
3069     454     1 bit   ENABLE_SGPR_PRIVATE_SEGMENT     *see above*
3070                     _SIZE
3071     457:455 3 bits                                  Reserved, must be 0.
3072     458     1 bit   ENABLE_WAVEFRONT_SIZE32         GFX6-9
3073                                                       Reserved, must be 0.
3074                                                     GFX10
3075                                                       - If 0 execute in
3076                                                         wavefront size 64 mode.
3077                                                       - If 1 execute in
3078                                                         native wavefront size
3079                                                         32 mode.
3080     463:459 5 bits                                  Reserved, must be 0.
3081     511:464 6 bytes                                 Reserved, must be 0.
3082     512     **Total size 64 bytes.**
3083     ======= ====================================================================
3084
3085..
3086
3087  .. table:: compute_pgm_rsrc1 for GFX6-GFX10
3088     :name: amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table
3089
3090     ======= ======= =============================== ===========================================================================
3091     Bits    Size    Field Name                      Description
3092     ======= ======= =============================== ===========================================================================
3093     5:0     6 bits  GRANULATED_WORKITEM_VGPR_COUNT  Number of vector register
3094                                                     blocks used by each work-item;
3095                                                     granularity is device
3096                                                     specific:
3097
3098                                                     GFX6-GFX9
3099                                                       - vgprs_used 0..256
3100                                                       - max(0, ceil(vgprs_used / 4) - 1)
3101                                                     GFX10 (wavefront size 64)
3102                                                       - max_vgpr 1..256
3103                                                       - max(0, ceil(vgprs_used / 4) - 1)
3104                                                     GFX10 (wavefront size 32)
3105                                                       - max_vgpr 1..256
3106                                                       - max(0, ceil(vgprs_used / 8) - 1)
3107
3108                                                     Where vgprs_used is defined
3109                                                     as the highest VGPR number
3110                                                     explicitly referenced plus
3111                                                     one.
3112
3113                                                     Used by CP to set up
3114                                                     ``COMPUTE_PGM_RSRC1.VGPRS``.
3115
3116                                                     The
3117                                                     :ref:`amdgpu-assembler`
3118                                                     calculates this
3119                                                     automatically for the
3120                                                     selected processor from
3121                                                     values provided to the
3122                                                     `.amdhsa_kernel` directive
3123                                                     by the
3124                                                     `.amdhsa_next_free_vgpr`
3125                                                     nested directive (see
3126                                                     :ref:`amdhsa-kernel-directives-table`).
3127     9:6     4 bits  GRANULATED_WAVEFRONT_SGPR_COUNT Number of scalar register
3128                                                     blocks used by a wavefront;
3129                                                     granularity is device
3130                                                     specific:
3131
3132                                                     GFX6-GFX8
3133                                                       - sgprs_used 0..112
3134                                                       - max(0, ceil(sgprs_used / 8) - 1)
3135                                                     GFX9
3136                                                       - sgprs_used 0..112
3137                                                       - 2 * max(0, ceil(sgprs_used / 16) - 1)
3138                                                     GFX10
3139                                                       Reserved, must be 0.
3140                                                       (128 SGPRs always
3141                                                       allocated.)
3142
3143                                                     Where sgprs_used is
3144                                                     defined as the highest
3145                                                     SGPR number explicitly
3146                                                     referenced plus one, plus
3147                                                     a target specific number
3148                                                     of additional special
3149                                                     SGPRs for VCC,
3150                                                     FLAT_SCRATCH (GFX7+) and
3151                                                     XNACK_MASK (GFX8+), and
3152                                                     any additional
3153                                                     target specific
3154                                                     limitations. It does not
3155                                                     include the 16 SGPRs added
3156                                                     if a trap handler is
3157                                                     enabled.
3158
3159                                                     The target specific
3160                                                     limitations and special
3161                                                     SGPR layout are defined in
3162                                                     the hardware
3163                                                     documentation, which can
3164                                                     be found in the
3165                                                     :ref:`amdgpu-processors`
3166                                                     table.
3167
3168                                                     Used by CP to set up
3169                                                     ``COMPUTE_PGM_RSRC1.SGPRS``.
3170
3171                                                     The
3172                                                     :ref:`amdgpu-assembler`
3173                                                     calculates this
3174                                                     automatically for the
3175                                                     selected processor from
3176                                                     values provided to the
3177                                                     `.amdhsa_kernel` directive
3178                                                     by the
3179                                                     `.amdhsa_next_free_sgpr`
3180                                                     and `.amdhsa_reserve_*`
3181                                                     nested directives (see
3182                                                     :ref:`amdhsa-kernel-directives-table`).
3183     11:10   2 bits  PRIORITY                        Must be 0.
3184
3185                                                     Start executing wavefront
3186                                                     at the specified priority.
3187
3188                                                     CP is responsible for
3189                                                     filling in
3190                                                     ``COMPUTE_PGM_RSRC1.PRIORITY``.
3191     13:12   2 bits  FLOAT_ROUND_MODE_32             Wavefront starts execution
3192                                                     with specified rounding
3193                                                     mode for single (32
3194                                                     bit) floating point
3195                                                     precision floating point
3196                                                     operations.
3197
3198                                                     Floating point rounding
3199                                                     mode values are defined in
3200                                                     :ref:`amdgpu-amdhsa-floating-point-rounding-mode-enumeration-values-table`.
3201
3202                                                     Used by CP to set up
3203                                                     ``COMPUTE_PGM_RSRC1.FLOAT_MODE``.
3204     15:14   2 bits  FLOAT_ROUND_MODE_16_64          Wavefront starts execution
3205                                                     with specified rounding
3206                                                     denorm mode for half/double (16
3207                                                     and 64-bit) floating point
3208                                                     precision floating point
3209                                                     operations.
3210
3211                                                     Floating point rounding
3212                                                     mode values are defined in
3213                                                     :ref:`amdgpu-amdhsa-floating-point-rounding-mode-enumeration-values-table`.
3214
3215                                                     Used by CP to set up
3216                                                     ``COMPUTE_PGM_RSRC1.FLOAT_MODE``.
3217     17:16   2 bits  FLOAT_DENORM_MODE_32            Wavefront starts execution
3218                                                     with specified denorm mode
3219                                                     for single (32
3220                                                     bit)  floating point
3221                                                     precision floating point
3222                                                     operations.
3223
3224                                                     Floating point denorm mode
3225                                                     values are defined in
3226                                                     :ref:`amdgpu-amdhsa-floating-point-denorm-mode-enumeration-values-table`.
3227
3228                                                     Used by CP to set up
3229                                                     ``COMPUTE_PGM_RSRC1.FLOAT_MODE``.
3230     19:18   2 bits  FLOAT_DENORM_MODE_16_64         Wavefront starts execution
3231                                                     with specified denorm mode
3232                                                     for half/double (16
3233                                                     and 64-bit) floating point
3234                                                     precision floating point
3235                                                     operations.
3236
3237                                                     Floating point denorm mode
3238                                                     values are defined in
3239                                                     :ref:`amdgpu-amdhsa-floating-point-denorm-mode-enumeration-values-table`.
3240
3241                                                     Used by CP to set up
3242                                                     ``COMPUTE_PGM_RSRC1.FLOAT_MODE``.
3243     20      1 bit   PRIV                            Must be 0.
3244
3245                                                     Start executing wavefront
3246                                                     in privilege trap handler
3247                                                     mode.
3248
3249                                                     CP is responsible for
3250                                                     filling in
3251                                                     ``COMPUTE_PGM_RSRC1.PRIV``.
3252     21      1 bit   ENABLE_DX10_CLAMP               Wavefront starts execution
3253                                                     with DX10 clamp mode
3254                                                     enabled. Used by the vector
3255                                                     ALU to force DX10 style
3256                                                     treatment of NaN's (when
3257                                                     set, clamp NaN to zero,
3258                                                     otherwise pass NaN
3259                                                     through).
3260
3261                                                     Used by CP to set up
3262                                                     ``COMPUTE_PGM_RSRC1.DX10_CLAMP``.
3263     22      1 bit   DEBUG_MODE                      Must be 0.
3264
3265                                                     Start executing wavefront
3266                                                     in single step mode.
3267
3268                                                     CP is responsible for
3269                                                     filling in
3270                                                     ``COMPUTE_PGM_RSRC1.DEBUG_MODE``.
3271     23      1 bit   ENABLE_IEEE_MODE                Wavefront starts execution
3272                                                     with IEEE mode
3273                                                     enabled. Floating point
3274                                                     opcodes that support
3275                                                     exception flag gathering
3276                                                     will quiet and propagate
3277                                                     signaling-NaN inputs per
3278                                                     IEEE 754-2008. Min_dx10 and
3279                                                     max_dx10 become IEEE
3280                                                     754-2008 compliant due to
3281                                                     signaling-NaN propagation
3282                                                     and quieting.
3283
3284                                                     Used by CP to set up
3285                                                     ``COMPUTE_PGM_RSRC1.IEEE_MODE``.
3286     24      1 bit   BULKY                           Must be 0.
3287
3288                                                     Only one work-group allowed
3289                                                     to execute on a compute
3290                                                     unit.
3291
3292                                                     CP is responsible for
3293                                                     filling in
3294                                                     ``COMPUTE_PGM_RSRC1.BULKY``.
3295     25      1 bit   CDBG_USER                       Must be 0.
3296
3297                                                     Flag that can be used to
3298                                                     control debugging code.
3299
3300                                                     CP is responsible for
3301                                                     filling in
3302                                                     ``COMPUTE_PGM_RSRC1.CDBG_USER``.
3303     26      1 bit   FP16_OVFL                       GFX6-GFX8
3304                                                       Reserved, must be 0.
3305                                                     GFX9-GFX10
3306                                                       Wavefront starts execution
3307                                                       with specified fp16 overflow
3308                                                       mode.
3309
3310                                                       - If 0, fp16 overflow generates
3311                                                         +/-INF values.
3312                                                       - If 1, fp16 overflow that is the
3313                                                         result of an +/-INF input value
3314                                                         or divide by 0 produces a +/-INF,
3315                                                         otherwise clamps computed
3316                                                         overflow to +/-MAX_FP16 as
3317                                                         appropriate.
3318
3319                                                       Used by CP to set up
3320                                                       ``COMPUTE_PGM_RSRC1.FP16_OVFL``.
3321     28:27   2 bits                                  Reserved, must be 0.
3322     29      1 bit    WGP_MODE                       GFX6-GFX9
3323                                                       Reserved, must be 0.
3324                                                     GFX10
3325                                                       - If 0 execute work-groups in
3326                                                         CU wavefront execution mode.
3327                                                       - If 1 execute work-groups on
3328                                                         in WGP wavefront execution mode.
3329
3330                                                       See :ref:`amdgpu-amdhsa-memory-model`.
3331
3332                                                       Used by CP to set up
3333                                                       ``COMPUTE_PGM_RSRC1.WGP_MODE``.
3334     30      1 bit    MEM_ORDERED                    GFX6-9
3335                                                       Reserved, must be 0.
3336                                                     GFX10
3337                                                       Controls the behavior of the
3338                                                       waitcnt's vmcnt and vscnt
3339                                                       counters.
3340
3341                                                       - If 0 vmcnt reports completion
3342                                                         of load and atomic with return
3343                                                         out of order with sample
3344                                                         instructions, and the vscnt
3345                                                         reports the completion of
3346                                                         store and atomic without
3347                                                         return in order.
3348                                                       - If 1 vmcnt reports completion
3349                                                         of load, atomic with return
3350                                                         and sample instructions in
3351                                                         order, and the vscnt reports
3352                                                         the completion of store and
3353                                                         atomic without return in order.
3354
3355                                                       Used by CP to set up
3356                                                       ``COMPUTE_PGM_RSRC1.MEM_ORDERED``.
3357     31      1 bit    FWD_PROGRESS                   GFX6-9
3358                                                       Reserved, must be 0.
3359                                                     GFX10
3360                                                       - If 0 execute SIMD wavefronts
3361                                                         using oldest first policy.
3362                                                       - If 1 execute SIMD wavefronts to
3363                                                         ensure wavefronts will make some
3364                                                         forward progress.
3365
3366                                                       Used by CP to set up
3367                                                       ``COMPUTE_PGM_RSRC1.FWD_PROGRESS``.
3368     32      **Total size 4 bytes**
3369     ======= ===================================================================================================================
3370
3371..
3372
3373  .. table:: compute_pgm_rsrc2 for GFX6-GFX10
3374     :name: amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table
3375
3376     ======= ======= =============================== ===========================================================================
3377     Bits    Size    Field Name                      Description
3378     ======= ======= =============================== ===========================================================================
3379     0       1 bit   ENABLE_SGPR_PRIVATE_SEGMENT     Enable the setup of the
3380                     _WAVEFRONT_OFFSET               SGPR wavefront scratch offset
3381                                                     system register (see
3382                                                     :ref:`amdgpu-amdhsa-initial-kernel-execution-state`).
3383
3384                                                     Used by CP to set up
3385                                                     ``COMPUTE_PGM_RSRC2.SCRATCH_EN``.
3386     5:1     5 bits  USER_SGPR_COUNT                 The total number of SGPR
3387                                                     user data registers
3388                                                     requested. This number must
3389                                                     match the number of user
3390                                                     data registers enabled.
3391
3392                                                     Used by CP to set up
3393                                                     ``COMPUTE_PGM_RSRC2.USER_SGPR``.
3394     6       1 bit   ENABLE_TRAP_HANDLER             Must be 0.
3395
3396                                                     This bit represents
3397                                                     ``COMPUTE_PGM_RSRC2.TRAP_PRESENT``,
3398                                                     which is set by the CP if
3399                                                     the runtime has installed a
3400                                                     trap handler.
3401     7       1 bit   ENABLE_SGPR_WORKGROUP_ID_X      Enable the setup of the
3402                                                     system SGPR register for
3403                                                     the work-group id in the X
3404                                                     dimension (see
3405                                                     :ref:`amdgpu-amdhsa-initial-kernel-execution-state`).
3406
3407                                                     Used by CP to set up
3408                                                     ``COMPUTE_PGM_RSRC2.TGID_X_EN``.
3409     8       1 bit   ENABLE_SGPR_WORKGROUP_ID_Y      Enable the setup of the
3410                                                     system SGPR register for
3411                                                     the work-group id in the Y
3412                                                     dimension (see
3413                                                     :ref:`amdgpu-amdhsa-initial-kernel-execution-state`).
3414
3415                                                     Used by CP to set up
3416                                                     ``COMPUTE_PGM_RSRC2.TGID_Y_EN``.
3417     9       1 bit   ENABLE_SGPR_WORKGROUP_ID_Z      Enable the setup of the
3418                                                     system SGPR register for
3419                                                     the work-group id in the Z
3420                                                     dimension (see
3421                                                     :ref:`amdgpu-amdhsa-initial-kernel-execution-state`).
3422
3423                                                     Used by CP to set up
3424                                                     ``COMPUTE_PGM_RSRC2.TGID_Z_EN``.
3425     10      1 bit   ENABLE_SGPR_WORKGROUP_INFO      Enable the setup of the
3426                                                     system SGPR register for
3427                                                     work-group information (see
3428                                                     :ref:`amdgpu-amdhsa-initial-kernel-execution-state`).
3429
3430                                                     Used by CP to set up
3431                                                     ``COMPUTE_PGM_RSRC2.TGID_SIZE_EN``.
3432     12:11   2 bits  ENABLE_VGPR_WORKITEM_ID         Enable the setup of the
3433                                                     VGPR system registers used
3434                                                     for the work-item ID.
3435                                                     :ref:`amdgpu-amdhsa-system-vgpr-work-item-id-enumeration-values-table`
3436                                                     defines the values.
3437
3438                                                     Used by CP to set up
3439                                                     ``COMPUTE_PGM_RSRC2.TIDIG_CMP_CNT``.
3440     13      1 bit   ENABLE_EXCEPTION_ADDRESS_WATCH  Must be 0.
3441
3442                                                     Wavefront starts execution
3443                                                     with address watch
3444                                                     exceptions enabled which
3445                                                     are generated when L1 has
3446                                                     witnessed a thread access
3447                                                     an *address of
3448                                                     interest*.
3449
3450                                                     CP is responsible for
3451                                                     filling in the address
3452                                                     watch bit in
3453                                                     ``COMPUTE_PGM_RSRC2.EXCP_EN_MSB``
3454                                                     according to what the
3455                                                     runtime requests.
3456     14      1 bit   ENABLE_EXCEPTION_MEMORY         Must be 0.
3457
3458                                                     Wavefront starts execution
3459                                                     with memory violation
3460                                                     exceptions exceptions
3461                                                     enabled which are generated
3462                                                     when a memory violation has
3463                                                     occurred for this wavefront from
3464                                                     L1 or LDS
3465                                                     (write-to-read-only-memory,
3466                                                     mis-aligned atomic, LDS
3467                                                     address out of range,
3468                                                     illegal address, etc.).
3469
3470                                                     CP sets the memory
3471                                                     violation bit in
3472                                                     ``COMPUTE_PGM_RSRC2.EXCP_EN_MSB``
3473                                                     according to what the
3474                                                     runtime requests.
3475     23:15   9 bits  GRANULATED_LDS_SIZE             Must be 0.
3476
3477                                                     CP uses the rounded value
3478                                                     from the dispatch packet,
3479                                                     not this value, as the
3480                                                     dispatch may contain
3481                                                     dynamically allocated group
3482                                                     segment memory. CP writes
3483                                                     directly to
3484                                                     ``COMPUTE_PGM_RSRC2.LDS_SIZE``.
3485
3486                                                     Amount of group segment
3487                                                     (LDS) to allocate for each
3488                                                     work-group. Granularity is
3489                                                     device specific:
3490
3491                                                     GFX6:
3492                                                       roundup(lds-size / (64 * 4))
3493                                                     GFX7-GFX10:
3494                                                       roundup(lds-size / (128 * 4))
3495
3496     24      1 bit   ENABLE_EXCEPTION_IEEE_754_FP    Wavefront starts execution
3497                     _INVALID_OPERATION              with specified exceptions
3498                                                     enabled.
3499
3500                                                     Used by CP to set up
3501                                                     ``COMPUTE_PGM_RSRC2.EXCP_EN``
3502                                                     (set from bits 0..6).
3503
3504                                                     IEEE 754 FP Invalid
3505                                                     Operation
3506     25      1 bit   ENABLE_EXCEPTION_FP_DENORMAL    FP Denormal one or more
3507                     _SOURCE                         input operands is a
3508                                                     denormal number
3509     26      1 bit   ENABLE_EXCEPTION_IEEE_754_FP    IEEE 754 FP Division by
3510                     _DIVISION_BY_ZERO               Zero
3511     27      1 bit   ENABLE_EXCEPTION_IEEE_754_FP    IEEE 754 FP FP Overflow
3512                     _OVERFLOW
3513     28      1 bit   ENABLE_EXCEPTION_IEEE_754_FP    IEEE 754 FP Underflow
3514                     _UNDERFLOW
3515     29      1 bit   ENABLE_EXCEPTION_IEEE_754_FP    IEEE 754 FP Inexact
3516                     _INEXACT
3517     30      1 bit   ENABLE_EXCEPTION_INT_DIVIDE_BY  Integer Division by Zero
3518                     _ZERO                           (rcp_iflag_f32 instruction
3519                                                     only)
3520     31      1 bit                                   Reserved, must be 0.
3521     32      **Total size 4 bytes.**
3522     ======= ===================================================================================================================
3523
3524..
3525
3526  .. table:: compute_pgm_rsrc3 for GFX10
3527     :name: amdgpu-amdhsa-compute_pgm_rsrc3-gfx10-table
3528
3529     ======= ======= =============================== ===========================================================================
3530     Bits    Size    Field Name                      Description
3531     ======= ======= =============================== ===========================================================================
3532     3:0     4 bits  SHARED_VGPR_COUNT               Number of shared VGPRs for wavefront size 64. Granularity 8. Value 0-120.
3533                                                     compute_pgm_rsrc1.vgprs + shared_vgpr_cnt cannot exceed 64.
3534     31:4    28                                      Reserved, must be 0.
3535             bits
3536     32      **Total size 4 bytes.**
3537     ======= ===================================================================================================================
3538
3539..
3540
3541  .. table:: Floating Point Rounding Mode Enumeration Values
3542     :name: amdgpu-amdhsa-floating-point-rounding-mode-enumeration-values-table
3543
3544     ====================================== ===== ==============================
3545     Enumeration Name                       Value Description
3546     ====================================== ===== ==============================
3547     FLOAT_ROUND_MODE_NEAR_EVEN             0     Round Ties To Even
3548     FLOAT_ROUND_MODE_PLUS_INFINITY         1     Round Toward +infinity
3549     FLOAT_ROUND_MODE_MINUS_INFINITY        2     Round Toward -infinity
3550     FLOAT_ROUND_MODE_ZERO                  3     Round Toward 0
3551     ====================================== ===== ==============================
3552
3553..
3554
3555  .. table:: Floating Point Denorm Mode Enumeration Values
3556     :name: amdgpu-amdhsa-floating-point-denorm-mode-enumeration-values-table
3557
3558     ====================================== ===== ==============================
3559     Enumeration Name                       Value Description
3560     ====================================== ===== ==============================
3561     FLOAT_DENORM_MODE_FLUSH_SRC_DST        0     Flush Source and Destination
3562                                                  Denorms
3563     FLOAT_DENORM_MODE_FLUSH_DST            1     Flush Output Denorms
3564     FLOAT_DENORM_MODE_FLUSH_SRC            2     Flush Source Denorms
3565     FLOAT_DENORM_MODE_FLUSH_NONE           3     No Flush
3566     ====================================== ===== ==============================
3567
3568..
3569
3570  .. table:: System VGPR Work-Item ID Enumeration Values
3571     :name: amdgpu-amdhsa-system-vgpr-work-item-id-enumeration-values-table
3572
3573     ======================================== ===== ============================
3574     Enumeration Name                         Value Description
3575     ======================================== ===== ============================
3576     SYSTEM_VGPR_WORKITEM_ID_X                0     Set work-item X dimension
3577                                                    ID.
3578     SYSTEM_VGPR_WORKITEM_ID_X_Y              1     Set work-item X and Y
3579                                                    dimensions ID.
3580     SYSTEM_VGPR_WORKITEM_ID_X_Y_Z            2     Set work-item X, Y and Z
3581                                                    dimensions ID.
3582     SYSTEM_VGPR_WORKITEM_ID_UNDEFINED        3     Undefined.
3583     ======================================== ===== ============================
3584
3585.. _amdgpu-amdhsa-initial-kernel-execution-state:
3586
3587Initial Kernel Execution State
3588~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
3589
3590This section defines the register state that will be set up by the packet
3591processor prior to the start of execution of every wavefront. This is limited by
3592the constraints of the hardware controllers of CP/ADC/SPI.
3593
3594The order of the SGPR registers is defined, but the compiler can specify which
3595ones are actually setup in the kernel descriptor using the ``enable_sgpr_*`` bit
3596fields (see :ref:`amdgpu-amdhsa-kernel-descriptor`). The register numbers used
3597for enabled registers are dense starting at SGPR0: the first enabled register is
3598SGPR0, the next enabled register is SGPR1 etc.; disabled registers do not have
3599an SGPR number.
3600
3601The initial SGPRs comprise up to 16 User SRGPs that are set by CP and apply to
3602all wavefronts of the grid. It is possible to specify more than 16 User SGPRs
3603using the ``enable_sgpr_*`` bit fields, in which case only the first 16 are
3604actually initialized. These are then immediately followed by the System SGPRs
3605that are set up by ADC/SPI and can have different values for each wavefront of
3606the grid dispatch.
3607
3608SGPR register initial state is defined in
3609:ref:`amdgpu-amdhsa-sgpr-register-set-up-order-table`.
3610
3611  .. table:: SGPR Register Set Up Order
3612     :name: amdgpu-amdhsa-sgpr-register-set-up-order-table
3613
3614     ========== ========================== ====== ==============================
3615     SGPR Order Name                       Number Description
3616                (kernel descriptor enable  of
3617                field)                     SGPRs
3618     ========== ========================== ====== ==============================
3619     First      Private Segment Buffer     4      V# that can be used, together
3620                (enable_sgpr_private              with Scratch Wavefront Offset
3621                _segment_buffer)                  as an offset, to access the
3622                                                  private address space using a
3623                                                  segment address.
3624
3625                                                  CP uses the value provided by
3626                                                  the runtime.
3627     then       Dispatch Ptr               2      64-bit address of AQL dispatch
3628                (enable_sgpr_dispatch_ptr)        packet for kernel dispatch
3629                                                  actually executing.
3630     then       Queue Ptr                  2      64-bit address of amd_queue_t
3631                (enable_sgpr_queue_ptr)           object for AQL queue on which
3632                                                  the dispatch packet was
3633                                                  queued.
3634     then       Kernarg Segment Ptr        2      64-bit address of Kernarg
3635                (enable_sgpr_kernarg              segment. This is directly
3636                _segment_ptr)                     copied from the
3637                                                  kernarg_address in the kernel
3638                                                  dispatch packet.
3639
3640                                                  Having CP load it once avoids
3641                                                  loading it at the beginning of
3642                                                  every wavefront.
3643     then       Dispatch Id                2      64-bit Dispatch ID of the
3644                (enable_sgpr_dispatch_id)         dispatch packet being
3645                                                  executed.
3646     then       Flat Scratch Init          2      This is 2 SGPRs:
3647                (enable_sgpr_flat_scratch
3648                _init)                            GFX6
3649                                                    Not supported.
3650                                                  GFX7-GFX8
3651                                                    The first SGPR is a 32-bit
3652                                                    byte offset from
3653                                                    ``SH_HIDDEN_PRIVATE_BASE_VIMID``
3654                                                    to per SPI base of memory
3655                                                    for scratch for the queue
3656                                                    executing the kernel
3657                                                    dispatch. CP obtains this
3658                                                    from the runtime. (The
3659                                                    Scratch Segment Buffer base
3660                                                    address is
3661                                                    ``SH_HIDDEN_PRIVATE_BASE_VIMID``
3662                                                    plus this offset.) The value
3663                                                    of Scratch Wavefront Offset must
3664                                                    be added to this offset by
3665                                                    the kernel machine code,
3666                                                    right shifted by 8, and
3667                                                    moved to the FLAT_SCRATCH_HI
3668                                                    SGPR register.
3669                                                    FLAT_SCRATCH_HI corresponds
3670                                                    to SGPRn-4 on GFX7, and
3671                                                    SGPRn-6 on GFX8 (where SGPRn
3672                                                    is the highest numbered SGPR
3673                                                    allocated to the wavefront).
3674                                                    FLAT_SCRATCH_HI is
3675                                                    multiplied by 256 (as it is
3676                                                    in units of 256 bytes) and
3677                                                    added to
3678                                                    ``SH_HIDDEN_PRIVATE_BASE_VIMID``
3679                                                    to calculate the per wavefront
3680                                                    FLAT SCRATCH BASE in flat
3681                                                    memory instructions that
3682                                                    access the scratch
3683                                                    aperture.
3684
3685                                                    The second SGPR is 32-bit
3686                                                    byte size of a single
3687                                                    work-item's scratch memory
3688                                                    usage. CP obtains this from
3689                                                    the runtime, and it is
3690                                                    always a multiple of DWORD.
3691                                                    CP checks that the value in
3692                                                    the kernel dispatch packet
3693                                                    Private Segment Byte Size is
3694                                                    not larger and requests the
3695                                                    runtime to increase the
3696                                                    queue's scratch size if
3697                                                    necessary. The kernel code
3698                                                    must move it to
3699                                                    FLAT_SCRATCH_LO which is
3700                                                    SGPRn-3 on GFX7 and SGPRn-5
3701                                                    on GFX8. FLAT_SCRATCH_LO is
3702                                                    used as the FLAT SCRATCH
3703                                                    SIZE in flat memory
3704                                                    instructions. Having CP load
3705                                                    it once avoids loading it at
3706                                                    the beginning of every
3707                                                    wavefront.
3708                                                  GFX9-GFX10
3709                                                    This is the
3710                                                    64-bit base address of the
3711                                                    per SPI scratch backing
3712                                                    memory managed by SPI for
3713                                                    the queue executing the
3714                                                    kernel dispatch. CP obtains
3715                                                    this from the runtime (and
3716                                                    divides it if there are
3717                                                    multiple Shader Arrays each
3718                                                    with its own SPI). The value
3719                                                    of Scratch Wavefront Offset must
3720                                                    be added by the kernel
3721                                                    machine code and the result
3722                                                    moved to the FLAT_SCRATCH
3723                                                    SGPR which is SGPRn-6 and
3724                                                    SGPRn-5. It is used as the
3725                                                    FLAT SCRATCH BASE in flat
3726                                                    memory instructions.
3727     then       Private Segment Size       1      The 32-bit byte size of a
3728                                                  (enable_sgpr_private single
3729                                                  work-item's
3730                                                  scratch_segment_size) memory
3731                                                  allocation. This is the
3732                                                  value from the kernel
3733                                                  dispatch packet Private
3734                                                  Segment Byte Size rounded up
3735                                                  by CP to a multiple of
3736                                                  DWORD.
3737
3738                                                  Having CP load it once avoids
3739                                                  loading it at the beginning of
3740                                                  every wavefront.
3741
3742                                                  This is not used for
3743                                                  GFX7-GFX8 since it is the same
3744                                                  value as the second SGPR of
3745                                                  Flat Scratch Init. However, it
3746                                                  may be needed for GFX9-GFX10 which
3747                                                  changes the meaning of the
3748                                                  Flat Scratch Init value.
3749     then       Grid Work-Group Count X    1      32-bit count of the number of
3750                (enable_sgpr_grid                 work-groups in the X dimension
3751                _workgroup_count_X)               for the grid being
3752                                                  executed. Computed from the
3753                                                  fields in the kernel dispatch
3754                                                  packet as ((grid_size.x +
3755                                                  workgroup_size.x - 1) /
3756                                                  workgroup_size.x).
3757     then       Grid Work-Group Count Y    1      32-bit count of the number of
3758                (enable_sgpr_grid                 work-groups in the Y dimension
3759                _workgroup_count_Y &&             for the grid being
3760                less than 16 previous             executed. Computed from the
3761                SGPRs)                            fields in the kernel dispatch
3762                                                  packet as ((grid_size.y +
3763                                                  workgroup_size.y - 1) /
3764                                                  workgroupSize.y).
3765
3766                                                  Only initialized if <16
3767                                                  previous SGPRs initialized.
3768     then       Grid Work-Group Count Z    1      32-bit count of the number of
3769                (enable_sgpr_grid                 work-groups in the Z dimension
3770                _workgroup_count_Z &&             for the grid being
3771                less than 16 previous             executed. Computed from the
3772                SGPRs)                            fields in the kernel dispatch
3773                                                  packet as ((grid_size.z +
3774                                                  workgroup_size.z - 1) /
3775                                                  workgroupSize.z).
3776
3777                                                  Only initialized if <16
3778                                                  previous SGPRs initialized.
3779     then       Work-Group Id X            1      32-bit work-group id in X
3780                (enable_sgpr_workgroup_id         dimension of grid for
3781                _X)                               wavefront.
3782     then       Work-Group Id Y            1      32-bit work-group id in Y
3783                (enable_sgpr_workgroup_id         dimension of grid for
3784                _Y)                               wavefront.
3785     then       Work-Group Id Z            1      32-bit work-group id in Z
3786                (enable_sgpr_workgroup_id         dimension of grid for
3787                _Z)                               wavefront.
3788     then       Work-Group Info            1      {first_wavefront, 14'b0000,
3789                (enable_sgpr_workgroup            ordered_append_term[10:0],
3790                _info)                            threadgroup_size_in_wavefronts[5:0]}
3791     then       Scratch Wavefront Offset   1      32-bit byte offset from base
3792                (enable_sgpr_private              of scratch base of queue
3793                _segment_wavefront_offset)        executing the kernel
3794                                                  dispatch. Must be used as an
3795                                                  offset with Private
3796                                                  segment address when using
3797                                                  Scratch Segment Buffer. It
3798                                                  must be used to set up FLAT
3799                                                  SCRATCH for flat addressing
3800                                                  (see
3801                                                  :ref:`amdgpu-amdhsa-kernel-prolog-flat-scratch`).
3802     ========== ========================== ====== ==============================
3803
3804The order of the VGPR registers is defined, but the compiler can specify which
3805ones are actually setup in the kernel descriptor using the ``enable_vgpr*`` bit
3806fields (see :ref:`amdgpu-amdhsa-kernel-descriptor`). The register numbers used
3807for enabled registers are dense starting at VGPR0: the first enabled register is
3808VGPR0, the next enabled register is VGPR1 etc.; disabled registers do not have a
3809VGPR number.
3810
3811VGPR register initial state is defined in
3812:ref:`amdgpu-amdhsa-vgpr-register-set-up-order-table`.
3813
3814  .. table:: VGPR Register Set Up Order
3815     :name: amdgpu-amdhsa-vgpr-register-set-up-order-table
3816
3817     ========== ========================== ====== ==============================
3818     VGPR Order Name                       Number Description
3819                (kernel descriptor enable  of
3820                field)                     VGPRs
3821     ========== ========================== ====== ==============================
3822     First      Work-Item Id X             1      32-bit work item id in X
3823                (Always initialized)              dimension of work-group for
3824                                                  wavefront lane.
3825     then       Work-Item Id Y             1      32-bit work item id in Y
3826                (enable_vgpr_workitem_id          dimension of work-group for
3827                > 0)                              wavefront lane.
3828     then       Work-Item Id Z             1      32-bit work item id in Z
3829                (enable_vgpr_workitem_id          dimension of work-group for
3830                > 1)                              wavefront lane.
3831     ========== ========================== ====== ==============================
3832
3833The setting of registers is done by GPU CP/ADC/SPI hardware as follows:
3834
38351. SGPRs before the Work-Group Ids are set by CP using the 16 User Data
3836   registers.
38372. Work-group Id registers X, Y, Z are set by ADC which supports any
3838   combination including none.
38393. Scratch Wavefront Offset is set by SPI in a per wavefront basis which is why
3840   its value cannot be included with the flat scratch init value which is per
3841   queue.
38424. The VGPRs are set by SPI which only supports specifying either (X), (X, Y)
3843   or (X, Y, Z).
3844
3845Flat Scratch register pair are adjacent SGPRs so they can be moved as a 64-bit
3846value to the hardware required SGPRn-3 and SGPRn-4 respectively.
3847
3848The global segment can be accessed either using buffer instructions (GFX6 which
3849has V# 64-bit address support), flat instructions (GFX7-GFX10), or global
3850instructions (GFX9-GFX10).
3851
3852If buffer operations are used, then the compiler can generate a V# with the
3853following properties:
3854
3855* base address of 0
3856* no swizzle
3857* ATC: 1 if IOMMU present (such as APU)
3858* ptr64: 1
3859* MTYPE set to support memory coherence that matches the runtime (such as CC for
3860  APU and NC for dGPU).
3861
3862.. _amdgpu-amdhsa-kernel-prolog:
3863
3864Kernel Prolog
3865~~~~~~~~~~~~~
3866
3867The compiler performs initialization in the kernel prologue depending on the
3868target and information about things like stack usage in the kernel and called
3869functions. Some of this initialization requires the compiler to request certain
3870User and System SGPRs be present in the
3871:ref:`amdgpu-amdhsa-initial-kernel-execution-state` via the
3872:ref:`amdgpu-amdhsa-kernel-descriptor`.
3873
3874.. _amdgpu-amdhsa-kernel-prolog-cfi:
3875
3876CFI
3877+++
3878
38791.  The CFI return address is undefined.
3880
38812.  The CFI CFA is defined using an expression which evaluates to a location
3882    description that comprises one memory location description for the
3883    ``DW_ASPACE_AMDGPU_private_lane`` address space address ``0``.
3884
3885.. _amdgpu-amdhsa-kernel-prolog-m0:
3886
3887M0
3888++
3889
3890GFX6-GFX8
3891  The M0 register must be initialized with a value at least the total LDS size
3892  if the kernel may access LDS via DS or flat operations. Total LDS size is
3893  available in dispatch packet. For M0, it is also possible to use maximum
3894  possible value of LDS for given target (0x7FFF for GFX6 and 0xFFFF for
3895  GFX7-GFX8).
3896GFX9-GFX10
3897  The M0 register is not used for range checking LDS accesses and so does not
3898  need to be initialized in the prolog.
3899
3900.. _amdgpu-amdhsa-kernel-prolog-stack-pointer:
3901
3902Stack Pointer
3903+++++++++++++
3904
3905If the kernel has function calls it must set up the ABI stack pointer described
3906in :ref:`amdgpu-amdhsa-function-call-convention-non-kernel-functions` by setting
3907SGPR32 to the unswizzled scratch offset of the address past the last local
3908allocation.
3909
3910.. _amdgpu-amdhsa-kernel-prolog-frame-pointer:
3911
3912Frame Pointer
3913+++++++++++++
3914
3915If the kernel needs a frame pointer for the reasons defined in
3916``SIFrameLowering`` then SGPR33 is used and is always set to ``0`` in the
3917kernel prolog. If a frame pointer is not required then all uses of the frame
3918pointer are replaced with immediate ``0`` offsets.
3919
3920.. _amdgpu-amdhsa-kernel-prolog-flat-scratch:
3921
3922Flat Scratch
3923++++++++++++
3924
3925If the kernel or any function it calls may use flat operations to access
3926scratch memory, the prolog code must set up the FLAT_SCRATCH register pair
3927(FLAT_SCRATCH_LO/FLAT_SCRATCH_HI which are in SGPRn-4/SGPRn-3). Initialization
3928uses Flat Scratch Init and Scratch Wavefront Offset SGPR registers (see
3929:ref:`amdgpu-amdhsa-initial-kernel-execution-state`):
3930
3931GFX6
3932  Flat scratch is not supported.
3933
3934GFX7-GFX8
3935
3936  1. The low word of Flat Scratch Init is 32-bit byte offset from
3937     ``SH_HIDDEN_PRIVATE_BASE_VIMID`` to the base of scratch backing memory
3938     being managed by SPI for the queue executing the kernel dispatch. This is
3939     the same value used in the Scratch Segment Buffer V# base address. The
3940     prolog must add the value of Scratch Wavefront Offset to get the
3941     wavefront's byte scratch backing memory offset from
3942     ``SH_HIDDEN_PRIVATE_BASE_VIMID``. Since FLAT_SCRATCH_LO is in units of 256
3943     bytes, the offset must be right shifted by 8 before moving into
3944     FLAT_SCRATCH_LO.
3945  2. The second word of Flat Scratch Init is 32-bit byte size of a single
3946     work-items scratch memory usage. This is directly loaded from the kernel
3947     dispatch packet Private Segment Byte Size and rounded up to a multiple of
3948     DWORD. Having CP load it once avoids loading it at the beginning of every
3949     wavefront. The prolog must move it to FLAT_SCRATCH_LO for use as FLAT
3950     SCRATCH SIZE.
3951
3952GFX9-GFX10
3953  The Flat Scratch Init is the 64-bit address of the base of scratch backing
3954  memory being managed by SPI for the queue executing the kernel dispatch. The
3955  prolog must add the value of Scratch Wavefront Offset and moved to the
3956  FLAT_SCRATCH pair for use as the flat scratch base in flat memory
3957  instructions.
3958
3959.. _amdgpu-amdhsa-kernel-prolog-private-segment-buffer:
3960
3961Private Segment Buffer
3962++++++++++++++++++++++
3963
3964A set of four SGPRs beginning at a four-aligned SGPR index are always selected
3965to serve as the scratch V# for the kernel as follows:
3966
3967  - If it is known during instruction selection that there is stack usage,
3968    SGPR0-3 is reserved for use as the scratch V#.  Stack usage is assumed if
3969    optimizations are disabled (``-O0``), if stack objects already exist (for
3970    locals, etc.), or if there are any function calls.
3971
3972  - Otherwise, four high numbered SGPRs beginning at a four-aligned SGPR index
3973    are reserved for the tentative scratch V#. These will be used if it is
3974    determined that spilling is needed.
3975
3976    - If no use is made of the tentative scratch V#, then it is unreserved,
3977      and the register count is determined ignoring it.
3978    - If use is made of the tentative scratch V#, then its register numbers
3979      are shifted to the first four-aligned SGPR index after the highest one
3980      allocated by the register allocator, and all uses are updated. The
3981      register count includes them in the shifted location.
3982    - In either case, if the processor has the SGPR allocation bug, the
3983      tentative allocation is not shifted or unreserved in order to ensure
3984      the register count is higher to workaround the bug.
3985
3986    .. note::
3987
3988      This approach of using a tentative scratch V# and shifting the register
3989      numbers if used avoids having to perform register allocation a second
3990      time if the tentative V# is eliminated. This is more efficient and
3991      avoids the problem that the second register allocation may perform
3992      spilling which will fail as there is no longer a scratch V#.
3993
3994When the kernel prolog code is being emitted it is known whether the scratch V#
3995described above is actually used. If it is, the prolog code must set it up by
3996copying the Private Segment Buffer to the scratch V# registers and then adding
3997the Private Segment Wavefront Offset to the queue base address in the V#. The
3998result is a V# with a base address pointing to the beginning of the wavefront
3999scratch backing memory.
4000
4001The Private Segment Buffer is always requested, but the Private Segment
4002Wavefront Offset is only requested if it is used (see
4003:ref:`amdgpu-amdhsa-initial-kernel-execution-state`).
4004
4005.. _amdgpu-amdhsa-memory-model:
4006
4007Memory Model
4008~~~~~~~~~~~~
4009
4010This section describes the mapping of LLVM memory model onto AMDGPU machine code
4011(see :ref:`memmodel`).
4012
4013The AMDGPU backend supports the memory synchronization scopes specified in
4014:ref:`amdgpu-memory-scopes`.
4015
4016The code sequences used to implement the memory model are defined in table
4017:ref:`amdgpu-amdhsa-memory-model-code-sequences-gfx6-gfx10-table`.
4018
4019The sequences specify the order of instructions that a single thread must
4020execute. The ``s_waitcnt`` and ``buffer_wbinvl1_vol`` are defined with respect
4021to other memory instructions executed by the same thread. This allows them to be
4022moved earlier or later which can allow them to be combined with other instances
4023of the same instruction, or hoisted/sunk out of loops to improve
4024performance. Only the instructions related to the memory model are given;
4025additional ``s_waitcnt`` instructions are required to ensure registers are
4026defined before being used. These may be able to be combined with the memory
4027model ``s_waitcnt`` instructions as described above.
4028
4029The AMDGPU backend supports the following memory models:
4030
4031  HSA Memory Model [HSA]_
4032    The HSA memory model uses a single happens-before relation for all address
4033    spaces (see :ref:`amdgpu-address-spaces`).
4034  OpenCL Memory Model [OpenCL]_
4035    The OpenCL memory model which has separate happens-before relations for the
4036    global and local address spaces. Only a fence specifying both global and
4037    local address space, and seq_cst instructions join the relationships. Since
4038    the LLVM ``memfence`` instruction does not allow an address space to be
4039    specified the OpenCL fence has to conservatively assume both local and
4040    global address space was specified. However, optimizations can often be
4041    done to eliminate the additional ``s_waitcnt`` instructions when there are
4042    no intervening memory instructions which access the corresponding address
4043    space. The code sequences in the table indicate what can be omitted for the
4044    OpenCL memory. The target triple environment is used to determine if the
4045    source language is OpenCL (see :ref:`amdgpu-opencl`).
4046
4047``ds/flat_load/store/atomic`` instructions to local memory are termed LDS
4048operations.
4049
4050``buffer/global/flat_load/store/atomic`` instructions to global memory are
4051termed vector memory operations.
4052
4053For GFX6-GFX9:
4054
4055* Each agent has multiple shader arrays (SA).
4056* Each SA has multiple compute units (CU).
4057* Each CU has multiple SIMDs that execute wavefronts.
4058* The wavefronts for a single work-group are executed in the same CU but may be
4059  executed by different SIMDs.
4060* Each CU has a single LDS memory shared by the wavefronts of the work-groups
4061  executing on it.
4062* All LDS operations of a CU are performed as wavefront wide operations in a
4063  global order and involve no caching. Completion is reported to a wavefront in
4064  execution order.
4065* The LDS memory has multiple request queues shared by the SIMDs of a
4066  CU. Therefore, the LDS operations performed by different wavefronts of a
4067  work-group can be reordered relative to each other, which can result in
4068  reordering the visibility of vector memory operations with respect to LDS
4069  operations of other wavefronts in the same work-group. A ``s_waitcnt
4070  lgkmcnt(0)`` is required to ensure synchronization between LDS operations and
4071  vector memory operations between wavefronts of a work-group, but not between
4072  operations performed by the same wavefront.
4073* The vector memory operations are performed as wavefront wide operations and
4074  completion is reported to a wavefront in execution order. The exception is
4075  that for GFX7-GFX9 ``flat_load/store/atomic`` instructions can report out of
4076  vector memory order if they access LDS memory, and out of LDS operation order
4077  if they access global memory.
4078* The vector memory operations access a single vector L1 cache shared by all
4079  SIMDs a CU. Therefore, no special action is required for coherence between the
4080  lanes of a single wavefront, or for coherence between wavefronts in the same
4081  work-group. A ``buffer_wbinvl1_vol`` is required for coherence between
4082  wavefronts executing in different work-groups as they may be executing on
4083  different CUs.
4084* The scalar memory operations access a scalar L1 cache shared by all wavefronts
4085  on a group of CUs. The scalar and vector L1 caches are not coherent. However,
4086  scalar operations are used in a restricted way so do not impact the memory
4087  model. See :ref:`amdgpu-address-spaces`.
4088* The vector and scalar memory operations use an L2 cache shared by all CUs on
4089  the same agent.
4090* The L2 cache has independent channels to service disjoint ranges of virtual
4091  addresses.
4092* Each CU has a separate request queue per channel. Therefore, the vector and
4093  scalar memory operations performed by wavefronts executing in different
4094  work-groups (which may be executing on different CUs) of an agent can be
4095  reordered relative to each other. A ``s_waitcnt vmcnt(0)`` is required to
4096  ensure synchronization between vector memory operations of different CUs. It
4097  ensures a previous vector memory operation has completed before executing a
4098  subsequent vector memory or LDS operation and so can be used to meet the
4099  requirements of acquire and release.
4100* The L2 cache can be kept coherent with other agents on some targets, or ranges
4101  of virtual addresses can be set up to bypass it to ensure system coherence.
4102
4103For GFX10:
4104
4105* Each agent has multiple shader arrays (SA).
4106* Each SA has multiple work-group processors (WGP).
4107* Each WGP has multiple compute units (CU).
4108* Each CU has multiple SIMDs that execute wavefronts.
4109* The wavefronts for a single work-group are executed in the same
4110  WGP. In CU wavefront execution mode the wavefronts may be executed by
4111  different SIMDs in the same CU. In WGP wavefront execution mode the
4112  wavefronts may be executed by different SIMDs in different CUs in the same
4113  WGP.
4114* Each WGP has a single LDS memory shared by the wavefronts of the work-groups
4115  executing on it.
4116* All LDS operations of a WGP are performed as wavefront wide operations in a
4117  global order and involve no caching. Completion is reported to a wavefront in
4118  execution order.
4119* The LDS memory has multiple request queues shared by the SIMDs of a
4120  WGP. Therefore, the LDS operations performed by different wavefronts of a
4121  work-group can be reordered relative to each other, which can result in
4122  reordering the visibility of vector memory operations with respect to LDS
4123  operations of other wavefronts in the same work-group. A ``s_waitcnt
4124  lgkmcnt(0)`` is required to ensure synchronization between LDS operations and
4125  vector memory operations between wavefronts of a work-group, but not between
4126  operations performed by the same wavefront.
4127* The vector memory operations are performed as wavefront wide operations.
4128  Completion of load/store/sample operations are reported to a wavefront in
4129  execution order of other load/store/sample operations performed by that
4130  wavefront.
4131* The vector memory operations access a vector L0 cache. There is a single L0
4132  cache per CU. Each SIMD of a CU accesses the same L0 cache. Therefore, no
4133  special action is required for coherence between the lanes of a single
4134  wavefront. However, a ``BUFFER_GL0_INV`` is required for coherence between
4135  wavefronts executing in the same work-group as they may be executing on SIMDs
4136  of different CUs that access different L0s. A ``BUFFER_GL0_INV`` is also
4137  required for coherence between wavefronts executing in different work-groups
4138  as they may be executing on different WGPs.
4139* The scalar memory operations access a scalar L0 cache shared by all wavefronts
4140  on a WGP. The scalar and vector L0 caches are not coherent. However, scalar
4141  operations are used in a restricted way so do not impact the memory model. See
4142  :ref:`amdgpu-address-spaces`.
4143* The vector and scalar memory L0 caches use an L1 cache shared by all WGPs on
4144  the same SA. Therefore, no special action is required for coherence between
4145  the wavefronts of a single work-group. However, a ``BUFFER_GL1_INV`` is
4146  required for coherence between wavefronts executing in different work-groups
4147  as they may be executing on different SAs that access different L1s.
4148* The L1 caches have independent quadrants to service disjoint ranges of virtual
4149  addresses.
4150* Each L0 cache has a separate request queue per L1 quadrant. Therefore, the
4151  vector and scalar memory operations performed by different wavefronts, whether
4152  executing in the same or different work-groups (which may be executing on
4153  different CUs accessing different L0s), can be reordered relative to each
4154  other. A ``s_waitcnt vmcnt(0) & vscnt(0)`` is required to ensure
4155  synchronization between vector memory operations of different wavefronts. It
4156  ensures a previous vector memory operation has completed before executing a
4157  subsequent vector memory or LDS operation and so can be used to meet the
4158  requirements of acquire, release and sequential consistency.
4159* The L1 caches use an L2 cache shared by all SAs on the same agent.
4160* The L2 cache has independent channels to service disjoint ranges of virtual
4161  addresses.
4162* Each L1 quadrant of a single SA accesses a different L2 channel. Each L1
4163  quadrant has a separate request queue per L2 channel. Therefore, the vector
4164  and scalar memory operations performed by wavefronts executing in different
4165  work-groups (which may be executing on different SAs) of an agent can be
4166  reordered relative to each other. A ``s_waitcnt vmcnt(0) & vscnt(0)`` is
4167  required to ensure synchronization between vector memory operations of
4168  different SAs. It ensures a previous vector memory operation has completed
4169  before executing a subsequent vector memory and so can be used to meet the
4170  requirements of acquire, release and sequential consistency.
4171* The L2 cache can be kept coherent with other agents on some targets, or ranges
4172  of virtual addresses can be set up to bypass it to ensure system coherence.
4173
4174Private address space uses ``buffer_load/store`` using the scratch V#
4175(GFX6-GFX8), or ``scratch_load/store`` (GFX9-GFX10). Since only a single thread
4176is accessing the memory, atomic memory orderings are not meaningful, and all
4177accesses are treated as non-atomic.
4178
4179Constant address space uses ``buffer/global_load`` instructions (or equivalent
4180scalar memory instructions). Since the constant address space contents do not
4181change during the execution of a kernel dispatch it is not legal to perform
4182stores, and atomic memory orderings are not meaningful, and all access are
4183treated as non-atomic.
4184
4185A memory synchronization scope wider than work-group is not meaningful for the
4186group (LDS) address space and is treated as work-group.
4187
4188The memory model does not support the region address space which is treated as
4189non-atomic.
4190
4191Acquire memory ordering is not meaningful on store atomic instructions and is
4192treated as non-atomic.
4193
4194Release memory ordering is not meaningful on load atomic instructions and is
4195treated a non-atomic.
4196
4197Acquire-release memory ordering is not meaningful on load or store atomic
4198instructions and is treated as acquire and release respectively.
4199
4200AMDGPU backend only uses scalar memory operations to access memory that is
4201proven to not change during the execution of the kernel dispatch. This includes
4202constant address space and global address space for program scope const
4203variables. Therefore, the kernel machine code does not have to maintain the
4204scalar L1 cache to ensure it is coherent with the vector L1 cache. The scalar
4205and vector L1 caches are invalidated between kernel dispatches by CP since
4206constant address space data may change between kernel dispatch executions. See
4207:ref:`amdgpu-address-spaces`.
4208
4209The one exception is if scalar writes are used to spill SGPR registers. In this
4210case the AMDGPU backend ensures the memory location used to spill is never
4211accessed by vector memory operations at the same time. If scalar writes are used
4212then a ``s_dcache_wb`` is inserted before the ``s_endpgm`` and before a function
4213return since the locations may be used for vector memory instructions by a
4214future wavefront that uses the same scratch area, or a function call that
4215creates a frame at the same address, respectively. There is no need for a
4216``s_dcache_inv`` as all scalar writes are write-before-read in the same thread.
4217
4218For GFX6-GFX9, scratch backing memory (which is used for the private address
4219space) is accessed with MTYPE NC_NV (non-coherent non-volatile). Since the
4220private address space is only accessed by a single thread, and is always
4221write-before-read, there is never a need to invalidate these entries from the L1
4222cache. Hence all cache invalidates are done as ``*_vol`` to only invalidate the
4223volatile cache lines.
4224
4225For GFX10, scratch backing memory (which is used for the private address space)
4226is accessed with MTYPE NC (non-coherent). Since the private address space is
4227only accessed by a single thread, and is always write-before-read, there is
4228never a need to invalidate these entries from the L0 or L1 caches.
4229
4230For GFX10, wavefronts are executed in native mode with in-order reporting of
4231loads and sample instructions. In this mode vmcnt reports completion of load,
4232atomic with return and sample instructions in order, and the vscnt reports the
4233completion of store and atomic without return in order. See ``MEM_ORDERED``
4234field in :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`.
4235
4236In GFX10, wavefronts can be executed in WGP or CU wavefront execution mode:
4237
4238* In WGP wavefront execution mode the wavefronts of a work-group are executed
4239  on the SIMDs of both CUs of the WGP. Therefore, explicit management of the per
4240  CU L0 caches is required for work-group synchronization. Also accesses to L1
4241  at work-group scope need to be explicitly ordered as the accesses from
4242  different CUs are not ordered.
4243* In CU wavefront execution mode the wavefronts of a work-group are executed on
4244  the SIMDs of a single CU of the WGP. Therefore, all global memory access by
4245  the work-group access the same L0 which in turn ensures L1 accesses are
4246  ordered and so do not require explicit management of the caches for
4247  work-group synchronization.
4248
4249See ``WGP_MODE`` field in
4250:ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table` and
4251:ref:`amdgpu-target-features`.
4252
4253On dGPU the kernarg backing memory is accessed as UC (uncached) to avoid needing
4254to invalidate the L2 cache. For GFX6-GFX9, this also causes it to be treated as
4255non-volatile and so is not invalidated by ``*_vol``. On APU it is accessed as CC
4256(cache coherent) and so the L2 cache will be coherent with the CPU and other
4257agents.
4258
4259  .. table:: AMDHSA Memory Model Code Sequences GFX6-GFX10
4260     :name: amdgpu-amdhsa-memory-model-code-sequences-gfx6-gfx10-table
4261
4262     ============ ============ ============== ========== =============================== ==================================
4263     LLVM Instr   LLVM Memory  LLVM Memory    AMDGPU     AMDGPU Machine Code             AMDGPU Machine Code
4264                  Ordering     Sync Scope     Address    GFX6-9                          GFX10
4265                                              Space
4266     ============ ============ ============== ========== =============================== ==================================
4267     **Non-Atomic**
4268     ----------------------------------------------------------------------------------------------------------------------
4269     load         *none*       *none*         - global   - !volatile & !nontemporal      - !volatile & !nontemporal
4270                                              - generic
4271                                              - private    1. buffer/global/flat_load      1. buffer/global/flat_load
4272                                              - constant
4273                                                         - volatile & !nontemporal       - volatile & !nontemporal
4274
4275                                                           1. buffer/global/flat_load      1. buffer/global/flat_load
4276                                                              glc=1                           glc=1 dlc=1
4277
4278                                                         - nontemporal                   - nontemporal
4279
4280                                                           1. buffer/global/flat_load      1. buffer/global/flat_load
4281                                                              glc=1 slc=1                     slc=1
4282
4283     load         *none*       *none*         - local    1. ds_load                      1. ds_load
4284     store        *none*       *none*         - global   - !nontemporal                  - !nontemporal
4285                                              - generic
4286                                              - private    1. buffer/global/flat_store     1. buffer/global/flat_store
4287                                              - constant
4288                                                         - nontemporal                   - nontemporal
4289
4290                                                           1. buffer/global/flat_store      1. buffer/global/flat_store
4291                                                              glc=1 slc=1                      slc=1
4292
4293     store        *none*       *none*         - local    1. ds_store                     1. ds_store
4294     **Unordered Atomic**
4295     ----------------------------------------------------------------------------------------------------------------------
4296     load atomic  unordered    *any*          *any*      *Same as non-atomic*.           *Same as non-atomic*.
4297     store atomic unordered    *any*          *any*      *Same as non-atomic*.           *Same as non-atomic*.
4298     atomicrmw    unordered    *any*          *any*      *Same as monotonic              *Same as monotonic
4299                                                         atomic*.                        atomic*.
4300     **Monotonic Atomic**
4301     ----------------------------------------------------------------------------------------------------------------------
4302     load atomic  monotonic    - singlethread - global   1. buffer/global/flat_load      1. buffer/global/flat_load
4303                               - wavefront    - generic
4304     load atomic  monotonic    - workgroup    - global   1. buffer/global/flat_load      1. buffer/global/flat_load
4305                                              - generic                                     glc=1
4306
4307                                                                                           - If CU wavefront execution mode, omit glc=1.
4308
4309     load atomic  monotonic    - singlethread - local    1. ds_load                      1. ds_load
4310                               - wavefront
4311                               - workgroup
4312     load atomic  monotonic    - agent        - global   1. buffer/global/flat_load      1. buffer/global/flat_load
4313                               - system       - generic     glc=1                           glc=1 dlc=1
4314     store atomic monotonic    - singlethread - global   1. buffer/global/flat_store     1. buffer/global/flat_store
4315                               - wavefront    - generic
4316                               - workgroup
4317                               - agent
4318                               - system
4319     store atomic monotonic    - singlethread - local    1. ds_store                     1. ds_store
4320                               - wavefront
4321                               - workgroup
4322     atomicrmw    monotonic    - singlethread - global   1. buffer/global/flat_atomic    1. buffer/global/flat_atomic
4323                               - wavefront    - generic
4324                               - workgroup
4325                               - agent
4326                               - system
4327     atomicrmw    monotonic    - singlethread - local    1. ds_atomic                    1. ds_atomic
4328                               - wavefront
4329                               - workgroup
4330     **Acquire Atomic**
4331     ----------------------------------------------------------------------------------------------------------------------
4332     load atomic  acquire      - singlethread - global   1. buffer/global/ds/flat_load   1. buffer/global/ds/flat_load
4333                               - wavefront    - local
4334                                              - generic
4335     load atomic  acquire      - workgroup    - global   1. buffer/global/flat_load      1. buffer/global_load glc=1
4336
4337                                                                                           - If CU wavefront execution mode, omit glc=1.
4338
4339                                                                                         2. s_waitcnt vmcnt(0)
4340
4341                                                                                           - If CU wavefront execution mode, omit.
4342                                                                                           - Must happen before
4343                                                                                             the following buffer_gl0_inv
4344                                                                                             and before any following
4345                                                                                             global/generic
4346                                                                                             load/load
4347                                                                                             atomic/store/store
4348                                                                                             atomic/atomicrmw.
4349
4350                                                                                         3. buffer_gl0_inv
4351
4352                                                                                           - If CU wavefront execution mode, omit.
4353                                                                                           - Ensures that
4354                                                                                             following
4355                                                                                             loads will not see
4356                                                                                             stale data.
4357
4358     load atomic  acquire      - workgroup    - local    1. ds_load                      1. ds_load
4359                                                         2. s_waitcnt lgkmcnt(0)         2. s_waitcnt lgkmcnt(0)
4360
4361                                                           - If OpenCL, omit.              - If OpenCL, omit.
4362                                                           - Must happen before            - Must happen before
4363                                                             any following                   the following buffer_gl0_inv
4364                                                             global/generic                  and before any following
4365                                                             load/load                       global/generic load/load
4366                                                             atomic/store/store              atomic/store/store
4367                                                             atomic/atomicrmw.               atomic/atomicrmw.
4368                                                           - Ensures any                   - Ensures any
4369                                                             following global                following global
4370                                                             data read is no                 data read is no
4371                                                             older than the load             older than the load
4372                                                             atomic value being              atomic value being
4373                                                             acquired.                       acquired.
4374
4375                                                                                         3. buffer_gl0_inv
4376
4377                                                                                           - If CU wavefront execution mode, omit.
4378                                                                                           - If OpenCL, omit.
4379                                                                                           - Ensures that
4380                                                                                             following
4381                                                                                             loads will not see
4382                                                                                             stale data.
4383
4384     load atomic  acquire      - workgroup    - generic  1. flat_load                    1. flat_load glc=1
4385
4386                                                                                           - If CU wavefront execution mode, omit glc=1.
4387
4388                                                         2. s_waitcnt lgkmcnt(0)         2. s_waitcnt lgkmcnt(0) &
4389                                                                                            vmcnt(0)
4390
4391                                                                                           - If CU wavefront execution mode, omit vmcnt.
4392                                                           - If OpenCL, omit.              - If OpenCL, omit
4393                                                                                             lgkmcnt(0).
4394                                                           - Must happen before            - Must happen before
4395                                                             any following                   the following
4396                                                             global/generic                  buffer_gl0_inv and any
4397                                                             load/load                       following global/generic
4398                                                             atomic/store/store              load/load
4399                                                             atomic/atomicrmw.               atomic/store/store
4400                                                                                             atomic/atomicrmw.
4401                                                           - Ensures any                   - Ensures any
4402                                                             following global                following global
4403                                                             data read is no                 data read is no
4404                                                             older than the load             older than the load
4405                                                             atomic value being              atomic value being
4406                                                             acquired.                       acquired.
4407
4408                                                                                         3. buffer_gl0_inv
4409
4410                                                                                           - If CU wavefront execution mode, omit.
4411                                                                                           - Ensures that
4412                                                                                             following
4413                                                                                             loads will not see
4414                                                                                             stale data.
4415
4416     load atomic  acquire      - agent        - global   1. buffer/global/flat_load      1. buffer/global_load
4417                               - system                     glc=1                           glc=1 dlc=1
4418                                                         2. s_waitcnt vmcnt(0)           2. s_waitcnt vmcnt(0)
4419
4420                                                           - Must happen before            - Must happen before
4421                                                             following                       following
4422                                                             buffer_wbinvl1_vol.             buffer_gl*_inv.
4423                                                           - Ensures the load              - Ensures the load
4424                                                             has completed                   has completed
4425                                                             before invalidating             before invalidating
4426                                                             the cache.                      the caches.
4427
4428                                                         3. buffer_wbinvl1_vol           3. buffer_gl0_inv;
4429                                                                                            buffer_gl1_inv
4430
4431                                                           - Must happen before            - Must happen before
4432                                                             any following                   any following
4433                                                             global/generic                  global/generic
4434                                                             load/load                       load/load
4435                                                             atomic/atomicrmw.               atomic/atomicrmw.
4436                                                           - Ensures that                  - Ensures that
4437                                                             following                       following
4438                                                             loads will not see              loads will not see
4439                                                             stale global data.              stale global data.
4440
4441     load atomic  acquire      - agent        - generic  1. flat_load glc=1              1. flat_load glc=1 dlc=1
4442                               - system                  2. s_waitcnt vmcnt(0) &         2. s_waitcnt vmcnt(0) &
4443                                                            lgkmcnt(0)                      lgkmcnt(0)
4444
4445                                                           - If OpenCL omit                - If OpenCL omit
4446                                                             lgkmcnt(0).                     lgkmcnt(0).
4447                                                           - Must happen before            - Must happen before
4448                                                             following                       following
4449                                                             buffer_wbinvl1_vol.             buffer_gl*_invl.
4450                                                           - Ensures the flat_load         - Ensures the flat_load
4451                                                             has completed                   has completed
4452                                                             before invalidating             before invalidating
4453                                                             the cache.                      the caches.
4454
4455                                                         3. buffer_wbinvl1_vol           3. buffer_gl0_inv;
4456                                                                                            buffer_gl1_inv
4457
4458                                                           - Must happen before            - Must happen before
4459                                                             any following                   any following
4460                                                             global/generic                  global/generic
4461                                                             load/load                       load/load
4462                                                             atomic/atomicrmw.               atomic/atomicrmw.
4463                                                           - Ensures that                  - Ensures that
4464                                                             following loads                 following loads
4465                                                             will not see stale              will not see stale
4466                                                             global data.                    global data.
4467
4468     atomicrmw    acquire      - singlethread - global   1. buffer/global/ds/flat_atomic 1. buffer/global/ds/flat_atomic
4469                               - wavefront    - local
4470                                              - generic
4471     atomicrmw    acquire      - workgroup    - global   1. buffer/global/flat_atomic    1. buffer/global_atomic
4472                                                                                         2. s_waitcnt vm/vscnt(0)
4473
4474                                                                                           - If CU wavefront execution mode, omit.
4475                                                                                           - Use vmcnt if atomic with
4476                                                                                             return and vscnt if atomic
4477                                                                                             with no-return.
4478                                                                                           - Must happen before
4479                                                                                             the following buffer_gl0_inv
4480                                                                                             and before any following
4481                                                                                             global/generic
4482                                                                                             load/load
4483                                                                                             atomic/store/store
4484                                                                                             atomic/atomicrmw.
4485
4486                                                                                         3. buffer_gl0_inv
4487
4488                                                                                           - If CU wavefront execution mode, omit.
4489                                                                                           - Ensures that
4490                                                                                             following
4491                                                                                             loads will not see
4492                                                                                             stale data.
4493
4494     atomicrmw    acquire      - workgroup    - local    1. ds_atomic                    1. ds_atomic
4495                                                         2. waitcnt lgkmcnt(0)           2. waitcnt lgkmcnt(0)
4496
4497                                                           - If OpenCL, omit.              - If OpenCL, omit.
4498                                                           - Must happen before            - Must happen before
4499                                                             any following                   the following
4500                                                             global/generic                  buffer_gl0_inv.
4501                                                             load/load
4502                                                             atomic/store/store
4503                                                             atomic/atomicrmw.
4504                                                           - Ensures any                   - Ensures any
4505                                                             following global                following global
4506                                                             data read is no                 data read is no
4507                                                             older than the                  older than the
4508                                                             atomicrmw value                 atomicrmw value
4509                                                             being acquired.                 being acquired.
4510
4511                                                                                         3. buffer_gl0_inv
4512
4513                                                                                           - If OpenCL omit.
4514                                                                                           - Ensures that
4515                                                                                             following
4516                                                                                             loads will not see
4517                                                                                             stale data.
4518
4519     atomicrmw    acquire      - workgroup    - generic  1. flat_atomic                  1. flat_atomic
4520                                                         2. waitcnt lgkmcnt(0)           2. waitcnt lgkmcnt(0) &
4521                                                                                            vm/vscnt(0)
4522
4523                                                                                           - If CU wavefront execution mode, omit vm/vscnt.
4524                                                           - If OpenCL, omit.              - If OpenCL, omit
4525                                                                                             waitcnt lgkmcnt(0)..
4526                                                                                           - Use vmcnt if atomic with
4527                                                                                             return and vscnt if atomic
4528                                                                                             with no-return.
4529                                                                                             waitcnt lgkmcnt(0).
4530                                                           - Must happen before            - Must happen before
4531                                                             any following                   the following
4532                                                             global/generic                  buffer_gl0_inv.
4533                                                             load/load
4534                                                             atomic/store/store
4535                                                             atomic/atomicrmw.
4536                                                           - Ensures any                   - Ensures any
4537                                                             following global                following global
4538                                                             data read is no                 data read is no
4539                                                             older than the                  older than the
4540                                                             atomicrmw value                 atomicrmw value
4541                                                             being acquired.                 being acquired.
4542
4543                                                                                         3. buffer_gl0_inv
4544
4545                                                                                           - If CU wavefront execution mode, omit.
4546                                                                                           - Ensures that
4547                                                                                             following
4548                                                                                             loads will not see
4549                                                                                             stale data.
4550
4551     atomicrmw    acquire      - agent        - global   1. buffer/global/flat_atomic    1. buffer/global_atomic
4552                               - system                  2. s_waitcnt vmcnt(0)           2. s_waitcnt vm/vscnt(0)
4553
4554                                                                                           - Use vmcnt if atomic with
4555                                                                                             return and vscnt if atomic
4556                                                                                             with no-return.
4557                                                                                             waitcnt lgkmcnt(0).
4558                                                           - Must happen before            - Must happen before
4559                                                             following                       following
4560                                                             buffer_wbinvl1_vol.             buffer_gl*_inv.
4561                                                           - Ensures the                   - Ensures the
4562                                                             atomicrmw has                   atomicrmw has
4563                                                             completed before                completed before
4564                                                             invalidating the                invalidating the
4565                                                             cache.                          caches.
4566
4567                                                         3. buffer_wbinvl1_vol           3. buffer_gl0_inv;
4568                                                                                            buffer_gl1_inv
4569
4570                                                           - Must happen before            - Must happen before
4571                                                             any following                   any following
4572                                                             global/generic                  global/generic
4573                                                             load/load                       load/load
4574                                                             atomic/atomicrmw.               atomic/atomicrmw.
4575                                                           - Ensures that                  - Ensures that
4576                                                             following loads                 following loads
4577                                                             will not see stale              will not see stale
4578                                                             global data.                    global data.
4579
4580     atomicrmw    acquire      - agent        - generic  1. flat_atomic                  1. flat_atomic
4581                               - system                  2. s_waitcnt vmcnt(0) &         2. s_waitcnt vm/vscnt(0) &
4582                                                            lgkmcnt(0)                      lgkmcnt(0)
4583
4584                                                           - If OpenCL, omit               - If OpenCL, omit
4585                                                             lgkmcnt(0).                     lgkmcnt(0).
4586                                                                                           - Use vmcnt if atomic with
4587                                                                                             return and vscnt if atomic
4588                                                                                             with no-return.
4589                                                           - Must happen before            - Must happen before
4590                                                             following                       following
4591                                                             buffer_wbinvl1_vol.             buffer_gl*_inv.
4592                                                           - Ensures the                   - Ensures the
4593                                                             atomicrmw has                   atomicrmw has
4594                                                             completed before                completed before
4595                                                             invalidating the                invalidating the
4596                                                             cache.                          caches.
4597
4598                                                         3. buffer_wbinvl1_vol           3. buffer_gl0_inv;
4599                                                                                            buffer_gl1_inv
4600
4601                                                           - Must happen before            - Must happen before
4602                                                             any following                   any following
4603                                                             global/generic                  global/generic
4604                                                             load/load                       load/load
4605                                                             atomic/atomicrmw.               atomic/atomicrmw.
4606                                                           - Ensures that                  - Ensures that
4607                                                             following loads                 following loads
4608                                                             will not see stale              will not see stale
4609                                                             global data.                    global data.
4610
4611     fence        acquire      - singlethread *none*     *none*                          *none*
4612                               - wavefront
4613     fence        acquire      - workgroup    *none*     1. s_waitcnt lgkmcnt(0)         1. s_waitcnt lgkmcnt(0) &
4614                                                                                            vmcnt(0) & vscnt(0)
4615
4616                                                                                           - If CU wavefront execution mode, omit vmcnt and
4617                                                                                             vscnt.
4618                                                           - If OpenCL and                 - If OpenCL and
4619                                                             address space is                address space is
4620                                                             not generic, omit.              not generic, omit
4621                                                                                             lgkmcnt(0).
4622                                                                                           - If OpenCL and
4623                                                                                             address space is
4624                                                                                             local, omit
4625                                                                                             vmcnt(0) and vscnt(0).
4626                                                           - However, since LLVM           - However, since LLVM
4627                                                             currently has no                currently has no
4628                                                             address space on                address space on
4629                                                             the fence need to               the fence need to
4630                                                             conservatively                  conservatively
4631                                                             always generate. If             always generate. If
4632                                                             fence had an                    fence had an
4633                                                             address space then              address space then
4634                                                             set to address                  set to address
4635                                                             space of OpenCL                 space of OpenCL
4636                                                             fence flag, or to               fence flag, or to
4637                                                             generic if both                 generic if both
4638                                                             local and global                local and global
4639                                                             flags are                       flags are
4640                                                             specified.                      specified.
4641                                                           - Must happen after
4642                                                             any preceding
4643                                                             local/generic load
4644                                                             atomic/atomicrmw
4645                                                             with an equal or
4646                                                             wider sync scope
4647                                                             and memory ordering
4648                                                             stronger than
4649                                                             unordered (this is
4650                                                             termed the
4651                                                             fence-paired-atomic).
4652                                                           - Must happen before
4653                                                             any following
4654                                                             global/generic
4655                                                             load/load
4656                                                             atomic/store/store
4657                                                             atomic/atomicrmw.
4658                                                           - Ensures any
4659                                                             following global
4660                                                             data read is no
4661                                                             older than the
4662                                                             value read by the
4663                                                             fence-paired-atomic.
4664                                                                                           - Could be split into
4665                                                                                             separate s_waitcnt
4666                                                                                             vmcnt(0), s_waitcnt
4667                                                                                             vscnt(0) and s_waitcnt
4668                                                                                             lgkmcnt(0) to allow
4669                                                                                             them to be
4670                                                                                             independently moved
4671                                                                                             according to the
4672                                                                                             following rules.
4673                                                                                           - s_waitcnt vmcnt(0)
4674                                                                                             must happen after
4675                                                                                             any preceding
4676                                                                                             global/generic load
4677                                                                                             atomic/
4678                                                                                             atomicrmw-with-return-value
4679                                                                                             with an equal or
4680                                                                                             wider sync scope
4681                                                                                             and memory ordering
4682                                                                                             stronger than
4683                                                                                             unordered (this is
4684                                                                                             termed the
4685                                                                                             fence-paired-atomic).
4686                                                                                           - s_waitcnt vscnt(0)
4687                                                                                             must happen after
4688                                                                                             any preceding
4689                                                                                             global/generic
4690                                                                                             atomicrmw-no-return-value
4691                                                                                             with an equal or
4692                                                                                             wider sync scope
4693                                                                                             and memory ordering
4694                                                                                             stronger than
4695                                                                                             unordered (this is
4696                                                                                             termed the
4697                                                                                             fence-paired-atomic).
4698                                                                                           - s_waitcnt lgkmcnt(0)
4699                                                                                             must happen after
4700                                                                                             any preceding
4701                                                                                             local/generic load
4702                                                                                             atomic/atomicrmw
4703                                                                                             with an equal or
4704                                                                                             wider sync scope
4705                                                                                             and memory ordering
4706                                                                                             stronger than
4707                                                                                             unordered (this is
4708                                                                                             termed the
4709                                                                                             fence-paired-atomic).
4710                                                                                           - Must happen before
4711                                                                                             the following
4712                                                                                             buffer_gl0_inv.
4713                                                                                           - Ensures that the
4714                                                                                             fence-paired atomic
4715                                                                                             has completed
4716                                                                                             before invalidating
4717                                                                                             the
4718                                                                                             cache. Therefore
4719                                                                                             any following
4720                                                                                             locations read must
4721                                                                                             be no older than
4722                                                                                             the value read by
4723                                                                                             the
4724                                                                                             fence-paired-atomic.
4725
4726                                                                                         3. buffer_gl0_inv
4727
4728                                                                                           - If CU wavefront execution mode, omit.
4729                                                                                           - Ensures that
4730                                                                                             following
4731                                                                                             loads will not see
4732                                                                                             stale data.
4733
4734     fence        acquire      - agent        *none*     1. s_waitcnt lgkmcnt(0) &       1. s_waitcnt lgkmcnt(0) &
4735                               - system                     vmcnt(0)                        vmcnt(0) & vscnt(0)
4736
4737                                                           - If OpenCL and                 - If OpenCL and
4738                                                             address space is                address space is
4739                                                             not generic, omit               not generic, omit
4740                                                             lgkmcnt(0).                     lgkmcnt(0).
4741                                                                                           - If OpenCL and
4742                                                                                             address space is
4743                                                                                             local, omit
4744                                                                                             vmcnt(0) and vscnt(0).
4745                                                           - However, since LLVM           - However, since LLVM
4746                                                             currently has no                currently has no
4747                                                             address space on                address space on
4748                                                             the fence need to               the fence need to
4749                                                             conservatively                  conservatively
4750                                                             always generate                 always generate
4751                                                             (see comment for                (see comment for
4752                                                             previous fence).                previous fence).
4753                                                           - Could be split into
4754                                                             separate s_waitcnt
4755                                                             vmcnt(0) and
4756                                                             s_waitcnt
4757                                                             lgkmcnt(0) to allow
4758                                                             them to be
4759                                                             independently moved
4760                                                             according to the
4761                                                             following rules.
4762                                                           - s_waitcnt vmcnt(0)
4763                                                             must happen after
4764                                                             any preceding
4765                                                             global/generic load
4766                                                             atomic/atomicrmw
4767                                                             with an equal or
4768                                                             wider sync scope
4769                                                             and memory ordering
4770                                                             stronger than
4771                                                             unordered (this is
4772                                                             termed the
4773                                                             fence-paired-atomic).
4774                                                           - s_waitcnt lgkmcnt(0)
4775                                                             must happen after
4776                                                             any preceding
4777                                                             local/generic load
4778                                                             atomic/atomicrmw
4779                                                             with an equal or
4780                                                             wider sync scope
4781                                                             and memory ordering
4782                                                             stronger than
4783                                                             unordered (this is
4784                                                             termed the
4785                                                             fence-paired-atomic).
4786                                                           - Must happen before
4787                                                             the following
4788                                                             buffer_wbinvl1_vol.
4789                                                           - Ensures that the
4790                                                             fence-paired atomic
4791                                                             has completed
4792                                                             before invalidating
4793                                                             the
4794                                                             cache. Therefore
4795                                                             any following
4796                                                             locations read must
4797                                                             be no older than
4798                                                             the value read by
4799                                                             the
4800                                                             fence-paired-atomic.
4801                                                                                           - Could be split into
4802                                                                                             separate s_waitcnt
4803                                                                                             vmcnt(0), s_waitcnt
4804                                                                                             vscnt(0) and s_waitcnt
4805                                                                                             lgkmcnt(0) to allow
4806                                                                                             them to be
4807                                                                                             independently moved
4808                                                                                             according to the
4809                                                                                             following rules.
4810                                                                                           - s_waitcnt vmcnt(0)
4811                                                                                             must happen after
4812                                                                                             any preceding
4813                                                                                             global/generic load
4814                                                                                             atomic/
4815                                                                                             atomicrmw-with-return-value
4816                                                                                             with an equal or
4817                                                                                             wider sync scope
4818                                                                                             and memory ordering
4819                                                                                             stronger than
4820                                                                                             unordered (this is
4821                                                                                             termed the
4822                                                                                             fence-paired-atomic).
4823                                                                                           - s_waitcnt vscnt(0)
4824                                                                                             must happen after
4825                                                                                             any preceding
4826                                                                                             global/generic
4827                                                                                             atomicrmw-no-return-value
4828                                                                                             with an equal or
4829                                                                                             wider sync scope
4830                                                                                             and memory ordering
4831                                                                                             stronger than
4832                                                                                             unordered (this is
4833                                                                                             termed the
4834                                                                                             fence-paired-atomic).
4835                                                                                           - s_waitcnt lgkmcnt(0)
4836                                                                                             must happen after
4837                                                                                             any preceding
4838                                                                                             local/generic load
4839                                                                                             atomic/atomicrmw
4840                                                                                             with an equal or
4841                                                                                             wider sync scope
4842                                                                                             and memory ordering
4843                                                                                             stronger than
4844                                                                                             unordered (this is
4845                                                                                             termed the
4846                                                                                             fence-paired-atomic).
4847                                                                                           - Must happen before
4848                                                                                             the following
4849                                                                                             buffer_gl*_inv.
4850                                                                                           - Ensures that the
4851                                                                                             fence-paired atomic
4852                                                                                             has completed
4853                                                                                             before invalidating
4854                                                                                             the
4855                                                                                             caches. Therefore
4856                                                                                             any following
4857                                                                                             locations read must
4858                                                                                             be no older than
4859                                                                                             the value read by
4860                                                                                             the
4861                                                                                             fence-paired-atomic.
4862
4863                                                         2. buffer_wbinvl1_vol           2. buffer_gl0_inv;
4864                                                                                            buffer_gl1_inv
4865
4866                                                           - Must happen before any        - Must happen before any
4867                                                             following global/generic        following global/generic
4868                                                             load/load                       load/load
4869                                                             atomic/store/store              atomic/store/store
4870                                                             atomic/atomicrmw.               atomic/atomicrmw.
4871                                                           - Ensures that                  - Ensures that
4872                                                             following loads                 following loads
4873                                                             will not see stale              will not see stale
4874                                                             global data.                    global data.
4875
4876     **Release Atomic**
4877     ----------------------------------------------------------------------------------------------------------------------
4878     store atomic release      - singlethread - global   1. buffer/global/ds/flat_store  1. buffer/global/ds/flat_store
4879                               - wavefront    - local
4880                                              - generic
4881     store atomic release      - workgroup    - global   1. s_waitcnt lgkmcnt(0)         1. s_waitcnt lgkmcnt(0) &
4882                                                                                            vmcnt(0) & vscnt(0)
4883
4884                                                                                           - If CU wavefront execution mode, omit vmcnt and
4885                                                                                             vscnt.
4886                                                           - If OpenCL, omit.              - If OpenCL, omit
4887                                                                                             lgkmcnt(0).
4888                                                           - Must happen after
4889                                                             any preceding
4890                                                             local/generic
4891                                                             load/store/load
4892                                                             atomic/store
4893                                                             atomic/atomicrmw.
4894                                                                                           - Could be split into
4895                                                                                             separate s_waitcnt
4896                                                                                             vmcnt(0), s_waitcnt
4897                                                                                             vscnt(0) and s_waitcnt
4898                                                                                             lgkmcnt(0) to allow
4899                                                                                             them to be
4900                                                                                             independently moved
4901                                                                                             according to the
4902                                                                                             following rules.
4903                                                                                           - s_waitcnt vmcnt(0)
4904                                                                                             must happen after
4905                                                                                             any preceding
4906                                                                                             global/generic load/load
4907                                                                                             atomic/
4908                                                                                             atomicrmw-with-return-value.
4909                                                                                           - s_waitcnt vscnt(0)
4910                                                                                             must happen after
4911                                                                                             any preceding
4912                                                                                             global/generic
4913                                                                                             store/store
4914                                                                                             atomic/
4915                                                                                             atomicrmw-no-return-value.
4916                                                                                           - s_waitcnt lgkmcnt(0)
4917                                                                                             must happen after
4918                                                                                             any preceding
4919                                                                                             local/generic
4920                                                                                             load/store/load
4921                                                                                             atomic/store
4922                                                                                             atomic/atomicrmw.
4923                                                           - Must happen before            - Must happen before
4924                                                             the following                   the following
4925                                                             store.                          store.
4926                                                           - Ensures that all              - Ensures that all
4927                                                             memory operations               memory operations
4928                                                             to local have                   have
4929                                                             completed before                completed before
4930                                                             performing the                  performing the
4931                                                             store that is being             store that is being
4932                                                             released.                       released.
4933
4934                                                         2. buffer/global/flat_store     2. buffer/global_store
4935     store atomic release      - workgroup    - local                                    1. waitcnt vmcnt(0) & vscnt(0)
4936
4937                                                                                           - If CU wavefront execution mode, omit.
4938                                                                                           - If OpenCL, omit.
4939                                                                                           - Could be split into
4940                                                                                             separate s_waitcnt
4941                                                                                             vmcnt(0) and s_waitcnt
4942                                                                                             vscnt(0) to allow
4943                                                                                             them to be
4944                                                                                             independently moved
4945                                                                                             according to the
4946                                                                                             following rules.
4947                                                                                           - s_waitcnt vmcnt(0)
4948                                                                                             must happen after
4949                                                                                             any preceding
4950                                                                                             global/generic load/load
4951                                                                                             atomic/
4952                                                                                             atomicrmw-with-return-value.
4953                                                                                           - s_waitcnt vscnt(0)
4954                                                                                             must happen after
4955                                                                                             any preceding
4956                                                                                             global/generic
4957                                                                                             store/store atomic/
4958                                                                                             atomicrmw-no-return-value.
4959                                                                                           - Must happen before
4960                                                                                             the following
4961                                                                                             store.
4962                                                                                           - Ensures that all
4963                                                                                             global memory
4964                                                                                             operations have
4965                                                                                             completed before
4966                                                                                             performing the
4967                                                                                             store that is being
4968                                                                                             released.
4969
4970                                                         1. ds_store                     2. ds_store
4971     store atomic release      - workgroup    - generic  1. s_waitcnt lgkmcnt(0)         1. s_waitcnt lgkmcnt(0) &
4972                                                                                            vmcnt(0) & vscnt(0)
4973
4974                                                                                           - If CU wavefront execution mode, omit vmcnt and
4975                                                                                             vscnt.
4976                                                           - If OpenCL, omit.              - If OpenCL, omit
4977                                                                                             lgkmcnt(0).
4978                                                           - Must happen after
4979                                                             any preceding
4980                                                             local/generic
4981                                                             load/store/load
4982                                                             atomic/store
4983                                                             atomic/atomicrmw.
4984                                                                                           - Could be split into
4985                                                                                             separate s_waitcnt
4986                                                                                             vmcnt(0), s_waitcnt
4987                                                                                             vscnt(0) and s_waitcnt
4988                                                                                             lgkmcnt(0) to allow
4989                                                                                             them to be
4990                                                                                             independently moved
4991                                                                                             according to the
4992                                                                                             following rules.
4993                                                                                           - s_waitcnt vmcnt(0)
4994                                                                                             must happen after
4995                                                                                             any preceding
4996                                                                                             global/generic load/load
4997                                                                                             atomic/
4998                                                                                             atomicrmw-with-return-value.
4999                                                                                           - s_waitcnt vscnt(0)
5000                                                                                             must happen after
5001                                                                                             any preceding
5002                                                                                             global/generic
5003                                                                                             store/store
5004                                                                                             atomic/
5005                                                                                             atomicrmw-no-return-value.
5006                                                                                           - s_waitcnt lgkmcnt(0)
5007                                                                                             must happen after
5008                                                                                             any preceding
5009                                                                                             local/generic load/store/load
5010                                                                                             atomic/store atomic/atomicrmw.
5011                                                           - Must happen before            - Must happen before
5012                                                             the following                   the following
5013                                                             store.                          store.
5014                                                           - Ensures that all              - Ensures that all
5015                                                             memory operations               memory operations
5016                                                             to local have                   have
5017                                                             completed before                completed before
5018                                                             performing the                  performing the
5019                                                             store that is being             store that is being
5020                                                             released.                       released.
5021
5022                                                         2. flat_store                   2. flat_store
5023     store atomic release      - agent        - global   1. s_waitcnt lgkmcnt(0) &         1. s_waitcnt lgkmcnt(0) &
5024                               - system       - generic     vmcnt(0)                          vmcnt(0) & vscnt(0)
5025
5026                                                           - If OpenCL, omit               - If OpenCL, omit
5027                                                             lgkmcnt(0).                     lgkmcnt(0).
5028                                                           - Could be split into           - Could be split into
5029                                                             separate s_waitcnt              separate s_waitcnt
5030                                                             vmcnt(0) and                    vmcnt(0), s_waitcnt vscnt(0)
5031                                                             s_waitcnt                       and s_waitcnt
5032                                                             lgkmcnt(0) to allow             lgkmcnt(0) to allow
5033                                                             them to be                      them to be
5034                                                             independently moved             independently moved
5035                                                             according to the                according to the
5036                                                             following rules.                following rules.
5037                                                           - s_waitcnt vmcnt(0)            - s_waitcnt vmcnt(0)
5038                                                             must happen after               must happen after
5039                                                             any preceding                   any preceding
5040                                                             global/generic                  global/generic
5041                                                             load/store/load                 load/load
5042                                                             atomic/store                    atomic/
5043                                                             atomic/atomicrmw.               atomicrmw-with-return-value.
5044                                                                                           - s_waitcnt vscnt(0)
5045                                                                                             must happen after
5046                                                                                             any preceding
5047                                                                                             global/generic
5048                                                                                             store/store atomic/
5049                                                                                             atomicrmw-no-return-value.
5050                                                           - s_waitcnt lgkmcnt(0)          - s_waitcnt lgkmcnt(0)
5051                                                             must happen after               must happen after
5052                                                             any preceding                   any preceding
5053                                                             local/generic                   local/generic
5054                                                             load/store/load                 load/store/load
5055                                                             atomic/store                    atomic/store
5056                                                             atomic/atomicrmw.               atomic/atomicrmw.
5057                                                           - Must happen before            - Must happen before
5058                                                             the following                   the following
5059                                                             store.                          store.
5060                                                           - Ensures that all              - Ensures that all
5061                                                             memory operations               memory operations
5062                                                             to memory have                  to memory have
5063                                                             completed before                completed before
5064                                                             performing the                  performing the
5065                                                             store that is being             store that is being
5066                                                             released.                       released.
5067
5068                                                         2. buffer/global/ds/flat_store  2. buffer/global/ds/flat_store
5069     atomicrmw    release      - singlethread - global   1. buffer/global/ds/flat_atomic 1. buffer/global/ds/flat_atomic
5070                               - wavefront    - local
5071                                              - generic
5072     atomicrmw    release      - workgroup    - global   1. s_waitcnt lgkmcnt(0)         1. s_waitcnt lgkmcnt(0) &
5073                                                                                            vmcnt(0) & vscnt(0)
5074
5075                                                                                           - If CU wavefront execution mode, omit vmcnt and
5076                                                                                             vscnt.
5077                                                           - If OpenCL, omit.
5078
5079                                                           - Must happen after
5080                                                             any preceding
5081                                                             local/generic
5082                                                             load/store/load
5083                                                             atomic/store
5084                                                             atomic/atomicrmw.
5085                                                                                           - Could be split into
5086                                                                                             separate s_waitcnt
5087                                                                                             vmcnt(0), s_waitcnt
5088                                                                                             vscnt(0) and s_waitcnt
5089                                                                                             lgkmcnt(0) to allow
5090                                                                                             them to be
5091                                                                                             independently moved
5092                                                                                             according to the
5093                                                                                             following rules.
5094                                                                                           - s_waitcnt vmcnt(0)
5095                                                                                             must happen after
5096                                                                                             any preceding
5097                                                                                             global/generic load/load
5098                                                                                             atomic/
5099                                                                                             atomicrmw-with-return-value.
5100                                                                                           - s_waitcnt vscnt(0)
5101                                                                                             must happen after
5102                                                                                             any preceding
5103                                                                                             global/generic
5104                                                                                             store/store
5105                                                                                             atomic/
5106                                                                                             atomicrmw-no-return-value.
5107                                                                                           - s_waitcnt lgkmcnt(0)
5108                                                                                             must happen after
5109                                                                                             any preceding
5110                                                                                             local/generic
5111                                                                                             load/store/load
5112                                                                                             atomic/store
5113                                                                                             atomic/atomicrmw.
5114                                                           - Must happen before            - Must happen before
5115                                                             the following                   the following
5116                                                             atomicrmw.                      atomicrmw.
5117                                                           - Ensures that all              - Ensures that all
5118                                                             memory operations               memory operations
5119                                                             to local have                   have
5120                                                             completed before                completed before
5121                                                             performing the                  performing the
5122                                                             atomicrmw that is               atomicrmw that is
5123                                                             being released.                 being released.
5124
5125                                                         2. buffer/global/flat_atomic    2. buffer/global_atomic
5126     atomicrmw    release      - workgroup    - local                                    1. waitcnt vmcnt(0) & vscnt(0)
5127
5128                                                                                           - If CU wavefront execution mode, omit.
5129                                                                                           - If OpenCL, omit.
5130                                                                                           - Could be split into
5131                                                                                             separate s_waitcnt
5132                                                                                             vmcnt(0) and s_waitcnt
5133                                                                                             vscnt(0) to allow
5134                                                                                             them to be
5135                                                                                             independently moved
5136                                                                                             according to the
5137                                                                                             following rules.
5138                                                                                           - s_waitcnt vmcnt(0)
5139                                                                                             must happen after
5140                                                                                             any preceding
5141                                                                                             global/generic load/load
5142                                                                                             atomic/
5143                                                                                             atomicrmw-with-return-value.
5144                                                                                           - s_waitcnt vscnt(0)
5145                                                                                             must happen after
5146                                                                                             any preceding
5147                                                                                             global/generic
5148                                                                                             store/store atomic/
5149                                                                                             atomicrmw-no-return-value.
5150                                                                                           - Must happen before
5151                                                                                             the following
5152                                                                                             store.
5153                                                                                           - Ensures that all
5154                                                                                             global memory
5155                                                                                             operations have
5156                                                                                             completed before
5157                                                                                             performing the
5158                                                                                             store that is being
5159                                                                                             released.
5160
5161                                                         1. ds_atomic                    2. ds_atomic
5162     atomicrmw    release      - workgroup    - generic  1. s_waitcnt lgkmcnt(0)         1. s_waitcnt lgkmcnt(0) &
5163                                                                                            vmcnt(0) & vscnt(0)
5164
5165                                                                                           - If CU wavefront execution mode, omit vmcnt and
5166                                                                                             vscnt.
5167                                                           - If OpenCL, omit.              - If OpenCL, omit
5168                                                                                             waitcnt lgkmcnt(0).
5169                                                           - Must happen after
5170                                                             any preceding
5171                                                             local/generic
5172                                                             load/store/load
5173                                                             atomic/store
5174                                                             atomic/atomicrmw.
5175                                                                                           - Could be split into
5176                                                                                             separate s_waitcnt
5177                                                                                             vmcnt(0), s_waitcnt
5178                                                                                             vscnt(0) and s_waitcnt
5179                                                                                             lgkmcnt(0) to allow
5180                                                                                             them to be
5181                                                                                             independently moved
5182                                                                                             according to the
5183                                                                                             following rules.
5184                                                                                           - s_waitcnt vmcnt(0)
5185                                                                                             must happen after
5186                                                                                             any preceding
5187                                                                                             global/generic load/load
5188                                                                                             atomic/
5189                                                                                             atomicrmw-with-return-value.
5190                                                                                           - s_waitcnt vscnt(0)
5191                                                                                             must happen after
5192                                                                                             any preceding
5193                                                                                             global/generic
5194                                                                                             store/store
5195                                                                                             atomic/
5196                                                                                             atomicrmw-no-return-value.
5197                                                                                           - s_waitcnt lgkmcnt(0)
5198                                                                                             must happen after
5199                                                                                             any preceding
5200                                                                                             local/generic load/store/load
5201                                                                                             atomic/store atomic/atomicrmw.
5202                                                           - Must happen before            - Must happen before
5203                                                             the following                   the following
5204                                                             atomicrmw.                      atomicrmw.
5205                                                           - Ensures that all              - Ensures that all
5206                                                             memory operations               memory operations
5207                                                             to local have                   have
5208                                                             completed before                completed before
5209                                                             performing the                  performing the
5210                                                             atomicrmw that is               atomicrmw that is
5211                                                             being released.                 being released.
5212
5213                                                         2. flat_atomic                  2. flat_atomic
5214     atomicrmw    release      - agent        - global   1. s_waitcnt lgkmcnt(0) &       1. s_waitcnt lkkmcnt(0) &
5215                               - system       - generic     vmcnt(0)                         vmcnt(0) & vscnt(0)
5216
5217                                                           - If OpenCL, omit               - If OpenCL, omit
5218                                                             lgkmcnt(0).                     lgkmcnt(0).
5219                                                           - Could be split into           - Could be split into
5220                                                             separate s_waitcnt              separate s_waitcnt
5221                                                             vmcnt(0) and                    vmcnt(0), s_waitcnt
5222                                                             s_waitcnt                       vscnt(0) and s_waitcnt
5223                                                             lgkmcnt(0) to allow             lgkmcnt(0) to allow
5224                                                             them to be                      them to be
5225                                                             independently moved             independently moved
5226                                                             according to the                according to the
5227                                                             following rules.                following rules.
5228                                                           - s_waitcnt vmcnt(0)            - s_waitcnt vmcnt(0)
5229                                                             must happen after               must happen after
5230                                                             any preceding                   any preceding
5231                                                             global/generic                  global/generic
5232                                                             load/store/load                 load/load atomic/
5233                                                             atomic/store                    atomicrmw-with-return-value.
5234                                                             atomic/atomicrmw.
5235                                                                                           - s_waitcnt vscnt(0)
5236                                                                                             must happen after
5237                                                                                             any preceding
5238                                                                                             global/generic
5239                                                                                             store/store atomic/
5240                                                                                             atomicrmw-no-return-value.
5241                                                           - s_waitcnt lgkmcnt(0)          - s_waitcnt lgkmcnt(0)
5242                                                             must happen after               must happen after
5243                                                             any preceding                   any preceding
5244                                                             local/generic                   local/generic
5245                                                             load/store/load                 load/store/load
5246                                                             atomic/store                    atomic/store
5247                                                             atomic/atomicrmw.               atomic/atomicrmw.
5248                                                           - Must happen before            - Must happen before
5249                                                             the following                   the following
5250                                                             atomicrmw.                      atomicrmw.
5251                                                           - Ensures that all              - Ensures that all
5252                                                             memory operations               memory operations
5253                                                             to global and local             to global and local
5254                                                             have completed                  have completed
5255                                                             before performing               before performing
5256                                                             the atomicrmw that              the atomicrmw that
5257                                                             is being released.              is being released.
5258
5259                                                         2. buffer/global/ds/flat_atomic 2. buffer/global/ds/flat_atomic
5260     fence        release      - singlethread *none*     *none*                          *none*
5261                               - wavefront
5262     fence        release      - workgroup    *none*     1. s_waitcnt lgkmcnt(0)         1. s_waitcnt lgkmcnt(0) &
5263                                                                                            vmcnt(0) & vscnt(0)
5264
5265                                                                                           - If CU wavefront execution mode, omit vmcnt and
5266                                                                                             vscnt.
5267                                                           - If OpenCL and                 - If OpenCL and
5268                                                             address space is                address space is
5269                                                             not generic, omit.              not generic, omit
5270                                                                                             lgkmcnt(0).
5271                                                                                           - If OpenCL and
5272                                                                                             address space is
5273                                                                                             local, omit
5274                                                                                             vmcnt(0) and vscnt(0).
5275                                                           - However, since LLVM           - However, since LLVM
5276                                                             currently has no                currently has no
5277                                                             address space on                address space on
5278                                                             the fence need to               the fence need to
5279                                                             conservatively                  conservatively
5280                                                             always generate. If             always generate. If
5281                                                             fence had an                    fence had an
5282                                                             address space then              address space then
5283                                                             set to address                  set to address
5284                                                             space of OpenCL                 space of OpenCL
5285                                                             fence flag, or to               fence flag, or to
5286                                                             generic if both                 generic if both
5287                                                             local and global                local and global
5288                                                             flags are                       flags are
5289                                                             specified.                      specified.
5290                                                           - Must happen after
5291                                                             any preceding
5292                                                             local/generic
5293                                                             load/load
5294                                                             atomic/store/store
5295                                                             atomic/atomicrmw.
5296                                                                                           - Could be split into
5297                                                                                             separate s_waitcnt
5298                                                                                             vmcnt(0), s_waitcnt
5299                                                                                             vscnt(0) and s_waitcnt
5300                                                                                             lgkmcnt(0) to allow
5301                                                                                             them to be
5302                                                                                             independently moved
5303                                                                                             according to the
5304                                                                                             following rules.
5305                                                                                           - s_waitcnt vmcnt(0)
5306                                                                                             must happen after
5307                                                                                             any preceding
5308                                                                                             global/generic
5309                                                                                             load/load
5310                                                                                             atomic/
5311                                                                                             atomicrmw-with-return-value.
5312                                                                                           - s_waitcnt vscnt(0)
5313                                                                                             must happen after
5314                                                                                             any preceding
5315                                                                                             global/generic
5316                                                                                             store/store atomic/
5317                                                                                             atomicrmw-no-return-value.
5318                                                                                           - s_waitcnt lgkmcnt(0)
5319                                                                                             must happen after
5320                                                                                             any preceding
5321                                                                                             local/generic
5322                                                                                             load/store/load
5323                                                                                             atomic/store atomic/
5324                                                                                             atomicrmw.
5325                                                           - Must happen before            - Must happen before
5326                                                             any following store             any following store
5327                                                             atomic/atomicrmw                atomic/atomicrmw
5328                                                             with an equal or                with an equal or
5329                                                             wider sync scope                wider sync scope
5330                                                             and memory ordering             and memory ordering
5331                                                             stronger than                   stronger than
5332                                                             unordered (this is              unordered (this is
5333                                                             termed the                      termed the
5334                                                             fence-paired-atomic).           fence-paired-atomic).
5335                                                           - Ensures that all              - Ensures that all
5336                                                             memory operations               memory operations
5337                                                             to local have                   have
5338                                                             completed before                completed before
5339                                                             performing the                  performing the
5340                                                             following                       following
5341                                                             fence-paired-atomic.            fence-paired-atomic.
5342
5343     fence        release      - agent        *none*     1. s_waitcnt lgkmcnt(0) &       1. s_waitcnt lgkmcnt(0) &
5344                               - system                     vmcnt(0)                        vmcnt(0) & vscnt(0)
5345
5346                                                           - If OpenCL and                 - If OpenCL and
5347                                                             address space is                address space is
5348                                                             not generic, omit               not generic, omit
5349                                                             lgkmcnt(0).                     lgkmcnt(0).
5350                                                           - If OpenCL and                 - If OpenCL and
5351                                                             address space is                address space is
5352                                                             local, omit                     local, omit
5353                                                             vmcnt(0).                       vmcnt(0) and vscnt(0).
5354                                                           - However, since LLVM           - However, since LLVM
5355                                                             currently has no                currently has no
5356                                                             address space on                address space on
5357                                                             the fence need to               the fence need to
5358                                                             conservatively                  conservatively
5359                                                             always generate. If             always generate. If
5360                                                             fence had an                    fence had an
5361                                                             address space then              address space then
5362                                                             set to address                  set to address
5363                                                             space of OpenCL                 space of OpenCL
5364                                                             fence flag, or to               fence flag, or to
5365                                                             generic if both                 generic if both
5366                                                             local and global                local and global
5367                                                             flags are                       flags are
5368                                                             specified.                      specified.
5369                                                           - Could be split into           - Could be split into
5370                                                             separate s_waitcnt              separate s_waitcnt
5371                                                             vmcnt(0) and                    vmcnt(0), s_waitcnt
5372                                                             s_waitcnt                       vscnt(0) and s_waitcnt
5373                                                             lgkmcnt(0) to allow             lgkmcnt(0) to allow
5374                                                             them to be                      them to be
5375                                                             independently moved             independently moved
5376                                                             according to the                according to the
5377                                                             following rules.                following rules.
5378                                                           - s_waitcnt vmcnt(0)            - s_waitcnt vmcnt(0)
5379                                                             must happen after               must happen after
5380                                                             any preceding                   any preceding
5381                                                             global/generic                  global/generic
5382                                                             load/store/load                 load/load atomic/
5383                                                             atomic/store                    atomicrmw-with-return-value.
5384                                                             atomic/atomicrmw.
5385                                                                                           - s_waitcnt vscnt(0)
5386                                                                                             must happen after
5387                                                                                             any preceding
5388                                                                                             global/generic
5389                                                                                             store/store atomic/
5390                                                                                             atomicrmw-no-return-value.
5391                                                           - s_waitcnt lgkmcnt(0)          - s_waitcnt lgkmcnt(0)
5392                                                             must happen after               must happen after
5393                                                             any preceding                   any preceding
5394                                                             local/generic                   local/generic
5395                                                             load/store/load                 load/store/load
5396                                                             atomic/store                    atomic/store
5397                                                             atomic/atomicrmw.               atomic/atomicrmw.
5398                                                           - Must happen before            - Must happen before
5399                                                             any following store             any following store
5400                                                             atomic/atomicrmw                atomic/atomicrmw
5401                                                             with an equal or                with an equal or
5402                                                             wider sync scope                wider sync scope
5403                                                             and memory ordering             and memory ordering
5404                                                             stronger than                   stronger than
5405                                                             unordered (this is              unordered (this is
5406                                                             termed the                      termed the
5407                                                             fence-paired-atomic).           fence-paired-atomic).
5408                                                           - Ensures that all              - Ensures that all
5409                                                             memory operations               memory operations
5410                                                             have                            have
5411                                                             completed before                completed before
5412                                                             performing the                  performing the
5413                                                             following                       following
5414                                                             fence-paired-atomic.            fence-paired-atomic.
5415
5416     **Acquire-Release Atomic**
5417     ----------------------------------------------------------------------------------------------------------------------
5418     atomicrmw    acq_rel      - singlethread - global   1. buffer/global/ds/flat_atomic 1. buffer/global/ds/flat_atomic
5419                               - wavefront    - local
5420                                              - generic
5421     atomicrmw    acq_rel      - workgroup    - global   1. s_waitcnt lgkmcnt(0)         1. s_waitcnt lgkmcnt(0) &
5422                                                                                            vmcnt(0) & vscnt(0)
5423
5424                                                                                           - If CU wavefront execution mode, omit vmcnt and
5425                                                                                             vscnt.
5426                                                           - If OpenCL, omit.              - If OpenCL, omit
5427                                                                                             s_waitcnt lgkmcnt(0).
5428                                                           - Must happen after             - Must happen after
5429                                                             any preceding                   any preceding
5430                                                             local/generic                   local/generic
5431                                                             load/store/load                 load/store/load
5432                                                             atomic/store                    atomic/store
5433                                                             atomic/atomicrmw.               atomic/atomicrmw.
5434                                                                                           - Could be split into
5435                                                                                             separate s_waitcnt
5436                                                                                             vmcnt(0), s_waitcnt
5437                                                                                             vscnt(0) and s_waitcnt
5438                                                                                             lgkmcnt(0) to allow
5439                                                                                             them to be
5440                                                                                             independently moved
5441                                                                                             according to the
5442                                                                                             following rules.
5443                                                                                           - s_waitcnt vmcnt(0)
5444                                                                                             must happen after
5445                                                                                             any preceding
5446                                                                                             global/generic load/load
5447                                                                                             atomic/
5448                                                                                             atomicrmw-with-return-value.
5449                                                                                           - s_waitcnt vscnt(0)
5450                                                                                             must happen after
5451                                                                                             any preceding
5452                                                                                             global/generic
5453                                                                                             store/store
5454                                                                                             atomic/
5455                                                                                             atomicrmw-no-return-value.
5456                                                                                           - s_waitcnt lgkmcnt(0)
5457                                                                                             must happen after
5458                                                                                             any preceding
5459                                                                                             local/generic load/store/load
5460                                                                                             atomic/store atomic/atomicrmw.
5461                                                           - Must happen before            - Must happen before
5462                                                             the following                   the following
5463                                                             atomicrmw.                      atomicrmw.
5464                                                           - Ensures that all              - Ensures that all
5465                                                             memory operations               memory operations
5466                                                             to local have                   have
5467                                                             completed before                completed before
5468                                                             performing the                  performing the
5469                                                             atomicrmw that is               atomicrmw that is
5470                                                             being released.                 being released.
5471
5472                                                         2. buffer/global/flat_atomic    2. buffer/global_atomic
5473                                                                                         3. s_waitcnt vm/vscnt(0)
5474
5475                                                                                           - If CU wavefront execution mode, omit vm/vscnt.
5476                                                                                           - Use vmcnt if atomic with
5477                                                                                             return and vscnt if atomic
5478                                                                                             with no-return.
5479                                                                                             waitcnt lgkmcnt(0).
5480                                                                                           - Must happen before
5481                                                                                             the following
5482                                                                                             buffer_gl0_inv.
5483                                                                                           - Ensures any
5484                                                                                             following global
5485                                                                                             data read is no
5486                                                                                             older than the
5487                                                                                             atomicrmw value
5488                                                                                             being acquired.
5489
5490                                                                                         4. buffer_gl0_inv
5491
5492                                                                                           - If CU wavefront execution mode, omit.
5493                                                                                           - Ensures that
5494                                                                                             following
5495                                                                                             loads will not see
5496                                                                                             stale data.
5497
5498     atomicrmw    acq_rel      - workgroup    - local                                    1. waitcnt vmcnt(0) & vscnt(0)
5499
5500                                                                                           - If CU wavefront execution mode, omit.
5501                                                                                           - If OpenCL, omit.
5502                                                                                           - Could be split into
5503                                                                                             separate s_waitcnt
5504                                                                                             vmcnt(0) and s_waitcnt
5505                                                                                             vscnt(0) to allow
5506                                                                                             them to be
5507                                                                                             independently moved
5508                                                                                             according to the
5509                                                                                             following rules.
5510                                                                                           - s_waitcnt vmcnt(0)
5511                                                                                             must happen after
5512                                                                                             any preceding
5513                                                                                             global/generic load/load
5514                                                                                             atomic/
5515                                                                                             atomicrmw-with-return-value.
5516                                                                                           - s_waitcnt vscnt(0)
5517                                                                                             must happen after
5518                                                                                             any preceding
5519                                                                                             global/generic
5520                                                                                             store/store atomic/
5521                                                                                             atomicrmw-no-return-value.
5522                                                                                           - Must happen before
5523                                                                                             the following
5524                                                                                             store.
5525                                                                                           - Ensures that all
5526                                                                                             global memory
5527                                                                                             operations have
5528                                                                                             completed before
5529                                                                                             performing the
5530                                                                                             store that is being
5531                                                                                             released.
5532
5533                                                         1. ds_atomic                    2. ds_atomic
5534                                                         2. s_waitcnt lgkmcnt(0)         3. s_waitcnt lgkmcnt(0)
5535
5536                                                           - If OpenCL, omit.              - If OpenCL, omit.
5537                                                           - Must happen before            - Must happen before
5538                                                             any following                   the following
5539                                                             global/generic                  buffer_gl0_inv.
5540                                                             load/load
5541                                                             atomic/store/store
5542                                                             atomic/atomicrmw.
5543                                                           - Ensures any                   - Ensures any
5544                                                             following global                following global
5545                                                             data read is no                 data read is no
5546                                                             older than the load             older than the load
5547                                                             atomic value being              atomic value being
5548                                                             acquired.                       acquired.
5549
5550                                                                                         4. buffer_gl0_inv
5551
5552                                                                                           - If CU wavefront execution mode, omit.
5553                                                                                           - If OpenCL omit.
5554                                                                                           - Ensures that
5555                                                                                             following
5556                                                                                             loads will not see
5557                                                                                             stale data.
5558
5559     atomicrmw    acq_rel      - workgroup    - generic  1. s_waitcnt lgkmcnt(0)         1. s_waitcnt lgkmcnt(0) &
5560                                                                                            vmcnt(0) & vscnt(0)
5561
5562                                                                                           - If CU wavefront execution mode, omit vmcnt and
5563                                                                                             vscnt.
5564                                                           - If OpenCL, omit.              - If OpenCL, omit
5565                                                                                             waitcnt lgkmcnt(0).
5566                                                           - Must happen after
5567                                                             any preceding
5568                                                             local/generic
5569                                                             load/store/load
5570                                                             atomic/store
5571                                                             atomic/atomicrmw.
5572                                                                                           - Could be split into
5573                                                                                             separate s_waitcnt
5574                                                                                             vmcnt(0), s_waitcnt
5575                                                                                             vscnt(0) and s_waitcnt
5576                                                                                             lgkmcnt(0) to allow
5577                                                                                             them to be
5578                                                                                             independently moved
5579                                                                                             according to the
5580                                                                                             following rules.
5581                                                                                           - s_waitcnt vmcnt(0)
5582                                                                                             must happen after
5583                                                                                             any preceding
5584                                                                                             global/generic load/load
5585                                                                                             atomic/
5586                                                                                             atomicrmw-with-return-value.
5587                                                                                           - s_waitcnt vscnt(0)
5588                                                                                             must happen after
5589                                                                                             any preceding
5590                                                                                             global/generic
5591                                                                                             store/store
5592                                                                                             atomic/
5593                                                                                             atomicrmw-no-return-value.
5594                                                                                           - s_waitcnt lgkmcnt(0)
5595                                                                                             must happen after
5596                                                                                             any preceding
5597                                                                                             local/generic load/store/load
5598                                                                                             atomic/store atomic/atomicrmw.
5599                                                           - Must happen before            - Must happen before
5600                                                             the following                   the following
5601                                                             atomicrmw.                      atomicrmw.
5602                                                           - Ensures that all              - Ensures that all
5603                                                             memory operations               memory operations
5604                                                             to local have                   have
5605                                                             completed before                completed before
5606                                                             performing the                  performing the
5607                                                             atomicrmw that is               atomicrmw that is
5608                                                             being released.                 being released.
5609
5610                                                         2. flat_atomic                  2. flat_atomic
5611                                                         3. s_waitcnt lgkmcnt(0)         3. s_waitcnt lgkmcnt(0) &
5612                                                                                            vm/vscnt(0)
5613
5614                                                                                           - If CU wavefront execution mode, omit vm/vscnt.
5615                                                           - If OpenCL, omit.              - If OpenCL, omit
5616                                                                                             waitcnt lgkmcnt(0).
5617                                                           - Must happen before            - Must happen before
5618                                                             any following                   the following
5619                                                             global/generic                  buffer_gl0_inv.
5620                                                             load/load
5621                                                             atomic/store/store
5622                                                             atomic/atomicrmw.
5623                                                           - Ensures any                   - Ensures any
5624                                                             following global                following global
5625                                                             data read is no                 data read is no
5626                                                             older than the load             older than the load
5627                                                             atomic value being              atomic value being
5628                                                             acquired.                       acquired.
5629
5630                                                                                         3. buffer_gl0_inv
5631
5632                                                                                           - If CU wavefront execution mode, omit.
5633                                                                                           - Ensures that
5634                                                                                             following
5635                                                                                             loads will not see
5636                                                                                             stale data.
5637
5638     atomicrmw    acq_rel      - agent        - global   1. s_waitcnt lgkmcnt(0) &       1. s_waitcnt lgkmcnt(0) &
5639                               - system                     vmcnt(0)                        vmcnt(0) & vscnt(0)
5640
5641                                                           - If OpenCL, omit               - If OpenCL, omit
5642                                                             lgkmcnt(0).                     lgkmcnt(0).
5643                                                           - Could be split into           - Could be split into
5644                                                             separate s_waitcnt              separate s_waitcnt
5645                                                             vmcnt(0) and                    vmcnt(0), s_waitcnt
5646                                                             s_waitcnt                       vscnt(0) and s_waitcnt
5647                                                             lgkmcnt(0) to allow             lgkmcnt(0) to allow
5648                                                             them to be                      them to be
5649                                                             independently moved             independently moved
5650                                                             according to the                according to the
5651                                                             following rules.                following rules.
5652                                                           - s_waitcnt vmcnt(0)            - s_waitcnt vmcnt(0)
5653                                                             must happen after               must happen after
5654                                                             any preceding                   any preceding
5655                                                             global/generic                  global/generic
5656                                                             load/store/load                 load/load atomic/
5657                                                             atomic/store                    atomicrmw-with-return-value.
5658                                                             atomic/atomicrmw.
5659                                                                                           - s_waitcnt vscnt(0)
5660                                                                                             must happen after
5661                                                                                             any preceding
5662                                                                                             global/generic
5663                                                                                             store/store atomic/
5664                                                                                             atomicrmw-no-return-value.
5665                                                           - s_waitcnt lgkmcnt(0)          - s_waitcnt lgkmcnt(0)
5666                                                             must happen after               must happen after
5667                                                             any preceding                   any preceding
5668                                                             local/generic                   local/generic
5669                                                             load/store/load                 load/store/load
5670                                                             atomic/store                    atomic/store
5671                                                             atomic/atomicrmw.               atomic/atomicrmw.
5672                                                           - Must happen before            - Must happen before
5673                                                             the following                   the following
5674                                                             atomicrmw.                      atomicrmw.
5675                                                           - Ensures that all              - Ensures that all
5676                                                             memory operations               memory operations
5677                                                             to global have                  to global have
5678                                                             completed before                completed before
5679                                                             performing the                  performing the
5680                                                             atomicrmw that is               atomicrmw that is
5681                                                             being released.                 being released.
5682
5683                                                         2. buffer/global/flat_atomic    2. buffer/global_atomic
5684                                                         3. s_waitcnt vmcnt(0)           3. s_waitcnt vm/vscnt(0)
5685
5686                                                                                           - Use vmcnt if atomic with
5687                                                                                             return and vscnt if atomic
5688                                                                                             with no-return.
5689                                                                                             waitcnt lgkmcnt(0).
5690                                                           - Must happen before            - Must happen before
5691                                                             following                       following
5692                                                             buffer_wbinvl1_vol.             buffer_gl*_inv.
5693                                                           - Ensures the                   - Ensures the
5694                                                             atomicrmw has                   atomicrmw has
5695                                                             completed before                completed before
5696                                                             invalidating the                invalidating the
5697                                                             cache.                          caches.
5698
5699                                                         4. buffer_wbinvl1_vol           4. buffer_gl0_inv;
5700                                                                                            buffer_gl1_inv
5701
5702                                                           - Must happen before            - Must happen before
5703                                                             any following                   any following
5704                                                             global/generic                  global/generic
5705                                                             load/load                       load/load
5706                                                             atomic/atomicrmw.               atomic/atomicrmw.
5707                                                           - Ensures that                  - Ensures that
5708                                                             following loads                 following loads
5709                                                             will not see stale              will not see stale
5710                                                             global data.                    global data.
5711
5712     atomicrmw    acq_rel      - agent        - generic  1. s_waitcnt lgkmcnt(0) &       1. s_waitcnt lgkmcnt(0) &
5713                               - system                     vmcnt(0)                        vmcnt(0) & vscnt(0)
5714
5715                                                           - If OpenCL, omit               - If OpenCL, omit
5716                                                             lgkmcnt(0).                     lgkmcnt(0).
5717                                                           - Could be split into           - Could be split into
5718                                                             separate s_waitcnt              separate s_waitcnt
5719                                                             vmcnt(0) and                    vmcnt(0), s_waitcnt
5720                                                             s_waitcnt                       vscnt(0) and s_waitcnt
5721                                                             lgkmcnt(0) to allow             lgkmcnt(0) to allow
5722                                                             them to be                      them to be
5723                                                             independently moved             independently moved
5724                                                             according to the                according to the
5725                                                             following rules.                following rules.
5726                                                           - s_waitcnt vmcnt(0)            - s_waitcnt vmcnt(0)
5727                                                             must happen after               must happen after
5728                                                             any preceding                   any preceding
5729                                                             global/generic                  global/generic
5730                                                             load/store/load                 load/load atomic
5731                                                             atomic/store                    atomicrmw-with-return-value.
5732                                                             atomic/atomicrmw.
5733                                                                                           - s_waitcnt vscnt(0)
5734                                                                                             must happen after
5735                                                                                             any preceding
5736                                                                                             global/generic
5737                                                                                             store/store atomic/
5738                                                                                             atomicrmw-no-return-value.
5739                                                           - s_waitcnt lgkmcnt(0)          - s_waitcnt lgkmcnt(0)
5740                                                             must happen after               must happen after
5741                                                             any preceding                   any preceding
5742                                                             local/generic                   local/generic
5743                                                             load/store/load                 load/store/load
5744                                                             atomic/store                    atomic/store
5745                                                             atomic/atomicrmw.               atomic/atomicrmw.
5746                                                           - Must happen before            - Must happen before
5747                                                             the following                   the following
5748                                                             atomicrmw.                      atomicrmw.
5749                                                           - Ensures that all              - Ensures that all
5750                                                             memory operations               memory operations
5751                                                             to global have                  have
5752                                                             completed before                completed before
5753                                                             performing the                  performing the
5754                                                             atomicrmw that is               atomicrmw that is
5755                                                             being released.                 being released.
5756
5757                                                         2. flat_atomic                  2. flat_atomic
5758                                                         3. s_waitcnt vmcnt(0) &         3. s_waitcnt vm/vscnt(0) &
5759                                                            lgkmcnt(0)                      lgkmcnt(0)
5760
5761                                                           - If OpenCL, omit               - If OpenCL, omit
5762                                                             lgkmcnt(0).                     lgkmcnt(0).
5763                                                                                           - Use vmcnt if atomic with
5764                                                                                             return and vscnt if atomic
5765                                                                                             with no-return.
5766                                                           - Must happen before            - Must happen before
5767                                                             following                       following
5768                                                             buffer_wbinvl1_vol.             buffer_gl*_inv.
5769                                                           - Ensures the                   - Ensures the
5770                                                             atomicrmw has                   atomicrmw has
5771                                                             completed before                completed before
5772                                                             invalidating the                invalidating the
5773                                                             cache.                          caches.
5774
5775                                                         4. buffer_wbinvl1_vol           4. buffer_gl0_inv;
5776                                                                                            buffer_gl1_inv
5777
5778                                                           - Must happen before            - Must happen before
5779                                                             any following                   any following
5780                                                             global/generic                  global/generic
5781                                                             load/load                       load/load
5782                                                             atomic/atomicrmw.               atomic/atomicrmw.
5783                                                           - Ensures that                  - Ensures that
5784                                                             following loads                 following loads
5785                                                             will not see stale              will not see stale
5786                                                             global data.                    global data.
5787
5788     fence        acq_rel      - singlethread *none*     *none*                          *none*
5789                               - wavefront
5790     fence        acq_rel      - workgroup    *none*     1. s_waitcnt lgkmcnt(0)         1. s_waitcnt lgkmcnt(0) &
5791                                                                                            vmcnt(0) & vscnt(0)
5792
5793                                                                                           - If CU wavefront execution mode, omit vmcnt and
5794                                                                                             vscnt.
5795                                                           - If OpenCL and                 - If OpenCL and
5796                                                             address space is                address space is
5797                                                             not generic, omit.              not generic, omit
5798                                                                                             lgkmcnt(0).
5799                                                                                           - If OpenCL and
5800                                                                                             address space is
5801                                                                                             local, omit
5802                                                                                             vmcnt(0) and vscnt(0).
5803                                                           - However,                      - However,
5804                                                             since LLVM                      since LLVM
5805                                                             currently has no                currently has no
5806                                                             address space on                address space on
5807                                                             the fence need to               the fence need to
5808                                                             conservatively                  conservatively
5809                                                             always generate                 always generate
5810                                                             (see comment for                (see comment for
5811                                                             previous fence).                previous fence).
5812                                                           - Must happen after
5813                                                             any preceding
5814                                                             local/generic
5815                                                             load/load
5816                                                             atomic/store/store
5817                                                             atomic/atomicrmw.
5818                                                                                           - Could be split into
5819                                                                                             separate s_waitcnt
5820                                                                                             vmcnt(0), s_waitcnt
5821                                                                                             vscnt(0) and s_waitcnt
5822                                                                                             lgkmcnt(0) to allow
5823                                                                                             them to be
5824                                                                                             independently moved
5825                                                                                             according to the
5826                                                                                             following rules.
5827                                                                                           - s_waitcnt vmcnt(0)
5828                                                                                             must happen after
5829                                                                                             any preceding
5830                                                                                             global/generic
5831                                                                                             load/load
5832                                                                                             atomic/
5833                                                                                             atomicrmw-with-return-value.
5834                                                                                           - s_waitcnt vscnt(0)
5835                                                                                             must happen after
5836                                                                                             any preceding
5837                                                                                             global/generic
5838                                                                                             store/store atomic/
5839                                                                                             atomicrmw-no-return-value.
5840                                                                                           - s_waitcnt lgkmcnt(0)
5841                                                                                             must happen after
5842                                                                                             any preceding
5843                                                                                             local/generic
5844                                                                                             load/store/load
5845                                                                                             atomic/store atomic/
5846                                                                                             atomicrmw.
5847                                                           - Must happen before            - Must happen before
5848                                                             any following                   any following
5849                                                             global/generic                  global/generic
5850                                                             load/load                       load/load
5851                                                             atomic/store/store              atomic/store/store
5852                                                             atomic/atomicrmw.               atomic/atomicrmw.
5853                                                           - Ensures that all              - Ensures that all
5854                                                             memory operations               memory operations
5855                                                             to local have                   have
5856                                                             completed before                completed before
5857                                                             performing any                  performing any
5858                                                             following global                following global
5859                                                             memory operations.              memory operations.
5860                                                           - Ensures that the              - Ensures that the
5861                                                             preceding                       preceding
5862                                                             local/generic load              local/generic load
5863                                                             atomic/atomicrmw                atomic/atomicrmw
5864                                                             with an equal or                with an equal or
5865                                                             wider sync scope                wider sync scope
5866                                                             and memory ordering             and memory ordering
5867                                                             stronger than                   stronger than
5868                                                             unordered (this is              unordered (this is
5869                                                             termed the                      termed the
5870                                                             acquire-fence-paired-atomic     acquire-fence-paired-atomic
5871                                                             ) has completed                 ) has completed
5872                                                             before following                before following
5873                                                             global memory                   global memory
5874                                                             operations. This                operations. This
5875                                                             satisfies the                   satisfies the
5876                                                             requirements of                 requirements of
5877                                                             acquire.                        acquire.
5878                                                           - Ensures that all              - Ensures that all
5879                                                             previous memory                 previous memory
5880                                                             operations have                 operations have
5881                                                             completed before a              completed before a
5882                                                             following                       following
5883                                                             local/generic store             local/generic store
5884                                                             atomic/atomicrmw                atomic/atomicrmw
5885                                                             with an equal or                with an equal or
5886                                                             wider sync scope                wider sync scope
5887                                                             and memory ordering             and memory ordering
5888                                                             stronger than                   stronger than
5889                                                             unordered (this is              unordered (this is
5890                                                             termed the                      termed the
5891                                                             release-fence-paired-atomic     release-fence-paired-atomic
5892                                                             ). This satisfies the           ). This satisfies the
5893                                                             requirements of                 requirements of
5894                                                             release.                        release.
5895                                                                                           - Must happen before
5896                                                                                             the following
5897                                                                                             buffer_gl0_inv.
5898                                                                                           - Ensures that the
5899                                                                                             acquire-fence-paired
5900                                                                                             atomic has completed
5901                                                                                             before invalidating
5902                                                                                             the
5903                                                                                             cache. Therefore
5904                                                                                             any following
5905                                                                                             locations read must
5906                                                                                             be no older than
5907                                                                                             the value read by
5908                                                                                             the
5909                                                                                             acquire-fence-paired-atomic.
5910
5911                                                                                         3. buffer_gl0_inv
5912
5913                                                                                           - If CU wavefront execution mode, omit.
5914                                                                                           - Ensures that
5915                                                                                             following
5916                                                                                             loads will not see
5917                                                                                             stale data.
5918
5919     fence        acq_rel      - agent        *none*     1. s_waitcnt lgkmcnt(0) &       1. s_waitcnt lgkmcnt(0) &
5920                               - system                     vmcnt(0)                        vmcnt(0) & vscnt(0)
5921
5922                                                           - If OpenCL and                 - If OpenCL and
5923                                                             address space is                address space is
5924                                                             not generic, omit               not generic, omit
5925                                                             lgkmcnt(0).                     lgkmcnt(0).
5926                                                                                           - If OpenCL and
5927                                                                                             address space is
5928                                                                                             local, omit
5929                                                                                             vmcnt(0) and vscnt(0).
5930                                                           - However, since LLVM           - However, since LLVM
5931                                                             currently has no                currently has no
5932                                                             address space on                address space on
5933                                                             the fence need to               the fence need to
5934                                                             conservatively                  conservatively
5935                                                             always generate                 always generate
5936                                                             (see comment for                (see comment for
5937                                                             previous fence).                previous fence).
5938                                                           - Could be split into           - Could be split into
5939                                                             separate s_waitcnt              separate s_waitcnt
5940                                                             vmcnt(0) and                    vmcnt(0), s_waitcnt
5941                                                             s_waitcnt                       vscnt(0) and s_waitcnt
5942                                                             lgkmcnt(0) to allow             lgkmcnt(0) to allow
5943                                                             them to be                      them to be
5944                                                             independently moved             independently moved
5945                                                             according to the                according to the
5946                                                             following rules.                following rules.
5947                                                           - s_waitcnt vmcnt(0)            - s_waitcnt vmcnt(0)
5948                                                             must happen after               must happen after
5949                                                             any preceding                   any preceding
5950                                                             global/generic                  global/generic
5951                                                             load/store/load                 load/load
5952                                                             atomic/store                    atomic/
5953                                                             atomic/atomicrmw.               atomicrmw-with-return-value.
5954                                                                                           - s_waitcnt vscnt(0)
5955                                                                                             must happen after
5956                                                                                             any preceding
5957                                                                                             global/generic
5958                                                                                             store/store atomic/
5959                                                                                             atomicrmw-no-return-value.
5960                                                           - s_waitcnt lgkmcnt(0)          - s_waitcnt lgkmcnt(0)
5961                                                             must happen after               must happen after
5962                                                             any preceding                   any preceding
5963                                                             local/generic                   local/generic
5964                                                             load/store/load                 load/store/load
5965                                                             atomic/store                    atomic/store
5966                                                             atomic/atomicrmw.               atomic/atomicrmw.
5967                                                           - Must happen before            - Must happen before
5968                                                             the following                   the following
5969                                                             buffer_wbinvl1_vol.             buffer_gl*_inv.
5970                                                           - Ensures that the              - Ensures that the
5971                                                             preceding                       preceding
5972                                                             global/local/generic            global/local/generic
5973                                                             load                            load
5974                                                             atomic/atomicrmw                atomic/atomicrmw
5975                                                             with an equal or                with an equal or
5976                                                             wider sync scope                wider sync scope
5977                                                             and memory ordering             and memory ordering
5978                                                             stronger than                   stronger than
5979                                                             unordered (this is              unordered (this is
5980                                                             termed the                      termed the
5981                                                             acquire-fence-paired-atomic     acquire-fence-paired-atomic
5982                                                             ) has completed                 ) has completed
5983                                                             before invalidating             before invalidating
5984                                                             the cache. This                 the caches. This
5985                                                             satisfies the                   satisfies the
5986                                                             requirements of                 requirements of
5987                                                             acquire.                        acquire.
5988                                                           - Ensures that all              - Ensures that all
5989                                                             previous memory                 previous memory
5990                                                             operations have                 operations have
5991                                                             completed before a              completed before a
5992                                                             following                       following
5993                                                             global/local/generic            global/local/generic
5994                                                             store                           store
5995                                                             atomic/atomicrmw                atomic/atomicrmw
5996                                                             with an equal or                with an equal or
5997                                                             wider sync scope                wider sync scope
5998                                                             and memory ordering             and memory ordering
5999                                                             stronger than                   stronger than
6000                                                             unordered (this is              unordered (this is
6001                                                             termed the                      termed the
6002                                                             release-fence-paired-atomic     release-fence-paired-atomic
6003                                                             ). This satisfies the           ). This satisfies the
6004                                                             requirements of                 requirements of
6005                                                             release.                        release.
6006
6007                                                         2. buffer_wbinvl1_vol           2. buffer_gl0_inv;
6008                                                                                            buffer_gl1_inv
6009
6010                                                           - Must happen before            - Must happen before
6011                                                             any following                   any following
6012                                                             global/generic                  global/generic
6013                                                             load/load                       load/load
6014                                                             atomic/store/store              atomic/store/store
6015                                                             atomic/atomicrmw.               atomic/atomicrmw.
6016                                                           - Ensures that                  - Ensures that
6017                                                             following loads                 following loads
6018                                                             will not see stale              will not see stale
6019                                                             global data. This               global data. This
6020                                                             satisfies the                   satisfies the
6021                                                             requirements of                 requirements of
6022                                                             acquire.                        acquire.
6023
6024     **Sequential Consistent Atomic**
6025     ----------------------------------------------------------------------------------------------------------------------
6026     load atomic  seq_cst      - singlethread - global   *Same as corresponding          *Same as corresponding
6027                               - wavefront    - local    load atomic acquire,            load atomic acquire,
6028                                              - generic  except must generated           except must generated
6029                                                         all instructions even           all instructions even
6030                                                         for OpenCL.*                    for OpenCL.*
6031     load atomic  seq_cst      - workgroup    - global   1. s_waitcnt lgkmcnt(0)         1. s_waitcnt lgkmcnt(0) &
6032                                              - generic                                     vmcnt(0) & vscnt(0)
6033
6034                                                                                           - If CU wavefront execution mode, omit vmcnt and
6035                                                                                             vscnt.
6036                                                                                           - Could be split into
6037                                                                                             separate s_waitcnt
6038                                                                                             vmcnt(0), s_waitcnt
6039                                                                                             vscnt(0) and s_waitcnt
6040                                                                                             lgkmcnt(0) to allow
6041                                                                                             them to be
6042                                                                                             independently moved
6043                                                                                             according to the
6044                                                                                             following rules.
6045                                                           - Must                          - waitcnt lgkmcnt(0) must
6046                                                             happen after                    happen after
6047                                                             preceding                       preceding
6048                                                             global/generic load             local load
6049                                                             atomic/store                    atomic/store
6050                                                             atomic/atomicrmw                atomic/atomicrmw
6051                                                             with memory                     with memory
6052                                                             ordering of seq_cst             ordering of seq_cst
6053                                                             and with equal or               and with equal or
6054                                                             wider sync scope.               wider sync scope.
6055                                                             (Note that seq_cst              (Note that seq_cst
6056                                                             fences have their               fences have their
6057                                                             own s_waitcnt                   own s_waitcnt
6058                                                             lgkmcnt(0) and so do            lgkmcnt(0) and so do
6059                                                             not need to be                  not need to be
6060                                                             considered.)                    considered.)
6061                                                                                           - waitcnt vmcnt(0)
6062                                                                                             Must happen after
6063                                                                                             preceding
6064                                                                                             global/generic load
6065                                                                                             atomic/
6066                                                                                             atomicrmw-with-return-value
6067                                                                                             with memory
6068                                                                                             ordering of seq_cst
6069                                                                                             and with equal or
6070                                                                                             wider sync scope.
6071                                                                                             (Note that seq_cst
6072                                                                                             fences have their
6073                                                                                             own s_waitcnt
6074                                                                                             vmcnt(0) and so do
6075                                                                                             not need to be
6076                                                                                             considered.)
6077                                                                                           - waitcnt vscnt(0)
6078                                                                                             Must happen after
6079                                                                                             preceding
6080                                                                                             global/generic store
6081                                                                                             atomic/
6082                                                                                             atomicrmw-no-return-value
6083                                                                                             with memory
6084                                                                                             ordering of seq_cst
6085                                                                                             and with equal or
6086                                                                                             wider sync scope.
6087                                                                                             (Note that seq_cst
6088                                                                                             fences have their
6089                                                                                             own s_waitcnt
6090                                                                                             vscnt(0) and so do
6091                                                                                             not need to be
6092                                                                                             considered.)
6093                                                           - Ensures any                   - Ensures any
6094                                                             preceding                       preceding
6095                                                             sequential                      sequential
6096                                                             consistent local                consistent global/local
6097                                                             memory instructions             memory instructions
6098                                                             have completed                  have completed
6099                                                             before executing                before executing
6100                                                             this sequentially               this sequentially
6101                                                             consistent                      consistent
6102                                                             instruction. This               instruction. This
6103                                                             prevents reordering             prevents reordering
6104                                                             a seq_cst store                 a seq_cst store
6105                                                             followed by a                   followed by a
6106                                                             seq_cst load. (Note             seq_cst load. (Note
6107                                                             that seq_cst is                 that seq_cst is
6108                                                             stronger than                   stronger than
6109                                                             acquire/release as              acquire/release as
6110                                                             the reordering of               the reordering of
6111                                                             load acquire                    load acquire
6112                                                             followed by a store             followed by a store
6113                                                             release is                      release is
6114                                                             prevented by the                prevented by the
6115                                                             waitcnt of                      waitcnt of
6116                                                             the release, but                the release, but
6117                                                             there is nothing                there is nothing
6118                                                             preventing a store              preventing a store
6119                                                             release followed by             release followed by
6120                                                             load acquire from               load acquire from
6121                                                             competing out of                competing out of
6122                                                             order.)                         order.)
6123
6124                                                         2. *Following                   2. *Following
6125                                                            instructions same as            instructions same as
6126                                                            corresponding load              corresponding load
6127                                                            atomic acquire,                 atomic acquire,
6128                                                            except must generated           except must generated
6129                                                            all instructions even           all instructions even
6130                                                            for OpenCL.*                    for OpenCL.*
6131     load atomic  seq_cst      - workgroup    - local    *Same as corresponding
6132                                                         load atomic acquire,
6133                                                         except must generated
6134                                                         all instructions even
6135                                                         for OpenCL.*
6136
6137                                                                                         1. s_waitcnt vmcnt(0) & vscnt(0)
6138
6139                                                                                           - If CU wavefront execution mode, omit.
6140                                                                                           - Could be split into
6141                                                                                             separate s_waitcnt
6142                                                                                             vmcnt(0) and s_waitcnt
6143                                                                                             vscnt(0) to allow
6144                                                                                             them to be
6145                                                                                             independently moved
6146                                                                                             according to the
6147                                                                                             following rules.
6148                                                                                           - waitcnt vmcnt(0)
6149                                                                                             Must happen after
6150                                                                                             preceding
6151                                                                                             global/generic load
6152                                                                                             atomic/
6153                                                                                             atomicrmw-with-return-value
6154                                                                                             with memory
6155                                                                                             ordering of seq_cst
6156                                                                                             and with equal or
6157                                                                                             wider sync scope.
6158                                                                                             (Note that seq_cst
6159                                                                                             fences have their
6160                                                                                             own s_waitcnt
6161                                                                                             vmcnt(0) and so do
6162                                                                                             not need to be
6163                                                                                             considered.)
6164                                                                                           - waitcnt vscnt(0)
6165                                                                                             Must happen after
6166                                                                                             preceding
6167                                                                                             global/generic store
6168                                                                                             atomic/
6169                                                                                             atomicrmw-no-return-value
6170                                                                                             with memory
6171                                                                                             ordering of seq_cst
6172                                                                                             and with equal or
6173                                                                                             wider sync scope.
6174                                                                                             (Note that seq_cst
6175                                                                                             fences have their
6176                                                                                             own s_waitcnt
6177                                                                                             vscnt(0) and so do
6178                                                                                             not need to be
6179                                                                                             considered.)
6180                                                                                           - Ensures any
6181                                                                                             preceding
6182                                                                                             sequential
6183                                                                                             consistent global
6184                                                                                             memory instructions
6185                                                                                             have completed
6186                                                                                             before executing
6187                                                                                             this sequentially
6188                                                                                             consistent
6189                                                                                             instruction. This
6190                                                                                             prevents reordering
6191                                                                                             a seq_cst store
6192                                                                                             followed by a
6193                                                                                             seq_cst load. (Note
6194                                                                                             that seq_cst is
6195                                                                                             stronger than
6196                                                                                             acquire/release as
6197                                                                                             the reordering of
6198                                                                                             load acquire
6199                                                                                             followed by a store
6200                                                                                             release is
6201                                                                                             prevented by the
6202                                                                                             waitcnt of
6203                                                                                             the release, but
6204                                                                                             there is nothing
6205                                                                                             preventing a store
6206                                                                                             release followed by
6207                                                                                             load acquire from
6208                                                                                             competing out of
6209                                                                                             order.)
6210
6211                                                                                         2. *Following
6212                                                                                            instructions same as
6213                                                                                            corresponding load
6214                                                                                            atomic acquire,
6215                                                                                            except must generated
6216                                                                                            all instructions even
6217                                                                                            for OpenCL.*
6218
6219     load atomic  seq_cst      - agent        - global   1. s_waitcnt lgkmcnt(0) &       1. s_waitcnt lgkmcnt(0) &
6220                               - system       - generic     vmcnt(0)                        vmcnt(0) & vscnt(0)
6221
6222                                                           - Could be split into           - Could be split into
6223                                                             separate s_waitcnt              separate s_waitcnt
6224                                                             vmcnt(0)                        vmcnt(0), s_waitcnt
6225                                                             and s_waitcnt                   vscnt(0) and s_waitcnt
6226                                                             lgkmcnt(0) to allow             lgkmcnt(0) to allow
6227                                                             them to be                      them to be
6228                                                             independently moved             independently moved
6229                                                             according to the                according to the
6230                                                             following rules.                following rules.
6231                                                           - waitcnt lgkmcnt(0)            - waitcnt lgkmcnt(0)
6232                                                             must happen after               must happen after
6233                                                             preceding                       preceding
6234                                                             global/generic load             local load
6235                                                             atomic/store                    atomic/store
6236                                                             atomic/atomicrmw                atomic/atomicrmw
6237                                                             with memory                     with memory
6238                                                             ordering of seq_cst             ordering of seq_cst
6239                                                             and with equal or               and with equal or
6240                                                             wider sync scope.               wider sync scope.
6241                                                             (Note that seq_cst              (Note that seq_cst
6242                                                             fences have their               fences have their
6243                                                             own s_waitcnt                   own s_waitcnt
6244                                                             lgkmcnt(0) and so do            lgkmcnt(0) and so do
6245                                                             not need to be                  not need to be
6246                                                             considered.)                    considered.)
6247                                                           - waitcnt vmcnt(0)              - waitcnt vmcnt(0)
6248                                                             must happen after               must happen after
6249                                                             preceding                       preceding
6250                                                             global/generic load             global/generic load
6251                                                             atomic/store                    atomic/
6252                                                             atomic/atomicrmw                atomicrmw-with-return-value
6253                                                             with memory                     with memory
6254                                                             ordering of seq_cst             ordering of seq_cst
6255                                                             and with equal or               and with equal or
6256                                                             wider sync scope.               wider sync scope.
6257                                                             (Note that seq_cst              (Note that seq_cst
6258                                                             fences have their               fences have their
6259                                                             own s_waitcnt                   own s_waitcnt
6260                                                             vmcnt(0) and so do              vmcnt(0) and so do
6261                                                             not need to be                  not need to be
6262                                                             considered.)                    considered.)
6263                                                                                           - waitcnt vscnt(0)
6264                                                                                             Must happen after
6265                                                                                             preceding
6266                                                                                             global/generic store
6267                                                                                             atomic/
6268                                                                                             atomicrmw-no-return-value
6269                                                                                             with memory
6270                                                                                             ordering of seq_cst
6271                                                                                             and with equal or
6272                                                                                             wider sync scope.
6273                                                                                             (Note that seq_cst
6274                                                                                             fences have their
6275                                                                                             own s_waitcnt
6276                                                                                             vscnt(0) and so do
6277                                                                                             not need to be
6278                                                                                             considered.)
6279                                                           - Ensures any                   - Ensures any
6280                                                             preceding                       preceding
6281                                                             sequential                      sequential
6282                                                             consistent global               consistent global
6283                                                             memory instructions             memory instructions
6284                                                             have completed                  have completed
6285                                                             before executing                before executing
6286                                                             this sequentially               this sequentially
6287                                                             consistent                      consistent
6288                                                             instruction. This               instruction. This
6289                                                             prevents reordering             prevents reordering
6290                                                             a seq_cst store                 a seq_cst store
6291                                                             followed by a                   followed by a
6292                                                             seq_cst load. (Note             seq_cst load. (Note
6293                                                             that seq_cst is                 that seq_cst is
6294                                                             stronger than                   stronger than
6295                                                             acquire/release as              acquire/release as
6296                                                             the reordering of               the reordering of
6297                                                             load acquire                    load acquire
6298                                                             followed by a store             followed by a store
6299                                                             release is                      release is
6300                                                             prevented by the                prevented by the
6301                                                             waitcnt of                      waitcnt of
6302                                                             the release, but                the release, but
6303                                                             there is nothing                there is nothing
6304                                                             preventing a store              preventing a store
6305                                                             release followed by             release followed by
6306                                                             load acquire from               load acquire from
6307                                                             competing out of                competing out of
6308                                                             order.)                         order.)
6309
6310                                                         2. *Following                   2. *Following
6311                                                            instructions same as            instructions same as
6312                                                            corresponding load              corresponding load
6313                                                            atomic acquire,                 atomic acquire,
6314                                                            except must generated           except must generated
6315                                                            all instructions even           all instructions even
6316                                                            for OpenCL.*                    for OpenCL.*
6317     store atomic seq_cst      - singlethread - global   *Same as corresponding          *Same as corresponding
6318                               - wavefront    - local    store atomic release,           store atomic release,
6319                               - workgroup    - generic  except must generated           except must generated
6320                                                         all instructions even           all instructions even
6321                                                         for OpenCL.*                    for OpenCL.*
6322     store atomic seq_cst      - agent        - global   *Same as corresponding          *Same as corresponding
6323                               - system       - generic  store atomic release,           store atomic release,
6324                                                         except must generated           except must generated
6325                                                         all instructions even           all instructions even
6326                                                         for OpenCL.*                    for OpenCL.*
6327     atomicrmw    seq_cst      - singlethread - global   *Same as corresponding          *Same as corresponding
6328                               - wavefront    - local    atomicrmw acq_rel,              atomicrmw acq_rel,
6329                               - workgroup    - generic  except must generated           except must generated
6330                                                         all instructions even           all instructions even
6331                                                         for OpenCL.*                    for OpenCL.*
6332     atomicrmw    seq_cst      - agent        - global   *Same as corresponding          *Same as corresponding
6333                               - system       - generic  atomicrmw acq_rel,              atomicrmw acq_rel,
6334                                                         except must generated           except must generated
6335                                                         all instructions even           all instructions even
6336                                                         for OpenCL.*                    for OpenCL.*
6337     fence        seq_cst      - singlethread *none*     *Same as corresponding          *Same as corresponding
6338                               - wavefront               fence acq_rel,                  fence acq_rel,
6339                               - workgroup               except must generated           except must generated
6340                               - agent                   all instructions even           all instructions even
6341                               - system                  for OpenCL.*                    for OpenCL.*
6342     ============ ============ ============== ========== =============================== ==================================
6343
6344The memory order also adds the single thread optimization constrains defined in
6345table
6346:ref:`amdgpu-amdhsa-memory-model-single-thread-optimization-constraints-gfx6-gfx10-table`.
6347
6348  .. table:: AMDHSA Memory Model Single Thread Optimization Constraints GFX6-GFX10
6349     :name: amdgpu-amdhsa-memory-model-single-thread-optimization-constraints-gfx6-gfx10-table
6350
6351     ============ ==============================================================
6352     LLVM Memory  Optimization Constraints
6353     Ordering
6354     ============ ==============================================================
6355     unordered    *none*
6356     monotonic    *none*
6357     acquire      - If a load atomic/atomicrmw then no following load/load
6358                    atomic/store/ store atomic/atomicrmw/fence instruction can
6359                    be moved before the acquire.
6360                  - If a fence then same as load atomic, plus no preceding
6361                    associated fence-paired-atomic can be moved after the fence.
6362     release      - If a store atomic/atomicrmw then no preceding load/load
6363                    atomic/store/ store atomic/atomicrmw/fence instruction can
6364                    be moved after the release.
6365                  - If a fence then same as store atomic, plus no following
6366                    associated fence-paired-atomic can be moved before the
6367                    fence.
6368     acq_rel      Same constraints as both acquire and release.
6369     seq_cst      - If a load atomic then same constraints as acquire, plus no
6370                    preceding sequentially consistent load atomic/store
6371                    atomic/atomicrmw/fence instruction can be moved after the
6372                    seq_cst.
6373                  - If a store atomic then the same constraints as release, plus
6374                    no following sequentially consistent load atomic/store
6375                    atomic/atomicrmw/fence instruction can be moved before the
6376                    seq_cst.
6377                  - If an atomicrmw/fence then same constraints as acq_rel.
6378     ============ ==============================================================
6379
6380Trap Handler ABI
6381~~~~~~~~~~~~~~~~
6382
6383For code objects generated by AMDGPU backend for HSA [HSA]_ compatible runtimes
6384(such as ROCm [AMD-ROCm]_), the runtime installs a trap handler that supports
6385the ``s_trap`` instruction with the following usage:
6386
6387  .. table:: AMDGPU Trap Handler for AMDHSA OS
6388     :name: amdgpu-trap-handler-for-amdhsa-os-table
6389
6390     =================== =============== =============== =======================
6391     Usage               Code Sequence   Trap Handler    Description
6392                                         Inputs
6393     =================== =============== =============== =======================
6394     reserved            ``s_trap 0x00``                 Reserved by hardware.
6395     ``debugtrap(arg)``  ``s_trap 0x01`` ``SGPR0-1``:    Reserved for HSA
6396                                           ``queue_ptr`` ``debugtrap``
6397                                         ``VGPR0``:      intrinsic (not
6398                                           ``arg``       implemented).
6399     ``llvm.trap``       ``s_trap 0x02`` ``SGPR0-1``:    Causes dispatch to be
6400                                           ``queue_ptr`` terminated and its
6401                                                         associated queue put
6402                                                         into the error state.
6403     ``llvm.debugtrap``  ``s_trap 0x03``                 - If debugger not
6404                                                           installed then
6405                                                           behaves as a
6406                                                           no-operation. The
6407                                                           trap handler is
6408                                                           entered and
6409                                                           immediately returns
6410                                                           to continue
6411                                                           execution of the
6412                                                           wavefront.
6413                                                         - If the debugger is
6414                                                           installed, causes
6415                                                           the debug trap to be
6416                                                           reported by the
6417                                                           debugger and the
6418                                                           wavefront is put in
6419                                                           the halt state until
6420                                                           resumed by the
6421                                                           debugger.
6422     reserved            ``s_trap 0x04``                 Reserved.
6423     reserved            ``s_trap 0x05``                 Reserved.
6424     reserved            ``s_trap 0x06``                 Reserved.
6425     debugger breakpoint ``s_trap 0x07``                 Reserved for debugger
6426                                                         breakpoints.
6427     reserved            ``s_trap 0x08``                 Reserved.
6428     reserved            ``s_trap 0xfe``                 Reserved.
6429     reserved            ``s_trap 0xff``                 Reserved.
6430     =================== =============== =============== =======================
6431
6432.. _amdgpu-amdhsa-function-call-convention:
6433
6434Call Convention
6435~~~~~~~~~~~~~~~
6436
6437.. note::
6438
6439  This section is currently incomplete and has inakkuracies. It is WIP that will
6440  be updated as information is determined.
6441
6442See :ref:`amdgpu-dwarf-address-space-identifier` for information on swizzled
6443addresses. Unswizzled addresses are normal linear addresses.
6444
6445.. _amdgpu-amdhsa-function-call-convention-kernel-functions:
6446
6447Kernel Functions
6448++++++++++++++++
6449
6450This section describes the call convention ABI for the outer kernel function.
6451
6452See :ref:`amdgpu-amdhsa-initial-kernel-execution-state` for the kernel call
6453convention.
6454
6455The following is not part of the AMDGPU kernel calling convention but describes
6456how the AMDGPU implements function calls:
6457
64581.  Clang decides the kernarg layout to match the *HSA Programmer's Language
6459    Reference* [HSA]_.
6460
6461    - All structs are passed directly.
6462    - Lambda values are passed *TBA*.
6463
6464    .. TODO::
6465
6466      - Does this really follow HSA rules? Or are structs >16 bytes passed
6467        by-value struct?
6468      - What is ABI for lambda values?
6469
64704.  The kernel performs certain setup in its prolog, as described in
6471    :ref:`amdgpu-amdhsa-kernel-prolog`.
6472
6473.. _amdgpu-amdhsa-function-call-convention-non-kernel-functions:
6474
6475Non-Kernel Functions
6476++++++++++++++++++++
6477
6478This section describes the call convention ABI for functions other than the
6479outer kernel function.
6480
6481If a kernel has function calls then scratch is always allocated and used for
6482the call stack which grows from low address to high address using the swizzled
6483scratch address space.
6484
6485On entry to a function:
6486
64871.  SGPR0-3 contain a V# with the following properties (see
6488    :ref:`amdgpu-amdhsa-kernel-prolog-private-segment-buffer`):
6489
6490    * Base address pointing to the beginning of the wavefront scratch backing
6491      memory.
6492    * Swizzled with dword element size and stride of wavefront size elements.
6493
64942.  The FLAT_SCRATCH register pair is setup. See
6495    :ref:`amdgpu-amdhsa-kernel-prolog-flat-scratch`.
64963.  GFX6-8: M0 register set to the size of LDS in bytes. See
6497    :ref:`amdgpu-amdhsa-kernel-prolog-m0`.
64984.  The EXEC register is set to the lanes active on entry to the function.
64995.  MODE register: *TBD*
65006.  VGPR0-31 and SGPR4-29 are used to pass function input arguments as described
6501    below.
65027.  SGPR30-31 return address (RA). The code address that the function must
6503    return to when it completes. The value is undefined if the function is *no
6504    return*.
65058.  SGPR32 is used for the stack pointer (SP). It is an unswizzled scratch
6506    offset relative to the beginning of the wavefront scratch backing memory.
6507
6508    The unswizzled SP can be used with buffer instructions as an unswizzled SGPR
6509    offset with the scratch V# in SGPR0-3 to access the stack in a swizzled
6510    manner.
6511
6512    The unswizzled SP value can be converted into the swizzled SP value by:
6513
6514      | swizzled SP = unswizzled SP / wavefront size
6515
6516    This may be used to obtain the private address space address of stack
6517    objects and to convert this address to a flat address by adding the flat
6518    scratch aperture base address.
6519
6520    The swizzled SP value is always 4 bytes aligned for the ``r600``
6521    architecture and 16 byte aligned for the ``amdgcn`` architecture.
6522
6523    .. note::
6524
6525      The ``amdgcn`` value is selected to avoid dynamic stack alignment for the
6526      OpenCL language which has the largest base type defined as 16 bytes.
6527
6528    On entry, the swizzled SP value is the address of the first function
6529    argument passed on the stack. Other stack passed arguments are positive
6530    offsets from the entry swizzled SP value.
6531
6532    The function may use positive offsets beyond the last stack passed argument
6533    for stack allocated local variables and register spill slots. If necessary,
6534    the function may align these to greater alignment than 16 bytes. After these
6535    the function may dynamically allocate space for such things as runtime sized
6536    ``alloca`` local allocations.
6537
6538    If the function calls another function, it will place any stack allocated
6539    arguments after the last local allocation and adjust SGPR32 to the address
6540    after the last local allocation.
6541
65429.  All other registers are unspecified.
654310. Any necessary ``waitcnt`` has been performed to ensure memory is available
6544    to the function.
6545
6546On exit from a function:
6547
65481.  VGPR0-31 and SGPR4-29 are used to pass function result arguments as
6549    described below. Any registers used are considered clobbered registers.
65502.  The following registers are preserved and have the same value as on entry:
6551
6552    * FLAT_SCRATCH
6553    * EXEC
6554    * GFX6-8: M0
6555    * All SGPR registers except the clobbered registers of SGPR4-31.
6556    * VGPR40-47
6557      VGPR56-63
6558      VGPR72-79
6559      VGPR88-95
6560      VGPR104-111
6561      VGPR120-127
6562      VGPR136-143
6563      VGPR152-159
6564      VGPR168-175
6565      VGPR184-191
6566      VGPR200-207
6567      VGPR216-223
6568      VGPR232-239
6569      VGPR248-255
6570
6571        *Except the argument registers, the VGPR cloberred and the preserved
6572        registers are intermixed at regular intervals in order to
6573        get a better occupancy.*
6574
6575      For the AMDGPU backend, an inter-procedural register allocation (IPRA)
6576      optimization may mark some of clobbered SGPR and VGPR registers as
6577      preserved if it can be determined that the called function does not change
6578      their value.
6579
65802.  The PC is set to the RA provided on entry.
65813.  MODE register: *TBD*.
65824.  All other registers are clobbered.
65835.  Any necessary ``waitcnt`` has been performed to ensure memory accessed by
6584    function is available to the caller.
6585
6586.. TODO::
6587
6588  - On gfx908 are all ACC registers clobbered?
6589
6590  - How are function results returned? The address of structured types is passed
6591    by reference, but what about other types?
6592
6593The function input arguments are made up of the formal arguments explicitly
6594declared by the source language function plus the implicit input arguments used
6595by the implementation.
6596
6597The source language input arguments are:
6598
65991. Any source language implicit ``this`` or ``self`` argument comes first as a
6600   pointer type.
66012. Followed by the function formal arguments in left to right source order.
6602
6603The source language result arguments are:
6604
66051. The function result argument.
6606
6607The source language input or result struct type arguments that are less than or
6608equal to 16 bytes, are decomposed recursively into their base type fields, and
6609each field is passed as if a separate argument. For input arguments, if the
6610called function requires the struct to be in memory, for example because its
6611address is taken, then the function body is responsible for allocating a stack
6612location and copying the field arguments into it. Clang terms this *direct
6613struct*.
6614
6615The source language input struct type arguments that are greater than 16 bytes,
6616are passed by reference. The caller is responsible for allocating a stack
6617location to make a copy of the struct value and pass the address as the input
6618argument. The called function is responsible to perform the dereference when
6619accessing the input argument. Clang terms this *by-value struct*.
6620
6621A source language result struct type argument that is greater than 16 bytes, is
6622returned by reference. The caller is responsible for allocating a stack location
6623to hold the result value and passes the address as the last input argument
6624(before the implicit input arguments). In this case there are no result
6625arguments. The called function is responsible to perform the dereference when
6626storing the result value. Clang terms this *structured return (sret)*.
6627
6628*TODO: correct the ``sret`` definition.*
6629
6630.. TODO::
6631
6632  Is this definition correct? Or is ``sret`` only used if passing in registers, and
6633  pass as non-decomposed struct as stack argument? Or something else? Is the
6634  memory location in the caller stack frame, or a stack memory argument and so
6635  no address is passed as the caller can directly write to the argument stack
6636  location? But then the stack location is still live after return. If an
6637  argument stack location is it the first stack argument or the last one?
6638
6639Lambda argument types are treated as struct types with an implementation defined
6640set of fields.
6641
6642.. TODO::
6643
6644  Need to specify the ABI for lambda types for AMDGPU.
6645
6646For AMDGPU backend all source language arguments (including the decomposed
6647struct type arguments) are passed in VGPRs unless marked ``inreg`` in which case
6648they are passed in SGPRs.
6649
6650The AMDGPU backend walks the function call graph from the leaves to determine
6651which implicit input arguments are used, propagating to each caller of the
6652function. The used implicit arguments are appended to the function arguments
6653after the source language arguments in the following order:
6654
6655.. TODO::
6656
6657  Is recursion or external functions supported?
6658
66591.  Work-Item ID (1 VGPR)
6660
6661    The X, Y and Z work-item ID are packed into a single VGRP with the following
6662    layout. Only fields actually used by the function are set. The other bits
6663    are undefined.
6664
6665    The values come from the initial kernel execution state. See
6666    :ref:`amdgpu-amdhsa-vgpr-register-set-up-order-table`.
6667
6668    .. table:: Work-item implicit argument layout
6669      :name: amdgpu-amdhsa-workitem-implicit-argument-layout-table
6670
6671      ======= ======= ==============
6672      Bits    Size    Field Name
6673      ======= ======= ==============
6674      9:0     10 bits X Work-Item ID
6675      19:10   10 bits Y Work-Item ID
6676      29:20   10 bits Z Work-Item ID
6677      31:30   2 bits  Unused
6678      ======= ======= ==============
6679
66802.  Dispatch Ptr (2 SGPRs)
6681
6682    The value comes from the initial kernel execution state. See
6683    :ref:`amdgpu-amdhsa-sgpr-register-set-up-order-table`.
6684
66853.  Queue Ptr (2 SGPRs)
6686
6687    The value comes from the initial kernel execution state. See
6688    :ref:`amdgpu-amdhsa-sgpr-register-set-up-order-table`.
6689
66904.  Kernarg Segment Ptr (2 SGPRs)
6691
6692    The value comes from the initial kernel execution state. See
6693    :ref:`amdgpu-amdhsa-sgpr-register-set-up-order-table`.
6694
66955.  Dispatch id (2 SGPRs)
6696
6697    The value comes from the initial kernel execution state. See
6698    :ref:`amdgpu-amdhsa-sgpr-register-set-up-order-table`.
6699
67006.  Work-Group ID X (1 SGPR)
6701
6702    The value comes from the initial kernel execution state. See
6703    :ref:`amdgpu-amdhsa-sgpr-register-set-up-order-table`.
6704
67057.  Work-Group ID Y (1 SGPR)
6706
6707    The value comes from the initial kernel execution state. See
6708    :ref:`amdgpu-amdhsa-sgpr-register-set-up-order-table`.
6709
67108.  Work-Group ID Z (1 SGPR)
6711
6712    The value comes from the initial kernel execution state. See
6713    :ref:`amdgpu-amdhsa-sgpr-register-set-up-order-table`.
6714
67159.  Implicit Argument Ptr (2 SGPRs)
6716
6717    The value is computed by adding an offset to Kernarg Segment Ptr to get the
6718    global address space pointer to the first kernarg implicit argument.
6719
6720The input and result arguments are assigned in order in the following manner:
6721
6722.. note::
6723
6724  There are likely some errors and omissions in the following description that
6725  need correction.
6726
6727  .. TODO::
6728
6729    Check the clang source code to decipher how function arguments and return
6730    results are handled. Also see the AMDGPU specific values used.
6731
6732* VGPR arguments are assigned to consecutive VGPRs starting at VGPR0 up to
6733  VGPR31.
6734
6735  If there are more arguments than will fit in these registers, the remaining
6736  arguments are allocated on the stack in order on naturally aligned
6737  addresses.
6738
6739  .. TODO::
6740
6741    How are overly aligned structures allocated on the stack?
6742
6743* SGPR arguments are assigned to consecutive SGPRs starting at SGPR0 up to
6744  SGPR29.
6745
6746  If there are more arguments than will fit in these registers, the remaining
6747  arguments are allocated on the stack in order on naturally aligned
6748  addresses.
6749
6750Note that decomposed struct type arguments may have some fields passed in
6751registers and some in memory.
6752
6753.. TODO::
6754
6755  So, a struct which can pass some fields as decomposed register arguments, will
6756  pass the rest as decomposed stack elements? But an argument that will not start
6757  in registers will not be decomposed and will be passed as a non-decomposed
6758  stack value?
6759
6760The following is not part of the AMDGPU function calling convention but
6761describes how the AMDGPU implements function calls:
6762
67631.  SGPR33 is used as a frame pointer (FP) if necessary. Like the SP it is an
6764    unswizzled scratch address. It is only needed if runtime sized ``alloca``
6765    are used, or for the reasons defined in ``SIFrameLowering``.
67662.  Runtime stack alignment is supported. SGPR34 is used as a base pointer (BP)
6767    to access the incoming stack arguments in the function. The BP is needed
6768    only when the function requires the runtime stack alignment.
6769
67703.  Allocating SGPR arguments on the stack are not supported.
6771
67724.  No CFI is currently generated. See
6773    :ref:`amdgpu-dwarf-call-frame-information`.
6774
6775    .. note::
6776
6777      CFI will be generated that defines the CFA as the unswizzled address
6778      relative to the wave scratch base in the unswizzled private address space
6779      of the lowest address stack allocated local variable.
6780
6781      ``DW_AT_frame_base`` will be defined as the swizzled address in the
6782      swizzled private address space by dividing the CFA by the wavefront size
6783      (since CFA is always at least dword aligned which matches the scratch
6784      swizzle element size).
6785
6786      If no dynamic stack alignment was performed, the stack allocated arguments
6787      are accessed as negative offsets relative to ``DW_AT_frame_base``, and the
6788      local variables and register spill slots are accessed as positive offsets
6789      relative to ``DW_AT_frame_base``.
6790
67915.  Function argument passing is implemented by copying the input physical
6792    registers to virtual registers on entry. The register allocator can spill if
6793    necessary. These are copied back to physical registers at call sites. The
6794    net effect is that each function call can have these values in entirely
6795    distinct locations. The IPRA can help avoid shuffling argument registers.
67966.  Call sites are implemented by setting up the arguments at positive offsets
6797    from SP. Then SP is incremented to account for the known frame size before
6798    the call and decremented after the call.
6799
6800    .. note::
6801
6802      The CFI will reflect the changed calculation needed to compute the CFA
6803      from SP.
6804
68057.  4 byte spill slots are used in the stack frame. One slot is allocated for an
6806    emergency spill slot. Buffer instructions are used for stack accesses and
6807    not the ``flat_scratch`` instruction.
6808
6809    .. TODO::
6810
6811      Explain when the emergency spill slot is used.
6812
6813.. TODO::
6814
6815  Possible broken issues:
6816
6817  - Stack arguments must be aligned to required alignment.
6818  - Stack is aligned to max(16, max formal argument alignment)
6819  - Direct argument < 64 bits should check register budget.
6820  - Register budget calculation should respect ``inreg`` for SGPR.
6821  - SGPR overflow is not handled.
6822  - struct with 1 member unpeeling is not checking size of member.
6823  - ``sret`` is after ``this`` pointer.
6824  - Caller is not implementing stack realignment: need an extra pointer.
6825  - Should say AMDGPU passes FP rather than SP.
6826  - Should CFI define CFA as address of locals or arguments. Difference is
6827    apparent when have implemented dynamic alignment.
6828  - If ``SCRATCH`` instruction could allow negative offsets, then can make FP be
6829    highest address of stack frame and use negative offset for locals. Would
6830    allow SP to be the same as FP and could support signal-handler-like as now
6831    have a real SP for the top of the stack.
6832  - How is ``sret`` passed on the stack? In argument stack area? Can it overlay
6833    arguments?
6834
6835AMDPAL
6836------
6837
6838This section provides code conventions used when the target triple OS is
6839``amdpal`` (see :ref:`amdgpu-target-triples`) for passing runtime parameters
6840from the application/runtime to each invocation of a hardware shader. These
6841parameters include both generic, application-controlled parameters called
6842*user data* as well as system-generated parameters that are a product of the
6843draw or dispatch execution.
6844
6845User Data
6846~~~~~~~~~
6847
6848Each hardware stage has a set of 32-bit *user data registers* which can be
6849written from a command buffer and then loaded into SGPRs when waves are launched
6850via a subsequent dispatch or draw operation. This is the way most arguments are
6851passed from the application/runtime to a hardware shader.
6852
6853Compute User Data
6854~~~~~~~~~~~~~~~~~
6855
6856Compute shader user data mappings are simpler than graphics shaders and have a
6857fixed mapping.
6858
6859Note that there are always 10 available *user data entries* in registers -
6860entries beyond that limit must be fetched from memory (via the spill table
6861pointer) by the shader.
6862
6863  .. table:: PAL Compute Shader User Data Registers
6864     :name: pal-compute-user-data-registers
6865
6866     ============= ================================
6867     User Register Description
6868     ============= ================================
6869     0             Global Internal Table (32-bit pointer)
6870     1             Per-Shader Internal Table (32-bit pointer)
6871     2 - 11        Application-Controlled User Data (10 32-bit values)
6872     12            Spill Table (32-bit pointer)
6873     13 - 14       Thread Group Count (64-bit pointer)
6874     15            GDS Range
6875     ============= ================================
6876
6877Graphics User Data
6878~~~~~~~~~~~~~~~~~~
6879
6880Graphics pipelines support a much more flexible user data mapping:
6881
6882  .. table:: PAL Graphics Shader User Data Registers
6883     :name: pal-graphics-user-data-registers
6884
6885     ============= ================================
6886     User Register Description
6887     ============= ================================
6888     0             Global Internal Table (32-bit pointer)
6889     +             Per-Shader Internal Table (32-bit pointer)
6890     + 1-15        Application Controlled User Data
6891                   (1-15 Contiguous 32-bit Values in Registers)
6892     +             Spill Table (32-bit pointer)
6893     +             Draw Index (First Stage Only)
6894     +             Vertex Offset (First Stage Only)
6895     +             Instance Offset (First Stage Only)
6896     ============= ================================
6897
6898  The placement of the global internal table remains fixed in the first *user
6899  data SGPR register*. Otherwise all parameters are optional, and can be mapped
6900  to any desired *user data SGPR register*, with the following restrictions:
6901
6902  * Draw Index, Vertex Offset, and Instance Offset can only be used by the first
6903    active hardware stage in a graphics pipeline (i.e. where the API vertex
6904    shader runs).
6905
6906  * Application-controlled user data must be mapped into a contiguous range of
6907    user data registers.
6908
6909  * The application-controlled user data range supports compaction remapping, so
6910    only *entries* that are actually consumed by the shader must be assigned to
6911    corresponding *registers*. Note that in order to support an efficient runtime
6912    implementation, the remapping must pack *registers* in the same order as
6913    *entries*, with unused *entries* removed.
6914
6915.. _pal_global_internal_table:
6916
6917Global Internal Table
6918~~~~~~~~~~~~~~~~~~~~~
6919
6920The global internal table is a table of *shader resource descriptors* (SRDs)
6921that define how certain engine-wide, runtime-managed resources should be
6922accessed from a shader. The majority of these resources have HW-defined formats,
6923and it is up to the compiler to write/read data as required by the target
6924hardware.
6925
6926The following table illustrates the required format:
6927
6928  .. table:: PAL Global Internal Table
6929     :name: pal-git-table
6930
6931     ============= ================================
6932     Offset        Description
6933     ============= ================================
6934     0-3           Graphics Scratch SRD
6935     4-7           Compute Scratch SRD
6936     8-11          ES/GS Ring Output SRD
6937     12-15         ES/GS Ring Input SRD
6938     16-19         GS/VS Ring Output #0
6939     20-23         GS/VS Ring Output #1
6940     24-27         GS/VS Ring Output #2
6941     28-31         GS/VS Ring Output #3
6942     32-35         GS/VS Ring Input SRD
6943     36-39         Tessellation Factor Buffer SRD
6944     40-43         Off-Chip LDS Buffer SRD
6945     44-47         Off-Chip Param Cache Buffer SRD
6946     48-51         Sample Position Buffer SRD
6947     52            vaRange::ShadowDescriptorTable High Bits
6948     ============= ================================
6949
6950  The pointer to the global internal table passed to the shader as user data
6951  is a 32-bit pointer. The top 32 bits should be assumed to be the same as
6952  the top 32 bits of the pipeline, so the shader may use the program
6953  counter's top 32 bits.
6954
6955Unspecified OS
6956--------------
6957
6958This section provides code conventions used when the target triple OS is
6959empty (see :ref:`amdgpu-target-triples`).
6960
6961Trap Handler ABI
6962~~~~~~~~~~~~~~~~
6963
6964For code objects generated by AMDGPU backend for non-amdhsa OS, the runtime does
6965not install a trap handler. The ``llvm.trap`` and ``llvm.debugtrap``
6966instructions are handled as follows:
6967
6968  .. table:: AMDGPU Trap Handler for Non-AMDHSA OS
6969     :name: amdgpu-trap-handler-for-non-amdhsa-os-table
6970
6971     =============== =============== ===========================================
6972     Usage           Code Sequence   Description
6973     =============== =============== ===========================================
6974     llvm.trap       s_endpgm        Causes wavefront to be terminated.
6975     llvm.debugtrap  *none*          Compiler warning given that there is no
6976                                     trap handler installed.
6977     =============== =============== ===========================================
6978
6979Source Languages
6980================
6981
6982.. _amdgpu-opencl:
6983
6984OpenCL
6985------
6986
6987When the language is OpenCL the following differences occur:
6988
69891. The OpenCL memory model is used (see :ref:`amdgpu-amdhsa-memory-model`).
69902. The AMDGPU backend appends additional arguments to the kernel's explicit
6991   arguments for the AMDHSA OS (see
6992   :ref:`opencl-kernel-implicit-arguments-appended-for-amdhsa-os-table`).
69933. Additional metadata is generated
6994   (see :ref:`amdgpu-amdhsa-code-object-metadata`).
6995
6996  .. table:: OpenCL kernel implicit arguments appended for AMDHSA OS
6997     :name: opencl-kernel-implicit-arguments-appended-for-amdhsa-os-table
6998
6999     ======== ==== ========= ===========================================
7000     Position Byte Byte      Description
7001              Size Alignment
7002     ======== ==== ========= ===========================================
7003     1        8    8         OpenCL Global Offset X
7004     2        8    8         OpenCL Global Offset Y
7005     3        8    8         OpenCL Global Offset Z
7006     4        8    8         OpenCL address of printf buffer
7007     5        8    8         OpenCL address of virtual queue used by
7008                             enqueue_kernel.
7009     6        8    8         OpenCL address of AqlWrap struct used by
7010                             enqueue_kernel.
7011     7        8    8         Pointer argument used for Multi-gird
7012                             synchronization.
7013     ======== ==== ========= ===========================================
7014
7015.. _amdgpu-hcc:
7016
7017HCC
7018---
7019
7020When the language is HCC the following differences occur:
7021
70221. The HSA memory model is used (see :ref:`amdgpu-amdhsa-memory-model`).
7023
7024.. _amdgpu-assembler:
7025
7026Assembler
7027---------
7028
7029AMDGPU backend has LLVM-MC based assembler which is currently in development.
7030It supports AMDGCN GFX6-GFX10.
7031
7032This section describes general syntax for instructions and operands.
7033
7034Instructions
7035~~~~~~~~~~~~
7036
7037An instruction has the following :doc:`syntax<AMDGPUInstructionSyntax>`:
7038
7039  | ``<``\ *opcode*\ ``> <``\ *operand0*\ ``>, <``\ *operand1*\ ``>,...
7040    <``\ *modifier0*\ ``> <``\ *modifier1*\ ``>...``
7041
7042:doc:`Operands<AMDGPUOperandSyntax>` are comma-separated while
7043:doc:`modifiers<AMDGPUModifierSyntax>` are space-separated.
7044
7045The order of operands and modifiers is fixed.
7046Most modifiers are optional and may be omitted.
7047
7048Links to detailed instruction syntax description may be found in the following
7049table. Note that features under development are not included
7050in this description.
7051
7052    =================================== =======================================
7053    Core ISA                            ISA Extensions
7054    =================================== =======================================
7055    :doc:`GFX7<AMDGPU/AMDGPUAsmGFX7>`   \-
7056    :doc:`GFX8<AMDGPU/AMDGPUAsmGFX8>`   \-
7057    :doc:`GFX9<AMDGPU/AMDGPUAsmGFX9>`   :doc:`gfx900<AMDGPU/AMDGPUAsmGFX900>`
7058
7059                                        :doc:`gfx902<AMDGPU/AMDGPUAsmGFX900>`
7060
7061                                        :doc:`gfx904<AMDGPU/AMDGPUAsmGFX904>`
7062
7063                                        :doc:`gfx906<AMDGPU/AMDGPUAsmGFX906>`
7064
7065                                        :doc:`gfx908<AMDGPU/AMDGPUAsmGFX908>`
7066
7067                                        :doc:`gfx909<AMDGPU/AMDGPUAsmGFX900>`
7068
7069    :doc:`GFX10<AMDGPU/AMDGPUAsmGFX10>` :doc:`gfx1011<AMDGPU/AMDGPUAsmGFX1011>`
7070
7071                                        :doc:`gfx1012<AMDGPU/AMDGPUAsmGFX1011>`
7072    =================================== =======================================
7073
7074For more information about instructions, their semantics and supported
7075combinations of operands, refer to one of instruction set architecture manuals
7076[AMD-GCN-GFX6]_, [AMD-GCN-GFX7]_, [AMD-GCN-GFX8]_, [AMD-GCN-GFX9]_ and
7077[AMD-GCN-GFX10]_.
7078
7079Operands
7080~~~~~~~~
7081
7082Detailed description of operands may be found :doc:`here<AMDGPUOperandSyntax>`.
7083
7084Modifiers
7085~~~~~~~~~
7086
7087Detailed description of modifiers may be found
7088:doc:`here<AMDGPUModifierSyntax>`.
7089
7090Instruction Examples
7091~~~~~~~~~~~~~~~~~~~~
7092
7093DS
7094++
7095
7096.. code-block:: nasm
7097
7098  ds_add_u32 v2, v4 offset:16
7099  ds_write_src2_b64 v2 offset0:4 offset1:8
7100  ds_cmpst_f32 v2, v4, v6
7101  ds_min_rtn_f64 v[8:9], v2, v[4:5]
7102
7103For full list of supported instructions, refer to "LDS/GDS instructions" in ISA
7104Manual.
7105
7106FLAT
7107++++
7108
7109.. code-block:: nasm
7110
7111  flat_load_dword v1, v[3:4]
7112  flat_store_dwordx3 v[3:4], v[5:7]
7113  flat_atomic_swap v1, v[3:4], v5 glc
7114  flat_atomic_cmpswap v1, v[3:4], v[5:6] glc slc
7115  flat_atomic_fmax_x2 v[1:2], v[3:4], v[5:6] glc
7116
7117For full list of supported instructions, refer to "FLAT instructions" in ISA
7118Manual.
7119
7120MUBUF
7121+++++
7122
7123.. code-block:: nasm
7124
7125  buffer_load_dword v1, off, s[4:7], s1
7126  buffer_store_dwordx4 v[1:4], v2, ttmp[4:7], s1 offen offset:4 glc tfe
7127  buffer_store_format_xy v[1:2], off, s[4:7], s1
7128  buffer_wbinvl1
7129  buffer_atomic_inc v1, v2, s[8:11], s4 idxen offset:4 slc
7130
7131For full list of supported instructions, refer to "MUBUF Instructions" in ISA
7132Manual.
7133
7134SMRD/SMEM
7135+++++++++
7136
7137.. code-block:: nasm
7138
7139  s_load_dword s1, s[2:3], 0xfc
7140  s_load_dwordx8 s[8:15], s[2:3], s4
7141  s_load_dwordx16 s[88:103], s[2:3], s4
7142  s_dcache_inv_vol
7143  s_memtime s[4:5]
7144
7145For full list of supported instructions, refer to "Scalar Memory Operations" in
7146ISA Manual.
7147
7148SOP1
7149++++
7150
7151.. code-block:: nasm
7152
7153  s_mov_b32 s1, s2
7154  s_mov_b64 s[0:1], 0x80000000
7155  s_cmov_b32 s1, 200
7156  s_wqm_b64 s[2:3], s[4:5]
7157  s_bcnt0_i32_b64 s1, s[2:3]
7158  s_swappc_b64 s[2:3], s[4:5]
7159  s_cbranch_join s[4:5]
7160
7161For full list of supported instructions, refer to "SOP1 Instructions" in ISA
7162Manual.
7163
7164SOP2
7165++++
7166
7167.. code-block:: nasm
7168
7169  s_add_u32 s1, s2, s3
7170  s_and_b64 s[2:3], s[4:5], s[6:7]
7171  s_cselect_b32 s1, s2, s3
7172  s_andn2_b32 s2, s4, s6
7173  s_lshr_b64 s[2:3], s[4:5], s6
7174  s_ashr_i32 s2, s4, s6
7175  s_bfm_b64 s[2:3], s4, s6
7176  s_bfe_i64 s[2:3], s[4:5], s6
7177  s_cbranch_g_fork s[4:5], s[6:7]
7178
7179For full list of supported instructions, refer to "SOP2 Instructions" in ISA
7180Manual.
7181
7182SOPC
7183++++
7184
7185.. code-block:: nasm
7186
7187  s_cmp_eq_i32 s1, s2
7188  s_bitcmp1_b32 s1, s2
7189  s_bitcmp0_b64 s[2:3], s4
7190  s_setvskip s3, s5
7191
7192For full list of supported instructions, refer to "SOPC Instructions" in ISA
7193Manual.
7194
7195SOPP
7196++++
7197
7198.. code-block:: nasm
7199
7200  s_barrier
7201  s_nop 2
7202  s_endpgm
7203  s_waitcnt 0 ; Wait for all counters to be 0
7204  s_waitcnt vmcnt(0) & expcnt(0) & lgkmcnt(0) ; Equivalent to above
7205  s_waitcnt vmcnt(1) ; Wait for vmcnt counter to be 1.
7206  s_sethalt 9
7207  s_sleep 10
7208  s_sendmsg 0x1
7209  s_sendmsg sendmsg(MSG_INTERRUPT)
7210  s_trap 1
7211
7212For full list of supported instructions, refer to "SOPP Instructions" in ISA
7213Manual.
7214
7215Unless otherwise mentioned, little verification is performed on the operands
7216of SOPP Instructions, so it is up to the programmer to be familiar with the
7217range or acceptable values.
7218
7219VALU
7220++++
7221
7222For vector ALU instruction opcodes (VOP1, VOP2, VOP3, VOPC, VOP_DPP, VOP_SDWA),
7223the assembler will automatically use optimal encoding based on its operands. To
7224force specific encoding, one can add a suffix to the opcode of the instruction:
7225
7226* _e32 for 32-bit VOP1/VOP2/VOPC
7227* _e64 for 64-bit VOP3
7228* _dpp for VOP_DPP
7229* _sdwa for VOP_SDWA
7230
7231VOP1/VOP2/VOP3/VOPC examples:
7232
7233.. code-block:: nasm
7234
7235  v_mov_b32 v1, v2
7236  v_mov_b32_e32 v1, v2
7237  v_nop
7238  v_cvt_f64_i32_e32 v[1:2], v2
7239  v_floor_f32_e32 v1, v2
7240  v_bfrev_b32_e32 v1, v2
7241  v_add_f32_e32 v1, v2, v3
7242  v_mul_i32_i24_e64 v1, v2, 3
7243  v_mul_i32_i24_e32 v1, -3, v3
7244  v_mul_i32_i24_e32 v1, -100, v3
7245  v_addc_u32 v1, s[0:1], v2, v3, s[2:3]
7246  v_max_f16_e32 v1, v2, v3
7247
7248VOP_DPP examples:
7249
7250.. code-block:: nasm
7251
7252  v_mov_b32 v0, v0 quad_perm:[0,2,1,1]
7253  v_sin_f32 v0, v0 row_shl:1 row_mask:0xa bank_mask:0x1 bound_ctrl:0
7254  v_mov_b32 v0, v0 wave_shl:1
7255  v_mov_b32 v0, v0 row_mirror
7256  v_mov_b32 v0, v0 row_bcast:31
7257  v_mov_b32 v0, v0 quad_perm:[1,3,0,1] row_mask:0xa bank_mask:0x1 bound_ctrl:0
7258  v_add_f32 v0, v0, |v0| row_shl:1 row_mask:0xa bank_mask:0x1 bound_ctrl:0
7259  v_max_f16 v1, v2, v3 row_shl:1 row_mask:0xa bank_mask:0x1 bound_ctrl:0
7260
7261VOP_SDWA examples:
7262
7263.. code-block:: nasm
7264
7265  v_mov_b32 v1, v2 dst_sel:BYTE_0 dst_unused:UNUSED_PRESERVE src0_sel:DWORD
7266  v_min_u32 v200, v200, v1 dst_sel:WORD_1 dst_unused:UNUSED_PAD src0_sel:BYTE_1 src1_sel:DWORD
7267  v_sin_f32 v0, v0 dst_unused:UNUSED_PAD src0_sel:WORD_1
7268  v_fract_f32 v0, |v0| dst_sel:DWORD dst_unused:UNUSED_PAD src0_sel:WORD_1
7269  v_cmpx_le_u32 vcc, v1, v2 src0_sel:BYTE_2 src1_sel:WORD_0
7270
7271For full list of supported instructions, refer to "Vector ALU instructions".
7272
7273.. TODO::
7274
7275  Remove once we switch to code object v3 by default.
7276
7277.. _amdgpu-amdhsa-assembler-predefined-symbols-v2:
7278
7279Code Object V2 Predefined Symbols (-mattr=-code-object-v3)
7280~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
7281
7282.. warning:: Code Object V2 is not the default code object version emitted by
7283  this version of LLVM. For a description of the predefined symbols available
7284  with the default configuration (Code Object V3) see
7285  :ref:`amdgpu-amdhsa-assembler-predefined-symbols-v3`.
7286
7287The AMDGPU assembler defines and updates some symbols automatically. These
7288symbols do not affect code generation.
7289
7290.option.machine_version_major
7291+++++++++++++++++++++++++++++
7292
7293Set to the GFX major generation number of the target being assembled for. For
7294example, when assembling for a "GFX9" target this will be set to the integer
7295value "9". The possible GFX major generation numbers are presented in
7296:ref:`amdgpu-processors`.
7297
7298.option.machine_version_minor
7299+++++++++++++++++++++++++++++
7300
7301Set to the GFX minor generation number of the target being assembled for. For
7302example, when assembling for a "GFX810" target this will be set to the integer
7303value "1". The possible GFX minor generation numbers are presented in
7304:ref:`amdgpu-processors`.
7305
7306.option.machine_version_stepping
7307++++++++++++++++++++++++++++++++
7308
7309Set to the GFX stepping generation number of the target being assembled for.
7310For example, when assembling for a "GFX704" target this will be set to the
7311integer value "4". The possible GFX stepping generation numbers are presented
7312in :ref:`amdgpu-processors`.
7313
7314.kernel.vgpr_count
7315++++++++++++++++++
7316
7317Set to zero each time a
7318:ref:`amdgpu-amdhsa-assembler-directive-amdgpu_hsa_kernel` directive is
7319encountered. At each instruction, if the current value of this symbol is less
7320than or equal to the maximum VPGR number explicitly referenced within that
7321instruction then the symbol value is updated to equal that VGPR number plus
7322one.
7323
7324.kernel.sgpr_count
7325++++++++++++++++++
7326
7327Set to zero each time a
7328:ref:`amdgpu-amdhsa-assembler-directive-amdgpu_hsa_kernel` directive is
7329encountered. At each instruction, if the current value of this symbol is less
7330than or equal to the maximum VPGR number explicitly referenced within that
7331instruction then the symbol value is updated to equal that SGPR number plus
7332one.
7333
7334.. _amdgpu-amdhsa-assembler-directives-v2:
7335
7336Code Object V2 Directives (-mattr=-code-object-v3)
7337~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
7338
7339.. warning:: Code Object V2 is not the default code object version emitted by
7340  this version of LLVM. For a description of the directives supported with
7341  the default configuration (Code Object V3) see
7342  :ref:`amdgpu-amdhsa-assembler-directives-v3`.
7343
7344AMDGPU ABI defines auxiliary data in output code object. In assembly source,
7345one can specify them with assembler directives.
7346
7347.hsa_code_object_version major, minor
7348+++++++++++++++++++++++++++++++++++++
7349
7350*major* and *minor* are integers that specify the version of the HSA code
7351object that will be generated by the assembler.
7352
7353.hsa_code_object_isa [major, minor, stepping, vendor, arch]
7354+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
7355
7356
7357*major*, *minor*, and *stepping* are all integers that describe the instruction
7358set architecture (ISA) version of the assembly program.
7359
7360*vendor* and *arch* are quoted strings. *vendor* should always be equal to
7361"AMD" and *arch* should always be equal to "AMDGPU".
7362
7363By default, the assembler will derive the ISA version, *vendor*, and *arch*
7364from the value of the -mcpu option that is passed to the assembler.
7365
7366.. _amdgpu-amdhsa-assembler-directive-amdgpu_hsa_kernel:
7367
7368.amdgpu_hsa_kernel (name)
7369+++++++++++++++++++++++++
7370
7371This directives specifies that the symbol with given name is a kernel entry
7372point (label) and the object should contain corresponding symbol of type
7373STT_AMDGPU_HSA_KERNEL.
7374
7375.amd_kernel_code_t
7376++++++++++++++++++
7377
7378This directive marks the beginning of a list of key / value pairs that are used
7379to specify the amd_kernel_code_t object that will be emitted by the assembler.
7380The list must be terminated by the *.end_amd_kernel_code_t* directive. For any
7381amd_kernel_code_t values that are unspecified a default value will be used. The
7382default value for all keys is 0, with the following exceptions:
7383
7384- *amd_code_version_major* defaults to 1.
7385- *amd_kernel_code_version_minor* defaults to 2.
7386- *amd_machine_kind* defaults to 1.
7387- *amd_machine_version_major*, *machine_version_minor*, and
7388  *amd_machine_version_stepping* are derived from the value of the -mcpu option
7389  that is passed to the assembler.
7390- *kernel_code_entry_byte_offset* defaults to 256.
7391- *wavefront_size* defaults 6 for all targets before GFX10. For GFX10 onwards
7392  defaults to 6 if target feature ``wavefrontsize64`` is enabled, otherwise 5.
7393  Note that wavefront size is specified as a power of two, so a value of **n**
7394  means a size of 2^ **n**.
7395- *call_convention* defaults to -1.
7396- *kernarg_segment_alignment*, *group_segment_alignment*, and
7397  *private_segment_alignment* default to 4. Note that alignments are specified
7398  as a power of 2, so a value of **n** means an alignment of 2^ **n**.
7399- *enable_wgp_mode* defaults to 1 if target feature ``cumode`` is disabled for
7400  GFX10 onwards.
7401- *enable_mem_ordered* defaults to 1 for GFX10 onwards.
7402
7403The *.amd_kernel_code_t* directive must be placed immediately after the
7404function label and before any instructions.
7405
7406For a full list of amd_kernel_code_t keys, refer to AMDGPU ABI document,
7407comments in lib/Target/AMDGPU/AmdKernelCodeT.h and test/CodeGen/AMDGPU/hsa.s.
7408
7409.. _amdgpu-amdhsa-assembler-example-v2:
7410
7411Code Object V2 Example Source Code (-mattr=-code-object-v3)
7412~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
7413
7414.. warning:: Code Object V2 is not the default code object version emitted by
7415  this version of LLVM. For a description of the directives supported with
7416  the default configuration (Code Object V3) see
7417  :ref:`amdgpu-amdhsa-assembler-example-v3`.
7418
7419Here is an example of a minimal assembly source file, defining one HSA kernel:
7420
7421.. code::
7422   :number-lines:
7423
7424   .hsa_code_object_version 1,0
7425   .hsa_code_object_isa
7426
7427   .hsatext
7428   .globl  hello_world
7429   .p2align 8
7430   .amdgpu_hsa_kernel hello_world
7431
7432   hello_world:
7433
7434      .amd_kernel_code_t
7435         enable_sgpr_kernarg_segment_ptr = 1
7436         is_ptr64 = 1
7437         compute_pgm_rsrc1_vgprs = 0
7438         compute_pgm_rsrc1_sgprs = 0
7439         compute_pgm_rsrc2_user_sgpr = 2
7440         compute_pgm_rsrc1_wgp_mode = 0
7441         compute_pgm_rsrc1_mem_ordered = 0
7442         compute_pgm_rsrc1_fwd_progress = 1
7443     .end_amd_kernel_code_t
7444
7445     s_load_dwordx2 s[0:1], s[0:1] 0x0
7446     v_mov_b32 v0, 3.14159
7447     s_waitcnt lgkmcnt(0)
7448     v_mov_b32 v1, s0
7449     v_mov_b32 v2, s1
7450     flat_store_dword v[1:2], v0
7451     s_endpgm
7452   .Lfunc_end0:
7453        .size   hello_world, .Lfunc_end0-hello_world
7454
7455.. _amdgpu-amdhsa-assembler-predefined-symbols-v3:
7456
7457Code Object V3 Predefined Symbols (-mattr=+code-object-v3)
7458~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
7459
7460The AMDGPU assembler defines and updates some symbols automatically. These
7461symbols do not affect code generation.
7462
7463.amdgcn.gfx_generation_number
7464+++++++++++++++++++++++++++++
7465
7466Set to the GFX major generation number of the target being assembled for. For
7467example, when assembling for a "GFX9" target this will be set to the integer
7468value "9". The possible GFX major generation numbers are presented in
7469:ref:`amdgpu-processors`.
7470
7471.amdgcn.gfx_generation_minor
7472++++++++++++++++++++++++++++
7473
7474Set to the GFX minor generation number of the target being assembled for. For
7475example, when assembling for a "GFX810" target this will be set to the integer
7476value "1". The possible GFX minor generation numbers are presented in
7477:ref:`amdgpu-processors`.
7478
7479.amdgcn.gfx_generation_stepping
7480+++++++++++++++++++++++++++++++
7481
7482Set to the GFX stepping generation number of the target being assembled for.
7483For example, when assembling for a "GFX704" target this will be set to the
7484integer value "4". The possible GFX stepping generation numbers are presented
7485in :ref:`amdgpu-processors`.
7486
7487.. _amdgpu-amdhsa-assembler-symbol-next_free_vgpr:
7488
7489.amdgcn.next_free_vgpr
7490++++++++++++++++++++++
7491
7492Set to zero before assembly begins. At each instruction, if the current value
7493of this symbol is less than or equal to the maximum VGPR number explicitly
7494referenced within that instruction then the symbol value is updated to equal
7495that VGPR number plus one.
7496
7497May be used to set the `.amdhsa_next_free_vpgr` directive in
7498:ref:`amdhsa-kernel-directives-table`.
7499
7500May be set at any time, e.g. manually set to zero at the start of each kernel.
7501
7502.. _amdgpu-amdhsa-assembler-symbol-next_free_sgpr:
7503
7504.amdgcn.next_free_sgpr
7505++++++++++++++++++++++
7506
7507Set to zero before assembly begins. At each instruction, if the current value
7508of this symbol is less than or equal the maximum SGPR number explicitly
7509referenced within that instruction then the symbol value is updated to equal
7510that SGPR number plus one.
7511
7512May be used to set the `.amdhsa_next_free_spgr` directive in
7513:ref:`amdhsa-kernel-directives-table`.
7514
7515May be set at any time, e.g. manually set to zero at the start of each kernel.
7516
7517.. _amdgpu-amdhsa-assembler-directives-v3:
7518
7519Code Object V3 Directives (-mattr=+code-object-v3)
7520~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
7521
7522Directives which begin with ``.amdgcn`` are valid for all ``amdgcn``
7523architecture processors, and are not OS-specific. Directives which begin with
7524``.amdhsa`` are specific to ``amdgcn`` architecture processors when the
7525``amdhsa`` OS is specified. See :ref:`amdgpu-target-triples` and
7526:ref:`amdgpu-processors`.
7527
7528.amdgcn_target <target>
7529+++++++++++++++++++++++
7530
7531Optional directive which declares the target supported by the containing
7532assembler source file. Valid values are described in
7533:ref:`amdgpu-amdhsa-code-object-target-identification`. Used by the assembler
7534to validate command-line options such as ``-triple``, ``-mcpu``, and those
7535which specify target features.
7536
7537.amdhsa_kernel <name>
7538+++++++++++++++++++++
7539
7540Creates a correctly aligned AMDHSA kernel descriptor and a symbol,
7541``<name>.kd``, in the current location of the current section. Only valid when
7542the OS is ``amdhsa``. ``<name>`` must be a symbol that labels the first
7543instruction to execute, and does not need to be previously defined.
7544
7545Marks the beginning of a list of directives used to generate the bytes of a
7546kernel descriptor, as described in :ref:`amdgpu-amdhsa-kernel-descriptor`.
7547Directives which may appear in this list are described in
7548:ref:`amdhsa-kernel-directives-table`. Directives may appear in any order, must
7549be valid for the target being assembled for, and cannot be repeated. Directives
7550support the range of values specified by the field they reference in
7551:ref:`amdgpu-amdhsa-kernel-descriptor`. If a directive is not specified, it is
7552assumed to have its default value, unless it is marked as "Required", in which
7553case it is an error to omit the directive. This list of directives is
7554terminated by an ``.end_amdhsa_kernel`` directive.
7555
7556  .. table:: AMDHSA Kernel Assembler Directives
7557     :name: amdhsa-kernel-directives-table
7558
7559     ======================================================== =================== ============ ===================
7560     Directive                                                Default             Supported On Description
7561     ======================================================== =================== ============ ===================
7562     ``.amdhsa_group_segment_fixed_size``                     0                   GFX6-GFX10   Controls GROUP_SEGMENT_FIXED_SIZE in
7563                                                                                               :ref:`amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table`.
7564     ``.amdhsa_private_segment_fixed_size``                   0                   GFX6-GFX10   Controls PRIVATE_SEGMENT_FIXED_SIZE in
7565                                                                                               :ref:`amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table`.
7566     ``.amdhsa_user_sgpr_private_segment_buffer``             0                   GFX6-GFX10   Controls ENABLE_SGPR_PRIVATE_SEGMENT_BUFFER in
7567                                                                                               :ref:`amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table`.
7568     ``.amdhsa_user_sgpr_dispatch_ptr``                       0                   GFX6-GFX10   Controls ENABLE_SGPR_DISPATCH_PTR in
7569                                                                                               :ref:`amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table`.
7570     ``.amdhsa_user_sgpr_queue_ptr``                          0                   GFX6-GFX10   Controls ENABLE_SGPR_QUEUE_PTR in
7571                                                                                               :ref:`amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table`.
7572     ``.amdhsa_user_sgpr_kernarg_segment_ptr``                0                   GFX6-GFX10   Controls ENABLE_SGPR_KERNARG_SEGMENT_PTR in
7573                                                                                               :ref:`amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table`.
7574     ``.amdhsa_user_sgpr_dispatch_id``                        0                   GFX6-GFX10   Controls ENABLE_SGPR_DISPATCH_ID in
7575                                                                                               :ref:`amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table`.
7576     ``.amdhsa_user_sgpr_flat_scratch_init``                  0                   GFX6-GFX10   Controls ENABLE_SGPR_FLAT_SCRATCH_INIT in
7577                                                                                               :ref:`amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table`.
7578     ``.amdhsa_user_sgpr_private_segment_size``               0                   GFX6-GFX10   Controls ENABLE_SGPR_PRIVATE_SEGMENT_SIZE in
7579                                                                                               :ref:`amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table`.
7580     ``.amdhsa_wavefront_size32``                             Target              GFX10        Controls ENABLE_WAVEFRONT_SIZE32 in
7581                                                              Feature                          :ref:`amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table`.
7582                                                              Specific
7583                                                              (-wavefrontsize64)
7584     ``.amdhsa_system_sgpr_private_segment_wavefront_offset`` 0                   GFX6-GFX10   Controls ENABLE_SGPR_PRIVATE_SEGMENT_WAVEFRONT_OFFSET in
7585                                                                                               :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`.
7586     ``.amdhsa_system_sgpr_workgroup_id_x``                   1                   GFX6-GFX10   Controls ENABLE_SGPR_WORKGROUP_ID_X in
7587                                                                                               :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`.
7588     ``.amdhsa_system_sgpr_workgroup_id_y``                   0                   GFX6-GFX10   Controls ENABLE_SGPR_WORKGROUP_ID_Y in
7589                                                                                               :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`.
7590     ``.amdhsa_system_sgpr_workgroup_id_z``                   0                   GFX6-GFX10   Controls ENABLE_SGPR_WORKGROUP_ID_Z in
7591                                                                                               :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`.
7592     ``.amdhsa_system_sgpr_workgroup_info``                   0                   GFX6-GFX10   Controls ENABLE_SGPR_WORKGROUP_INFO in
7593                                                                                               :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`.
7594     ``.amdhsa_system_vgpr_workitem_id``                      0                   GFX6-GFX10   Controls ENABLE_VGPR_WORKITEM_ID in
7595                                                                                               :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`.
7596                                                                                               Possible values are defined in
7597                                                                                               :ref:`amdgpu-amdhsa-system-vgpr-work-item-id-enumeration-values-table`.
7598     ``.amdhsa_next_free_vgpr``                               Required            GFX6-GFX10   Maximum VGPR number explicitly referenced, plus one.
7599                                                                                               Used to calculate GRANULATED_WORKITEM_VGPR_COUNT in
7600                                                                                               :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`.
7601     ``.amdhsa_next_free_sgpr``                               Required            GFX6-GFX10   Maximum SGPR number explicitly referenced, plus one.
7602                                                                                               Used to calculate GRANULATED_WAVEFRONT_SGPR_COUNT in
7603                                                                                               :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`.
7604     ``.amdhsa_reserve_vcc``                                  1                   GFX6-GFX10   Whether the kernel may use the special VCC SGPR.
7605                                                                                               Used to calculate GRANULATED_WAVEFRONT_SGPR_COUNT in
7606                                                                                               :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`.
7607     ``.amdhsa_reserve_flat_scratch``                         1                   GFX7-GFX10   Whether the kernel may use flat instructions to access
7608                                                                                               scratch memory. Used to calculate
7609                                                                                               GRANULATED_WAVEFRONT_SGPR_COUNT in
7610                                                                                               :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`.
7611     ``.amdhsa_reserve_xnack_mask``                           Target              GFX8-GFX10   Whether the kernel may trigger XNACK replay.
7612                                                              Feature                          Used to calculate GRANULATED_WAVEFRONT_SGPR_COUNT in
7613                                                              Specific                         :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`.
7614                                                              (+xnack)
7615     ``.amdhsa_float_round_mode_32``                          0                   GFX6-GFX10   Controls FLOAT_ROUND_MODE_32 in
7616                                                                                               :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`.
7617                                                                                               Possible values are defined in
7618                                                                                               :ref:`amdgpu-amdhsa-floating-point-rounding-mode-enumeration-values-table`.
7619     ``.amdhsa_float_round_mode_16_64``                       0                   GFX6-GFX10   Controls FLOAT_ROUND_MODE_16_64 in
7620                                                                                               :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`.
7621                                                                                               Possible values are defined in
7622                                                                                               :ref:`amdgpu-amdhsa-floating-point-rounding-mode-enumeration-values-table`.
7623     ``.amdhsa_float_denorm_mode_32``                         0                   GFX6-GFX10   Controls FLOAT_DENORM_MODE_32 in
7624                                                                                               :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`.
7625                                                                                               Possible values are defined in
7626                                                                                               :ref:`amdgpu-amdhsa-floating-point-denorm-mode-enumeration-values-table`.
7627     ``.amdhsa_float_denorm_mode_16_64``                      3                   GFX6-GFX10   Controls FLOAT_DENORM_MODE_16_64 in
7628                                                                                               :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`.
7629                                                                                               Possible values are defined in
7630                                                                                               :ref:`amdgpu-amdhsa-floating-point-denorm-mode-enumeration-values-table`.
7631     ``.amdhsa_dx10_clamp``                                   1                   GFX6-GFX10   Controls ENABLE_DX10_CLAMP in
7632                                                                                               :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`.
7633     ``.amdhsa_ieee_mode``                                    1                   GFX6-GFX10   Controls ENABLE_IEEE_MODE in
7634                                                                                               :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`.
7635     ``.amdhsa_fp16_overflow``                                0                   GFX9-GFX10   Controls FP16_OVFL in
7636                                                                                               :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`.
7637     ``.amdhsa_workgroup_processor_mode``                     Target              GFX10        Controls ENABLE_WGP_MODE in
7638                                                              Feature                          :ref:`amdgpu-amdhsa-kernel-descriptor-gfx6-gfx10-table`.
7639                                                              Specific
7640                                                              (-cumode)
7641     ``.amdhsa_memory_ordered``                               1                   GFX10        Controls MEM_ORDERED in
7642                                                                                               :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`.
7643     ``.amdhsa_forward_progress``                             0                   GFX10        Controls FWD_PROGRESS in
7644                                                                                               :ref:`amdgpu-amdhsa-compute_pgm_rsrc1-gfx6-gfx10-table`.
7645     ``.amdhsa_exception_fp_ieee_invalid_op``                 0                   GFX6-GFX10   Controls ENABLE_EXCEPTION_IEEE_754_FP_INVALID_OPERATION in
7646                                                                                               :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`.
7647     ``.amdhsa_exception_fp_denorm_src``                      0                   GFX6-GFX10   Controls ENABLE_EXCEPTION_FP_DENORMAL_SOURCE in
7648                                                                                               :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`.
7649     ``.amdhsa_exception_fp_ieee_div_zero``                   0                   GFX6-GFX10   Controls ENABLE_EXCEPTION_IEEE_754_FP_DIVISION_BY_ZERO in
7650                                                                                               :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`.
7651     ``.amdhsa_exception_fp_ieee_overflow``                   0                   GFX6-GFX10   Controls ENABLE_EXCEPTION_IEEE_754_FP_OVERFLOW in
7652                                                                                               :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`.
7653     ``.amdhsa_exception_fp_ieee_underflow``                  0                   GFX6-GFX10   Controls ENABLE_EXCEPTION_IEEE_754_FP_UNDERFLOW in
7654                                                                                               :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`.
7655     ``.amdhsa_exception_fp_ieee_inexact``                    0                   GFX6-GFX10   Controls ENABLE_EXCEPTION_IEEE_754_FP_INEXACT in
7656                                                                                               :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`.
7657     ``.amdhsa_exception_int_div_zero``                       0                   GFX6-GFX10   Controls ENABLE_EXCEPTION_INT_DIVIDE_BY_ZERO in
7658                                                                                               :ref:`amdgpu-amdhsa-compute_pgm_rsrc2-gfx6-gfx10-table`.
7659     ======================================================== =================== ============ ===================
7660
7661.amdgpu_metadata
7662++++++++++++++++
7663
7664Optional directive which declares the contents of the ``NT_AMDGPU_METADATA``
7665note record (see :ref:`amdgpu-elf-note-records-table-v3`).
7666
7667The contents must be in the [YAML]_ markup format, with the same structure and
7668semantics described in :ref:`amdgpu-amdhsa-code-object-metadata-v3`.
7669
7670This directive is terminated by an ``.end_amdgpu_metadata`` directive.
7671
7672.. _amdgpu-amdhsa-assembler-example-v3:
7673
7674Code Object V3 Example Source Code (-mattr=+code-object-v3)
7675~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
7676
7677Here is an example of a minimal assembly source file, defining one HSA kernel:
7678
7679.. code::
7680   :number-lines:
7681
7682   .amdgcn_target "amdgcn-amd-amdhsa--gfx900+xnack" // optional
7683
7684   .text
7685   .globl hello_world
7686   .p2align 8
7687   .type hello_world,@function
7688   hello_world:
7689     s_load_dwordx2 s[0:1], s[0:1] 0x0
7690     v_mov_b32 v0, 3.14159
7691     s_waitcnt lgkmcnt(0)
7692     v_mov_b32 v1, s0
7693     v_mov_b32 v2, s1
7694     flat_store_dword v[1:2], v0
7695     s_endpgm
7696   .Lfunc_end0:
7697     .size   hello_world, .Lfunc_end0-hello_world
7698
7699   .rodata
7700   .p2align 6
7701   .amdhsa_kernel hello_world
7702     .amdhsa_user_sgpr_kernarg_segment_ptr 1
7703     .amdhsa_next_free_vgpr .amdgcn.next_free_vgpr
7704     .amdhsa_next_free_sgpr .amdgcn.next_free_sgpr
7705   .end_amdhsa_kernel
7706
7707   .amdgpu_metadata
7708   ---
7709   amdhsa.version:
7710     - 1
7711     - 0
7712   amdhsa.kernels:
7713     - .name: hello_world
7714       .symbol: hello_world.kd
7715       .kernarg_segment_size: 48
7716       .group_segment_fixed_size: 0
7717       .private_segment_fixed_size: 0
7718       .kernarg_segment_align: 4
7719       .wavefront_size: 64
7720       .sgpr_count: 2
7721       .vgpr_count: 3
7722       .max_flat_workgroup_size: 256
7723   ...
7724   .end_amdgpu_metadata
7725
7726If an assembly source file contains multiple kernels and/or functions, the
7727:ref:`amdgpu-amdhsa-assembler-symbol-next_free_vgpr` and
7728:ref:`amdgpu-amdhsa-assembler-symbol-next_free_sgpr` symbols may be reset using
7729the ``.set <symbol>, <expression>`` directive. For example, in the case of two
7730kernels, where ``function1`` is only called from ``kernel1`` it is sufficient
7731to group the function with the kernel that calls it and reset the symbols
7732between the two connected components:
7733
7734.. code::
7735   :number-lines:
7736
7737   .amdgcn_target "amdgcn-amd-amdhsa--gfx900+xnack" // optional
7738
7739   // gpr tracking symbols are implicitly set to zero
7740
7741   .text
7742   .globl kern0
7743   .p2align 8
7744   .type kern0,@function
7745   kern0:
7746     // ...
7747     s_endpgm
7748   .Lkern0_end:
7749     .size   kern0, .Lkern0_end-kern0
7750
7751   .rodata
7752   .p2align 6
7753   .amdhsa_kernel kern0
7754     // ...
7755     .amdhsa_next_free_vgpr .amdgcn.next_free_vgpr
7756     .amdhsa_next_free_sgpr .amdgcn.next_free_sgpr
7757   .end_amdhsa_kernel
7758
7759   // reset symbols to begin tracking usage in func1 and kern1
7760   .set .amdgcn.next_free_vgpr, 0
7761   .set .amdgcn.next_free_sgpr, 0
7762
7763   .text
7764   .hidden func1
7765   .global func1
7766   .p2align 2
7767   .type func1,@function
7768   func1:
7769     // ...
7770     s_setpc_b64 s[30:31]
7771   .Lfunc1_end:
7772   .size func1, .Lfunc1_end-func1
7773
7774   .globl kern1
7775   .p2align 8
7776   .type kern1,@function
7777   kern1:
7778     // ...
7779     s_getpc_b64 s[4:5]
7780     s_add_u32 s4, s4, func1@rel32@lo+4
7781     s_addc_u32 s5, s5, func1@rel32@lo+4
7782     s_swappc_b64 s[30:31], s[4:5]
7783     // ...
7784     s_endpgm
7785   .Lkern1_end:
7786     .size   kern1, .Lkern1_end-kern1
7787
7788   .rodata
7789   .p2align 6
7790   .amdhsa_kernel kern1
7791     // ...
7792     .amdhsa_next_free_vgpr .amdgcn.next_free_vgpr
7793     .amdhsa_next_free_sgpr .amdgcn.next_free_sgpr
7794   .end_amdhsa_kernel
7795
7796These symbols cannot identify connected components in order to automatically
7797track the usage for each kernel. However, in some cases careful organization of
7798the kernels and functions in the source file means there is minimal additional
7799effort required to accurately calculate GPR usage.
7800
7801Additional Documentation
7802========================
7803
7804.. [AMD-GCN-GFX6] `AMD Southern Islands Series ISA <http://developer.amd.com/wordpress/media/2012/12/AMD_Southern_Islands_Instruction_Set_Architecture.pdf>`__
7805.. [AMD-GCN-GFX7] `AMD Sea Islands Series ISA <http://developer.amd.com/wordpress/media/2013/07/AMD_Sea_Islands_Instruction_Set_Architecture.pdf>`_
7806.. [AMD-GCN-GFX8] `AMD GCN3 Instruction Set Architecture <http://amd-dev.wpengine.netdna-cdn.com/wordpress/media/2013/12/AMD_GCN3_Instruction_Set_Architecture_rev1.1.pdf>`__
7807.. [AMD-GCN-GFX9] `AMD "Vega" Instruction Set Architecture <http://developer.amd.com/wordpress/media/2013/12/Vega_Shader_ISA_28July2017.pdf>`__
7808.. [AMD-GCN-GFX10] `AMD "RDNA 1.0" Instruction Set Architecture <https://gpuopen.com/wp-content/uploads/2019/08/RDNA_Shader_ISA_5August2019.pdf>`__
7809.. [AMD-RADEON-HD-2000-3000] `AMD R6xx shader ISA <http://developer.amd.com/wordpress/media/2012/10/R600_Instruction_Set_Architecture.pdf>`__
7810.. [AMD-RADEON-HD-4000] `AMD R7xx shader ISA <http://developer.amd.com/wordpress/media/2012/10/R700-Family_Instruction_Set_Architecture.pdf>`__
7811.. [AMD-RADEON-HD-5000] `AMD Evergreen shader ISA <http://developer.amd.com/wordpress/media/2012/10/AMD_Evergreen-Family_Instruction_Set_Architecture.pdf>`__
7812.. [AMD-RADEON-HD-6000] `AMD Cayman/Trinity shader ISA <http://developer.amd.com/wordpress/media/2012/10/AMD_HD_6900_Series_Instruction_Set_Architecture.pdf>`__
7813.. [AMD-ROCm] `AMD ROCm Platform <https://rocm-documentation.readthedocs.io>`__
7814.. [AMD-ROCm-github] `ROCm github <http://github.com/RadeonOpenCompute>`__
7815.. [CLANG-ATTR] `Attributes in Clang <https://clang.llvm.org/docs/AttributeReference.html>`__
7816.. [DWARF] `DWARF Debugging Information Format <http://dwarfstd.org/>`__
7817.. [ELF] `Executable and Linkable Format (ELF) <http://www.sco.com/developers/gabi/>`__
7818.. [HRF] `Heterogeneous-race-free Memory Models <http://benedictgaster.org/wp-content/uploads/2014/01/asplos269-FINAL.pdf>`__
7819.. [HSA] `Heterogeneous System Architecture (HSA) Foundation <http://www.hsafoundation.com/>`__
7820.. [MsgPack] `Message Pack <http://www.msgpack.org/>`__
7821.. [OpenCL] `The OpenCL Specification Version 2.0 <http://www.khronos.org/registry/cl/specs/opencl-2.0.pdf>`__
7822.. [SEMVER] `Semantic Versioning <https://semver.org/>`__
7823.. [YAML] `YAML Ain't Markup Language (YAML™) Version 1.2 <http://www.yaml.org/spec/1.2/spec.html>`__
7824