1# MLIR Language Reference 2 3MLIR (Multi-Level IR) is a compiler intermediate representation with 4similarities to traditional three-address SSA representations (like 5[LLVM IR](http://llvm.org/docs/LangRef.html) or 6[SIL](https://github.com/apple/swift/blob/master/docs/SIL.rst)), but which 7introduces notions from polyhedral loop optimization as first-class concepts. 8This hybrid design is optimized to represent, analyze, and transform high level 9dataflow graphs as well as target-specific code generated for high performance 10data parallel systems. Beyond its representational capabilities, its single 11continuous design provides a framework to lower from dataflow graphs to 12high-performance target-specific code. 13 14This document defines and describes the key concepts in MLIR, and is intended 15to be a dry reference document - the [rationale 16documentation](Rationale/Rationale.md), 17[glossary](../getting_started/Glossary.md), and other content are hosted 18elsewhere. 19 20MLIR is designed to be used in three different forms: a human-readable textual 21form suitable for debugging, an in-memory form suitable for programmatic 22transformations and analysis, and a compact serialized form suitable for 23storage and transport. The different forms all describe the same semantic 24content. This document describes the human-readable textual form. 25 26[TOC] 27 28## High-Level Structure 29 30MLIR is fundamentally based on a graph-like data structure of nodes, called 31*Operations*, and edges, called *Values*. Each Value is the result of exactly 32one Operation or Block Argument, and has a *Value Type* defined by the [type 33system](#type-system). [Operations](#operations) are contained in 34[Blocks](#blocks) and Blocks are contained in [Regions](#regions). Operations 35are also ordered within their containing block and Blocks are ordered in their 36containing region, although this order may or may not be semantically 37meaningful in a given [kind of region](Interfaces.md#regionkindinterfaces)). 38Operations may also contain regions, enabling hierarchical structures to be 39represented. 40 41Operations can represent many different concepts, from higher-level concepts 42like function definitions, function calls, buffer allocations, view or slices 43of buffers, and process creation, to lower-level concepts like 44target-independent arithmetic, target-specific instructions, configuration 45registers, and logic gates. These different concepts are represented by 46different operations in MLIR and the set of operations usable in MLIR can be 47arbitrarily extended. 48 49MLIR also provides an extensible framework for transformations on operations, 50using familiar concepts of compiler [Passes](Passes.md). Enabling an arbitrary 51set of passes on an arbitrary set of operations results in a significant 52scaling challenge, since each transformation must potentially take into 53account the semantics of any operation. MLIR addresses this complexity by 54allowing operation semantics to be described abstractly using 55[Traits](Traits.md) and [Interfaces](Interfaces.md), enabling transformations 56to operate on operations more generically. Traits often describe verification 57constraints on valid IR, enabling complex invariants to be captured and 58checked. (see [Op vs 59Operation](docs/Tutorials/Toy/Ch-2/#op-vs-operation-using-mlir-operations)) 60 61One obvious application of MLIR is to represent an 62[SSA-based](https://en.wikipedia.org/wiki/Static_single_assignment_form) IR, 63like the LLVM core IR, with appropriate choice of Operation Types to define 64[Modules](#module), [Functions](#functions), Branches, Allocations, and 65verification constraints to ensure the SSA Dominance property. MLIR includes a 66'standard' dialect which defines just such structures. However, MLIR is 67intended to be general enough to represent other compiler-like data 68structures, such as Abstract Syntax Trees in a language frontend, generated 69instructions in a target-specific backend, or circuits in a High-Level 70Synthesis tool. 71 72Here's an example of an MLIR module: 73 74```mlir 75// Compute A*B using an implementation of multiply kernel and print the 76// result using a TensorFlow op. The dimensions of A and B are partially 77// known. The shapes are assumed to match. 78func @mul(%A: tensor<100x?xf32>, %B: tensor<?x50xf32>) -> (tensor<100x50xf32>) { 79 // Compute the inner dimension of %A using the dim operation. 80 %n = dim %A, 1 : tensor<100x?xf32> 81 82 // Allocate addressable "buffers" and copy tensors %A and %B into them. 83 %A_m = alloc(%n) : memref<100x?xf32> 84 tensor_store %A to %A_m : memref<100x?xf32> 85 86 %B_m = alloc(%n) : memref<?x50xf32> 87 tensor_store %B to %B_m : memref<?x50xf32> 88 89 // Call function @multiply passing memrefs as arguments, 90 // and getting returned the result of the multiplication. 91 %C_m = call @multiply(%A_m, %B_m) 92 : (memref<100x?xf32>, memref<?x50xf32>) -> (memref<100x50xf32>) 93 94 dealloc %A_m : memref<100x?xf32> 95 dealloc %B_m : memref<?x50xf32> 96 97 // Load the buffer data into a higher level "tensor" value. 98 %C = tensor_load %C_m : memref<100x50xf32> 99 dealloc %C_m : memref<100x50xf32> 100 101 // Call TensorFlow built-in function to print the result tensor. 102 "tf.Print"(%C){message: "mul result"} 103 : (tensor<100x50xf32) -> (tensor<100x50xf32>) 104 105 return %C : tensor<100x50xf32> 106} 107 108// A function that multiplies two memrefs and returns the result. 109func @multiply(%A: memref<100x?xf32>, %B: memref<?x50xf32>) 110 -> (memref<100x50xf32>) { 111 // Compute the inner dimension of %A. 112 %n = dim %A, 1 : memref<100x?xf32> 113 114 // Allocate memory for the multiplication result. 115 %C = alloc() : memref<100x50xf32> 116 117 // Multiplication loop nest. 118 affine.for %i = 0 to 100 { 119 affine.for %j = 0 to 50 { 120 store 0 to %C[%i, %j] : memref<100x50xf32> 121 affine.for %k = 0 to %n { 122 %a_v = load %A[%i, %k] : memref<100x?xf32> 123 %b_v = load %B[%k, %j] : memref<?x50xf32> 124 %prod = mulf %a_v, %b_v : f32 125 %c_v = load %C[%i, %j] : memref<100x50xf32> 126 %sum = addf %c_v, %prod : f32 127 store %sum, %C[%i, %j] : memref<100x50xf32> 128 } 129 } 130 } 131 return %C : memref<100x50xf32> 132} 133``` 134 135## Notation 136 137MLIR has a simple and unambiguous grammar, allowing it to reliably round-trip 138through a textual form. This is important for development of the compiler - 139e.g. for understanding the state of code as it is being transformed and 140writing test cases. 141 142This document describes the grammar using 143[Extended Backus-Naur Form (EBNF)](https://en.wikipedia.org/wiki/Extended_Backus%E2%80%93Naur_form). 144 145This is the EBNF grammar used in this document, presented in yellow boxes. 146 147``` 148alternation ::= expr0 | expr1 | expr2 // Either expr0 or expr1 or expr2. 149sequence ::= expr0 expr1 expr2 // Sequence of expr0 expr1 expr2. 150repetition0 ::= expr* // 0 or more occurrences. 151repetition1 ::= expr+ // 1 or more occurrences. 152optionality ::= expr? // 0 or 1 occurrence. 153grouping ::= (expr) // Everything inside parens is grouped together. 154literal ::= `abcd` // Matches the literal `abcd`. 155``` 156 157Code examples are presented in blue boxes. 158 159```mlir 160// This is an example use of the grammar above: 161// This matches things like: ba, bana, boma, banana, banoma, bomana... 162example ::= `b` (`an` | `om`)* `a` 163``` 164 165### Common syntax 166 167The following core grammar productions are used in this document: 168 169``` 170// TODO: Clarify the split between lexing (tokens) and parsing (grammar). 171digit ::= [0-9] 172hex_digit ::= [0-9a-fA-F] 173letter ::= [a-zA-Z] 174id-punct ::= [$._-] 175 176integer-literal ::= decimal-literal | hexadecimal-literal 177decimal-literal ::= digit+ 178hexadecimal-literal ::= `0x` hex_digit+ 179float-literal ::= [-+]?[0-9]+[.][0-9]*([eE][-+]?[0-9]+)? 180string-literal ::= `"` [^"\n\f\v\r]* `"` TODO: define escaping rules 181``` 182 183Not listed here, but MLIR does support comments. They use standard BCPL syntax, 184starting with a `//` and going until the end of the line. 185 186### Identifiers and keywords 187 188Syntax: 189 190``` 191// Identifiers 192bare-id ::= (letter|[_]) (letter|digit|[_$.])* 193bare-id-list ::= bare-id (`,` bare-id)* 194value-id ::= `%` suffix-id 195suffix-id ::= (digit+ | ((letter|id-punct) (letter|id-punct|digit)*)) 196 197symbol-ref-id ::= `@` (suffix-id | string-literal) 198value-id-list ::= value-id (`,` value-id)* 199 200// Uses of value, e.g. in an operand list to an operation. 201value-use ::= value-id 202value-use-list ::= value-use (`,` value-use)* 203``` 204 205Identifiers name entities such as values, types and functions, and are 206chosen by the writer of MLIR code. Identifiers may be descriptive (e.g. 207`%batch_size`, `@matmul`), or may be non-descriptive when they are 208auto-generated (e.g. `%23`, `@func42`). Identifier names for values may be 209used in an MLIR text file but are not persisted as part of the IR - the printer 210will give them anonymous names like `%42`. 211 212MLIR guarantees identifiers never collide with keywords by prefixing identifiers 213with a sigil (e.g. `%`, `#`, `@`, `^`, `!`). In certain unambiguous contexts 214(e.g. affine expressions), identifiers are not prefixed, for brevity. New 215keywords may be added to future versions of MLIR without danger of collision 216with existing identifiers. 217 218Value identifiers are only [in scope](#value-scoping) for the (nested) 219region in which they are defined and cannot be accessed or referenced 220outside of that region. Argument identifiers in mapping functions are 221in scope for the mapping body. Particular operations may further limit 222which identifiers are in scope in their regions. For instance, the 223scope of values in a region with [SSA control flow 224semantics](#control-flow-and-ssacfg-regions) is constrained according 225to the standard definition of [SSA 226dominance](https://en.wikipedia.org/wiki/Dominator_\(graph_theory\)). Another 227example is the [IsolatedFromAbove trait](Traits.md#isolatedfromabove), 228which restricts directly accessing values defined in containing 229regions. 230 231Function identifiers and mapping identifiers are associated with 232[Symbols](SymbolsAndSymbolTables) and have scoping rules dependent on 233symbol attributes. 234 235## Dialects 236 237Dialects are the mechanism by which to engage with and extend the MLIR 238ecosystem. They allow for defining new [operations](#operations), as well as 239[attributes](#attributes) and [types](#type-system). Each dialect is given a 240unique `namespace` that is prefixed to each defined attribute/operation/type. 241For example, the [Affine dialect](Dialects/Affine.md) defines the namespace: 242`affine`. 243 244MLIR allows for multiple dialects, even those outside of the main tree, to 245co-exist together within one module. Dialects are produced and consumed by 246certain passes. MLIR provides a [framework](DialectConversion.md) to convert 247between, and within, different dialects. 248 249A few of the dialects supported by MLIR: 250 251* [Affine dialect](Dialects/Affine.md) 252* [GPU dialect](Dialects/GPU.md) 253* [LLVM dialect](Dialects/LLVM.md) 254* [SPIR-V dialect](Dialects/SPIR-V.md) 255* [Standard dialect](Dialects/Standard.md) 256* [Vector dialect](Dialects/Vector.md) 257 258### Target specific operations 259 260Dialects provide a modular way in which targets can expose target-specific 261operations directly through to MLIR. As an example, some targets go through 262LLVM. LLVM has a rich set of intrinsics for certain target-independent 263operations (e.g. addition with overflow check) as well as providing access to 264target-specific operations for the targets it supports (e.g. vector 265permutation operations). LLVM intrinsics in MLIR are represented via 266operations that start with an "llvm." name. 267 268Example: 269 270```mlir 271// LLVM: %x = call {i16, i1} @llvm.sadd.with.overflow.i16(i16 %a, i16 %b) 272%x:2 = "llvm.sadd.with.overflow.i16"(%a, %b) : (i16, i16) -> (i16, i1) 273``` 274 275These operations only work when targeting LLVM as a backend (e.g. for CPUs and 276GPUs), and are required to align with the LLVM definition of these intrinsics. 277 278## Operations 279 280Syntax: 281 282``` 283operation ::= op-result-list? (generic-operation | custom-operation) 284 trailing-location? 285generic-operation ::= string-literal `(` value-use-list? `)` successor-list? 286 (`(` region-list `)`)? dictionary-attribute? `:` function-type 287custom-operation ::= bare-id custom-operation-format 288op-result-list ::= op-result (`,` op-result)* `=` 289op-result ::= value-id (`:` integer-literal) 290successor-list ::= successor (`,` successor)* 291successor ::= caret-id (`:` bb-arg-list)? 292region-list ::= region (`,` region)* 293trailing-location ::= (`loc` `(` location `)`)? 294``` 295 296MLIR introduces a uniform concept called _operations_ to enable describing 297many different levels of abstractions and computations. Operations in MLIR are 298fully extensible (there is no fixed list of operations) and have 299application-specific semantics. For example, MLIR supports [target-independent 300operations](Dialects/Standard.md#memory-operations), [affine 301operations](Dialects/Affine.md), and [target-specific machine 302operations](#target-specific-operations). 303 304The internal representation of an operation is simple: an operation is 305identified by a unique string (e.g. `dim`, `tf.Conv2d`, `x86.repmovsb`, 306`ppc.eieio`, etc), can return zero or more results, take zero or more 307operands, has a dictionary of [attributes](#attributes), has zero or more 308successors, and zero or more enclosed [regions](#regions). The generic printing 309form includes all these elements literally, with a function type to indicate the 310types of the results and operands. 311 312Example: 313 314```mlir 315// An operation that produces two results. 316// The results of %result can be accessed via the <name> `#` <opNo> syntax. 317%result:2 = "foo_div"() : () -> (f32, i32) 318 319// Pretty form that defines a unique name for each result. 320%foo, %bar = "foo_div"() : () -> (f32, i32) 321 322// Invoke a TensorFlow function called tf.scramble with two inputs 323// and an attribute "fruit". 324%2 = "tf.scramble"(%result#0, %bar) {fruit = "banana"} : (f32, i32) -> f32 325``` 326 327In addition to the basic syntax above, dialects may register known operations. 328This allows those dialects to support _custom assembly form_ for parsing and 329printing operations. In the operation sets listed below, we show both forms. 330 331### Terminator Operations 332 333These are a special category of operations that *must* terminate a block, e.g. 334[branches](Dialects/Standard.md#terminator-operations). These operations may 335also have a list of successors ([blocks](#blocks) and their arguments). 336 337Example: 338 339```mlir 340// Branch to ^bb1 or ^bb2 depending on the condition %cond. 341// Pass value %v to ^bb2, but not to ^bb1. 342"cond_br"(%cond)[^bb1, ^bb2(%v : index)] : (i1) -> () 343``` 344 345### Module 346 347``` 348module ::= `module` symbol-ref-id? (`attributes` dictionary-attribute)? region 349``` 350 351An MLIR Module represents a top-level container operation. It contains a single 352[SSACFG region](#control-flow-and-ssacfg-regions) containing a single block 353which can contain any operations. Operations within this region cannot 354implicitly capture values defined outside the module, i.e. Modules are 355[IsolatedFromAbove](Traits.md#isolatedfromabove). Modules have an optional 356[symbol name](SymbolsAndSymbolTables.md) which can be used to refer to them in 357operations. 358 359### Functions 360 361An MLIR Function is an operation with a name containing a single [SSACFG 362region](#control-flow-and-ssacfg-regions). Operations within this region 363cannot implicitly capture values defined outside of the function, 364i.e. Functions are [IsolatedFromAbove](Traits.md#isolatedfromabove). All 365external references must use function arguments or attributes that establish a 366symbolic connection (e.g. symbols referenced by name via a string attribute 367like [SymbolRefAttr](#symbol-reference-attribute)): 368 369``` 370function ::= `func` function-signature function-attributes? function-body? 371 372function-signature ::= symbol-ref-id `(` argument-list `)` 373 (`->` function-result-list)? 374 375argument-list ::= (named-argument (`,` named-argument)*) | /*empty*/ 376argument-list ::= (type dictionary-attribute? (`,` type dictionary-attribute?)*) 377 | /*empty*/ 378named-argument ::= value-id `:` type dictionary-attribute? 379 380function-result-list ::= function-result-list-parens 381 | non-function-type 382function-result-list-parens ::= `(` `)` 383 | `(` function-result-list-no-parens `)` 384function-result-list-no-parens ::= function-result (`,` function-result)* 385function-result ::= type dictionary-attribute? 386 387function-attributes ::= `attributes` dictionary-attribute 388function-body ::= region 389``` 390 391An external function declaration (used when referring to a function declared 392in some other module) has no body. While the MLIR textual form provides a nice 393inline syntax for function arguments, they are internally represented as 394"block arguments" to the first block in the region. 395 396Only dialect attribute names may be specified in the attribute dictionaries 397for function arguments, results, or the function itself. 398 399Examples: 400 401```mlir 402// External function definitions. 403func @abort() 404func @scribble(i32, i64, memref<? x 128 x f32, #layout_map0>) -> f64 405 406// A function that returns its argument twice: 407func @count(%x: i64) -> (i64, i64) 408 attributes {fruit: "banana"} { 409 return %x, %x: i64, i64 410} 411 412// A function with an argument attribute 413func @example_fn_arg(%x: i32 {swift.self = unit}) 414 415// A function with a result attribute 416func @example_fn_result() -> (f64 {dialectName.attrName = 0 : i64}) 417 418// A function with an attribute 419func @example_fn_attr() attributes {dialectName.attrName = false} 420``` 421 422## Blocks 423 424Syntax: 425 426``` 427block ::= block-label operation+ 428block-label ::= block-id block-arg-list? `:` 429block-id ::= caret-id 430caret-id ::= `^` suffix-id 431value-id-and-type ::= value-id `:` type 432 433// Non-empty list of names and types. 434value-id-and-type-list ::= value-id-and-type (`,` value-id-and-type)* 435 436block-arg-list ::= `(` value-id-and-type-list? `)` 437``` 438 439A *Block* is an ordered list of operations, concluding with a single 440[terminator operation](#terminator-operations). In [SSACFG 441regions](#control-flow-and-ssacfg-regions), each block represents a compiler 442[basic block](https://en.wikipedia.org/wiki/Basic_block) where instructions 443inside the block are executed in order and terminator operations implement 444control flow branches between basic blocks. 445 446Blocks in MLIR take a list of block arguments, notated in a function-like 447way. Block arguments are bound to values specified by the semantics of 448individual operations. Block arguments of the entry block of a region are also 449arguments to the region and the values bound to these arguments are determined 450by the semantics of the containing operation. Block arguments of other blocks 451are determined by the semantics of terminator operations, e.g. Branches, which 452have the block as a successor. In regions with [control 453flow](#control-flow-and-ssacfg-regions), MLIR leverages this structure to 454implicitly represent the passage of control-flow dependent values without the 455complex nuances of PHI nodes in traditional SSA representations. Note that 456values which are not control-flow dependent can be referenced directly and do 457not need to be passed through block arguments. 458 459Here is a simple example function showing branches, returns, and block 460arguments: 461 462```mlir 463func @simple(i64, i1) -> i64 { 464^bb0(%a: i64, %cond: i1): // Code dominated by ^bb0 may refer to %a 465 cond_br %cond, ^bb1, ^bb2 466 467^bb1: 468 br ^bb3(%a: i64) // Branch passes %a as the argument 469 470^bb2: 471 %b = addi %a, %a : i64 472 br ^bb3(%b: i64) // Branch passes %b as the argument 473 474// ^bb3 receives an argument, named %c, from predecessors 475// and passes it on to bb4 along with %a. %a is referenced 476// directly from its defining operation and is not passed through 477// an argument of ^bb3. 478^bb3(%c: i64): 479 br ^bb4(%c, %a : i64, i64) 480 481^bb4(%d : i64, %e : i64): 482 %0 = addi %d, %e : i64 483 return %0 : i64 // Return is also a terminator. 484} 485``` 486 487**Context:** The "block argument" representation eliminates a number 488of special cases from the IR compared to traditional "PHI nodes are 489operations" SSA IRs (like LLVM). For example, the [parallel copy 490semantics](http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.524.5461&rep=rep1&type=pdf) 491of SSA is immediately apparent, and function arguments are no longer a 492special case: they become arguments to the entry block [[more 493rationale](Rationale/Rationale.md#block-arguments-vs-phi-nodes)]. Blocks 494are also a fundamental concept that cannot be represented by 495operations because values defined in an operation cannot be accessed 496outside the operation. 497 498## Regions 499 500### Definition 501 502A region is an ordered list of MLIR [Blocks](#blocks). The semantics within a 503region is not imposed by the IR. Instead, the containing operation defines the 504semantics of the regions it contains. MLIR currently defines two kinds of 505regions: [SSACFG regions](#control-flow-and-ssacfg-regions), which describe 506control flow between blocks, and [Graph regions](#graph-regions), which do not 507require control flow between block. The kinds of regions within an operation 508are described using the 509[RegionKindInterface](Interfaces.md#regionkindinterfaces). 510 511Regions do not have a name or an address, only the blocks contained in a 512region do. Regions must be contained within operations and have no type or 513attributes. The first block in the region is a special block called the 'entry 514block'. The arguments to the entry block are also the arguments of the region 515itself. The entry block cannot be listed as a successor of any other 516block. The syntax for a region is as follows: 517 518``` 519region ::= `{` block* `}` 520``` 521 522A function body is an example of a region: it consists of a CFG of blocks and 523has additional semantic restrictions that other types of regions may not have. 524For example, in a function body, block terminators must either branch to a 525different block, or return from a function where the types of the `return` 526arguments must match the result types of the function signature. Similarly, 527the function arguments must match the types and count of the region arguments. 528In general, operations with regions can define these correspondances 529arbitrarily. 530 531### Value Scoping 532 533Regions provide hierarchical encapsulation of programs: it is impossible to 534reference, i.e. branch to, a block which is not in the same region as the 535source of the reference, i.e. a terminator operation. Similarly, regions 536provides a natural scoping for value visibility: values defined in a region 537don't escape to the enclosing region, if any. By default, operations inside a 538region can reference values defined outside of the region whenever it would 539have been legal for operands of the enclosing operation to reference those 540values, but this can be restricted using traits, such as 541[OpTrait::IsolatedFromAbove](Traits.md#isolatedfromabove), or a custom 542verifier. 543 544Example: 545 546```mlir 547 "any_op"(%a) ({ // if %a is in-scope in the containing region... 548 // then %a is in-scope here too. 549 %new_value = "another_op"(%a) : (i64) -> (i64) 550 }) : (i64) -> (i64) 551``` 552 553MLIR defines a generalized 'hierarchical dominance' concept that operates 554across hierarchy and defines whether a value is 'in scope' and can be used by 555a particular operation. Whether a value can be used by another operation in 556the same region is defined by the kind of region. A value defined in a region 557can be used by an operation which has a parent in the same region, if and only 558if the parent could use the value. A value defined by an argument to a region 559can always be used by any operation deeply contained in the region. A value 560defined in a region can never be used outside of the region. 561 562### Control Flow and SSACFG Regions 563 564In MLIR, control flow semantics of a region is indicated by 565[RegionKind::SSACFG](Interfaces.md#regionkindinterfaces). Informally, these 566regions support semantics where operations in a region 'execute 567sequentially'. Before an operation executes, its operands have well-defined 568values. After an operation executes, the operands have the same values and 569results also have well-defined values. After an operation executes, the next 570operation in the block executes until the operation is the terminator operation 571at the end of a block, in which case some other operation will execute. The 572determination of the next instruction to execute is the 'passing of control 573flow'. 574 575In general, when control flow is passed to an operation, MLIR does not 576restrict when control flow enters or exits the regions contained in that 577operation. However, when control flow enters a region, it always begins in the 578first block of the region, called the *entry* block. Terminator operations 579ending each block represent control flow by explicitly specifying the 580successor blocks of the block. Control flow can only pass to one of the 581specified successor blocks as in a `branch` operation, or back to the 582containing operation as in a `return` operation. Terminator operations without 583successors can only pass control back to the containing operation. Within 584these restrictions, the particular semantics of terminator operations is 585determined by the specific dialect operations involved. Blocks (other than the 586entry block) that are not listed as a successor of a terminator operation are 587defined to be unreachable and can be removed without affecting the semantics 588of the containing operation. 589 590Although control flow always enters a region through the entry block, control 591flow may exit a region through any block with an appropriate terminator. The 592standard dialect leverages this capability to define operations with 593Single-Entry-Multiple-Exit (SEME) regions, possibly flowing through different 594blocks in the region and exiting through any block with a `return` 595operation. This behavior is similar to that of a function body in most 596programming languages. In addition, control flow may also not reach the end of 597a block or region, for example if a function call does not return. 598 599Example: 600 601```mlir 602func @accelerator_compute(i64, i1) -> i64 { // An SSACFG region 603^bb0(%a: i64, %cond: i1): // Code dominated by ^bb0 may refer to %a 604 cond_br %cond, ^bb1, ^bb2 605 606^bb1: 607 // This def for %value does not dominate ^bb2 608 %value = "op.convert"(%a) : (i64) -> i64 609 br ^bb3(%a: i64) // Branch passes %a as the argument 610 611^bb2: 612 accelerator.launch() { // An SSACFG region 613 ^bb0: 614 // Region of code nested under "accelerator.launch", it can reference %a but 615 // not %value. 616 %new_value = "accelerator.do_something"(%a) : (i64) -> () 617 } 618 // %new_value cannot be referenced outside of the region 619 620^bb3: 621 ... 622} 623``` 624 625#### Operations with Multiple Regions 626 627An operation containing multiple regions also completely determines the 628semantics of those regions. In particular, when control flow is passed to an 629operation, it may transfer control flow to any contained region. When control 630flow exits a region and is returned to the containing operation, the 631containing operation may pass control flow to any region in the same 632operation. An operation may also pass control flow to multiple contained 633regions concurrently. An operation may also pass control flow into regions 634that were specified in other operations, in particular those that defined the 635values or symbols the given operation uses as in a call operation. This 636passage of control is generally independent of passage of control flow through 637the basic blocks of the containing region. 638 639#### Closure 640 641Regions allow defining an operation that creates a closure, for example by 642“boxing” the body of the region into a value they produce. It remains up to the 643operation to define its semantics. Note that if an operation triggers 644asynchronous execution of the region, it is under the responsibility of the 645operation caller to wait for the region to be executed guaranteeing that any 646directly used values remain live. 647 648### Graph Regions 649 650In MLIR, graph-like semantics in a region is indicated by 651[RegionKind::Graph](Interfaces.md#regionkindinterfaces). Graph regions are 652appropriate for concurrent semantics without control flow, or for modeling 653generic directed graph data structures. Graph regions are appropriate for 654representing cyclic relationships between coupled values where there is no 655fundamental order to the relationships. For instance, operations in a graph 656region may represent independent threads of control with values representing 657streams of data. As usual in MLIR, the particular semantics of a region is 658completely determined by its containing operation. Graph regions may only 659contain a single basic block (the entry block). 660 661**Rationale:** Currently graph regions are arbitrarily limited to a single 662basic block, although there is no particular semantic reason for this 663limitation. This limitation has been added to make it easier to stabilize the 664pass infrastructure and commonly used passes for processing graph regions to 665properly handle feedback loops. Multi-block regions may be allowed in the 666future if use cases that require it arise. 667 668In graph regions, MLIR operations naturally represent nodes, while each MLIR 669value represents a multi-edge connecting a single source node and multiple 670destination nodes. All values defined in the region as results of operations 671are in scope within the region and can be accessed by any other operation in 672the region. In graph regions, the order of operations within a block and the 673order of blocks in a region is not semantically meaningful and non-terminator 674operations may be freely reordered, for instance, by canonicalization. Other 675kinds of graphs, such as graphs with multiple source nodes and multiple 676destination nodes, can also be represented by representing graph edges as MLIR 677operations. 678 679Note that cycles can occur within a single block in a graph region, or between 680basic blocks. 681 682```mlir 683"test.graph_region"() ({ // A Graph region 684 %1 = "op1"(%1, %3) : (i32, i32) -> (i32) // OK: %1, %3 allowed here 685 %2 = "test.ssacfg_region"() ({ 686 %5 = "op2"(%1, %2, %3, %4) : (i32, i32, i32, i32) -> (i32) // OK: %1, %2, %3, %4 all defined in the containing region 687 }) : () -> (i32) 688 %3 = "op2"(%1, %4) : (i32, i32) -> (i32) // OK: %4 allowed here 689 %4 = "op3"(%1) : (i32) -> (i32) 690}) : () -> () 691``` 692 693### Arguments and Results 694 695The arguments of the first block of a region are treated as arguments of the 696region. The source of these arguments is defined by the semantics of the parent 697operation. They may correspond to some of the values the operation itself uses. 698 699Regions produce a (possibly empty) list of values. The operation semantics 700defines the relation between the region results and the operation results. 701 702## Type System 703 704Each value in MLIR has a type defined by the type system below. There are a 705number of primitive types (like integers) and also aggregate types for tensors 706and memory buffers. MLIR [builtin types](#builtin-types) do not include 707structures, arrays, or dictionaries. 708 709MLIR has an open type system (i.e. there is no fixed list of types), and types 710may have application-specific semantics. For example, MLIR supports a set of 711[dialect types](#dialect-types). 712 713``` 714type ::= type-alias | dialect-type | builtin-type 715 716type-list-no-parens ::= type (`,` type)* 717type-list-parens ::= `(` `)` 718 | `(` type-list-no-parens `)` 719 720// This is a common way to refer to a value with a specified type. 721ssa-use-and-type ::= ssa-use `:` type 722 723// Non-empty list of names and types. 724ssa-use-and-type-list ::= ssa-use-and-type (`,` ssa-use-and-type)* 725``` 726 727### Type Aliases 728 729``` 730type-alias-def ::= '!' alias-name '=' 'type' type 731type-alias ::= '!' alias-name 732``` 733 734MLIR supports defining named aliases for types. A type alias is an identifier 735that can be used in the place of the type that it defines. These aliases *must* 736be defined before their uses. Alias names may not contain a '.', since those 737names are reserved for [dialect types](#dialect-types). 738 739Example: 740 741```mlir 742!avx_m128 = type vector<4 x f32> 743 744// Using the original type. 745"foo"(%x) : vector<4 x f32> -> () 746 747// Using the type alias. 748"foo"(%x) : !avx_m128 -> () 749``` 750 751### Dialect Types 752 753Similarly to operations, dialects may define custom extensions to the type 754system. 755 756``` 757dialect-namespace ::= bare-id 758 759opaque-dialect-item ::= dialect-namespace '<' string-literal '>' 760 761pretty-dialect-item ::= dialect-namespace '.' pretty-dialect-item-lead-ident 762 pretty-dialect-item-body? 763 764pretty-dialect-item-lead-ident ::= '[A-Za-z][A-Za-z0-9._]*' 765pretty-dialect-item-body ::= '<' pretty-dialect-item-contents+ '>' 766pretty-dialect-item-contents ::= pretty-dialect-item-body 767 | '(' pretty-dialect-item-contents+ ')' 768 | '[' pretty-dialect-item-contents+ ']' 769 | '{' pretty-dialect-item-contents+ '}' 770 | '[^[<({>\])}\0]+' 771 772dialect-type ::= '!' opaque-dialect-item 773dialect-type ::= '!' pretty-dialect-item 774``` 775 776Dialect types can be specified in a verbose form, e.g. like this: 777 778```mlir 779// LLVM type that wraps around llvm IR types. 780!llvm<"i32*"> 781 782// Tensor flow string type. 783!tf.string 784 785// Complex type 786!foo<"something<abcd>"> 787 788// Even more complex type 789!foo<"something<a%%123^^^>>>"> 790``` 791 792Dialect types that are simple enough can use the pretty format, which is a 793lighter weight syntax that is equivalent to the above forms: 794 795```mlir 796// Tensor flow string type. 797!tf.string 798 799// Complex type 800!foo.something<abcd> 801``` 802 803Sufficiently complex dialect types are required to use the verbose form for 804generality. For example, the more complex type shown above wouldn't be valid in 805the lighter syntax: `!foo.something<a%%123^^^>>>` because it contains characters 806that are not allowed in the lighter syntax, as well as unbalanced `<>` 807characters. 808 809See [here](Tutorials/DefiningAttributesAndTypes.md) to learn how to define dialect types. 810 811### Builtin Types 812 813Builtin types are a core set of [dialect types](#dialect-types) that are defined 814in a builtin dialect and thus available to all users of MLIR. 815 816``` 817builtin-type ::= complex-type 818 | float-type 819 | function-type 820 | index-type 821 | integer-type 822 | memref-type 823 | none-type 824 | tensor-type 825 | tuple-type 826 | vector-type 827``` 828 829#### Complex Type 830 831Syntax: 832 833``` 834complex-type ::= `complex` `<` type `>` 835``` 836 837The value of `complex` type represents a complex number with a parameterized 838element type, which is composed of a real and imaginary value of that element 839type. The element must be a floating point or integer scalar type. 840 841Examples: 842 843```mlir 844complex<f32> 845complex<i32> 846``` 847 848#### Floating Point Types 849 850Syntax: 851 852``` 853// Floating point. 854float-type ::= `f16` | `bf16` | `f32` | `f64` | `f80` | `f128` 855``` 856 857MLIR supports float types of certain widths that are widely used as indicated 858above. 859 860#### Function Type 861 862Syntax: 863 864``` 865// MLIR functions can return multiple values. 866function-result-type ::= type-list-parens 867 | non-function-type 868 869function-type ::= type-list-parens `->` function-result-type 870``` 871 872MLIR supports first-class functions: for example, the 873[`constant` operation](Dialects/Standard.md#stdconstant-constantop) produces the 874address of a function as a value. This value may be passed to and 875returned from functions, merged across control flow boundaries with 876[block arguments](#blocks), and called with the 877[`call_indirect` operation](Dialects/Standard.md#call-indirect-operation). 878 879Function types are also used to indicate the arguments and results of 880[operations](#operations). 881 882#### Index Type 883 884Syntax: 885 886``` 887// Target word-sized integer. 888index-type ::= `index` 889``` 890 891The `index` type is a signless integer whose size is equal to the natural 892machine word of the target 893([rationale](Rationale/Rationale.md#integer-signedness-semantics)) and is used 894by the affine constructs in MLIR. Unlike fixed-size integers, it cannot be used 895as an element of vector 896([rationale](Rationale/Rationale.md#index-type-disallowed-in-vector-types)). 897 898**Rationale:** integers of platform-specific bit widths are practical to express 899sizes, dimensionalities and subscripts. 900 901#### Integer Type 902 903Syntax: 904 905``` 906// Sized integers like i1, i4, i8, i16, i32. 907signed-integer-type ::= `si` [1-9][0-9]* 908unsigned-integer-type ::= `ui` [1-9][0-9]* 909signless-integer-type ::= `i` [1-9][0-9]* 910integer-type ::= signed-integer-type | 911 unsigned-integer-type | 912 signless-integer-type 913``` 914 915MLIR supports arbitrary precision integer types. Integer types have a designated 916width and may have signedness semantics. 917 918**Rationale:** low precision integers (like `i2`, `i4` etc) are useful for 919low-precision inference chips, and arbitrary precision integers are useful for 920hardware synthesis (where a 13 bit multiplier is a lot cheaper/smaller than a 16 921bit one). 922 923TODO: Need to decide on a representation for quantized integers 924([initial thoughts](Rationale/Rationale.md#quantized-integer-operations)). 925 926#### Memref Type 927 928Syntax: 929 930``` 931memref-type ::= ranked-memref-type | unranked-memref-type 932 933ranked-memref-type ::= `memref` `<` dimension-list-ranked type 934 (`,` layout-specification)? (`,` memory-space)? `>` 935 936unranked-memref-type ::= `memref` `<*x` type (`,` memory-space)? `>` 937 938stride-list ::= `[` (dimension (`,` dimension)*)? `]` 939strided-layout ::= `offset:` dimension `,` `strides: ` stride-list 940semi-affine-map-composition ::= (semi-affine-map `,` )* semi-affine-map 941layout-specification ::= semi-affine-map-composition | strided-layout 942memory-space ::= integer-literal /* | TODO: address-space-id */ 943``` 944 945A `memref` type is a reference to a region of memory (similar to a buffer 946pointer, but more powerful). The buffer pointed to by a memref can be allocated, 947aliased and deallocated. A memref can be used to read and write data from/to the 948memory region which it references. Memref types use the same shape specifier as 949tensor types. Note that `memref<f32>`, `memref<0 x f32>`, `memref<1 x 0 x f32>`, 950and `memref<0 x 1 x f32>` are all different types. 951 952A `memref` is allowed to have an unknown rank (e.g. `memref<*xf32>`). The 953purpose of unranked memrefs is to allow external library functions to receive 954memref arguments of any rank without versioning the functions based on the rank. 955Other uses of this type are disallowed or will have undefined behavior. 956 957##### Codegen of Unranked Memref 958 959Using unranked memref in codegen besides the case mentioned above is highly 960discouraged. Codegen is concerned with generating loop nests and specialized 961instructions for high-performance, unranked memref is concerned with hiding the 962rank and thus, the number of enclosing loops required to iterate over the data. 963However, if there is a need to code-gen unranked memref, one possible path is to 964cast into a static ranked type based on the dynamic rank. Another possible path 965is to emit a single while loop conditioned on a linear index and perform 966delinearization of the linear index to a dynamic array containing the (unranked) 967indices. While this is possible, it is expected to not be a good idea to perform 968this during codegen as the cost of the translations is expected to be 969prohibitive and optimizations at this level are not expected to be worthwhile. 970If expressiveness is the main concern, irrespective of performance, passing 971unranked memrefs to an external C++ library and implementing rank-agnostic logic 972there is expected to be significantly simpler. 973 974Unranked memrefs may provide expressiveness gains in the future and help bridge 975the gap with unranked tensors. Unranked memrefs will not be expected to be 976exposed to codegen but one may query the rank of an unranked memref (a special 977op will be needed for this purpose) and perform a switch and cast to a ranked 978memref as a prerequisite to codegen. 979 980Example: 981 982```mlir 983// With static ranks, we need a function for each possible argument type 984%A = alloc() : memref<16x32xf32> 985%B = alloc() : memref<16x32x64xf32> 986call @helper_2D(%A) : (memref<16x32xf32>)->() 987call @helper_3D(%B) : (memref<16x32x64xf32>)->() 988 989// With unknown rank, the functions can be unified under one unranked type 990%A = alloc() : memref<16x32xf32> 991%B = alloc() : memref<16x32x64xf32> 992// Remove rank info 993%A_u = memref_cast %A : memref<16x32xf32> -> memref<*xf32> 994%B_u = memref_cast %B : memref<16x32x64xf32> -> memref<*xf32> 995// call same function with dynamic ranks 996call @helper(%A_u) : (memref<*xf32>)->() 997call @helper(%B_u) : (memref<*xf32>)->() 998``` 999 1000The core syntax and representation of a layout specification is a 1001[semi-affine map](Dialects/Affine.md#semi-affine-maps). Additionally, syntactic 1002sugar is supported to make certain layout specifications more intuitive to read. 1003For the moment, a `memref` supports parsing a strided form which is converted to 1004a semi-affine map automatically. 1005 1006The memory space of a memref is specified by a target-specific integer index. If 1007no memory space is specified, then the default memory space (0) is used. The 1008default space is target specific but always at index 0. 1009 1010TODO: MLIR will eventually have target-dialects which allow symbolic use of 1011memory hierarchy names (e.g. L3, L2, L1, ...) but we have not spec'd the details 1012of that mechanism yet. Until then, this document pretends that it is valid to 1013refer to these memories by `bare-id`. 1014 1015The notionally dynamic value of a memref value includes the address of the 1016buffer allocated, as well as the symbols referred to by the shape, layout map, 1017and index maps. 1018 1019Examples of memref static type 1020 1021```mlir 1022// Identity index/layout map 1023#identity = affine_map<(d0, d1) -> (d0, d1)> 1024 1025// Column major layout. 1026#col_major = affine_map<(d0, d1, d2) -> (d2, d1, d0)> 1027 1028// A 2-d tiled layout with tiles of size 128 x 256. 1029#tiled_2d_128x256 = affine_map<(d0, d1) -> (d0 div 128, d1 div 256, d0 mod 128, d1 mod 256)> 1030 1031// A tiled data layout with non-constant tile sizes. 1032#tiled_dynamic = affine_map<(d0, d1)[s0, s1] -> (d0 floordiv s0, d1 floordiv s1, 1033 d0 mod s0, d1 mod s1)> 1034 1035// A layout that yields a padding on two at either end of the minor dimension. 1036#padded = affine_map<(d0, d1) -> (d0, (d1 + 2) floordiv 2, (d1 + 2) mod 2)> 1037 1038 1039// The dimension list "16x32" defines the following 2D index space: 1040// 1041// { (i, j) : 0 <= i < 16, 0 <= j < 32 } 1042// 1043memref<16x32xf32, #identity> 1044 1045// The dimension list "16x4x?" defines the following 3D index space: 1046// 1047// { (i, j, k) : 0 <= i < 16, 0 <= j < 4, 0 <= k < N } 1048// 1049// where N is a symbol which represents the runtime value of the size of 1050// the third dimension. 1051// 1052// %N here binds to the size of the third dimension. 1053%A = alloc(%N) : memref<16x4x?xf32, #col_major> 1054 1055// A 2-d dynamic shaped memref that also has a dynamically sized tiled layout. 1056// The memref index space is of size %M x %N, while %B1 and %B2 bind to the 1057// symbols s0, s1 respectively of the layout map #tiled_dynamic. Data tiles of 1058// size %B1 x %B2 in the logical space will be stored contiguously in memory. 1059// The allocation size will be (%M ceildiv %B1) * %B1 * (%N ceildiv %B2) * %B2 1060// f32 elements. 1061%T = alloc(%M, %N) [%B1, %B2] : memref<?x?xf32, #tiled_dynamic> 1062 1063// A memref that has a two-element padding at either end. The allocation size 1064// will fit 16 * 64 float elements of data. 1065%P = alloc() : memref<16x64xf32, #padded> 1066 1067// Affine map with symbol 's0' used as offset for the first dimension. 1068#imapS = affine_map<(d0, d1) [s0] -> (d0 + s0, d1)> 1069// Allocate memref and bind the following symbols: 1070// '%n' is bound to the dynamic second dimension of the memref type. 1071// '%o' is bound to the symbol 's0' in the affine map of the memref type. 1072%n = ... 1073%o = ... 1074%A = alloc (%n)[%o] : <16x?xf32, #imapS> 1075``` 1076 1077##### Index Space 1078 1079A memref dimension list defines an index space within which the memref can be 1080indexed to access data. 1081 1082##### Index 1083 1084Data is accessed through a memref type using a multidimensional index into the 1085multidimensional index space defined by the memref's dimension list. 1086 1087Examples 1088 1089```mlir 1090// Allocates a memref with 2D index space: 1091// { (i, j) : 0 <= i < 16, 0 <= j < 32 } 1092%A = alloc() : memref<16x32xf32, #imapA> 1093 1094// Loads data from memref '%A' using a 2D index: (%i, %j) 1095%v = load %A[%i, %j] : memref<16x32xf32, #imapA> 1096``` 1097 1098##### Index Map 1099 1100An index map is a one-to-one 1101[semi-affine map](Dialects/Affine.md#semi-affine-maps) that transforms a 1102multidimensional index from one index space to another. For example, the 1103following figure shows an index map which maps a 2-dimensional index from a 2x2 1104index space to a 3x3 index space, using symbols `S0` and `S1` as offsets. 1105 1106 1107 1108The number of domain dimensions and range dimensions of an index map can be 1109different, but must match the number of dimensions of the input and output index 1110spaces on which the map operates. The index space is always non-negative and 1111integral. In addition, an index map must specify the size of each of its range 1112dimensions onto which it maps. Index map symbols must be listed in order with 1113symbols for dynamic dimension sizes first, followed by other required symbols. 1114 1115##### Layout Map 1116 1117A layout map is a [semi-affine map](Dialects/Affine.md#semi-affine-maps) which 1118encodes logical to physical index space mapping, by mapping input dimensions to 1119their ordering from most-major (slowest varying) to most-minor (fastest 1120varying). Therefore, an identity layout map corresponds to a row-major layout. 1121Identity layout maps do not contribute to the MemRef type identification and are 1122discarded on construction. That is, a type with an explicit identity map is 1123`memref<?x?xf32, (i,j)->(i,j)>` is strictly the same as the one without layout 1124maps, `memref<?x?xf32>`. 1125 1126Layout map examples: 1127 1128```mlir 1129// MxN matrix stored in row major layout in memory: 1130#layout_map_row_major = (i, j) -> (i, j) 1131 1132// MxN matrix stored in column major layout in memory: 1133#layout_map_col_major = (i, j) -> (j, i) 1134 1135// MxN matrix stored in a 2-d blocked/tiled layout with 64x64 tiles. 1136#layout_tiled = (i, j) -> (i floordiv 64, j floordiv 64, i mod 64, j mod 64) 1137``` 1138 1139##### Affine Map Composition 1140 1141A memref specifies a semi-affine map composition as part of its type. A 1142semi-affine map composition is a composition of semi-affine maps beginning with 1143zero or more index maps, and ending with a layout map. The composition must be 1144conformant: the number of dimensions of the range of one map, must match the 1145number of dimensions of the domain of the next map in the composition. 1146 1147The semi-affine map composition specified in the memref type, maps from accesses 1148used to index the memref in load/store operations to other index spaces (i.e. 1149logical to physical index mapping). Each of the 1150[semi-affine maps](Dialects/Affine.md) and thus its composition is required to 1151be one-to-one. 1152 1153The semi-affine map composition can be used in dependence analysis, memory 1154access pattern analysis, and for performance optimizations like vectorization, 1155copy elision and in-place updates. If an affine map composition is not specified 1156for the memref, the identity affine map is assumed. 1157 1158##### Strided MemRef 1159 1160A memref may specify strides as part of its type. A stride specification is a 1161list of integer values that are either static or `?` (dynamic case). Strides 1162encode the distance, in number of elements, in (linear) memory between 1163successive entries along a particular dimension. A stride specification is 1164syntactic sugar for an equivalent strided memref representation using 1165semi-affine maps. For example, `memref<42x16xf32, offset: 33, strides: [1, 64]>` 1166specifies a non-contiguous memory region of `42` by `16` `f32` elements such 1167that: 1168 11691. the minimal size of the enclosing memory region must be `33 + 42 * 1 + 16 * 1170 64 = 1066` elements; 11712. the address calculation for accessing element `(i, j)` computes `33 + i + 1172 64 * j` 11733. the distance between two consecutive elements along the inner dimension is 1174 `1` element and the distance between two consecutive elements along the 1175 outer dimension is `64` elements. 1176 1177This corresponds to a column major view of the memory region and is internally 1178represented as the type `memref<42x16xf32, (i, j) -> (33 + i + 64 * j)>`. 1179 1180The specification of strides must not alias: given an n-D strided memref, 1181indices `(i1, ..., in)` and `(j1, ..., jn)` may not refer to the same memory 1182address unless `i1 == j1, ..., in == jn`. 1183 1184Strided memrefs represent a view abstraction over preallocated data. They are 1185constructed with special ops, yet to be introduced. Strided memrefs are a 1186special subclass of memrefs with generic semi-affine map and correspond to a 1187normalized memref descriptor when lowering to LLVM. 1188 1189#### None Type 1190 1191Syntax: 1192 1193``` 1194none-type ::= `none` 1195``` 1196 1197The `none` type is a unit type, i.e. a type with exactly one possible value, 1198where its value does not have a defined dynamic representation. 1199 1200#### Tensor Type 1201 1202Syntax: 1203 1204``` 1205tensor-type ::= `tensor` `<` dimension-list type `>` 1206 1207dimension-list ::= dimension-list-ranked | (`*` `x`) 1208dimension-list-ranked ::= (dimension `x`)* 1209dimension ::= `?` | decimal-literal 1210``` 1211 1212Values with tensor type represents aggregate N-dimensional data values, and 1213have a known element type. It may have an unknown rank (indicated by `*`) or may 1214have a fixed rank with a list of dimensions. Each dimension may be a static 1215non-negative decimal constant or be dynamically determined (indicated by `?`). 1216 1217The runtime representation of the MLIR tensor type is intentionally abstracted - 1218you cannot control layout or get a pointer to the data. For low level buffer 1219access, MLIR has a [`memref` type](#memref-type). This abstracted runtime 1220representation holds both the tensor data values as well as information about 1221the (potentially dynamic) shape of the tensor. The 1222[`dim` operation](Dialects/Standard.md#dim-operation) returns the size of a 1223dimension from a value of tensor type. 1224 1225Note: hexadecimal integer literals are not allowed in tensor type declarations 1226to avoid confusion between `0xf32` and `0 x f32`. Zero sizes are allowed in 1227tensors and treated as other sizes, e.g., `tensor<0 x 1 x i32>` and `tensor<1 x 12280 x i32>` are different types. Since zero sizes are not allowed in some other 1229types, such tensors should be optimized away before lowering tensors to vectors. 1230 1231Examples: 1232 1233```mlir 1234// Tensor with unknown rank. 1235tensor<* x f32> 1236 1237// Known rank but unknown dimensions. 1238tensor<? x ? x ? x ? x f32> 1239 1240// Partially known dimensions. 1241tensor<? x ? x 13 x ? x f32> 1242 1243// Full static shape. 1244tensor<17 x 4 x 13 x 4 x f32> 1245 1246// Tensor with rank zero. Represents a scalar. 1247tensor<f32> 1248 1249// Zero-element dimensions are allowed. 1250tensor<0 x 42 x f32> 1251 1252// Zero-element tensor of f32 type (hexadecimal literals not allowed here). 1253tensor<0xf32> 1254``` 1255 1256#### Tuple Type 1257 1258Syntax: 1259 1260``` 1261tuple-type ::= `tuple` `<` (type ( `,` type)*)? `>` 1262``` 1263 1264The value of `tuple` type represents a fixed-size collection of elements, where 1265each element may be of a different type. 1266 1267**Rationale:** Though this type is first class in the type system, MLIR provides 1268no standard operations for operating on `tuple` types 1269([rationale](Rationale/Rationale.md#tuple-types)). 1270 1271Examples: 1272 1273```mlir 1274// Empty tuple. 1275tuple<> 1276 1277// Single element 1278tuple<f32> 1279 1280// Many elements. 1281tuple<i32, f32, tensor<i1>, i5> 1282``` 1283 1284#### Vector Type 1285 1286Syntax: 1287 1288``` 1289vector-type ::= `vector` `<` static-dimension-list vector-element-type `>` 1290vector-element-type ::= float-type | integer-type 1291 1292static-dimension-list ::= (decimal-literal `x`)+ 1293``` 1294 1295The vector type represents a SIMD style vector, used by target-specific 1296operation sets like AVX. While the most common use is for 1D vectors (e.g. 1297vector<16 x f32>) we also support multidimensional registers on targets that 1298support them (like TPUs). 1299 1300Vector shapes must be positive decimal integers. 1301 1302Note: hexadecimal integer literals are not allowed in vector type declarations, 1303`vector<0x42xi32>` is invalid because it is interpreted as a 2D vector with 1304shape `(0, 42)` and zero shapes are not allowed. 1305 1306## Attributes 1307 1308Syntax: 1309 1310``` 1311attribute-entry ::= (bare-id | string-literal) `=` attribute-value 1312attribute-value ::= attribute-alias | dialect-attribute | builtin-attribute 1313``` 1314 1315Attributes are the mechanism for specifying constant data on operations in 1316places where a variable is never allowed - e.g. the comparison predicate of a 1317[`cmpi` operation](Dialects/Standard.md#stdcmpi-cmpiop). Each operation has an 1318attribute dictionary, which associates a set of attribute names to attribute 1319values. MLIR's builtin dialect provides a rich set of 1320[builtin attribute values](#builtin-attribute-values) out of the box (such as 1321arrays, dictionaries, strings, etc.). Additionally, dialects can define their 1322own [dialect attribute values](#dialect-attribute-values). 1323 1324The top-level attribute dictionary attached to an operation has special 1325semantics. The attribute entries are considered to be of two different kinds 1326based on whether their dictionary key has a dialect prefix: 1327 1328- *inherent attributes* are inherent to the definition of an operation's 1329 semantics. The operation itself is expected to verify the consistency of these 1330 attributes. An example is the `predicate` attribute of the `std.cmpi` op. 1331 These attributes must have names that do not start with a dialect prefix. 1332 1333- *discardable attributes* have semantics defined externally to the operation 1334 itself, but must be compatible with the operations's semantics. These 1335 attributes must have names that start with a dialect prefix. The dialect 1336 indicated by the dialect prefix is expected to verify these attributes. An 1337 example is the `gpu.container_module` attribute. 1338 1339Note that attribute values are allowed to themselves be dictionary attributes, 1340but only the top-level dictionary attribute attached to the operation is subject 1341to the classification above. 1342 1343### Attribute Value Aliases 1344 1345``` 1346attribute-alias-def ::= '#' alias-name '=' attribute-value 1347attribute-alias ::= '#' alias-name 1348``` 1349 1350MLIR supports defining named aliases for attribute values. An attribute alias is 1351an identifier that can be used in the place of the attribute that it defines. 1352These aliases *must* be defined before their uses. Alias names may not contain a 1353'.', since those names are reserved for 1354[dialect attributes](#dialect-attribute-values). 1355 1356Example: 1357 1358```mlir 1359#map = affine_map<(d0) -> (d0 + 10)> 1360 1361// Using the original attribute. 1362%b = affine.apply affine_map<(d0) -> (d0 + 10)> (%a) 1363 1364// Using the attribute alias. 1365%b = affine.apply #map(%a) 1366``` 1367 1368### Dialect Attribute Values 1369 1370Similarly to operations, dialects may define custom attribute values. The 1371syntactic structure of these values is identical to custom dialect type values, 1372except that dialect attribute values are distinguished with a leading '#', while 1373dialect types are distinguished with a leading '!'. 1374 1375``` 1376dialect-attribute-value ::= '#' opaque-dialect-item 1377dialect-attribute-value ::= '#' pretty-dialect-item 1378``` 1379 1380Dialect attribute values can be specified in a verbose form, e.g. like this: 1381 1382```mlir 1383// Complex attribute value. 1384#foo<"something<abcd>"> 1385 1386// Even more complex attribute value. 1387#foo<"something<a%%123^^^>>>"> 1388``` 1389 1390Dialect attribute values that are simple enough can use the pretty format, which 1391is a lighter weight syntax that is equivalent to the above forms: 1392 1393```mlir 1394// Complex attribute 1395#foo.something<abcd> 1396``` 1397 1398Sufficiently complex dialect attribute values are required to use the verbose 1399form for generality. For example, the more complex type shown above would not be 1400valid in the lighter syntax: `#foo.something<a%%123^^^>>>` because it contains 1401characters that are not allowed in the lighter syntax, as well as unbalanced 1402`<>` characters. 1403 1404See [here](Tutorials/DefiningAttributesAndTypes.md) on how to define dialect 1405attribute values. 1406 1407### Builtin Attribute Values 1408 1409Builtin attributes are a core set of 1410[dialect attribute values](#dialect-attribute-values) that are defined in a 1411builtin dialect and thus available to all users of MLIR. 1412 1413``` 1414builtin-attribute ::= affine-map-attribute 1415 | array-attribute 1416 | bool-attribute 1417 | dictionary-attribute 1418 | elements-attribute 1419 | float-attribute 1420 | integer-attribute 1421 | integer-set-attribute 1422 | string-attribute 1423 | symbol-ref-attribute 1424 | type-attribute 1425 | unit-attribute 1426``` 1427 1428#### AffineMap Attribute 1429 1430Syntax: 1431 1432``` 1433affine-map-attribute ::= `affine_map` `<` affine-map `>` 1434``` 1435 1436An affine-map attribute is an attribute that represents an affine-map object. 1437 1438#### Array Attribute 1439 1440Syntax: 1441 1442``` 1443array-attribute ::= `[` (attribute-value (`,` attribute-value)*)? `]` 1444``` 1445 1446An array attribute is an attribute that represents a collection of attribute 1447values. 1448 1449#### Boolean Attribute 1450 1451Syntax: 1452 1453``` 1454bool-attribute ::= bool-literal 1455``` 1456 1457A boolean attribute is a literal attribute that represents a one-bit boolean 1458value, true or false. 1459 1460#### Dictionary Attribute 1461 1462Syntax: 1463 1464``` 1465dictionary-attribute ::= `{` (attribute-entry (`,` attribute-entry)*)? `}` 1466``` 1467 1468A dictionary attribute is an attribute that represents a sorted collection of 1469named attribute values. The elements are sorted by name, and each name must be 1470unique within the collection. 1471 1472#### Elements Attributes 1473 1474Syntax: 1475 1476``` 1477elements-attribute ::= dense-elements-attribute 1478 | opaque-elements-attribute 1479 | sparse-elements-attribute 1480``` 1481 1482An elements attribute is a literal attribute that represents a constant 1483[vector](#vector-type) or [tensor](#tensor-type) value. 1484 1485##### Dense Elements Attribute 1486 1487Syntax: 1488 1489``` 1490dense-elements-attribute ::= `dense` `<` attribute-value `>` `:` 1491 ( tensor-type | vector-type ) 1492``` 1493 1494A dense elements attribute is an elements attribute where the storage for the 1495constant vector or tensor value has been densely packed. The attribute supports 1496storing integer or floating point elements, with integer/index/floating element 1497types. It also support storing string elements with a custom dialect string 1498element type. 1499 1500##### Opaque Elements Attribute 1501 1502Syntax: 1503 1504``` 1505opaque-elements-attribute ::= `opaque` `<` dialect-namespace `,` 1506 hex-string-literal `>` `:` 1507 ( tensor-type | vector-type ) 1508``` 1509 1510An opaque elements attribute is an elements attribute where the content of the 1511value is opaque. The representation of the constant stored by this elements 1512attribute is only understood, and thus decodable, by the dialect that created 1513it. 1514 1515Note: The parsed string literal must be in hexadecimal form. 1516 1517##### Sparse Elements Attribute 1518 1519Syntax: 1520 1521``` 1522sparse-elements-attribute ::= `sparse` `<` attribute-value `,` attribute-value 1523 `>` `:` ( tensor-type | vector-type ) 1524``` 1525 1526A sparse elements attribute is an elements attribute that represents a sparse 1527vector or tensor object. This is where very few of the elements are non-zero. 1528 1529The attribute uses COO (coordinate list) encoding to represent the sparse 1530elements of the elements attribute. The indices are stored via a 2-D tensor of 153164-bit integer elements with shape [N, ndims], which specifies the indices of 1532the elements in the sparse tensor that contains non-zero values. The element 1533values are stored via a 1-D tensor with shape [N], that supplies the 1534corresponding values for the indices. 1535 1536Example: 1537 1538```mlir 1539 sparse<[[0, 0], [1, 2]], [1, 5]> : tensor<3x4xi32> 1540 1541// This represents the following tensor: 1542/// [[1, 0, 0, 0], 1543/// [0, 0, 5, 0], 1544/// [0, 0, 0, 0]] 1545``` 1546 1547#### Float Attribute 1548 1549Syntax: 1550 1551``` 1552float-attribute ::= (float-literal (`:` float-type)?) 1553 | (hexadecimal-literal `:` float-type) 1554``` 1555 1556A float attribute is a literal attribute that represents a floating point value 1557of the specified [float type](#floating-point-types). It can be represented in 1558the hexadecimal form where the hexadecimal value is interpreted as bits of the 1559underlying binary representation. This form is useful for representing infinity 1560and NaN floating point values. To avoid confusion with integer attributes, 1561hexadecimal literals _must_ be followed by a float type to define a float 1562attribute. 1563 1564Examples: 1565 1566``` 156742.0 // float attribute defaults to f64 type 156842.0 : f32 // float attribute of f32 type 15690x7C00 : f16 // positive infinity 15700x7CFF : f16 // NaN (one of possible values) 157142 : f32 // Error: expected integer type 1572``` 1573 1574#### Integer Attribute 1575 1576Syntax: 1577 1578``` 1579integer-attribute ::= integer-literal ( `:` (index-type | integer-type) )? 1580``` 1581 1582An integer attribute is a literal attribute that represents an integral value of 1583the specified integer or index type. The default type for this attribute, if one 1584is not specified, is a 64-bit integer. 1585 1586##### Integer Set Attribute 1587 1588Syntax: 1589 1590``` 1591integer-set-attribute ::= `affine_set` `<` integer-set `>` 1592``` 1593 1594An integer-set attribute is an attribute that represents an integer-set object. 1595 1596#### String Attribute 1597 1598Syntax: 1599 1600``` 1601string-attribute ::= string-literal (`:` type)? 1602``` 1603 1604A string attribute is an attribute that represents a string literal value. 1605 1606#### Symbol Reference Attribute 1607 1608Syntax: 1609 1610``` 1611symbol-ref-attribute ::= symbol-ref-id (`::` symbol-ref-id)* 1612``` 1613 1614A symbol reference attribute is a literal attribute that represents a named 1615reference to an operation that is nested within an operation with the 1616`OpTrait::SymbolTable` trait. As such, this reference is given meaning by the 1617nearest parent operation containing the `OpTrait::SymbolTable` trait. It may 1618optionally contain a set of nested references that further resolve to a symbol 1619nested within a different symbol table. 1620 1621This attribute can only be held internally by 1622[array attributes](#array-attribute) and 1623[dictionary attributes](#dictionary-attribute)(including the top-level operation 1624attribute dictionary), i.e. no other attribute kinds such as Locations or 1625extended attribute kinds. 1626 1627**Rationale:** Identifying accesses to global data is critical to 1628enabling efficient multi-threaded compilation. Restricting global 1629data access to occur through symbols and limiting the places that can 1630legally hold a symbol reference simplifies reasoning about these data 1631accesses. 1632 1633See [`Symbols And SymbolTables`](SymbolsAndSymbolTables.md) for more 1634information. 1635 1636#### Type Attribute 1637 1638Syntax: 1639 1640``` 1641type-attribute ::= type 1642``` 1643 1644A type attribute is an attribute that represents a [type object](#type-system). 1645 1646#### Unit Attribute 1647 1648``` 1649unit-attribute ::= `unit` 1650``` 1651 1652A unit attribute is an attribute that represents a value of `unit` type. The 1653`unit` type allows only one value forming a singleton set. This attribute value 1654is used to represent attributes that only have meaning from their existence. 1655 1656One example of such an attribute could be the `swift.self` attribute. This 1657attribute indicates that a function parameter is the self/context parameter. It 1658could be represented as a [boolean attribute](#boolean-attribute)(true or 1659false), but a value of false doesn't really bring any value. The parameter 1660either is the self/context or it isn't. 1661 1662```mlir 1663// A unit attribute defined with the `unit` value specifier. 1664func @verbose_form(i1) attributes {dialectName.unitAttr = unit} 1665 1666// A unit attribute can also be defined without the value specifier. 1667func @simple_form(i1) attributes {dialectName.unitAttr} 1668``` 1669