175897234Sdrh<html> 275897234Sdrh<head> 375897234Sdrh<title>The Lemon Parser Generator</title> 475897234Sdrh</head> 575897234Sdrh<body bgcolor=white> 675897234Sdrh<h1 align=center>The Lemon Parser Generator</h1> 775897234Sdrh 89bccde3dSdrh<p>Lemon is an LALR(1) parser generator for C. 99bccde3dSdrhIt does the same job as "bison" and "yacc". 109bccde3dSdrhBut lemon is not a bison or yacc clone. Lemon 1175897234Sdrhuses a different grammar syntax which is designed to 129bccde3dSdrhreduce the number of coding errors. Lemon also uses a 139bccde3dSdrhparsing engine that is faster than yacc and 149bccde3dSdrhbison and which is both reentrant and threadsafe. 159bccde3dSdrh(Update: Since the previous sentence was written, bison 169bccde3dSdrhhas also been updated so that it too can generate a 179bccde3dSdrhreentrant and threadsafe parser.) 189bccde3dSdrhLemon also implements features that can be used 1975897234Sdrhto eliminate resource leaks, making is suitable for use 2075897234Sdrhin long-running programs such as graphical user interfaces 2175897234Sdrhor embedded controllers.</p> 2275897234Sdrh 2375897234Sdrh<p>This document is an introduction to the Lemon 2475897234Sdrhparser generator.</p> 2575897234Sdrh 26*c5e56b34Sdrh<h2>Security Note</h2> 27*c5e56b34Sdrh 28*c5e56b34Sdrh<p>The language parser code created by Lemon is very robust and 29*c5e56b34Sdrhis well-suited for use in internet-facing applications that need to 30*c5e56b34Sdrhsafely process maliciously crafted inputs. 31*c5e56b34Sdrh 32*c5e56b34Sdrh<p>The "lemon.exe" command-line tool itself works great when given a valid 33*c5e56b34Sdrhinput grammar file and almost always gives helpful 34*c5e56b34Sdrherror messages for malformed inputs. However, it is possible for 35*c5e56b34Sdrha malicious user to craft a grammar file that will cause 36*c5e56b34Sdrhlemon.exe to crash. 37*c5e56b34SdrhWe do not see this as a problem, as lemon.exe is not intended to be used 38*c5e56b34Sdrhwith hostile inputs. 39*c5e56b34SdrhTo summarize:</p> 40*c5e56b34Sdrh 41*c5e56b34Sdrh<ul> 42*c5e56b34Sdrh<li>Parser code generated by lemon → Robust and secure 43*c5e56b34Sdrh<li>The "lemon.exe" command line tool itself → Not so much 44*c5e56b34Sdrh</ul> 45*c5e56b34Sdrh 4675897234Sdrh<h2>Theory of Operation</h2> 4775897234Sdrh 4875897234Sdrh<p>The main goal of Lemon is to translate a context free grammar (CFG) 4975897234Sdrhfor a particular language into C code that implements a parser for 5075897234Sdrhthat language. 5175897234SdrhThe program has two inputs: 5275897234Sdrh<ul> 5375897234Sdrh<li>The grammar specification. 5475897234Sdrh<li>A parser template file. 5575897234Sdrh</ul> 5675897234SdrhTypically, only the grammar specification is supplied by the programmer. 5775897234SdrhLemon comes with a default parser template which works fine for most 5875897234Sdrhapplications. But the user is free to substitute a different parser 5975897234Sdrhtemplate if desired.</p> 6075897234Sdrh 6175897234Sdrh<p>Depending on command-line options, Lemon will generate between 6275897234Sdrhone and three files of outputs. 6375897234Sdrh<ul> 6475897234Sdrh<li>C code to implement the parser. 6575897234Sdrh<li>A header file defining an integer ID for each terminal symbol. 6675897234Sdrh<li>An information file that describes the states of the generated parser 6775897234Sdrh automaton. 6875897234Sdrh</ul> 6975897234SdrhBy default, all three of these output files are generated. 709bccde3dSdrhThe header file is suppressed if the "-m" command-line option is 719bccde3dSdrhused and the report file is omitted when "-q" is selected.</p> 7275897234Sdrh 739bccde3dSdrh<p>The grammar specification file uses a ".y" suffix, by convention. 7475897234SdrhIn the examples used in this document, we'll assume the name of the 759bccde3dSdrhgrammar file is "gram.y". A typical use of Lemon would be the 7675897234Sdrhfollowing command: 7775897234Sdrh<pre> 7875897234Sdrh lemon gram.y 7975897234Sdrh</pre> 809bccde3dSdrhThis command will generate three output files named "gram.c", 819bccde3dSdrh"gram.h" and "gram.out". 8275897234SdrhThe first is C code to implement the parser. The second 8375897234Sdrhis the header file that defines numerical values for all 8475897234Sdrhterminal symbols, and the last is the report that explains 8575897234Sdrhthe states used by the parser automaton.</p> 8675897234Sdrh 8775897234Sdrh<h3>Command Line Options</h3> 8875897234Sdrh 8975897234Sdrh<p>The behavior of Lemon can be modified using command-line options. 9075897234SdrhYou can obtain a list of the available command-line options together 9175897234Sdrhwith a brief explanation of what each does by typing 9275897234Sdrh<pre> 9375897234Sdrh lemon -? 9475897234Sdrh</pre> 9575897234SdrhAs of this writing, the following command-line options are supported: 9675897234Sdrh<ul> 979bccde3dSdrh<li><b>-b</b> 989bccde3dSdrhShow only the basis for each parser state in the report file. 999bccde3dSdrh<li><b>-c</b> 1009bccde3dSdrhDo not compress the generated action tables. 1019bccde3dSdrh<li><b>-D<i>name</i></b> 1029bccde3dSdrhDefine C preprocessor macro <i>name</i>. This macro is useable by 1039bccde3dSdrh"%ifdef" lines in the grammar file. 1049bccde3dSdrh<li><b>-g</b> 1059bccde3dSdrhDo not generate a parser. Instead write the input grammar to standard 1069bccde3dSdrhoutput with all comments, actions, and other extraneous text removed. 1079bccde3dSdrh<li><b>-l</b> 108dfe4e6bbSdrhOmit "#line" directives in the generated parser C code. 1099bccde3dSdrh<li><b>-m</b> 1109bccde3dSdrhCause the output C source code to be compatible with the "makeheaders" 1119bccde3dSdrhprogram. 1129bccde3dSdrh<li><b>-p</b> 1139bccde3dSdrhDisplay all conflicts that are resolved by 1149bccde3dSdrh<a href='#precrules'>precedence rules</a>. 1159bccde3dSdrh<li><b>-q</b> 1169bccde3dSdrhSuppress generation of the report file. 1179bccde3dSdrh<li><b>-r</b> 1189bccde3dSdrhDo not sort or renumber the parser states as part of optimization. 1199bccde3dSdrh<li><b>-s</b> 1209bccde3dSdrhShow parser statistics before existing. 1219bccde3dSdrh<li><b>-T<i>file</i></b> 1229bccde3dSdrhUse <i>file</i> as the template for the generated C-code parser implementation. 1239bccde3dSdrh<li><b>-x</b> 1249bccde3dSdrhPrint the Lemon version number. 12575897234Sdrh</ul> 12675897234Sdrh 12775897234Sdrh<h3>The Parser Interface</h3> 12875897234Sdrh 12975897234Sdrh<p>Lemon doesn't generate a complete, working program. It only generates 13075897234Sdrha few subroutines that implement a parser. This section describes 13175897234Sdrhthe interface to those subroutines. It is up to the programmer to 13275897234Sdrhcall these subroutines in an appropriate way in order to produce a 13375897234Sdrhcomplete system.</p> 13475897234Sdrh 13575897234Sdrh<p>Before a program begins using a Lemon-generated parser, the program 13675897234Sdrhmust first create the parser. 13775897234SdrhA new parser is created as follows: 13875897234Sdrh<pre> 13975897234Sdrh void *pParser = ParseAlloc( malloc ); 14075897234Sdrh</pre> 14175897234SdrhThe ParseAlloc() routine allocates and initializes a new parser and 14275897234Sdrhreturns a pointer to it. 1439bccde3dSdrhThe actual data structure used to represent a parser is opaque — 14475897234Sdrhits internal structure is not visible or usable by the calling routine. 14575897234SdrhFor this reason, the ParseAlloc() routine returns a pointer to void 14675897234Sdrhrather than a pointer to some particular structure. 14775897234SdrhThe sole argument to the ParseAlloc() routine is a pointer to the 1489bccde3dSdrhsubroutine used to allocate memory. Typically this means malloc().</p> 14975897234Sdrh 15075897234Sdrh<p>After a program is finished using a parser, it can reclaim all 15175897234Sdrhmemory allocated by that parser by calling 15275897234Sdrh<pre> 15375897234Sdrh ParseFree(pParser, free); 15475897234Sdrh</pre> 15575897234SdrhThe first argument is the same pointer returned by ParseAlloc(). The 15675897234Sdrhsecond argument is a pointer to the function used to release bulk 15775897234Sdrhmemory back to the system.</p> 15875897234Sdrh 15975897234Sdrh<p>After a parser has been allocated using ParseAlloc(), the programmer 16075897234Sdrhmust supply the parser with a sequence of tokens (terminal symbols) to 16175897234Sdrhbe parsed. This is accomplished by calling the following function 16275897234Sdrhonce for each token: 16375897234Sdrh<pre> 16475897234Sdrh Parse(pParser, hTokenID, sTokenData, pArg); 16575897234Sdrh</pre> 16675897234SdrhThe first argument to the Parse() routine is the pointer returned by 16775897234SdrhParseAlloc(). 16875897234SdrhThe second argument is a small positive integer that tells the parse the 16975897234Sdrhtype of the next token in the data stream. 17075897234SdrhThere is one token type for each terminal symbol in the grammar. 17175897234SdrhThe gram.h file generated by Lemon contains #define statements that 17275897234Sdrhmap symbolic terminal symbol names into appropriate integer values. 1739bccde3dSdrhA value of 0 for the second argument is a special flag to the 1749bccde3dSdrhparser to indicate that the end of input has been reached. 17575897234SdrhThe third argument is the value of the given token. By default, 17675897234Sdrhthe type of the third argument is integer, but the grammar will 17775897234Sdrhusually redefine this type to be some kind of structure. 17875897234SdrhTypically the second argument will be a broad category of tokens 1799bccde3dSdrhsuch as "identifier" or "number" and the third argument will 18075897234Sdrhbe the name of the identifier or the value of the number.</p> 18175897234Sdrh 18275897234Sdrh<p>The Parse() function may have either three or four arguments, 18345f31be8Sdrhdepending on the grammar. If the grammar specification file requests 18445f31be8Sdrhit (via the <a href='#extraarg'><tt>extra_argument</tt> directive</a>), 18545f31be8Sdrhthe Parse() function will have a fourth parameter that can be 18675897234Sdrhof any type chosen by the programmer. The parser doesn't do anything 18775897234Sdrhwith this argument except to pass it through to action routines. 18875897234SdrhThis is a convenient mechanism for passing state information down 18975897234Sdrhto the action routines without having to use global variables.</p> 19075897234Sdrh 19175897234Sdrh<p>A typical use of a Lemon parser might look something like the 19275897234Sdrhfollowing: 19375897234Sdrh<pre> 19475897234Sdrh 01 ParseTree *ParseFile(const char *zFilename){ 19575897234Sdrh 02 Tokenizer *pTokenizer; 19675897234Sdrh 03 void *pParser; 19775897234Sdrh 04 Token sToken; 19875897234Sdrh 05 int hTokenId; 19975897234Sdrh 06 ParserState sState; 20075897234Sdrh 07 20175897234Sdrh 08 pTokenizer = TokenizerCreate(zFilename); 20275897234Sdrh 09 pParser = ParseAlloc( malloc ); 20375897234Sdrh 10 InitParserState(&sState); 20475897234Sdrh 11 while( GetNextToken(pTokenizer, &hTokenId, &sToken) ){ 20575897234Sdrh 12 Parse(pParser, hTokenId, sToken, &sState); 20675897234Sdrh 13 } 20775897234Sdrh 14 Parse(pParser, 0, sToken, &sState); 20875897234Sdrh 15 ParseFree(pParser, free ); 20975897234Sdrh 16 TokenizerFree(pTokenizer); 21075897234Sdrh 17 return sState.treeRoot; 21175897234Sdrh 18 } 21275897234Sdrh</pre> 21375897234SdrhThis example shows a user-written routine that parses a file of 21475897234Sdrhtext and returns a pointer to the parse tree. 2159bccde3dSdrh(All error-handling code is omitted from this example to keep it 21675897234Sdrhsimple.) 21775897234SdrhWe assume the existence of some kind of tokenizer which is created 21875897234Sdrhusing TokenizerCreate() on line 8 and deleted by TokenizerFree() 21975897234Sdrhon line 16. The GetNextToken() function on line 11 retrieves the 22075897234Sdrhnext token from the input file and puts its type in the 22175897234Sdrhinteger variable hTokenId. The sToken variable is assumed to be 22275897234Sdrhsome kind of structure that contains details about each token, 22375897234Sdrhsuch as its complete text, what line it occurs on, etc. </p> 22475897234Sdrh 22575897234Sdrh<p>This example also assumes the existence of structure of type 22675897234SdrhParserState that holds state information about a particular parse. 22775897234SdrhAn instance of such a structure is created on line 6 and initialized 22875897234Sdrhon line 10. A pointer to this structure is passed into the Parse() 22975897234Sdrhroutine as the optional 4th argument. 23075897234SdrhThe action routine specified by the grammar for the parser can use 23175897234Sdrhthe ParserState structure to hold whatever information is useful and 23275897234Sdrhappropriate. In the example, we note that the treeRoot field of 23375897234Sdrhthe ParserState structure is left pointing to the root of the parse 23475897234Sdrhtree.</p> 23575897234Sdrh 23675897234Sdrh<p>The core of this example as it relates to Lemon is as follows: 23775897234Sdrh<pre> 23875897234Sdrh ParseFile(){ 23975897234Sdrh pParser = ParseAlloc( malloc ); 24075897234Sdrh while( GetNextToken(pTokenizer,&hTokenId, &sToken) ){ 24175897234Sdrh Parse(pParser, hTokenId, sToken); 24275897234Sdrh } 24375897234Sdrh Parse(pParser, 0, sToken); 24475897234Sdrh ParseFree(pParser, free ); 24575897234Sdrh } 24675897234Sdrh</pre> 24775897234SdrhBasically, what a program has to do to use a Lemon-generated parser 24875897234Sdrhis first create the parser, then send it lots of tokens obtained by 24975897234Sdrhtokenizing an input source. When the end of input is reached, the 25075897234SdrhParse() routine should be called one last time with a token type 25175897234Sdrhof 0. This step is necessary to inform the parser that the end of 25275897234Sdrhinput has been reached. Finally, we reclaim memory used by the 25375897234Sdrhparser by calling ParseFree().</p> 25475897234Sdrh 25575897234Sdrh<p>There is one other interface routine that should be mentioned 25675897234Sdrhbefore we move on. 25775897234SdrhThe ParseTrace() function can be used to generate debugging output 25875897234Sdrhfrom the parser. A prototype for this routine is as follows: 25975897234Sdrh<pre> 26075897234Sdrh ParseTrace(FILE *stream, char *zPrefix); 26175897234Sdrh</pre> 26275897234SdrhAfter this routine is called, a short (one-line) message is written 26375897234Sdrhto the designated output stream every time the parser changes states 26475897234Sdrhor calls an action routine. Each such message is prefaced using 26575897234Sdrhthe text given by zPrefix. This debugging output can be turned off 26675897234Sdrhby calling ParseTrace() again with a first argument of NULL (0).</p> 26775897234Sdrh 26875897234Sdrh<h3>Differences With YACC and BISON</h3> 26975897234Sdrh 27075897234Sdrh<p>Programmers who have previously used the yacc or bison parser 27175897234Sdrhgenerator will notice several important differences between yacc and/or 27275897234Sdrhbison and Lemon. 27375897234Sdrh<ul> 27475897234Sdrh<li>In yacc and bison, the parser calls the tokenizer. In Lemon, 27575897234Sdrh the tokenizer calls the parser. 27675897234Sdrh<li>Lemon uses no global variables. Yacc and bison use global variables 27775897234Sdrh to pass information between the tokenizer and parser. 27875897234Sdrh<li>Lemon allows multiple parsers to be running simultaneously. Yacc 27975897234Sdrh and bison do not. 28075897234Sdrh</ul> 28175897234SdrhThese differences may cause some initial confusion for programmers 28275897234Sdrhwith prior yacc and bison experience. 28375897234SdrhBut after years of experience using Lemon, I firmly 28475897234Sdrhbelieve that the Lemon way of doing things is better.</p> 28575897234Sdrh 28645f31be8Sdrh<p><i>Updated as of 2016-02-16:</i> 28745f31be8SdrhThe text above was written in the 1990s. 28845f31be8SdrhWe are told that Bison has lately been enhanced to support the 28945f31be8Sdrhtokenizer-calls-parser paradigm used by Lemon, and to obviate the 29045f31be8Sdrhneed for global variables.</p> 29145f31be8Sdrh 29275897234Sdrh<h2>Input File Syntax</h2> 29375897234Sdrh 29475897234Sdrh<p>The main purpose of the grammar specification file for Lemon is 29575897234Sdrhto define the grammar for the parser. But the input file also 29675897234Sdrhspecifies additional information Lemon requires to do its job. 29775897234SdrhMost of the work in using Lemon is in writing an appropriate 29875897234Sdrhgrammar file.</p> 29975897234Sdrh 30075897234Sdrh<p>The grammar file for lemon is, for the most part, free format. 30175897234SdrhIt does not have sections or divisions like yacc or bison. Any 30275897234Sdrhdeclaration can occur at any point in the file. 30375897234SdrhLemon ignores whitespace (except where it is needed to separate 30475897234Sdrhtokens) and it honors the same commenting conventions as C and C++.</p> 30575897234Sdrh 30675897234Sdrh<h3>Terminals and Nonterminals</h3> 30775897234Sdrh 30875897234Sdrh<p>A terminal symbol (token) is any string of alphanumeric 3099bccde3dSdrhand/or underscore characters 31075897234Sdrhthat begins with an upper case letter. 311c8eee5e5SdrhA terminal can contain lowercase letters after the first character, 31275897234Sdrhbut the usual convention is to make terminals all upper case. 31375897234SdrhA nonterminal, on the other hand, is any string of alphanumeric 31475897234Sdrhand underscore characters than begins with a lower case letter. 31575897234SdrhAgain, the usual convention is to make nonterminals use all lower 31675897234Sdrhcase letters.</p> 31775897234Sdrh 31875897234Sdrh<p>In Lemon, terminal and nonterminal symbols do not need to 31975897234Sdrhbe declared or identified in a separate section of the grammar file. 32075897234SdrhLemon is able to generate a list of all terminals and nonterminals 32175897234Sdrhby examining the grammar rules, and it can always distinguish a 32275897234Sdrhterminal from a nonterminal by checking the case of the first 32375897234Sdrhcharacter of the name.</p> 32475897234Sdrh 32575897234Sdrh<p>Yacc and bison allow terminal symbols to have either alphanumeric 32675897234Sdrhnames or to be individual characters included in single quotes, like 32775897234Sdrhthis: ')' or '$'. Lemon does not allow this alternative form for 32875897234Sdrhterminal symbols. With Lemon, all symbols, terminals and nonterminals, 32975897234Sdrhmust have alphanumeric names.</p> 33075897234Sdrh 33175897234Sdrh<h3>Grammar Rules</h3> 33275897234Sdrh 33375897234Sdrh<p>The main component of a Lemon grammar file is a sequence of grammar 33475897234Sdrhrules. 33575897234SdrhEach grammar rule consists of a nonterminal symbol followed by 3369bccde3dSdrhthe special symbol "::=" and then a list of terminals and/or nonterminals. 33775897234SdrhThe rule is terminated by a period. 33875897234SdrhThe list of terminals and nonterminals on the right-hand side of the 33975897234Sdrhrule can be empty. 34075897234SdrhRules can occur in any order, except that the left-hand side of the 34175897234Sdrhfirst rule is assumed to be the start symbol for the grammar (unless 34275897234Sdrhspecified otherwise using the <tt>%start</tt> directive described below.) 34375897234SdrhA typical sequence of grammar rules might look something like this: 34475897234Sdrh<pre> 34575897234Sdrh expr ::= expr PLUS expr. 34675897234Sdrh expr ::= expr TIMES expr. 34775897234Sdrh expr ::= LPAREN expr RPAREN. 34875897234Sdrh expr ::= VALUE. 34975897234Sdrh</pre> 35075897234Sdrh</p> 35175897234Sdrh 3529bccde3dSdrh<p>There is one non-terminal in this example, "expr", and five 3539bccde3dSdrhterminal symbols or tokens: "PLUS", "TIMES", "LPAREN", 3549bccde3dSdrh"RPAREN" and "VALUE".</p> 35575897234Sdrh 35675897234Sdrh<p>Like yacc and bison, Lemon allows the grammar to specify a block 35775897234Sdrhof C code that will be executed whenever a grammar rule is reduced 35875897234Sdrhby the parser. 35975897234SdrhIn Lemon, this action is specified by putting the C code (contained 36075897234Sdrhwithin curly braces <tt>{...}</tt>) immediately after the 36175897234Sdrhperiod that closes the rule. 36275897234SdrhFor example: 36375897234Sdrh<pre> 36475897234Sdrh expr ::= expr PLUS expr. { printf("Doing an addition...\n"); } 36575897234Sdrh</pre> 36675897234Sdrh</p> 36775897234Sdrh 36875897234Sdrh<p>In order to be useful, grammar actions must normally be linked to 36975897234Sdrhtheir associated grammar rules. 3709bccde3dSdrhIn yacc and bison, this is accomplished by embedding a "$$" in the 37175897234Sdrhaction to stand for the value of the left-hand side of the rule and 3729bccde3dSdrhsymbols "$1", "$2", and so forth to stand for the value of 37375897234Sdrhthe terminal or nonterminal at position 1, 2 and so forth on the 37475897234Sdrhright-hand side of the rule. 37575897234SdrhThis idea is very powerful, but it is also very error-prone. The 37675897234Sdrhsingle most common source of errors in a yacc or bison grammar is 37775897234Sdrhto miscount the number of symbols on the right-hand side of a grammar 3789bccde3dSdrhrule and say "$7" when you really mean "$8".</p> 37975897234Sdrh 38075897234Sdrh<p>Lemon avoids the need to count grammar symbols by assigning symbolic 38175897234Sdrhnames to each symbol in a grammar rule and then using those symbolic 38275897234Sdrhnames in the action. 38375897234SdrhIn yacc or bison, one would write this: 38475897234Sdrh<pre> 38575897234Sdrh expr -> expr PLUS expr { $$ = $1 + $3; }; 38675897234Sdrh</pre> 38775897234SdrhBut in Lemon, the same rule becomes the following: 38875897234Sdrh<pre> 38975897234Sdrh expr(A) ::= expr(B) PLUS expr(C). { A = B+C; } 39075897234Sdrh</pre> 39175897234SdrhIn the Lemon rule, any symbol in parentheses after a grammar rule 39275897234Sdrhsymbol becomes a place holder for that symbol in the grammar rule. 39375897234SdrhThis place holder can then be used in the associated C action to 39475897234Sdrhstand for the value of that symbol.<p> 39575897234Sdrh 39675897234Sdrh<p>The Lemon notation for linking a grammar rule with its reduce 39775897234Sdrhaction is superior to yacc/bison on several counts. 39875897234SdrhFirst, as mentioned above, the Lemon method avoids the need to 39975897234Sdrhcount grammar symbols. 40075897234SdrhSecondly, if a terminal or nonterminal in a Lemon grammar rule 40175897234Sdrhincludes a linking symbol in parentheses but that linking symbol 40275897234Sdrhis not actually used in the reduce action, then an error message 40375897234Sdrhis generated. 40475897234SdrhFor example, the rule 40575897234Sdrh<pre> 40675897234Sdrh expr(A) ::= expr(B) PLUS expr(C). { A = B; } 40775897234Sdrh</pre> 4089bccde3dSdrhwill generate an error because the linking symbol "C" is used 40975897234Sdrhin the grammar rule but not in the reduce action.</p> 41075897234Sdrh 41175897234Sdrh<p>The Lemon notation for linking grammar rules to reduce actions 41275897234Sdrhalso facilitates the use of destructors for reclaiming memory 41375897234Sdrhallocated by the values of terminals and nonterminals on the 41475897234Sdrhright-hand side of a rule.</p> 41575897234Sdrh 4169bccde3dSdrh<a name='precrules'></a> 41775897234Sdrh<h3>Precedence Rules</h3> 41875897234Sdrh 41975897234Sdrh<p>Lemon resolves parsing ambiguities in exactly the same way as 42075897234Sdrhyacc and bison. A shift-reduce conflict is resolved in favor 42175897234Sdrhof the shift, and a reduce-reduce conflict is resolved by reducing 42275897234Sdrhwhichever rule comes first in the grammar file.</p> 42375897234Sdrh 42475897234Sdrh<p>Just like in 42575897234Sdrhyacc and bison, Lemon allows a measure of control 42675897234Sdrhover the resolution of paring conflicts using precedence rules. 42775897234SdrhA precedence value can be assigned to any terminal symbol 4289bccde3dSdrhusing the 4299bccde3dSdrh<a href='#pleft'>%left</a>, 4309bccde3dSdrh<a href='#pright'>%right</a> or 4319bccde3dSdrh<a href='#pnonassoc'>%nonassoc</a> directives. Terminal symbols 43275897234Sdrhmentioned in earlier directives have a lower precedence that 43375897234Sdrhterminal symbols mentioned in later directives. For example:</p> 43475897234Sdrh 43575897234Sdrh<p><pre> 43675897234Sdrh %left AND. 43775897234Sdrh %left OR. 43875897234Sdrh %nonassoc EQ NE GT GE LT LE. 43975897234Sdrh %left PLUS MINUS. 44075897234Sdrh %left TIMES DIVIDE MOD. 44175897234Sdrh %right EXP NOT. 44275897234Sdrh</pre></p> 44375897234Sdrh 44475897234Sdrh<p>In the preceding sequence of directives, the AND operator is 44575897234Sdrhdefined to have the lowest precedence. The OR operator is one 44675897234Sdrhprecedence level higher. And so forth. Hence, the grammar would 44775897234Sdrhattempt to group the ambiguous expression 44875897234Sdrh<pre> 44975897234Sdrh a AND b OR c 45075897234Sdrh</pre> 45175897234Sdrhlike this 45275897234Sdrh<pre> 45375897234Sdrh a AND (b OR c). 45475897234Sdrh</pre> 45575897234SdrhThe associativity (left, right or nonassoc) is used to determine 45675897234Sdrhthe grouping when the precedence is the same. AND is left-associative 45775897234Sdrhin our example, so 45875897234Sdrh<pre> 45975897234Sdrh a AND b AND c 46075897234Sdrh</pre> 46175897234Sdrhis parsed like this 46275897234Sdrh<pre> 46375897234Sdrh (a AND b) AND c. 46475897234Sdrh</pre> 46575897234SdrhThe EXP operator is right-associative, though, so 46675897234Sdrh<pre> 46775897234Sdrh a EXP b EXP c 46875897234Sdrh</pre> 46975897234Sdrhis parsed like this 47075897234Sdrh<pre> 47175897234Sdrh a EXP (b EXP c). 47275897234Sdrh</pre> 47375897234SdrhThe nonassoc precedence is used for non-associative operators. 47475897234SdrhSo 47575897234Sdrh<pre> 47675897234Sdrh a EQ b EQ c 47775897234Sdrh</pre> 47875897234Sdrhis an error.</p> 47975897234Sdrh 48075897234Sdrh<p>The precedence of non-terminals is transferred to rules as follows: 48175897234SdrhThe precedence of a grammar rule is equal to the precedence of the 48275897234Sdrhleft-most terminal symbol in the rule for which a precedence is 48375897234Sdrhdefined. This is normally what you want, but in those cases where 48475897234Sdrhyou want to precedence of a grammar rule to be something different, 48575897234Sdrhyou can specify an alternative precedence symbol by putting the 48675897234Sdrhsymbol in square braces after the period at the end of the rule and 48775897234Sdrhbefore any C-code. For example:</p> 48875897234Sdrh 48975897234Sdrh<p><pre> 49075897234Sdrh expr = MINUS expr. [NOT] 49175897234Sdrh</pre></p> 49275897234Sdrh 49375897234Sdrh<p>This rule has a precedence equal to that of the NOT symbol, not the 49475897234SdrhMINUS symbol as would have been the case by default.</p> 49575897234Sdrh 49675897234Sdrh<p>With the knowledge of how precedence is assigned to terminal 49775897234Sdrhsymbols and individual 49875897234Sdrhgrammar rules, we can now explain precisely how parsing conflicts 49975897234Sdrhare resolved in Lemon. Shift-reduce conflicts are resolved 50075897234Sdrhas follows: 50175897234Sdrh<ul> 50275897234Sdrh<li> If either the token to be shifted or the rule to be reduced 50375897234Sdrh lacks precedence information, then resolve in favor of the 50475897234Sdrh shift, but report a parsing conflict. 50575897234Sdrh<li> If the precedence of the token to be shifted is greater than 50675897234Sdrh the precedence of the rule to reduce, then resolve in favor 50775897234Sdrh of the shift. No parsing conflict is reported. 50875897234Sdrh<li> If the precedence of the token it be shifted is less than the 50975897234Sdrh precedence of the rule to reduce, then resolve in favor of the 51075897234Sdrh reduce action. No parsing conflict is reported. 51175897234Sdrh<li> If the precedences are the same and the shift token is 51275897234Sdrh right-associative, then resolve in favor of the shift. 51375897234Sdrh No parsing conflict is reported. 514d5578433Smistachkin<li> If the precedences are the same the shift token is 51575897234Sdrh left-associative, then resolve in favor of the reduce. 51675897234Sdrh No parsing conflict is reported. 51775897234Sdrh<li> Otherwise, resolve the conflict by doing the shift and 51875897234Sdrh report the parsing conflict. 51975897234Sdrh</ul> 52075897234SdrhReduce-reduce conflicts are resolved this way: 52175897234Sdrh<ul> 52275897234Sdrh<li> If either reduce rule 52375897234Sdrh lacks precedence information, then resolve in favor of the 52475897234Sdrh rule that appears first in the grammar and report a parsing 52575897234Sdrh conflict. 52675897234Sdrh<li> If both rules have precedence and the precedence is different 52775897234Sdrh then resolve the dispute in favor of the rule with the highest 52875897234Sdrh precedence and do not report a conflict. 52975897234Sdrh<li> Otherwise, resolve the conflict by reducing by the rule that 53075897234Sdrh appears first in the grammar and report a parsing conflict. 53175897234Sdrh</ul> 53275897234Sdrh 53375897234Sdrh<h3>Special Directives</h3> 53475897234Sdrh 53575897234Sdrh<p>The input grammar to Lemon consists of grammar rules and special 53675897234Sdrhdirectives. We've described all the grammar rules, so now we'll 53775897234Sdrhtalk about the special directives.</p> 53875897234Sdrh 53975897234Sdrh<p>Directives in lemon can occur in any order. You can put them before 54075897234Sdrhthe grammar rules, or after the grammar rules, or in the mist of the 54175897234Sdrhgrammar rules. It doesn't matter. The relative order of 54275897234Sdrhdirectives used to assign precedence to terminals is important, but 54375897234Sdrhother than that, the order of directives in Lemon is arbitrary.</p> 54475897234Sdrh 54575897234Sdrh<p>Lemon supports the following special directives: 54675897234Sdrh<ul> 547f2340fc7Sdrh<li><tt>%code</tt> 548f2340fc7Sdrh<li><tt>%default_destructor</tt> 549f2340fc7Sdrh<li><tt>%default_type</tt> 55075897234Sdrh<li><tt>%destructor</tt> 5519bccde3dSdrh<li><tt>%endif</tt> 55275897234Sdrh<li><tt>%extra_argument</tt> 5539bccde3dSdrh<li><tt>%fallback</tt> 5549bccde3dSdrh<li><tt>%ifdef</tt> 5559bccde3dSdrh<li><tt>%ifndef</tt> 55675897234Sdrh<li><tt>%include</tt> 55775897234Sdrh<li><tt>%left</tt> 55875897234Sdrh<li><tt>%name</tt> 55975897234Sdrh<li><tt>%nonassoc</tt> 56075897234Sdrh<li><tt>%parse_accept</tt> 56175897234Sdrh<li><tt>%parse_failure </tt> 56275897234Sdrh<li><tt>%right</tt> 56375897234Sdrh<li><tt>%stack_overflow</tt> 56475897234Sdrh<li><tt>%stack_size</tt> 56575897234Sdrh<li><tt>%start_symbol</tt> 56675897234Sdrh<li><tt>%syntax_error</tt> 5679bccde3dSdrh<li><tt>%token_class</tt> 56875897234Sdrh<li><tt>%token_destructor</tt> 56975897234Sdrh<li><tt>%token_prefix</tt> 57075897234Sdrh<li><tt>%token_type</tt> 57175897234Sdrh<li><tt>%type</tt> 5729bccde3dSdrh<li><tt>%wildcard</tt> 57375897234Sdrh</ul> 57475897234SdrhEach of these directives will be described separately in the 57575897234Sdrhfollowing sections:</p> 57675897234Sdrh 5779bccde3dSdrh<a name='pcode'></a> 578f2340fc7Sdrh<h4>The <tt>%code</tt> directive</h4> 579f2340fc7Sdrh 5809bccde3dSdrh<p>The %code directive is used to specify addition C code that 581f2340fc7Sdrhis added to the end of the main output file. This is similar to 5829bccde3dSdrhthe <a href='#pinclude'>%include</a> directive except that %include 5839bccde3dSdrhis inserted at the beginning of the main output file.</p> 584f2340fc7Sdrh 585f2340fc7Sdrh<p>%code is typically used to include some action routines or perhaps 5869bccde3dSdrha tokenizer or even the "main()" function 5879bccde3dSdrhas part of the output file.</p> 588f2340fc7Sdrh 5899bccde3dSdrh<a name='default_destructor'></a> 590f2340fc7Sdrh<h4>The <tt>%default_destructor</tt> directive</h4> 591f2340fc7Sdrh 592f2340fc7Sdrh<p>The %default_destructor directive specifies a destructor to 593f2340fc7Sdrhuse for non-terminals that do not have their own destructor 594f2340fc7Sdrhspecified by a separate %destructor directive. See the documentation 5959bccde3dSdrhon the <a name='#destructor'>%destructor</a> directive below for 5969bccde3dSdrhadditional information.</p> 597f2340fc7Sdrh 598f2340fc7Sdrh<p>In some grammers, many different non-terminal symbols have the 599f2340fc7Sdrhsame datatype and hence the same destructor. This directive is 600f2340fc7Sdrha convenience way to specify the same destructor for all those 601f2340fc7Sdrhnon-terminals using a single statement.</p> 602f2340fc7Sdrh 6039bccde3dSdrh<a name='default_type'></a> 604f2340fc7Sdrh<h4>The <tt>%default_type</tt> directive</h4> 605f2340fc7Sdrh 606f2340fc7Sdrh<p>The %default_type directive specifies the datatype of non-terminal 607f2340fc7Sdrhsymbols that do no have their own datatype defined using a separate 6089bccde3dSdrh<a href='#ptype'>%type</a> directive. 6099bccde3dSdrh</p> 610f2340fc7Sdrh 6119bccde3dSdrh<a name='destructor'></a> 61275897234Sdrh<h4>The <tt>%destructor</tt> directive</h4> 61375897234Sdrh 61475897234Sdrh<p>The %destructor directive is used to specify a destructor for 61575897234Sdrha non-terminal symbol. 6169bccde3dSdrh(See also the <a href='#token_destructor'>%token_destructor</a> 6179bccde3dSdrhdirective which is used to specify a destructor for terminal symbols.)</p> 61875897234Sdrh 61975897234Sdrh<p>A non-terminal's destructor is called to dispose of the 62075897234Sdrhnon-terminal's value whenever the non-terminal is popped from 62175897234Sdrhthe stack. This includes all of the following circumstances: 62275897234Sdrh<ul> 62375897234Sdrh<li> When a rule reduces and the value of a non-terminal on 62475897234Sdrh the right-hand side is not linked to C code. 62575897234Sdrh<li> When the stack is popped during error processing. 62675897234Sdrh<li> When the ParseFree() function runs. 62775897234Sdrh</ul> 62875897234SdrhThe destructor can do whatever it wants with the value of 62975897234Sdrhthe non-terminal, but its design is to deallocate memory 63075897234Sdrhor other resources held by that non-terminal.</p> 63175897234Sdrh 63275897234Sdrh<p>Consider an example: 63375897234Sdrh<pre> 63475897234Sdrh %type nt {void*} 63575897234Sdrh %destructor nt { free($$); } 63675897234Sdrh nt(A) ::= ID NUM. { A = malloc( 100 ); } 63775897234Sdrh</pre> 63875897234SdrhThis example is a bit contrived but it serves to illustrate how 63975897234Sdrhdestructors work. The example shows a non-terminal named 6409bccde3dSdrh"nt" that holds values of type "void*". When the rule for 6419bccde3dSdrhan "nt" reduces, it sets the value of the non-terminal to 64275897234Sdrhspace obtained from malloc(). Later, when the nt non-terminal 64375897234Sdrhis popped from the stack, the destructor will fire and call 64475897234Sdrhfree() on this malloced space, thus avoiding a memory leak. 6459bccde3dSdrh(Note that the symbol "$$" in the destructor code is replaced 64675897234Sdrhby the value of the non-terminal.)</p> 64775897234Sdrh 64875897234Sdrh<p>It is important to note that the value of a non-terminal is passed 64975897234Sdrhto the destructor whenever the non-terminal is removed from the 65075897234Sdrhstack, unless the non-terminal is used in a C-code action. If 65175897234Sdrhthe non-terminal is used by C-code, then it is assumed that the 6529bccde3dSdrhC-code will take care of destroying it. 6539bccde3dSdrhMore commonly, the value is used to build some 65475897234Sdrhlarger structure and we don't want to destroy it, which is why 65575897234Sdrhthe destructor is not called in this circumstance.</p> 65675897234Sdrh 6579bccde3dSdrh<p>Destructors help avoid memory leaks by automatically freeing 6589bccde3dSdrhallocated objects when they go out of scope. 65975897234SdrhTo do the same using yacc or bison is much more difficult.</p> 66075897234Sdrh 66145f31be8Sdrh<a name="extraarg"></a> 66275897234Sdrh<h4>The <tt>%extra_argument</tt> directive</h4> 66375897234Sdrh 66475897234SdrhThe %extra_argument directive instructs Lemon to add a 4th parameter 66575897234Sdrhto the parameter list of the Parse() function it generates. Lemon 66675897234Sdrhdoesn't do anything itself with this extra argument, but it does 66775897234Sdrhmake the argument available to C-code action routines, destructors, 66875897234Sdrhand so forth. For example, if the grammar file contains:</p> 66975897234Sdrh 67075897234Sdrh<p><pre> 67175897234Sdrh %extra_argument { MyStruct *pAbc } 67275897234Sdrh</pre></p> 67375897234Sdrh 67475897234Sdrh<p>Then the Parse() function generated will have an 4th parameter 6759bccde3dSdrhof type "MyStruct*" and all action routines will have access to 6769bccde3dSdrha variable named "pAbc" that is the value of the 4th parameter 67775897234Sdrhin the most recent call to Parse().</p> 67875897234Sdrh 6799bccde3dSdrh<a name='pfallback'></a> 6809bccde3dSdrh<h4>The <tt>%fallback</tt> directive</h4> 6819bccde3dSdrh 6829bccde3dSdrh<p>The %fallback directive specifies an alternative meaning for one 6839bccde3dSdrhor more tokens. The alternative meaning is tried if the original token 6849bccde3dSdrhwould have generated a syntax error. 6859bccde3dSdrh 6869bccde3dSdrh<p>The %fallback directive was added to support robust parsing of SQL 6879bccde3dSdrhsyntax in <a href="https://www.sqlite.org/">SQLite</a>. 6889bccde3dSdrhThe SQL language contains a large assortment of keywords, each of which 6899bccde3dSdrhappears as a different token to the language parser. SQL contains so 6909bccde3dSdrhmany keywords, that it can be difficult for programmers to keep up with 6919bccde3dSdrhthem all. Programmers will, therefore, sometimes mistakenly use an 6929bccde3dSdrhobscure language keyword for an identifier. The %fallback directive 6939bccde3dSdrhprovides a mechanism to tell the parser: "If you are unable to parse 6949bccde3dSdrhthis keyword, try treating it as an identifier instead." 6959bccde3dSdrh 6969bccde3dSdrh<p>The syntax of %fallback is as follows: 6979bccde3dSdrh 6989bccde3dSdrh<blockquote> 6999bccde3dSdrh<tt>%fallback</tt> <i>ID</i> <i>TOKEN...</i> <b>.</b> 7009bccde3dSdrh</blockquote> 7019bccde3dSdrh 7029bccde3dSdrh<p>In words, the %fallback directive is followed by a list of token names 7039bccde3dSdrhterminated by a period. The first token name is the fallback token - the 7049bccde3dSdrhtoken to which all the other tokens fall back to. The second and subsequent 7059bccde3dSdrharguments are tokens which fall back to the token identified by the first 7069bccde3dSdrhargument. 7079bccde3dSdrh 7089bccde3dSdrh<a name='pifdef'></a> 7099bccde3dSdrh<h4>The <tt>%ifdef</tt>, <tt>%ifndef</tt>, and <tt>%endif</tt> directives.</h4> 7109bccde3dSdrh 7119bccde3dSdrh<p>The %ifdef, %ifndef, and %endif directives are similar to 7129bccde3dSdrh#ifdef, #ifndef, and #endif in the C-preprocessor, just not as general. 7139bccde3dSdrhEach of these directives must begin at the left margin. No whitespace 7149bccde3dSdrhis allowed between the "%" and the directive name. 7159bccde3dSdrh 7169bccde3dSdrh<p>Grammar text in between "%ifdef MACRO" and the next nested "%endif" is 7179bccde3dSdrhignored unless the "-DMACRO" command-line option is used. Grammar text 7189bccde3dSdrhbetwen "%ifndef MACRO" and the next nested "%endif" is included except when 7199bccde3dSdrhthe "-DMACRO" command-line option is used. 7209bccde3dSdrh 7219bccde3dSdrh<p>Note that the argument to %ifdef and %ifndef must be a single 7229bccde3dSdrhpreprocessor symbol name, not a general expression. There is no "%else" 7239bccde3dSdrhdirective. 7249bccde3dSdrh 7259bccde3dSdrh 7269bccde3dSdrh<a name='pinclude'></a> 72775897234Sdrh<h4>The <tt>%include</tt> directive</h4> 72875897234Sdrh 72975897234Sdrh<p>The %include directive specifies C code that is included at the 73075897234Sdrhtop of the generated parser. You can include any text you want -- 731f2340fc7Sdrhthe Lemon parser generator copies it blindly. If you have multiple 7329bccde3dSdrh%include directives in your grammar file, their values are concatenated 7339bccde3dSdrhso that all %include code ultimately appears near the top of the 7349bccde3dSdrhgenerated parser, in the same order as it appeared in the grammer.</p> 73575897234Sdrh 73675897234Sdrh<p>The %include directive is very handy for getting some extra #include 73775897234Sdrhpreprocessor statements at the beginning of the generated parser. 73875897234SdrhFor example:</p> 73975897234Sdrh 74075897234Sdrh<p><pre> 74175897234Sdrh %include {#include <unistd.h>} 74275897234Sdrh</pre></p> 74375897234Sdrh 74475897234Sdrh<p>This might be needed, for example, if some of the C actions in the 74575897234Sdrhgrammar call functions that are prototyed in unistd.h.</p> 74675897234Sdrh 7479bccde3dSdrh<a name='pleft'></a> 74875897234Sdrh<h4>The <tt>%left</tt> directive</h4> 74975897234Sdrh 7509bccde3dSdrhThe %left directive is used (along with the <a href='#pright'>%right</a> and 7519bccde3dSdrh<a href='#pnonassoc'>%nonassoc</a> directives) to declare precedences of 7529bccde3dSdrhterminal symbols. Every terminal symbol whose name appears after 7539bccde3dSdrha %left directive but before the next period (".") is 75475897234Sdrhgiven the same left-associative precedence value. Subsequent 75575897234Sdrh%left directives have higher precedence. For example:</p> 75675897234Sdrh 75775897234Sdrh<p><pre> 75875897234Sdrh %left AND. 75975897234Sdrh %left OR. 76075897234Sdrh %nonassoc EQ NE GT GE LT LE. 76175897234Sdrh %left PLUS MINUS. 76275897234Sdrh %left TIMES DIVIDE MOD. 76375897234Sdrh %right EXP NOT. 76475897234Sdrh</pre></p> 76575897234Sdrh 76675897234Sdrh<p>Note the period that terminates each %left, %right or %nonassoc 76775897234Sdrhdirective.</p> 76875897234Sdrh 76975897234Sdrh<p>LALR(1) grammars can get into a situation where they require 77075897234Sdrha large amount of stack space if you make heavy use or right-associative 77175897234Sdrhoperators. For this reason, it is recommended that you use %left 77275897234Sdrhrather than %right whenever possible.</p> 77375897234Sdrh 7749bccde3dSdrh<a name='pname'></a> 77575897234Sdrh<h4>The <tt>%name</tt> directive</h4> 77675897234Sdrh 77775897234Sdrh<p>By default, the functions generated by Lemon all begin with the 7789bccde3dSdrhfive-character string "Parse". You can change this string to something 77975897234Sdrhdifferent using the %name directive. For instance:</p> 78075897234Sdrh 78175897234Sdrh<p><pre> 78275897234Sdrh %name Abcde 78375897234Sdrh</pre></p> 78475897234Sdrh 78575897234Sdrh<p>Putting this directive in the grammar file will cause Lemon to generate 78675897234Sdrhfunctions named 78775897234Sdrh<ul> 78875897234Sdrh<li> AbcdeAlloc(), 78975897234Sdrh<li> AbcdeFree(), 79075897234Sdrh<li> AbcdeTrace(), and 79175897234Sdrh<li> Abcde(). 79275897234Sdrh</ul> 79375897234SdrhThe %name directive allows you to generator two or more different 79475897234Sdrhparsers and link them all into the same executable. 79575897234Sdrh</p> 79675897234Sdrh 7979bccde3dSdrh<a name='pnonassoc'></a> 79875897234Sdrh<h4>The <tt>%nonassoc</tt> directive</h4> 79975897234Sdrh 80075897234Sdrh<p>This directive is used to assign non-associative precedence to 8019bccde3dSdrhone or more terminal symbols. See the section on 8029bccde3dSdrh<a href='#precrules'>precedence rules</a> 8039bccde3dSdrhor on the <a href='#pleft'>%left</a> directive for additional information.</p> 80475897234Sdrh 8059bccde3dSdrh<a name='parse_accept'></a> 80675897234Sdrh<h4>The <tt>%parse_accept</tt> directive</h4> 80775897234Sdrh 80875897234Sdrh<p>The %parse_accept directive specifies a block of C code that is 8099bccde3dSdrhexecuted whenever the parser accepts its input string. To "accept" 81075897234Sdrhan input string means that the parser was able to process all tokens 81175897234Sdrhwithout error.</p> 81275897234Sdrh 81375897234Sdrh<p>For example:</p> 81475897234Sdrh 81575897234Sdrh<p><pre> 81675897234Sdrh %parse_accept { 81775897234Sdrh printf("parsing complete!\n"); 81875897234Sdrh } 81975897234Sdrh</pre></p> 82075897234Sdrh 8219bccde3dSdrh<a name='parse_failure'></a> 82275897234Sdrh<h4>The <tt>%parse_failure</tt> directive</h4> 82375897234Sdrh 82475897234Sdrh<p>The %parse_failure directive specifies a block of C code that 82575897234Sdrhis executed whenever the parser fails complete. This code is not 82675897234Sdrhexecuted until the parser has tried and failed to resolve an input 82775897234Sdrherror using is usual error recovery strategy. The routine is 82875897234Sdrhonly invoked when parsing is unable to continue.</p> 82975897234Sdrh 83075897234Sdrh<p><pre> 83175897234Sdrh %parse_failure { 83275897234Sdrh fprintf(stderr,"Giving up. Parser is hopelessly lost...\n"); 83375897234Sdrh } 83475897234Sdrh</pre></p> 83575897234Sdrh 8369bccde3dSdrh<a name='pright'></a> 83775897234Sdrh<h4>The <tt>%right</tt> directive</h4> 83875897234Sdrh 83975897234Sdrh<p>This directive is used to assign right-associative precedence to 8409bccde3dSdrhone or more terminal symbols. See the section on 8419bccde3dSdrh<a href='#precrules'>precedence rules</a> 8429bccde3dSdrhor on the <a href='#pleft'>%left</a> directive for additional information.</p> 84375897234Sdrh 8449bccde3dSdrh<a name='stack_overflow'></a> 84575897234Sdrh<h4>The <tt>%stack_overflow</tt> directive</h4> 84675897234Sdrh 84775897234Sdrh<p>The %stack_overflow directive specifies a block of C code that 84875897234Sdrhis executed if the parser's internal stack ever overflows. Typically 84975897234Sdrhthis just prints an error message. After a stack overflow, the parser 85075897234Sdrhwill be unable to continue and must be reset.</p> 85175897234Sdrh 85275897234Sdrh<p><pre> 85375897234Sdrh %stack_overflow { 85475897234Sdrh fprintf(stderr,"Giving up. Parser stack overflow\n"); 85575897234Sdrh } 85675897234Sdrh</pre></p> 85775897234Sdrh 85875897234Sdrh<p>You can help prevent parser stack overflows by avoiding the use 85975897234Sdrhof right recursion and right-precedence operators in your grammar. 86075897234SdrhUse left recursion and and left-precedence operators instead, to 86175897234Sdrhencourage rules to reduce sooner and keep the stack size down. 86275897234SdrhFor example, do rules like this: 86375897234Sdrh<pre> 86475897234Sdrh list ::= list element. // left-recursion. Good! 86575897234Sdrh list ::= . 86675897234Sdrh</pre> 86775897234SdrhNot like this: 86875897234Sdrh<pre> 86975897234Sdrh list ::= element list. // right-recursion. Bad! 87075897234Sdrh list ::= . 87175897234Sdrh</pre> 87275897234Sdrh 8739bccde3dSdrh<a name='stack_size'></a> 87475897234Sdrh<h4>The <tt>%stack_size</tt> directive</h4> 87575897234Sdrh 87675897234Sdrh<p>If stack overflow is a problem and you can't resolve the trouble 87775897234Sdrhby using left-recursion, then you might want to increase the size 87875897234Sdrhof the parser's stack using this directive. Put an positive integer 87975897234Sdrhafter the %stack_size directive and Lemon will generate a parse 88075897234Sdrhwith a stack of the requested size. The default value is 100.</p> 88175897234Sdrh 88275897234Sdrh<p><pre> 88375897234Sdrh %stack_size 2000 88475897234Sdrh</pre></p> 88575897234Sdrh 8869bccde3dSdrh<a name='start_symbol'></a> 88775897234Sdrh<h4>The <tt>%start_symbol</tt> directive</h4> 88875897234Sdrh 88975897234Sdrh<p>By default, the start-symbol for the grammar that Lemon generates 89075897234Sdrhis the first non-terminal that appears in the grammar file. But you 89175897234Sdrhcan choose a different start-symbol using the %start_symbol directive.</p> 89275897234Sdrh 89375897234Sdrh<p><pre> 89475897234Sdrh %start_symbol prog 89575897234Sdrh</pre></p> 89675897234Sdrh 8979bccde3dSdrh<a name='token_destructor'></a> 89875897234Sdrh<h4>The <tt>%token_destructor</tt> directive</h4> 89975897234Sdrh 90075897234Sdrh<p>The %destructor directive assigns a destructor to a non-terminal 90175897234Sdrhsymbol. (See the description of the %destructor directive above.) 90275897234SdrhThis directive does the same thing for all terminal symbols.</p> 90375897234Sdrh 90475897234Sdrh<p>Unlike non-terminal symbols which may each have a different data type 90575897234Sdrhfor their values, terminals all use the same data type (defined by 90675897234Sdrhthe %token_type directive) and so they use a common destructor. Other 90775897234Sdrhthan that, the token destructor works just like the non-terminal 90875897234Sdrhdestructors.</p> 90975897234Sdrh 9109bccde3dSdrh<a name='token_prefix'></a> 91175897234Sdrh<h4>The <tt>%token_prefix</tt> directive</h4> 91275897234Sdrh 91375897234Sdrh<p>Lemon generates #defines that assign small integer constants 91475897234Sdrhto each terminal symbol in the grammar. If desired, Lemon will 91575897234Sdrhadd a prefix specified by this directive 91675897234Sdrhto each of the #defines it generates. 91775897234SdrhSo if the default output of Lemon looked like this: 91875897234Sdrh<pre> 91975897234Sdrh #define AND 1 92075897234Sdrh #define MINUS 2 92175897234Sdrh #define OR 3 92275897234Sdrh #define PLUS 4 92375897234Sdrh</pre> 92475897234SdrhYou can insert a statement into the grammar like this: 92575897234Sdrh<pre> 92675897234Sdrh %token_prefix TOKEN_ 92775897234Sdrh</pre> 92875897234Sdrhto cause Lemon to produce these symbols instead: 92975897234Sdrh<pre> 93075897234Sdrh #define TOKEN_AND 1 93175897234Sdrh #define TOKEN_MINUS 2 93275897234Sdrh #define TOKEN_OR 3 93375897234Sdrh #define TOKEN_PLUS 4 93475897234Sdrh</pre> 93575897234Sdrh 9369bccde3dSdrh<a name='token_type'></a><a name='ptype'></a> 93775897234Sdrh<h4>The <tt>%token_type</tt> and <tt>%type</tt> directives</h4> 93875897234Sdrh 93975897234Sdrh<p>These directives are used to specify the data types for values 94075897234Sdrhon the parser's stack associated with terminal and non-terminal 94175897234Sdrhsymbols. The values of all terminal symbols must be of the same 94275897234Sdrhtype. This turns out to be the same data type as the 3rd parameter 94375897234Sdrhto the Parse() function generated by Lemon. Typically, you will 94475897234Sdrhmake the value of a terminal symbol by a pointer to some kind of 94575897234Sdrhtoken structure. Like this:</p> 94675897234Sdrh 94775897234Sdrh<p><pre> 94875897234Sdrh %token_type {Token*} 94975897234Sdrh</pre></p> 95075897234Sdrh 95175897234Sdrh<p>If the data type of terminals is not specified, the default value 952dfe4e6bbSdrhis "void*".</p> 95375897234Sdrh 95475897234Sdrh<p>Non-terminal symbols can each have their own data types. Typically 95575897234Sdrhthe data type of a non-terminal is a pointer to the root of a parse-tree 95675897234Sdrhstructure that contains all information about that non-terminal. 95775897234SdrhFor example:</p> 95875897234Sdrh 95975897234Sdrh<p><pre> 96075897234Sdrh %type expr {Expr*} 96175897234Sdrh</pre></p> 96275897234Sdrh 96375897234Sdrh<p>Each entry on the parser's stack is actually a union containing 96475897234Sdrhinstances of all data types for every non-terminal and terminal symbol. 96575897234SdrhLemon will automatically use the correct element of this union depending 96675897234Sdrhon what the corresponding non-terminal or terminal symbol is. But 96775897234Sdrhthe grammar designer should keep in mind that the size of the union 96875897234Sdrhwill be the size of its largest element. So if you have a single 96975897234Sdrhnon-terminal whose data type requires 1K of storage, then your 100 97075897234Sdrhentry parser stack will require 100K of heap space. If you are willing 97175897234Sdrhand able to pay that price, fine. You just need to know.</p> 97275897234Sdrh 9739bccde3dSdrh<a name='pwildcard'></a> 9749bccde3dSdrh<h4>The <tt>%wildcard</tt> directive</h4> 9759bccde3dSdrh 9769bccde3dSdrh<p>The %wildcard directive is followed by a single token name and a 9779bccde3dSdrhperiod. This directive specifies that the identified token should 9789bccde3dSdrhmatch any input token. 9799bccde3dSdrh 9809bccde3dSdrh<p>When the generated parser has the choice of matching an input against 9819bccde3dSdrhthe wildcard token and some other token, the other token is always used. 9829bccde3dSdrhThe wildcard token is only matched if there are no other alternatives. 9839bccde3dSdrh 98475897234Sdrh<h3>Error Processing</h3> 98575897234Sdrh 98675897234Sdrh<p>After extensive experimentation over several years, it has been 98775897234Sdrhdiscovered that the error recovery strategy used by yacc is about 98875897234Sdrhas good as it gets. And so that is what Lemon uses.</p> 98975897234Sdrh 99075897234Sdrh<p>When a Lemon-generated parser encounters a syntax error, it 99175897234Sdrhfirst invokes the code specified by the %syntax_error directive, if 99275897234Sdrhany. It then enters its error recovery strategy. The error recovery 99375897234Sdrhstrategy is to begin popping the parsers stack until it enters a 99475897234Sdrhstate where it is permitted to shift a special non-terminal symbol 9959bccde3dSdrhnamed "error". It then shifts this non-terminal and continues 99675897234Sdrhparsing. But the %syntax_error routine will not be called again 99775897234Sdrhuntil at least three new tokens have been successfully shifted.</p> 99875897234Sdrh 99975897234Sdrh<p>If the parser pops its stack until the stack is empty, and it still 100075897234Sdrhis unable to shift the error symbol, then the %parse_failed routine 100175897234Sdrhis invoked and the parser resets itself to its start state, ready 100275897234Sdrhto begin parsing a new file. This is what will happen at the very 100375897234Sdrhfirst syntax error, of course, if there are no instances of the 10049bccde3dSdrh"error" non-terminal in your grammar.</p> 100575897234Sdrh 100675897234Sdrh</body> 100775897234Sdrh</html> 1008