Class: PGN::Lexer
- Inherits:
-
Object
- Object
- PGN::Lexer
- Defined in:
- lib/pgn/lexer.rb
Overview
Lexer is a StringScanner-based tokenizer for PGN. It reuses the same terminal patterns as the legacy whittle parser so tokenization is byte-compatible, but performs far fewer Ruby allocations because the scanning happens in the C-backed StringScanner.
The lexer also records, per game, the byte offset of the game's first non-discarded token (#game_starts). The parser uses these offsets to slice the verbatim Game#pgn raw text out of the original input, reproducing the legacy accumulator's output exactly.
All offsets are byte offsets (StringScanner works in bytes); slicing is done with String#byteslice so multibyte (UTF-8) input stays intact.
Defined Under Namespace
Classes: Token
Constant Summary collapse
- WSP =
Discarded: insignificant whitespace.
/\s+/- PGN_COMMENT =
Discarded: a PGN "rest of line" comment beginning with
%. /% .*/- STRING =
A tag value string. Allows unescaped double-quotes inside the value (a form seen in real-world PGN files) — only a bare backslash starts an escape. Matches the legacy parser's relaxed string rule.
/ " # beginning of string ( [[:print:]&&[^\\]] | # printing characters except backslash \\\\ | # escaped backslashes \\" # escaped quotation marks )* # zero or more of the above " # end of string /x- COMMENT =
A brace-delimited comment, with recursive nesting via \g<1>.
/ ( \{ # beginning of comment ( [[:print:]&&[^\\{}]] | # printing characters except brace and backslash \n | \\\\ | # escaped backslashes \\\{|\\\} | # escaped braces \n | # newlines \g<1> # recursive )* # zero or more of the above \} # end of comment ) /x- GAME_TERMINATION =
Game termination marker.
%r{ 1-0 | # white wins 0-1 | # black wins 1/2-1/2 | # draw \* # ? }x- SAN_MOVE =
A move in standard algebraic notation (incl. castling, promotion, check/mate, the
--"don't care" move). %r{ ( -- | # "don't care" move (used in variations) [O0](-[O0]){1,2} | # castling (O-O, O-O-O) [a-h][1-8] | # pawn moves (e4, d7) [BKNQR][a-h1-8]?x?[a-h][1-8] | # major piece moves w/ optional specifier [a-h][1-8]?x[a-h][1-8] # pawn captures ) ( =[BNQR] # optional promotion (d8=Q) )? ( \+ | # check (g5+) \# # checkmate (Qe7#) )? }x- MOVE_NUMBER =
A move number indication, e.g.
1.,12.,1.... /[[:digit:]]+\.*/- TAG_NAME =
A tag name (letters, digits, underscores).
/[A-Za-z0-9_]+/- NAG =
A numeric annotation glyph (
$1) or a punctuation annotation (?!,!?,??, ...). / \$\d+ | # dollar sign followed by an integer [?!][?!]? # support the most used annotations directly /x- RULES =
Order matters: more specific / longer tokens are tried first so that e.g.
1-0(termination) wins over1(move number), and0-0(castling) wins over0(move number). Whitespace and%comments are discarded (consumed but not emitted). Beyond that constraint, rules are ordered most- to least-frequent (one san_move/move_number per ply/full-move vs. a handful of comments/strings per game) so the common case fails the fewest regexes before matching. [ [:wsp, WSP, true], # discarded [:pgn_comment, PGN_COMMENT, true], # discarded [:game_termination, GAME_TERMINATION, false], [:san_move, SAN_MOVE, false], [:move_number, MOVE_NUMBER, false], [:nag, NAG, false], [:comment, COMMENT, false], [:string, STRING, false], [:tag_name, TAG_NAME, false] ].freeze.each(&:freeze)
- BYTE_DISPATCH =
Byte-dispatch table: maps the leading byte of the next token to the ordered list of RULES indices that could possibly match it. This lets #scan_one try one or two regexes for the common tokens instead of walking all nine rules in order (StringScanner#scan was ~23% of parse CPU and the rule loop ~38% inclusive per profiling). The order within each list mirrors RULES, so tokenization is byte-compatible with the linear scan. Indices: 0 wsp, 1 pgn_comment, 2 game_termination, 3 san_move, 4 move_number, 5 nag, 6 comment, 7 string, 8 tag_name.
begin h = { 9 => [0], 10 => [0], 11 => [0], 12 => [0], 13 => [0], 32 => [0], # whitespace 37 => [1], # % pgn_comment 42 => [2], # * game_termination 34 => [7], # " string 123 => [6], # { comment 36 => [5], 63 => [5], 33 => [5], # $ ? ! nag 48 => [2, 3, 4, 8], 49 => [2, 4, 8], # 0, 1 (term/castle/num/tag) 95 => [8] # _ tag_name } (50..57).each { |b| h[b] = [4, 8] } # 2..9 move_number/tag_name [66, 75, 78, 79, 81, 82].each { |b| h[b] = [3, 8] } # B K N O Q R san_move/tag (97..104).each { |b| h[b] = [3, 8] } # a-h pawn san_move/tag # other tag_name letters ((65..90).to_a + (105..122).to_a - [66, 75, 78, 79, 81, 82]).each do |b| h[b] = [8] end h.each_value(&:freeze) h.freeze end
- ALL_RULES =
Fallback for bytes not in BYTE_DISPATCH: try every rule in order.
(0...RULES.length).to_a.freeze
- LITERAL_BYTES =
Single-character literals, matched by their byte value: [type, frozen value].
{ 91 => [:lbracket, '['], # [ 93 => [:rbracket, ']'], # ] 40 => [:lparen, '('], # ( 41 => [:rparen, ')'] # ) }.freeze
Instance Attribute Summary collapse
-
#game_starts ⇒ Object
readonly
Returns the value of attribute game_starts.
Instance Method Summary collapse
-
#initialize(input) ⇒ Lexer
constructor
A new instance of Lexer.
-
#next_token ⇒ Object
Returns the next Token, or
nilat end of input. -
#next_token_pair ⇒ Object
Fast path for the parser: returns [type, value] for the next non-discarded token, or
nilat end of input. -
#tokens ⇒ Object
The list of Tokens for the whole input.
Constructor Details
#initialize(input) ⇒ Lexer
Returns a new instance of Lexer.
161 162 163 164 165 166 167 |
# File 'lib/pgn/lexer.rb', line 161 def initialize(input) @input = input @ss = StringScanner.new(input) @line = 1 @game_starts = [] @between_games = true # at start we are "between" games end |
Instance Attribute Details
#game_starts ⇒ Object (readonly)
Returns the value of attribute game_starts.
169 170 171 |
# File 'lib/pgn/lexer.rb', line 169 def game_starts @game_starts end |
Instance Method Details
#next_token ⇒ Object
Returns the next Token, or nil at end of input.
181 182 183 184 185 186 |
# File 'lib/pgn/lexer.rb', line 181 def next_token type, value = next_token_pair return nil unless type Token.new(type: type, value: value, offset: @last_offset, line: @line) end |
#next_token_pair ⇒ Object
Fast path for the parser: returns [type, value] for the next
non-discarded token, or nil at end of input. Does not allocate a
Token Struct, and scan_one returns the matched string directly
(stashing its type/discarded flag in ivars) so the only array
allocated per token is the [type, value] pair Racc requires.
193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 |
# File 'lib/pgn/lexer.rb', line 193 def next_token_pair until @ss.eos? off = @ss.pos if (lit = LITERAL_BYTES[@input.getbyte(off)]) type, value = lit @ss.pos = off + 1 note_token(type, off) @last_offset = off return [type, value] end value = scan_one advance_line(value) next if @scan_discarded note_token(@scan_type, off) @last_offset = off return [@scan_type, value] end nil end |
#tokens ⇒ Object
The list of Tokens for the whole input. Convenience for specs.
172 173 174 175 176 177 178 |
# File 'lib/pgn/lexer.rb', line 172 def tokens result = [] while (t = next_token) result << t end result end |