手写 tokenizer:把源码切成 token 流

本节目标

解析的第一步是词法分析(Lexical Analysis):把一长串字符,切成有语义的“词”(token)。

例如源码 let total = price * 3 会被切成:

[let, total, =, price, *, 3]
 关键字 标识符  运算符 标识符 运算符 数字

每个 token 长这样:{ type: 'ident', value: 'total' }。统一形状很关键——后面的 Parser 只认 type,不关心原文。

为什么用手写状态机,而不是一个正则搞定? 正则适合简单语言;但 JS 这种有字符串、注释、多字符运算符(===、&&)的语言,状态机更可控:逐字符读,记住“我现在在攒一个什么类型的词”,遇到边界就把攒下的内容吐出来。

动手:一个能跑的 tokenizer

下面实现:跳过空白与 // 注释;识别数字、标识符/关键字、运算符、()、字符串。

// 运行环境:Node.js 18+(无需任何依赖)
// 运行方式:把整段保存为 app-l3.cjs,终端执行  node app-l3.cjs

// 词法分析器:逐字符扫描,按“当前在攒什么”产出 token
function tokenize(src) {
  const tokens = []
  let i = 0, buf = '', state = 'idle'
  const isDigit = (c) => c >= '0' && c <= '9'
  const isIdentStart = (c) => /[A-Za-z_$]/.test(c)
  const isIdentPart = (c) => /[A-Za-z0-9_$]/.test(c)
  const isSpace = (c) => c === ' ' || c === '\n' || c === '\t' || c === '\r'

  // 把当前累积的 buf 作为一个 token 吐出(type 由 state 决定)
  const flush = (type) => {
    if (buf) { tokens.push({ type: type, value: buf }); buf = '' }
  }

  while (i < src.length) {
    const c = src[i]
    // 碰到 // 注释:一直吃到行尾
    if (c === '/' && src[i + 1] === '/') {
      flush(state === 'idle' ? null : state)
      while (i < src.length && src[i] !== '\n') i++
      state = 'idle'
      continue
    }
    if (isSpace(c)) { flush(state === 'idle' ? null : state); state = 'idle'; i++; continue }
    if (c === '"' || c === "'") {           // 字符串:读到配对的引号结束
      flush(state === 'idle' ? null : state)
      const quote = c; i++
      let str = ''
      while (i < src.length && src[i] !== quote) { str += src[i]; i++ }
      i++ // 跳过结束引号
      tokens.push({ type: 'string', value: str })
      state = 'idle'
      continue
    }
    if (isDigit(c)) { state = 'num'; buf += c; i++; continue }
    if (isIdentStart(c)) { state = 'ident'; buf += c; i++; continue }
    if (isIdentPart(c) && state !== 'idle') { buf += c; i++; continue }
    // 多字符运算符(== === && || <= >= != !==)优先合并
    const two = src.slice(i, i + 2)
    const multiOps = ['==', '===', '&&', '||', '<=', '>=', '!=', '!==']
    if (multiOps.includes(two)) { flush(state === 'idle' ? null : state); tokens.push({ type: 'op', value: two }); state = 'idle'; i += 2; continue }
    // 单字符符号
    if ('+-*/%=<>!'.includes(c)) { flush(state === 'idle' ? null : state); tokens.push({ type: 'op', value: c }); state = 'idle'; i++; continue }
    if (c === '(' || c === ')') { flush(state === 'idle' ? null : state); tokens.push({ type: 'paren', value: c }); state = 'idle'; i++; continue }
    // 兜底:未知字符直接当 op
    flush(state === 'idle' ? null : state); tokens.push({ type: 'op', value: c }); state = 'idle'; i++; continue
  }
  flush(state === 'idle' ? null : state)

  // 把 ident 里的关键字单独标记,方便后面 Parser 识别
  const KEYWORDS = ['let', 'const', 'var', 'function', 'return', 'if', 'else']
  return tokens.map((t) =>
    t.type === 'ident' && KEYWORDS.includes(t.value) ? { type: 'keyword', value: t.value } : t
  )
}

// === 调用示例 ===
const code = 'let total = price * 3 + 1 // 算总价'
const toks = tokenize(code)
console.log(JSON.stringify(toks, null, 0))
// 输出:
// [{"type":"keyword","value":"let"},{"type":"ident","value":"total"},
//  {"type":"op","value":"="},{"type":"ident","value":"price"},
//  {"type":"op","value":"*"},{"type":"num","value":"3"},
//  {"type":"op","value":"+"},{"type":"num","value":"1"}]

名词解释

token(词法单元):源码被“切词”后得到的最小“有语义的词”,统一形状是 { type, value }。比如 let total = price * 3 会被切成 let、total、=、price、*、3 共 6 个 token。它是“字符流”和“语法树”之间的中间产品——先有 token,Parser 才方便拼树。

状态机(State Machine):一个“记住自己当前处于什么状态、根据新输入决定下一步”的模型。本节 tokenizer 就是状态机:逐字符扫描,用 state 变量记住“我正在攒一个什么类型的词”,遇到边界就把攒下的内容吐成一个 token、再回到空闲态。

课后练习

练习 1:源码 const n = 10 经过本节的 tokenize 后,会按顺序得到哪些 token?(写出每个 token 的 type 与 value)

答案:{type:'keyword', value:'const'}、{type:'ident', value:'n'}、{type:'op', value:'='}、{type:'num', value:'10'}。注意 10 被识别成数字、不是标识符,因为 tokenizer 先判断 isDigit。

练习 2:为什么多字符运算符 === 必须“优先合并”成一个 token,而不能被拆成两个 = token?

答案:如果拆成两个 = token,后面的 Parser 会误以为出现了两次赋值(a == b 被理解成 a = = b),语法和语义都错。合并成单个 === token 才能正确表达“严格相等”这一个运算符。

本节小结(观点与完整描述)

本节你亲手写的不是一个“玩具函数”,而是所有解析器的第一步真身:词法分析 = 把字符流切成有类型的词。你学到的关键不是那几十行代码,而是两个认知:第一,token 形状统一为 { type, value },后面的 Parser 只认 type,不关心原文长什么样;第二,状态机的价值在于“可控”——要支持字符串转义、模板字符串、注释,只要多添几个状态即可,而一个超大正则几乎无法扩展。真实项目里你当然不需要手写 tokenizer(acorn / @babel/parser 已经写好),但亲手写一遍,你会彻底看懂“源码到底是怎么被读懂的”,这层理解在调试语法报错时非常值钱。