Skip to content

How to define a language module

This is a walk through writing a language pack for rainbow-fmt from nothing to an installable plugin, using a small Scheme as the example. Scheme is a good first language: its whole grammar is parentheses, atoms, strings and ; comments, so nothing about parsing gets in the way of learning how a pack is put together. The finished pack is in rainbow-lang-scheme/ next to this page (about 250 lines, tests included) and is kept working by the main test suite.

The same steps apply to a language with a tree-sitter grammar; the differences are called out where they occur.

What a pack is

A pack gives rainbow-fmt three things: a parser that turns bytes into a tree, rules that turn tree nodes into a layout description (a Doc), and a Language record that ties them together with the file extensions and the pack's options. Everything else — reading configuration, printing the Doc to fit the width, verifying that nothing but whitespace changed, caching, parallel runs, the command line — is shared.

flowchart LR
    src[/"source bytes"/] --> P["parser<br/>(yours)"]
    P --> T["CST nodes"]
    T --> R["rules<br/>(yours)"]
    R --> D["Doc"]
    D --> Pr["printer<br/>(shared)"]
    Pr --> out[/"formatted text"/]
    out --> V["verifier<br/>(shared)"]
    src --> V
    V -->|"same tree, same comments,<br/>stable"| ok([written])
    V -->|"else"| err([reported, file unchanged])

Because a node with no rule is printed exactly as written, a pack is safe to run on real files from the moment it parses: a half-finished pack re-indents what it understands and leaves the rest alone (extending.md, "Levels of support").

Step 1 — a parser

rainbow-fmt talks to parsers through two small protocols in rainbow_fmt.core.parser: a Parser has a language name and parse(bytes) -> CstNode, and a CstNode has a type, byte and (row, column) ranges, children, and four flags (is_named, is_extra for comments, is_error, is_missing). Whitespace is not in the tree: it is whatever lies between two nodes' byte ranges, which is how the formatter knows where the source had blank lines or none.

With a tree-sitter grammar, this step is one line: TreeSitterParser("css", tree_sitter.Language(tree_sitter_css.language())) (rainbow_fmt.core.treesitter). For Scheme no grammar is on PyPI, and the language is small, so the pack has its own parser in parser.py: a frozen dataclass Node with the protocol's fields, and one loop over the bytes with an explicit stack of open lists (a recursive descent would hit Python's recursion limit a few hundred parentheses deep; the shared pipeline runs rules with a large stack, but a parser runs where it is called). It produces these node types:

Type Children Example
program top-level forms and comments the file
list (, items, ) (* x x)
quote ', the datum '(1 2)
atom — define, 3.14, #t, #\a
string — "hello"
comment (extra) — ; note
graph TD
    P[program] --> L1["list  (define (square x) (* x x))"]
    L1 --> o1["("]
    L1 --> a1[atom define]
    L1 --> L2["list  (square x)"]
    L1 --> L3["list  (* x x)"]
    L1 --> c1[")"]
    L2 --> o2["("] --> a2[atom square]
    L2 --> a3[atom x]
    L2 --> c2[")"]
    L3 --> o3["("] --> a4["atom *"]
    L3 --> a5[atom x]
    L3 --> a6[atom x]
    L3 --> c3[")"]

Three details matter for the rest of the pipeline:

  • Tokens that are not nodes. Parentheses and the quote mark are nodes with is_named = False. The verifier compares them like any other token, so the formatter cannot lose one.
  • Comments are extras. is_extra = True lets the shared comment attachment (rainbow_fmt.core.comments.attach_comments) decide whether a comment trails the item before it or leads the next one, by rows.
  • Errors are nodes. A stray ) becomes an ERROR node and an unclosed ( gets a zero-width ) with is_missing = True. The shared pipeline then leaves the file unchanged, or raises SourceSyntaxError in strict mode, and the verifier refuses any output for it. The parser never raises.

Check the parser with the shared helpers before writing a rule:

>>> from rainbow_fmt.core.parser import has_errors, iter_nodes
>>> root = SchemeParser().parse(b"(define (square x) (* x x))")
>>> [n.type for n in iter_nodes(root)][:6]
['program', 'list', '(', 'atom', 'list', '(']
>>> has_errors(SchemeParser().parse(b"(f a b"))
True

Step 2 — the Language, and a pack that does nothing

from rainbow_fmt.languages.base import Language
from rainbow_fmt.languages.pack import format_with_rules

def format_scheme(source, options, *, strict=False):
    return format_with_rules(LANGUAGE, source, options, RULES, HELPERS, _context, strict=strict)

LANGUAGE = Language("scheme", (".scm", ".ss"), format_scheme, OPTIONS, PARSER)
RULES = {}

format_with_rules (rainbow_fmt.languages.pack) is the shared driver: it resolves the options, parses, runs the root rule with a large stack (a rule recurses once per nesting level, so (((((…))))) five thousand deep is fine), prints the Doc with the core options (max_width, indent_size, line_ending …), keeps a byte order mark, and handles syntax errors. With RULES empty every node is printed as written: this is already a working, safe pack.

The _context function builds the object every rule receives besides its node. A context needs doc(node) (format a child with its rule) and text(node) (its source text); the Scheme one also carries the resolved options in a small frozen dataclass, so rules read ctx.options.argument_alignment rather than a mapping:

@dataclass(frozen=True)
class Context:
    source: bytes
    options: SchemeOptions
    rules: Mapping[str, Rule["Context"]]

    def doc(self, node):
        return self.rules.get(node.type, source_text)(node, self)

    def text(self, node):
        return self.source[node.start_byte : node.end_byte].decode("utf-8")

Step 3 — golden fixtures first

Decide what the output should look like before writing rules, as small input.scm/expected.scm pairs in fixtures/<nn>_<name>/, with an options.toml where a case needs non-default options. The pack's fixtures cover: forms that fit on one line, special forms, the two argument alignments, comments, blank lines and a syntax error. The shared harness rainbow_fmt.testing loads them and diffs the output, and the pack's tests/test_scheme.py runs each case three ways — the output matches, formatting the output again changes nothing, and the verifier accepts it:

@pytest.mark.parametrize("case", CASES, ids=lambda c: c.name)
def test_fixture_output_verifies(case):
    options = resolve_options(case.options, LANGUAGE)
    verify(LANGUAGE, case.source, format_scheme(case.source, options), options)

The third test is the one that catches real mistakes: a rule that drops a token, reorders a comment, or produces output that formats differently the second time fails here before anyone runs the pack on a file.

Step 4 — rules

A rule is (node, ctx) -> Doc. A Doc is not text: it is a description of text with possible line breaks, and the shared printer picks the breaks so the result fits the width (doc-ir.md). The builders a Scheme pack needs:

Builder Meaning
concat(*parts) parts in order
group(*parts) on one line if it fits, otherwise every line inside breaks
line a space, or a line break when the group breaks
hardline always a line break
align(n, *parts) lines inside start n columns further right
join(sep, items) items with sep between them
line_suffix(…) moved to the end of the line (trailing comments)

The list rule is the whole formatter. A form that fits stays on one line; one that does not breaks after its head, and the pack decides what the following lines align with:

def list_(node, ctx):
    children = node.children[1:-1]            # between the parentheses
    head, *arguments = [c for c in children if not c.is_extra]
    name = ctx.text(head) if head.type == "atom" else None
    docs = [ctx.doc(a) for a in arguments]
    if name in SPECIAL_FORMS and len(arguments) > 1:
        first, *body = docs                   # (define (f x)
        rest = align(2, line, join(line, body))   #   body…)
        return group("(", ctx.doc(head), " ", first, rest, ")")
    width = len(name) + 2                     # "(" + head + " "
    return group("(", ctx.doc(head), " ", align(width, join(line, docs)), ")")
(define (long-function-name first-argument second-argument)
  (let ((a (compute first-argument)) (b (compute second-argument)))
    (combine a b)))
(some-function-with-a-long-name (first argument)
                                (second argument)
                                (third argument)
                                (fourth argument))

Two things the shared layer does for the rule: the inner (let …) is a group of its own, so it stays on one line when it fits even though the define broke; and align counts in columns of the current line, so nested forms line up under their own heads.

Comments. attach_comments(children, separators=()) returns each member with its leading and trailing comments and the blank lines before it; the program and list rules print leading comments on their own line, trailing ones as a line_suffix, and keep blank lines up to core.max_blank_lines:

(define (f x) ; trailing the head
  ; a comment on its own line inside the form
  (* x 2)) ; trailing the form

What is never changed. Atoms and strings have no rule, so #\a, 3.14, "hello, world" are printed as written. This is the pack's contract with the verifier: tokens equal, tree equal, only whitespace differs. (A pack that does normalize tokens, such as JavaScript semicolons, declares them on its Language — optional_tokens, optional_trailing and friends in extending.md.)

Step 5 — options

An option is declared once, on the Language, and shows up in [language.scheme], --set language.scheme.…, rainbow-fmt options and the rainbow: set directive without further work:

OPTIONS = (
    Option(
        "argument_alignment",
        Choice("first", "indent"),
        "first",
        "Arguments of a broken form: aligned under the first argument, or indented by two.",
    ),
)
[language.scheme]
argument_alignment = "indent"
(some-function-with-a-long-name
  (first argument)
  (second argument)
  (third argument)
  (fourth argument))

Core options need nothing from the pack: max_width and line_ending are applied by the printer, max_blank_lines is read by the rules through the context.

Step 6 — install it

The pack's pyproject.toml names its Language in the rainbow_fmt.languages entry-point group:

[project.entry-points."rainbow_fmt.languages"]
scheme = "rainbow_lang_scheme:LANGUAGE"
$ pip install -e docs/howto/define-language-module/rainbow-lang-scheme
$ rainbow-fmt diff program.scm
$ rainbow-fmt options program.scm
# program.scm (scheme)
core.max_width = 80                     # default
…
language.scheme.argument_alignment = "first" # default
$ rainbow-fmt format src/ --set language.scheme.argument_alignment=indent

Built-in packs are found first, so a plugin cannot claim .py; a plugin that fails to import stops the run with an error naming its entry point.

Where to go from here

  • Directives (; rainbow: off, skip-next, set): describe each level of members as Slots and let rainbow_fmt.core.directives.plan decide which are printed as written; the CSS pack is the shortest example.
  • Embedding: a pack that contains other languages (HTML, Svelte) splices their Docs into its own through rainbow_fmt.inject; every pack, this one included, exposes Language.doc so it can be embedded.
  • A tree-sitter grammar: replace SchemeParser with TreeSitterParser and the node types in RULES with the grammar's; everything else stays. python -c "import tree_sitter_x" plus a dump of a parse tree is the first thing to do (TASKS.md T20's "Grammar check" shows what to record).
  • A corpus: before calling a pack done, run rainbow-fmt check over a few thousand real files; the verifier turns every mistake into an error message rather than a damaged file.