How to define a language module¶
This is a walk through writing a language pack for rainbow-fmt from
nothing to an installable plugin, using a small Scheme as the example.
Scheme is a good first language: its whole grammar is parentheses, atoms,
strings and ; comments, so nothing about parsing gets in the way of
learning how a pack is put together. The finished pack is in
rainbow-lang-scheme/ next to this page (about
250 lines, tests included) and is kept working by the main test suite.
The same steps apply to a language with a tree-sitter grammar; the differences are called out where they occur.
What a pack is¶
A pack gives rainbow-fmt three things: a parser that turns bytes into a
tree, rules that turn tree nodes into a layout description (a Doc),
and a Language record that ties them together with the file
extensions and the pack's options. Everything else — reading
configuration, printing the Doc to fit the width, verifying that nothing
but whitespace changed, caching, parallel runs, the command line — is
shared.
flowchart LR
src[/"source bytes"/] --> P["parser<br/>(yours)"]
P --> T["CST nodes"]
T --> R["rules<br/>(yours)"]
R --> D["Doc"]
D --> Pr["printer<br/>(shared)"]
Pr --> out[/"formatted text"/]
out --> V["verifier<br/>(shared)"]
src --> V
V -->|"same tree, same comments,<br/>stable"| ok([written])
V -->|"else"| err([reported, file unchanged])
Because a node with no rule is printed exactly as written, a pack is safe
to run on real files from the moment it parses: a half-finished pack
re-indents what it understands and leaves the rest alone
(extending.md, "Levels of support").
Step 1 — a parser¶
rainbow-fmt talks to parsers through two small protocols in
rainbow_fmt.core.parser: a Parser has a language name and
parse(bytes) -> CstNode, and a CstNode has a type, byte and
(row, column) ranges, children, and four flags (is_named,
is_extra for comments, is_error, is_missing). Whitespace is not in
the tree: it is whatever lies between two nodes' byte ranges, which is
how the formatter knows where the source had blank lines or none.
With a tree-sitter grammar, this step is one line:
TreeSitterParser("css", tree_sitter.Language(tree_sitter_css.language()))
(rainbow_fmt.core.treesitter). For Scheme no grammar is on PyPI, and
the language is small, so the pack has its own parser in
parser.py: a
frozen dataclass Node with the protocol's fields, and one loop over the
bytes with an explicit stack of open lists (a recursive descent would hit
Python's recursion limit a few hundred parentheses deep; the shared
pipeline runs rules with a large stack, but a parser runs where it is
called). It produces these node types:
| Type | Children | Example |
|---|---|---|
program |
top-level forms and comments | the file |
list |
(, items, ) |
(* x x) |
quote |
', the datum |
'(1 2) |
atom |
— | define, 3.14, #t, #\a |
string |
— | "hello" |
comment (extra) |
— | ; note |
graph TD
P[program] --> L1["list (define (square x) (* x x))"]
L1 --> o1["("]
L1 --> a1[atom define]
L1 --> L2["list (square x)"]
L1 --> L3["list (* x x)"]
L1 --> c1[")"]
L2 --> o2["("] --> a2[atom square]
L2 --> a3[atom x]
L2 --> c2[")"]
L3 --> o3["("] --> a4["atom *"]
L3 --> a5[atom x]
L3 --> a6[atom x]
L3 --> c3[")"]
Three details matter for the rest of the pipeline:
- Tokens that are not nodes. Parentheses and the quote mark are
nodes with
is_named = False. The verifier compares them like any other token, so the formatter cannot lose one. - Comments are extras.
is_extra = Truelets the shared comment attachment (rainbow_fmt.core.comments.attach_comments) decide whether a comment trails the item before it or leads the next one, by rows. - Errors are nodes. A stray
)becomes anERRORnode and an unclosed(gets a zero-width)withis_missing = True. The shared pipeline then leaves the file unchanged, or raisesSourceSyntaxErrorin strict mode, and the verifier refuses any output for it. The parser never raises.
Check the parser with the shared helpers before writing a rule:
>>> from rainbow_fmt.core.parser import has_errors, iter_nodes
>>> root = SchemeParser().parse(b"(define (square x) (* x x))")
>>> [n.type for n in iter_nodes(root)][:6]
['program', 'list', '(', 'atom', 'list', '(']
>>> has_errors(SchemeParser().parse(b"(f a b"))
True
Step 2 — the Language, and a pack that does nothing¶
from rainbow_fmt.languages.base import Language
from rainbow_fmt.languages.pack import format_with_rules
def format_scheme(source, options, *, strict=False):
return format_with_rules(LANGUAGE, source, options, RULES, HELPERS, _context, strict=strict)
LANGUAGE = Language("scheme", (".scm", ".ss"), format_scheme, OPTIONS, PARSER)
RULES = {}
format_with_rules (rainbow_fmt.languages.pack) is the shared driver:
it resolves the options, parses, runs the root rule with a large stack (a
rule recurses once per nesting level, so (((((…))))) five thousand deep
is fine), prints the Doc with the core options (max_width,
indent_size, line_ending …), keeps a byte order mark, and handles
syntax errors. With RULES empty every node is printed as written: this
is already a working, safe pack.
The _context function builds the object every rule receives besides its
node. A context needs doc(node) (format a child with its rule) and
text(node) (its source text); the Scheme one also carries the resolved
options in a small frozen dataclass, so rules read ctx.options.argument_alignment
rather than a mapping:
@dataclass(frozen=True)
class Context:
source: bytes
options: SchemeOptions
rules: Mapping[str, Rule["Context"]]
def doc(self, node):
return self.rules.get(node.type, source_text)(node, self)
def text(self, node):
return self.source[node.start_byte : node.end_byte].decode("utf-8")
Step 3 — golden fixtures first¶
Decide what the output should look like before writing rules, as small
input.scm/expected.scm pairs in fixtures/<nn>_<name>/, with an
options.toml where a case needs non-default options. The pack's
fixtures cover: forms that fit on one
line, special forms, the two argument alignments, comments, blank lines
and a syntax error. The shared harness rainbow_fmt.testing loads them
and diffs the output, and the pack's
tests/test_scheme.py runs
each case three ways — the output matches, formatting the output again
changes nothing, and the verifier accepts it:
@pytest.mark.parametrize("case", CASES, ids=lambda c: c.name)
def test_fixture_output_verifies(case):
options = resolve_options(case.options, LANGUAGE)
verify(LANGUAGE, case.source, format_scheme(case.source, options), options)
The third test is the one that catches real mistakes: a rule that drops a token, reorders a comment, or produces output that formats differently the second time fails here before anyone runs the pack on a file.
Step 4 — rules¶
A rule is (node, ctx) -> Doc. A Doc is not text: it is a description
of text with possible line breaks, and the shared printer picks the
breaks so the result fits the width (doc-ir.md).
The builders a Scheme pack needs:
| Builder | Meaning |
|---|---|
concat(*parts) |
parts in order |
group(*parts) |
on one line if it fits, otherwise every line inside breaks |
line |
a space, or a line break when the group breaks |
hardline |
always a line break |
align(n, *parts) |
lines inside start n columns further right |
join(sep, items) |
items with sep between them |
line_suffix(…) |
moved to the end of the line (trailing comments) |
The list rule is the whole formatter. A form that fits stays on one
line; one that does not breaks after its head, and the pack decides what
the following lines align with:
def list_(node, ctx):
children = node.children[1:-1] # between the parentheses
head, *arguments = [c for c in children if not c.is_extra]
name = ctx.text(head) if head.type == "atom" else None
docs = [ctx.doc(a) for a in arguments]
if name in SPECIAL_FORMS and len(arguments) > 1:
first, *body = docs # (define (f x)
rest = align(2, line, join(line, body)) # body…)
return group("(", ctx.doc(head), " ", first, rest, ")")
width = len(name) + 2 # "(" + head + " "
return group("(", ctx.doc(head), " ", align(width, join(line, docs)), ")")
(define (long-function-name first-argument second-argument)
(let ((a (compute first-argument)) (b (compute second-argument)))
(combine a b)))
(some-function-with-a-long-name (first argument)
(second argument)
(third argument)
(fourth argument))
Two things the shared layer does for the rule: the inner (let …) is a
group of its own, so it stays on one line when it fits even though the
define broke; and align counts in columns of the current line, so
nested forms line up under their own heads.
Comments. attach_comments(children, separators=()) returns each
member with its leading and trailing comments and the blank lines before
it; the program and list rules print leading comments on their own
line, trailing ones as a line_suffix, and keep blank lines up to
core.max_blank_lines:
(define (f x) ; trailing the head
; a comment on its own line inside the form
(* x 2)) ; trailing the form
What is never changed. Atoms and strings have no rule, so #\a,
3.14, "hello, world" are printed as written. This is the pack's
contract with the verifier: tokens equal, tree equal, only whitespace
differs. (A pack that does normalize tokens, such as JavaScript
semicolons, declares them on its Language — optional_tokens,
optional_trailing and friends in
extending.md.)
Step 5 — options¶
An option is declared once, on the Language, and shows up in
[language.scheme], --set language.scheme.…, rainbow-fmt options
and the rainbow: set directive without further work:
OPTIONS = (
Option(
"argument_alignment",
Choice("first", "indent"),
"first",
"Arguments of a broken form: aligned under the first argument, or indented by two.",
),
)
[language.scheme]
argument_alignment = "indent"
(some-function-with-a-long-name
(first argument)
(second argument)
(third argument)
(fourth argument))
Core options need nothing from the pack: max_width and line_ending
are applied by the printer, max_blank_lines is read by the rules through
the context.
Step 6 — install it¶
The pack's pyproject.toml names its Language in the
rainbow_fmt.languages entry-point group:
[project.entry-points."rainbow_fmt.languages"]
scheme = "rainbow_lang_scheme:LANGUAGE"
$ pip install -e docs/howto/define-language-module/rainbow-lang-scheme
$ rainbow-fmt diff program.scm
$ rainbow-fmt options program.scm
# program.scm (scheme)
core.max_width = 80 # default
…
language.scheme.argument_alignment = "first" # default
$ rainbow-fmt format src/ --set language.scheme.argument_alignment=indent
Built-in packs are found first, so a plugin cannot claim .py; a plugin
that fails to import stops the run with an error naming its entry point.
Where to go from here¶
- Directives (
; rainbow: off,skip-next,set): describe each level of members asSlots and letrainbow_fmt.core.directives.plandecide which are printed as written; the CSS pack is the shortest example. - Embedding: a pack that contains other languages (HTML, Svelte)
splices their Docs into its own through
rainbow_fmt.inject; every pack, this one included, exposesLanguage.docso it can be embedded. - A tree-sitter grammar: replace
SchemeParserwithTreeSitterParserand the node types inRULESwith the grammar's; everything else stays.python -c "import tree_sitter_x"plus a dump of a parse tree is the first thing to do (TASKS.md T20's "Grammar check" shows what to record). - A corpus: before calling a pack done, run
rainbow-fmt checkover a few thousand real files; the verifier turns every mistake into an error message rather than a damaged file.