Contributing

Testing strategy

Besides the unit tests of each package, four checks each guard one property of the formatter.

Identity

With every rule off, the tree gives the source back byte for byte.

Classification

The parser reads the first line of a block the way Asciidoctor does.

Render equivalence

No rule changes what the rendered document shows.

Golden files

A rule does what its cases say.

The checks run over two sets of cases. Golden cases are written for a rule, and a new or changed rule comes with them. The Asciidoctor cases are the test inputs of the Asciidoctor test suite, stored under testdata/asciidoctor-cases. Identity and render equivalence run over both sets, which makes it more likely that a rule’s mistake is caught even where no case was written for it.

To run the tests, follow the development setup.

Identity

Printing the block tree with every rule off reproduces the source byte for byte. Render equivalence cannot replace this check, because a lost blank line renders the same.

The printer has a raw mode for this check. Raw mode walks the tree the same way as formatting mode, but prints every node as it stands. Format never offers raw mode, because turning the rules off is a test tool, not a mode of the formatter.

TestPrintIsIdentity prints every Asciidoctor case and every golden case back and compares bytes. TestParsePartitions checks that the nodes cover the source without gaps or overlaps, so a failure names the node that lost or repeated bytes. FuzzPrintIsIdentity searches for source that the Asciidoctor cases do not cover. Its seeds are fixed inputs that run as ordinary cases with every go test. The search mutates the seeds and runs on demand:

go test -fuzz=FuzzPrintIsIdentity ./internal/printer

An input that the search finds breaking the test is saved under testdata/fuzz and runs from then on.

Classification

The parser decides what each block is, such as a heading, a list or a code block, and every rule builds on that decision. Asciidoctor decides what kind of block starts from the first line of the block, and so does the parser. TestClassifyAgreesWithAsciidoctor compares that decision line by line.

The test collects every distinct line of the Asciidoctor cases. It also edits each line in a few ways that can move it from one kind to another, such as a leading // or a tab instead of a space. adocfmt’s parser reads each line as the first line of a block. So does line_kinds.rb, which asks Asciidoctor’s own parser what block the line starts. The test fails on every line where the two disagree, and on any Asciidoctor release other than the pinned one.

So a misread line shows up before any rule rewrites its block, not only once a rule does and the rendering changes.

What depends on more than one line is out of the test’s reach, such as where a list ends or whether two lines form a title. A wrong decision there does harm only where a rule rewrites the misread block, and there the render equivalence checks fail. A misread block that no rule rewrites cannot change the rendered document.

Tip
When you fix a misread block, add a case to parse_test.go that pins the tree the parser now produces.

Render equivalence

Formatting is correct when it does not change what the reader of the rendered document sees. The rendering that counts is Asciidoctor’s, the reference implementation of AsciiDoc. The checks render both sides and compare three things.

Rendering

Both HTML renderings, with every run of whitespace collapsed to a single space. Where a paragraph wraps is invisible in the rendered page, because HTML shows any run of whitespace as one space.

Verbatim content

The content of every code and literal block, byte for byte. Whitespace carries meaning there. Because this check compares it, the rendering check can ignore whitespace everywhere else. Asciidoctor strips trailing whitespace from those lines before it renders. So this check cannot see a rule that changes trailing whitespace there.

Comments

The comment lines of both sources, in order, because Asciidoctor drops them before they reach the HTML. The check reads raw lines, not the block tree. So a comment that the parser took for prose still counts.

TestChecksCatchCorruption feeds each check a document broken in the way the check is meant to catch, to prove that the check can fail at all.

Where the checks run

The Asciidoctor cases are the main input, because breadth gives confidence. There is one file per test input of the Asciidoctor test suite. Those inputs sit inside Ruby test files, so tools/fetch-asciidoctor-cases extracts them. corpus.go pins the Asciidoctor release that the tool extracts from, which is also the release that the checks run with. After you change it, regenerate the cases:

go run ./tools/fetch-asciidoctor-cases

The parser mirrors some of Asciidoctor’s patterns and constants, and its comments name each one, such as BlockTitleRx or PARAGRAPH_STYLES. Compare them between the two releases in lib/asciidoctor/rx.rb and lib/asciidoctor.rb, because the classification check sees only the lines that the cases hold.

TestAsciidoctorCases formats every case and compares the result with the original. TestGoldenCasesRenderEquivalent does the same for every golden case, comparing the input with the expectation.

Some Asciidoctor cases are documents that adocfmt refuses, as decision record 0003 describes. The refused list in cases_test.go pins which ones, grouped by cause. If a case on that list formats after all, TestAsciidoctorCases fails.

The Asciidoctor cases have two blind spots: they hold little prose, and each conditional in them is rendered in one reading only. Golden cases cover both with expectations written for them.

Golden files

Render equivalence shows that a rule does no harm; a golden case shows that it does what it should. A case is a directory under testdata/golden that holds two files:

testdata/golden/
  <case>/
    input.adoc
    golden.adoc
  <group>/<case>/
    input.adoc
    golden.adoc

input.adoc goes through the formatter, and the result must equal golden.adoc byte for byte. Directories can nest, so cases can be grouped by rule. The path below testdata/golden is the test name.

The cases run with every rule on, so cut the input until only the rule under test has anything to do. The exception is website/home-example, the before and after example on the home page of the website, which shows several rules at once.

To add a case, write input.adoc, then let the formatter write the expectation:

go run ./tools/update-golden
Important
Read the diff before you commit it. The tool records whatever the formatter produces, so an unreviewed update turns a bug into the expectation.

The render checks are blind to some changes, such as trailing whitespace inside verbatim blocks or comments. For those, write golden.adoc by hand.

Every rule also needs a case whose input is already in the rule’s target state, so input.adoc and golden.adoc are identical. That case catches a rule that keeps reformatting its own output.

Git stores every file in the repository with LF line endings. A case that tests CRLF line endings therefore needs a line in .gitattributes that exempts its files, such as testdata/golden/list-markers/crlf/* -text.

View the source on GitHub