Skip to content

packt-prose CLI reference

The complete command surface, exit-code contract, and JSON schemas of the packt-prose binary. This page is for engineers and agents calling the tool; editors want the user guide instead.

The binary contains no model and makes no network calls. Every command is a pure function of its inputs and the embedded house style, so the same inputs always produce the same output, byte for byte.

Invocation

packt-prose <command> [flags]

Flags always follow the command name. There are no true global flags: --json, --profile, and --fail-on are per-command flags that several commands share, so a flag before the command name is a usage error. The exceptions are -h, --help, and help, which print the command list, and -v/--version, which is a synonym for packt-prose version. packt-prose <command> -h prints one command's flags. The Go flag parser accepts one or two dashes interchangeably: -json and --json are the same flag.

The book decisions file

A file named packt-book.yml in the same directory as the input document is discovered automatically by every command that loads rules, the way the auto profile is: it is a fact about the manuscript's home, not something to retype per command. It records the choices a single chapter cannot see — the editor's style sheet, machine-readable:

schema: packt-prose/book@1
decisions:
  caption-separator: endash   # or colon
  data-verb: singular         # or plural
  methodology-casing: lowercase  # or capitalized

With a decision recorded, the matching consistency rule (packt.mech.caption-separator, cmos.consist.data-agreement, cmos.consist.methodology-casing) stops counting the chapter's own majority and holds every chapter to the book's choice — including a chapter that is internally consistent the other way, which is the point. Unknown keys and invalid values are errors, so a typo cannot silently record nothing.

Input and output conventions:

  • - means standard input for lint -i, for --chapter, and for --edits. In every case the content read from stdin is JSON — chapter JSON or an edit list. A .docx must be a real path, so extract -i -, autofix -i -, and apply -i - all fail with exit 3.
  • -o - (and, for extract and corpus baseline, an omitted -o) writes to standard output. .docx outputs require a real path.
  • JSON goes to stdout; diagnostics, notes, and errors go to stderr. A command that writes JSON writes nothing else to stdout.
  • An output .docx must not be the input .docx. The input file is never modified.
  • On a command that writes a document (apply, accept, autofix, styles apply), -o names the document and may be given once; repeating it is a usage error, because the last value used to win silently and two editors wrote their Word document over their report's path with exit 0. Those commands take --report <file> for the JSON report instead.
  • Reads classify a mis-styled paragraph as what it is — the same decision the style pass writes — so judging a chapter before the write agrees with the document after it. --raw on extract, lint, and review shows the document's own labels, for inspection only.

Offsets are rune offsets into normalized text

This is the one thing to get right, and the most common interop mistake.

Every span and protected range in every schema is:

  • a rune offset — a count of Unicode code points, not bytes and not UTF-16 units;
  • half-open, [start, end), so end - start is the length in runes;
  • measured into that paragraph's normalized text, not into the original Word document.

Normalized means the paragraph has been flattened to one line, every run of whitespace collapsed to a single space, and both ends trimmed. The text field of the chapter JSON is the exact string those offsets index, so the safe way to resolve a span is always to slice the text you were given, as runes:

runes := []rune(paragraph.Text)
found := string(runes[span[0]:span[1]])

In Python, list(text)[start:end]; in JavaScript, Array.from(text).slice(start, end) — plain String.prototype.slice is wrong for any paragraph containing a character outside the Basic Multilingual Plane.

An edit's find string is matched against the same normalized text, so a find copied out of a .docx with a double space or a line break in it will not match.

Exit codes

Code Meaning
0 Clean run with nothing to report.
1 Ran successfully, but reported findings or refused something.
2 The invocation was wrong.
3 I/O failure or internal error.

The distinction that matters to a caller is 1 against 2 and 3: exit 1 is an answer, exit 2 and exit 3 mean no answer was produced.

What each command returns:

Command 0 1 2 3
extract always, on success never no -i, bad --format cannot read or parse the input
styles report nothing to re-style paragraphs to re-style, or ones needing a person bad or missing subcommand cannot read the input
styles apply the document was written never no -i, no -o, -o equals -i input has tracked changes, or I/O failure
review no blocker, and no finding at or above --fail-on a blocker, or a finding at or above --fail-on no -i, bad --fail-on cannot read the input
comments always, on success never unknown flag cannot read the input
show always, on success never no -i, no --pid, a pid outside the chapter cannot read the input
lint no finding at or above --fail-on any finding at or above --fail-on no -i, bad --format, bad --fail-on, --pid not in the chapter cannot read the input, or the profile does not exist
autofix fixes written, or nothing to fix --dry-run found fixes no -i, no -o without --dry-run, -o equals -i input has tracked changes, or I/O failure
check-edit every edit passed any edit failed a gate missing --chapter or --edits cannot read either input
apply every edit applied one or more edits refused missing -i, -o, or --edits, or -o equals -i input has tracked changes, or I/O failure
voice no metric flagged one or more metrics flagged neither --before/--after nor --chapter/--edits cannot read an input
rules always, on success never bad --format, or --id names no enabled rule the profile does not exist
corpus always, on success never missing or unknown subcommand cannot read the corpus file
doctor the install is healthy the embedded house style failed to load unknown flag never
version always never unknown flag never

apply still writes its output document when edits are refused; exit 1 means "read the report", not "no file was produced".

Commands

extract

Reads a .docx into the canonical chapter JSON that every other command consumes. Text is read with tracked changes accepted; a document that carries revisions is extracted anyway, with a note on stderr.

packt-prose extract -i chapter.docx -o chapter.chapter.json
Flag Default Effect
-i <path> — Input .docx. Required.
-o <path> - Output file, or - for stdout.
--format json\|text json json emits the chapter schema; text emits the one-line-per-paragraph view.
--json false Synonym for --format json, so "always pass --json" holds here too.
--raw false Tag paragraphs from their own styles, without classifying mis-styled ones. Inspection only.

The text format is the compact form a language model reads: [pid|TAG] followed by the normalized paragraph text, with empty paragraphs omitted.

[0|H1] Tuning the storage layer
[1|P] By default the server reserves 32GB of RAM for the cache, and IO throughput is the limit you will hit first.

accept

Accepts every tracked change a document already carries — the author's own revision history — and writes the resolved document. Raw drafts routinely arrive mid-revision, and every writing command refuses such a file; this is the tool-native way past that refusal when accepting all of the changes is what you want. When it is not — when some of those revisions should be rejected — resolve them selectively in Word instead.

packt-prose accept -i chapter.docx -o chapter_accepted.docx
Flag Default Effect
-i <path> — Input .docx. Required.
-o <path> — Output .docx. Required, never the input, and given once.
--report <path> — Also write the JSON report to this file.
--json false Emit the machine-readable report.

Accepting is structural, not editorial: inserted content is unwrapped, deleted content is dropped, a deleted paragraph mark merges its paragraph into the one that follows, deleted table rows vanish, and format-change markers resolve in favour of the formatting the document shows. A document with no revisions passes through unchanged, so a pipeline can run accept unconditionally.

The JSON report (packt-prose/accept-report@1):

Field Type Notes
output string The file written.
insertions int Insertions kept: unwrapped content, inserted paragraph marks and rows.
deletions int Deletions removed, including deleted table rows.
paragraphsMerged int Deleted paragraph marks resolved by merging.
formatChanges int Format-revision markers dropped.

styles

Reports or applies the house Word styles. Packt manuscripts arrive with the right words in the wrong paragraph styles — a heading typed as bold body text, a note that is a paragraph rather than an info box — and every downstream rule reads a paragraph's tag, so the style pass has to come first.

packt-prose styles report -i chapter.docx
packt-prose styles apply  -i chapter.docx -o chapter_styled.docx
Flag Default Effect
-i <path> — Input .docx. Required.
-o <path> — Output .docx. Required for apply.
--author <name> Packt CE Name shown against each tracked paragraph revision.
--install-styles true Add any house style definitions the document lacks.
--spacing true Close up double spaces and spaces before punctuation.
--review false Also apply the calls the tool is not confident about.
--json false Emit the machine-readable report.

The style pass is a production task that runs on its own request — it is not part of a copyedit, whose apply never passes --styles. Use styles when the styles are the whole job, and apply --styles when one deliverable needs the style layer and an edit list in a single write.

A style change is written as a tracked paragraph revision (w:pPrChange), which Word shows in the margin and which does not block the text edits that follow. Installing a missing style definition is not trackable at all — Word has no revision type for it — so apply says which ones it installed.

Furniture keeps only its style's look. A raw draft hand-paints what the styles provide — an orange 20pt chapter line, a shaded tip box — and Word renders direct formatting over the style, so a re-styled heading would keep wearing the author's hand. When a paragraph is re-styled into furniture (a chapter line, heading, caption, figure or callout), hand formatting comes off with it: uniform run color, size, face, highlight, shading and underline as tracked run revisions, and direct alignment, spacing, indents and shading under the same paragraph revision — so Reject restores the author's look whole. A property carried by only some runs is a choice, not a style substitute, and stays; so does a monospaced face, which is code wherever it sits. The report lists each paragraph and what came off under furnitureFormatting.

One look across deliverables. The chapter furniture and heading ladder carry an explicit face in the catalog (Jost* Medium / Jost*, the 86% dominant across 3,246 delivered books) rather than inheriting whatever base font a draft arrived with. And a manuscript that arrives defining a house style name with its own designer's look — one real book set the chapter furniture in Trebuchet at 72pt under the same ids — has those definitions rewritten to the catalog's. Like installing a missing definition, a redefinition is untracked by nature, so the report names the ids under stylesRedefined. A style's own list numbering survives the rewrite.

The chapter opening is reshaped to what copyeditors deliver. A paragraph reading Chapter N: Title among the first few becomes two: the bare number styled HS - ChapterNumber and the bare title styled HS - ChapterTitle — the corpus carries hundreds of tracked re-styles into exactly this pair. The word "Chapter" and the delimiter leave as tracked deletions, the new paragraph mark is a tracked insertion, and rejecting it all restores the author's line. Title-shaped lines above the chapter opening — a chapter file that arrives carrying the book's own title line — are the book's furniture, not the chapter's, and each leaves as a whole tracked deletion, named in the report under chapterFurniture.removed. It runs after every other channel has written, because it reshapes paragraphs and nothing may still be measuring. A document that opens any other way is left alone.

Re-styling a list item moves its numbering along with its look. Word counts a list through numbering instances in a separate part, and a manuscript restarts each list by giving its first item its own instance; a bare style swap would leave that first item counting in the old template's list while its siblings join the house one, chaining every list in the chapter onto the last. The styles pass instead rebuilds the manuscript's instances on the house numbering, restarts included, so each list still begins at 1 — and a list item re-styled into prose stops being counted at all. Rejecting the change restores the old numbering with the old style.

--review exists because the classifier separates the paragraphs it can prove from the ones it is guessing at. By default only the confident calls are written and the rest are listed under needsReview for a person.

Whitespace, and why it is untracked. --spacing closes up double spaces and spaces before punctuation. It cannot be a rule: the canonical text a rule reads has already had its whitespace collapsed, which is what lets a quoted find string match text a human typed, so no rule can see a double space. It is also deliberately not tracked — a revision mark on one space renders wider than the space it deletes, and a chapter carrying three hundred of them buries the edits an author needs to read.

That makes the report the only record of it, so the report carries both a count and the repairs themselves:

Field Type Notes
spacingDefects int Individual defects closed up.
spacingFixes[] array One entry per paragraph stretch repaired.
spacingFixes[].pid int Paragraph.
spacingFixes[].count int Defects in that stretch.
spacingFixes[].before string The text either side of the first repair, before.
spacingFixes[].after string The same window, after.

Spacing inside a code listing is content, not a defect, and is left alone.

review

One read of a document that answers what a copyeditor needs to know before starting: what kind of chapter it is, which profile applies, how big an edit it can take, what the findings are per rule, what comments it already carries, and anything that will block a write.

packt-prose review -i chapter.docx --chapter chapter.json --mech mech.json
Flag Default Effect
-i <path> — Input .docx, or a chapter JSON. Required.
--chapter <path> — Also write the chapter JSON here, so the extract is not repeated.
--raw false Tag paragraphs from their own styles, unclassified (docx input only). Inspection only.
--mech <path> — Also write the mechanical fix plan here, as an edit list.
-o <path> - Write the report here.
--profile <name> auto Style profile name or path.
--json false Emit the machine-readable report.

It replaces the four commands a copyedit used to open with, and the two side files mean the extract and the mechanical plan are on disk before the first edit is written.

Blocking items are the point of it: existing tracked changes (with the names of the reviewers who made them), comments a previous reviewer left, a document whose paragraphs are not in house styles. Each names the command that clears it.

The exit code gates on them. A blocker, or any finding at or above --fail-on, exits 1; a blocker exits 1 even under --fail-on never, because nothing can be written until it is cleared and that is a fact about the document rather than a severity to filter. Most raw drafts therefore exit 1 on the first read, which is the honest answer — exit 1 means "read the report", not "the command failed".

comments

Reads back the comments a document carries — the author's, a previous editor's, or the ones this tool just wrote — with the text each is anchored to.

packt-prose comments -i chapter_CE.docx --json
Flag Default Effect
-i <path> — Input .docx. Required.
-o <path> - Write the report here.
--json false Emit the machine-readable report.

Each comment reports an id, author, initials, date, text, the pid it sits in, and the anchor — the words it marks. A comment whose anchor spans paragraphs reports the first. This is how to check what a document already says before adding to it, and how to verify what apply wrote.

show

Prints chosen paragraphs with everything a copyedit needs to know about them: the exact canonical text (quoted, so whitespace is visible), tag, style, list and cell position, protected and bold spans, and — with --findings — every finding on those paragraphs.

packt-prose show -i chapter.json --pid 164 --findings
packt-prose show -i chapter.json --pid 12,40-45 --json
Flag Default Effect
-i <path> — Input .docx or chapter JSON. Required.
--pid <list> — Paragraphs to show: 12, 12,40, 12-18, or a mix. Required.
--findings false Run the rules and include each paragraph's findings.
--profile <name\|path> auto Style profile, for --findings.
-o <path> - Write the output here.
--json false Emit the show schema (packt-prose/show@1) instead of text.

This is the random-access read. extract --format text is the sequential one — reading the manuscript — and this one answers "show me paragraph 164 again, exactly": the text it prints is the canonical string an edit's find must quote.

lint

Reports house-style findings for a chapter.

packt-prose lint -i chapter.docx --format md
Flag Default Effect
-i <path> — Input .docx, chapter JSON, or - for chapter JSON on stdin. Required.
--format text\|json\|md text Output format.
--json false Forces --format json. Takes precedence over --format.
--pid <n> -1 Lint only this paragraph. Document-wide rules are skipped.
--rule <id,glob> — Report only these rules' findings: comma-separated ids or globs, e.g. packt.caps.*.
--severity <level> — Report only this severity and above: error, warning, or advice.
--summary false One line per rule with a count and an example, instead of every finding.
--off <id,glob> — Silence these rules for this run.
--raw false Tag paragraphs from their own styles, unclassified (docx input only). Inspection only.
--profile <name\|path> auto Style profile.
--fail-on error\|warning\|never error Lowest severity that sets exit code 1.

-i chapter.docx extracts in process. -i chapter.chapter.json skips extraction, which is why an agent extracts once and reuses the JSON.

Findings are sorted by pid, then span start, then rule id, then span end. The order is total and stable, so two runs are diffable.

--format text prints file:pid:spanStart severity: message, one finding per line, with an indented fix: line where a mechanical replacement exists, and a count as the last line. --format md groups findings by rule and renders each group as a table.

autofix

Applies the mechanically safe subset of the rules — the ones marked with * in packt-prose rules — as tracked changes. No model is involved and no judgment is exercised. The fixes never touch a protected span, never change a digit, and are idempotent.

packt-prose autofix -i chapter.docx -o chapter_fixed.docx
Flag Default Effect
-i <path> — Input .docx. Required.
-o <path> — Output .docx. Required unless --dry-run. Must differ from -i.
--dry-run false List the fixes and write no document.
--author <name> Packt CE Name shown against each tracked change.
--json false With --dry-run, emit the planned fixes as an edit list; in write mode, emit the apply-report schema.
--profile <name\|path> auto Style profile.
--fail-on <level> error Accepted, but the exit code is driven by whether fixes were found, not by severity.

A document that already contains tracked changes is refused with exit 3 unless --dry-run is given. When there is nothing to fix, -o is still written — a copy of the input — so a pipeline can rely on the file existing.

autofix -o … --json still reports in text: the JSON payload exists only for --dry-run, where the planned fixes are an edit list. To get a machine-readable record of a write, plan with autofix --dry-run --json, then pass that edit list to apply --json.

check-edit

Runs every safety gate against a proposed edit list and reports one verdict per edit, in input order. Pure: no document is written and nothing is read except the two inputs.

packt-prose check-edit --chapter chapter.chapter.json --edits edits.json --json
Flag Default Effect
--chapter <path> — Chapter JSON or .docx the edits apply to. Required.
--edits <path> — Edit list JSON, or - for stdin. Required.
--budget <n> 0 Largest number of edits allowed for the chapter. 0 means no limit.
--mechanical true Verify the batch against the mechanical plan apply --mechanical will run (below).
--profile <name> auto Profile for the mechanical plan; match the apply run's.
--off <ids> — Rules silenced for the apply run's mechanical plan; match the apply run's.
--json false Emit the verdicts schema instead of the text summary.

An edit that overlaps a planned mechanical fix silences that fix at apply, on the theory that the rewrite subsumes it. The theory fails exactly when the edit's replacement carries the defect forward — a caption rewritten with its period separator intact — so check-edit re-plans the mechanical fixes on the text as the batch leaves it and fails the run when a silenced fix is still needed. The report's mechanicalOwed array quotes each owed fix against the post-batch text: fold it into the edit that covers it. apply refuses the same condition, so nothing ships either way; this check exists so the refusal lands before the write is spent.

Every check that can be evaluated is evaluated, even after one has failed, so a caller can fix all of an edit's problems in one revision. Checks that cannot be evaluated — anything needing a located span, when find could not be found — are omitted from the report rather than reported as passing, so the number of checks per verdict varies.

Comments in the edit list are checked too, though only their anchors: a comment changes no text, so the only way it can go wrong is to attach to the wrong words. A comment whose quote cannot be located, is ambiguous, or whose text is empty is reported under commentsRejected, and any such failure gives exit 1 exactly as a failed edit does.

Text output reports only the failures, plus a passed of total count.

apply

Writes an edit list into a .docx as native tracked changes, and its comments as native Word comments, then reports what was applied and what was refused.

packt-prose apply -i chapter.docx --edits edits.json -o chapter_CE.docx
Flag Default Effect
-i <path> — Input .docx. Required.
-o <path> — Output .docx. Required. Must differ from -i.
--edits <path> — Edit list JSON, or - for stdin. Required.
--chapter <path> — Chapter JSON the edits were written against. Checked against -i.
--author <name> Packt CE Name shown against each tracked change.
--budget <n> 0 Largest number of edits allowed for the chapter. 0 means no limit.
--styles false Run the house-style pass first, in this same write.
--spacing true With --styles or --mechanical, close up double spaces and spaces before punctuation (untracked, itemised in the report).
--report <path> — Also write the apply report to this file. -o is the document and may be given once.
--mechanical false Plan the mechanical fixes and apply them alongside the edit list; the whitespace closes ride this pass too unless --spacing=false.
--accept false Accept the input's existing tracked changes first, so a re-run can proceed.
--layer false Write on top of the input's existing tracked changes, keeping them pending under their own authors.
--profile <name> auto Style profile name or path.
--off <id,glob> Drop named rules from the mechanical plan for this run.
--json false Emit the apply-report schema instead of the text summary.

--styles and --mechanical exist so a document takes one write. Run them as separate commands and each writes tracked changes that the next one refuses; the obvious way round that — an accept in between — silently takes the earlier layer with it. --styles also re-extracts the chapter before the edits are measured, because re-styling changes how paragraphs are classified.

With --styles, the whitespace repairs described under styles run too, and there is no flag to stop them here. They are untracked, so read spacingDefects and spacingFixes out of the report.

--mechanical drops any planned fix that overlaps an edit already in the list, and reports the count as mechanicalSuperseded with each dropped fix named in mechanicalSupersededFixes. Overlap is measured by span, not by text: the same wrong word can appear three times in a paragraph, and matching on the string alone silently discarded the two occurrences nobody was fixing.

apply re-runs the full gate set itself and never assumes the caller ran check-edit. It is the last point before bytes are written, so it is the only one that can guarantee nothing unsafe reaches the file. A refused edit is reported with the name of the first check it failed.

Layering. By default a document that already carries tracked changes is refused, and --accept resolves them into the baseline — which bakes the earlier editor's work into the text under nobody's name. --layer keeps it instead: edits are measured against the accepted view (the text extract already reads), new revisions go in under --author, and the existing ones stay pending under theirs. Word renders the two layers in different colors and every reject path survives — deleting text another editor inserted nests the deletion inside their insertion, and writing into the middle of one splits it around the new text, the same shapes Word itself records. Extract the chapter from the revised document, not from an earlier version: the pids and the text are the revised document's.

Three things still refuse or step aside under --layer. A pending structural revision — an inserted or deleted paragraph mark, a table row in flux — is refused outright, because paragraph positions depend on how it resolves. A split is rejected on a paragraph carrying any revision (split_on_revised_paragraph), since reshaping it would disturb them. And the styles pass skips a paragraph holding another editor's style revision, reporting the pids: Word keeps one style revision per paragraph, so writing ours would erase their reject path. The untracked whitespace repair also leaves revised paragraphs alone.

Comments ride the same edit list. An edit carrying a comment field gets a Word comment anchored on its revision, written only when the edit itself passes the gates; the comments array anchors margin notes on text that is not being changed. Comments never alter a document's text — a reviewer who rejects everything is left with the author's exact words — and a comment that cannot be anchored is reported and skipped rather than guessed at.

--chapter guards against a stale extract. Paragraph ids are positions in one specific file, so an edit list built against an older copy of a document can name paragraphs that have since moved. When the flag is given, the chapter's sourceSha256 is compared with the input document and a mismatch stops the run with exit 2 and writes nothing. Pass it whenever the chapter JSON and the document are separate arguments in a pipeline; the check costs nothing and the failure it prevents is silent.

A document that already contains tracked changes is refused with exit 3. In a partly revised paragraph the offsets an edit was measured against depend on which revisions a reviewer later accepts, so a new change could land in the wrong place. There is no override.

voice

Measures author voice before and after an edit and flags the metrics that moved further than a professional copyedit moves them.

packt-prose voice --chapter chapter.chapter.json --edits edits.json --json
Flag Default Effect
--before <path> — The author's text: .docx, chapter JSON, or plain text.
--after <path> — The edited text, same forms.
--chapter <path> — Chapter to measure the effect of --edits on.
--edits <path> — Edit list to measure, with --chapter.
--json false Emit the voice schema instead of the table.

Give either --before and --after, or --chapter and --edits; --chapter/--edits wins if both pairs are present. With --chapter/--edits the edits are applied to the text in memory and no document is written, so voice can be measured before anything is committed. An edit that cannot be located is skipped rather than failing the run.

Voice is measured over editable prose only: headings, captions, and code say nothing about an author's register. Inputs with a .docx or .json extension are read as documents; anything else is read as plain text.

Metrics are compared against styles/voice-baseline.json, derived from aligned author-to-published sentence pairs in the reference corpus. A flag means a metric fell outside expected ± tolerance, where tolerance is two standard deviations of the observed corpus deltas.

rules

Lists the house-style rules this binary enforces, under a given profile.

packt-prose rules --format md
Flag Default Effect
--format table\|json\|md table Output format.
--json false Forces --format json.
--id <ruleId> — Show only this rule. Exit 2 if no enabled rule has that id.
--profile <name\|path> auto Style profile.
--fail-on <level> error Accepted and ignored: rules reports no findings.

--format table marks autofix-capable rules with *. --format md generates the rules reference document, including a section listing the rules the profile turns off.

corpus

Corpus statistics, voice baseline derivation, and scoring of the deterministic rules against real copyedits. These are build-time tools: the corpus is source material for developing the house style, not something an editor needs, so it is not embedded in the binary and the default path is repository-relative.

packt-prose corpus stats

A subcommand is required; corpus alone, or an unknown subcommand, is exit 2. There is no corpus -h.

Subcommand Flag Default Effect
stats --corpus <path> tools/packt-prose/reference/corpus/b32128_edit_corpus.json Corpus JSON to read.
stats --json false Emit JSON.
baseline --corpus <path> same default Corpus JSON to read.
baseline -o <path> - Output file, or - for stdout. Always JSON.
score --corpus <path> same default Corpus JSON to read.
score --json false Emit JSON.
score --profile <name\|path> auto Style profile to score.
score --fail-on <level> error Accepted and ignored.

baseline regenerates the contents of styles/voice-baseline.json. score runs the rule set over every aligned pair and reports how often the rules alone moved the author's sentence toward the published one — the mechanical floor any model has to beat.

doctor

Reports the version, the platform, and what is embedded, and verifies the embedded house style by loading and compiling every rule.

packt-prose doctor
Flag Default Effect
--json false Emit the doctor schema instead of the text report.

version

packt-prose version --json
Flag Default Effect
--json false Emit the version schema instead of one line of text.

Style profiles

--profile takes either the name of a built-in profile or a path to a YAML file. packt-prose doctor lists the built-in names. An unknown name is exit 3; a path that does not parse is exit 3.

A profile sets a name, the per-book options, and per-rule enablement:

name: B18088
options:
  contractions: neutral
rules:
  packt.advice.passive: false
  packt.advice.*: false
  packt.advice.referent: true

contractions is neutral by default — no findings in either direction, because the corpus's contraction edits are nearly balanced and a diff against print-accepted chapters corroborated none of the prompted expansions. The other policies: expand flags each contraction as a warning carrying the full form as its suggestion, enforcing the Writers' Guide letter of the law for a book that wants it; house reports a full form as a candidate for contraction, as advice. The built-in conversational profile is house, for a book written that way on purpose.

Under no policy is the rule mechanical: autofix and --mechanical never touch a contraction. Expanding one changes the author's voice, and the one day this ran mechanically a memoir chapter took 160 unreviewed expansions — two-thirds of its tracked changes. An editor decides each instance; the suggestion makes accepting one a single action.

The built-in names are default, conversational, appendix, frontmatter, preface, parts, questions, flashcards and example-book. --profile auto picks one by reading the document.

Rule keys accept a trailing * to match a family; the longest matching pattern wins, so a family can be silenced and one member re-enabled. The profile name appears in the diagnostics and rules schemas.

JSON schemas

Every payload carries a schema field naming the contract and its version, except the two corpus payloads noted below. Field names are camelCase. All output is indented with two spaces, HTML escaping off, with a trailing newline.

Two-element ranges — span and each entry of protected — are [start, end) rune offsets into the paragraph's normalized text, as set out under offsets above.

chapter

packt-prose/chapter@1. Produced by extract; consumed by lint, check-edit, voice, and anything else taking a chapter. The excerpt below is four paragraphs of a 662-paragraph extract, with paragraph 11's seven protected spans cut to two. Nothing else is changed.

packt-prose extract -i chapter.docx -o chapter.chapter.json
{
  "schema": "packt-prose/chapter@1",
  "sourcePath": "chapter.docx",
  "sourceSha256": "e4a1a19cce5cc5f1cfd759b110d3adf03729a2964f4f3eea3cbee3634a70a29c",
  "paragraphs": [
    {
      "pid": 11,
      "tag": "P",
      "style": "P-Regular",
      "text": "This chapter's example uses the xgboost, pandas, sklearn, solas-ai, matplotlib, fairlearn, and seaborn libraries. Instructions on how to install all these libraries can be found in the Preface.",
      "protected": [
        [
          32,
          39
        ],
        [
          41,
          47
        ]
      ],
      "list": null,
      "editable": true
    },
    {
      "pid": 24,
      "tag": "P",
      "style": "L-Numbers",
      "text": "To run this example, we will first install and load the libraries like this:",
      "protected": null,
      "list": {
        "level": 0,
        "ordered": true
      },
      "editable": true
    },
    {
      "pid": 25,
      "tag": "CODE",
      "style": "L-Source",
      "text": "%%capture",
      "protected": [
        [
          0,
          9
        ]
      ],
      "list": null,
      "editable": false
    },
    {
      "pid": 61,
      "tag": "CAPTION",
      "style": "IMG-Caption",
      "text": "Figure 13.1 – A sample output from the loaded dataset",
      "protected": null,
      "list": null,
      "editable": true
    }
  ]
}
Field Type Notes
sourcePath string The path as given on the command line.
sourceSha256 string SHA-256 of the source .docx bytes.
paragraphs[].pid int Index in document order, counting every paragraph including empty ones. Meaningful only against this source file.
paragraphs[].tag string P, H1–H6, CODE, CAPTION, or INFOBOX. Bullets and numbered items are P with a non-null list.
paragraphs[].style string The underlying Word style name, for diagnosis. May be empty.
paragraphs[].text string Normalized paragraph text. The coordinate system for every offset.
paragraphs[].protected array or null Rune spans no edit may touch: code-formatted runs, code-shaped tokens in prose, links, fields, tabs, breaks. null, not [], when there are none.
paragraphs[].bold array Rune spans set in bold. Omitted when empty.
paragraphs[].list object or null {"level": int, "ordered": bool} for list paragraphs.
paragraphs[].editable bool false for paragraphs that are not prose at all. Rules never fire in them and gates reject edits on them. Describes the kind of paragraph, not whether it is empty.

The paragraph excerpt above shows the two shapes that trip up consumers: protected is null rather than an empty array when a paragraph has nothing protected, and a two-element range is rendered across multiple lines by the indenting encoder.

diagnostics

packt-prose/diagnostics@1. Produced by lint --json. This is the whole output of packt-prose lint -i chapter.docx --json --pid 1 on a chapter whose paragraph 1 reads By default the server reserves 32GB of RAM for the cache, and IO throughput is the limit you will hit first.

{
  "schema": "packt-prose/diagnostics@1",
  "profile": "default",
  "counts": {
    "error": 2,
    "warning": 1,
    "advice": 0
  },
  "diagnostics": [
    {
      "pid": 1,
      "span": [
        0,
        14
      ],
      "ruleId": "packt.mech.intro-comma",
      "severity": "warning",
      "message": "Put a comma after the introductory word or phrase. (found \"By default\")",
      "styleRef": "STYLE_ANALYSIS §4 Punctuation",
      "fix": {
        "find": "By default the",
        "replace": "By default, the"
      }
    },
    {
      "pid": 1,
      "span": [
        31,
        35
      ],
      "ruleId": "packt.mech.unit-space",
      "severity": "error",
      "message": "Put a space between a number and its unit. (found \"32GB\")",
      "styleRef": "STYLE_ANALYSIS §4 Numbers, units, capitalization",
      "fix": {
        "find": "32GB",
        "replace": "32 GB"
      }
    },
    {
      "pid": 1,
      "span": [
        62,
        64
      ],
      "ruleId": "packt.mech.io",
      "severity": "error",
      "message": "Write I/O, not IO. (found \"IO\")",
      "styleRef": "STYLE_ANALYSIS §4 Hyphenation and compounds",
      "fix": {
        "find": "IO",
        "replace": "I/O"
      }
    }
  ]
}
Field Type Notes
profile string Name of the profile in force.
counts object error, warning, advice tallies over diagnostics.
diagnostics[].pid int Paragraph the finding is in.
diagnostics[].span [int, int] Rune span of the flagged text.
diagnostics[].ruleId string Stable rule identifier, for example packt.mech.io.
diagnostics[].severity string error, warning, or advice.
diagnostics[].message string What to do, in an editor's terms.
diagnostics[].styleRef string Section of the house style the rule enforces.
diagnostics[].fix object Present only when a deterministic, always-safe replacement exists — the autofix set. {"find": …, "replace": …}.

diagnostics is always present, [] when there is nothing to report. A fix is a find/replace pair against the paragraph's normalized text, in the same form an edit takes.

edits

packt-prose/edits@1. Consumed by check-edit, apply, and voice --chapter/--edits; produced by autofix --dry-run --json. This is an excerpt of the output of packt-prose autofix -i chapter.docx --dry-run --json.

{
  "schema": "packt-prose/edits@1",
  "edits": [
    {
      "pid": 1,
      "find": "By default the",
      "replace": "By default, the",
      "category": "mechanics",
      "rule": "packt.mech.intro-comma"
    },
    {
      "pid": 2,
      "find": "file system",
      "replace": "filesystem",
      "category": "mechanics",
      "rule": "packt.consist.compounds"
    },
    {
      "pid": 3,
      "find": "behaviour",
      "replace": "behavior",
      "category": "mechanics",
      "rule": "packt.consist.us-spelling"
    }
  ],
  "comments": [
    {
      "pid": 5,
      "find": "at least eight edits per chapter",
      "text": "Left as-is: this is the author's claim, not a house figure."
    }
  ]
}
Field Type Notes
pid int Paragraph the edit applies to. Required.
find string Must appear in that paragraph's normalized text exactly. Required.
replace string The replacement. Required; may be empty only where a gate allows it.
occurrence int 1-based match selector. Required when find appears more than once. Omitted otherwise.
category string One of rewrite, referent, grammar, acronym, voice, heading, mechanics. Reviewer metadata; never changes how an edit is applied. Optional.
rule string Rule id the edit satisfies, when it came from one. Optional.
note string Free-text rationale. Reaches reports only, never the document. Optional.
comment string Written into the document as a Word comment anchored on this edit's revision: the margin note saying why the change was made. Optional.
intent string Declares an edit that is not a substitution: rewrite, delete, or insert. Changes which gates judge the edit — see below. Optional; omitted means a plain substitution, which most edits are.

Declared intents

The substitution gates hold every edit to substitution scale, which is the right default and the wrong ceiling: recasting a sentence, deleting a redundant one, or adding a bridging one are all real copyedits that no substitution can express. An edit may declare what it is, and the declaration is a trade — the substitution-scale bounds come off, and that class's own checks go on in their place. The quotation must be at least four words in every class; below that, a plain edit expresses the same change under tighter checks.

rewrite recasts whole sentences. The quotation must start where a sentence starts and end where one ends, the replacement must stay under 300 characters and within half-to-double the quotation's length, most of the quotation's content words must survive, and every number, link, identifier and acronym must come through intact (entity_drift applies unchanged).

delete removes prose; replace must be empty. Up to eight words the deletion may take a comma-bounded clause; past that, whole sentences only. Three checks read what the chapter loses: an entity whose only occurrence is in the deleted text (delete_last_entity), an acronym definition the chapter still uses (delete_expansion), and a whole numbered step (delete_ordered_item) are all refused — those are questions for the author, raised as comments.

insert adds prose beside an anchor. The quotation must survive intact at the start or the end of the replacement, the addition is capped at 60 words, and every entity the addition mentions must already occur somewhere in the chapter (insert_entities) — new facts cannot arrive by insertion.

An undeclared edit doing a declared class's work at scale is refused: a plain edit that deletes four or more words fails undeclared_deletion and the message names the intent to set.

The comments array carries margin notes on text that is not being changed — why something was left as-is, or a question for the author — anchored by quotation exactly as an edit is:

Field Type Notes
pid int Paragraph the comment attaches to. Required.
find string The text to anchor on; must appear in that paragraph's normalized text exactly. Required.
occurrence int 1-based match selector, as for an edit.
text string The comment itself. Required.

The splits array is the structural channel: one paragraph becoming a lead-in plus a bulleted list, all tracked. {"pid", "items": [...], "style"?} — items quoted verbatim and in order; the lead-in must end with a colon; between items only separators may remain (deleted as tracked changes); the new paragraph marks are tracked insertions, so a reviewer's Reject restores the paragraph whole; items take L - Bullets unless style names another, and they share a fresh numbering instance so the list renders its bullets and starts at 1 wherever it lands. A split must have its paragraph to itself in the batch (split_collision otherwise), and residue that is not a separator is refused (split_residue) — a split may never delete words nobody quoted. check-edit validates splits with the same code path apply uses, reported under splitsChecked/splitsRejected, and the apply report carries splitsApplied/splitsRejected.

On input, both the envelope above and a bare JSON array of edit objects are accepted, because both shapes turn up in practice. Output always uses the envelope.

Quoting rather than describing is the point: an edit that quotes text which is not there is rejected instead of being applied somewhere plausible, and a comment that quotes text which is not there is rejected instead of being pinned somewhere plausible.

verdicts

packt-prose/verdicts@1. Produced by check-edit --json. One verdict per input edit, in input order. This is the whole output of a single-edit run of packt-prose check-edit --chapter chapter.chapter.json --edits edits.json --json, abbreviated only where noted.

{
  "schema": "packt-prose/verdicts@1",
  "pass": false,
  "verdicts": [
    {
      "edit": {
        "pid": 296,
        "find": "The 60% cutoff value is used for all datasets.",
        "replace": "The same cutoff value is used for all datasets.",
        "category": "rewrite"
      },
      "pass": false,
      "checks": [
        {
          "name": "pid_range",
          "pass": true
        },
        {
          "name": "code_paragraph",
          "pass": true
        },
        {
          "name": "find_missing",
          "pass": true
        },
        {
          "name": "find_ambiguous",
          "pass": true
        },
        {
          "name": "protected_span",
          "pass": true
        },
        {
          "name": "entity_drift",
          "pass": false,
          "detail": "the replacement drops a number: \"60\""
        },
        {
          "name": "overlap",
          "pass": true
        }
      ]
    }
  ]
}

The real run reports 21 checks for this edit; the fourteen that passed between protected_span and entity_drift are elided above for length. Nothing else is changed.

Field Type Notes
pass bool True when every verdict passed.
verdicts[].edit object The input edit, echoed verbatim including occurrence, category, rule, and note.
verdicts[].pass bool True when every check in checks passed.
verdicts[].checks[].name string Gate name. Stable API; appears in review records.
verdicts[].checks[].pass bool Outcome.
verdicts[].checks[].detail string Why it failed. Omitted when the check passed.
commentsChecked int Number of comments in the input. Omitted when zero.
commentsRejected array Comments whose anchor failed, each with the comment echoed and a reason (pid_range, find_missing, find_ambiguous, comment_empty) and detail. Omitted when empty.
formatsChecked int Number of formatting requests in the input. Omitted when zero.
formatsRejected array Formatting requests that would be refused, each echoed with a reason and detail. Always present when formats were checked — empty means checked and clean.
stylesChecked int Number of style overrides in the input. Omitted when zero.
stylesRejected array Style overrides that would be refused (pid_range, unknown_style). Always present when styles were checked.
stylesResolved array Accepted overrides whose name maps to a different style id — {pid, style, resolvesTo}. "heading 5" resolves to H4-Subheading because the heading ladder stops at four; an unexpected entry here means the style you asked for does not exist. Omitted when every name is a plain spelling of its style.
voice object A voice report measuring what the batch's passing edits do to the chapter's prose, against the copyedit baseline. Advisory: its flags never fail the check. Present whenever any edit passed.

The gates, in report order:

Check Fails when
pid_range pid is not in the chapter.
code_paragraph The target paragraph is not editable prose.
heading_deletion A heading paragraph's replace is empty. Reported only for headings.
heading_degraded A heading's find is 3+ words and replace is 2 or fewer. Reported only for headings.
find_missing find does not appear in the paragraph after normalization.
find_ambiguous find matches more than once and occurrence was not set.
protected_span The located span intersects a protected rune, or the replacement fails to carry a protected token through unchanged.
noop_edit find equals replace after normalization.
degenerate_find find is a single character, or contains no letter or digit — a quotation that identifies nothing.
oversize_replace replace is more than 4× find in runes and more than 120 runes longer.
large_insertion replace adds 12 or more words, or 8 or more novel content words.
collapsed_replace find is 4+ words and replace is 2 or fewer.
case_mismatch_start replace starts as a mid-word or wrong-case fragment against the find context.
artifact_parens replace introduces ().
artifact_dots replace introduces a run of three or more dots.
bullet_artifact replace introduces a literal hyphen-and-space bullet prefix.
reversed_insertion replace embeds the whole of find behind a connective — the tell of a reversed edit pair.
and_removed The edit removes a serial , and.
markdown_introduced replace adds Markdown syntax that find did not have.
control_chars replace contains control characters.
quote_churn find and replace differ only in quote or apostrophe glyphs.
entity_drift The multiset of digits, numbers with units, URLs, code-shaped tokens, or all-caps acronyms differs between find and replace.
undeclared_deletion A plain edit deletes 4 or more words; a deletion at that scale must declare "intent": "delete".
intent_unknown intent names a class the gates do not know.
rewrite_scope A declared rewrite quotes fewer than 4 words, replaces with nothing, or replaces with more than 300 characters.
rewrite_bounds A declared rewrite starts or stops mid-sentence, or its replacement does not close a sentence.
rewrite_ratio A declared rewrite's replacement falls below half or rises above double the quotation's length, beyond a 40-rune slack.
rewrite_retention A declared rewrite keeps under 40% of the quotation's content words, or over half its replacement's content words are new.
delete_shape A declared delete carries replacement text.
delete_scope A declared delete quotes fewer than 4 words.
delete_bounds A declared delete of more than 8 words starts or stops mid-sentence.
delete_last_entity The deleted text holds the chapter's only occurrence of an entity.
delete_expansion The deleted text holds an acronym's only definition while the chapter keeps using the acronym.
delete_ordered_item The deletion takes the whole of a numbered list item.
insert_anchor A declared insert's quotation does not survive intact at the start or the end of the replacement.
insert_scope A declared insert adds fewer than 4 or more than 60 words.
insert_entities The addition mentions an entity the chapter holds nowhere.
overlap Two located spans overlap within one paragraph. The first edit wins.
budget_exceeded The batch's applicable edits exceed --budget. Reported only when --budget is non-zero.

A declared intent swaps check sets rather than adding one: a rewrite is not judged by oversize_replace, large_insertion, or collapsed_replace — its own scope, bounds, ratio, and retention checks are the versions of those bounds that make sense at sentence scale, and entity_drift holds unchanged. A delete is judged against the chapter instead of against a replacement. An insert is judged by its anchor and its whitelist instead of by large_insertion.

Checks are pure functions of (chapter, edit) with no I/O, so the same library gives the same verdict inside check-edit, inside apply, and in any future service.

overlap and budget_exceeded are the two checks that depend on the rest of the batch, and both apply only to an edit that is otherwise sound. An edit already refused for its own reasons is never going to reach the document, so it claims neither a span nor a place in the budget, and it carries neither check in its verdict. The consequence worth knowing: a caller can submit a long list in priority order and get the first budget's worth of applicable edits, rather than losing a good edit because a bad one preceded it.

review

packt-prose/review@1. Produced by review. The orientation bundle: what the document is, what blocks a write, and the shape of its findings. The byRule entries share the summary shape lint --summary emits.

Field Type Notes
schema string packt-prose/review@1.
input string The file reviewed.
sourceSha256 string Checksum of the document; pids are positions in this exact file.
docType string What the document reads as (chapter, preface, …).
detected bool Whether docType was detected or merely the default.
profile string The profile in force.
proseWords int Words in editable prose paragraphs.
suggestedBudget int Edits this chapter should attract, at the corpus midpoint (25/1k; the professional band is 20–35, and a typical draft ending far below it was under-edited).
existingComments int Comments already in the document. Read them first.
blocking []string What must be dealt with before anything is written. A blocker sets exit 1 whatever --fail-on says.
findings object Counts by severity: {"error": n, "warning": n, "advice": n}.
byRule []object One entry per firing rule: ruleId, severity, count, fixable, example, pid.
styles object {"unstyled": n, "toRestyle": n, "needsAPerson": n} — what the style pass would do.
mechanical []edit The mechanical fix plan, in the edits schema shape.

styles report

Emitted by styles report --json (and, with applied, installedStyles and output added, by styles apply --json).

Field Type Notes
schema string Report version.
sourcePath string The document read.
generation string Which template generation the document is in.
paragraphs int Paragraphs inspected.
unstyled int Paragraphs in no recognised style.
missingStyles []string House styles the document does not define.
confident []object The re-stylings the pass would write: pid, from, to, reason, quote.
needsReview []object The calls it will not make alone; same shape, to may be empty.
spacingDefects int Present only when the spacing pass ran: 0 means ran-and-clean, absent means not run.
spacingFixes []object One entry per repaired stretch: pid, count, before, after.

show

packt-prose/show@1. Produced by show --json: the requested paragraphs in the chapter schema's paragraph shape, each with an optional findings array in the diagnostics shape.

apply-report

packt-prose/apply-report@1. Produced by apply --json. The audit trail of a run: what was written, and what was refused with the reason. The JSON below is the whole output of this command, on an edit list of two edits where the second drops a number.

packt-prose apply -i chapter.docx --edits edits.json -o chapter_CE.docx --json
{
  "schema": "packt-prose/apply-report@1",
  "output": "chapter_CE.docx",
  "author": "Packt CE",
  "applied": 1,
  "rejected": [
    {
      "edit": {
        "pid": 296,
        "find": "The 60% cutoff value is used for all datasets.",
        "replace": "The same cutoff value is used for all datasets.",
        "category": "rewrite"
      },
      "reason": "entity_drift",
      "detail": "the replacement drops a number: \"60\""
    }
  ]
}
Field Type Notes
output string Path written, as given on the command line.
author string Name attributed to every tracked change and comment.
applied int Number of tracked changes written.
rejected array Always present, [] when nothing was refused.
formatsApplied int Formatting requests that changed anything. Omitted when zero.
formatRunsChanged int Runs those requests touched — one request over a phrase set in two runs changes both.
mechanicalSuperseded int Planned mechanical fixes a judged edit already covers.
mechanicalSupersededFixes array Each dropped fix, echoed as an edit with its rule.
rejected[].reason string Name of the first gate the edit failed.
rejected[].detail string That gate's detail. Omitted when empty.
commentsApplied int Number of Word comments written: edit-attached plus standalone. Omitted when zero.
commentsRejected array Comments whose anchor failed, in the check-edit shape. Omitted when empty.
voice object A voice report measuring what the applied edits did to the chapter's prose, against the copyedit baseline. Advisory: the write has already happened, so a flag is a reason to reread the output in Word. Present whenever any edit was applied.

voice

packt-prose/voice@1. Produced by voice --json. This is the whole output of packt-prose voice --before before.txt --after after.txt --json on a short passage rewritten into the passive.

{
  "schema": "packt-prose/voice@1",
  "before": {
    "words": 50,
    "sentences": 5,
    "contractionRate": 0,
    "wePer1k": 240,
    "youPer1k": 0,
    "exclamations": 0,
    "questions": 0,
    "meanSentenceLen": 10,
    "hedgePer1k": 0,
    "passiveRate": 0.2
  },
  "after": {
    "words": 44,
    "sentences": 5,
    "contractionRate": 0,
    "wePer1k": 0,
    "youPer1k": 0,
    "exclamations": 0,
    "questions": 0,
    "meanSentenceLen": 8.8,
    "hedgePer1k": 0,
    "passiveRate": 1.4
  },
  "delta": {
    "contractionRate": 0,
    "exclamations": 0,
    "hedgePer1k": 0,
    "meanSentenceLen": -0.11999999999999993,
    "passiveRate": 5.999999999999999,
    "questions": 0,
    "wePer1k": -1,
    "words": -0.12,
    "youPer1k": 0
  },
  "flags": [
    {
      "metric": "passiveRate",
      "observed": 5.999999999999999,
      "expected": 0.13,
      "message": "passiveRate rose by 600% against an expected 13%"
    }
  ]
}
Field Type Notes
before, after object The same nine measurements over each text.
delta object after relative to before, as a ratio: -1 is a 100% fall, 6 a 600% rise. Keys are sorted.
flags array Metrics outside expected ± tolerance. Always present, [] when none.
flags[].observed number The delta ratio seen.
flags[].expected number The baseline's expected delta ratio.

Rates are per word (contractionRate, passiveRate) or per 1,000 words (wePer1k, youPer1k, hedgePer1k). exclamations and questions are raw counts. Deltas are floating point and not rounded; compare with a tolerance rather than for equality.

rules

packt-prose/rules@1. Produced by rules --json. This is the whole output of packt-prose rules --json --id packt.voice.contractions.

{
  "schema": "packt-prose/rules@1",
  "profile": "default",
  "rules": [
    {
      "id": "packt.voice.contractions",
      "type": "policy",
      "severity": "advice",
      "scope": "prose",
      "autofix": false,
      "message": "House voice uses contractions; consider contracting this.",
      "styleRef": "STYLE_ANALYSIS §2 Voice",
      "docs": "Contractions are house voice: the published corpus uses \"we'll\" 59 times against \"we will\" 10, and \"don't\" 62 times against \"do not\" 8. The default policy, house, therefore flags a full form as a candidate for contraction, as advice. A book sets the profile's contractions option to expand for the opposite policy, or to neutral for no findings either way. Under every policy \"cannot\" is preferred over \"can't\", and \"let's\" survives in a walkthrough."
    }
  ]
}
Field Type Notes
profile string Profile the listing was filtered by.
rules[].id string Stable rule identifier.
rules[].type string Matcher kind: advisory, capitalization, conditional, consistency, document, existence, policy, regex, or substitution.
rules[].severity string error, warning, or advice.
rules[].scope string prose, heading, or all.
rules[].autofix bool Whether autofix can apply this rule unaided.
rules[].message string The message a finding carries.
rules[].styleRef string Section of the house style.
rules[].docs string Long-form explanation, including any profile options the rule reads.

Only rules enabled under the profile are listed, so the listing is the authoritative answer to "what will this run check".

doctor

packt-prose/doctor@1. Produced by doctor --json.

{
  "schema": "packt-prose/doctor@1",
  "version": "2026.08.05.1432",
  "commit": "9f2b7c1e5a8d43069f2b7c1e5a8d43069f2b7c1e",
  "platform": "darwin/arm64",
  "goVersion": "go1.26.1",
  "rules": 42,
  "profiles": [
    "default",
    "example-book"
  ],
  "styleChecksum": "1a8d6a75cab57374",
  "problems": []
}
Field Type Notes
version string Stamped at build time as the UTC timestamp of the release; dev in an unstamped build.
commit string The source the binary was built from. Omitted in a development build.
platform string GOOS/GOARCH.
rules int Rules that loaded and compiled under the default profile.
profiles array Built-in profile names accepted by --profile.
styleChecksum string Checksum of the embedded house style. Two binaries agreeing here enforce identical rules.
problems array Always present, [] when healthy. Non-empty means exit 1.

styleChecksum is the field to compare when two machines disagree about a chapter.

version

packt-prose/version@1. Produced by version --json.

{
  "schema": "packt-prose/version@1",
  "version": "2026.08.05.1432",
  "commit": "9f2b7c1e5a8d43069f2b7c1e5a8d43069f2b7c1e",
  "os": "darwin",
  "arch": "arm64",
  "goVersion": "go1.26.1"
}

A release version is the UTC timestamp it was built at, to the minute. commit is the source it was built from, and is omitted from a development build, where version reads dev. Releases carry no git tag, so commit is what maps a version back to the code that produced it.

corpus payloads

corpus stats --json and corpus score --json carry no schema field; they are build-time reports rather than an interop contract. corpus baseline does carry one, packt-prose/voice-baseline@1, because the voice comparison reads it.

corpus stats --json:

{
  "pairs": 949,
  "changed": 949,
  "unchanged": 0,
  "changedRate": 1,
  "authorWords": 26993,
  "publishedWords": 28819,
  "meanAuthorSentenceLen": 28.443624868282402,
  "meanPublishedSentenceLen": 30.367755532139093
}

corpus score --json, with byRule truncated:

{
  "pairs": 949,
  "attempted": 83,
  "exactMatch": 1,
  "improved": 50,
  "worsened": 18,
  "unchanged": 866,
  "meanTokenF1Before": 0.7397284751283753,
  "meanTokenF1After": 0.7416928760990683,
  "byRule": {
    "packt.mech.intro-comma": 16,
    "packt.mech.unit-space": 18,
    "packt.words.usage": 11
  }
}

corpus baseline, which is what styles/voice-baseline.json contains, with metrics truncated:

{
  "schema": "packt-prose/voice-baseline@1",
  "source": "b32128_edit_corpus.json",
  "pairs": 949,
  "metrics": {
    "meanSentenceLen": {
      "expected": 0.04,
      "tolerance": 0.17
    },
    "wePer1k": {
      "expected": -0.33,
      "tolerance": 0.94
    }
  }
}

For agents

Four habits make the difference between a fast loop and a slow one.

Always pass --json. Text output is for humans and its layout is not a contract. --json overrides --format wherever both exist, so it is enough on its own. The one exception is autofix in write mode, which reports in text either way; plan with autofix --dry-run --json and hand the resulting edit list to apply --json when you need a machine-readable record.

Extract once, then reuse the chapter JSON. lint -i chapter.docx re-parses and re-normalizes the whole document on every call. Run extract once, keep chapter.chapter.json, and pass that to lint, check-edit, and voice. The pid values and every offset are only meaningful against the file recorded in sourceSha256, so re-extract if the document changes.

Re-check one paragraph with lint --pid N. After proposing a fix, lint --pid N --json re-runs just that paragraph's rules and skips the document-wide ones. It is the fast inner loop; a full lint between every edit is wasted work.

Batch every edit into one check-edit call. The gates report all failures for all edits in a single pass, and every check that can be evaluated is evaluated even after one has failed. One call gives you everything you need to revise the whole list; per-edit calls give you the same information one round trip at a time. Then call apply once — it re-runs the gates itself, so an edit that slipped past your reasoning still cannot reach the file.

Two further notes. Offsets are rune offsets into normalized text, which is the mistake most likely to cost an hour; read that section before writing a consumer. And prefer quoting a long, distinctive find over a short one: degenerate_find refuses anything under 8 runes, and a short quotation is what makes find_ambiguous fire.