~bigbes/sr-ht-spec

ref: 658bae75f9714af212b463dbba6bb2de565112f9 sr-ht-spec/prosediff d---------
658bae75 — Eugene Blikh 24 days ago
feat(prosediff): recover the source line each word edit sits on (spec-by6.3.5)

The review UI is moving to a line-numbered unified diff, which needs to know
which line a word-level change happened on. The differ does not keep that.
Tokenize drops whitespace — "\n" and " " both collapse to Token.Space — and
that is precisely what makes a rewrapped paragraph produce a byte-identical
token stream and therefore no diff at all. The property is load-bearing, so the
line is recovered here rather than retained there.

It is recoverable because the script is ordered: the equal and deleted runs
reproduce the old block's tokens in sequence, and the equal and inserted ones
the new block's. Walking each side in step with that side's re-tokenized lines
says which line every token belongs to, and a run crossing a line break is cut
at the boundary.

The script supplies only the operation per token; text and spacing come from
re-tokenizing the source line. Taking text from the spans instead drops
separators — Span.Space is false on an insertion that directly replaces a
deletion, because in a combined rendering the deletion before it carried the
space, and split onto one side that deletion is gone. Caught by a test:
"delta CHANGED zeta" rendered as "deltaCHANGED zeta".

Reports ok=false rather than guessing when a block's Lines and Text disagree
about token count. A caller that cannot split falls back to rendering the block
as one old/new pair labelled by line range: a wrong line number is worse than
an honest range, because it invites a comment onto text that was never there.

spec-by6.3.5
3510f9c3 — Eugene Blikh 27 days ago
feat: prosediff — word-level prose diff over markdown block structure

The Phase 0 de-risk gate. Segments a document into blocks with goldmark
(headings, paragraphs, list items, code fences, table rows, block quotes,
frontmatter), aligns the two block sequences with Myers over content
hashes, and diffs word-by-word inside modified prose blocks and
line-by-line inside modified code fences.

Whitespace and line wrapping alone produce no diff in prose, and always
do in code — that split is the whole point. Moves are detected by
verbatim anchor and grown over their neighbours, so a relocated section
does not explode into add+remove; a move that also bridges one edited
block is recognised, a section rewritten while moving is not, and that
limit is pinned by a test rather than papered over.

SPIKE.md reports the verdict against twelve real revisions of
docs/DESIGN.md: rewrapping the whole 1139-line document produces 1563
changed lines for git and zero changes here; 77% of real prose
modifications read as small edits; 13% shred and want a two-column
fallback in the web layer, which BlockChange.Similarity already gates.
Verdict: the approach works, build the review UI on it.