Free Handy Tools

Text Diff Checker

Detail level
Try an example

Differences

3 changes

2 changed · 1 added · 0 removed · 40% the same, compared line by line

11The quick brown fox jumps over the lazy dog.The quick brown fox leaps over the lazy dog.
22Delivery is expected within 30 days.Delivery is expected within 45 days.
33Payment terms are net 30.Payment terms are net 30.
4Late payment carries interest at 4% a month.
45Signed on 4 March.Signed on 4 March.
How this was worked out
  • Unit comparedwhole lines, trailing spaces included
  • Pieces on each side4 original, 5 changed
  • Treated as equalcapitalisation counts; spacing counts; CRLF and LF are the same line ending
  • Matched up2 identical, 2 replaced, 1 added, 0 removed
  • Same, as a share2 ÷ 5 = 40%

Both texts stay in this page — nothing is uploaded, which matters when the thing being compared is a contract. Lines are matched with a longest-common-subsequence diff, the same approach diff and every code review tool use, so a line inserted in the middle shifts everything below it without being reported as a hundred changes.

The unit compared is a whole line, or a word with the space that follows it. Line endings are normalised first, so a file saved on Windows compared against the same text saved on a Mac or Linux reports no differences rather than reporting every line as changed. Everything else in the text counts: a trailing space, a tab where spaces were, and a final newline are all differences unless “Ignore spacing” is on, and “Ignore spacing” changes what counts as equal without changing what is displayed — a line that differs only in indentation is reported as unchanged and still shown indented.

Finding what changed between two drafts

Two versions of the same document and no record of what happened in between: a contract returned by the other side, a paragraph a colleague rewrote without tracking anything. This checker takes both texts and reports the difference, line by line or word by word, with counts of how many lines changed, were added and were removed, and a percentage of the document that stayed the same.

The options cover the differences that are not really differences. Capitalisation and spacing can each be ignored, and unchanged lines can be hidden so that a long file collapses to the parts worth reading. The result can be copied out as a unified diff for pasting into a ticket or a commit message.

A worked comparison: four lines against five

Comparing a four-line document against a five-line revision of it reported 2 changed, 1 added, 0 removed, and 40% the same.

The second line is the interesting one. "jumps over the lazy dog" against "leaps over the lazy dog" comes back as a single changed line with only the first word marked, rather than as one deletion followed immediately by one insertion. Reading an edited line as a pair of unrelated lines is what makes a lot of diff output so tiring to read.

The unified diff produced for that same comparison is byte-identical to what diff -u prints for the two files.

How the two texts are lined up

The matching is a longest-common-subsequence diff, the same family of algorithm behind diff(1) and the review screens in code hosting tools, described in Myers, "An O(ND) Difference Algorithm and Its Variations", 1986. It looks for the longest run of material the two versions genuinely share and treats whatever is left over as inserted or deleted, rather than comparing line 1 against line 1 and falling apart the moment the two get out of step.

Identical material at the start and end of both documents is matched off before the comparison table is built. That is what makes two long drafts of one document cheap to compare: where a couple of paragraphs moved in the middle of something long, most of the work never happens.

Comparison and display are kept apart. Ignoring case or ignoring spacing changes what counts as equal while the matching runs, and never changes what appears on screen — the lines shown are the lines you pasted.

Privacy, and where the tool stops short

Neither text is uploaded. Both stay in the page, which is the point when the thing being compared is a contract or an unreleased draft. The alternative is posting two versions of a confidential document to somebody else’s server to discover that one comma moved.

Very large pairs are not lined up exactly. Past a certain size the tool declines the exhaustive match and says so on screen: a stated approximation is worth more than a precise answer that arrives once the tab has frozen.

Line against word, and what the percentage means

Should I compare by line or by word?

Line mode for code, configuration and anything where a line is the unit of meaning. Word mode for prose, where a rewritten sentence would otherwise show as a whole line replaced when the real edit was three words.

What does the percentage the same actually measure?

The share of lines that matched, not a measure of meaning. Reformatting a document without altering a single word can drop the figure sharply, and swapping one crucial "not" can leave it in the high nineties.

Why does ignoring spacing still show the spacing?

Because matching and rendering are separate steps. The option tells the comparison to treat two lines with different spacing as equal; the panel still shows each line exactly as pasted, so nothing is quietly rewritten on your behalf.

Does hiding unchanged lines lose anything?

Only context. The counts and the percentage are worked out across the whole document either way, so collapsing the matched sections changes what is on the screen and not what is reported above it.