/tools/remove-duplicate-lines

Remove Duplicate Lines

Clean a list down to unique entries. Useful for email lists, keyword sets, log lines and anything pasted together from several sources.

Remove Duplicate LinesRuns in your browser

How to use the Remove Duplicate Lines

  1. Paste your list with one item per line.
  2. Set the options — ignoring case and trimming whitespace catches near-duplicates that look identical but are not.
  3. Press Process, then copy or download the cleaned list.

Why "identical" lines often are not

Two lines that look the same on screen frequently differ in ways you cannot see. The usual culprits:

  • Trailing spaces, added when text is copied from a table or a PDF.
  • Case differences — Ada@example.com and ada@example.com are the same mailbox but different strings.
  • Non-breaking spaces (U+00A0) pasted from web pages, which look identical to ordinary spaces.
  • Line ending style — Windows files use CRLF, leaving an invisible carriage return at the end of each line.

Trimming whitespace and ignoring case are both on by default because these account for most missed duplicates. Turn them off when exact matching matters, as it does for passwords, tokens and case-sensitive identifiers.

The three output modes

Unique lines keeps the first occurrence of each value and drops the rest — the standard deduplication.

Only duplicated lines inverts the question: it shows which values appeared more than once. This is the mode for auditing rather than cleaning — finding double-booked records, repeated log entries or accidentally duplicated rows.

Only lines that appeared once gives the true singletons, dropping every value that repeated at all. Useful for finding entries present in one list but not another after concatenating both.

Frequently asked questions

Does it keep my original order?

Yes, unless you switch on sorting. The first occurrence of each value stays in the position it appeared.

How large a list can it handle?

Hundreds of thousands of lines are fine — deduplication uses a hash map, so it scales linearly. Very large pastes are limited by the browser textarea rather than the algorithm.

Can I deduplicate on part of each line?

Not directly. Sort or split the data so the key you care about is the whole line, or use a spreadsheet for column-based deduplication.

Is the sort case-sensitive?

No. Sorting uses natural ordering that ignores case and handles embedded numbers sensibly, so item2 comes before item10.

Related tools