Remove Duplicate Lines

Remove duplicate lines from a list, keeping the first or last copy. Optional case-insensitive and trimmed matching, with counts of what was removed.

Loading tool…

Worked examples

  • Exact duplicates, keep first

    The default: each distinct line is kept at the position where it first appears, so the original order is preserved.

  • Keep the last occurrence instead

    Keeping the last occurrence changes which position survives: useful for log or config lines where the latest entry wins.

  • Case-insensitive, ignoring stray spaces and blanks

    A typical mailing-list cleanup: addresses that differ only by case or surrounding spaces count as duplicates, and blank lines are dropped.

What this tool does

This tool removes repeated lines from a block of text. You choose whether to keep the first or the last copy of each line, whether case matters, whether to ignore surrounding whitespace, and whether to drop blank lines. It reports how many lines you started with, how many remain, and how many were removed as duplicates or as blanks.

When you need it

  • De-duplicating an email list, a list of URLs, hostnames, IPs or keywords before importing or processing them.
  • Cleaning a log excerpt or config file where the same line was added several times.
  • Producing a unique set of values from a pasted spreadsheet column.
  • Tidying an .gitignore or allow-list after merging two versions.

How matching works

Two lines are duplicates when their comparison keys are equal. The key is the line text after the options are applied:

  1. If trim whitespace is on, leading and trailing spaces and tabs are removed from every line first, and the trimmed version is what appears in the output as well. With it off, whitespace is significant and never altered: apple and apple (with a trailing space) are different lines.
  2. The text is normalized to Unicode NFC. This means an accented letter typed as one character and the same letter typed as a base letter plus a combining accent count as the same line, which they visually are. Without this, text pasted from different sources, some of which store accents in the decomposed form, would silently pass as distinct.
  3. If case-insensitive is on, the key is lowercased. Case affects only the comparison: the line that survives is output exactly as written.

The tool does not do fuzzy matching. Lines that differ by a non-breaking space, a different dash character or a typo are distinct.

Keep first or keep last

With keep first, each distinct line stays at the position where it first appeared, and later copies are removed, preserving original order. With keep last, the final copy survives at its own position and earlier ones are removed. For the input apple, banana, apple, cherry, banana, apple, keep first yields apple, banana, cherry while keep last yields cherry, banana, apple. Keep last is useful when later entries supersede earlier ones, such as a log of settings changes or an appended list of overrides. If case-insensitive matching is on and the copies differ in case, the one you keep is the one you asked for: Apple then apple gives Apple with keep first and apple with keep last.

Blank lines

Blank lines are lines like any other: by default, many empty lines in a row collapse to a single empty line, because they are all equal. If you turn on remove blank lines, empty and whitespace-only lines are dropped entirely and counted separately from duplicates, so the counts always add up: original lines = unique lines kept + duplicates removed + blank lines removed.

Order and sorting

This tool never reorders lines. That is the main difference from the Unix sort -u idiom, which sorts as a side effect. If you want a sorted unique list, run this tool first, then sort the result with the sort lines tool, or the other way around; the outcome is the same set of lines.

Line endings and limits

Lines may end in \n, \r\n or \r. Output always uses \n. A single trailing newline in your input is preserved in the output and does not count as an extra empty line. Input is capped at 200,000 characters.

Frequently asked questions

Does it keep the original order?
Yes. Unlike `sort -u`, which sorts as a side effect, this keeps surviving lines in their original order. With keep first, each line stays at its first position; with keep last, it stays at its last position. To sort afterwards, use the sort lines tool.
What does the trim whitespace option change in the output?
When it is on, leading and trailing spaces and tabs are removed from every line, both for comparison and in the result, so ' apple ' and 'apple' become the single line 'apple'. When it is off, lines must match exactly, spaces included, and whitespace is never altered.
Does case-insensitive matching change the lines it keeps?
No. Case only affects the comparison; the surviving line is output exactly as you typed it. If Apple comes first and apple later, keep first outputs Apple and keep last outputs apple.
Why are two lines that look identical not treated as duplicates, or vice versa?
Matching is on the exact characters, so a trailing space, a non-breaking space or a different hyphen character makes lines distinct (turn on trim for plain spaces). One exception goes the other way: an accented letter typed as a single character and as a letter plus combining accent look the same and are treated as equal, because lines are compared after Unicode NFC normalization.