Docuboxer
By Sergio Alonzo Piña··5 min read

Clean a list: duplicates, sorting and numbering

One pass to tidy a list pasted from anywhere: trim, drop blanks and duplicates, sort, number and join, with the effect of each setting visible.

To clean up a list of text, paste it with one item per line and switch on what you need: trim whitespace, remove blank lines, remove duplicates, sort, number, or join everything into one line. Text Cleaner does it in your browser without uploading the list, which is handy for emails, customer names or product codes. What trips people up is the order the steps run in and how sorting treats accented letters, and that's what this guide covers.

The order of operations is fixed

The tool always applies the same sequence, whichever order you tick the boxes:

  1. Trim whitespace at the start and end of each line.
  2. Collapse repeated spaces inside the line.
  3. Remove blank lines.
  4. Remove duplicates, or keep only the duplicates.
  5. Sort.
  6. Number, or add a prefix and suffix.
  7. Join everything into a single line with a separator.

The reason is practical: joining goes last because once the lines are glued there is nothing left to sort or de-duplicate, and trimming goes first because "apple " with a trailing space wouldn't match "apple".

A real example

Suppose this list, with a blank line, a stray space and repeats:

ñandú
zorro
árbol

Zorro
  arbol 
ñu
zorro

With trim, remove blanks, remove duplicates ignoring case and A-Z sorting, it becomes:

arbol
árbol
ñandú
ñu
zorro

Notice three things. The three zorro entries, lowercase and capitalized, collapsed to one. arbol and árbol remain two items, because the accent makes them different when comparing. And the ñ landed where a Spanish speaker expects it, after n and before o, rather than at the end of the alphabet.

Why accented names sort correctly

Programming languages often sort by the code of each character, which pushes ñ after z and accented words to the end. The tool follows the language's own collation rules, so ñandú and árbol sit at their real alphabetical position. If your lists contain Spanish or other accented names, you'll notice it.

Duplicate options

  • Remove duplicates: keeps the first occurrence of each line.
  • Keep only duplicates: shows the lines that repeat, once each. With the example list it returns zorro. It's useful for finding what repeats.
  • Ignore case and Ignore whitespace when comparing: these change how lines are compared, not the output text.

Sorting: the options

The tool offers A-Z, Z-A, numeric ascending and descending, by length, random, and reverse of the current order. Watch out with numbers: in A-Z, 10 cats and 100 parrots come before 9 dogs, because they're compared as text. For 9 to go first, use numeric sorting, which reads the number at the start of each line.

Numbering, prefix, suffix and joining

You can number with 1. 2. 3. or 1) 2) 3), and add text before and after every line. Combining prefix, suffix and joining builds paste-ready lists. With the lines one, two and three, an apostrophe as both prefix and suffix and a comma with a space as separator, you get 'one', 'two', 'three': the shape of a list of values in a SQL query, or a string for a spreadsheet formula.

Edge cases worth knowing

Windows-style line endings are normalized, so a list copied from a file saved on Windows behaves like any other. Lines that contain only spaces count as blank and are removed when you drop empty lines, whether or not trimming is on. And collapsing spaces turns any run of spaces or tabs inside a line into a single space, which is useful for text copied from tables.

Where it comes in handy

  • Emails from a document. After extracting emails from a document, remove duplicates, sort and join them with semicolons to paste into a recipients field.
  • Keywords. Paste a list of terms, drop repeats ignoring case and sort them.
  • Codes or SKUs. Strip stray spaces and repeats before loading a list into a system.
  • Guest and attendance lists. Sort by surname if each line already starts with it, and number the result.

In every case, check the output before you use it: the tool can't know whether two different lines refer to the same person.

What about a spreadsheet?

Excel and Google Sheets have their own options to remove duplicates and sort, and they work well once your data sits in columns. This tool earns its place when the text comes pasted from somewhere else: an email, a PDF, a chat. If the list is trapped in a PDF, extract the text from the PDF first. And to compare two sorted lists, see see what changed between two versions of a text.

Frequently asked questions

How do I remove duplicates from a list?

Paste the list with one item per line and choose Remove duplicates. The first occurrence of each line is kept. Switch on ignore case if "Zorro" and "zorro" should count as the same.

How do I sort a list alphabetically with accented letters?

Use the A-Z option, which follows the language's collation: ñ falls between n and o, and accented words sit with the rest.

Why does 10 sort before 9?

Because A-Z compares as text and "1" comes before "9". Use numeric sorting to order by the number at the start of each line.

How do I join all the lines into one?

Turn on join into a single line and choose the separator, for example a comma and a space. It's always the last step.

How do I find the repeated items in a list?

Use Keep only duplicates, which shows each repeated line once.

Is my list uploaded to a server?

No. It's processed in your browser.

Clean up your list now

Duplicates, sorting, numbering, prefixes and joining into one line. Runs locally, free, no signup.

Open Text Cleaner →

Related tools

You might also like: change text case in Word, Excel and Google Docs, 10 text utilities you'll actually use every week, Word count, reading time and character limits and Regex examples you can test: email, URL, IP.