Spreadsheet & CSV Tools

CSV Cleaner

Tidy a messy CSV before it causes an import error. Tick the problems you want fixed, such as stray spaces, empty rows, duplicate records and rows with missing columns, and download a clean file together with a report of exactly what was changed.

  • Runs in your browser
  • No sign-up
  • Free to use

How to use CSV Cleaner

  1. Drop the CSV or TSV file.
  2. Tick the cleaning rules to apply and choose a style for header names if you want them standardised.
  3. Select Clean file.
  4. Read the change report, check the preview and download the cleaned file.

CSV Cleaner features

Whitespace repair

Trims spaces around values and optionally collapses repeated spaces inside them.

Invisible characters

Removes zero-width characters, byte-order marks and control characters, and turns non-breaking spaces into ordinary ones.

Empty rows and columns

Deletes rows with no values and columns that are empty in every data row.

Duplicate removal

Keeps the first of each set of identical rows.

Even rows

Pads short rows and names extra columns so that every row has the same number of fields.

Change report

Lists each rule that did something, with counts, so nothing happens unnoticed.

When to use CSV Cleaner

  • Preparing a contact list for an email platform that rejects malformed files.
  • Cleaning data copied out of web pages, PDFs or old systems.
  • Removing duplicates after merging several exports.
  • Standardising header names before loading a file into a database.

CSV Cleaner FAQ

Which problems does the cleaner fix?

Spaces before and after values, invisible characters, empty rows, empty columns, exact duplicate rows and rows with too few or too many fields. Optional rules collapse repeated spaces, replace curly quotation marks with straight ones and rewrite header names in a consistent style.

What are invisible characters and why do they matter?

Characters such as the zero-width space, the byte-order mark and the non-breaking space look like nothing or like an ordinary space, yet they make two values that look identical compare as different. They are a frequent reason why lookups fail, duplicates are not detected and logins or codes do not match.

How are duplicates identified?

Rows are compared after the other cleaning rules have run, so two records that differ only by a trailing space are recognised as the same. Every value in the row must be identical. The first occurrence is kept, and the header row is never removed.

What happens to rows with a different number of columns?

Rows shorter than the header are filled up with empty values. If a row has more values than the header has names, the header is extended with names such as column_7 so that no data is lost. The report tells you how many rows were affected, which is a hint to look for unquoted separators in the source.

Does the tool change my values?

Only in the ways you tick. It never reformats numbers or dates, changes capitalisation of data, or guesses at corrections. The report lists every rule that modified the file.

Is the original file modified?

No. The original stays as it is. You download a new file whose name ends in -clean.

Why CSV files get dirty

Data rarely arrives clean, because it rarely comes from one careful source. A contact list has been typed by many people, pasted from emails, exported from one system and edited in a spreadsheet. Along the way it collects the debris of that history: a space after a surname, an empty line where someone deleted a record, the same person entered twice, a description containing a comma that pushed the remaining fields one column to the right.

Each of these is harmless to a human reader and troublesome to software. An importer that matches customers by email treats “ada@example.com ” with a trailing space as a new address. A database refuses a file in which row 300 has seven fields and the table has six. A deduplication step misses pairs that differ by an invisible character. Cleaning is the unglamorous work of removing such differences so that equal things are recognised as equal.

The order of operations matters, and the cleaner applies its rules in a deliberate sequence. Characters are normalised first and spaces trimmed, since those steps make more rows identical. Empty rows are then dropped, rows are brought to a uniform length, and empty columns removed. Duplicates are detected last, on the fully normalised data, which catches far more of them than a comparison of the raw rows would.

Automatic cleaning has limits that are worth stating. It cannot know that “Jon Smith” and “John Smith” are the same person, or that a phone number is missing a digit. What it does is clear away the mechanical noise, so that the remaining problems, the ones that need human judgement, become visible.

Other useful tools