HTML Table to CSV
Paste HTML containing a table and get RFC 4180 CSV. Handles th/td, colspan, rowspan, entities and multiple tables. Runs in your browser, nothing uploaded.
Worked examples
- A simple table with thead and tbody
The everyday case: th and td cells are treated alike, and the header row is simply the first CSV line.
- Entities, commas, quotes and a line break
Entities are decoded and whitespace collapsed first; then RFC 4180 quoting wraps fields holding a comma, quote or line break and doubles any inner quotes.
- colspan and rowspan
CSV has no merged cells, so a spanning cell keeps its text in its first position and the other covered positions are left blank (or repeated, if you choose repeat).
What this does
Paste the HTML of a web page, an email or a saved report, and this tool pulls out a table and writes it as CSV. It understands header cells (<th>) and data cells (<td>), decodes HTML entities such as & and , handles merged cells, and produces output that follows RFC 4180, the CSV specification, so it opens correctly in Excel, Google Sheets, pandas and the usual libraries.
How it works
The tool does not need a browser DOM. It tokenizes the markup itself, ignores comments, <script> and <style> content, and collects the rows of the table you choose.
- Which table. Tables are counted in document order starting at 1, counting only top-level tables. The result panel shows how many it found. A table nested inside a cell is not extracted separately; its text is flattened into the cell that contains it, with a space between its cells.
- Cell text. Entities are decoded, runs of whitespace (including newlines from pretty-printed markup) collapse to a single space, and the text is trimmed. A
<br>becomes a real line break inside the cell, which CSV can represent when the field is quoted. - Quoting. A field is wrapped in double quotes when it contains the delimiter, a double quote, a carriage return or a line feed, and any double quote inside it is doubled. Nothing else is quoted, which keeps the output readable.
- Merged cells. A CSV has no merged cells, so the tool lays the table out on a grid the way a browser does. A cell with
colspan="2"orrowspan="2"covers several grid positions. By default its text is repeated in each covered position, so every row is complete and sortable; choose blank mode to keep the text in the top-left position only. Each span is capped at 200, a table whose spans expand to more than 2 million grid cells is rejected with an error, androwspan="0"is treated as 1. - Row and column counts. Rows are written in the order they appear in the source (
thead,tbodyandtfootare all included in that order), and short rows are padded to the widest row.
Edge cases and limits
- Missing end tags are tolerated. An unclosed
<td>or<tr>ends when the next cell or row begins, the way browsers recover. - Only the entities in a built-in list are decoded by name, along with all numeric entities (
—,—). A rare named entity is left as written. - Tables built from
<div>elements or CSS grids are not tables. Only real<table>markup is read. - No JavaScript runs, so a table that a page fills in with scripts after load will be empty in the HTML source. Copy the rendered markup from your browser's developer tools (Inspect, then Copy outerHTML) instead of using View Source.
- Spreadsheet formula injection. A cell that starts with
=,+,-or@can be interpreted as a formula by Excel. This tool does not alter cell text, so review untrusted data before opening it in a spreadsheet. - Line endings default to LF. Choose CRLF if the consumer insists on the exact form the RFC specifies. Input is limited to 2 MB.
When to use it
- Getting a data table out of a web page or HTML email into a spreadsheet.
- Converting an HTML report generated by a build tool, such as a coverage or test report, into something you can diff or chart.
- Rescuing a table from a CMS page where only the rendered HTML survives.
After conversion, the CSV to JSON tool turns the result into objects, and the CSV Diff tool compares it with an earlier export.
Frequently asked questions
- How are colspan and rowspan handled?
- A merged cell cannot exist in CSV, so the tool lays the table out on a grid the way a browser would. By default the merged cell's text is repeated in every grid position it covers, which keeps each row complete for sorting and filtering. Choose blank to keep the text only in the top-left position and leave the other positions empty. A rowspan of 0 is treated as 1.
- Does it follow the CSV standard for quoting?
- Yes. Output follows RFC 4180: a field is wrapped in double quotes when it contains the delimiter, a double quote or a line break, and quotes inside it are doubled. Line endings are LF by default; pick CRLF if your consumer expects the RFC's exact form.
- What if the HTML has several tables?
- The tool reports how many top-level tables it found and converts the one you choose with the table number option, counted in document order starting at 1. A table nested inside a cell is not extracted separately; its text is flattened into that cell.
- Can it read a table from a live web page URL?
- No. This tool works on HTML you paste, entirely in your browser, and makes no network requests. Copy the table's HTML from your browser's developer tools (Inspect, then Copy outerHTML) and paste it in.