HTML to Markdown Converter

Convert HTML to Markdown in your browser: headings, links, images, lists, code blocks and quotes. Unknown tags keep their text. Nothing is uploaded or stored.

Loading tool…

Worked examples

  • An article fragment

    The common blog-post shape: a heading, a paragraph with inline formatting and a link, and a bullet list.

  • Nested list, quote, code block, image and entities

    Shows list nesting, a fenced block with the language taken from the class attribute, and the bullet and emphasis markers you can choose.

  • Tags Markdown has no syntax for

    Unknown tags such as span, u and mark are removed but their text is kept; script content is discarded, and the tool lists what it dropped.

What this does

Paste HTML, such as a blog post, a help-center article or an email body, and get Markdown back. The converter reads the tags that Markdown can express, writes them as CommonMark syntax, and tells you which other tags it removed. Everything runs in your browser, and nothing is uploaded.

How it works

The HTML is tokenized and built into a tree that recovers from sloppy markup: unclosed <p> and <li> tags end when the next sibling starts, stray end tags are ignored, and comments, <script>, <style> and <head> content are discarded. Entities such as &amp; and &#8212; are decoded. The tree is then written out like this:

| HTML | Markdown | |---|---| | h1 to h6 | # to ###### | | p | paragraph separated by a blank line | | strong, b | **bold** | | em, i | *italic* or _italic_, your choice | | code | `code`, with a longer fence if the text contains backticks | | pre > code | fenced block, with the language taken from a language-xxx or lang-xxx class | | a | [text](href "title") | | img | ![alt](src "title") | | ul, ol | nested lists, using your bullet marker, and an ol start value is honored | | blockquote | > prefixed lines | | hr | --- | | br | a line ending with two trailing spaces |

You can choose the bullet (-, * or +) and the emphasis marker (* or _). Strong text always uses **.

Unknown tags

A tag with no Markdown equivalent is dropped and its inner text is kept. That covers span, u, mark, font, small, sup and every custom element. Wrapper elements such as div, section, article and table rows are treated as block containers, so their contents become separate paragraphs. Table cells in a row are joined with spaces. Markdown tables are a GitHub extension, so tables are not converted. The second result line lists every tag name that was removed, which is the quickest way to see what formatting you lost.

Edge cases and limits

  • Escaping. Characters that Markdown would interpret are backslash-escaped: *, `, [, ], \, a < that looks like a tag, and underscores at word boundaries. Underscores inside a word, like snake_case, are left alone because CommonMark does not treat them as emphasis. A paragraph that starts with #, >, -, + or 1. is escaped so it does not turn into a heading, quote or list.
  • Whitespace is normalized. HTML collapses runs of spaces and newlines, so the Markdown does as well, apart from inside pre blocks, which are kept exactly.
  • Inline HTML is not preserved. If you need to keep an element, such as <kbd> or a <details> block, Markdown allows raw HTML, but this tool does not emit it.
  • Link and image URLs have spaces and parentheses percent-encoded so they cannot end the Markdown link early. Links and images that use a javascript:, vbscript:, data: or file: URL are not written as links: a link keeps only its text, and an image keeps only its alt text.
  • Styles and classes are ignored. Bold produced by font-weight in CSS, rather than strong or b, is not detected.
  • Input is limited to 2 MB. If the HTML contains no visible content, the tool says so instead of returning an empty result.

When to use it

  • Migrating content from a CMS or help desk into a docs-as-code repository.
  • Cleaning up HTML copied from a web page before pasting it into a Markdown editor, issue or wiki.
  • Producing a plain-text-friendly version of an HTML snippet for a README.

For the opposite direction, the Markdown to HTML converter implements the same subset, so a round trip of common content returns similar text.

Frequently asked questions

Which HTML elements are converted?
h1 to h6, p, br, strong/b, em/i, code, pre (with language-xxx classes), a, img, ul, ol (including the start attribute and nesting), blockquote and hr. Everything else is treated as unknown.
What happens to tags Markdown cannot represent?
The tag is removed and its inner text is kept, so <span>, <u>, <mark> and <font> simply disappear. Block-level wrappers such as div, section and table rows become separate paragraphs. The tool lists every such tag name it dropped, so you can see what formatting was lost. script, style and comment content is discarded entirely.
Does it convert tables?
No. Markdown tables are a GitHub extension, not part of CommonMark, so table cells are flattened into one line of text per row. For tabular data, use the HTML Table to CSV tool and then the CSV to Markdown Table tool.
Why do some characters get a backslash in front?
Characters that would otherwise be read as Markdown, such as * and [ ] in plain text, or a paragraph that starts with # or 1., are escaped so the output renders the same text as the original HTML. Underscores inside words like snake_case are left alone because CommonMark does not treat them as emphasis.