HTML Encoder & Decoder

Encode text to HTML entities and decode them back instantly

What Is an HTML Encoder?

An HTML encoder converts characters that have a special meaning in HTML into HTML entities — safe text references that a browser renders as literal characters instead of interpreting as markup. An HTML decoder does the reverse, turning entities back into the characters they represent. This tool does both directions instantly, entirely in your browser, so you can encode HTML online without installing anything or sending your data to a server.

You reach for an HTML encoder whenever you need to display code, user comments, or arbitrary text inside a web page without the browser mistaking angle brackets and ampersands for real tags. Paste your text, click Encode, and copy the result — or paste entity-laden text and click Decode to get the original back.

What Are HTML Entities?

HTML entities are character references used in HTML documents to represent characters that either have special meaning in HTML syntax or can’t be easily typed on a standard keyboard. Every HTML entity begins with an ampersand (&) and ends with a semicolon (;).

The five characters that almost always need encoding are &amp; for the ampersand (&), &lt; for the less-than sign (<), &gt; for the greater-than sign (>), &quot; for double quotes ("), and &#39; (or &apos;) for single quotes ('). These are the reserved characters of HTML — the ones that, left raw, change how the parser reads your document.

Why Encode HTML?

Encoding HTML serves three main purposes in web development:

Security (XSS prevention). The most important reason is stopping Cross-Site Scripting (XSS) attacks. If user input containing HTML or JavaScript is rendered without encoding, an attacker can inject a <script> tag that runs in your visitors’ browsers. Encoding turns that payload into inert text — <script> becomes &lt;script&gt; and is displayed, not executed.

Correct rendering. Characters like < and > would otherwise be interpreted as the start of HTML tags. Encoding them guarantees they appear on screen exactly as written — essential for tutorials, documentation, and anything that shows code samples.

Special and non-ASCII characters. Symbols outside the basic ASCII range — em dashes, copyright and trademark signs, currency symbols, accented letters — can be represented with entities to ensure consistent rendering regardless of the page’s character encoding.

Which Characters Do You Actually Need to Encode?

Not every character needs an entity, and over-encoding makes markup harder to read. Which characters are mandatory depends on where the text is going:

  • In HTML body text (between tags), only three characters are structurally dangerous: <, > and &. Encoding just these is enough to stop text from being parsed as markup.
  • Inside an attribute value, the quote character that delimits the attribute also becomes dangerous. If an attribute is written with double quotes, an un-encoded " in the value closes the attribute early — a classic injection vector. So inside attributes you must also encode " (as &quot;) and, for single-quoted attributes, ' (as &#39;).

Because a template can switch quote styles without warning, the safe universal set is the five reserved characters — & < > " ' — which is exactly what this encoder outputs. That covers both body and attribute context in one pass, so you never have to reason about which quote style the surrounding markup uses.

A word of caution about attribute context: HTML encoding protects quoted attribute values. If your template emits unquoted attributes, entity encoding alone is not enough — a space or > in the value can still break out. Always wrap attribute values in quotes and encode the reserved set; the two defenses work together.

How to Use the HTML Encoder / Decoder

  1. Paste your text or HTML entities into the input area.
  2. Click Encode to convert special characters to HTML entities, or Decode to convert entities back to characters. Press Ctrl+Enter to run the primary action from the keyboard.
  3. Copy the result with the Copy button or Ctrl+Shift+C. Use Ctrl+L to clear and start over.

Example: displaying a code snippet safely

Say you want to show this line of markup as text in a tutorial:

<a href="index.html" title="Home & About">Go</a>

Pasting it straight into your page would render a real link. Run it through Encode and you get:

&lt;a href=&quot;index.html&quot; title=&quot;Home &amp; About&quot;&gt;Go&lt;/a&gt;

Drop that into your HTML and the browser prints the tag literally instead of building a link. Notice all four reserved characters were handled at once — <, >, the & in “Home & About”, and the double quotes around the attribute values. Click Decode on the encoded string and you get the original snippet back byte-for-byte.

Named vs. Numeric vs. Hexadecimal Entities

The same character can be written three ways, and this tool decodes all of them:

  • Named — a human-readable alias, e.g. &copy; for ©. Easiest to read, but only characters with an assigned name can use this form.
  • Numeric (decimal) — &# followed by the Unicode code point in base 10 and a semicolon, e.g. &#169; for ©.
  • Numeric (hexadecimal) — &#x followed by the code point in base 16, e.g. &#xA9; for ©. Hex is common because Unicode charts list code points in hexadecimal.

Numeric references can encode any Unicode character — including emoji and rare symbols that have no named entity — which is why encoders often fall back to them.

Encoding Emoji and Non-ASCII Symbols

Most emoji and less common symbols have no named entity, so a numeric reference is the only entity form available. The key rule: a numeric reference uses the character’s actual Unicode code point, not its UTF-16 storage bytes.

Take the grinning-face emoji 😀, code point U+1F600. Its correct HTML entity is the full code point — &#128512; in decimal or &#x1F600; in hexadecimal. A common bug in hand-rolled JavaScript encoders is to iterate with charCodeAt, which returns UTF-16 surrogate halves (0xD83D and 0xDE00) for characters above U+FFFF, and then emit two broken &#55357;&#56832; references that no browser will reassemble into the emoji. Use codePointAt (or spread the string into an array with [...str]) so you read whole code points, and this tool handles that for you automatically.

In practice you rarely need entities for emoji at all: if your page is served as UTF-8 (<meta charset="utf-8">), you can paste the raw emoji directly and it renders fine. Reach for numeric entities only when a pipeline strips non-ASCII bytes or when you must keep the source file pure ASCII. When you need \uXXXX-style escapes for source-code string literals instead of HTML, use the Unicode escape tool — that is a different escaping context, covered below.

Common HTML Entities Reference

CharacterNamed EntityDecimalHexDescription
&&amp;&#38;&#x26;Ampersand
<&lt;&#60;&#x3C;Less than
>&gt;&#62;&#x3E;Greater than
"&quot;&#34;&#x22;Double quote
'&apos;&#39;&#x27;Single quote / apostrophe
(space)&nbsp;&#160;&#xA0;Non-breaking space
©&copy;&#169;&#xA9;Copyright
®&reg;&#174;&#xAE;Registered trademark
™&trade;&#8482;&#x2122;Trademark
—&mdash;&#8212;&#x2014;Em dash
€&euro;&#8364;&#x20AC;Euro sign

Encoding HTML Programmatically

Once you understand what the tool does, you can reproduce it in code. A few common approaches:

JavaScript (browser). Let the DOM escape for you:

const encode = (s) => { const d = document.createElement('div'); d.textContent = s; return d.innerHTML; };
const decode = (s) => { const d = document.createElement('div'); d.innerHTML = s; return d.textContent; };

For Node.js there’s no DOM, so use a maintained library such as he or html-entities, which ship the full named-entity table.

Python. The standard library has it built in:

import html
html.escape("<a href=\"x\">&")   # '&lt;a href=&quot;x&quot;&gt;&amp;'
html.unescape("&lt;p&gt;")        # '<p>'

PHP. Use htmlspecialchars() for the five reserved characters, or htmlentities() to encode everything with a named equivalent; html_entity_decode() reverses either.

Prefer these battle-tested functions over hand-rolled string replacement — subtle mistakes in encoding order (encoding & last instead of first) are a classic source of bugs.

HTML Encoding vs. URL Encoding vs. Unicode Escaping

These three are easy to confuse because they all “escape” characters, but they target different contexts:

  • HTML encoding makes text safe for display inside an HTML document. A space stays a space; < becomes &lt;.
  • URL encoding (percent-encoding) makes text safe inside a URL or query string. A space becomes %20, and & becomes %26. Use the URL encoder/decoder for that.
  • Unicode escaping (\uXXXX) makes characters safe inside source-code string literals — JSON, JavaScript, or Java. The Unicode escape tool handles that form.

Applying the wrong one — or applying two of them to the same string — produces garbled output, so pick the encoding that matches where the text will live.

Common Errors and How to Fix Them

Double-encoding. If content is already encoded and you encode it again, &amp; becomes &amp;amp; and shows up literally on the page. When you see stray amp; in your output, decode once and stop re-encoding already-safe strings.

Missing semicolons. An entity without its trailing semicolon (&lt instead of &lt;) may or may not be recognized depending on the parser. Always terminate entities with ; — decode here and re-encode cleanly if you inherit malformed input.

Mojibake from wrong charset. Entities like &#8217; rendering as ’ usually means the page is served with the wrong charset. Serve UTF-8 (<meta charset="utf-8">) and the underlying characters render without needing entities at all.

Encoding the ampersand last. When rolling your own encoder, always replace & first; otherwise you re-escape the ampersands you just introduced. This tool handles ordering for you.

Best Practices for HTML Encoding

  • Always encode user-generated content before rendering it in HTML — treat every external string as untrusted.
  • Use your framework’s built-in encoding functions (template auto-escaping, htmlspecialchars, html.escape) rather than manual string replacement.
  • Be aware of context-specific encoding — HTML body text, HTML attributes, JavaScript strings, and CSS values each require different escaping rules.
  • Don’t double-encode; if content is already encoded, decoding and re-encoding will corrupt it.
  • Layer a Content Security Policy (CSP) on top of encoding for defense in depth against XSS.

Frequently Asked Questions

How do I encode HTML online?

Paste your text into the input box and click Encode. The tool converts characters that have special meaning in HTML — &, <, >, " and ' — into their entity equivalents (&amp;, &lt;, &gt;, &quot;, &#39;) so they display as literal text instead of being parsed as markup. Everything runs in your browser, so no data is uploaded.

What are HTML entities?

HTML entities are special codes used to represent characters that have special meaning in HTML or that can't be easily typed on a keyboard. They start with an ampersand (&) and end with a semicolon (;). For example, &amp; represents the ampersand character (&), and &lt; represents the less-than sign (<).

Why do I need to encode HTML entities?

HTML encoding prevents XSS (Cross-Site Scripting) attacks, displays special characters correctly on web pages, and ensures that characters like <, >, & and quotes are rendered as text rather than interpreted as HTML markup. Any time you output untrusted or user-supplied text into a page, it should be HTML-encoded first.

What is the difference between named and numeric HTML entities?

Named entities use a descriptive name (like &amp; for &), while numeric entities use a Unicode code point in decimal (&#38;) or hexadecimal (&#x26;). Named entities are more readable, but there are only a few hundred of them; numeric entities can represent any Unicode character, which makes them useful for symbols and emoji that have no name.

How do I decode HTML entities?

Paste your text containing HTML entities into the input area and click Decode. The tool converts every entity back to its original character — named (&lt;), decimal (&#60;) and hexadecimal (&#x3C;) forms all resolve. For example, &lt;p&gt; becomes <p>.

What is the difference between HTML encoding and escaping?

The terms are used interchangeably. "HTML encoding" and "HTML escaping" both mean replacing reserved characters with entity references so they are treated as data, not markup. It is distinct from URL encoding (percent-encoding for URLs) and Unicode escaping (\uXXXX for source code), which solve different problems.

How do I encode HTML entities in JavaScript?

The DOM does it for you: set an element's textContent to your string and read back its innerHTML, and the browser escapes &, < and > automatically. For decoding, assign the entity string to innerHTML of a detached element and read textContent. For server-side Node, libraries like he or html-entities cover the full named-entity set.

Is my data safe when encoding?

Yes. All encoding and decoding happens entirely in your browser using JavaScript. Your text is never sent to any server, logged, or stored — close the tab and it's gone.

Why does my encoded output contain &amp;amp; or extra amp;?

That is double-encoding. The string was already HTML-encoded once, and encoding it a second time turns the ampersand of every existing entity into &amp;amp;, so &amp;lt; becomes &amp;amp;lt; and shows up literally on the page. Paste the text here and click Decode once to strip the extra layer, then encode only if the source really was raw. Encode each string exactly once.

Do I need to encode single and double quotes?

It depends on context. In ordinary HTML body text, < & > are enough. But when text is placed inside an attribute value, the quote that delimits the attribute must be encoded or it will terminate the attribute early and open an injection hole — encode double quotes as &quot; inside double-quoted attributes and single quotes as &#39; inside single-quoted ones. Because you rarely control which quote style a template uses, encoding both quotes (as this tool does) is the safe default.

What is the difference between htmlspecialchars() and htmlentities() in PHP?

htmlspecialchars() encodes only the reserved characters — & < > and, with ENT_QUOTES, both quote types — leaving every other character as-is. htmlentities() additionally converts every character that has a named entity (accented letters, symbols, dashes) into that entity. For XSS-safe output you only need htmlspecialchars(); htmlentities() is heavier and mostly useful for legacy non-UTF-8 pages. Both are reversed by html_entity_decode().