The character encoding of an HTML page is the rule that turns the bytes in the file back into letters. For every modern page the answer is the same: save the file as UTF-8 and declare it with <meta charset="utf-8"> as the first element inside <head>, within the first 1024 bytes of the document. The HTML standard requires UTF-8, and if your server also sends a Content-Type header with a charset, the two must agree.
Key takeaways
- Use UTF-8. The HTML standard says the document’s actual encoding “must be UTF-8”, and W3Techs reported in September 2026 that 99.1% of websites with a known encoding already use it.
- Declare it with
<meta charset="utf-8">right after<head>. It must fit in the first 1024 bytes. - Declaring is not enough: the file itself, your database and your server header all have to be UTF-8 too.
- When declarations disagree, a byte-order mark beats the HTTP header, which beats the
<meta>tag. - Garbled text like
éinstead oféalmost always means UTF-8 bytes are being read as a legacy encoding, or text was converted twice.
What a character encoding is
Computers store text as numbers. An encoding is the table that maps those numbers to characters. Older encodings such as ISO-8859-1 or Windows-1252 cover one byte per character and only a couple of hundred characters, enough for Western European languages and nothing else. UTF-8 encodes every character in Unicode, so English, Arabic, Hindi, Japanese and emoji can sit on the same page, using one to four bytes per character.
The W3C puts the risk plainly: if you don’t specify the encoding, “you risk that characters in your content are incorrectly interpreted”, and it adds that an encoding declaration is also needed to process non-ASCII text that users type into forms (W3C Internationalization).
The one line you need
<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="utf-8" />
<title>…</title>
</head>
</html>A few details people get wrong:
- Position. The HTML standard says the element containing the declaration “must be serialized completely within the first 1024 bytes of the document” (HTML Living Standard). Put it first in
<head>, before the title, scripts and long comments. - Only one. A document can have only one
<meta>-based encoding declaration. - Case doesn’t matter.
UTF-8andutf-8are both fine. The W3C notes the longer form,<meta http-equiv="Content-Type" content="text/html; charset=utf-8">, works the same way but is more to type. - No other value is valid for new pages. The Encoding Standard says “Authors must use the UTF-8 encoding” and must label it
utf-8(WHATWG Encoding Standard). Older labels likeiso-8859-1still work in browsers, but they aren’t conforming for new documents.
The declaration has to match the bytes
The <meta> tag describes the file. It doesn’t convert it. If your editor saves the file as Windows-1252 (which some Windows tools call “ANSI”) and the tag says UTF-8, every accented or non-Latin character will break. The W3C reminds authors that using UTF-8 means “you also need to save your content as UTF-8” (W3C).
Check your editor’s save format. In most code editors the current encoding is shown in the status bar, with an option to “Save with encoding” or “Reopen with encoding”. Choose UTF-8.
ANSI vs UTF-8. “ANSI” is not a single encoding. On Windows it means the system’s legacy code page, which on Western European systems is usually Windows-1252. It can’t represent Arabic, Chinese or most other scripts, and files saved that way will display incorrectly when the page declares UTF-8. For the web, always choose UTF-8.
Which declaration wins when they disagree
A browser can find the encoding in three places. The W3C lists the precedence:
- Byte-order mark (BOM). A few invisible bytes at the very start of the file. If present, it overrides everything else, “including the HTTP header”.
- HTTP
Content-Typeheader, for exampleContent-Type: text/html; charset=utf-8. It beats in-document declarations. - The
<meta charset>element.
So if your server sends charset=iso-8859-1 and your page says UTF-8, the server wins and the page breaks. The W3C’s advice is to declare the encoding in the document in any case, and if you also use the header, to make sure it says the same thing.
How to set the header on common servers:
- Apache:
AddDefaultCharset utf-8adds the charset totext/htmlandtext/plainresponses (Apache docs). Apache notes you should only use it when all the affected files really are in that encoding. - nginx: the
charset utf-8;directive in thehttp,serverorlocationblock adds the charset to theContent-Typeheader (nginx docs). - PHP:
header('Content-Type: text/html; charset=utf-8');before any output.
You can see what your server sends with curl -I https://your-site.example/ or in the browser’s developer tools under Network, response headers.
Fixing garbled characters (mojibake)
Garbled text follows recognisable patterns. Here is how to read them:
| What you see | What happened | Fix |
|---|---|---|
café instead of café | UTF-8 bytes read as Windows-1252 or ISO-8859-1 | Make the header and meta both say UTF-8 |
caf� (a diamond with a question mark) | Legacy bytes read as UTF-8 | Re-save the file, or convert the data, as UTF-8 |
café | Text converted to UTF-8 twice | Find the double conversion, usually on import or in the database connection |
??? in place of every non-Latin character | Characters lost when saved to a column or file that can’t hold them | Change the storage to UTF-8 and re-import from the source |
Work through the chain in order: the file on disk, the database, the connection between them, the HTTP header and the <meta> tag. Every link has to be UTF-8.
Databases. In MySQL, the old utf8 character set is an alias for utf8mb3, which stores at most three bytes per character and so cannot hold “supplementary” characters such as most emoji. MySQL documents utf8mb3 as deprecated (MySQL Reference Manual). Use utf8mb4 for tables and for the connection.
HTML entities: when you still need them
With UTF-8 you can type é, ü, ж or 中 directly into the source. You only need character references such as &, < and > for characters that have meaning in HTML markup. Entities are also handy for invisible characters that are hard to spot in source, such as a no-break space ( ) or directional marks in right-to-left text.
Encoding on multilingual sites
Encoding problems show up the moment a site adds languages. A page that has only ever displayed English can hide a wrong declaration for years, because plain English letters are the same in UTF-8 and in the legacy Western encodings. The first Polish, Greek or Japanese translation exposes it.
Before you add languages, check three things:
- Every template and static file is saved as UTF-8 and declares
<meta charset="utf-8">. - The
langattribute is set per language, for example<html lang="ja">, and right-to-left pages also getdir="rtl". The encoding says how to read the bytes;langsays what language the text is in, which matters for fonts, hyphenation and screen readers. - URLs with non-Latin characters are UTF-8 and percent-encoded. Google’s guidance on multilingual sites says localized words in URLs are fine, but to “use UTF-8 encoding in the URL (in fact, we recommend using UTF-8 wherever possible)” and to escape URLs properly when linking (Google Search Central).
ConveyThis translates the text on your pages into any of its 210 supported languages, including scripts such as Arabic, Chinese, Hindi and Hebrew, and serves translated pages from their own language URLs. It can’t fix a template that declares the wrong encoding, though, so run the checks above first. For right-to-left languages there is a setting to switch text direction on translated pages; our RTL design guide covers the layout side. See the full list of features and how translated pages are handled for multilingual SEO, or compare plans and pricing.
Quick checklist
- File saved as UTF-8 (without a BOM, unless you have a reason to keep one)
<meta charset="utf-8">is the first element in<head>- Server header says
charset=utf-8, or nothing, never a different charset - Database tables and connection use
utf8mb4(MySQL/MariaDB) - Forms and APIs send and accept UTF-8
lang(anddirfor RTL languages) set on each language version
Frequently asked questions
What is the correct character encoding for HTML?
UTF-8. The HTML Living Standard requires that the document’s actual encoding be UTF-8 and that any declaration use the utf-8 label.
Where should meta charset go?
As the first element inside head. The whole element must be within the first 1024 bytes of the document, so put it before the title, scripts and styles.
Is meta charset=“utf-8” still needed with an HTTP header?
The W3C recommends declaring the encoding inside the document even when the server sends it, because the header is lost when the file is saved or opened locally. If you use both, they must say the same thing.
Why does my page show é instead of é?
The text is UTF-8 but is being decoded as a legacy Western encoding such as Windows-1252. Usually the server header or meta tag declares the wrong charset. Make both say UTF-8.
What is the difference between ANSI and UTF-8?
ANSI usually refers to a Windows legacy code page such as Windows-1252, which covers only Western European characters. UTF-8 can represent every Unicode character in every language, and it is the only encoding the HTML standard allows for new documents.
Ready to add languages?
Once your pages are clean UTF-8, adding languages is mostly a content job. You can create a ConveyThis account to translate your site and publish each language on its own URL.