Find any symbol by name or paste — code point and block shown.
🔤 Curated set covering the characters people actually search for. Paste any character from it as the search term to identify its code point. Full Unicode has 150k+ characters — emoji pickers cover that space.
One catalog assigning a number (code point) to every character in every writing system — over 150,000 entries. Fonts decide appearance; Unicode decides identity. U+2014 is an em dash everywhere, however differently fonts draw it.
Homoglyphs (Cyrillic а vs Latin a) look identical but break passwords, domains, and code searches invisibly. When text "matches" but comparisons fail, paste both versions here and compare code points.
150k characters would be noise. These are the arrows, math operators, typographic quotes, currency marks, and box-drawing characters that searches actually target, each with its proper name and code point.
The Unicode Character Lookup handles unicode lookupdirectly in your browser. Paste or type your input, and the tool processes it instantly — no upload, no signup, no waiting. It's built for the moments when you need a quick transformation and don't want to leave your workflow.
Because the tool runs client-side, it's fast and private. Your text never touches a server, which makes it safe for sensitive content. The interface is keyboard-friendly and works on any device with a modern browser.
Common uses: people reach for this tool when they need to use a what is this symbol u2014, checkmark unicode copy paste, arrow character code points list, or emoji vs unicode difference.
Browser-based tools like this one have a few real advantages over installed software or manual methods:
The Unicode Character Lookup is based on the following formula:
code point = U+XXXX (hexadecimal, U+0000 to U+10FFFF) UTF-8 length: 1 byte for U+0000-007F, 2 bytes U+0080-07FF, 3 bytes U+0800-FFFF, 4 bytes U+10000-10FFFF
Variables: U+XXXX: Unicode code point in hexadecimal UTF-8: variable-width encoding using 1 to 4 bytes per code point block: named range the character belongs to, e.g. Latin-1 Supplement category: character class, e.g. Letter, Punctuation, Symbol
Every character has a numeric code point written U+ followed by hex digits, and UTF-8 stores it in as few bytes as the code point's magnitude requires: shorter for ASCII and Latin, longer for CJK and emoji. The leading bits of each UTF-8 byte mark how many bytes the sequence spans, which makes the encoding self-synchronizing.
Worked example: Step 1: Take é, code point U+00E9. Step 2: U+00E9 falls in U+0080-07FF, so UTF-8 uses 2 bytes. Step 3: Fill its 11 bits 000 1110 1001 into the 2-byte template 110xxxxx 10xxxxxx. Result: é (U+00E9) is stored as the 2 bytes 0xC3 0xA9.
More tools you might find useful