Estimate byte size of your text in various formats.
📏 Estimates storage size in different units. Useful before storing text in databases or APIs.
This tool estimates the storage size of your text in different formats — bytes (UTF-8), kilobytes, Base64, and character/word/line counts. Useful for planning database fields and API limits.
Characters take different byte counts in UTF-8: ASCII letters = 1 byte, accented Latin = 2 bytes, Chinese/CJK = 3 bytes, emoji = 4 bytes. A 1000-character Chinese text is ~3000 bytes, not 1000.
The Text Size Estimator lets you figure out text size estimatorinstantly, without reaching for a spreadsheet or doing the math by hand. Whether you're planning a budget, checking a loan, or working through homework, the tool applies the correct formula behind the scenes and shows the result the moment you enter your numbers.
Unlike a static chart or table, this calculator adapts to your exact inputs. You can adjust any value and see the outcome update in real time, which makes it easy to compare scenarios — for example, "what if the rate were 1% lower?" or "what if I paid an extra $50 a month?"
Common uses: people reach for this tool when they need to find a text byte size calculator, string size in bytes, utf 8 byte length, or text size for database column.
Browser-based tools like this one have a few real advantages over installed software or manual methods:
The Text Size Estimator is based on the following formula:
size ≈ Σ bytes per character UTF-8: ASCII 1 byte, Latin/Greek/Cyrillic/Arabic 2 bytes, CJK 3 bytes, emoji 4 bytes Base64 size ≈ 4 × ⌈bytes ÷ 3⌉
Variables: size: UTF-8 storage size of the text (bytes) ASCII: characters U+0000-007F, 1 byte each Latin/Greek/Cyrillic/Arabic: accented and non-Latin letters, 2 bytes each CJK: Chinese, Japanese, Korean characters, 3 bytes each emoji: 4 bytes each Base64: encoded size ≈ 4 × ⌈bytes ÷ 3⌉ characters
UTF-8 is variable-width: common characters cost fewer bytes, so storage size depends on the script mix, not the character count. Base64 packs every 3 bytes into 4 characters, which is why encoded text is always about a third larger than the raw bytes.
Worked example: Step 1: The text héllo has 5 characters. Step 2: h, l, l, o are ASCII at 1 byte each; é is Latin at 2 bytes. Step 3: size = 1 + 2 + 1 + 1 + 1 = 6 bytes. Step 4: Base64 ≈ 4 × ⌈6 ÷ 3⌉ = 8 characters. Result: 6 bytes in UTF-8, about 8 characters in Base64.
More tools you might find useful