Free Handy Tools

Shannon Entropy Calculator

Try an example

Nothing pasted here reaches the address bar. A link to this tool carries the example and the size of the table below, never the text in this box.

Entropy

3.84 bits per character

245.8 bits in total across 64 characters, drawn from 16 distinct symbols.

Bits per character3.84
Total bits245.8
Distinct symbols16
UTF-8 bytes64
How this was worked out
  • Characters counted64 code points, 16 of them distinct (64 bytes as UTF-8)
  • Commonest symbol“1” appears 9 times, a share of 14.1%, contributing −p·log2(p) = 0.398 bits
  • Summed over every symbolH = −Σ p·log2(p) = 3.84 bits per character
  • Across the whole string3.84 × 64 = 245.8 bits
  • Compared againsthexadecimal — 3.83 bits is what a random string of this length over that alphabet would be expected to measure

What that reading suggests

As high as a random string of this length over hexadecimal could be expected to measure, which is what an encoded key, a hash, a token or ciphertext looks like.

  • Smallest alphabet that fitsHexadecimal
  • Maximum for that alphabet4 bits
  • Expected maximum at this length3.83 bits
  • Share of that expected maximum100%

Where each character goes

  • 19 · 14.1%
  • 07 · 10.9%
  • a6 · 9.4%
  • f6 · 9.4%
  • 24 · 6.3%
  • 44 · 6.3%
  • 64 · 6.3%
  • b4 · 6.3%
  • 33 · 4.7%
  • 73 · 4.7%
  • d3 · 4.7%
  • e3 · 4.7%

And 4 further symbols, all rarer than those listed.

Between 4 and 40. Every symbol is counted whatever this says — it only decides how many of them are listed.

  • Lower case24
  • Digits40

Something to measure against

  • English prose, single-letter frequencies4.14 bits
  • Random hexadecimal4 bits
  • Random base325 bits
  • Random base646 bits
  • Random printable ASCII6.57 bits
  • Encrypted or compressed bytes8 bits

The exact figures are logarithms of the alphabet size. The English one is Shannon’s 1951 measurement of single-letter frequencies over the 26 letters and a space; the same paper puts English at nearer one bit per letter once the surrounding words are taken into account, which is a thing this page structurally cannot see.

This measures the string, not the way it was chosen, and it assumes nothing about an attacker at all. abababab scores exactly 1.000 bits per character, and so does a random string of eight coin flips — the first is guessed instantly and the second takes 256 attempts. A password’s strength comes from the size of the set it was drawn from and from what the attacker knows about how you drew it, neither of which any amount of staring at the result can recover. Frequency counting sees no order, no repetition and no words: it would give the same answer to these characters in any arrangement at all. Do not read a figure here as bits of security.

Where it does earn its keep is triage. A 44-character run at 5 bits over base64 is a key or a token; the same length of log message sits well below that. High entropy in a field that should hold a name, or low entropy in a field that should hold a secret, is the signal — and a compression test is the natural next one, because a compressor finds the structure this measure is blind to.

Bits per character, and what they count

Shannon entropy is the average surprise in a sequence of symbols. Count how often each character appears, turn the counts into proportions, and sum the negative of each proportion multiplied by its own base-two logarithm; the answer is how many bits of information one character carries on average, given the mix in front of you and nothing else.

The definition comes from Claude Shannon’s 1948 paper on communication, where it measures how much a message can be compressed before something is lost. Applied to one string, it is a fast way to ask whether the characters are spread evenly across a large alphabet or concentrated in a few — which is exactly the difference between a token and a sentence.

Why a short random string scores low

A string of forty-four characters can contain at most forty-four distinct symbols, so it can never demonstrate more than the base-two logarithm of its own length. Worse, counting frequencies from a small sample underestimates the true spread: with sixty-four possible characters and only forty-four drawn, several will be missing purely by chance.

The size of that undercount is known — roughly one less than the alphabet size, divided by twice the length and by the natural logarithm of two, a correction published by Miller in 1955. It matters here because a genuinely random base64 key of that length measures about five bits per character rather than six, and calling that eighty-three per cent of maximum would be an accusation of structure that the string does not deserve. This page compares against the corrected figure instead.

Telling an encoded blob from a line of prose

This is the measure’s best use. A field that should hold a name and instead holds forty characters near the ceiling for base64 is carrying a key, a token or ciphertext. A field that should hold a secret and reads like English is holding a placeholder somebody never replaced.

The alphabet matters as much as the number. English prose and a hexadecimal digest can land within a tenth of a bit of each other — around four bits per character — and mean entirely different things, because prose is drawn from a twenty-seven symbol alphabet it barely fills while the digest saturates a sixteen symbol one. That is why the reading here is always given as a share of what its own alphabet allows.

What frequency counting cannot see

Order is invisible to it. The same characters shuffled into any arrangement give an identical result, so a repeating pattern and a random draw over the same symbols are indistinguishable — a string alternating two letters scores exactly one bit per character, and so does a genuine sequence of coin flips.

Shannon measured this limitation himself in 1951: single-letter frequencies put printed English at 4.14 bits per letter, while accounting for the surrounding words brings the real figure down to roughly one. Everything between those two numbers is structure that a per-character count is blind to. When that structure is what you are after, a compressor is the better instrument, since finding repetition is precisely what it does.

Why a short random string scores badly

Does a high reading mean a strong password?

No, and this is the most common misreading. Entropy here describes the string you pasted; password strength describes the set it was drawn from. A password chosen by a person from a predictable pattern can score well and still be guessed in seconds.

My 16-character key only scores 3.9 bits. Is something wrong with it?

Almost certainly not. At that length the measurement is capped by the sample itself, and the expected maximum shown beside the reading is the number to compare against rather than the alphabet’s theoretical ceiling.

Can I use this to find secrets committed to a repository?

It is a reasonable first filter and a poor last one. Real scanners combine an entropy threshold with pattern rules for known key formats, because entropy alone flags minified code and base64 images while missing short credentials entirely.

What is the highest value I can get?

The base-two logarithm of the number of distinct symbols present: four for hexadecimal, six for base64, eight per byte for encrypted or compressed data. Reaching the exact figure requires a long string in which every symbol appears equally often.

Is this the same entropy a password meter reports?

No. A meter multiplies length by the logarithm of the pool it assumes you chose from, which is a claim about the process that produced the password. This measures the string in front of it, which is a claim about the sample. The two agree only for a long random draw.