Hangman Words

Why A beats E as a first guess on short words

By Aurangzeb Khan··4 min read

Everybody opens with E. It is the most common letter in English, it is the first thing anyone is told about the language, and for long words it is the right guess. For short ones it is not, and the reason is a confusion between two different things that both get called 'letter frequency'.

Two different measurements

The famous figure — E at about 12.7% — counts letters in running text. It is a statement about prose: if you take a page of English and count characters, roughly an eighth of them are E.

Hangman is not played against a page. It is played against one word, and the question is not how often E occurs but how many distinct words contain at least one E. Those are different questions and they have different answers, because E is concentrated: it turns up repeatedly in the words that have it, and in a great many common endings.

Counted the second way, across every word in the dictionary this site uses, the opening guess that touches the most 6,798 four-letter words is A (36.5%) — not E.

What the numbers actually say

Three letters: the best opening is A (26.6%). Four: A (36.5%). Five: A (44.1%). Six: E (57.5%). Seven: E (63.1%). Eight: E (66.9%).

The crossover sits around six letters. Below it A leads; at six and above E takes over and keeps the lead for every length after. That is one rule a player can actually hold: short word, open with A; long word, open with E.

Why the crossover happens where it does

English builds long words out of pieces, and the pieces are full of E — the -ed and -er and -ment and -ness endings, the re- and pre- beginnings. A word has to be long enough to carry one of those before E's advantage appears.

Short words have no room for morphology. They are mostly single syllables built around a single vowel, and across the whole set of them A is the vowel that turns up in the most.

What this is worth in a real game

Not as much as it sounds, and more than nothing. On a four-letter word the difference between the best and second-best opening is a few percentage points of the word set — one guess in twenty or thirty where you get a letter instead of a limb.

It compounds, though. A guesser who is right about the opening sees the board move, and a board that has moved makes the second guess easier. The advantage is in what the first letter tells you, not in the first letter itself.

The second guess matters more

Once the first letter has landed or missed, the useful question changes from 'which letter is commonest' to 'which letter best splits what is left'. That is a harder calculation and nobody does it at the table, but the rough version is easy: after a vowel lands, guess the consonants that pair with it; after a vowel misses, guess another vowel before you touch a consonant.

A word that has survived two vowel guesses is a word with very few vowels, and that is worth knowing early — it is the shape of a word that will end the game.

Why this is not in the strategy guides

Because the figure everybody quotes comes from cryptanalysis, where counting characters in intercepted text is exactly the right thing to do. Hangman inherited the number without inheriting the question it answers.

It is a good example of a statistic that is correct, widely known, and applied to the wrong problem. Anybody can check this one: take a word list, count how many distinct words contain each letter, and compare it with the frequency table you were taught.

The second-order effect nobody mentions

A first guess that lands does more than fill a blank: it tells you where the vowels are not. Learning that positions two and five are a vowel narrows the consonant possibilities around them sharply, because English does not permit most consonant clusters.

This is why a correct opening is worth more than its frequency advantage suggests, and why a guesser who opens well tends to finish several guesses ahead rather than one.

It also explains the opposite case. An opening that misses entirely is not a wasted guess — a word with no A in it is already an unusual word, and three missed vowels narrows the field more than most single correct guesses would.

Where this breaks

Named categories. The moment somebody says 'it's an animal', dictionary frequency stops being the right model, because the set of candidate words has collapsed to a few hundred with their own distribution. Play the category, not the alphabet.

And proper nouns, which obey no distribution at all. This is one of several reasons the lists here exclude them.

On this site

All guides