Function Word Frequencies

Imagine you are trying to identify a mystery sender by looking at their grocery list. You might notice they buy the same basic staples like milk, bread, and eggs regardless of what fancy meal they plan to cook. These common items act like a fingerprint because the sender cannot easily change their habits. In literature, we find this same pattern when we look at function words, which are the small, structural building blocks of any language. These words include common items like the, and, of, and with, which connect our thoughts together. While authors change their complex vocabulary to fit different themes, they rarely change how they use these small, frequent words. This makes them a perfect tool for identifying the unique style of a specific writer.
The Logic of Word Frequency
When we analyze a text, we look for the hidden patterns that exist beneath the surface of the story. Most writers do not consciously choose how often they use the word 'the' in a single paragraph. This unconscious behavior creates a reliable signature that is incredibly difficult for an author to fake or hide. Think of it like a person's walking gait, which stays consistent even when they try to change their shoes or clothing. By counting the frequency of these small words, we can build a numerical profile of an author. This profile helps us compare anonymous texts against known works to see if they share the same writer.
Key term: Function words — the small, grammatical words that provide structure to sentences rather than carrying specific meaning.
To understand this, consider how we might sort a large collection of random objects into specific groups. If you have a pile of mixed hardware, you might group the items by their shape or their size. In computational stylometry, we treat words like pieces of hardware that we sort into frequency bins. We count every instance of a word and divide that by the total number of words in the text. This gives us a percentage that represents the author's specific preference for that word. When we compare these percentages across different books, we can see if the author maintains a consistent style over time.
Calculating the Stylometric Signature
We must organize these frequencies carefully to ensure our math provides a clear result for our analysis. We typically build a table to track how often each specific function word appears in a sample. This allows us to visualize the data and spot any major differences between two suspected authors. The table below shows a sample frequency distribution for three common words found in two different short stories.
| Word | Author A Frequency | Author B Frequency | Difference |
|---|---|---|---|
| the | 0.072 | 0.061 | 0.011 |
| and | 0.035 | 0.042 | 0.007 |
| of | 0.028 | 0.029 | 0.001 |
This simple distribution helps us identify which author is more likely to have written an anonymous text. If the anonymous text shows frequencies closer to Author A, we have strong evidence for our claim. This process is much faster than reading thousands of pages to look for themes or character traits. By focusing on these tiny, frequent words, we turn the art of writing into a precise science of numbers.
Because these words are so common, they provide a large amount of data for our mathematical models. Even a short letter contains hundreds of these function words, which is enough to run a reliable test. We do not need a whole novel to find a signature, as small samples often provide enough information. This efficiency makes our work useful for historians and detectives who only have fragments of text to study. As we gather more data, our ability to confirm the identity of an author becomes more accurate and reliable over time.
Calculating the usage rates of common function words provides a stable numerical signature that reveals an author's identity regardless of their subject matter.
The next Station introduces vocabulary richness metrics, which determines how an author uses diverse words to build their unique narrative style.