Stylometric Signature
Definition
The set of measurable linguistic features : function-word frequencies, sentence length, punctuation patterns, vocabulary richness : that characterises a specific writer's output and allows their texts to be distinguished from others.
- Field
- Forensic linguistics, authorship attribution
- Core features
- Function-word frequency, sentence length, punctuation, vocabulary richness
- Use
- Distinguishing one writer's output from another's
- Emerging challenge
- AI-generated text lacks a stable individual signature
Common questions
Why are function words more useful than content words for this?+
Content words track topic, so they vary with subject matter. Function words such as prepositions and conjunctions are used unconsciously and stay fairly stable across a writer's topics, which makes their frequencies a better fingerprint of habit than a fingerprint of subject.
Can a stylometric signature be deliberately disguised?+
Yes. A writer aware of stylometric analysis can vary sentence length, swap synonyms, or imitate another author's habits, which lowers reliability. Analysts look for internal inconsistency across a disputed document as one sign of deliberate disguise or multiple authorship.
Related terms
- AI Detection Classifier
- A machine-learning system trained to discriminate between human-written and LLM-generated text by measuring features such as perplexity, burstiness, and n-gram probabilities. Current...
- Burstiness
- The variance in sentence-level complexity across a passage. Human writing tends to mix complex and simple sentences unevenly; LLM output tends to...
- Human-LLM Collaboration
- Text production in which a human contributes prompts, editing decisions, and intentional content while an LLM generates or transforms the prose. The...
- Large Language Model (LLM)
- A neural network trained on large text corpora to predict likely next tokens. At inference, it generates text by sampling from probability...
- Perplexity
- A measure of how surprising a sequence of words is to a language model. LLMs tend to generate low-perplexity text (predictable word...