Tweet density
Comparison of the information density of tweets in different languages
Under a uniform 140-character limit, how much actual information can different languages convey?
To measure this, I crawled tweets across various languages and compared their character lengths against their automated English translations to quantify relative "linguistic density."
Information Conveyed per 140 Characters
Using a sample of 100–1,200 tweets per language, each language's character-to-information ratio was computed relative to English (baseline 1.00).
- CJK Dominance: Chinese exceeded English density by over 3×, while Japanese and Korean scored more than 2×. Under the same character ceiling, CJK languages convey twice to three times as much content.
- Alphabetical Variations: While Thai and Serbian trended slightly denser than English, most European languages (Spanish, French, German, etc.) produced longer translations, reflecting lower relative character efficiency.
A quick data experiment quantifying structural language efficiency through the lens of character limits on social media.