Numbers
Frequency and distribution insights from the corpus. Each card ends with its verification status.
How the Qur'an is counted ●
The same text can be counted at different grains, and the numbers on this page use several. The key one is the word-unit — the corpus term is token — the smallest piece the corpus counts. Arabic often writes several word-units as a single word: bismi ("in the name of") is one written word but two word-units (bi + ism). That is why the Qur'an has far more word-units than written words. From the largest grain to the smallest:
● Leeds corpus v0.4.
ℹ How this is computed
● every figure recomputed from the bundled Leeds Quranic Arabic Corpus v0.4 morphology. "Word-unit" is this site's plain-language name for a corpus token.
Every figure here is recomputed from the corpus data bundled with the site. Where the numbers come from.
Jump to: Scale · Vocabulary · Structure · Across revelation
Scale
How long surahs and verses run.
Surah length
Mean: 54.7 verses. Median: 39. Longest: al-Baqarah (286). Shortest: al-Kawthar, al-'Asr, al-Nasr (3 each).
● from Tanzil (114 surahs, 6,236 verses by Cairo numbering).
Verse-length distribution
Most verses are short. Of 6,236 verses, the longest runs to 128 word-tokens (al-Baqarah 2:282).
● token-length of every verse (Leeds tokenization), grouped into bands.
Vocabulary
What the text is made of: the most frequent roots, everyday concepts, and the rare words.
Top 5 most frequent roots
| Root | Count | Per 1,000 tokens |
|---|---|---|
| a-l-h (God) | 2,851 | 36.82 |
| q-w-l (say) | 1,722 | 22.24 |
| k-w-n (be) | 1,390 | 17.95 |
| r-b-b (Lord) | 980 | 12.66 |
| a-m-n (believe) | 879 | 11.35 |
Verified Leeds corpus v0.4.
ℹ How this is computed
Verified all five from the Quranic Arabic Corpus at corpus.quran.com (Leeds, Dukes 2009–2017). Full top-10 table with verification status on Roots. "Per 1,000 tokens" is (count / 77,429) × 1,000, direct computation.
Concept frequency
| Concept | Count | Per 1,000 tokens |
|---|---|---|
| Justice (ʿ-d-l) | 28 | 0.36 |
| Patience (ṣ-b-r) | 103 | 1.33 |
| Knowledge (ʿ-l-m) | 854 | 11.03 |
| Gratitude (sh-k-r) | 75 | 0.97 |
| Taqwa (w-q-y) | 258 | 3.33 |
● root counts from the Leeds Quranic Arabic Corpus (Dukes 2009–2017).
Earth and sea
Root a-r-ḍ (earth, land): 461. Root b-ḥ-r (sea): 42.
Verified Leeds corpus v0.4.
ℹ How this is computed
Verified root counts from the Leeds Quranic Arabic Corpus (Dukes 2009–2017). A widely cited surface-form count (13 "land" / 32 "sea") uses specific noun senses only; see counting note above. Nuanced on the popular interpretation.
Day, month, year
Day (y-w-m, root): 405 occurrences (Leeds corpus). Widely circulated totals of 365 or 475 select particular surface or compound forms; this site has not reproduced them. Month (sh-h-r, root): 21. Claims of 12 for singular shahr likewise depend on a surface-form selection not reproduced here. Year: not a single canonical word in the Qur'an.
● Leeds corpus v0.4.
ℹ How this is computed
● root counts from Leeds. ~ only the root totals are verified here. The popular surface-form totals are shown as a methodological caution, not endorsed as findings. See Roots for the recomputable root detail.
Words that occur only once (hapax legomena)
12,004 surface forms (out of 18,993 unique forms) occur exactly once, alongside 1,994 once-occurring lemmas. At the root level, 395 roots (out of 1,642) occur exactly once; each once-root's single verse is listed below.
List the once-occurring roots
● Leeds corpus v0.4.
ℹ How this is computed
● from Leeds corpus word-form and root counts; once-occurring roots listed with the verse of their single occurrence.
Vocabulary variety by period (TTR, MATTR, MTLD)
Type-token ratio (distinct surface forms, or distinct lemmas, divided by all word-tokens) measures vocabulary variety versus repetition — but it falls mechanically as sample size grows, and these four periods differ greatly in token count (2,704 to 30,572). Two length-robust alternatives are shown alongside it: MATTR (a moving-average TTR over a fixed window, Covington & McFall 2010) and MTLD (a factor-count measure, McCarthy & Jarvis 2010). They tell a different story than raw form TTR does: MATTR is nearly flat across all four periods, and MTLD shows a real but much smaller decline. Raw TTR's steep decline is mostly the sample-size artifact, not a vocabulary difference this large.
| Period | Tokens | Form TTR | Lemma TTR | Form MATTR | Form MTLD |
|---|---|---|---|---|---|
| Early Meccan | 2,704 | 0.6202 | 0.3528 | 0.8707 | 673.77 |
| Middle Meccan | 14,163 | 0.4081 | 0.1466 | 0.8913 | 647.64 |
| Late Meccan | 30,572 | 0.3073 | 0.0926 | 0.8868 | 609.05 |
| Medinan | 29,990 | 0.3005 | 0.0902 | 0.86 | 395.98 |
● Leeds corpus v0.4.
ℹ How this is computed
●
distinct surface forms / lemmas ÷ word-tokens, by Cairo 1924
period. Form MATTR: mean TTR over a 100-token sliding window
(Covington & McFall 2010). Form MTLD: bidirectionally
averaged factor count at a 0.72 TTR threshold (McCarthy &
Jarvis 2010). Both computed by
scripts/lib/lexical-diversity.mjs.
~
type-token ratio is sensitive to sample size; these periods
range from 2,704 to 30,572 tokens, so the falling ratio partly
reflects length, not style alone.
Structure
Grammatical texture and structural features of the text.
Grammatical texture
Of 77,429 word-tokens, 25,135 are nouns and 19,356 are verbs. The noun and verb share shifts across the revelation:
| Period | Nouns | Verbs |
|---|---|---|
| Early Meccan | 35.5% | 24.9% |
| Middle Meccan | 32.6% | 26.4% |
| Late Meccan | 33% | 24.9% |
| Medinan | 31.6% | 24.5% |
● Leeds corpus v0.4.
ℹ How this is computed
● exact counts of the Leeds N (noun) and V (verb) tags as a share of all tokens, by Cairo 1924 period. Noun and verb are the two unambiguous single tags; the remaining tokens are particles and other function words, so no family taxonomy is asserted.
Root density across surahs ●
Normalized frequency (per 1,000 tokens) of the 40 most frequent roots across all 114 surahs. Darker cells mean the root is more locally concentrated in that surah, not more important.
Revelation order follows Cairo 1924 revelation order (Nöldeke-Bell four-period classification), one chronology among several scholarly proposals.
Fawatih (opening particles)
29 surahs open with disconnected letters (al-muqatta'at). 14 unique letter combinations across these 29 surahs.
List all 29
● Tanzil.
ℹ How this is computed
● from Tanzil text. ~ on interpretation; meaning is discussed across the tafsir tradition. ● The 29/14 count is computed directly from the bundled morphology, not asserted — see it listed above.
Across revelation
How these measures shift across the four traditional revelation periods.
Verse length by band
| Band | Avg. word-units/verse |
|---|---|
| Early Meccan | 4.5 |
| Middle Meccan | 10.1 |
| Late Meccan | 11.7 |
| Medinan | 18.5 |
● Leeds corpus v0.4.
ℹ How this is computed
● from Leeds word counts + Cairo 1924 chronological classification. The increasing-length pattern across periods is widely observed in academic literature.
Vocabulary growth across revelation
Following Cairo 1924 order, the first revealed surah already introduces 32 distinct roots; by the last, all 1,642 rooted roots have appeared. New roots enter fastest early and taper as the shared vocabulary saturates.
● Leeds corpus v0.4.
ℹ How this is computed
● count of roots first appearing at each revelation-order step. Order is a convention; the curve describes the corpus, not a claim about composition.
Roots unique to one period
547 roots appear only in Meccan surahs (zero Medinan occurrences). 198 roots appear only in Medinan surahs.
● Leeds corpus v0.4.
ℹ How this is computed
● from Leeds + Cairo 1924 four-period classification. ~: a handful of surahs have mixed-period classification in some traditions; the boundary roots may shift by a few in other schemes.
Distinctive vocabulary by period (keyness, G2) ~
Keyness (Dunning's G2) measures how distinctively a root's token frequency in one revelation period differs from the rest of the corpus. The top 15 roots by G2 are listed for each of the four periods, by Cairo 1924 revelation order (Nöldeke-Bell four-period classification). Periodization varies across scholarly chronologies; treat these rankings as Nuanced, not as a settled chronology.
● Leeds corpus v0.4.
ℹ How this is computed
●
G2 computed from root token counts per period
(data/roots-summary.json) against period token
totals (data/numbers.json), by
scripts/compute-association-stats.mjs.
~
on the period assignment itself (Cairo 1924 / Nöldeke-Bell).
Full root-frequency and keyness tables (CSV/JSON):
Export.
How evenly roots are spread across the Qur'an ●
Frequency alone cannot tell whether a root recurs across the whole text or clumps inside a few surahs telling one story. Gries's Deviation of Proportions (normalized, DPnorm) measures exactly that, for every root with at least 20 corpus-wide occurrences: 0 means the root's occurrences track surah size as closely as possible; 1 means total concentration in the smallest surah.
Per-root detail (DP, DPnorm, Juilland's D, range, adjusted frequency): open any root on the Roots page.
ℹ How this is computed
●
Computed by scripts/compute-dispersion.mjs over
the 114 surahs, weighted by token count (not treated as
equal-sized). DP = 0.5·Σ|v_i − s_i|; DPnorm = DP / (1 −
min(s_i)), the corrected normalization from Lijffijt &
Gries (2012). Roots below 20 corpus-wide occurrences are
excluded from this ranking (though every root's own DP/DPnorm
still appears on its Roots page detail) since a rare root is
trivially maximally concentrated without that being
informative. Full per-root table (CSV/JSON):
Export.