Divine Discourses

Toward direct engagement with the Qur'an

Numbers

Frequency and distribution insights from the corpus. Each card ends with its verification status.

Depth:
Keys 1 2 3

How the Qur'an is counted

The same text can be counted at different grains, and the numbers on this page use several. The key one is the word-unit — the corpus term is token — the smallest piece the corpus counts. Arabic often writes several word-units as a single word: bismi ("in the name of") is one written word but two word-units (bi + ism). That is why the Qur'an has far more word-units than written words. From the largest grain to the smallest:

6,236
verses (ayat)
the numbered verses, by Cairo counting
77,429
word-units (tokens)
every morphological piece of running text — the raw building blocks
18,993
distinct spellings (surface forms)
how many different written forms appear at least once
4,832
dictionary words (lemmas)
distinct headwords you would look up in a dictionary
1,642
roots
the triliteral roots the whole vocabulary grows from

Leeds corpus v0.4.

ℹ How this is computed

every figure recomputed from the bundled Leeds Quranic Arabic Corpus v0.4 morphology. "Word-unit" is this site's plain-language name for a corpus token.

Every figure here is recomputed from the corpus data bundled with the site. Where the numbers come from.

Jump to: Scale · Vocabulary · Structure · Across revelation

Scale

How long surahs and verses run.

Surah length

Mean: 54.7 verses. Median: 39. Longest: al-Baqarah (286). Shortest: al-Kawthar, al-'Asr, al-Nasr (3 each).

from Tanzil (114 surahs, 6,236 verses by Cairo numbering).

Verse-length distribution

Most verses are short. Of 6,236 verses, the longest runs to 128 word-tokens (al-Baqarah 2:282).

token-length of every verse (Leeds tokenization), grouped into bands.

Vocabulary

What the text is made of: the most frequent roots, everyday concepts, and the rare words.

Top 5 most frequent roots

Root Count Per 1,000 tokens
a-l-h (God) 2,851 36.82
q-w-l (say) 1,722 22.24
k-w-n (be) 1,390 17.95
r-b-b (Lord) 980 12.66
a-m-n (believe) 879 11.35

Verified Leeds corpus v0.4.

ℹ How this is computed

Verified all five from the Quranic Arabic Corpus at corpus.quran.com (Leeds, Dukes 2009–2017). Full top-10 table with verification status on Roots. "Per 1,000 tokens" is (count / 77,429) × 1,000, direct computation.

Concept frequency

Concept Count Per 1,000 tokens
Justice (ʿ-d-l) 28 0.36
Patience (ṣ-b-r) 103 1.33
Knowledge (ʿ-l-m) 854 11.03
Gratitude (sh-k-r) 75 0.97
Taqwa (w-q-y) 258 3.33

root counts from the Leeds Quranic Arabic Corpus (Dukes 2009–2017).

Earth and sea

Root a-r-ḍ (earth, land): 461. Root b-ḥ-r (sea): 42.

a-r-ḍ (land)
b-ḥ-r (sea)

Verified Leeds corpus v0.4.

ℹ How this is computed

Verified root counts from the Leeds Quranic Arabic Corpus (Dukes 2009–2017). A widely cited surface-form count (13 "land" / 32 "sea") uses specific noun senses only; see counting note above. Nuanced on the popular interpretation.

Day, month, year

Day (y-w-m, root): 405 occurrences (Leeds corpus). Widely circulated totals of 365 or 475 select particular surface or compound forms; this site has not reproduced them. Month (sh-h-r, root): 21. Claims of 12 for singular shahr likewise depend on a surface-form selection not reproduced here. Year: not a single canonical word in the Qur'an.

Leeds corpus v0.4.

ℹ How this is computed

root counts from Leeds. ~ only the root totals are verified here. The popular surface-form totals are shown as a methodological caution, not endorsed as findings. See Roots for the recomputable root detail.

Words that occur only once (hapax legomena)

12,004 surface forms (out of 18,993 unique forms) occur exactly once, alongside 1,994 once-occurring lemmas. At the root level, 395 roots (out of 1,642) occur exactly once; each once-root's single verse is listed below.

List the once-occurring roots

Leeds corpus v0.4.

ℹ How this is computed

from Leeds corpus word-form and root counts; once-occurring roots listed with the verse of their single occurrence.

Vocabulary variety by period (TTR, MATTR, MTLD)

Type-token ratio (distinct surface forms, or distinct lemmas, divided by all word-tokens) measures vocabulary variety versus repetition — but it falls mechanically as sample size grows, and these four periods differ greatly in token count (2,704 to 30,572). Two length-robust alternatives are shown alongside it: MATTR (a moving-average TTR over a fixed window, Covington & McFall 2010) and MTLD (a factor-count measure, McCarthy & Jarvis 2010). They tell a different story than raw form TTR does: MATTR is nearly flat across all four periods, and MTLD shows a real but much smaller decline. Raw TTR's steep decline is mostly the sample-size artifact, not a vocabulary difference this large.

Period Tokens Form TTR Lemma TTR Form MATTR Form MTLD
Early Meccan 2,704 0.6202 0.3528 0.8707 673.77
Middle Meccan 14,163 0.4081 0.1466 0.8913 647.64
Late Meccan 30,572 0.3073 0.0926 0.8868 609.05
Medinan 29,990 0.3005 0.0902 0.86 395.98

Leeds corpus v0.4.

ℹ How this is computed

distinct surface forms / lemmas ÷ word-tokens, by Cairo 1924 period. Form MATTR: mean TTR over a 100-token sliding window (Covington & McFall 2010). Form MTLD: bidirectionally averaged factor count at a 0.72 TTR threshold (McCarthy & Jarvis 2010). Both computed by scripts/lib/lexical-diversity.mjs. ~ type-token ratio is sensitive to sample size; these periods range from 2,704 to 30,572 tokens, so the falling ratio partly reflects length, not style alone.

Structure

Grammatical texture and structural features of the text.

Grammatical texture

Of 77,429 word-tokens, 25,135 are nouns and 19,356 are verbs. The noun and verb share shifts across the revelation:

Period Nouns Verbs
Early Meccan 35.5% 24.9%
Middle Meccan 32.6% 26.4%
Late Meccan 33% 24.9%
Medinan 31.6% 24.5%

Leeds corpus v0.4.

ℹ How this is computed

exact counts of the Leeds N (noun) and V (verb) tags as a share of all tokens, by Cairo 1924 period. Noun and verb are the two unambiguous single tags; the remaining tokens are particles and other function words, so no family taxonomy is asserted.

Root density across surahs

Normalized frequency (per 1,000 tokens) of the 40 most frequent roots across all 114 surahs. Darker cells mean the root is more locally concentrated in that surah, not more important.

Surah order:

Fawatih (opening particles)

29 surahs open with disconnected letters (al-muqatta'at). 14 unique letter combinations across these 29 surahs.

List all 29

Tanzil.

ℹ How this is computed

from Tanzil text. ~ on interpretation; meaning is discussed across the tafsir tradition. The 29/14 count is computed directly from the bundled morphology, not asserted — see it listed above.

Across revelation

How these measures shift across the four traditional revelation periods.

Verse length by band

Band Avg. word-units/verse
Early Meccan 4.5
Middle Meccan 10.1
Late Meccan 11.7
Medinan 18.5

Leeds corpus v0.4.

ℹ How this is computed

from Leeds word counts + Cairo 1924 chronological classification. The increasing-length pattern across periods is widely observed in academic literature.

Vocabulary growth across revelation

Following Cairo 1924 order, the first revealed surah already introduces 32 distinct roots; by the last, all 1,642 rooted roots have appeared. New roots enter fastest early and taper as the shared vocabulary saturates.

Leeds corpus v0.4.

ℹ How this is computed

count of roots first appearing at each revelation-order step. Order is a convention; the curve describes the corpus, not a claim about composition.

Roots unique to one period

547 roots appear only in Meccan surahs (zero Medinan occurrences). 198 roots appear only in Medinan surahs.

Leeds corpus v0.4.

ℹ How this is computed

from Leeds + Cairo 1924 four-period classification. ~: a handful of surahs have mixed-period classification in some traditions; the boundary roots may shift by a few in other schemes.

Distinctive vocabulary by period (keyness, G2) ~

Keyness (Dunning's G2) measures how distinctively a root's token frequency in one revelation period differs from the rest of the corpus. The top 15 roots by G2 are listed for each of the four periods, by Cairo 1924 revelation order (Nöldeke-Bell four-period classification). Periodization varies across scholarly chronologies; treat these rankings as Nuanced, not as a settled chronology.

Leeds corpus v0.4.

ℹ How this is computed

G2 computed from root token counts per period (data/roots-summary.json) against period token totals (data/numbers.json), by scripts/compute-association-stats.mjs. ~ on the period assignment itself (Cairo 1924 / Nöldeke-Bell). Full root-frequency and keyness tables (CSV/JSON): Export.

How evenly roots are spread across the Qur'an

Frequency alone cannot tell whether a root recurs across the whole text or clumps inside a few surahs telling one story. Gries's Deviation of Proportions (normalized, DPnorm) measures exactly that, for every root with at least 20 corpus-wide occurrences: 0 means the root's occurrences track surah size as closely as possible; 1 means total concentration in the smallest surah.

Per-root detail (DP, DPnorm, Juilland's D, range, adjusted frequency): open any root on the Roots page.

ℹ How this is computed

Computed by scripts/compute-dispersion.mjs over the 114 surahs, weighted by token count (not treated as equal-sized). DP = 0.5·Σ|v_i − s_i|; DPnorm = DP / (1 − min(s_i)), the corrected normalization from Lijffijt & Gries (2012). Roots below 20 corpus-wide occurrences are excluded from this ranking (though every root's own DP/DPnorm still appears on its Roots page detail) since a rare root is trivially maximally concentrated without that being informative. Full per-root table (CSV/JSON): Export.