A computational typology of dramatic voice, built from the complete speech of every speaking character in the early modern English repertoire — 6,466 characters from 494 plays, grouped into 25 register clusters and traceable across a century of theatre (1520–1641).
Every character as a point in archetype space. Hover for play, author, date and role; filter by decade, genre, author, company or theater.
Open the map →One page per cluster: when the type emerged, its prototype candidates, historical profile, signatures, and every member with speech excerpts.
Browse the 25 clusters →How every number and badge is computed — typicality, emergence, prototypes, year bases, signatures — and what the pages do and do not claim.
Read the methods →Each cluster page names an archetype and lays out its evidence: a historical profile against the corpus baseline, author and genre signatures, landmark members, and a full chronological roster under performance-first dating. An emergence line dates when the type became an established resource ("established by" = the year the first tenth of its dated members had appeared), and prototype badges mark the formative-era members that are also strongly typical of the voice — the best candidates for its founding exemplars. Typicality (cosine similarity to the cluster centre) shows how central any member is; low-typicality members are blends of several registers, not misfits. Every row links to the character's full speech, so each archetypal reading can be verified against the raw text.
Each point is a single character; color encodes its archetype cluster. Hover for play, author, date, genre, role, and the cluster's distinguishing words. Click a legend entry to hide that cluster, double-click to isolate it. The dropdown menus highlight a single decade, genre, author, play, play type, theater, or company — everything else dims. Characters speaking fewer than 150 words are not shown (too little text for a stable stylistic signature), and only one edition of each play is included. The 2-D layout is a UMAP projection for display only — cluster membership comes from the full embedding space.
Characters were extracted from EEBO-TCP drama transcriptions; play metadata is
merged from Wiggins & Richardson's British Drama: A Catalogue of Plays and
from DEEP.
Speech is typographically normalized, spelling-modernized (MorphAdorner), and proper
nouns are masked so clusters reflect register rather than shared names. Each
character's speech is embedded with gte-Qwen2-1.5B-instruct in
~1024-token windows, mean-pooled per character; play-level
signal is removed by leave-one-out play centering; the typology is spherical k-means
(k=25). The space is a continuum with dense regions — read clusters as density peaks,
not boxes; the methods page spells out every rule, and every
data-affecting decision is recorded in the project's
provenance log.