Early Modern Drama · Character Archetypes

Methods & labels

How every number and badge on these pages is computed. Partition: 2026-07-07 NOS-verified run (k=25, seed 42, 6,466 characters from 494 plays). Added 2026-07-10; every methodological decision is recorded in the repository's PROVENANCE_LOG.

Corpus and characters

The corpus is the drama subset of EEBO-TCP: 9,638 speaking characters extracted from 557 playbook transcriptions, each character represented by everything they speak. Speech prefixes were mapped to characters play-by-play (with an LLM-assisted, fully recorded speaker-mapping pass). Two kinds of row are excluded from clustering, not from the data: characters speaking fewer than 150 words (no stable stylistic signature), and duplicate editions of the same play — one edition is kept per work, confirmed by cast overlap, with New Oxford Shakespeare canonical choices for Shakespeare adopted under an explicit provenance note. That leaves the 6,466 clustered characters shown here.

From speeches to a character space

Character and place names in the speeches are masked (someone, that place, …) and spelling is modernized, so clusters reflect register rather than named content. Each character's masked speech is embedded with gte-Qwen2-1.5B-instruct in 1024-token chunks, token-weighted mean-pooled into one vector. Two corrections then make the space about characters rather than plays: leave-one-out play centering (each character minus the mean of its play-mates — "how does this voice differ from its own play?") and linear removal of the document-length direction (length R² 0.78 → 0.35). Spherical k-means with k=25 (seed 42) partitions the result.

Read clusters as regions, not boxes

The space is a continuum with dense regions: silhouette 0.021, seed-to-seed ARI 0.29 (0.26–0.35 over 5 refits; report 2026-07-08). Clusters are density peaks in that continuum — membership is graded, boundary members are blends, and cluster ids reshuffle on every re-run. Nothing on these pages should be read as a hard Theophrastan pigeonhole.

The labels, one by one

Typicality. Cosine similarity between a character and its cluster's centroid in the processed space. High = close to the heart of the register; low = a blend of several regions. It measures representativeness, never quality or importance — canonical protagonists are often low-typicality blends.

Three year bases. (1) Catalogued year — the year in the DEEP/TCP catalogue record, largely the print year: used by the historical-profile strip and its badges, matching the published report (6,040 dated members, median 1611). (2) Performance-first dating — the first-performance year where the catalogue records one (5,911 of 6,040 dated members), else the catalogued year: used by the roster, its Year column and sort, the emergence line, prototype badges, index year-ranges and sparklines. Performance years run a median of ~3 years earlier than print years, sometimes decades. (3) Display fills — 426 characters from metadata-bare collection items carry an approved parent-volume print year as a last-resort display date (marked in the provenance log); these never enter the emergence computation. Character pages show the catalogued year and, where the catalogue records one, the first-performance year.

First recorded / established by (emergence). Computed per cluster on performance-first dating, for clusters with ≥20 dated members. First recorded is the earliest dated member. Established by is the year the first tenth of the cluster's dated members had appeared — the point where the type stops being sporadic and becomes an available resource. (Fixed-count burst rules were tested and rejected: they collapse to the corpus onset for nearly every cluster, tracking survival of playbooks rather than the type's own take-off.)

Prototype. A member of the formative era (dated up to the establishment year) whose typicality is at or above the cluster median; the top 5 by typicality are badged. This replaces an earlier "early" badge that marked the five oldest dated members regardless of typicality (on average those sat at the 35th typicality percentile of their cluster — old, but mostly blends, and weak evidence for a prototype claim). Prototype badges assert kinship and priority within the corpus, not proven indebtedness: a directed temporal nearest-neighbour genealogy is planned as the real lineage instrument.

early-rooted / late register. Historical-profile badges on the catalogued-year basis, for clusters with ≥20 dated members: early-rooted = pre-1590 share ≥ 1.4× the corpus share; late register = post-1625 share ≥ 1.15× corpus and median year above the corpus median (1611).

Signatures (×lift). For authors, genres, companies, theaters and play types: the share of the cluster's members carrying the facet ÷ the same share over all 6,466 clustered characters. Shown only for facets carried by ≥5 members (authors ≥4) with lift ≥ 1.2; top 5. Multi-valued fields are split into atomic facets first.

Distinguishing vocabulary. c-TF-IDF over the masked, modernized speech text, with stoplists and a character-name blocklist. Name-like tokens that survive (e.g. mephastophilis) are residue of incomplete masking, not register — a second masking pass is a recorded open item. These words drive the excerpt highlighting.

Cluster names and families. Names are curation, not computation: all 25 are assigned by the author (Heejin Kim) against the cluster profiles, and the drafting provenance of each name is recorded in the names file and the provenance log (Entries 004 and 009) rather than displayed on the pages. c-TF-IDF fallback labels appear where no name is curated yet (e.g. after a re-clustering). Family groupings on the index are curated the same way.

What these pages do not claim

Survival is not production: dating and counts describe the extant, dated record. Priority is not influence: "prototype" marks the earliest typical exemplars in the corpus, not demonstrated imitation. Membership is not essence: a low-typicality member is evidence of blending, not misclassification. The partition is one defensible view of a continuous space, published with its instability measured and stated.

Provenance

Every data-affecting change — canonical-edition choices, approved metadata fills, year-basis decisions, label-rule changes — is logged with dates, approvals and commit hashes in PROVENANCE_LOG.md. The full pipeline (stages 01–07) and all derived tables are in the repository; re-running stages 04–07 regenerates this site.