6 Calibration🔗ℹ

Every parameter that has been checked against a real book was wrong when first guessed, usually by an order of magnitude. The record is worth keeping in full, because it is the only part of this document that is not inference:

parameter

   

in the real books

   

first guess

   

now

the fount

   

21,953 sorts (Okes)

   

60,000

   

31,200 incl. space

tilde abbreviations

   

1.01 / 1000 words

   

83

   

1.66

superscript y-t, w-ch

   

5.5 per million words

   

6,600

   

~0

foul case + turned letters

   

0.25 / 1000 words

   

11.57

   

0.87

word division

   

5.1 / 100 lines

   

0.0

   

5.3

medial apostrophes

   

9.58 / 1000 words

   

1.17

   

5.37

ampersand

   

3.18 / 1000 words

   

35

   

3.02

class spelling habits

   

57% (Blayney)

   

82-91%

   

57%

wrong-fount sorts

   

a handful a book

   

248

   

13

gaps a compositor could set

   

every one

   

14%

   

95%

The measurements come from the 1600 quarto of Much Ado About Nothing and the 1623 Folio text set from it: 11,990 words of real copy-text and a real setting from it, which is the only kind of evidence that can settle any of this.Transcriptions from Best (1996). Hinman (1963) establishes the descent; Halliwell-Phillipps, quoted in Furness (1899), inferred that the copy was “a play-house copy of the edition of 1600, an exemplar of it, with a few manuscript directions and notes” — which is the corrector’s stage exactly. In the same appendix P. A. Daniel judges that most of the variation is the printer’s rather than the annotator’s, which is the assumption this program had been making without warrant.

A caution about the size of this evidence. Eleven thousand words is enough to catch an error of an order of magnitude and not enough to settle a question of usage. Several confident claims made from it — see the footnote to plausible? — turned out to be true of the sample and false of the language. Where a figure below rests on the two Much Ado texts alone, it should be read as a bound rather than a measurement.

The instructive failure. The program used to produce implēētatiō for implementation, stacking tilde contractions on a single word, and label it a space-saving. It was neither. The Folio has no scribal contractions in these scenes, and some substituted forms were longer than what they replaced — an expansion mislabelled as a contraction, because the TEI marked both with <abbr>. Expansions are now <orig>/<reg>, only one alteration may be applied to a word, and the scribal signs are off by default. The genuine English space-saver was in the same data unnoticed: see The elided ending.

Forward test, Q1600 → F1623, on the spelling that attribution work depends on:

   

actual F1

   

simulated

here / heere

   

52%

   

51%

do / doe

   

61%

   

80%

The do/doe overshoot has a known cause rather than an excuse: the real scenes were set by more than one man, and the simulation ran a single man’s habit across all of them.