{"id":559,"date":"2014-04-26T01:34:07","date_gmt":"2014-04-25T22:34:07","guid":{"rendered":"http:\/\/blogs.helsinki.fi\/kaislani\/?p=559"},"modified":"2024-09-05T18:43:44","modified_gmt":"2024-09-05T18:43:44","slug":"spelling-variation","status":"publish","type":"post","link":"https:\/\/samuli.kaislaniemi.fi\/blog\/2014\/04\/26\/spelling-variation\/","title":{"rendered":"Did English spelling variation end in the 1630s?"},"content":{"rendered":"<h2>1. Early Modern English spelling variation<\/h2>\n<p>Yesterday, rather late in the evening, I followed a link on Twitter:<\/p>\n<blockquote class=\"twitter-tweet\" data-width=\"550\" data-dnt=\"true\">\n<p lang=\"en\" dir=\"ltr\">So there&#39;s an EEBO-TCP spelling variation google ngram browser <a href=\"http:\/\/t.co\/OLxUv5NLBQ\">http:\/\/t.co\/OLxUv5NLBQ<\/a> (via <a href=\"https:\/\/x.com\/dr_heil?ref_src=twsrc%5Etfw\">@dr_heil<\/a>)<\/p>\n<p>&mdash; heather froehlich (@heatherfro) <a href=\"https:\/\/x.com\/heatherfro\/status\/459361695677558784?ref_src=twsrc%5Etfw\">April 24, 2014<\/a><\/p><\/blockquote>\n<p><script async src=\"https:\/\/platform.x.com\/widgets.js\" charset=\"utf-8\"><\/script><\/p>\n<p>This led to the great <a href=\"http:\/\/earlyprint.wustl.edu\/\">Early Modern Print : Text Mining Early Printed English<\/a> website where there was an interface like the <a href=\"https:\/\/books.google.com\/ngrams\">Google Books Ngram Viewer<\/a> but for the <a href=\"http:\/\/www.textcreationpartnership.org\/tcp-eebo\/\">EEBO-TCP<\/a> corpus, called\u00a0<em>EEBO Spelling Browser<\/em> (or more technically, <em>EEBO-TCP Ngram Browser<\/em>).\u00a0With the delight of a researcher falling upon a new toy I started to play with it \u2013 but hadn&#8217;t even started when I was struck by the figure that is displayed when you navigate to the <a href=\"http:\/\/earlyprint.wustl.edu\/tooleebospellingbrowser.html?queryTerms=above,aboue&amp;spellings=original_spellings&amp;smoothing=True&amp;rollingAverage=15_year\">EEBO-TCP Ngram Browser<\/a> page. It looks like this:<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-726\" src=\"http:\/\/samuli.kaislaniemi.fi\/blog\/wp-content\/uploads\/2019\/04\/1630s-spelling-change-EEBO-1.png\" alt=\"\" width=\"940\" height=\"630\" srcset=\"https:\/\/samuli.kaislaniemi.fi\/blog\/wp-content\/uploads\/2019\/04\/1630s-spelling-change-EEBO-1.png 940w, https:\/\/samuli.kaislaniemi.fi\/blog\/wp-content\/uploads\/2019\/04\/1630s-spelling-change-EEBO-1-300x201.png 300w, https:\/\/samuli.kaislaniemi.fi\/blog\/wp-content\/uploads\/2019\/04\/1630s-spelling-change-EEBO-1-768x515.png 768w\" sizes=\"auto, (max-width: 940px) 100vw, 940px\" \/><\/p>\n<p><span style=\"line-height: 1.5em;\">The idea of an ngram viewer \u2013 as per Google \u2013\u00a0is to look at the frequency of occurrences over time, of a word (a 1-gram) or a phrase (2-, 3-, 4- \u2026 N-gram). Frequency here means the proportion of the search phrase to all the words in the corpus, plotted over time. So for instance, the frequency of the word &#8220;war&#8221; rises during wartime, and falls in peacetime. But things get much more interesting when you look at less obvious things.<\/span><\/p>\n<p>Anyway.\u00a0The point of the EEBO-TCP spelling variant ngram viewer is to compare the change and development of spelling variants over time: for instance, plotting &#8220;spell&#8221; against &#8220;spelle&#8221;, &#8220;spel&#8221;, etc:<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-733\" src=\"http:\/\/samuli.kaislaniemi.fi\/blog\/wp-content\/uploads\/2019\/04\/spel.png\" alt=\"\" width=\"963\" height=\"592\" srcset=\"https:\/\/samuli.kaislaniemi.fi\/blog\/wp-content\/uploads\/2019\/04\/spel.png 963w, https:\/\/samuli.kaislaniemi.fi\/blog\/wp-content\/uploads\/2019\/04\/spel-300x184.png 300w, https:\/\/samuli.kaislaniemi.fi\/blog\/wp-content\/uploads\/2019\/04\/spel-768x472.png 768w\" sizes=\"auto, (max-width: 963px) 100vw, 963px\" \/><\/p>\n<p>English spelling only became standardized in the 18th century, and anyone who wants to read earlier texts has to learn to deal with the fact that apparently all spellings were equally acceptable, and that writers haphazardly used the first spelling that came to their mind* \u2013 one of the most famous (or notorious) examples being how William Shakespeare <i><a href=\"http:\/\/en.wikipedia.org\/wiki\/Shakespeare%27s_handwriting#Shakespeare.27s_signatures\">signed his name<\/a><\/i>\u00a0in six different ways. Despite eventual standardization, spelling variation in English has not completely disappeared today, for although varying how you spell your name today sounds outrageous and unthinkable, all students of English as a foreign language have to learn that there are British and American spellings for many familiar words: <i>colour<\/i> and <i>color<\/i>, <i>standardize<\/i> and <i>standardise<\/i>, etc.<\/p>\n<h2>2. What the hell happened in 1625?<\/h2>\n<p>But to return to the EEBO-TCP Spelling Browser, what struck me was the dramatic change in the 1630s.\u00a0<span style=\"line-height: 1.5em;\">If you look back to the first figure above, you can see that of two spellings of the word <\/span><em style=\"line-height: 1.5em;\">above<\/em><span style=\"line-height: 1.5em;\">, the spelling &#8220;aboue&#8221; is essentially the given form until 1625, when it rapidly loses to the alternative spelling &#8220;above&#8221;, which is firmly established by about 1640.<\/span><\/p>\n<p>..Hang on, what? The centuries-old practice of not differentiating between the graphemes &lt;u&gt; and &lt;v<em>&gt;<\/em>\u00a0according to the phonemes they indicate \u2013 \/u\/ and \/v\/ \u2013 is replaced, over the stunningly short period of\u00a0<em>15 years<\/em> \u2013 across the board (!?) in printed texts by consistent mapping of &lt;u&gt; to \/u\/ and &lt;v&gt; to \/v\/..!?<\/p>\n<blockquote class=\"twitter-tweet\" data-width=\"550\" data-dnt=\"true\">\n<p lang=\"en\" dir=\"ltr\"><a href=\"https:\/\/x.com\/heatherfro?ref_src=twsrc%5Etfw\">@heatherfro<\/a> \u2026That was 90mins in the middle of the night playing with EModE spelling variation. What the hell happened in 1625?!?<\/p>\n<p>&mdash; samklai in the other place too (@samklai) <a href=\"https:\/\/x.com\/samklai\/status\/459442820546584576?ref_src=twsrc%5Etfw\">April 24, 2014<\/a><\/p><\/blockquote>\n<p><script async src=\"https:\/\/platform.x.com\/widgets.js\" charset=\"utf-8\"><\/script><\/p>\n<p>Sooo many questions.<\/p>\n<p>My very first thought was that it must be an artefact of the dataset. One word, of course, hardly tells the whole story. Did this change hold for other words that show u\/v variation? What about i\/j variation?\u00a0Or perhaps the EEBO-TCP material was somehow skewed?<\/p>\n<p>But however much I fiddled with the browser, the period between 1620 and 1640 remained the significant factor. And it also applied for i\/j-words:<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-727\" src=\"http:\/\/samuli.kaislaniemi.fi\/blog\/wp-content\/uploads\/2019\/04\/1630s-spelling-change-EEBO-2.png\" alt=\"\" width=\"938\" height=\"626\" srcset=\"https:\/\/samuli.kaislaniemi.fi\/blog\/wp-content\/uploads\/2019\/04\/1630s-spelling-change-EEBO-2.png 938w, https:\/\/samuli.kaislaniemi.fi\/blog\/wp-content\/uploads\/2019\/04\/1630s-spelling-change-EEBO-2-300x200.png 300w, https:\/\/samuli.kaislaniemi.fi\/blog\/wp-content\/uploads\/2019\/04\/1630s-spelling-change-EEBO-2-768x513.png 768w\" sizes=\"auto, (max-width: 938px) 100vw, 938px\" \/><\/p>\n<p>But I did also check <a href=\"http:\/\/eebo.chadwyck.com\/home\">EEBO proper<\/a> \u2013 knowing that the results may well be different from those of the EEBO-TCP Ngram Browser. However, it turned out once again that the Browser had been right:<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-730\" src=\"http:\/\/samuli.kaislaniemi.fi\/blog\/wp-content\/uploads\/2019\/04\/joy-ioy-in-EEBO-proper.png\" alt=\"\" width=\"755\" height=\"416\" srcset=\"https:\/\/samuli.kaislaniemi.fi\/blog\/wp-content\/uploads\/2019\/04\/joy-ioy-in-EEBO-proper.png 755w, https:\/\/samuli.kaislaniemi.fi\/blog\/wp-content\/uploads\/2019\/04\/joy-ioy-in-EEBO-proper-300x165.png 300w\" sizes=\"auto, (max-width: 755px) 100vw, 755px\" \/><\/p>\n<p>So what on earth happened in the 1620s and 1630s to explain this dramatic shift into standardized spelling?<\/p>\n<p>\u2026actually, I don&#8217;t know. Googling revealed that, on the one hand, this is a known phenomenon \u2013 although I so far have not found a definitive study of the phenomenon nor a good explanation. (Clearly it has something to do with what&#8217;s going on in printing houses). But for instance in\u00a0<a href=\"http:\/\/universitypublishingonline.org\/cambridge\/histories\/chapter.jsf?bid=CBO9781139053747&amp;cid=CBO9781139053747A005\">her article<\/a> in the <em>Cambridge History of the English Language<\/em> vol 3 (2000),\u00a0Vivian Salmon discusses historical variation in using &lt;u&gt; and &lt;v&gt; to indicate both \/u\/ and \/v\/, and then quite casually mentions how &#8220;the distinction was made in the 1630s&#8221; (p. 39). I think that a corpus-based study of this change remains to be done \u2013 although I could be wrong.<\/p>\n<p>Yet rather than starting to look at this point in more detail, I pursued another question that had come to mind: how did this shift in orthographical practices manifest in non-printed texts, such as letters?<\/p>\n<h2>3. Non-printed texts and manuscripts<\/h2>\n<p>Happily, I am in a perfect position to ask this question, being part of the team who have compiled the\u00a0<a href=\"http:\/\/www.helsinki.fi\/varieng\/CoRD\/corpora\/CEEC\/index.html\"><em>Corpus of Early English Correspondence<\/em> (CEEC)<\/a>. The CEEC is a corpus of English personal letters, spanning 1400-1800 and presently containing about 12,000 letters (5.2m words). It was designed for historical sociolinguistics \u2013 to apply modern sociolinguistics methods on historical texts.<\/p>\n<p>Of course, there&#8217;s a caveat: CEEC is based on printed editions of letters. &#8220;Hang on&#8221;, you might say, &#8220;is a corpus built from such sources linguistically reliable? Shouldn&#8217;t the corpus have been compiled from manuscript texts?&#8221; Well, yes \u2013 but we have been careful not to use editions that modernize the letter texts, as well as editions that normalize the texts extensively. For the kinds of linguistic queries that the corpus was designed for, the normalization of features such as u\/v variation was deemed acceptable. And we have always been careful to stress that the CEEC is\u00a0<em>not<\/em> suitable for studying English orthography.<\/p>\n<p>Anyway, I nonetheless rushed right in to see what the CEEC threw up. Not having fancy tools (like <a href=\"http:\/\/corpora.lancs.ac.uk\/dicer\/\">DICER<\/a>) to reveal the proper extent of variant spellings in the corpus, I used a short list of sample words (euer\/ever, ouer\/over, aboue\/above, vp\/up). But the results were underwhelming:<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-728\" src=\"http:\/\/samuli.kaislaniemi.fi\/blog\/wp-content\/uploads\/2019\/04\/CEEC-u.v.-17C.png\" alt=\"\" width=\"542\" height=\"362\" srcset=\"https:\/\/samuli.kaislaniemi.fi\/blog\/wp-content\/uploads\/2019\/04\/CEEC-u.v.-17C.png 542w, https:\/\/samuli.kaislaniemi.fi\/blog\/wp-content\/uploads\/2019\/04\/CEEC-u.v.-17C-300x200.png 300w\" sizes=\"auto, (max-width: 542px) 100vw, 542px\" \/><\/p>\n<p>In this figure, the ratio of the old form of u\/v-spelling variants was far too low through the whole period \u2013 it should have been at least around 80%, if the EEBO data was indicative of English spelling practices overall, rather than just those restricted to printed texts.<\/p>\n<p>In order to have better data, I spent some time extracting a subcorpus from the CEEC\u2020 consisting of texts only from editions of 17th-century letters in which I could find u\/v and i\/j-variation. This time, the results were more interesting:<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-729\" src=\"http:\/\/samuli.kaislaniemi.fi\/blog\/wp-content\/uploads\/2019\/04\/CEEC_part-u.v.-17C.png\" alt=\"\" width=\"790\" height=\"492\" srcset=\"https:\/\/samuli.kaislaniemi.fi\/blog\/wp-content\/uploads\/2019\/04\/CEEC_part-u.v.-17C.png 790w, https:\/\/samuli.kaislaniemi.fi\/blog\/wp-content\/uploads\/2019\/04\/CEEC_part-u.v.-17C-300x187.png 300w, https:\/\/samuli.kaislaniemi.fi\/blog\/wp-content\/uploads\/2019\/04\/CEEC_part-u.v.-17C-768x478.png 768w\" sizes=\"auto, (max-width: 790px) 100vw, 790px\" \/><\/p>\n<p>Although the ratio of old spelling variants is still much lower than I had expected, in this figure there is a sharp decline from the 1640s on \u2013 which would be in accordance to a prescribed change. (For example, if all schoolchildren are taught to spell according to certain rules, it takes a while for the older generations of writers to die out (or change their spelling habits). Similarly, it makes sense that the influence of a standard orthography in printed texts would reflect in manuscript texts with a slight time lag.)<\/p>\n<p>Yet I remained unhappy with this data. In EEBO, the shift is from nearly 100% old form to 100% new form. Clearly the texts of the editions used for CEEC were normalized more than I had thought. Even given that this was a quick pilot study, the discrepancy was simply too large to accept as a difference between orthographical practices of manuscript and print.<\/p>\n<p>I had one last trick up my sleeve: I did have a fairly good-sized corpus of letters from the first decade of the 1600s\u00a0transcribed from manuscript, which retained original spellings and other orthographical features. It wouldn&#8217;t show me change over time, but it would give me a control figure for how much, exactly, were letter-writers using the old forms in their letters.\u2021<\/p>\n<p>The result can be seen in the figure above \u2013 it is the red X, marking a whopping 87.9% old forms. Finally,\u00a0<em>something<\/em> resembling the situation in EEBO.<\/p>\n<p>There was a fair bit of variation between different words in the manuscript sources, and in some cases the new form was dominant:<\/p>\n<table style=\"border: 1px solid #cccccc; padding: 6px;\">\n<tbody>\n<tr>\n<td style=\"border: 1px solid #cccccc; padding: 6px;\"><\/td>\n<td style=\"border: 1px solid #cccccc; padding: 6px;\">euer\/ever<\/td>\n<td style=\"border: 1px solid #cccccc; padding: 6px;\">adu*\/adv*<\/td>\n<td style=\"border: 1px solid #cccccc; padding: 6px;\">haue\/have<\/td>\n<\/tr>\n<tr>\n<td style=\"border: 1px solid #cccccc; padding: 6px;\">old spelling<\/td>\n<td style=\"border: 1px solid #cccccc; padding: 6px;\">35<\/td>\n<td style=\"border: 1px solid #cccccc; padding: 6px;\">90<\/td>\n<td style=\"border: 1px solid #cccccc; padding: 6px;\">1017<\/td>\n<\/tr>\n<tr>\n<td style=\"border: 1px solid #cccccc; padding: 6px;\">new spelling<\/td>\n<td style=\"border: 1px solid #cccccc; padding: 6px;\">28<\/td>\n<td style=\"border: 1px solid #cccccc; padding: 6px;\">191<\/td>\n<td style=\"border: 1px solid #cccccc; padding: 6px;\">20<\/td>\n<\/tr>\n<tr>\n<td style=\"border: 1px solid #cccccc; padding: 6px;\">% old<\/td>\n<td style=\"border: 1px solid #cccccc; padding: 6px;\">56%<\/td>\n<td style=\"border: 1px solid #cccccc; padding: 6px;\">32%<\/td>\n<td style=\"border: 1px solid #cccccc; padding: 6px;\">98%<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>The greatest discrepancy between the manuscript sources and CEEC (namely the second extracted subcorpus) could be seen in the fact that in the manuscripts, words beginning \/un-\/ were spelled with a &lt;v&gt; 99.6% of the time (of 987 tokens), whereas in CEEC, the &lt;v&gt;-form occurred only 31% of the time (of 295 tokens). Even editions which claim to retain original spellings clearly cannot be taken at face value.<\/p>\n<h2>4. Summing up<\/h2>\n<p>So what can we say about that dramatic end to spelling variation in the 1630s seen in the figures from the EEBO-TCP Ngram Browser? Actually, not much.<\/p>\n<p>1. It would appear that in the EEBO corpus, u\/v variation became standardized between 1620 and 1640. However, without a comprehensive survey even this conclusion may be wrong \u2013 cf. for instance i\/j variation in the proper name James, where it takes longer for the &lt;i&gt;-form to start declining, nor is it gone by the end of the century (this might have to do with capitalisation):<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-732\" src=\"http:\/\/samuli.kaislaniemi.fi\/blog\/wp-content\/uploads\/2019\/04\/iames-james.png\" alt=\"\" width=\"953\" height=\"590\" srcset=\"https:\/\/samuli.kaislaniemi.fi\/blog\/wp-content\/uploads\/2019\/04\/iames-james.png 953w, https:\/\/samuli.kaislaniemi.fi\/blog\/wp-content\/uploads\/2019\/04\/iames-james-300x186.png 300w, https:\/\/samuli.kaislaniemi.fi\/blog\/wp-content\/uploads\/2019\/04\/iames-james-768x475.png 768w\" sizes=\"auto, (max-width: 953px) 100vw, 953px\" \/><\/p>\n<p>2. In manuscript texts, it\u00a0<em>looks like<\/em> the spelling standardization process occurred 20 or more years after it took place in print. But without a broader survey, even this estimate may be well wrong.<\/p>\n<p>3. The EEBO Spelling Browser is awesome!<\/p>\n<p>I do remain curious about what happened in the 1620s &amp; 30s. Particularly in whether the standardization of spelling was something more than a development in printing house practices. But I think I&#8217;ve done my share of midnight rabbit chasing for the moment.<\/p>\n<hr \/>\n<p>* Students of Early Modern English beware: this is not true! There are methods in the apparent madness, although the rules may be subtle, and they do vary between writers.<\/p>\n<p>\u2020 My first search of CEEC material was of c.\u00a05,000 letters\u00a0(2.2m words), finding\u00a06,657 tokens of which 700 were old spellings.\u00a0My second dataset consisted of c. 1,900 letters (just under 800k words), and\u00a03,613 tokens of which 887 were old forms \u00a0(types: euer\/ever, ouer\/over, aboue\/above, vp\/up, vs\/us).<\/p>\n<p>\u2021 This manuscript-based corpus contained about 200 letters (130k words). I expanded my sample word list \u00a0(types: euer\/ever, ouer\/over, aboue\/above, haue\/have, giue\/give, vp\/up, vs\/us, adu*\/adv*, vn*\/un*), extracting 2,712 tokens \u2013 of which 2,385 were old forms.<\/p>\n<p>&#8212;<\/p>\n<p>ETA 5.9.2024<\/p>\n<p>Belatedly realized there are links to this blog post out there, even in print! Here&#8217;s the old, now defunct, link; my blog used to be hosted by the University of Helsinki, while I was affiliated there:<\/p>\n<p>http:\/\/blogs.helsinki.fi\/kaislani\/2014\/04\/26\/spelling-variation<\/p>\n<p>(I&#8217;m hoping web crawlers will catch this so it becomes googlable, and leads here.)<\/p>\n","protected":false},"excerpt":{"rendered":"<p>1. Early Modern English spelling variation Yesterday, rather late in the evening, I followed a link on Twitter: So there&#39;s an EEBO-TCP spelling variation google ngram browser http:\/\/t.co\/OLxUv5NLBQ (via @dr_heil) &mdash; heather froehlich (@heatherfro) April 24, 2014 This led to the great Early Modern Print : Text Mining Early Printed English website where there was [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[5,10],"tags":[22,23,24,28,33,36,38,40,41,49,50],"class_list":["post-559","post","type-post","status-publish","format-standard","hentry","category-digital-humanities","category-survey","tag-cool-stuff","tag-corpora","tag-corpus-linguistics","tag-digital-humanities","tag-emode","tag-language","tag-letters","tag-linguistics","tag-manuscripts","tag-print","tag-printed-books","entry"],"_links":{"self":[{"href":"https:\/\/samuli.kaislaniemi.fi\/blog\/wp-json\/wp\/v2\/posts\/559","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/samuli.kaislaniemi.fi\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/samuli.kaislaniemi.fi\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/samuli.kaislaniemi.fi\/blog\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/samuli.kaislaniemi.fi\/blog\/wp-json\/wp\/v2\/comments?post=559"}],"version-history":[{"count":3,"href":"https:\/\/samuli.kaislaniemi.fi\/blog\/wp-json\/wp\/v2\/posts\/559\/revisions"}],"predecessor-version":[{"id":935,"href":"https:\/\/samuli.kaislaniemi.fi\/blog\/wp-json\/wp\/v2\/posts\/559\/revisions\/935"}],"wp:attachment":[{"href":"https:\/\/samuli.kaislaniemi.fi\/blog\/wp-json\/wp\/v2\/media?parent=559"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/samuli.kaislaniemi.fi\/blog\/wp-json\/wp\/v2\/categories?post=559"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/samuli.kaislaniemi.fi\/blog\/wp-json\/wp\/v2\/tags?post=559"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}