#d2e2015
Sorry, #d2e2015 .
January 6, 2025 at 1:27 AM
..whew, just made it to the opening plenary of #d2e2015
Packed room listening to @TonyMcEnery on "the corpus as social history"
January 6, 2025 at 1:22 AM
.@TonyMcEnery :
Lots of disciplines talk about language and make claims about language.
#d2e2015
January 6, 2025 at 1:17 AM
.@TonyMcEnery :
"Close reading is key – this is NOT culturomics, nor should it be"
#d2e2015
January 6, 2025 at 1:17 AM
.@TonyMcEnery points out OED has no frequency info on changing meanings of words. (the mind boggles, that would be incredible)
#d2e2015
January 6, 2025 at 1:17 AM
.@TonyMcEnery
Sometimes a modern ontology is just fundamentally wrong for a historical period
#d2e2015
January 6, 2025 at 1:17 AM
.@TonyMcEnery
pretty terrible examples from applying modern ontology on 17C text – errors in all or most semantic categories
#d2e2015
January 6, 2025 at 1:17 AM
Since @OxfordEEBOTCP came up, I can't help plugging the great EEBO ngram viewer by @anupam_basu
http://bit.ly/1CxlXfy
#d2e2015
January 6, 2025 at 1:17 AM
Here's a correspoding EEBO ngram graph of what @TonyMcEnery just showed in a slide
#d2e2015 http://t.co/dLhB8tCSBy
January 6, 2025 at 1:17 AM
Lancaster have been dividing the @OxfordEEBOTCP corpus into genres. "more strictures" will help with interpretation of data.
#d2e2015
January 6, 2025 at 1:12 AM
.@TonyMcEnery
Collocation linked to amount of data – more data = more collocates – but not straightforward correlation.
#d2e2015
January 6, 2025 at 1:12 AM
Revived by tea.
Now in session of number-crunching linguistics.
#d2e2015
January 6, 2025 at 1:12 AM
First up, James McCracken on integrating historical frequency data with the @OED – nicely picking up on comment in plenary!
#d2e2015
January 6, 2025 at 1:12 AM
McCracken:
@OED 's problem is representativeness
#d2e2015
January 6, 2025 at 1:12 AM
Fantastic animation of shifting proportions of HTOED categories over time by McCracken
#d2e2015
January 6, 2025 at 1:12 AM
McCracken:
"The @OED isn't Big Data, but it is Big Information" – useful to think of it as a network of features
#d2e2015
January 6, 2025 at 1:12 AM
Next up, Tyler Kendall on aggregating sociolinguistic datasets – on creating New sources from "old data"
#d2e2015
January 6, 2025 at 1:12 AM
Kendall's discussion of aggregating old data sets makes me think of the #interchange element of #TEIXML
#d2e2015
January 6, 2025 at 1:12 AM
Kendall: key issues:
Administrative: ethics, rights, ownership
Technical: formatting, metadata, discoverability, delivery
#d2e2015
January 6, 2025 at 1:07 AM
Kendall:
SLAAP is a digital archive collection socli data (interviews + transcripts) http://bit.ly/1NPFLA5
#d2e2015
Linguistics
Welcome to the NC State Linguistics program website. Ling...
bit.ly
January 6, 2025 at 1:07 AM
.@JWGrieve talking about work done on nearly 1 billion tweets – c. 8.9b words – by 7m different tweeters
#d2e2015
January 6, 2025 at 1:07 AM
.@JWGrieve showing how your tweet may reveal your location & we can find you in GMaps!
#d2e2015
"we do not use it at this level of detail"
January 6, 2025 at 1:07 AM
.@JWGrieve
on emerging words; mapping them; mapping uncommon grammatical constructions; ID:ing American cultural regions
#d2e2015
January 6, 2025 at 1:07 AM
I really like the idea of mining twitter for emerging words. I do similar stuff on 17C letters, but on a slightly smaller dataset..
#d2e2015
January 6, 2025 at 1:07 AM
.@JWGrieve
Top emerging words on twitter in 2014 – top 8 to do with profanity and intoxicants, many are acronyms
#d2e2015
January 6, 2025 at 1:07 AM