home.social

#corpora — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #corpora, aggregated by home.social.

fetched live
  1. Later today at #CHR2024, we are going to present our work on #Multilingual #Stylometry!

    We isolated the influence of #language on #authorship #attribution #accuracy by translating multiple #corpora into each others' languages while keeping #corpus composition stable.

    Interactive showcase: showcases.clsinfra.io/stylomet

    Full paper: ceur-ws.org/Vol-3834/paper9.pd

    This work was developed within the @CLSinfra project in #Trier, #Krakow and #Prague with Artjoms Šeļa, Evgeniia Fileva and Julia Dudar.

  2. Later today at #CHR2024, we are going to present our work on #Multilingual #Stylometry!

    We isolated the influence of #language on #authorship #attribution #accuracy by translating multiple #corpora into each others' languages while keeping #corpus composition stable.

    Interactive showcase: showcases.clsinfra.io/stylomet

    Full paper: ceur-ws.org/Vol-3834/paper9.pd

    This work was developed within the @CLSinfra project in #Trier, #Krakow and #Prague with Artjoms Šeļa, Evgeniia Fileva and Julia Dudar.

  3. So you wanna parse/manipulate some #PDF's, huh!?

    Well, you better #test your #software thoroughly or bad things will happen!🧪

    So how about "this corpus [which] contains nearly 8 million PDFs gathered from across the web in July/August of 2021":
    digitalcorpora.org/corpora/fil

    The entire corpus when uncompressed takes up nearly 8 TB!

    You can find some more links to different #corpora (even to ones deemed #unsafe!😬) at pdf-association's Github:

    github.com/pdf-association/pdf

    #Parsing #Testing

  4. So you wanna parse/manipulate some #PDF's, huh!?

    Well, you better #test your #software thoroughly or bad things will happen!🧪

    So how about "this corpus [which] contains nearly 8 million PDFs gathered from across the web in July/August of 2021":
    digitalcorpora.org/corpora/fil

    The entire corpus when uncompressed takes up nearly 8 TB!

    You can find some more links to different #corpora (even to ones deemed #unsafe!😬) at pdf-association's Github:

    github.com/pdf-association/pdf

    #Parsing #Testing

  5. Interestingly, very few psychologists are aware of #linguistic #corpora 📊 and their immense research potential. Platforms like CLARIN-PL offer invaluable data that can significantly enhance our understanding of human behaviour and social interactions. 🤝🗣️ It's time more of us psych folk tapped into these resources to advance our field! 🌟🔍

  6. Interestingly, very few psychologists are aware of #linguistic #corpora 📊 and their immense research potential. Platforms like CLARIN-PL offer invaluable data that can significantly enhance our understanding of human behaviour and social interactions. 🤝🗣️ It's time more of us psych folk tapped into these resources to advance our field! 🌟🔍

  7. And another one for fellow linguists interested in compiling #corpora of digital discourse: MastoScraper takes advantage of the Mastodon API to collect toots based on a keyword search.
    Here goes, feedback welcome!
    #linguistics @linguistics
    fmoncomble.github.io/mastoscra

  8. And another one for fellow linguists interested in compiling #corpora of digital discourse: MastoScraper takes advantage of the Mastodon API to collect toots based on a keyword search.
    Here goes, feedback welcome!
    #linguistics @linguistics
    fmoncomble.github.io/mastoscra

  9. Next week, we'll be discussing how to archive and research social media data on a large scale "After Twitter". Very excited to see what comes out of this conference, and also the following data sprint delving into huge German Twitter corpora.
    dnb.de/twittertagung
    #AfterTwitter #corpora #research

  10. Next week, we'll be discussing how to archive and research social media data on a large scale "After Twitter". Very excited to see what comes out of this conference, and also the following data sprint delving into huge German Twitter corpora.
    dnb.de/twittertagung
    #AfterTwitter #corpora #research

  11. interesting publication on medieval Latin text corpora by @TimGeelhaar : 🔖 Geelhaar, Tim. „Hamsterrad oder Himmelsleiter? Oder warum die Digitalisierung so endlos scheint“. Application/epub+zip,application/pdf, 2024. doi.org/10.15499/KDS-005-016.

    #Latin #Neolatin #Corpora #OpenAccess

  12. interesting publication on medieval Latin text corpora by @TimGeelhaar : 🔖 Geelhaar, Tim. „Hamsterrad oder Himmelsleiter? Oder warum die Digitalisierung so endlos scheint“. Application/epub+zip,application/pdf, 2024. doi.org/10.15499/KDS-005-016.

    #Latin #Neolatin #Corpora #OpenAccess

  13. #Eduhub days 2024 at #ZHAW and I cannot be there 😢 🩼

    If you go, stop by at the marketplace—in the afternoon my colleague Maren Runte will show our work on creating a learning space for working with linguistic #corpora (to be released later this year)

    #EduhubDays24 #DigitalLinguistics

    eduhubdays2024.events.switch.c

  14. #Eduhub days 2024 at #ZHAW and I cannot be there 😢 🩼

    If you go, stop by at the marketplace—in the afternoon my colleague Maren Runte will show our work on creating a learning space for working with linguistic #corpora (to be released later this year)

    #EduhubDays24 #DigitalLinguistics

    eduhubdays2024.events.switch.c

  15. CLS INFRA Training School Vienna 2024 June 10th–12th, 2024:

    ExploreCor: "Using Programmable #Corpora in #Computational #Literary #Studies"

    This intensive program covers some of the most important steps in the research cycle of CLS, focusing on “Programmable Corpora” – dynamic collections of literary texts manipulated programmatically.

    Apply now! pretix.eu/CLSINFRA-trainingsch

    Colocated with #CCLS2024, June 13-14, 2024: jcls.io/site/conference/

    @CLSinfra #CLSINFRA

  16. CLS INFRA Training School Vienna 2024 June 10th–12th, 2024:

    ExploreCor: "Using Programmable #Corpora in #Computational #Literary #Studies"

    This intensive program covers some of the most important steps in the research cycle of CLS, focusing on “Programmable Corpora” – dynamic collections of literary texts manipulated programmatically.

    Apply now! pretix.eu/CLSINFRA-trainingsch

    Colocated with #CCLS2024, June 13-14, 2024: jcls.io/site/conference/

    @CLSinfra #CLSINFRA

  17. @ZBW_MediaTalk Das Tagungsprogramm steht (und wird bald veröffentlicht), Bewerbungen für den folgenden Data Sprint sind noch möglich! #datasprint #twitter #dataScience #corpora
    dnb.de/twitterdatasprint

  18. @ZBW_MediaTalk Das Tagungsprogramm steht (und wird bald veröffentlicht), Bewerbungen für den folgenden Data Sprint sind noch möglich! #datasprint #twitter #dataScience #corpora
    dnb.de/twitterdatasprint

  19. Oliver Watteler and Ulrike Schneider are talking about "Can I publish my social media corpus" @ #cmc2023 #corpora #socialMedia #linguistics #gdpr

  20. Oliver Watteler and Ulrike Schneider are talking about "Can I publish my social media corpus" @ #cmc2023 #corpora #socialMedia #linguistics #gdpr

  21. One of the prettiest (if not very practical) university locations in Germany. 👋 from #cmc2023 at the University of Mannheim! #corpora #linguistics #socialMedia

  22. One of the prettiest (if not very practical) university locations in Germany. 👋 from #cmc2023 at the University of Mannheim! #corpora #linguistics #socialMedia

  23. #corpora I bet we all have some anxiety about them. How do you choose what texts to look at? How do you know when you have enough? The #DataSittersClub is back, asking those questions and more to corpus linguist Shelley Staples, while exploring pizza and the Newbery Award for youth literature. #DigitalHumanities datasittersclub.github.io/site

  24. #corpora I bet we all have some anxiety about them. How do you choose what texts to look at? How do you know when you have enough? The #DataSittersClub is back, asking those questions and more to corpus linguist Shelley Staples, while exploring pizza and the Newbery Award for youth literature. #DigitalHumanities datasittersclub.github.io/site

  25. Is there any better news to wake up to than the fact that Norway has digitized All The Books and it's no problem at all to get all their Baby-Sitters Club translations? 🤩 #DigitalHumanities #DataSittersClub #corpora

  26. Is there any better news to wake up to than the fact that Norway has digitized All The Books and it's no problem at all to get all their Baby-Sitters Club translations? 🤩 #DigitalHumanities #DataSittersClub #corpora

  27. If you need to wrangle with EEBO-TCP for your text analysis project, consider using the EarlyPrint project corpus. They've done a bunch of preprocessing to transform "the early English print record, from 1473 to the early 1700s, into a linguistically annotated and deeply searchable text archive." Documentation and tutorials are all really thorough. earlyprint.org/about/ humanitiesdata.com/resources/4 -tcp

  28. If you need to wrangle with EEBO-TCP for your text analysis project, consider using the EarlyPrint project corpus. They've done a bunch of preprocessing to transform "the early English print record, from 1473 to the early 1700s, into a linguistically annotated and deeply searchable text archive." Documentation and tutorials are all really thorough. earlyprint.org/about/ humanitiesdata.com/resources/4 #CulturalAnalytics #dh #opendata #corpora #eebo-tcp

  29. Update: The submission deadline has been extended to May 21 2023! 📣 #CfP for the 10th International Conference on CMC and Social Media Corpora for the Humanities 2023! More info: 👉uni-mannheim.de/cmc-corpora202 The conference will be held at the University of Mannheim (@unimannheim) in collaboration with the IDS, from 14–15 September 2023.
    #CallForPapers #conference #DigitalHumanities #Corpuslinguistics #Korpuslinguistik #Corpora #SocialMedia #Linguistics #Linguistik #IDSMannheim

  30. Have you explored the new #EighteenthCenturyPoetryArchive corpus builder yet? It allows you to quickly create and share collections of poems, editions, or lists of authors with a single link!

    eighteenthcenturypoetry.org/re

    #c18th #poetry #18thC #c18dh #ECPA #corpora #readinglists

  31. Have you explored the new #EighteenthCenturyPoetryArchive corpus builder yet? It allows you to quickly create and share collections of poems, editions, or lists of authors with a single link!

    eighteenthcenturypoetry.org/re

    #c18th #poetry #18thC #c18dh #ECPA #corpora #readinglists

  32. How many copies of Matthias's vacation message do we all get before someone at ELRA figures out how to filter them?

    #NLP #corpora #email

  33. How many copies of Matthias's vacation message do we all get before someone at ELRA figures out how to filter them?

    #NLP #corpora #email

  34. Hi Fediverse, #introduction
    Currently I'm spending a lot of my time on the computer researching into #music #corpora in order to finish my #phd @ #epfl by the end of 2023. My main subject is #musicTheory and I'm trying to measure stylistic differences between tonal languages of the last four centuries through #statistics on #harmony (#stylometry).
    I'm here to connect with people who are interested in #dh #DataScience #machinelearning #opendata #dataset #foss #privacy #musicianship #funk #techno