home.social

#interspeech — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #interspeech, aggregated by home.social.

fetched live
  1. @emilyskidsister hi, former linguistic engineer on the Common Voice project here, and I still work very closely with the project.

    We changed the optional gender label taxonomy in early 2024 with the growing recognition that genders are a multidimensional spectrum, not a binary classification, and with the explicit intent of ensuring better representation.

    mozillafoundation.org/en/blog/

    It's still not perfect, but IMHO better than it was.

    ---

    Labelling of any dataset used for machine learning is a double edged sword. On one hand, having more voices represented in ways that are accurate and truthful helps with ensuring these voices are better represented in speech and AI models.

    On the other hand, any sort of categorical label can be used for classification. For example, Common Voice data can include speaker self-labelled accent data, which can be used to create accent classification models - which can be harmful depending on how they're used.

    I wrote a paper that touches on this topic here:
    dl.acm.org/doi/10.1145/3617694

    ---

    IMHO one of the key *weaknesses* with CV is that it is released under a CC-0 license. When CV was created in 2017, it was appropriate - create and release as much data as possible to spur development of ASR models that were limited by data scarcity.

    Over time our understanding of data and its discontents - the rise of movements such as Data Feminism (Ignazio and Klein) and Design Justice (Costanza-Chock) has become more nuanced - but the mechanisms of control we have around how data is used for ML are not yet as sophisticated as the intellectual discourse around them.

    We're trying to change that at Moz Data Collective, where I work now (for total transparency).

    Data poisoning, IMHO, is a totally rational and reasonable response to the inability to have mechanisms that enforce how data contributors' data is used in downstream applications.

    ---

    If you're interested in this space, there's a lot of good work being exhibited at #Interspeech in Sydney in September that tackles genders representation in speech and voice technologies, too.

    ---

    Always interested in how we can do better and shepherd technology in ways that are beneficial, not harmful, and very open to chatting further.

  2. In typical fashion, I am working on 2 #Interspeech papers less than 3 days before deadline 😆

  3. Alongside my colleagues, Associate Professor Francis Tyers from Indiana University Bloomington and E.M. L., from Mozilla Data Collective, we're delighted to be presenting a Special Session at Interspeech 2026 in Sydney in September 2026 on:

    Safeguarding Synthetic #Speech: Ethical, Technical and Legal perspectives

    You can see the Special Session website at:

    safeguardingsyntheticspeech.or

    ---

    There are some outstanding Special Sessions this year, and I've taken the liberty of listing them - and their organisers. Do consider submitting to the #CfP - it's rare that #Interspeech makes it to Australia and we all want to make the most of it.

    #Indigenous voices in Speech Sciences and Technology

    Website at: indigenousvoicesinterspeech.gi

    Brought to you by Dr. Cat Kutay, Prof. Janet Wiles, A/Prof Celeste Rodríguez Louro, Dr. Ben Hutchinson, and Daan van Esch

    ECT-SpeechAI: #Explainability for #Compliance and #Trust in Speech AI

    Website at: sites.google.com/view/ect-spee

    Brought to you by Dr. Xiaoliang Wu, Sarah Kiden, Poppy Welch, Assistant Professor Sneha Das, Assistant Professor Zhengjun Yue, Dr. Sarenne Wallbridge, Assistant Professor Jennifer Williams, PhD

    Challenges in Speech Data Collection, Curation and Annotation

    Website at: sites.google.com/view/speech-d

    Brought to you by Mostafa Shahin, Tunde Szalay, Thomas Schaaf, Ahmed Ali, Jingyao (Jennifer) Wu, Mengyue Wu, Professor Tan Lee, Professor Mark Liberman, Professor Carlos Busso, A/Prof Beena Ahmed

    Post-Training of Speech Foundation Models

    Website at: sites.google.com/view/ptsfm

    Brought to you by Yang Xiao, Xiangyu Zhang, Ziyang Ma, Siyi Wang, Jiaheng Dong, Professor Eng Siong Chng, Dr Ting Dang

    CHILDSPACE: Child Home Interaction and Language Dynamics: Speech, Psychology, Affect, Computation and Environments

    Website at: sites.google.com/view/childspa

    Brought to you by Dr Kaveri Sheth, Dr Alejandrina Cristia, A/Prof Beena Ahmed, Assistant Professor Meg Cychosz, Dr Ting Dang, Kaya de Barbaro, Professor Carol Espy-Wilson, Rebecca Holt, Professor Herman Kamper, Marvin Lavechin, Assistant Professor Jialu Li, Ph.D., Professor Okko Rasanen, Professor Odette Scharenborg, Mostafa Shahin.

    Pacific Voices 2026: Speech Science and Technology for Languages of the Pacific Ocean

    Website at: speechresearch.auckland.ac.nz/

    Brought to you by: Associate Professor Peter Keegan, Dr Jesin James, Dr Isabella Shields, Dr Satwinder Singh, Dr Rolando Coto-Solano.

    Queer and Trans Speech Science and Technology

    Website at: sites.google.com/view/is-queer

    Brought to you by: Robin Netzorg, Nina Markl, Cliodhna Hughes, Juliana Francis, Francisca Pessanha.

    Speech AI for All: Challenging Deficit-based approaches to Speech Science and Technology

    Website at: sites.google.com/aimpower.org/

    #SpeakingTogether

  4. still hoping for ISCA or at least the ISCA-SAC to eventually also have a presence here, or at this point, to at least stop having one on Elon Musks X Dot Com

    ISCA have a bsky now, at least

    #isca #iscasac #interspeech

  5. Oh, #INTERSPEECH review assignments are out, guess I am just gonna do that today

  6. #interspeech reviews done, and before the first deadline, too!

    unfortunately I won’t be attending myself this year :(

  7. Lorenz Gutscher delivered a talk on #DialectEmbeddings, "Neural Speech Synthesis for Austrian Dialects with Standard German Grapheme-to-Phoneme Conversion and Dialect Embeddings", at the #Interspeech #SIGUL2023 workshop. Full paper and online demo: ofai.at/news/2023-08-18intersp

  8. InterSpeech 2023 is starting this week in sunny Dublin ! interspeech2023.org

    Congratulations to Sam Kotey presenting her work there on "Query Based Acoustic Summarization for Podcasts" #InterSpeech #InterSpeech2023

    doi.org/10.21437/Interspeech.2

    @AdaptCentre

  9. wonder if I'll see more fedi activity related to #INTERSPEECH this year

    not like it's usually a very social media active conference, but hey

  10. Today, we're releasing our #audio #speech quality assessment model for the audio packet-loss concealment ( #plc ) task, so you, too, can stop using mediocre metrics like PESQ and POLQA! #science

    github.com/microsoft/PLC-Chall

    Paper to appear at #INTERSPEECH2023 #interspeech , preprint for now: halcy.de/cites/papers/diener20 , arxiv soon

  11. What makes this a good #AI challenge?
    👉Natural code switching (no shuffled segments)
    👉Accented English & Mandarin
    👉Precision human annotation
    👉Various far field mics (laptops/tablets)
    👉Internet audio (Zoom)
    👉Adults speaking to kids

    #MerlionChallenge #Interspeech

  12. Participating teams have a chance to submit their papers at our special session at #Interspeech 2023 in Dublin (yes, Ireland) ☘️

    You can find out more about the #MerlionChallenge or sign up to take part by taking a look at our shiny new website!

    sites.google.com/view/merlion-

  13. For the #MerlionChallenge at #Interspeech we’ll be asking teams to train a #SpeechProc / #AI system that can guess which language is which (Task 1: Language ID) and when (Task 2: Language Diarization)!

    👉Challenge audio is Zoom recordings with English and/or Mandarin Chinese
    👉Audio for development matches audio for evaluation 😗👌

  14. The #MerlionChallenge is a 🌏 collab between psycholinguists at NTU in Singapore (me, Victoria Chua, Fei Ting Woon) and Engineers at JHU, TUD and NTU (Leibny Paola GarciaPerera, Sanjeev Khudanpur, Justin Dauwels, Hexin Liu, Andy Khong) for #Interspeech 2023

    interspeech2023.org

  15. Session about denoising now - paper about using perceptual loss as they do for images (think distance of last layer of vgg16) to train audio networks #interspeech #interspeech2019

  16. today’s keynote is on the physiology and physics of speech production! #interspeech

  17. no one visible to i.w has posted in the #interspeech hashtag since I did last year at Interspeech:|