home.social

#voicedata — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #voicedata, aggregated by home.social.

fetched live
  1. Estamos expandiendo nuestras capacidades en Datos para IA en #Pangeanic y lanzamos 2 proyectos masivos dirigidos a particulares España/ América Latina.
    Buscamos perfiles que quieran colaborar desde casa en:
    a) Grabación de voz en español (más de 20 horas de audio, móvil/PC).
    b) Videos egocéntricos (perspectiva en primera persona con smartphone) en LATAM.
    Detalles y registro directo:
    pangeanic.com/es/trabajos/
    #ArtificialIntelligence #VoiceData #GenerativeAI #TrabajoRemoto #DataCollection

  2. 2600: The 2600 Voice BBS Archives, Hope_16 Videos, And A Shower Curtain. “For those who don’t know, the 2600 Voice BBS was a unique hacker bulletin board system run by 2600 in the late 1990s. Its purpose was to act like a regular BBS, only with voice instead of text. We’ve been putting this archive together for years and it’s a real time capsule of the hacker world. There are literally […]

    https://rbfirehose.com/2026/01/13/2600-the-2600-voice-bbs-archives-hope_16-videos-and-a-shower-curtain/
  3. 🗣️ Your smart speaker’s always got one ear open.

    It listens 24/7 for wake words like “Hey Alexa”—then sends recordings to the cloud (yes, sometimes reviewed by real people 👀).

    📌 Hit mute when not in use
    📌 Turn on auto-delete (major companies allow it!)
    📌 Wake words = always listening locally

    #SmartHome #VoiceData #PrivacyTips #CyberSecurity

  4. For the past couple of years, as each new @mozilla #CommonVoice dataset of #voice #data is released, I've been using @observablehq to visualise the #metadata coverage across the 100+ languages in the dataset.

    Version 17 was released yesterday (big ups to the team - EM Lewis-Jong, @jessie, Gina Moape, Dmitrij Feller) and there's some super interesting insights from the visualisation:

    ➡ Catalan (ca) now has more data in Common Voice than English (en) (!)

    ➡ The language with the highest average audio utterance duration at nearly 7 seconds is Icelandic (is). Perhaps Icelandic words are longer? I suspect so!

    ➡ Spanish (es), Bangla (Bengali) (bn), Mandarin Chinese (zh-CN) and Japanese (ja) all have a lot of recorded utterances that have not yet been validated. Albanian (sq) has the highest percentage of validated utterances, followed closely by Erzya / Arisa (myv).

    ➡ Votic (vot) has the highest percentage of invalidated utterances, but with 76% of utterances invalidated, I wonder if this language has been the target of deliberate invalidation activity (invalidating valid sentences, or recording sentences to be deliberately invalid) given the geopolitical instability in Russia currently.

    See the visualisation here and let me know your thoughts below!

    observablehq.com/@kathyreid/mo

    #linguistics #languages #data #VoiceAI #VoiceData #SpeechAI #SpeechData #DataViz

  5. #Marketing Company Claims That It Actually Is #Listening to Your Phone and Smart Speakers to Target #Ads

    “What would it mean for your business if you could target potential clients who are actively discussing their need for your services in their day-to-day conversations? No, it's not a #BlackMirror episode—it's #VoiceData, and #CMG has the capabilities to use it to your business advantage.”

    404media.co/cmg-cox-media-actu

  6. This is a fascinating article on the increasing use of #subtitles, by Claudia Forsberg for ABC #Ballarat - the way that sound is designed for movies intended for cinema means that it doesn't play back optimally on mobile devices or streaming services - and this is one factor driving the adoption of #ClosedCaptioning or #Subtitles.

    But, these #Subtitles are often inaccurate or mis-transcribed. They use #ASR technology - and this is another case for having good #voice #data #voicedata

    abc.net.au/news/2023-02-12/sub

  7. Do you work with #voice and #voicedata, such as #data used for #ASR & #TTS tools? This might be collecting, cleaning, parsing data, or using data to train and evaluate #ML models 💻 🎤 📈

    As part of my #PhD at ANU (not on fediverse), I'm doing some exploratory interviews to help shape a survey. The survey and interviews will help us understand #ML practice in voice, and create tools to help make voice tech fairer for everyone.

    Ethics protocol: 2021/427

    Would appreciate boosts / RTs for reach 💖