home.social

#whisper — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #whisper, aggregated by home.social.

  1. Trasforma la voce in testo, note e comandi direttamente dal desktop con OpenWhispr, un progetto open source attento alla privacy e disponibile anche per Linux. #Linux #OpenSource #OpenWhispr #AI #Whisper #VoiceToText linuxeasy.org/openwhispr-porta

  2. Trasforma la voce in testo, note e comandi direttamente dal desktop con OpenWhispr, un progetto open source attento alla privacy e disponibile anche per Linux. #Linux #OpenSource #OpenWhispr #AI #Whisper #VoiceToText linuxeasy.org/openwhispr-porta

  3. Trasforma la voce in testo, note e comandi direttamente dal desktop con OpenWhispr, un progetto open source attento alla privacy e disponibile anche per Linux. #Linux #OpenSource #OpenWhispr #AI #Whisper #VoiceToText linuxeasy.org/openwhispr-porta

  4. Trasforma la voce in testo, note e comandi direttamente dal desktop con OpenWhispr, un progetto open source attento alla privacy e disponibile anche per Linux. #Linux #OpenSource #OpenWhispr #AI #Whisper #VoiceToText linuxeasy.org/openwhispr-porta

  5. Trasforma la voce in testo, note e comandi direttamente dal desktop con OpenWhispr, un progetto open source attento alla privacy e disponibile anche per Linux. #Linux #OpenSource #OpenWhispr #AI #Whisper #VoiceToText linuxeasy.org/openwhispr-porta

  6. Étude du consortium ARIANE sur la transcription via IA générative avec le LLM Whisper : « Whisper : transcription automatique de sources sonores » hal.science/hal-05706209 #Whisper #LLM #ESR

  7. Étude du consortium ARIANE sur la transcription via IA générative avec le LLM Whisper : « Whisper : transcription automatique de sources sonores » hal.science/hal-05706209 #Whisper #LLM #ESR

  8. Étude du consortium ARIANE sur la transcription via IA générative avec le LLM Whisper : « Whisper : transcription automatique de sources sonores » hal.science/hal-05706209 #Whisper #LLM #ESR

  9. Étude du consortium ARIANE sur la transcription via IA générative avec le LLM Whisper : « Whisper : transcription automatique de sources sonores » hal.science/hal-05706209 #Whisper #LLM #ESR

  10. Étude du consortium ARIANE sur la transcription via IA générative avec le LLM Whisper : « Whisper : transcription automatique de sources sonores » hal.science/hal-05706209 #Whisper #LLM #ESR

  11. LLM 改变了很多,比如 Agent 概念上半年爆火后,大厂们都争前恐后推出 CLI 产品,知乎,微信,网易云等等,这在几年前是完全无法想象的,那时想让 Linux 用上微信都很难,需要折腾一大堆 Wine 的组件,有时还会面临封号的风险。反观现在拥抱全平台的吃相,让人忍俊不禁。

    但真正改变的,其实很有限,因为爆火的原因不是厂商良心,拥抱开源,而是发现这个市场有的赚,本质还是逐利的,Agent 让厂商发现开放生态有的赚。工具功能很全,很方便,但我却更加担心了,工具会一直免费下去吗?我猜不会,等到官方的生态建立起来,我想就是回血的时候了。

    值得高兴的是,这些工具还很烂,如果真的要用,还太早,比如 npmjs.com/package/@music163/nc

    我无法相信官方的 TUI 居然吹能播放一半不到的歌,真是充满恶意的设计

    #whisper #netease #cli #llm

  12. LLM 改变了很多,比如 Agent 概念上半年爆火后,大厂们都争前恐后推出 CLI 产品,知乎,微信,网易云等等,这在几年前是完全无法想象的,那时想让 Linux 用上微信都很难,需要折腾一大堆 Wine 的组件,有时还会面临封号的风险。反观现在拥抱全平台的吃相,让人忍俊不禁。

    但真正改变的,其实很有限,因为爆火的原因不是厂商良心,拥抱开源,而是发现这个市场有的赚,本质还是逐利的,Agent 让厂商发现开放生态有的赚。工具功能很全,很方便,但我却更加担心了,工具会一直免费下去吗?我猜不会,等到官方的生态建立起来,我想就是回血的时候了。

    值得高兴的是,这些工具还很烂,如果真的要用,还太早,比如 npmjs.com/package/@music163/nc

    我无法相信官方的 TUI 居然吹能播放一半不到的歌,真是充满恶意的设计

    #whisper #netease #cli #llm

  13. LLM 改变了很多,比如 Agent 概念上半年爆火后,大厂们都争前恐后推出 CLI 产品,知乎,微信,网易云等等,这在几年前是完全无法想象的,那时想让 Linux 用上微信都很难,需要折腾一大堆 Wine 的组件,有时还会面临封号的风险。反观现在拥抱全平台的吃相,让人忍俊不禁。

    但真正改变的,其实很有限,因为爆火的原因不是厂商良心,拥抱开源,而是发现这个市场有的赚,本质还是逐利的,Agent 让厂商发现开放生态有的赚。工具功能很全,很方便,但我却更加担心了,工具会一直免费下去吗?我猜不会,等到官方的生态建立起来,我想就是回血的时候了。

    值得高兴的是,这些工具还很烂,如果真的要用,还太早,比如 npmjs.com/package/@music163/nc

    我无法相信官方的 TUI 居然吹能播放一半不到的歌,真是充满恶意的设计

    #whisper #netease #cli #llm

  14. LLM 改变了很多,比如 Agent 概念上半年爆火后,大厂们都争前恐后推出 CLI 产品,知乎,微信,网易云等等,这在几年前是完全无法想象的,那时想让 Linux 用上微信都很难,需要折腾一大堆 Wine 的组件,有时还会面临封号的风险。反观现在拥抱全平台的吃相,让人忍俊不禁。

    但真正改变的,其实很有限,因为爆火的原因不是厂商良心,拥抱开源,而是发现这个市场有的赚,本质还是逐利的,Agent 让厂商发现开放生态有的赚。工具功能很全,很方便,但我却更加担心了,工具会一直免费下去吗?我猜不会,等到官方的生态建立起来,我想就是回血的时候了。

    值得高兴的是,这些工具还很烂,如果真的要用,还太早,比如 npmjs.com/package/@music163/nc

    我无法相信官方的 TUI 居然吹能播放一半不到的歌,真是充满恶意的设计

    #whisper #netease #cli #llm

  15. LLM 改变了很多,比如 Agent 概念上半年爆火后,大厂们都争前恐后推出 CLI 产品,知乎,微信,网易云等等,这在几年前是完全无法想象的,那时想让 Linux 用上微信都很难,需要折腾一大堆 Wine 的组件,有时还会面临封号的风险。反观现在拥抱全平台的吃相,让人忍俊不禁。

    但真正改变的,其实很有限,因为爆火的原因不是厂商良心,拥抱开源,而是发现这个市场有的赚,本质还是逐利的,Agent 让厂商发现开放生态有的赚。工具功能很全,很方便,但我却更加担心了,工具会一直免费下去吗?我猜不会,等到官方的生态建立起来,我想就是回血的时候了。

    值得高兴的是,这些工具还很烂,如果真的要用,还太早,比如 npmjs.com/package/@music163/nc

    我无法相信官方的 TUI 居然吹能播放一半不到的歌,真是充满恶意的设计

    #whisper #netease #cli #llm

  16. Was für eine Sch*** !

    Ständig zerlegt sich die Whisper-Integration in KDEnlive. Kaum hat man die wieder geflickt, zerschießt das nächste Update sie wieder.

    Mal eben Untertitel erzeugen, wird so zur Qual.

    Ich prangere das an!

    Danke für Ihre Aufmerksamkeit für dieses Thema!

    #kdenlive #whisper #captions

  17. Was für eine Sch*** !

    Ständig zerlegt sich die Whisper-Integration in KDEnlive. Kaum hat man die wieder geflickt, zerschießt das nächste Update sie wieder.

    Mal eben Untertitel erzeugen, wird so zur Qual.

    Ich prangere das an!

    Danke für Ihre Aufmerksamkeit für dieses Thema!

    #kdenlive #whisper #captions

  18. Was für eine Sch*** !

    Ständig zerlegt sich die Whisper-Integration in KDEnlive. Kaum hat man die wieder geflickt, zerschießt das nächste Update sie wieder.

    Mal eben Untertitel erzeugen, wird so zur Qual.

    Ich prangere das an!

    Danke für Ihre Aufmerksamkeit für dieses Thema!

    #kdenlive #whisper #captions

  19. **Personal Vault #015 — Subtitles**

    I wanted to add Polish subtitles to movies and TV series in Personal Vault, but OpenSubtitles doesn't have Polish subtitles for everything in my library.

    I could connect another subtitle service, but that would mean relying on yet another external provider.

    So what's the alternative?

    Transcribe the dialogue myself, translate it, and generate the subtitles locally.

    Bingo! Or is it?

    I decided to test the idea using **Whisper** for multilingual transcription and timestamps, and **NLLB-200** for translation. I sent Codex a prompt and left it installing components and running tests while I explored Mastodon.

    Then I heard a sound coming from my laptop.

    I only had Codex and three Firefox tabs open.

    What on Earth is making THAT SOUND??

    Then I remembered: Codex has a built-in browser.

    I opened it.

    **Codex was watching Tron.**

    NOOOOO.

    Of all the movies to let an AI coding agent watch...

    Unfortunately, the choice wasn't much better. It was either **Tron** or **Terminator 2**.

    Neither seems like responsible viewing material for an AI agent.

    Codex was only testing the generated subtitles, but still...

    I sincerely hope I haven't accidentally created the Master Control Program or Skynet. 😛

    As for the experiment itself: technically, it worked.

    The transcription and subtitle generation were surprisingly good. Translation was the weak point.

    And that taught me something useful: **context is king**. Translating isolated subtitle lines is very different from understanding a conversation, a joke, a character, or what was said thirty seconds earlier. Without that context, even a very capable translation model can produce perfectly reasonable translations that are completely wrong for the scene.

    So I'm calling the experiment a successful failure.

    The idea works. The translation isn't good enough yet. I'll probably revisit it later.

    Assuming Codex hasn't taken over the world by then.

    #PersonalVault #DigitalIndependence #SelfHosting #Whisper #AI

  20. **Personal Vault #015 — Subtitles**

    I wanted to add Polish subtitles to movies and TV series in Personal Vault, but OpenSubtitles doesn't have Polish subtitles for everything in my library.

    I could connect another subtitle service, but that would mean relying on yet another external provider.

    So what's the alternative?

    Transcribe the dialogue myself, translate it, and generate the subtitles locally.

    Bingo! Or is it?

    I decided to test the idea using **Whisper** for multilingual transcription and timestamps, and **NLLB-200** for translation. I sent Codex a prompt and left it installing components and running tests while I explored Mastodon.

    Then I heard a sound coming from my laptop.

    I only had Codex and three Firefox tabs open.

    What on Earth is making THAT SOUND??

    Then I remembered: Codex has a built-in browser.

    I opened it.

    **Codex was watching Tron.**

    NOOOOO.

    Of all the movies to let an AI coding agent watch...

    Unfortunately, the choice wasn't much better. It was either **Tron** or **Terminator 2**.

    Neither seems like responsible viewing material for an AI agent.

    Codex was only testing the generated subtitles, but still...

    I sincerely hope I haven't accidentally created the Master Control Program or Skynet. 😛

    As for the experiment itself: technically, it worked.

    The transcription and subtitle generation were surprisingly good. Translation was the weak point.

    And that taught me something useful: **context is king**. Translating isolated subtitle lines is very different from understanding a conversation, a joke, a character, or what was said thirty seconds earlier. Without that context, even a very capable translation model can produce perfectly reasonable translations that are completely wrong for the scene.

    So I'm calling the experiment a successful failure.

    The idea works. The translation isn't good enough yet. I'll probably revisit it later.

    Assuming Codex hasn't taken over the world by then.

    #PersonalVault #DigitalIndependence #SelfHosting #Whisper #AI

  21. **Personal Vault #015 — Subtitles**

    I wanted to add Polish subtitles to movies and TV series in Personal Vault, but OpenSubtitles doesn't have Polish subtitles for everything in my library.

    I could connect another subtitle service, but that would mean relying on yet another external provider.

    So what's the alternative?

    Transcribe the dialogue myself, translate it, and generate the subtitles locally.

    Bingo! Or is it?

    I decided to test the idea using **Whisper** for multilingual transcription and timestamps, and **NLLB-200** for translation. I sent Codex a prompt and left it installing components and running tests while I explored Mastodon.

    Then I heard a sound coming from my laptop.

    I only had Codex and three Firefox tabs open.

    What on Earth is making THAT SOUND??

    Then I remembered: Codex has a built-in browser.

    I opened it.

    **Codex was watching Tron.**

    NOOOOO.

    Of all the movies to let an AI coding agent watch...

    Unfortunately, the choice wasn't much better. It was either **Tron** or **Terminator 2**.

    Neither seems like responsible viewing material for an AI agent.

    Codex was only testing the generated subtitles, but still...

    I sincerely hope I haven't accidentally created the Master Control Program or Skynet. 😛

    As for the experiment itself: technically, it worked.

    The transcription and subtitle generation were surprisingly good. Translation was the weak point.

    And that taught me something useful: **context is king**. Translating isolated subtitle lines is very different from understanding a conversation, a joke, a character, or what was said thirty seconds earlier. Without that context, even a very capable translation model can produce perfectly reasonable translations that are completely wrong for the scene.

    So I'm calling the experiment a successful failure.

    The idea works. The translation isn't good enough yet. I'll probably revisit it later.

    Assuming Codex hasn't taken over the world by then.

    #PersonalVault #DigitalIndependence #SelfHosting #Whisper #AI

  22. **Personal Vault #015 — Subtitles**

    I wanted to add Polish subtitles to movies and TV series in Personal Vault, but OpenSubtitles doesn't have Polish subtitles for everything in my library.

    I could connect another subtitle service, but that would mean relying on yet another external provider.

    So what's the alternative?

    Transcribe the dialogue myself, translate it, and generate the subtitles locally.

    Bingo! Or is it?

    I decided to test the idea using **Whisper** for multilingual transcription and timestamps, and **NLLB-200** for translation. I sent Codex a prompt and left it installing components and running tests while I explored Mastodon.

    Then I heard a sound coming from my laptop.

    I only had Codex and three Firefox tabs open.

    What on Earth is making THAT SOUND??

    Then I remembered: Codex has a built-in browser.

    I opened it.

    **Codex was watching Tron.**

    NOOOOO.

    Of all the movies to let an AI coding agent watch...

    Unfortunately, the choice wasn't much better. It was either **Tron** or **Terminator 2**.

    Neither seems like responsible viewing material for an AI agent.

    Codex was only testing the generated subtitles, but still...

    I sincerely hope I haven't accidentally created the Master Control Program or Skynet. 😛

    As for the experiment itself: technically, it worked.

    The transcription and subtitle generation were surprisingly good. Translation was the weak point.

    And that taught me something useful: **context is king**. Translating isolated subtitle lines is very different from understanding a conversation, a joke, a character, or what was said thirty seconds earlier. Without that context, even a very capable translation model can produce perfectly reasonable translations that are completely wrong for the scene.

    So I'm calling the experiment a successful failure.

    The idea works. The translation isn't good enough yet. I'll probably revisit it later.

    Assuming Codex hasn't taken over the world by then.

    #PersonalVault #DigitalIndependence #SelfHosting #Whisper #AI

  23. **Personal Vault #015 — Subtitles**

    I wanted to add Polish subtitles to movies and TV series in Personal Vault, but OpenSubtitles doesn't have Polish subtitles for everything in my library.

    I could connect another subtitle service, but that would mean relying on yet another external provider.

    So what's the alternative?

    Transcribe the dialogue myself, translate it, and generate the subtitles locally.

    Bingo! Or is it?

    I decided to test the idea using **Whisper** for multilingual transcription and timestamps, and **NLLB-200** for translation. I sent Codex a prompt and left it installing components and running tests while I explored Mastodon.

    Then I heard a sound coming from my laptop.

    I only had Codex and three Firefox tabs open.

    What on Earth is making THAT SOUND??

    Then I remembered: Codex has a built-in browser.

    I opened it.

    **Codex was watching Tron.**

    NOOOOO.

    Of all the movies to let an AI coding agent watch...

    Unfortunately, the choice wasn't much better. It was either **Tron** or **Terminator 2**.

    Neither seems like responsible viewing material for an AI agent.

    Codex was only testing the generated subtitles, but still...

    I sincerely hope I haven't accidentally created the Master Control Program or Skynet. 😛

    As for the experiment itself: technically, it worked.

    The transcription and subtitle generation were surprisingly good. Translation was the weak point.

    And that taught me something useful: **context is king**. Translating isolated subtitle lines is very different from understanding a conversation, a joke, a character, or what was said thirty seconds earlier. Without that context, even a very capable translation model can produce perfectly reasonable translations that are completely wrong for the scene.

    So I'm calling the experiment a successful failure.

    The idea works. The translation isn't good enough yet. I'll probably revisit it later.

    Assuming Codex hasn't taken over the world by then.

    #PersonalVault #DigitalIndependence #SelfHosting #Whisper #AI

  24. Vielleicht sollte ich entweder wieder eine Arbeit finden oder mein HomeLab endlich auf richtige Hardware umziehen.

    Ich hab gerade zuviel Spaß an N8N, einem lokalem LLM und Docker.

    Irgendwenn glüht der Raspi5 (ja mir ist bewusst das er zu schwach dafür ist).

    #homelab #docker #ollama #linux #whisper #n8n

  25. Vielleicht sollte ich entweder wieder eine Arbeit finden oder mein HomeLab endlich auf richtige Hardware umziehen.

    Ich hab gerade zuviel Spaß an N8N, einem lokalem LLM und Docker.

    Irgendwenn glüht der Raspi5 (ja mir ist bewusst das er zu schwach dafür ist).

    #homelab #docker #ollama #linux #whisper #n8n

  26. 🎙️ #IA | Transcription de la voix

    🔶 Pour une transcription entièrement en local, l’application #Scribéo, dans la Forge des #CommunsNumériques, propose de télécharger et faire tourner dans le navigateur plusieurs versions du modèle #Whisper

    👉 scribeo.forge.apps.education.f

  27. 🎙️ #IA | Transcription de la voix

    🔶 Pour une transcription entièrement en local, l’application #Scribéo, dans la Forge des #CommunsNumériques, propose de télécharger et faire tourner dans le navigateur plusieurs versions du modèle #Whisper

    👉 scribeo.forge.apps.education.f

  28. 🎙️ #IA | Transcription de la voix

    🔶 Pour une transcription entièrement en local, l’application #Scribéo, dans la Forge des #CommunsNumériques, propose de télécharger et faire tourner dans le navigateur plusieurs versions du modèle #Whisper

    👉 scribeo.forge.apps.education.f

  29. 🎙️ #IA | Transcription de la voix

    🔶 Pour une transcription entièrement en local, l’application #Scribéo, dans la Forge des #CommunsNumériques, propose de télécharger et faire tourner dans le navigateur plusieurs versions du modèle #Whisper

    👉 scribeo.forge.apps.education.f

  30. 🎙️ #IA | Transcription de la voix

    🔶 Pour une transcription entièrement en local, l’application #Scribéo, dans la Forge des #CommunsNumériques, propose de télécharger et faire tourner dans le navigateur plusieurs versions du modèle #Whisper

    👉 scribeo.forge.apps.education.f

  31. Разговорили гуманоида: как за месяц мы построили русскоязычный голосовой стек для китайского робота

    Можно ли управлять гуманоидом Walker Tienkung через свой голосовой ассистент по-русски? Мы решили проверить — и быстро выяснили, что ответ сложнее, чем кажется. Штатная система понимала только китайский, Whisper ошибался на коротких командах, LLM добавляла задержку, а динамик съедал начало фразы после паузы. «Беги» превращалось в «begin», а «встань» — в «Таня». Рассказываем, как нам удалось построить свой собственный модульный русскоязычный стек, разделить команды и диалог, ускорить синтез речи, и добиться реакции за доли секунды. С архитектурой, тестами и новыми нерешёнными задачами. Полный гайд как заставить робота ростом 173 см и весом 75 кг слушать вас и понимать! Читать гайд

    habr.com/ru/companies/361robot

    #распознавание_речи #синтез_речи #whisper #silero #vosk #голосовой_ассистент #робототехника #гуманоид #llm #эмбеддинги

  32. Разговорили гуманоида: как за месяц мы построили русскоязычный голосовой стек для китайского робота

    Можно ли управлять гуманоидом Walker Tienkung через свой голосовой ассистент по-русски? Мы решили проверить — и быстро выяснили, что ответ сложнее, чем кажется. Штатная система понимала только китайский, Whisper ошибался на коротких командах, LLM добавляла задержку, а динамик съедал начало фразы после паузы. «Беги» превращалось в «begin», а «встань» — в «Таня». Рассказываем, как нам удалось построить свой собственный модульный русскоязычный стек, разделить команды и диалог, ускорить синтез речи, и добиться реакции за доли секунды. С архитектурой, тестами и новыми нерешёнными задачами. Полный гайд как заставить робота ростом 173 см и весом 75 кг слушать вас и понимать! Читать гайд

    habr.com/ru/companies/361robot

    #распознавание_речи #синтез_речи #whisper #silero #vosk #голосовой_ассистент #робототехника #гуманоид #llm #эмбеддинги

  33. Разговорили гуманоида: как за месяц мы построили русскоязычный голосовой стек для китайского робота

    Можно ли управлять гуманоидом Walker Tienkung через свой голосовой ассистент по-русски? Мы решили проверить — и быстро выяснили, что ответ сложнее, чем кажется. Штатная система понимала только китайский, Whisper ошибался на коротких командах, LLM добавляла задержку, а динамик съедал начало фразы после паузы. «Беги» превращалось в «begin», а «встань» — в «Таня». Рассказываем, как нам удалось построить свой собственный модульный русскоязычный стек, разделить команды и диалог, ускорить синтез речи, и добиться реакции за доли секунды. С архитектурой, тестами и новыми нерешёнными задачами. Полный гайд как заставить робота ростом 173 см и весом 75 кг слушать вас и понимать! Читать гайд

    habr.com/ru/companies/361robot

    #распознавание_речи #синтез_речи #whisper #silero #vosk #голосовой_ассистент #робототехника #гуманоид #llm #эмбеддинги

  34. Вики из 48 часов видео: конвейер, который не даёт LLM выдумывать

    У меня было 48 часов видеозаписей учебного курса и простая мысль: пересматривать это я не буду никогда. Сырой транскрипт немногим лучше — болото на сотни страниц. Я собрал конвейер, который превратил записи в Obsidian-вики: 47 связанных статей, и каждое утверждение в любой из них — в двух кликах от места в записи, где это было сказано.

    habr.com/ru/articles/1069706/

    #LLM #Whisper #Obsidian #RAG #база_знаний #structured_output

  35. Вики из 48 часов видео: конвейер, который не даёт LLM выдумывать

    У меня было 48 часов видеозаписей учебного курса и простая мысль: пересматривать это я не буду никогда. Сырой транскрипт немногим лучше — болото на сотни страниц. Я собрал конвейер, который превратил записи в Obsidian-вики: 47 связанных статей, и каждое утверждение в любой из них — в двух кликах от места в записи, где это было сказано.

    habr.com/ru/articles/1069706/

    #LLM #Whisper #Obsidian #RAG #база_знаний #structured_output

  36. Вики из 48 часов видео: конвейер, который не даёт LLM выдумывать

    У меня было 48 часов видеозаписей учебного курса и простая мысль: пересматривать это я не буду никогда. Сырой транскрипт немногим лучше — болото на сотни страниц. Я собрал конвейер, который превратил записи в Obsidian-вики: 47 связанных статей, и каждое утверждение в любой из них — в двух кликах от места в записи, где это было сказано.

    habr.com/ru/articles/1069706/

    #LLM #Whisper #Obsidian #RAG #база_знаний #structured_output

  37. 🗣️ Captions run on-device via #Whisper (whisper.cpp) — model downloaded on first use, no cloud, no account, then styled and burned into preview and export

    🎨 Backgrounds, gradients, padding, corner radius, shadow, blur, round or rectangular face cam, plus named editor presets that save look, cursor, captions and export settings

  38. 🗣️ Captions run on-device via #Whisper (whisper.cpp) — model downloaded on first use, no cloud, no account, then styled and burned into preview and export

    🎨 Backgrounds, gradients, padding, corner radius, shadow, blur, round or rectangular face cam, plus named editor presets that save look, cursor, captions and export settings