home.social

#whisper — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #whisper, aggregated by home.social.

  1. Étude du consortium ARIANE sur la transcription via IA générative avec le LLM Whisper : « Whisper : transcription automatique de sources sonores » hal.science/hal-05706209 #Whisper #LLM #ESR

  2. LLM 改变了很多,比如 Agent 概念上半年爆火后,大厂们都争前恐后推出 CLI 产品,知乎,微信,网易云等等,这在几年前是完全无法想象的,那时想让 Linux 用上微信都很难,需要折腾一大堆 Wine 的组件,有时还会面临封号的风险。反观现在拥抱全平台的吃相,让人忍俊不禁。

    但真正改变的,其实很有限,因为爆火的原因不是厂商良心,拥抱开源,而是发现这个市场有的赚,本质还是逐利的,Agent 让厂商发现开放生态有的赚。工具功能很全,很方便,但我却更加担心了,工具会一直免费下去吗?我猜不会,等到官方的生态建立起来,我想就是回血的时候了。

    值得高兴的是,这些工具还很烂,如果真的要用,还太早,比如 npmjs.com/package/@music163/nc

    我无法相信官方的 TUI 居然吹能播放一半不到的歌,真是充满恶意的设计

    #whisper #netease #cli #llm

  3. LLM 改变了很多,比如 Agent 概念上半年爆火后,大厂们都争前恐后推出 CLI 产品,知乎,微信,网易云等等,这在几年前是完全无法想象的,那时想让 Linux 用上微信都很难,需要折腾一大堆 Wine 的组件,有时还会面临封号的风险。反观现在拥抱全平台的吃相,让人忍俊不禁。

    但真正改变的,其实很有限,因为爆火的原因不是厂商良心,拥抱开源,而是发现这个市场有的赚,本质还是逐利的,Agent 让厂商发现开放生态有的赚。工具功能很全,很方便,但我却更加担心了,工具会一直免费下去吗?我猜不会,等到官方的生态建立起来,我想就是回血的时候了。

    值得高兴的是,这些工具还很烂,如果真的要用,还太早,比如 npmjs.com/package/@music163/nc

    我无法相信官方的 TUI 居然吹能播放一半不到的歌,真是充满恶意的设计

    #whisper #netease #cli #llm

  4. LLM 改变了很多,比如 Agent 概念上半年爆火后,大厂们都争前恐后推出 CLI 产品,知乎,微信,网易云等等,这在几年前是完全无法想象的,那时想让 Linux 用上微信都很难,需要折腾一大堆 Wine 的组件,有时还会面临封号的风险。反观现在拥抱全平台的吃相,让人忍俊不禁。

    但真正改变的,其实很有限,因为爆火的原因不是厂商良心,拥抱开源,而是发现这个市场有的赚,本质还是逐利的,Agent 让厂商发现开放生态有的赚。工具功能很全,很方便,但我却更加担心了,工具会一直免费下去吗?我猜不会,等到官方的生态建立起来,我想就是回血的时候了。

    值得高兴的是,这些工具还很烂,如果真的要用,还太早,比如 npmjs.com/package/@music163/nc

    我无法相信官方的 TUI 居然吹能播放一半不到的歌,真是充满恶意的设计

    #whisper #netease #cli #llm

  5. LLM 改变了很多,比如 Agent 概念上半年爆火后,大厂们都争前恐后推出 CLI 产品,知乎,微信,网易云等等,这在几年前是完全无法想象的,那时想让 Linux 用上微信都很难,需要折腾一大堆 Wine 的组件,有时还会面临封号的风险。反观现在拥抱全平台的吃相,让人忍俊不禁。

    但真正改变的,其实很有限,因为爆火的原因不是厂商良心,拥抱开源,而是发现这个市场有的赚,本质还是逐利的,Agent 让厂商发现开放生态有的赚。工具功能很全,很方便,但我却更加担心了,工具会一直免费下去吗?我猜不会,等到官方的生态建立起来,我想就是回血的时候了。

    值得高兴的是,这些工具还很烂,如果真的要用,还太早,比如 npmjs.com/package/@music163/nc

    我无法相信官方的 TUI 居然吹能播放一半不到的歌,真是充满恶意的设计

    #whisper #netease #cli #llm

  6. LLM 改变了很多,比如 Agent 概念上半年爆火后,大厂们都争前恐后推出 CLI 产品,知乎,微信,网易云等等,这在几年前是完全无法想象的,那时想让 Linux 用上微信都很难,需要折腾一大堆 Wine 的组件,有时还会面临封号的风险。反观现在拥抱全平台的吃相,让人忍俊不禁。

    但真正改变的,其实很有限,因为爆火的原因不是厂商良心,拥抱开源,而是发现这个市场有的赚,本质还是逐利的,Agent 让厂商发现开放生态有的赚。工具功能很全,很方便,但我却更加担心了,工具会一直免费下去吗?我猜不会,等到官方的生态建立起来,我想就是回血的时候了。

    值得高兴的是,这些工具还很烂,如果真的要用,还太早,比如 npmjs.com/package/@music163/nc

    我无法相信官方的 TUI 居然吹能播放一半不到的歌,真是充满恶意的设计

    #whisper #netease #cli #llm

  7. Was für eine Sch*** !

    Ständig zerlegt sich die Whisper-Integration in KDEnlive. Kaum hat man die wieder geflickt, zerschießt das nächste Update sie wieder.

    Mal eben Untertitel erzeugen, wird so zur Qual.

    Ich prangere das an!

    Danke für Ihre Aufmerksamkeit für dieses Thema!

    #kdenlive #whisper #captions

  8. Was für eine Sch*** !

    Ständig zerlegt sich die Whisper-Integration in KDEnlive. Kaum hat man die wieder geflickt, zerschießt das nächste Update sie wieder.

    Mal eben Untertitel erzeugen, wird so zur Qual.

    Ich prangere das an!

    Danke für Ihre Aufmerksamkeit für dieses Thema!

    #kdenlive #whisper #captions

  9. Was für eine Sch*** !

    Ständig zerlegt sich die Whisper-Integration in KDEnlive. Kaum hat man die wieder geflickt, zerschießt das nächste Update sie wieder.

    Mal eben Untertitel erzeugen, wird so zur Qual.

    Ich prangere das an!

    Danke für Ihre Aufmerksamkeit für dieses Thema!

    #kdenlive #whisper #captions

  10. **Personal Vault #015 — Subtitles**

    I wanted to add Polish subtitles to movies and TV series in Personal Vault, but OpenSubtitles doesn't have Polish subtitles for everything in my library.

    I could connect another subtitle service, but that would mean relying on yet another external provider.

    So what's the alternative?

    Transcribe the dialogue myself, translate it, and generate the subtitles locally.

    Bingo! Or is it?

    I decided to test the idea using **Whisper** for multilingual transcription and timestamps, and **NLLB-200** for translation. I sent Codex a prompt and left it installing components and running tests while I explored Mastodon.

    Then I heard a sound coming from my laptop.

    I only had Codex and three Firefox tabs open.

    What on Earth is making THAT SOUND??

    Then I remembered: Codex has a built-in browser.

    I opened it.

    **Codex was watching Tron.**

    NOOOOO.

    Of all the movies to let an AI coding agent watch...

    Unfortunately, the choice wasn't much better. It was either **Tron** or **Terminator 2**.

    Neither seems like responsible viewing material for an AI agent.

    Codex was only testing the generated subtitles, but still...

    I sincerely hope I haven't accidentally created the Master Control Program or Skynet. 😛

    As for the experiment itself: technically, it worked.

    The transcription and subtitle generation were surprisingly good. Translation was the weak point.

    And that taught me something useful: **context is king**. Translating isolated subtitle lines is very different from understanding a conversation, a joke, a character, or what was said thirty seconds earlier. Without that context, even a very capable translation model can produce perfectly reasonable translations that are completely wrong for the scene.

    So I'm calling the experiment a successful failure.

    The idea works. The translation isn't good enough yet. I'll probably revisit it later.

    Assuming Codex hasn't taken over the world by then.

    #PersonalVault #DigitalIndependence #SelfHosting #Whisper #AI

  11. **Personal Vault #015 — Subtitles**

    I wanted to add Polish subtitles to movies and TV series in Personal Vault, but OpenSubtitles doesn't have Polish subtitles for everything in my library.

    I could connect another subtitle service, but that would mean relying on yet another external provider.

    So what's the alternative?

    Transcribe the dialogue myself, translate it, and generate the subtitles locally.

    Bingo! Or is it?

    I decided to test the idea using **Whisper** for multilingual transcription and timestamps, and **NLLB-200** for translation. I sent Codex a prompt and left it installing components and running tests while I explored Mastodon.

    Then I heard a sound coming from my laptop.

    I only had Codex and three Firefox tabs open.

    What on Earth is making THAT SOUND??

    Then I remembered: Codex has a built-in browser.

    I opened it.

    **Codex was watching Tron.**

    NOOOOO.

    Of all the movies to let an AI coding agent watch...

    Unfortunately, the choice wasn't much better. It was either **Tron** or **Terminator 2**.

    Neither seems like responsible viewing material for an AI agent.

    Codex was only testing the generated subtitles, but still...

    I sincerely hope I haven't accidentally created the Master Control Program or Skynet. 😛

    As for the experiment itself: technically, it worked.

    The transcription and subtitle generation were surprisingly good. Translation was the weak point.

    And that taught me something useful: **context is king**. Translating isolated subtitle lines is very different from understanding a conversation, a joke, a character, or what was said thirty seconds earlier. Without that context, even a very capable translation model can produce perfectly reasonable translations that are completely wrong for the scene.

    So I'm calling the experiment a successful failure.

    The idea works. The translation isn't good enough yet. I'll probably revisit it later.

    Assuming Codex hasn't taken over the world by then.

    #PersonalVault #DigitalIndependence #SelfHosting #Whisper #AI

  12. **Personal Vault #015 — Subtitles**

    I wanted to add Polish subtitles to movies and TV series in Personal Vault, but OpenSubtitles doesn't have Polish subtitles for everything in my library.

    I could connect another subtitle service, but that would mean relying on yet another external provider.

    So what's the alternative?

    Transcribe the dialogue myself, translate it, and generate the subtitles locally.

    Bingo! Or is it?

    I decided to test the idea using **Whisper** for multilingual transcription and timestamps, and **NLLB-200** for translation. I sent Codex a prompt and left it installing components and running tests while I explored Mastodon.

    Then I heard a sound coming from my laptop.

    I only had Codex and three Firefox tabs open.

    What on Earth is making THAT SOUND??

    Then I remembered: Codex has a built-in browser.

    I opened it.

    **Codex was watching Tron.**

    NOOOOO.

    Of all the movies to let an AI coding agent watch...

    Unfortunately, the choice wasn't much better. It was either **Tron** or **Terminator 2**.

    Neither seems like responsible viewing material for an AI agent.

    Codex was only testing the generated subtitles, but still...

    I sincerely hope I haven't accidentally created the Master Control Program or Skynet. 😛

    As for the experiment itself: technically, it worked.

    The transcription and subtitle generation were surprisingly good. Translation was the weak point.

    And that taught me something useful: **context is king**. Translating isolated subtitle lines is very different from understanding a conversation, a joke, a character, or what was said thirty seconds earlier. Without that context, even a very capable translation model can produce perfectly reasonable translations that are completely wrong for the scene.

    So I'm calling the experiment a successful failure.

    The idea works. The translation isn't good enough yet. I'll probably revisit it later.

    Assuming Codex hasn't taken over the world by then.

    #PersonalVault #DigitalIndependence #SelfHosting #Whisper #AI

  13. **Personal Vault #015 — Subtitles**

    I wanted to add Polish subtitles to movies and TV series in Personal Vault, but OpenSubtitles doesn't have Polish subtitles for everything in my library.

    I could connect another subtitle service, but that would mean relying on yet another external provider.

    So what's the alternative?

    Transcribe the dialogue myself, translate it, and generate the subtitles locally.

    Bingo! Or is it?

    I decided to test the idea using **Whisper** for multilingual transcription and timestamps, and **NLLB-200** for translation. I sent Codex a prompt and left it installing components and running tests while I explored Mastodon.

    Then I heard a sound coming from my laptop.

    I only had Codex and three Firefox tabs open.

    What on Earth is making THAT SOUND??

    Then I remembered: Codex has a built-in browser.

    I opened it.

    **Codex was watching Tron.**

    NOOOOO.

    Of all the movies to let an AI coding agent watch...

    Unfortunately, the choice wasn't much better. It was either **Tron** or **Terminator 2**.

    Neither seems like responsible viewing material for an AI agent.

    Codex was only testing the generated subtitles, but still...

    I sincerely hope I haven't accidentally created the Master Control Program or Skynet. 😛

    As for the experiment itself: technically, it worked.

    The transcription and subtitle generation were surprisingly good. Translation was the weak point.

    And that taught me something useful: **context is king**. Translating isolated subtitle lines is very different from understanding a conversation, a joke, a character, or what was said thirty seconds earlier. Without that context, even a very capable translation model can produce perfectly reasonable translations that are completely wrong for the scene.

    So I'm calling the experiment a successful failure.

    The idea works. The translation isn't good enough yet. I'll probably revisit it later.

    Assuming Codex hasn't taken over the world by then.

    #PersonalVault #DigitalIndependence #SelfHosting #Whisper #AI

  14. **Personal Vault #015 — Subtitles**

    I wanted to add Polish subtitles to movies and TV series in Personal Vault, but OpenSubtitles doesn't have Polish subtitles for everything in my library.

    I could connect another subtitle service, but that would mean relying on yet another external provider.

    So what's the alternative?

    Transcribe the dialogue myself, translate it, and generate the subtitles locally.

    Bingo! Or is it?

    I decided to test the idea using **Whisper** for multilingual transcription and timestamps, and **NLLB-200** for translation. I sent Codex a prompt and left it installing components and running tests while I explored Mastodon.

    Then I heard a sound coming from my laptop.

    I only had Codex and three Firefox tabs open.

    What on Earth is making THAT SOUND??

    Then I remembered: Codex has a built-in browser.

    I opened it.

    **Codex was watching Tron.**

    NOOOOO.

    Of all the movies to let an AI coding agent watch...

    Unfortunately, the choice wasn't much better. It was either **Tron** or **Terminator 2**.

    Neither seems like responsible viewing material for an AI agent.

    Codex was only testing the generated subtitles, but still...

    I sincerely hope I haven't accidentally created the Master Control Program or Skynet. 😛

    As for the experiment itself: technically, it worked.

    The transcription and subtitle generation were surprisingly good. Translation was the weak point.

    And that taught me something useful: **context is king**. Translating isolated subtitle lines is very different from understanding a conversation, a joke, a character, or what was said thirty seconds earlier. Without that context, even a very capable translation model can produce perfectly reasonable translations that are completely wrong for the scene.

    So I'm calling the experiment a successful failure.

    The idea works. The translation isn't good enough yet. I'll probably revisit it later.

    Assuming Codex hasn't taken over the world by then.

    #PersonalVault #DigitalIndependence #SelfHosting #Whisper #AI

  15. 56 минут созвона → текст за 5 минут на M4 без OBS и облака

    У меня от трёх до пяти созвонов в день, почти все в браузере: Meet, Яндекс Телемост, ktalk, реже всего Zoom. Детали встреч мне нужны в тексте, иначе договорённости расползаются. Писал я звонки через OBS с захватом экрана, и на самих звонках началось веселье: звук временами жёстко квакал, всё чуть-чуть подвисало. Причину до конца не выяснял, моё предположение - процессор зажирался захватом и перебивал сам разговор; ktalk и браузер тут ни при чём. А экран, как потом дошло, мне вообще был не нужен. Дальше вторая серия. Расшифровку я гонял своими Python-скриптами: конвертация → транскрибация → диаризация (разметка, кто когда говорил). На встрече с пятью участниками диаризация расставляла голоса крайне плохо, половину фраз всё равно восстанавливал по памяти. Пошёл искать вариант, чтобы запись не мешала звонку, результат был приличный и всё крутилось локально: не хотелось ни отдавать записи в облако, ни платить подписку за минуты. Нашёл платный Audio Hijack , он подкупил лёгкостью и простотой настройки. Потом узнал, что на M-чипе быстрее всего у меня отрабатывает mlx-whisper на MLX : веса large-v3-turbo с Hugging Face, инференс на GPU Mac. Это архитектура Whisper, но рантайм — пакет mlx-whisper и бинарь mlx_whisper , не pip install openai-whisper и не репозиторий openai/whisper . Так собрался стек: Hijack пишет один mp3, mlx_whisper делает VTT, pyannote раскладывает по SPEAKER_XX. На 56-минутном звонке транскрипция заняла около пяти минут. Платный только Hijack — разовая лицензия, без подписки (цену смотрите на сайте Rogue Amoeba). Скрипты выложил в mac-call-transcribe под MIT. Ниже сборка, замеры с двух реальных звонков и грабли; повторяется на любом Mac с M-чипом.

    habr.com/ru/articles/1062068/

    #whisper #mlx #pyannote #speaker_diarization #расшифровка_звонков #python #apple_silicon #транскрибация_звонков

  16. Tester Nasjonalbibliotekets Whisper språkmodell til automatisk teksting i PeerTube. Det fungerer bedre enn standardmodellen til Whisper... helt til det ikke gjør det.

    #PeerTube #Whisper #Nasjonalbiblioteket #KI

  17. Fairly valid use of #AI, near realtime voice to text #estonian -> English translation tool. I have spent about 2 hours gluing together OpenAI #whisper and Google Gemma 4 e2b and WebRTC VAD (voice activity detection) in Python. Both whisper and Gemma are running on GPU together utilizing about 50% of my Radeon.

  18. Голосовой КПТ-дневник с распознаванием речи на устройстве: Flutter и on-device Whisper

    Эта статья про то, как я сделал голосовой дневник мыслей для когнитивно-поведенческой терапии, почему распознавание речи у меня крутится прямо на телефоне, и какие на этом пути были технические развилки. Кода почти не будет, будет архитектура и обоснование решений. Я сам прошёл через тревожные расстройства, панические атаки и несколько депрессивных периодов. Из всего, что мне помогало, переломной стала КПТ, и у неё есть домашняя часть, дневник мыслей, который нужно вести между сессиями. Вести его текстом в момент тревоги у меня не получалось годами, и в какой-то момент я понял, что хочу диктовать его голосом. Так появился проект, который я тут и разбираю.

    habr.com/ru/articles/1043432/

    #Flutter #Whisper #whispercpp #ondevice #распознавание_речи #Dart #КПТ #мобильная_разработка

  19. Как я решил проблему русской диктовки для ИИ

    По мере погружения в ИИ и вайб‑кодинг, я столкнулся с одним неудобным моментом — отсутствием возможности диктовать на русском языке в некоторых программах. И если OpenAI в своем приложении позаботились об этом, то в Anthropic такой возможности на тот момент просто не оказалось. А мне уже так понравилось, откинувшись на спинку кресла с чашкой чая, надиктовывать промпты без клавиатуры. Но я быстро нашел выход, хоть и костыльный — просто диктовать свой текст в окошке GPT, потом копировать его и вставлять в Claude. Вроде несложно, но и удобным этот метод я бы не назвал. И я задумался, как этот процесс оптимизировать. И какая же идея могла прийти в голову в 3 часа ночи человеку, который полжизни занимается программированием? Ну конечно же — разработать свое приложение. Посоветовавшись с Claude и GPT, я набросал небольшой план и приступил к разработке. Поскольку я работаю на macOS, то для начала не стал заморачиваться с мультиплатформенностью и решил делать все на Swift.

    habr.com/ru/articles/1039248/

    #AI #OpenAI #Claude #Whisper #speechtotext #диктовка #voice_input #Apple_Silicon

  20. OpenAI, çağrışım dilini başarıyla sesli yapay zeka uygulamalarıyla birleştiriyor: GPT‑Realtime‑2, Translate ve Whisper elemanlarıyla gerçek zamanlı konuşma, çeviri ve transkripsiyon sistemini tek bir akışta sunuyor.

    🚩 #OpenAI #SesAI #GPTRealtime #Whisper #Translate

  21. SLAY-ASR, или как я перестал волноваться и полюбил тренировать модели

    Как добавить аудио-модальность в LLMку максимально экономно? Рассказываю про серию попыток добиться совместимости эмбеддингов разной природы Погрузиться

    habr.com/ru/articles/1009614/

    #representation_learning #multimodality #multimodal_llm #machine_learning #audiomodality #regularization #contrastive_learning #whisper #gemma3

  22. Free tools for creativity!

    , , , , by Meta, , & more: These , -powered tools make image editing, audio work, transcription, and Large Language Models (#LLMs) exploration fun and easy!

    Learn more about these top picks by mentor @morrolinux: lpi.org/zhya

    @LPI , , , , , @openai

  23. 🔊 Whisper has a serious challenger: Moshi STT

    Developed by the French research lab Kyutai, Moshi STT is a new open-source speech recognition system that’s blazingly fast, highly accurate, and optimized for Apple Silicon and CUDA — all designed with real-time performance in mind.

    scalastic.io/en/moshi-stt-vs-w

    #SpeechToText #Kyutai #Whisper #OpenSource #AI #macOS #Moshi #STT #AppleSilicon #Rust #CUDA

  24. @SuperOscar isn't there a Johnny Cash song about that? Don't take rpi's to town boy leave rpi's at home? :)

    It's not the first missing thing I have found, but not having venv, which is the bane of Python anyway, seemed kind of, well, extreme. No pip, no wheels, that's crazy.

    I installed #OpenAI #Whisper on a #Ubuntu x86; the hope is to provide both #Nextcloud and #OpenVoiceOS from the same in-house service.