home.social

#voicesynthesis — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #voicesynthesis, aggregated by home.social.

  1. StepFun Unveils "StepAudio 2.5 Realtime," Promising End-to-End Voice Synthesis

    StepFun launches StepAudio 2.5 Realtime, an end-to-end voice model for roleplaying. Learn about its features and how developers can use GELab-Zero-4B-preview.

    #StepAudio2.5, #VoiceSynthesis, #AIforGaming, #RealtimeAI, #GELabZero

    newsletter.tf/stepfun-stepaudi

  2. StepFun's new StepAudio 2.5 Realtime model can generate speech instantly, making roleplaying games more immersive. It uses the GELab-Zero-4B-preview model.

    #StepAudio2.5, #VoiceSynthesis, #AIforGaming, #RealtimeAI, #GELabZero
    newsletter.tf/stepfun-stepaudi

  3. Success!! I've completed a first reimplementation of Dr. Sbaitso voice synthesis in Godot!

    Try it out at Itchio: eibriel.itch.io/scp-079-voice

    Ive configure the accessibility settings so the app properly describes the input fields and buttons, but for some reason is not working for me (tested on Linux with Orca). Let me know if it works for you!

    #TTS #DrSbaitso #VoiceSynthesis #TextToSpeech #079 #SCP079 #SCP #Godot

  4. Success!! I've completed a first reimplementation of Dr. Sbaitso voice synthesis in Godot!

    Try it out at Itchio: eibriel.itch.io/scp-079-voice

    Ive configure the accessibility settings so the app properly describes the input fields and buttons, but for some reason is not working for me (tested on Linux with Orca). Let me know if it works for you!

    #TTS #DrSbaitso #VoiceSynthesis #TextToSpeech #079 #SCP079 #SCP #Godot

  5. Success!! I've completed a first reimplementation of Dr. Sbaitso voice synthesis in Godot!

    Try it out at Itchio: eibriel.itch.io/scp-079-voice

    Ive configure the accessibility settings so the app properly describes the input fields and buttons, but for some reason is not working for me (tested on Linux with Orca). Let me know if it works for you!

    #TTS #DrSbaitso #VoiceSynthesis #TextToSpeech #079 #SCP079 #SCP #Godot

  6. Success!! I've completed a first reimplementation of Dr. Sbaitso voice synthesis in Godot!

    Try it out at Itchio: eibriel.itch.io/scp-079-voice

    Ive configure the accessibility settings so the app properly describes the input fields and buttons, but for some reason is not working for me (tested on Linux with Orca). Let me know if it works for you!

    #TTS #DrSbaitso #VoiceSynthesis #TextToSpeech #079 #SCP079 #SCP #Godot

  7. Some improvements to the concatenation, prosody is still missing.

    Here is a well known phrase by SCP 079.

    The audio contains the same phrase first performed by Dr. Sbaitso TTS and the by Godot reimplementation.

    #TTS #DrSbaitso #VoiceSynthesis #TextToSpeech #079 #SCP079 #SCP #Godot

  8. Some improvements to the concatenation, prosody is still missing.

    Here is a well known phrase by SCP 079.

    The audio contains the same phrase first performed by Dr. Sbaitso TTS and the by Godot reimplementation.

    #TTS #DrSbaitso #VoiceSynthesis #TextToSpeech #079 #SCP079 #SCP #Godot

  9. Some improvements to the concatenation, prosody is still missing.

    Here is a well known phrase by SCP 079.

    The audio contains the same phrase first performed by Dr. Sbaitso TTS and the by Godot reimplementation.

    #TTS #DrSbaitso #VoiceSynthesis #TextToSpeech #079 #SCP079 #SCP #Godot

  10. Some improvements to the concatenation, prosody is still missing.

    Here is a well known phrase by SCP 079.

    The audio contains the same phrase first performed by Dr. Sbaitso TTS and the by Godot reimplementation.

    #TTS #DrSbaitso #VoiceSynthesis #TextToSpeech #079 #SCP079 #SCP #Godot

  11. Some improvements to the concatenation, prosody is still missing.

    Here is a well known phrase by SCP 079.

    The audio contains the same phrase first performed by Dr. Sbaitso TTS and the by Godot reimplementation.

    #TTS #DrSbaitso #VoiceSynthesis #TextToSpeech #079 #SCP079 #SCP #Godot

  12. Dr. Sbaitso compared to my reimplementation in Godot (Sbaitso first) :computer_explorer: :pc_color:

    Implemented: basic waveform concatenation
    Missing: Interpolation, pitch control, prosody, text to phonemes

    Im very happy with the progress, will be great to be able to run the voice without needing emulation.

    #TTS #DrSbaitso #VoiceSynthesis #TextToSpeech #079 #SCP079

  13. Dr. Sbaitso compared to my reimplementation in Godot (Sbaitso first) :computer_explorer: :pc_color:

    Implemented: basic waveform concatenation
    Missing: Interpolation, pitch control, prosody, text to phonemes

    Im very happy with the progress, will be great to be able to run the voice without needing emulation.

    #TTS #DrSbaitso #VoiceSynthesis #TextToSpeech #079 #SCP079

  14. Dr. Sbaitso compared to my reimplementation in Godot (Sbaitso first) :computer_explorer: :pc_color:

    Implemented: basic waveform concatenation
    Missing: Interpolation, pitch control, prosody, text to phonemes

    Im very happy with the progress, will be great to be able to run the voice without needing emulation.

    #TTS #DrSbaitso #VoiceSynthesis #TextToSpeech #079 #SCP079

  15. Dr. Sbaitso compared to my reimplementation in Godot (Sbaitso first) :computer_explorer: :pc_color:

    Implemented: basic waveform concatenation
    Missing: Interpolation, pitch control, prosody, text to phonemes

    Im very happy with the progress, will be great to be able to run the voice without needing emulation.

    #TTS #DrSbaitso #VoiceSynthesis #TextToSpeech #079 #SCP079

  16. Dr. Sbaitso compared to my reimplementation in Godot (Sbaitso first) :computer_explorer: :pc_color:

    Implemented: basic waveform concatenation
    Missing: Interpolation, pitch control, prosody, text to phonemes

    Im very happy with the progress, will be great to be able to run the voice without needing emulation.

    #TTS #DrSbaitso #VoiceSynthesis #TextToSpeech #079 #SCP079

  17. What I've learned so far while reverse engineering Dr Sbaitso's voice:
    - Reverse engineering is hard

    Also, the voice was made by very clever people. It's optimized to sound as good as possible, while consuming very few resources.

    Progress after 5 days: 10%

    #TTS #DrSbaitso #VoiceSynthesis #TextToSpeech

  18. What I've learned so far while reverse engineering Dr Sbaitso's voice:
    - Reverse engineering is hard

    Also, the voice was made by very clever people. It's optimized to sound as good as possible, while consuming very few resources.

    Progress after 5 days: 10%

    #TTS #DrSbaitso #VoiceSynthesis #TextToSpeech

  19. What I've learned so far while reverse engineering Dr Sbaitso's voice:
    - Reverse engineering is hard

    Also, the voice was made by very clever people. It's optimized to sound as good as possible, while consuming very few resources.

    Progress after 5 days: 10%

    #TTS #DrSbaitso #VoiceSynthesis #TextToSpeech

  20. What I've learned so far while reverse engineering Dr Sbaitso's voice:
    - Reverse engineering is hard

    Also, the voice was made by very clever people. It's optimized to sound as good as possible, while consuming very few resources.

    Progress after 5 days: 10%

    #TTS #DrSbaitso #VoiceSynthesis #TextToSpeech

  21. What I've learned so far while reverse engineering Dr Sbaitso's voice:
    - Reverse engineering is hard

    Also, the voice was made by very clever people. It's optimized to sound as good as possible, while consuming very few resources.

    Progress after 5 days: 10%

    #TTS #DrSbaitso #VoiceSynthesis #TextToSpeech

  22. Primera prueba del sintetizador de voz por difonos hecho en Godot.

    Tiene un millón de problemas, grabé la voz así nomás.

    #Godot #TTS #VoiceSynthesis #SintesisDeVoz #Capusotto

  23. Primera prueba del sintetizador de voz por difonos hecho en Godot.

    Tiene un millón de problemas, grabé la voz así nomás.

    #Godot #TTS #VoiceSynthesis #SintesisDeVoz #Capusotto

  24. Primera prueba del sintetizador de voz por difonos hecho en Godot.

    Tiene un millón de problemas, grabé la voz así nomás.

    #Godot #TTS #VoiceSynthesis #SintesisDeVoz #Capusotto

  25. Primera prueba del sintetizador de voz por difonos hecho en Godot.

    Tiene un millón de problemas, grabé la voz así nomás.

    #Godot #TTS #VoiceSynthesis #SintesisDeVoz #Capusotto

  26. Germany gets a new AI call assistant from Deutsche Telekom that works straight from the cellular network—no app required. Powered by ElevenLabs’ voice synthesis, it can translate languages on the fly. Unveiled at Mobile World Congress, it shows how open‑source‑friendly AI can reshape everyday calls. Curious how it works? #MagentaAI #DeutscheTelekom #ElevenLabs #VoiceSynthesis

    🔗 aidailypost.com/news/magenta-a

  27. Germany gets a new AI call assistant from Deutsche Telekom that works straight from the cellular network—no app required. Powered by ElevenLabs’ voice synthesis, it can translate languages on the fly. Unveiled at Mobile World Congress, it shows how open‑source‑friendly AI can reshape everyday calls. Curious how it works? #MagentaAI #DeutscheTelekom #ElevenLabs #VoiceSynthesis

    🔗 aidailypost.com/news/magenta-a

  28. Germany gets a new AI call assistant from Deutsche Telekom that works straight from the cellular network—no app required. Powered by ElevenLabs’ voice synthesis, it can translate languages on the fly. Unveiled at Mobile World Congress, it shows how open‑source‑friendly AI can reshape everyday calls. Curious how it works? #MagentaAI #DeutscheTelekom #ElevenLabs #VoiceSynthesis

    🔗 aidailypost.com/news/magenta-a

  29. Qwen3-TTS ra mắt với độ trễ siêu thấp chỉ 97ms, hỗ trợ nhân bản giọng nói và API tương thích OpenAI. Công nghệ tổng hợp giọng nói tiên tiến, lý tưởng cho ứng dụng thời gian thực. #Qwen3TTS #VoiceSynthesis #AI #TextToSpeech #TríTuệNhânTạo #TTS #OpenAI

    reddit.com/r/ollama/comments/1

  30. ElevenLabs appoints Karthik Rajaram as India Country Head to accelerate AI voice growth. His leadership will boost multilingual audio, voice synthesis and conversational AI for creators and brands across the Indian market. Discover how this move could reshape digital content creation. #AIvoice #VoiceSynthesis #ElevenLabs #MultilingualAudio

    🔗 aidailypost.com/news/elevenlab

  31. Mệt mỏi với phí “SaaS Tax” khi tạo giọng nói? Nếu có GPU NVIDIA, bạn có thể chuyển sang **run mô hình VITS/Transformer hoàn toàn offline** – không giới hạn ký tự, không phí hàng tháng, dữ liệu riêng tư và độ trễ bằng 0. Hãy tận dụng toàn bộ sức mạnh phần cứng của mình! #AI #VoiceSynthesis #NVIDIA #SaaS #TechVietnam #AIVietnam #OfflineAI #Công_nghệ #GPU #SaaSTax

    reddit.com/r/selfhosted/commen

  32. **Sonya TTS - Mô hình chuyển văn bản thành giọng nói nhanh, biểu cảm và chạy mọi nơi!**

    Sonya TTS là mô hình TTS tiếng Anh nhỏ gọn, sử dụng VITS, mang đến giọng nói tự nhiên, giàu cảm xúc và nhịp điệu. Điểm nổi bật:
    ✅ Tốc độ xử lý cực nhanh, latency thấp
    ✅ Chế độ Audiobook cho văn bản dài
    ✅ Điều chỉnh cảm xúc, tốc độ và nhịp điệu
    ✅ Chạy trên mọi thiết bị: GPU, CPU, edge devices

    Dù chưa hoàn hảo nhưng đã rất ấn tượng với khả năng biểu cảm và hiệu suất!

    #TTS #AI #VoiceSynthesis #TechNews #CôngN

  33. Designing with the KT142C voice chip? Remember, its user-accessible memory is 320KB. If your audio files are larger, you'll need a different solution.

    The KT142F chip allows for an external Flash memory, letting you scale your voice storage capacity as needed. The trade-off is a slight increase in component cost and board space, but you gain total flexibility.

    linkedin.com/pulse/built-in-32

    #OpenHardware #Embedded #Electronics #VoiceSynthesis #PCBDesign #KT142 #Maker

  34. Designing with the KT142C voice chip? Remember, its user-accessible memory is 320KB. If your audio files are larger, you'll need a different solution.

    The KT142F chip allows for an external Flash memory, letting you scale your voice storage capacity as needed. The trade-off is a slight increase in component cost and board space, but you gain total flexibility.

    linkedin.com/pulse/built-in-32

    #OpenHardware #Embedded #Electronics #VoiceSynthesis #PCBDesign #KT142 #Maker

  35. Designing with the KT142C voice chip? Remember, its user-accessible memory is 320KB. If your audio files are larger, you'll need a different solution.

    The KT142F chip allows for an external Flash memory, letting you scale your voice storage capacity as needed. The trade-off is a slight increase in component cost and board space, but you gain total flexibility.

    linkedin.com/pulse/built-in-32

    #OpenHardware #Embedded #Electronics #VoiceSynthesis #PCBDesign #KT142 #Maker

  36. Designing with the KT142C voice chip? Remember, its user-accessible memory is 320KB. If your audio files are larger, you'll need a different solution.

    The KT142F chip allows for an external Flash memory, letting you scale your voice storage capacity as needed. The trade-off is a slight increase in component cost and board space, but you gain total flexibility.

    linkedin.com/pulse/built-in-32

    #OpenHardware #Embedded #Electronics #VoiceSynthesis #PCBDesign #KT142 #Maker

  37. Ever wondered how AI voices are becoming so human-like? We've moved beyond simple "voice packs" (pre-recorded clips) to true AI "voice clones" that generate new speech from text.

    The magic is in the details: AI models learn a voice's unique pitch and cadence. The secret sauce? "Emotional tuning," which adds happiness, sadness, or empathy to the performance. It's a game-changer for accessibility and content creation. #AIVoice #VoiceSynthesis #Tech

  38. Ever wondered how AI voices are becoming so human-like? We've moved beyond simple "voice packs" (pre-recorded clips) to true AI "voice clones" that generate new speech from text.

    The magic is in the details: AI models learn a voice's unique pitch and cadence. The secret sauce? "Emotional tuning," which adds happiness, sadness, or empathy to the performance. It's a game-changer for accessibility and content creation. #AIVoice #VoiceSynthesis #Tech

  39. Ever wondered how AI voices are becoming so human-like? We've moved beyond simple "voice packs" (pre-recorded clips) to true AI "voice clones" that generate new speech from text.

    The magic is in the details: AI models learn a voice's unique pitch and cadence. The secret sauce? "Emotional tuning," which adds happiness, sadness, or empathy to the performance. It's a game-changer for accessibility and content creation. #AIVoice #VoiceSynthesis #Tech

  40. Ever wondered how AI voices are becoming so human-like? We've moved beyond simple "voice packs" (pre-recorded clips) to true AI "voice clones" that generate new speech from text.

    The magic is in the details: AI models learn a voice's unique pitch and cadence. The secret sauce? "Emotional tuning," which adds happiness, sadness, or empathy to the performance. It's a game-changer for accessibility and content creation. #AIVoice #VoiceSynthesis #Tech

  41. Project update: LingoFreq #Mandarin

    Generates themed example sentences from language word-frequency list (Chinese HSK). Builds a printable book & (soon) voice audio files. This theme is "technology".

    We'll be building this for #English & #Spanish.

    Tech: #Python, OpenAI ChatGPT, ElevenLabs #VoiceSynthesis (coming soon), SQLite database, Jinja templates

    Source code:
    - Sentence builder: codeberg.org/jro/LingoFreq-app
    - Book builder: codeberg.org/jro/LingoFreq-app

    #China

  42. Project update: LingoFreq #Mandarin

    Generates themed example sentences from language word-frequency list (Chinese HSK). Builds a printable book & (soon) voice audio files. This theme is "technology".

    We'll be building this for #English & #Spanish.

    Tech: #Python, OpenAI ChatGPT, ElevenLabs #VoiceSynthesis (coming soon), SQLite database, Jinja templates

    Source code:
    - Sentence builder: codeberg.org/jro/LingoFreq-app
    - Book builder: codeberg.org/jro/LingoFreq-app

    #China

  43. Project update: LingoFreq #Mandarin

    Generates themed example sentences from language word-frequency list (Chinese HSK). Builds a printable book & (soon) voice audio files. This theme is "technology".

    We'll be building this for #English & #Spanish.

    Tech: #Python, OpenAI ChatGPT, ElevenLabs #VoiceSynthesis (coming soon), SQLite database, Jinja templates

    Source code:
    - Sentence builder: codeberg.org/jro/LingoFreq-app
    - Book builder: codeberg.org/jro/LingoFreq-app

    #China

  44. Project update: LingoFreq #Mandarin

    Generates themed example sentences from language word-frequency list (Chinese HSK). Builds a printable book & (soon) voice audio files. This theme is "technology".

    We'll be building this for #English & #Spanish.

    Tech: #Python, OpenAI ChatGPT, ElevenLabs #VoiceSynthesis (coming soon), SQLite database, Jinja templates

    Source code:
    - Sentence builder: codeberg.org/jro/LingoFreq-app
    - Book builder: codeberg.org/jro/LingoFreq-app

    #China