home.social

#failover — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #failover, aggregated by home.social.

fetched live
  1. Redis Sentinel: `sentinel monitor <name> <host> <port> <quorum>` with quorum=2 ensures 2 Sentinels agree before promoting a replica. Config on Redis 5.0+/Ubuntu 20.04. Automatic failover without manual intervention. #sentinel #high-availability #failover #ValtersIT

    valtersit.com/vault/automated-

  2. This happens. The money I've spent to allow this to happen? Worth every penny. #resilient #failover

  3. This happens. The money I've spent to allow this to happen? Worth every penny. #resilient #failover

  4. Das Einzige, was ich noch nicht rausbekommen habe: Wie erhalte ich bei #Failover auf 5G eine Nachricht (per Push oder als SMS)?

  5. Die Situation mit separatem 4G-#Failover vor meinem #Router war mir schon lange ein Dorn im Auge - und funktionierte auch nicht zuverlässig. Also habe ich erst versucht, mein #Netgear Orbi RBK752/LM1200-Kombi durch einen Netgear NBK752 zu ersetzen … nur, um nach Lieferung festzustellen, dass das Gerät komplett veraltet (und daher so preiswert zu haben) ist.

    ⬇️

    #SmartHome #WLAN

  6. Почему ваш DR-план не сработает при реальной аварии (и что с этим делать)

    Недавно на вебинаре Хайстекс разбирались реальные сценарии DR на практике. В чате трансляции зрители задали ряд технических вопросов, наиболее интересные из них легли в основу этого разбора.

    habr.com/ru/companies/hstx/art

    #аварийное_восстановление #disaster_recovery #RTO #RPO #репликация #failover #failback #резервное_копирование #виртуализация #хайстекс_акура

  7. Три задачи discovery при работе с PostgreSQL master/replica — и как их решить

    Когда у приложения появляется несколько хостов PostgreSQL, начинается головная боль: нужно динамически находить мастера после failover, выбирать реплику с нужным отставанием и гарантировать что пользователь не увидит устаревшие данные после своей же записи. DNS кешируется минутами, libpq не знает про lag, HAProxy не слышал про LSN. Разбираем как устроены существующие решения и как закрыть все три задачи через лёгкий HTTP сервис — pg-status .

    habr.com/ru/articles/1047374/

    #postgresql #репликация #failover #replication #python #sql #discovery #high_availability

  8. @PC_Fluesterer Auf jeden Fall unprofessionell, schlecht designed, schlecht gemacht und sowas von nicht zeitgemäß. Dabei weiß man doch schon lange, wie man hohe Verfügbarkeit, Redundanz und Ausfallsicherheit in der IT architekturell erlangt und, daß diese Komponenten nicht optional sind.

    #itsec #performance #ausfallsicherheit #redundanz #failover #fallback #verfügbarkeit #devops #itarchitektur #togaf #rechenzentrum #spof #singlepointoffailure #bottleneck

  9. @PC_Fluesterer Auf jeden Fall unprofessionell, schlecht designed, schlecht gemacht und sowas von nicht zeitgemäß. Dabei weiß man doch schon lange, wie man hohe Verfügbarkeit, Redundanz und Ausfallsicherheit in der IT architekturell erlangt und, daß diese Komponenten nicht optional sind.

    #itsec #performance #ausfallsicherheit #redundanz #failover #fallback #verfügbarkeit #devops #itarchitektur #togaf #rechenzentrum #spof #singlepointoffailure #bottleneck

  10. Строим шину данных для микросервисов на ZeroMQ: failover, гарантии доставки и E2E-шифрование

    Асинхронная клиент-серверная библиотека для обмена сообщениями между микросервисами на базе ZeroMQ. Реализует гарантированную доставку сообщений (At-Least-Once) с персистентной файловой очередью при обрывах связи, автоматический failover сервера переадресации (клиенты могут подхватывать роль сервера на лету) и два уровня защиты: шифрование канала (CurveZMQ) и сквозное шифрование сообщений (HMAC). Лёгкая альтернатива брокерам вроде RabbitMQ, не требующая отдельного сервера.

    habr.com/ru/articles/1030020/

    #python #zeromq #zmq #failover #atleastonce #endtoend_шифрование #микросервисы #распределенные_системы #hmac #криптография

  11. Turns out, failover success is subjective. Apparently, being ‘active’ just means you get tested harder. Ever wondered how ‘best intentions’ can invent new incidents? Let’s talk IT wisdom in the replies.

    Find out more in Episode 12 : The Failover That Failed Successfully

    youtube.com/shorts/L3s3K2E4-1I

    Listen here : ithorrorstories.eu/#ep12

    All other things : links.ithorrorstories.eu/

    #podcasts #failover #drtest #disaster #technology #operationalreadyness #tech

  12. Turns out, failover success is subjective. Apparently, being ‘active’ just means you get tested harder. Ever wondered how ‘best intentions’ can invent new incidents? Let’s talk IT wisdom in the replies.

    Find out more in Episode 12 : The Failover That Failed Successfully

    youtube.com/shorts/L3s3K2E4-1I

    Listen here : ithorrorstories.eu/#ep12

    All other things : links.ithorrorstories.eu/

  13. Ever run a failover test that worked perfectly… and still felt like everything was falling apart?

    In Episode 12, we take you into a disaster recovery test during a busy release weekend — where the tech held up, but communication didn’t.

    Subcontractors weren’t aligned, assumptions didn’t match reality, and suddenly a ‘simple test’ turned into a full coordination puzzle.

    No production impact — but plenty of lessons.

    Because resilience isn’t just about systems… it’s about people, timing, and actually talking to each other.

    Listen now to IT Horror Stories with Jack Smith
    You can find us on Spotify, Apple Music, Youtube, Deezer and of course at ITHorrorStories.eu

    You are one of us.

    youtube.com/shorts/k_SyFbQ71TU

    #podcast #technology #failover #failure #techlife

  14. Ever run a failover test that worked perfectly… and still felt like everything was falling apart?

    In Episode 12, we take you into a disaster recovery test during a busy release weekend — where the tech held up, but communication didn’t.

    Subcontractors weren’t aligned, assumptions didn’t match reality, and suddenly a ‘simple test’ turned into a full coordination puzzle.

    No production impact — but plenty of lessons.

    Because resilience isn’t just about systems… it’s about people, timing, and actually talking to each other.

    Listen now to IT Horror Stories with Jack Smith
    You can find us on Spotify, Apple Music, Youtube, Deezer and of course at ITHorrorStories.eu

    You are one of us.

    youtube.com/shorts/k_SyFbQ71TU

  15. For those who run #ProsodyIM as #xmpp server, I did something simple but effective in my failover architecture:

    • 2 Prosody instances in two different regions in a datacenter
    • lsyncd syncing from primary to stand by instance all data
    • an entrypoint script supervising Prosody execution
    • a lock file controlling if entrypoint script can up Prosody
    • a daemon checking if floating ip is linked to hosts and controlling the lock file and the lsyncd execution and configuration to primary/standby modes

    Perfect solution? Of course not.
    Effective solution? Hell yeah.

    :isacloud: :isacloudim:

    #xmpp #failover #container #vrrp #prosodyim #prosodyim

  16. For those who run #ProsodyIM as #xmpp server, I did something simple but effective in my failover architecture:

    • 2 Prosody instances in two different regions in a datacenter
    • lsyncd syncing from primary to stand by instance all data
    • an entrypoint script supervising Prosody execution
    • a lock file controlling if entrypoint script can up Prosody
    • a daemon checking if floating ip is linked to hosts and controlling the lock file and the lsyncd execution and configuration to primary/standby modes

    Perfect solution? Of course not.
    Effective solution? Hell yeah.

    :isacloud: :isacloudim:

    #xmpp #failover #container #vrrp #prosodyim #prosodyim

  17. During a busy release weekend, a planned failover exposed not technical flaws, but something more familiar: misaligned teams, unclear responsibilities, and communication that didn’t quite arrive when it should have. Production stayed safe — but confidence took a hit.

    This episode dives into how a “simple test” turned into a coordination challenge, and why resilience is just as much about people and processes as it is about systems.

    Find all links to listen on our website : ithorrorstories.eu/#ep12

    #podcast #datarecovery #failover #test #technology

    You can find our podcast on :

    Spotify : open.spotify.com/show/7LqbtykS
    Apple Music : podcasts.apple.com/us/podcast/
    YouTube : music.youtube.com/playlist?lis
    Deezer : link.deezer.com/s/30dyH3RoKvN8

  18. During a busy release weekend, a planned failover exposed not technical flaws, but something more familiar: misaligned teams, unclear responsibilities, and communication that didn’t quite arrive when it should have. Production stayed safe — but confidence took a hit.

    This episode dives into how a “simple test” turned into a coordination challenge, and why resilience is just as much about people and processes as it is about systems.

    Find all links to listen on our website : ithorrorstories.eu/#ep12

    You can find our podcast on :

    Spotify : open.spotify.com/show/7LqbtykS
    Apple Music : podcasts.apple.com/us/podcast/
    YouTube : music.youtube.com/playlist?lis
    Deezer : link.deezer.com/s/30dyH3RoKvN8

  19. [Перевод] Осваиваем replication slots в Postgres: как предотвратить разрастание WAL и другие проблемы в продакшене

    Логическая репликация в Postgres редко ломает прод внезапно — чаще она долго и методично копит проблему, пока replication slot удерживает всё больше WAL, потребитель отстаёт, а свободное место на диске начинает таять. В этой статье разбирается именно такая зона риска: как устроена работа replication slots, почему одних базовых настроек здесь недостаточно и какие практики реально помогают держать под контролем WAL, публикации, heartbeats, failover и мониторинг. Материал особенно полезен тем, кто работает с CDC, Debezium и production-инстансами Postgres, где цена ошибки измеряется уже не теорией, а стабильностью системы. Разбор PostgreSQL

    habr.com/ru/companies/otus/art

    #PostgreSQL #replication_slots #логическая_репликация #WAL #CDC #Debezium #pgoutput #failover #мониторинг_Postgres

  20. Мониторинг SQL Server Always On в Zabbix

    Если у вас стоит Always On Availability Groups, вы наверняка бывали в такой ситуации: в SSMS всё зелёное, дашборд показывает «Synchronized», а пользователи звонят с жалобами на тормоза. Смотришь на secondary — а там redo_queue_size 600 МБ, реплика отстаёт на полчаса. Ни одного алерта. У нас это случилось на продуктивном кластере с 1С: secondary молча отвалился в SYNCHRONIZING, а мы узнали только при плановом переключении. Полтора часа redo queue. Стало понятно, что встроенный дашборд SSMS — это не мониторинг. Дальше — как мы это закрыли Zabbix'ом за вечер.

    habr.com/ru/companies/cloud4y/

    #SQL_Server #Always_On #Zabbix #мониторинг #DMV #WSFC #кворум #failover #DBA

  21. And suddenly the NAS had switched itself off. It's up and running again but this is not good. Glad I have Nastig set up and ready to take over if the NAS is dying.

    #recovery #NAS #failover

  22. And suddenly the NAS had switched itself off. It's up and running again but this is not good. Glad I have Nastig set up and ready to take over if the NAS is dying.

    #recovery #NAS #failover

  23. #throwback What really happens inside a PostgreSQL cluster after failover? 🔄 David Pech dives deep into failover, switchover, split-brain scenarios, and recovery strategies—manually breaking down how tools like Patroni work.

    ▶️ Watch the video now! lnkd.in/dQiiBvXX

    #PostgreSQL #PGDay #PPDD #Failover #HighAvailability

  24. #throwback What really happens inside a PostgreSQL cluster after failover? 🔄 David Pech dives deep into failover, switchover, split-brain scenarios, and recovery strategies—manually breaking down how tools like Patroni work.

    ▶️ Watch the video now! lnkd.in/dQiiBvXX

    #PostgreSQL #PGDay #PPDD #Failover #HighAvailability

  25. Я наконец-то понял, как открытость может помешать — и отчёт об аварии

    В прошлый понедельник у нас случилась очередная крайне идиотская авария. Идиоты тут мы, если что, и сейчас я расскажу детали. Пострадало четыре сервера из всего ЦОДа — и все наши публичные коммуникации. Потому что владельцы виртуальных машин пришли под все посты и везде оставили комментарии. Параллельно была ещё одна история — под статьёй про то, что случалось за год, написал человек, мол, чего у вас всё постоянно ломается. Я вот размещаюсь у регионального провайдера, и у него за 7 лет ни одной проблемы. Так вот. Разница в том, что мы про всё это рассказываем. Тот провайдер наверняка уже раз 10 падал, останавливался и оставался без сети, но грамотно заталкивал косяки под ковёр. Это значит — никаких блогов на Хабре, никаких публичных коммуникаций с комментариями (типа канала в Телеграме), никаких объяснений кроме лицемерных ответов от службы поддержки и т.п. И тогда, внезапно, вас будут воспринимать более стабильным и надёжным. Наверное. Ну а я продолжаю рассказывать, что у нас происходило. Добро пожаловать в очередной RCA, где главное в поиске root cause было не выйти на самих себя. Но мы вышли!

    habr.com/ru/companies/ruvds/ar

    #ruvds_статьи #цод #авария #rca #ибп #резервное_питание #дизельгенераторные_установки #клиентский_сервис #failover

  26. Patroni и логическая реплика в PostgreSQL: как не потерять данные при failover’е

    Если вы используете nofailover: true (а многие так и делают), Patroni не синхронизирует слоты логической репликации — и при переходе на реплику часть данных может исчезнуть навсегда. Рассказываем, почему и как фиксить.

    habr.com/ru/companies/flant/ar

    #patroni #postgresql #sql #бд #failover #репликация #асинхронная_репликация #логическая_репликация #nofailover #потеря_данных

  27. Bevor ich jetzt Geld für Hardware ausgebe: Weiß jemand, ob ich meinen Netgear LTE Router unter dem Dach als WAN Failover nutzen kann, obwohl er über LAN (mit Switches dazwischen) mit dem Unifi Cloud Gateway im Keller verbunden ist? Über VLAN oder so, also ohne dedicated Port? Kann die Admin Software von Unifi das, oder wollen die, dass ich das überteuerte LTE Backup Pro kaufe?

    Das Unifi LTE Backup (Pro) funktioniert ja scheinbar ähnlich…

    Dach:
    - Netgear LTE Router
    - Unifi WIFI Access Point
    - Unifi Switch

    Keller:
    - Unifi Switch
    - Unifi Cloud Gateway (UCG)
    - Gasfaser ONT an WAN 1 von UCG

    #unifi #network #netzwerk #setup #wifi #5g #4g #failover #sysadmin #backup #internet #schwarmintelligenz #glasfaser #ausfallsicherheit

  28. Bevor ich jetzt Geld für Hardware ausgebe: Weiß jemand, ob ich meinen Netgear LTE Router unter dem Dach als WAN Failover nutzen kann, obwohl er über LAN (mit Switches dazwischen) mit dem Unifi Cloud Gateway im Keller verbunden ist? Über VLAN oder so, also ohne dedicated Port? Kann die Admin Software von Unifi das, oder wollen die, dass ich das überteuerte LTE Backup Pro kaufe?

    Das Unifi LTE Backup (Pro) funktioniert ja scheinbar ähnlich…

    Dach:
    - Netgear LTE Router
    - Unifi WIFI Access Point
    - Unifi Switch

    Keller:
    - Unifi Switch
    - Unifi Cloud Gateway (UCG)
    - Gasfaser ONT an WAN 1 von UCG

    #unifi #network #netzwerk #setup #wifi #5g #4g #failover #sysadmin #backup #internet #schwarmintelligenz #glasfaser #ausfallsicherheit

  29. We just made the new VDCF Free Edition 9.0.19 available.
    Try it in your environment.
    We are happy to support you by email ([email protected])

    Download and Register to our eMail Newsletter on
    jomasoft.ch/downloads/#vdcf-fr

    #solaris #oraclesolaris #virtualization #monitoring #operations #failover #vdcf

  30. We just made the new VDCF Free Edition 9.0.19 available.
    Try it in your environment.
    We are happy to support you by email ([email protected])

    Download and Register to our eMail Newsletter on
    jomasoft.ch/downloads/#vdcf-fr

    #solaris #oraclesolaris #virtualization #monitoring #operations #failover #vdcf

  31. Hey folks, is there an equivalent to the AWS JDBC Driver (github.com/aws/aws-adva...) for #golang? Specifically looking for support for fast(er) failover on Amazon Aurora #PostgreSQL. #aws #aws_rds #Aurora #AmazonWebServices #jdbc #failover

    GitHub - aws/aws-advanced-jdbc...

  32. Hey folks, is there an equivalent to the AWS JDBC Driver (github.com/aws/aws-advanced-jd) for #golang?
    Specifically looking for support for fast(er) failover on Amazon Aurora #PostgreSQL.

    #aws #aws_rds #Aurora #AmazonWebServices #jdbc #failover