home.social

#sitereliability — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #sitereliability, aggregated by home.social.

fetched live
  1. WordPress 7 is live, bringing AI‑assisted content tools, improved site editing, and performance gains.

    It’s an exciting update — but also one that deserves caution, especially for WooCommerce stores, WPML users, and page‑builder‑based sites.

    This post covers what’s new and how to update safely:
    👉 meraksystems.com/blog/2026/05/

    #WordPress7 #OpenWeb #WebDevelopment #SiteReliability

  2. 😊 Happy with your current #DNS provider? Fantastic. 👨‍💻 But you should be 100% sure you have multiple DNS providers so your site doesn’t go down if something happens.
    That’s why Secondary DNS is so important and setting it up is easier than you might think 👇
    🎥 Watch here: youtu.be/NPlkDqLL2Vo

    #DNS #SecondaryDNS #Networking #WebInfrastructure #SiteReliability

  3. 😊 Happy with your current #DNS provider? Fantastic. 👨‍💻 But you should be 100% sure you have multiple DNS providers so your site doesn’t go down if something happens.
    That’s why Secondary DNS is so important and setting it up is easier than you might think 👇
    🎥 Watch here: youtu.be/NPlkDqLL2Vo

    #DNS #SecondaryDNS #Networking #WebInfrastructure #SiteReliability

  4. Every system works perfectly until it meets DNS, timezones, certificates, or humans.
    Usually at the same time.
    In production.
    On a Friday.

    Experience is just pattern recognition with better alerts.

    #Production #DevOps #SiteReliability #EngineeringHumor #IncidentResponse #OnCall #TechReality #ByernNotes

  5. Today's AWS outage was a stark reminder: what happens when the tools you rely on to manage incidents... are part of the incident?

    When Slack, Zoom, PagerDuty, and even Statuspage are impacted, how do you get your response team re-connected to solve the underlying problem? Once they're talking to each other, they can improvise a response, but that first step of re-establishing contact is critical.

    This isn't just a hypothetical. It's a real-world scenario that can paralyze even the most prepared organizations. Relying on a plan that's tucked away in a long-forgotten document is a recipe for disaster.

    Here's what I recommend to the leaders I advise:

    🔹 Have a "Rally Point" Plan: Don't just have a backup concept; have a pre-defined, communicated, and accessible fallback plan. Every second counts in an incident, and you can't waste time figuring out where to communicate. If you normally use Slack and Zoom, then think Google Meet or Microsoft Teams for your backup, and vice versa. Maybe even an old-fashioned conference call bridge. The key is that everyone knows where to go, when the normal places aren't working.

    🔹 Make it Accessible: Your plan is useless if it's on a server that nobody can get to at the moment. Laminated wallet cards, a shared password vault with offline access, or a regularly updated file on every employee's laptop are all viable options.

    🔹 Practice, Practice, Practice: Fire drills aren't just for fires. Run drills for your fallback communication plan. This ensures everyone remembers it exists and that the mechanisms still work.

    🔹 Don't Forget Security: Assume that your fallback channel is compromised, and that outsiders are listening in. Use it just as a rendezvous point to direct responders to more secure, authenticated channels, where you can validate every participant. Don't discuss sensitive information in the open.

    Incidents are costly, not just in revenue, but in reputation and team morale. Proactive preparation isn't a luxury; it's a necessity.

    What's your team's communication fallback plan? Share your thoughts in the comments below. 👇

    #IncidentManagement #BusinessContinuity #SiteReliability #DevOps #AWSOutage

  6. Today's AWS outage was a stark reminder: what happens when the tools you rely on to manage incidents... are part of the incident?

    When Slack, Zoom, PagerDuty, and even Statuspage are impacted, how do you get your response team re-connected to solve the underlying problem? Once they're talking to each other, they can improvise a response, but that first step of re-establishing contact is critical.

    This isn't just a hypothetical. It's a real-world scenario that can paralyze even the most prepared organizations. Relying on a plan that's tucked away in a long-forgotten document is a recipe for disaster.

    Here's what I recommend to the leaders I advise:

    🔹 Have a "Rally Point" Plan: Don't just have a backup concept; have a pre-defined, communicated, and accessible fallback plan. Every second counts in an incident, and you can't waste time figuring out where to communicate. If you normally use Slack and Zoom, then think Google Meet or Microsoft Teams for your backup, and vice versa. Maybe even an old-fashioned conference call bridge. The key is that everyone knows where to go, when the normal places aren't working.

    🔹 Make it Accessible: Your plan is useless if it's on a server that nobody can get to at the moment. Laminated wallet cards, a shared password vault with offline access, or a regularly updated file on every employee's laptop are all viable options.

    🔹 Practice, Practice, Practice: Fire drills aren't just for fires. Run drills for your fallback communication plan. This ensures everyone remembers it exists and that the mechanisms still work.

    🔹 Don't Forget Security: Assume that your fallback channel is compromised, and that outsiders are listening in. Use it just as a rendezvous point to direct responders to more secure, authenticated channels, where you can validate every participant. Don't discuss sensitive information in the open.

    Incidents are costly, not just in revenue, but in reputation and team morale. Proactive preparation isn't a luxury; it's a necessity.

    What's your team's communication fallback plan? Share your thoughts in the comments below. 👇

    #IncidentManagement #BusinessContinuity #SiteReliability #DevOps #AWSOutage

  7. 🔍 Transforming Telecom with Automation

    OSS/BSS are the backbone of telecom operations.

    With automation, RELIANOID helps telecoms achieve:
    🔒 Enhanced security with SNMPv3.
    ⚡ 99.999% uptime via high availability.
    ⏱️ 70% faster issue resolution with automated workflows.

    Discover how we optimized OSS/BSS for a global telecom giant.
    ➡️ Request a demo to boost reliability with RELIANOID!


    relianoid.com/blog/oss-bss-rel

  8. Hannaford's recent weeklong outage has me wondering: Do companies truly understand the cost of cutting corners on engineering talent?
    These unacceptably long outages which are more frequently occurring at major retailers highlights a common problem I'm seeing in tech: undervaluing highly experienced & knowledgeable engineers. It's way past time for companies to rethink their hiring priorities... stop cheaping out on your Ops and Sec talent, it's going to cost you far more in the end!
    I'm exceptionally good at building reliable & resilient systems & teams, so it's super frustrating to be unemployed while witnessing preventable outages for which I could have made a difference. Yes, it's true, 30+ years of engineering experience doesn't come cheap, but I'm damn sure my price is far less than the loss in revenue from a weeklong eComm outage at a major business!
    Anyway, if yer looking for a decent engineer/leader, please reach out...
    #open_to_work #engineering #siteReliability #Technology

    mainepublic.org/business-and-e

  9. Hannaford's recent weeklong outage has me wondering: Do companies truly understand the cost of cutting corners on engineering talent?
    These unacceptably long outages which are more frequently occurring at major retailers highlights a common problem I'm seeing in tech: undervaluing highly experienced & knowledgeable engineers. It's way past time for companies to rethink their hiring priorities... stop cheaping out on your Ops and Sec talent, it's going to cost you far more in the end!
    I'm exceptionally good at building reliable & resilient systems & teams, so it's super frustrating to be unemployed while witnessing preventable outages for which I could have made a difference. Yes, it's true, 30+ years of engineering experience doesn't come cheap, but I'm damn sure my price is far less than the loss in revenue from a weeklong eComm outage at a major business!
    Anyway, if yer looking for a decent engineer/leader, please reach out...

    mainepublic.org/business-and-e

  10. No, I did not want to have a system-wide outage this morning, thankyouverymuch 😰

    (but we recovered, although not without some sweating. Aren't new and different failure modes fun?)

    (no, I'm not an SRE but we're a small shop)

    #onCall #siteReliability #SRE

  11. No, I did not want to have a system-wide outage this morning, thankyouverymuch 😰

    (but we recovered, although not without some sweating. Aren't new and different failure modes fun?)

    (no, I'm not an SRE but we're a small shop)

    #onCall #siteReliability #SRE

  12. "What should I monitor? Am I tracking the right metrics?" 📈📊
    Common industry metrics frameworks provide useful monitoring guidance for and .
    Here's a good overview for the different methods:
    logz.io/blog/evops-sre-metrics

  13. "What should I monitor? Am I tracking the right metrics?" 📈📊
    Common industry metrics frameworks provide useful monitoring guidance for #DevOps and #SRE.
    Here's a good overview for the different methods:
    logz.io/blog/evops-sre-metrics
    #monitoring #observability #sitereliability

  14. No one ever complains about #steam going down or being slow, despite tens of millions of concurrent users at all times. I'd like to know more about how Valve manages that. The service itself is practically transparent. #sitereliability #devops #cloud #CloudComputing #videogames

  15. No one ever complains about #steam going down or being slow, despite tens of millions of concurrent users at all times. I'd like to know more about how Valve manages that. The service itself is practically transparent. #sitereliability #devops #cloud #CloudComputing #videogames

  16. Here are the steps to enable #http3/#quic in #caddy:
    ....

    It takes 0, zero, nil lines to enable and configure #http3/#quic in #CaddyServer! You don't need to do anything special to keep up with the industry standard and progress. Caddy takes care of keeping your services up-to-date.

    #systemadministration #sysadmin #devops #sre #web #linux #unix #windows #sitereliability

  17. Here are the steps to enable #http3/#quic in #caddy:
    ....

    It takes 0, zero, nil lines to enable and configure #http3/#quic in #CaddyServer! You don't need to do anything special to keep up with the industry standard and progress. Caddy takes care of keeping your services up-to-date.

    #systemadministration #sysadmin #devops #sre #web #linux #unix #windows #sitereliability

  18. In 2022, we saw a massive increase in interest for continuous profiling using eBPF, and also some really interesting tools like Parca (and Polar Signals), as well as @grafana's Phlare,

    What do you think will be the next big thing for and in 2023?

  19. #Introduction 👋 Hello World!

    I’m a proud #dogMom that loves to overshare photos of my #rescue #dog (Cassie).

    Bringing #diversityEquitiyInclusion to #tech motivates me.

    Professionally, I’ve had a long career in #softwareEngineering, but am now on a journey in the world of #siteReliability #engineering.

    Sometimes I’ll also post things about #food, #coffee, #whiskey / #whisky, #wine, #travel, #nba #basketball, and #snowboarding.

    #introductions #dei #womenwhocode #sre #developer #dogs

  20. #Introduction 👋 Hello World!

    I’m a proud #dogMom that loves to overshare photos of my #rescue #dog (Cassie).

    Bringing #diversityEquitiyInclusion to #tech motivates me.

    Professionally, I’ve had a long career in #softwareEngineering, but am now on a journey in the world of #siteReliability #engineering.

    Sometimes I’ll also post things about #food, #coffee, #whiskey / #whisky, #wine, #travel, #nba #basketball, and #snowboarding.

    #introductions #dei #womenwhocode #sre #developer #dogs

  21. Questions to Engineering Managers:
    • How long were you an individual contributing engineer before making the switch to management?
    • What was your motivation to change your path?
    • If you switched back to an IC role, what made you leave management?
    • What advice would you give to your younger self now to put them at ease about making the switch?

    #engineeringmanagement #engineeringleadership #softwareengineering #techlead #developer #management #sre #sitereliability #tech #techcareer #career

  22. Questions to Engineering Managers:
    • How long were you an individual contributing engineer before making the switch to management?
    • What was your motivation to change your path?
    • If you switched back to an IC role, what made you leave management?
    • What advice would you give to your younger self now to put them at ease about making the switch?

    #engineeringmanagement #engineeringleadership #softwareengineering #techlead #developer #management #sre #sitereliability #tech #techcareer #career

  23. Aaand issue fixed before user came.

    Love the early alerting system and anomaly detection! Making sure the system is reliable and let us know before report happens.

    Hope this can make us having few steps closer to proper SRE practices~

    #SRE #SiteReliability #FioTechRant

  24. FireHydrant announces $23M Series B to grow disaster management platform - No matter how carefully you set up your systems, sooner or later something fundame... - feedproxy.google.com/~r/Techcr #disasterrecovery #sitereliability #recentfunding #engineering #firehydrant #developer #startups #funding #cloud #sres #tc

  25. StackPulse announces $28M investment to help developers manage outages - When a system outage happens, chaos can ensue as the team tries to figure out what’s happening and h... - feedproxy.google.com/~r/Techcr #bessemerventurepartners #disasterrecovery #sitereliability #recentfunding #enterprise #stackpulse #developer #startups #funding #ggv

  26. Lead #SiteReliability Engineer shares lessons learned implementing #NGINX at @Dell EMC. Note the last lesson.. it leads very nicely into the next portion of our livestream starting at 11:35 AM PT ;) nginx.com/livestream