#incidentmanagement — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #incidentmanagement, aggregated by home.social.
-
A cardboard factory doesn't sound like critical infrastructure...
...until every hour of downtime costs serious money.
In Episode 17, Kayne McGladrey shares a real manufacturing incident, how the team handled it, and the lessons every IT professional can take away.
Sometimes the least glamorous systems turn out to be the most business critical.
Listen here : https://ithorrorstories.eu/#ep17
#ITHorrorStories #Manufacturing #CyberSecurity #Podcast #IncidentManagement
-
How inDrive built 24/7 incident support across Kazakhstan, Brazil, and Malaysia, reducing SLA from 72 hours to 2—without night shifts https://hackernoon.com/follow-the-sun-support-how-indrive-cut-incident-sla-from-72-hours-to-2-hours #incidentmanagement
-
🚀🚨 IncidentRelay: O nouă platformă open-source promite să revoluționeze managementul incidentelor și serviciul on-call
Pentru echipele de DevOps și inginerii de site reliability (SRE), gestionarea alertelor și a programului de permanență (on-call) este adesea o sursă majoră de stres. Piața actuală este dominată de soluții comerciale proprietare scumpe, precum PagerDuty sau Opsgenie. În acest context, lansarea IncidentRelay vine ca o gură de aer proaspăt pentru comunitatea open-source, oferind o platformă modernă, puternică și complet transparentă pentru managementul incidentelor critice.
IncidentRelay își propune să unifice fluxul de alerte dintr-o infrastructură și să se asigure că persoana potrivită este notificată la momentul potrivit, fără bătăi de cap legate de licențiere.
Iată principalele caracteristici și avantaje introduse de IncidentRelay:
🔹 Gestionarea inteligentă a programului On-Call (Rotations & Schedules):
Inima platformei este motorul său flexibil de planificare. IncidentRelay permite administratorilor să creeze calendare complexe de permanență, rotații automate între membrii echipei și reguli stricte de escaladare. Dacă o alertă critică nu primește răspuns de la inginerul de serviciu într-un interval stabilit de minute, platforma o trimite automat către următorul nivel de suport.🔹 Integrare nativă cu uneltele de monitorizare:
Platforma acționează ca un hub centralizat pentru toate sistemele tale de monitorizare. IncidentRelay vine cu suport pentru recepționarea alertelor prin Webhooks de la utilitare populare precum:Prometheus și Grafana
Zabbix și Datadog
Alerte din infrastructuri cloud (AWS, Google Cloud, Azure)
🔹 Canale de notificare flexibile:
Când serverele pică la ora 3 dimineața, notificarea trebuie să fie eficientă. IncidentRelay oferă opțiuni multiple pentru a alerta inginerii: de la mesaje instantanee pe platforme de chat precum Slack, Discord sau Matrix, până la trimiterea de SMS-uri sau apeluri telefonice automatizate, asigurându-se că nicio alertă critică nu este trecută cu vederea.🔹 Control total și independență (Self-Hosting):
Spre deosebire de alternativele SaaS, IncidentRelay poate fi auto-găzduită în totalitate în propria infrastructură securizată. Acest lucru oferă companiilor un control absolut asupra jurnalelor de incidente și a datelor sensibile despre infrastructură, eliminând totodată riscul ca instrumentul de alertare să devină indisponibil din cauza unei pene de curent din cloud-ul public extern.Lansarea IncidentRelay marchează un pas important către democratizarea uneltelor de tip Enterprise DevOps, oferind companiilor de toate dimensiunile o soluție gratuită, extrem de robustă și personalizabilă pentru a menține sistemele online și funcționale.
#OpenSource #IncidentRelay #DevOps #SRE #OnCall #IncidentManagement #SysAdmin #TechNews #Linuxiac
-
The Treachery of Postmortems is a really great cure for the post incident review blues. Thank you @gallego !
https://resilienceinsoftware.org/news/11547831
#RISF #ResilienceInSoftwareFoundation #ResilienceInSoftware #Resilience #Postmortem #PIRWriteup #IncidentManagement #IncidentResponse #LearningReview
-
One laptop.
One warehouse.
One very expensive sleep mode.Episode 14:
Sleep Mode in ProductionListen here : https://ithorrorstories.eu/#ep14
#technology #podcast #shadowIT #operations #incidentmanagement
-
Production incident and the first question is: have we seen this before?
I built a Quarkus service that turns Java incidents into vectors with deterministic feature hashing — no embedding model, no LLM. Store them in Qdrant, search by failure shape, filter by service and environment.
The vectors are inspectable and repeatable. The scores are explainable from the input.
New tutorial on The Main Thread:
-
How realistic incident simulations helped product engineers build confidence, reduce mitigation time, improve communication, and strengthen blameless culture. https://hackernoon.com/beyond-on-call-how-we-taught-product-engineers-to-own-their-incidents #incidentmanagement
-
New from me today: A roundup of Datadog #DASH2026 livestreamed engineering breakout sessions, which all touched on a common theme: that AI-driven #incidentmanagement tools only work if human #platformengineers have first designed a solid infrastructure and set of workflows.
https://www.techtarget.com/searchitoperations/news/366644443/Datadog-shops-AI-incident-management-needs-platform-engineers #datadog #o11y #AI
-
The Engineering Leadership Crisis Nobody Talks About 🚨 #EngineeringLeadership #SoftwareEngineering #PlatformEngineering #TechLeadership #Microservices #SRE
Modern engineering teams are collapsing under platform complexity, AI chaos, organizational scaling failures, and unreliable architectures. This deep technical leadership guide explains how elite engineering leaders manage platform rewrites, reliability crises, organizational chaos, and large-scale modernization without destroying delivery velocity. #SoftwareArchitecture #EngineeringManagement #DevOps #CloudComputing #Leadership -
#Development #Findings
The Pragmatic Engineer 2025 Survey (Part 3) · Which tools do software engineers use today? https://ilo.im/167n2s_____
#Observability #IncidentManagement #Experimentation #TechStack #Tooling #Frameworks #DevOps #WebDev #Frontend -
Auch 2026 findet wieder ein #GI-SPRING-Graduiertenworkshop der Fachgruppe Security - Intrusion Detection and Response (SIDAR) statt. Diesmal am 21. und 22.04.2026 in #Heidelberg.
Zu den Themen gehören #VulnerabilityAssessment, #ThreatIntelligence, #IntrusionDetection, #Malware, #IncidentManagement, #WirelessSecurity, #DigitalForensics usw.
Einreichungen werden bis zum 15.03.2026 angenommen.
-
Today's AWS outage was a stark reminder: what happens when the tools you rely on to manage incidents... are part of the incident?
When Slack, Zoom, PagerDuty, and even Statuspage are impacted, how do you get your response team re-connected to solve the underlying problem? Once they're talking to each other, they can improvise a response, but that first step of re-establishing contact is critical.
This isn't just a hypothetical. It's a real-world scenario that can paralyze even the most prepared organizations. Relying on a plan that's tucked away in a long-forgotten document is a recipe for disaster.
Here's what I recommend to the leaders I advise:
🔹 Have a "Rally Point" Plan: Don't just have a backup concept; have a pre-defined, communicated, and accessible fallback plan. Every second counts in an incident, and you can't waste time figuring out where to communicate. If you normally use Slack and Zoom, then think Google Meet or Microsoft Teams for your backup, and vice versa. Maybe even an old-fashioned conference call bridge. The key is that everyone knows where to go, when the normal places aren't working.
🔹 Make it Accessible: Your plan is useless if it's on a server that nobody can get to at the moment. Laminated wallet cards, a shared password vault with offline access, or a regularly updated file on every employee's laptop are all viable options.
🔹 Practice, Practice, Practice: Fire drills aren't just for fires. Run drills for your fallback communication plan. This ensures everyone remembers it exists and that the mechanisms still work.
🔹 Don't Forget Security: Assume that your fallback channel is compromised, and that outsiders are listening in. Use it just as a rendezvous point to direct responders to more secure, authenticated channels, where you can validate every participant. Don't discuss sensitive information in the open.
Incidents are costly, not just in revenue, but in reputation and team morale. Proactive preparation isn't a luxury; it's a necessity.
What's your team's communication fallback plan? Share your thoughts in the comments below. 👇
#IncidentManagement #BusinessContinuity #SiteReliability #DevOps #AWSOutage
-
In DevOps, the real differentiator at 2 AM isn’t just the tech stack—it’s the soft skills that hold the line. Explore actual incident stories, unexpected lessons, and the human side of DevOps in “DevOps Soft Skills That Save You at 2 AM”.
Read more: https://shorturl.at/NabPS -
Some folks may recall my anger on August 18 over a vendor who wasn't responding to alerts about exposing their clients' data. The data included court files or records that were confidential or even sealed. At the time, researchers had discovered two entities that were exposed. They subsequently discovered more.
Yesterday, the vendor -- who had even ignored a call from the FBI -- finally secured one of the two after the client finally reached them on the phone.
The vendor told them they had fixed the problem. But did they?
[SPOILER ALERT: No.]
You won't believe what happened next, or maybe you will, but you'll have to stay tuned for this story, which has now gotten astronomically bigger because not only were the data still not secured but the vendor -- after claiming that the researchers had used hacking techniques to access unsecured data -- inexplicably sent the client a list of ALL of vendor's clients with their technical details AND ALL OF THEIR LOGIN CREDENTIALS.
[WTF!?]
I have never been as tempted to issue an actual press release warning all entities about a specific vendor, but... wow.
Stay tuned. Eventually, I will write this all up, but first, I want to hear what the client's lawyers and insurers decide to do to hold the vendor accountable.
(August 18 post: https://infosec.exchange/deck/@PogoWasRight/115033245331860859)
#databreach #dataleak #incidentresponse #incidentmanagement #thirdparty #vendor #accountability
-
🚀 Behold, the ultimate library for the technical leader who can’t lead without a script! 🌟 With over 1,000 #resources, you can now master the art of telling others what to do while pretending to manage incidents like a pro. 📚 Perfect for those who need a step-by-step guide to breathe in the world of tech leadership. 🦆
https://debuggingleadership.com/stdlib #techleadership #managementguide #incidentmanagement #leadershipskills #HackerNews #ngated -
Agile ITSM turns rigid processes into rapid value—what’s your next move? #AgileITSM #DigitalTransformation #ITLeadership #ModernIT #DevOps #ITOps #AgileMindset #ServiceExcellence #IncidentManagement #ContinuousImprovement #Automation #SelfService #Collaboration #Swarming #MTTR #MTTD #Metrics #Innovation #CustomerSatisfaction
https://medium.com/@sanjay.mohindroo66/beyond-the-ticket-agile-itsm-for-speed-clarity-and-impact-550a98882cb1 -
In August 2020, @SchizoDuckie and I published what was to become the first of a series of articles or posts called "No Need to Hack When It's Leaking."
In today's installment, I bring you "No Need to Hack When It's Leaking: Brandt Kettwick Defense Edition." It chronicles efforts by @JayeLTee, @masek, and I to alert a Minnesota law firm to lock down their exposed files, some of which were quite sensitive.
Read the post and see how even the state's Bureau of Criminal Apprehension had trouble getting this law firm to respond appropriately.
Great thanks to the Minnesota Bureau of Criminal Apprehension for their help on this one, and to @TonyYarusso and @bkoehn for their efforts.
#dataleak #misconfiguration #incidentresponse #incidentmanagement #responsibledisclosure #securityalert #infosec
-
The Information and Privacy Commissioner of Ontario has completed a review into Daixin Team's massive cyberattack on five regional hospitals in 2023 and found hospital officials acted “adequately.”
Perhaps the most notable aspect of the report (from my perspective) was that the IPC said the hospitals were obligated to notify patients whose data had been encrypted (and not just those whose data had been exfiltrated). They saw no point in requiring that now, but wanted it noted that it should have happened.
So that seems to be making PHIPA's interpretation clearer for future victims of encryption incidents.
The full report makes an interesting read.
PHIPA Decision 284:
https://decisions.ipc.on.ca/ipc-cipvp/phipa/en/item/521986/index.do#PHIPA #notification #incidentmanagement #databreach #ransomware
-
🚨 Cyber threats are evolving fast! 74% of CISOs are increasing their crisis simulation budgets in 2025 to stay ahead. With high-profile breaches on the rise, organizations must test and refine their response strategies.
At RELIANOID, we provide the tools to enhance cyber resilience and ensure businesses are always prepared. 🛡️
#CyberSecurity #CrisisResponse #IncidentManagement #CISO #RELIANOID
https://www.relianoid.com/blog/cisos-are-increasing-crisis-simulation-budgets/ -
Bradford Health Systems detected abnormal network activity in December 2023. They first sent out breach notices this week.
#databreach #ransomware #IncidentManagement #disclosure #transparency #healthsec #HIPAA
-
"If you focus too narrowly on preventing the specific details of the last incident, you’ll fail to identify the more general patterns that will enable your future incidents."
Great blog post from @norootcause
-
B.C. health authority faces class-action lawsuit over 2009 data breach
Let's see... they didn't prevent breaches, they didn't detect breaches on their own, and they didn't notify 20,000 employees timely or provide any mitigation services timely or at all.
But can plaintiffs prevail?
#databreach #infosec #cybersecurity #incidentmanagement #litigation
-
Mastering #TelemetryPipelines ensures high #ApplicationPerformance, cost efficiency, and security compliance. Implement best practices and stay ahead in #Observability & #Monitoring. #CloudComputing #DevOps #AI #Cybersecurity #ITGovernance #DigitalTransformation #DataAnalytics #Logging #IncidentManagement
https://medium.com/@sanjay.mohindroo66/how-to-use-telemetry-pipelines-to-maintain-application-performance-9d0972585d81 -
Just blogged: The Opiates of Root Cause and Counterfactual Reasoning