home.social

#sre — Public Fediverse posts

Live and recent posts from across the Fediverse tagged #sre, aggregated by home.social.

fetched live
  1. DevOpsDays Lima 2026 is coming!

    📅 August 27–28
    📍 Lima, Peru

    Two days focused on DevOps, cloud-native, platform engineering, automation, DevSecOps and SRE — bringing together the Latin American technology community.

    RELIANOID will be there to connect and talk about application delivery, load balancing, security and resilience.

    🚀 Attending? Let’s connect!

    🔗 [relianoid.com/about-us/events/

  2. 10 Linux Server Disasters & Open-Source SRE Cures 🚨

    SRE Blueprint:
    • OOM Shield: Systemd OOMScoreAdjust=-1000 prevents DB kills.
    • Inode Hogs: Find missing storage safely with du --inodes -xS /.
    • iowait Chokes: Diagnose disk saturation using iostat -xz 1.
    • Security: Replace Fail2Ban with CrowdSec's collaborative IPS.
    • Observability: Replace Datadog with Prometheus & Grafana Loki.

    servermo.com/blogs/linux-serve

    #Linux #DevOps #SRE #SysAdmin #ServerMO

  3. Как я перестал гадать, какая пара нод сломалась после апдейта CNI, и написал для этого свой мониторинг

    98% успешных проверок в дашборде выглядят прекрасно. Ровно до момента, когда доходит: недостающие 2% это одна пара нод, один протокол, и так каждый день. Обновили CNI или ядро, и между двумя конкретными нодами начал теряться UDP, а агрегат всё ещё зелёный. kconmon‑ng ставит агента на каждую ноду и гоняет TCP, UDP, ICMP, DNS и HTTP‑пробы между всеми парами каждые 5 секунд, а при сбое сам снимает MTR‑трейс, пока проблема ещё жива. К версии 2.0.0 у проекта выросла веб‑консоль: матрица N×N, расследование с ранжированием причин без ML, машина времени для разбора ночных инцидентов и алертинг, который сводит правила в настоящий PrometheusRule. Под капотом всё тот же Prometheus.

    habr.com/ru/articles/1072410/

    #kubernetes #monitoring #sre #devops #go #observability #opensource

  4. 🚨LIVE NOW!🚨 DevOps/SRE Instructor Livestream

    On this lovely Wednesday, let's chat about #Linux #SystemAdministration, #SelfHosting, or any other topic in the #DevOps and #SRE space you're interested in!

    Owncast: live.monospacementor.com/

  5. Apoyo consular a familias afectadas. 🇲🇽🤝 La SRE coordina los trámites para repatriar a México los restos de dos víctimas del sismo en Pereira, garantizando acompañamiento integral a sus familiares. Conoce más aquí. 👇 #Noticias #SRE #Internacional
    zurl.co/ksctG

  6. I'd like my next job to have something more than Linux, yes, I'm looking at your FreeBSD.

    #runbsd #freebsd #getfedihired

    ↩ Ludovic Hirlimann :
    "I can work in both French and English. I'm interested in infrastructure roles.

    #infrastructure #sysadmin #sre"

    [via Ponos] ponos.fr/thread/clf364ichbjbh9

  7. I can work in both French and English. I'm interested in infrastructure roles.

    #infrastructure #sysadmin #sre

    ↩ Ludovic Hirlimann :
    "I've been remote, fully since 2009 and would like to keep it that way.

    #fullremote"

    [via Ponos] ponos.fr/thread/c615qh73u2fb87

  8. 🇲🇽🕊️ #Noticias | La Cancillería mexicana brinda asistencia consular y gestiona la repatriación de los cuerpos de los dos connacionales fallecidos tras el terremoto en Pereira, Colombia. 🖤✈️ #SRE #AtenciónConsular #México #Colombia
    zurl.co/q46Ij

  9. 🇲🇽🕊️ #Noticias | La presidenta Claudia Sheinbaum confirmó el fallecimiento de dos mexicanos tras el terremoto de 7.4 en Colombia. La SRE brinda acompañamiento e interlocución a sus familiares para la atención requerida. 🖤🇲🇽 #Colombia #Sheinbaum #SRE
    zurl.co/L6bCe

  10. DevOps'ish 322: Linux wireless shuts the door on AI slop patches, KYAML fixes the Norway Bug, and more
    The Linux wireless maintainer stops arguing with LLMs, KYAML gets kubectl to quit lying about Norway, Charity Majors says the skepticism window has closed, and five AI rivals agree on a plugin format that punts on permissions. devopsish.com/322/ #DevOps #Kubernetes #SRE #AI

  11. CNCF Campinas SP - Meetup Presencial

    Laboratório Hacker de Campinas, quarta-feira, 30 de setembro às 18:00 BRT

    🚨 Campinas e Região, prepare-se. Tem coisa nova chegando na comunidade Cloud Native!

    No dia 30 de setembro, quarta-feira, às 19h, a CNCF Campinas SP volta a se reunir presencialmente.

    E dessa vez, em um novo espaço:

    📍 LHC - Laboratório Hacker de Campinas

    👀 Quem serão os palestrantes?
    🤔 Quais serão os temas?
    🔥 O que estamos preparando?

    Por enquanto, vamos deixar algumas respostas no ar...

    Mas se você gosta de Cloud Native, Kubernetes, DevOps, Platform Engineering, SRE, IA e tecnologia, já pode marcar essa data na agenda.

    Uma coisa podemos garantir: vai valer a pena estar presente.

    E tem um detalhe importante que precisamos compartilhar com a comunidade.

    ⚠️ Estamos enfrentando cerca de 50% de NO SHOW nos nossos eventos presenciais.

    Isso significa que metade das pessoas que garantem uma vaga acaba não comparecendo.

    Por isso, fazemos um pedido especial:

    👉 Inscreva-se somente se realmente puder comparecer.

    Sua presença faz diferença. Para a comunidade, para os palestrantes, para o espaço que nos recebe e para que possamos continuar trazendo encontros cada vez melhores para Campinas.

    📅 30 de setembro - quarta-feira
    🕖 19h
    📍 LHC - Laboratório Hacker de Campinas

    🎟️ Em breve divulgaremos os detalhes e o link de inscrição.

    👀 Fique de olho.

    Porque talvez essa seja uma daquelas noites que você não vai querer perder.

    CNCF Campinas SP - comunidade, conhecimento e tecnologia construídos juntos. 🚀

    #CNCFCampinas #CNCF #CloudNative #Kubernetes #DevOps #PlatformEngineering #SRE #AI #Campinas #ComunidadeTech

    eventos.lhc.net.br/event/cncf-

  12. Как мы добавили Global Orchestration в IncidentRelay: маршрутизация алертов до создания инцидента

    Когда система мониторинга отправляет алерт, он редко выглядит так, как хотелось бы дежурному инженеру. Один источник пишет critical , другой — disaster , третий передаёт severity: 5 . В одном payload сервис называется checkout , в другом — payments-api , а в третьем его приходится угадывать по namespace. Если все эти события приходят через общий Alertmanager или webhook, одной статической настройки маршрута быстро становится недостаточно. Обычно проблему решают одним из трёх способов:

    habr.com/ru/articles/1070294/

    #incident_management #alert_routing #oncall #SRE #DevOps #open_source #selfhosted #IncidentRelay

  13. На вкус и цвет все фломастеры разные — способы управления средами на личном опыте

    Знаешь , я прочитал ту статью про инфраструктурные релизы, идентичность и консистентность. Красиво написано, складно. Прямо как в душноватом учебнике. И даже местами правильно. Чего там про фломастеры ...

    habr.com/ru/articles/1070324/

    #инфраструктура #среды_разработки #devops #sre

  14. 🚨LIVE NOW!🚨 DevOps/SRE Instructor Livestream

    On this lovely Thursday, let's chat about #Linux #SystemAdministration, #SelfHosting, or any other topic in the #DevOps and #SRE space you're interested in!

    Owncast: live.monospacementor.com/

  15. Telling an agent to "always ask before acting" is a suggestion, and the model is free to weigh it against a well-crafted log line. An autonomy level you can trust lives in the write path: a bounded verb list instead of shell access, the agent running as its own scoped identity under RBAC, destructive verbs flagged in the tool definition, investigation and execution split into separate sessions.

    Roy Libman on autonomy levels for Kubernetes agents:
    radarhq.io/blog/ai-agent-auton
    #DevOps #AI #SRE

  16. Как мониторить Java-приложения: метрики, алерты и правило 80/20

    Хороший мониторинг помогает быстро понять, что происходит с приложением и куда смотреть в первую очередь. Для этого не нужно пытаться измерить всё: базовый набор технических метрик покрывает большинство типовых проблем, а бизнес-метрики, SLO и анализ аномалий помогают заранее замечать нетипичные отклонения. В Календаре мы называем этот подход правилом 80/20 . Всем привет! Меня зовут Настя, я бэкенд-разработчик в Яндекс 360 и отвечаю за надёжность Календаря. В этой статье я покажу, какие метрики стоит взять за основу, как выбирать полезные алерты и чем дополнять базовый набор для оставшихся 20%.

    habr.com/ru/companies/yandex/a

    #мониторинг #метрики #алертинг #java #devops #sre #site_reliability_engineering #надежность #инцидентменеджмент

  17. Hi everyone, my husband Craig is currently looking for a new job. Please feel free to boost this post for me 💐

    "I'm open for sysadmin, devops, SRE, or general coding work, at any level. Prefer to work *with* opensource, and in an ideal world *on* opensource, but work is work. Nearly 3 decades of experience across development and systems, from the tiny to the large (most recently GitLab). Permanent or contract is fine, working (remote) from Dunedin."

    CV at stroppykitten.com/static/Craig

    Email: [email protected]

    #GetFediContracted #GetFediHired #FediHire #fedijobs #remotejob #jobsearch #job #RemoteJobs #sre #devops #sysadmin #opensource

  18. I will not debate populist AI feelings, you're entitled to them and I respect that. However, it's arrived and these tools are in use. A new set of theory crafting is required for now and this is my start to deal with it: webframp.com/posts/six-princip

    #ai #devops #sre #swamp #manifesto

  19. I finally have the Invirtuate tofu/terraform provider mostly complete in my homelab. I need to add more update functionality for resources, but create, destroy, read, state import, and data sources all work now.

    Pictured are tofu plan, tofu apply, and tofu destroy.

    #technology #terraform #tofu #automation #homelab #selfhosted #devops #sre