#etcd — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #etcd, aggregated by home.social.
-
I asked this question over in the docker forum as well: https://forums.docker.com/t/how-to-run-etcd-cluster-as-service-with-docker-swarm-mode/152440
-
I asked this question over in the docker forum as well: https://forums.docker.com/t/how-to-run-etcd-cluster-as-service-with-docker-swarm-mode/152440
-
I asked this question over in the docker forum as well: https://forums.docker.com/t/how-to-run-etcd-cluster-as-service-with-docker-swarm-mode/152440
-
I asked this question over in the docker forum as well: https://forums.docker.com/t/how-to-run-etcd-cluster-as-service-with-docker-swarm-mode/152440
-
Happy weekend dear fellow fedizens :hamsterdance:
Could someone help me figure out a specific #docker #dockerswarm networking thing?
I want to deploy a service with multiple replicas on a docker swarm mode cluster.
The container replicas / tasks need to be able to reach each other individually at predictable/templateable urls, without manual configuration.I'm trying to make a #coopcloud recipe for #etcd (distributed key-value store). This will only be useful for multi-node co-op cloud, but would facilitate e.g. high availability postgresql via Patroni, and other interesting distributed computing stuffs.
[Here's](https://codeberg.org/papiris/etcd-recipe) the draft of the recipe, showing my attempts to template. In its current form, the individual replicas are reachable via `<service_name>.<task_slot (aka replica nr)>.<random hash which changes every container restart>`, but I don't know how to template the hash part or make a more stable alias per replica.
For a busybox service with multiple replicas, setting hostname like this in compose.yml lets the replicas resolve each other by e.g. busybox.1 ;
hostname: {{.Service.Name}}.{{.Task.Slot}}
But doing the same doesn't let the etcd replicas connect with each other... Does etcd need some special consideration?Advice appreciated!
#homelab #etcd #selfhosting #CollectiveHosting #distributedComputing #fediAsk
-
Happy weekend dear fellow fedizens :hamsterdance:
Could someone help me figure out a specific #docker #dockerswarm networking thing?
I want to deploy a service with multiple replicas on a docker swarm mode cluster.
The container replicas / tasks need to be able to reach each other individually at predictable/templateable urls, without manual configuration.I'm trying to make a #coopcloud recipe for #etcd (distributed key-value store). This will only be useful for multi-node co-op cloud, but would facilitate e.g. high availability postgresql via Patroni, and other interesting distributed computing stuffs.
[Here's](https://codeberg.org/papiris/etcd-recipe) the draft of the recipe, showing my attempts to template. In its current form, the individual replicas are reachable via `<service_name>.<task_slot (aka replica nr)>.<random hash which changes every container restart>`, but I don't know how to template the hash part or make a more stable alias per replica.
For a busybox service with multiple replicas, setting hostname like this in compose.yml lets the replicas resolve each other by e.g. busybox.1 ;
hostname: {{.Service.Name}}.{{.Task.Slot}}
But doing the same doesn't let the etcd replicas connect with each other... Does etcd need some special consideration?Advice appreciated!
#homelab #etcd #selfhosting #CollectiveHosting #distributedComputing #fediAsk
-
Happy weekend dear fellow fedizens :hamsterdance:
Could someone help me figure out a specific #docker #dockerswarm networking thing?
I want to deploy a service with multiple replicas on a docker swarm mode cluster.
The container replicas / tasks need to be able to reach each other individually at predictable/templateable urls, without manual configuration.I'm trying to make a #coopcloud recipe for #etcd (distributed key-value store). This will only be useful for multi-node co-op cloud, but would facilitate e.g. high availability postgresql via Patroni, and other interesting distributed computing stuffs.
[Here's](https://codeberg.org/papiris/etcd-recipe) the draft of the recipe, showing my attempts to template. In its current form, the individual replicas are reachable via `<service_name>.<task_slot (aka replica nr)>.<random hash which changes every container restart>`, but I don't know how to template the hash part or make a more stable alias per replica.
For a busybox service with multiple replicas, setting hostname like this in compose.yml lets the replicas resolve each other by e.g. busybox.1 ;
hostname: {{.Service.Name}}.{{.Task.Slot}}
But doing the same doesn't let the etcd replicas connect with each other... Does etcd need some special consideration?Advice appreciated!
#homelab #etcd #selfhosting #CollectiveHosting #distributedComputing #fediAsk
-
Happy weekend dear fellow fedizens :hamsterdance:
Could someone help me figure out a specific #docker #dockerswarm networking thing?
I want to deploy a service with multiple replicas on a docker swarm mode cluster.
The container replicas / tasks need to be able to reach each other individually at predictable/templateable urls, without manual configuration.I'm trying to make a #coopcloud recipe for #etcd (distributed key-value store). This will only be useful for multi-node co-op cloud, but would facilitate e.g. high availability postgresql via Patroni, and other interesting distributed computing stuffs.
[Here's](https://codeberg.org/papiris/etcd-recipe) the draft of the recipe, showing my attempts to template. In its current form, the individual replicas are reachable via `<service_name>.<task_slot (aka replica nr)>.<random hash which changes every container restart>`, but I don't know how to template the hash part or make a more stable alias per replica.
For a busybox service with multiple replicas, setting hostname like this in compose.yml lets the replicas resolve each other by e.g. busybox.1 ;
hostname: {{.Service.Name}}.{{.Task.Slot}}
But doing the same doesn't let the etcd replicas connect with each other... Does etcd need some special consideration?Advice appreciated!
#homelab #etcd #selfhosting #CollectiveHosting #distributedComputing #fediAsk
-
GFM: децентрализованная альтернатива Patroni и etcd для PostgreSQL на базе P2P-меша с In-Memory управлением
Если вы когда-нибудь настраивали Patroni, то знаете это чувство: чтобы просто следить за одной базой, вам нужно построить вокруг неё целый город. Тут у нас etcd, там Consul, здесь зоопарк зависимостей, а вон там отдельный бюджет на железо под систему управления. Это напоминает попытку установить охранную систему, которая потребляет больше электричества, чем защищаемый объект. Классика HA строится на «централизованном консенсусе». Это когда три сервера управления спорят между собой, кто из двух серверов базы данных сейчас главный. Я решил, что пора перестать плодить сущности. Познакомьтесь с GFM (Gorgona Failover Manager) . Это менеджер отказоустойчивости, который работает на принципах P2P-меша.
-
GFM: децентрализованная альтернатива Patroni и etcd для PostgreSQL на базе P2P-меша с In-Memory управлением
Если вы когда-нибудь настраивали Patroni, то знаете это чувство: чтобы просто следить за одной базой, вам нужно построить вокруг неё целый город. Тут у нас etcd, там Consul, здесь зоопарк зависимостей, а вон там отдельный бюджет на железо под систему управления. Это напоминает попытку установить охранную систему, которая потребляет больше электричества, чем защищаемый объект. Классика HA строится на «централизованном консенсусе». Это когда три сервера управления спорят между собой, кто из двух серверов базы данных сейчас главный. Я решил, что пора перестать плодить сущности. Познакомьтесь с GFM (Gorgona Failover Manager) . Это менеджер отказоустойчивости, который работает на принципах P2P-меша.
-
etcd для самых маленьких: гайд по хранилищу kubernetes
Разбираемся, что за зверь хранит весь ваш Kubernetes, учимся читать его ошибки и честно отвечаем на вопрос, можно ли ему доверять.
-
etcd для самых маленьких: гайд по хранилищу kubernetes
Разбираемся, что за зверь хранит весь ваш Kubernetes, учимся читать его ошибки и честно отвечаем на вопрос, можно ли ему доверять.
-
IP не валяй, или как я переобувал кластер kubernetes на ходу
В статье рассматривается опыт миграции сети Kubernetes-кластера на новый IP-пул (Pool IP) с использованием CNI Cilium. Особое внимание уделено процессу смены IP-адресов для Pod'ов и Service'ов в работающем кластере без длительной остановки, а также подводных камнях, с которыми пришлось столкнуться при обновлении до версии 1.20.0. Материал будет полезен инженерам по эксплуатации Kubernetes использующих уже Cilium, а так же желающим познакомиться с этим инструментом.
-
IP не валяй, или как я переобувал кластер kubernetes на ходу
В статье рассматривается опыт миграции сети Kubernetes-кластера на новый IP-пул (Pool IP) с использованием CNI Cilium. Особое внимание уделено процессу смены IP-адресов для Pod'ов и Service'ов в работающем кластере без длительной остановки, а также подводных камнях, с которыми пришлось столкнуться при обновлении до версии 1.20.0. Материал будет полезен инженерам по эксплуатации Kubernetes использующих уже Cilium, а так же желающим познакомиться с этим инструментом.
-
[Перевод] Может ли Kubernetes выдержать миллион узлов? Эксперимент, графики, выводы
Часто первое, что хочется сделать при неполадках в кластере, — уменьшить его и надеяться, что всё решится само. Получается, большой кластер = куча проблем? Автор статьи решил зайти максимально далеко: поднять Kubernetes-кластер с миллионом узлов и посмотреть, что будет. Спойлер: получился вовсе не выстрел в ногу, а серьёзный эксперимент — много подготовки, обход ограничений, графики производительности и, конечно, ценные выводы. И даже инструкция в конце для желающих повторить.
https://habr.com/ru/companies/flant/articles/1064962/
#kubernetes #etcd #kubeapiserver #кластер_на_миллион_узлов #TxnPut #lease #inmemory_etcd #resourceVersion #watchзапросы #gomemlimit
-
[Перевод] Может ли Kubernetes выдержать миллион узлов? Эксперимент, графики, выводы
Часто первое, что хочется сделать при неполадках в кластере, — уменьшить его и надеяться, что всё решится само. Получается, большой кластер = куча проблем? Автор статьи решил зайти максимально далеко: поднять Kubernetes-кластер с миллионом узлов и посмотреть, что будет. Спойлер: получился вовсе не выстрел в ногу, а серьёзный эксперимент — много подготовки, обход ограничений, графики производительности и, конечно, ценные выводы. И даже инструкция в конце для желающих повторить.
https://habr.com/ru/companies/flant/articles/1064962/
#kubernetes #etcd #kubeapiserver #кластер_на_миллион_узлов #TxnPut #lease #inmemory_etcd #resourceVersion #watchзапросы #gomemlimit
-
We've released #etcd v3.7.0! The new version includes RangeStream, multiple performance improvements, and runs entirely from v3store, and more: https://etcd.io/blog/2026/announcing-etcd-3.7/
-
We've released #etcd v3.7.0! The new version includes RangeStream, multiple performance improvements, and runs entirely from v3store, and more: https://etcd.io/blog/2026/announcing-etcd-3.7/
-
We've released #etcd v3.7.0! The new version includes RangeStream, multiple performance improvements, and runs entirely from v3store, and more: https://etcd.io/blog/2026/announcing-etcd-3.7/
-
We've released #etcd v3.7.0! The new version includes RangeStream, multiple performance improvements, and runs entirely from v3store, and more: https://etcd.io/blog/2026/announcing-etcd-3.7/
-
Latest etcd patch release is out. Aside from updating a few dependencies, mainly it fixes upgrades for folks with old custom v2store data.
-
Latest etcd patch release is out. Aside from updating a few dependencies, mainly it fixes upgrades for folks with old custom v2store data.
-
Latest etcd patch release is out. Aside from updating a few dependencies, mainly it fixes upgrades for folks with old custom v2store data.
-
Latest etcd patch release is out. Aside from updating a few dependencies, mainly it fixes upgrades for folks with old custom v2store data.
-
IP подов кончились, а обычные решения не подошли: как мы расширили сеть на проде, не пересоздавая кластер (кейс + гайд)
Штатная ситуация оказалась задачей со звёздочкой: кластер кинул алерт о том, что заканчивается сеть подов, но ни одно решение «из методички» не подходило, а вытаскивать кластер из прода было нельзя. В статье расскажу, как мы не просто расширили подсеть подов, но сделали это на работающем кластере и не потеряли при этом данные. Что важно — трюк сработает на любом дистрибутиве Kubernetes и CNI.
-
IP подов кончились, а обычные решения не подошли: как мы расширили сеть на проде, не пересоздавая кластер (кейс + гайд)
Штатная ситуация оказалась задачей со звёздочкой: кластер кинул алерт о том, что заканчивается сеть подов, но ни одно решение «из методички» не подходило, а вытаскивать кластер из прода было нельзя. В статье расскажу, как мы не просто расширили подсеть подов, но сделали это на работающем кластере и не потеряли при этом данные. Что важно — трюк сработает на любом дистрибутиве Kubernetes и CNI.
-
Проект Cozystack представил переработанный etcd-operator с новым API
В рамках проекта etcd-operator сообщество развивает оператор для развёртывания и сопровождения кластеров etcd в Kubernetes. На днях он был передан проекту Cozystack (CNCF Sandbox). Перед этим команда опубликовала написанную с нуля реализацию оператора с новой версией API — etcd-operator.cozystack.io/v1alpha2 . Эта версия пришла на смену etcd.aenix.io/v1alpha1 . Вместо управления узлами через StatefulSet новый оператор напрямую задействует штатный Membership API etcd (операции MemberAdd, MemberPromote и MemberRemove), что позволяет ему полностью контролировать состав кластера. Автор новой реализации — Тимофей Ларкин , один из мейнтейнеров прежнего оператора (старый код остался в ветке v1alpha1 ). Проект написан на Go и распространяется под лицензией Apache 2.0.
https://habr.com/ru/companies/aenix/articles/1047170/
#aenix #cozystack #devops #etcd #cncf #open_source #kubernetes #kubernetes_operator #kubernetes_cluster
-
Проект Cozystack представил переработанный etcd-operator с новым API
В рамках проекта etcd-operator сообщество развивает оператор для развёртывания и сопровождения кластеров etcd в Kubernetes. На днях он был передан проекту Cozystack (CNCF Sandbox). Перед этим команда опубликовала написанную с нуля реализацию оператора с новой версией API — etcd-operator.cozystack.io/v1alpha2 . Эта версия пришла на смену etcd.aenix.io/v1alpha1 . Вместо управления узлами через StatefulSet новый оператор напрямую задействует штатный Membership API etcd (операции MemberAdd, MemberPromote и MemberRemove), что позволяет ему полностью контролировать состав кластера. Автор новой реализации — Тимофей Ларкин , один из мейнтейнеров прежнего оператора (старый код остался в ветке v1alpha1 ). Проект написан на Go и распространяется под лицензией Apache 2.0.
https://habr.com/ru/companies/aenix/articles/1047170/
#aenix #cozystack #devops #etcd #cncf #open_source #kubernetes #kubernetes_operator #kubernetes_cluster
-
Как строить отказоустойчивые кластеры Kubernetes: краткий разбор от команды VK Cloud
Миграция в облако и переход к микросервисной архитектуре сделали Kubernetes (k8s) де-факто стандартом для управления контейнерами. По данным 2025 года, технологию уже применяют 60% крупных российских компаний, а ещё 15% планируют внедрение в будущем. Причем 59% компаний называют отказоустойчивость ключевым критерием при выборе Kubernetes, но лишь единицы реализуют его на практике. Проблема кроется в недооценке системных рисков — от отсутствия резервирования control plane до некорректных таймингов readiness-проб, пропускающих «полуживые» поды в балансировщик. В этой статье мы кратко разберем ключевые принципы проектирования и эксплуатации отказоустойчивых кластеров, типовые сценарии сбоев и рекомендации по исключению рисков на всех уровнях.
https://habr.com/ru/companies/vktech/articles/1042084/
#vk_cloud #kubernetes #отказоустойчивость #high_availability #devops #etcd #storage #statefulset #gitops #backup
-
Как строить отказоустойчивые кластеры Kubernetes: краткий разбор от команды VK Cloud
Миграция в облако и переход к микросервисной архитектуре сделали Kubernetes (k8s) де-факто стандартом для управления контейнерами. По данным 2025 года, технологию уже применяют 60% крупных российских компаний, а ещё 15% планируют внедрение в будущем. Причем 59% компаний называют отказоустойчивость ключевым критерием при выборе Kubernetes, но лишь единицы реализуют его на практике. Проблема кроется в недооценке системных рисков — от отсутствия резервирования control plane до некорректных таймингов readiness-проб, пропускающих «полуживые» поды в балансировщик. В этой статье мы кратко разберем ключевые принципы проектирования и эксплуатации отказоустойчивых кластеров, типовые сценарии сбоев и рекомендации по исключению рисков на всех уровнях.
https://habr.com/ru/companies/vktech/articles/1042084/
#vk_cloud #kubernetes #отказоустойчивость #high_availability #devops #etcd #storage #statefulset #gitops #backup
-
Friends don't let friends run production etcd on SATA disks (or basically anywhere other than local NVMe)!
I was deploying 50 kubernetes virtual clusters (with vcluster) on top of a bare metal cluster running Talos.
All the nodes had NVMe disks, except one, which had SATA SSD... And it did not go well at all.
I was expecting to see a difference, but not that big. The SATA disks were saturated while the NVMe were hovering between 1-5% of I/O load.
The SATA disks were so overloaded, that I had to kick out their node from the etcd cluster (because the API server was extremely slow).
Oh well, it'll be 100% NVMe on this cluster from now on!
-
Friends don't let friends run production etcd on SATA disks (or basically anywhere other than local NVMe)!
I was deploying 50 kubernetes virtual clusters (with vcluster) on top of a bare metal cluster running Talos.
All the nodes had NVMe disks, except one, which had SATA SSD... And it did not go well at all.
I was expecting to see a difference, but not that big. The SATA disks were saturated while the NVMe were hovering between 1-5% of I/O load.
The SATA disks were so overloaded, that I had to kick out their node from the etcd cluster (because the API server was extremely slow).
Oh well, it'll be 100% NVMe on this cluster from now on!
-
Friends don't let friends run production etcd on SATA disks (or basically anywhere other than local NVMe)!
I was deploying 50 kubernetes virtual clusters (with vcluster) on top of a bare metal cluster running Talos.
All the nodes had NVMe disks, except one, which had SATA SSD... And it did not go well at all.
I was expecting to see a difference, but not that big. The SATA disks were saturated while the NVMe were hovering between 1-5% of I/O load.
The SATA disks were so overloaded, that I had to kick out their node from the etcd cluster (because the API server was extremely slow).
Oh well, it'll be 100% NVMe on this cluster from now on!
-
Friends don't let friends run production etcd on SATA disks (or basically anywhere other than local NVMe)!
I was deploying 50 kubernetes virtual clusters (with vcluster) on top of a bare metal cluster running Talos.
All the nodes had NVMe disks, except one, which had SATA SSD... And it did not go well at all.
I was expecting to see a difference, but not that big. The SATA disks were saturated while the NVMe were hovering between 1-5% of I/O load.
The SATA disks were so overloaded, that I had to kick out their node from the etcd cluster (because the API server was extremely slow).
Oh well, it'll be 100% NVMe on this cluster from now on!
-
Ваш Kubernetes упал: найдёте root cause за 15 минут?
Вторник, 14:00. Кластер Kubernetes перестал отвечать, команда в панике, а вам нужно за 15 минут найти первопричину. В этой статье пройдём диагностику реального отказа вместе с SRE: увидим логи, манифест etcd и ошибки, которые совершают даже опытные инженеры. Попробуйте сначала решить задачу сами, а потом сверьтесь с пошаговым разбором и проверьте, насколько вы готовы к такому инциденту.
https://habr.com/ru/companies/otus/articles/1031260/
#Kubernetes #etcd #kubelet #SRE #DevOps #productionинцидент #отказ_кластера #root_cause #control_plane #runbook
-
Ваш Kubernetes упал: найдёте root cause за 15 минут?
Вторник, 14:00. Кластер Kubernetes перестал отвечать, команда в панике, а вам нужно за 15 минут найти первопричину. В этой статье пройдём диагностику реального отказа вместе с SRE: увидим логи, манифест etcd и ошибки, которые совершают даже опытные инженеры. Попробуйте сначала решить задачу сами, а потом сверьтесь с пошаговым разбором и проверьте, насколько вы готовы к такому инциденту.
https://habr.com/ru/companies/otus/articles/1031260/
#Kubernetes #etcd #kubelet #SRE #DevOps #productionинцидент #отказ_кластера #root_cause #control_plane #runbook
-
We've released #etcd v3.7.0-beta.0!
https://etcd.io/blog/2026/etcd-370-beta/
This release includes RangeStream queries and more.
It also represents several milestones for our project: second regular annual release, our first beta in years, and the addition of long-requested user-visible features instead of just focusing on stability.
Please test it out and let us know how it works for you!
-
We've released #etcd v3.7.0-beta.0!
https://etcd.io/blog/2026/etcd-370-beta/
This release includes RangeStream queries and more.
It also represents several milestones for our project: second regular annual release, our first beta in years, and the addition of long-requested user-visible features instead of just focusing on stability.
Please test it out and let us know how it works for you!
-
We've released #etcd v3.7.0-beta.0!
https://etcd.io/blog/2026/etcd-370-beta/
This release includes RangeStream queries and more.
It also represents several milestones for our project: second regular annual release, our first beta in years, and the addition of long-requested user-visible features instead of just focusing on stability.
Please test it out and let us know how it works for you!
-
We've released #etcd v3.7.0-beta.0!
https://etcd.io/blog/2026/etcd-370-beta/
This release includes RangeStream queries and more.
It also represents several milestones for our project: second regular annual release, our first beta in years, and the addition of long-requested user-visible features instead of just focusing on stability.
Please test it out and let us know how it works for you!
-
Распределенное KV-хранилище на базе etcd
Я постараюсь, не углубляясь в технические дебри, в научно-популярном ключе рассказать о распределенных KV-хранилищах: что это вообще такое, где применяется и почему мы выбрали именно etcd.
-
Распределенное KV-хранилище на базе etcd
Я постараюсь, не углубляясь в технические дебри, в научно-популярном ключе рассказать о распределенных KV-хранилищах: что это вообще такое, где применяется и почему мы выбрали именно etcd.
-
etcd operator 0.2 has been released!
https://etcd.io/blog/2026/announcing-etcd-operator-v0.2.0/
This now makes the operator useful for production, or at least staging, use-cases, and brings it up to the functionality of the old operator -- plus better handling of TLS.
Take a look!
-
etcd operator 0.2 has been released!
https://etcd.io/blog/2026/announcing-etcd-operator-v0.2.0/
This now makes the operator useful for production, or at least staging, use-cases, and brings it up to the functionality of the old operator -- plus better handling of TLS.
Take a look!
-
etcd operator 0.2 has been released!
https://etcd.io/blog/2026/announcing-etcd-operator-v0.2.0/
This now makes the operator useful for production, or at least staging, use-cases, and brings it up to the functionality of the old operator -- plus better handling of TLS.
Take a look!
-
etcd operator 0.2 has been released!
https://etcd.io/blog/2026/announcing-etcd-operator-v0.2.0/
This now makes the operator useful for production, or at least staging, use-cases, and brings it up to the functionality of the old operator -- plus better handling of TLS.
Take a look!
-
⚠️ NEW: Kubernetes Swap & etcd Stability!
Prevent control plane hangs with proper swap configuration. etcd performance tuning & swapfile best practices for production K8s.
-
Swap on K8s nodes? Containers hang instead of OOM-killing—etcd suffers, control plane cascades. 2 fixes: resource limits + etcd HAProxy LB. Protect your cluster! 👇
https://devopstales.github.io/kubernetes/k8s-swap-etcd-stability/
-
What reason would there be to seperate etcd out of the Kubernetes manifests of the control plane nodes but keep it as a native service installed on the same machines running the control plane?
There's nothing to gain in terms of high availability there.
You still have X amount of control plane nodes that also run etcd as cluster nodes.
I'm trying to figure out what my predecessor thought while building this Kubernetes environment.
The two etcd topologies mentioned in official K8S docs are:
- integrated etcd (etcd as a Kubernetes Manifest, started as containers together with coredns, kube-apiserver and so on)
- seperated etcd nodes (X amount of machines that host etcd as a native service on the OS and the control plane is configured to use them.
-
In my (short) dad time this morning, I've tried to install mgmt [1] to run a distributed hello world on my main machine running on Ubuntu LTS. The built-in binaries depend on augeas which was easy to fix. But also libvirt which is surprisingly old on Ubuntu compared to Debian (latest). I tried to build it myself but I couldn't install nex (the lexer). I then built the binary using Docker thanks to the quick start guide.
I first started to run mgmt in standalone mode. It's nice to see etcd embedded in the binary (at least for testing). Then I tried to deploy multi mgmt nodes with a standalone etcd using docker-compose. I've lost a lot of time trying to override the command because I didn't remember the expected syntax.
I was trying to make etcd listen to all interfaces so mgmt could connect when my daughter showed up.
[1] https://github.com/purpleidea/mgmt (@purpleidea)
#mgmt #homelab #selfhosting #etcd #docker #libvirt #ubuntu #debian
-
In my (short) dad time this morning, I've tried to install mgmt [1] to run a distributed hello world on my main machine running on Ubuntu LTS. The built-in binaries depend on augeas which was easy to fix. But also libvirt which is surprisingly old on Ubuntu compared to Debian (latest). I tried to build it myself but I couldn't install nex (the lexer). I then built the binary using Docker thanks to the quick start guide.
I first started to run mgmt in standalone mode. It's nice to see etcd embedded in the binary (at least for testing). Then I tried to deploy multi mgmt nodes with a standalone etcd using docker-compose. I've lost a lot of time trying to override the command because I didn't remember the expected syntax.
I was trying to make etcd listen to all interfaces so mgmt could connect when my daughter showed up.
[1] https://github.com/purpleidea/mgmt (@purpleidea)
#mgmt #homelab #selfhosting #etcd #docker #libvirt #ubuntu #debian
-
In my (short) dad time this morning, I've tried to install mgmt [1] to run a distributed hello world on my main machine running on Ubuntu LTS. The built-in binaries depend on augeas which was easy to fix. But also libvirt which is surprisingly old on Ubuntu compared to Debian (latest). I tried to build it myself but I couldn't install nex (the lexer). I then built the binary using Docker thanks to the quick start guide.
I first started to run mgmt in standalone mode. It's nice to see etcd embedded in the binary (at least for testing). Then I tried to deploy multi mgmt nodes with a standalone etcd using docker-compose. I've lost a lot of time trying to override the command because I didn't remember the expected syntax.
I was trying to make etcd listen to all interfaces so mgmt could connect when my daughter showed up.
[1] https://github.com/purpleidea/mgmt (@purpleidea)
#mgmt #homelab #selfhosting #etcd #docker #libvirt #ubuntu #debian
-
In my (short) dad time this morning, I've tried to install mgmt [1] to run a distributed hello world on my main machine running on Ubuntu LTS. The built-in binaries depend on augeas which was easy to fix. But also libvirt which is surprisingly old on Ubuntu compared to Debian (latest). I tried to build it myself but I couldn't install nex (the lexer). I then built the binary using Docker thanks to the quick start guide.
I first started to run mgmt in standalone mode. It's nice to see etcd embedded in the binary (at least for testing). Then I tried to deploy multi mgmt nodes with a standalone etcd using docker-compose. I've lost a lot of time trying to override the command because I didn't remember the expected syntax.
I was trying to make etcd listen to all interfaces so mgmt could connect when my daughter showed up.
[1] https://github.com/purpleidea/mgmt (@purpleidea)
#mgmt #homelab #selfhosting #etcd #docker #libvirt #ubuntu #debian
-
Эволюция сбора flow-статистики в Яндексе: архитектура, грабли и оптимизации
Привет, Хабр! На связи Саша Лопинцев, SRE в группе разработки сетевой инфраструктуры и мониторинга Yandex Infrastructure. Я очень люблю мониторинг — а когда дело касается видимости сетевого трафика, нам не обойтись без анализа flow‑данных. Сегодня расскажу, как и почему мы переехали с устаревшего flow‑коллектора на GoFlow2, реализовали запись в БД и через etcd решили проблемы с шаблонами. Новая система обрабатывает 85 тысяч пакетов статистики в секунду, обеспечивает отказоустойчивость и помогает создавать отчёты. Если вам интересно узнать чуть больше об архитектуре, экспериментах, ошибках и решениях, полезных для инфраструктурного мониторинга в продакшн‑среде, читайте далее.
-
Эволюция сбора flow-статистики в Яндексе: архитектура, грабли и оптимизации
Привет, Хабр! На связи Саша Лопинцев, SRE в группе разработки сетевой инфраструктуры и мониторинга Yandex Infrastructure. Я очень люблю мониторинг — а когда дело касается видимости сетевого трафика, нам не обойтись без анализа flow‑данных. Сегодня расскажу, как и почему мы переехали с устаревшего flow‑коллектора на GoFlow2, реализовали запись в БД и через etcd решили проблемы с шаблонами. Новая система обрабатывает 85 тысяч пакетов статистики в секунду, обеспечивает отказоустойчивость и помогает создавать отчёты. Если вам интересно узнать чуть больше об архитектуре, экспериментах, ошибках и решениях, полезных для инфраструктурного мониторинга в продакшн‑среде, читайте далее.
-
etcd và Consul giới hạn kích thước giá trị để tránh tắc nghẽn, nhưng dữ liệu hiện đại (vector AI, JSON khổng lồ) thường vượt quá. UnisonDB đề xuất WAL dạng đồ thị liên kết ngược (linked‑list) cho phép ghi lớn mà không làm giảm hiệu năng replication, heartbeat và bầu cử leader. #KV #Database #UnisonDB #etcd #Consul #AI #CôngNghệ
-
Finally completed the upgrade of all of the five @midgaard #Kubernetes nodes to what I colloquially refer to as midgaard-v3.
Same standard #Hetzner nodes with a couple of #NVMe sticks and around 10TB of spinning metal for bulk storage. Much simplified partition layout. NVMes used for caching. Scheduled backups of #etcd. Continuous SMART disk testing and reporting.
The four other nodes have been running smoothly for about a month so I don't expect any surprises at this point. Distro is #Debian Trixie. Kubernetes is version 1.33.
All is well.
So far.
-
Finally completed the upgrade of all of the five @midgaard #Kubernetes nodes to what I colloquially refer to as midgaard-v3.
Same standard #Hetzner nodes with a couple of #NVMe sticks and around 10TB of spinning metal for bulk storage. Much simplified partition layout. NVMes used for caching. Scheduled backups of #etcd. Continuous SMART disk testing and reporting.
The four other nodes have been running smoothly for about a month so I don't expect any surprises at this point. Distro is #Debian Trixie. Kubernetes is version 1.33.
All is well.
So far.
-
Finally completed the upgrade of all of the five @midgaard #Kubernetes nodes to what I colloquially refer to as midgaard-v3.
Same standard #Hetzner nodes with a couple of #NVMe sticks and around 10TB of spinning metal for bulk storage. Much simplified partition layout. NVMes used for caching. Scheduled backups of #etcd. Continuous SMART disk testing and reporting.
The four other nodes have been running smoothly for about a month so I don't expect any surprises at this point. Distro is #Debian Trixie. Kubernetes is version 1.33.
All is well.
So far.
-
Finally completed the upgrade of all of the five @midgaard #Kubernetes nodes to what I colloquially refer to as midgaard-v3.
Same standard #Hetzner nodes with a couple of #NVMe sticks and around 10TB of spinning metal for bulk storage. Much simplified partition layout. NVMes used for caching. Scheduled backups of #etcd. Continuous SMART disk testing and reporting.
The four other nodes have been running smoothly for about a month so I don't expect any surprises at this point. Distro is #Debian Trixie. Kubernetes is version 1.33.
All is well.
So far.