#perf — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #perf, aggregated by home.social.
-
Добавили потоков, стало медленнее: разбираемся с ложным разделением кеш‑линии
Добавили потоков, а программа стала медленнее? Причина может скрываться в ложном разделении кеш‑линий: потоки работают с разными переменными, но процессор всё равно заставляет ядра конкурировать за один участок памяти. Разберём, как воспроизвести такую просадку, найти её через perf c2c и исправить без лишнего раздувания структур.
https://habr.com/ru/companies/otus/articles/1067152/
#ложное_разделение_кешлинии #false_sharing #многопоточность #кеш_процессора #кешлиния #когерентность_кеша #производительность #perf #NUMA #выравнивание_памяти
-
Track page faults per process with perf stat and uprobes. Attach probes to mmap/munmap in libc, use --per-thread for per-thread breakdown. Requires debug symbols. #linux #snippet #perf #ValtersIT
https://www.valtersit.com/vault/page-fault-distribution-by-process-with-perf-stat-and-uprobe-25930b/
-
[Перевод] Перевополщение Stable Values в JDK 26
В новом переводе от команды Spring АйО рассмотрим ленивую инициализацию в Java , которая почти всегда значит: поле сначала null , потом double-checked locking, volatile, синхронизация. Ошибиться легко, а final не поставить. Итог - код хрупче и JVM хуже делает constant folding. В JDK 26 (preview, JEP 526) добавили LazyConstant<T> : final поле, рецепт вычисления через Supplier , значение берёте login.get() . Supplier выполнится при первом get и только один раз успешно, даже при гонке потоков. Кроме этого значение помечается как @Stable - JVM может считать его константой и агрессивнее оптимизировать. Граничные случаи: null нельзя; не сериализуется; исключение из Supplier пробросится и следующая попытка снова пересчитает; equals у LazyConstant - только identity. Для 1:n есть List.ofLazy и Map.ofLazy : элементы/значения считаются по индексу/ключу по требованию и кэшируются.
https://habr.com/ru/companies/spring_aio/articles/1042294/
#java #kotlin #jdk #jdk_26 #perf #performance #performance_optimization
-
Lancez votre magasin de jeux vidéo avec ce guide exclusif et simplifié ! Lancez et développez avec succès votre magasin de vente au détail de jeux vidéo est votre guide incontournable pour transformer votre passion en entreprise lucrative ! Découvrez des stratégies de développement robustes, des astuces SEO, et des techniques d’expérience utilisateur. Lancez-vous maintenant et propulsez votre magasin vers le succès !
Tags : #Ebook #Livres #JeuxVidéo #OptimisationSEO #StratégiesDeContenu #Perf... -
If you want to know who's taller, you don't measure people hours apart with a precise ruler - you line them up side by side
Denis Bazhenov (JetBrains) applies the same logic to microbenchmarking: instead of running implementations separately and comparing results, run them simultaneously on the same machine. Background noise affects both equally, and you measure relative performance directly.
🔗 https://oxidizeconf.com/sessions/just_stand_them_next_to_each_other
#Oxidize2026 #RustLang #Benchmarking #Perf #SystemsProgramming
-
If you want to know who's taller, you don't measure people hours apart with a precise ruler - you line them up side by side
Denis Bazhenov (JetBrains) applies the same logic to microbenchmarking: instead of running implementations separately and comparing results, run them simultaneously on the same machine. Background noise affects both equally, and you measure relative performance directly.
🔗 https://oxidizeconf.com/sessions/just_stand_them_next_to_each_other
#Oxidize2026 #RustLang #Benchmarking #Perf #SystemsProgramming
-
Lancez votre magasin de jeux vidéo avec ce guide exclusif et simplifié ! Lancez et développez avec succès votre magasin de vente au détail de jeux vidéo est votre guide incontournable pour transformer votre passion en entreprise lucrative ! Découvrez des stratégies de développement robustes, des astuces SEO, et des techniques d’expérience utilisateur. Lancez-vous maintenant et propulsez votre magasin vers le succès !
Tags : #Ebook #Livres #JeuxVidéo #OptimisationSEO #StratégiesDeContenu #Perf... -
Great perf deep-dive
"The Cost of Concurrency Coordination with Jon Gjengset"
https://www.youtube.com/watch?v=tND-wBBZ8RYmutex, caching, cache-lines, L1,L2, byte alignment
-
Hotspot v1.6.0 is out! The #Linux perf GUI for #performance analysis adds support for archives, tracepoints, and regex filtering. It also improves file handling and settings visibility, plus several bug fixes and stability improvements. #Hotspot #Perf #GUI
-
Hotspot v1.6.0 is out! The #Linux perf GUI for #performance analysis adds support for archives, tracepoints, and regex filtering. It also improves file handling and settings visibility, plus several bug fixes and stability improvements. #Hotspot #Perf #GUI
-
The Kirby Staticache plugin can now pre-compress ahead. That sounds interesting. High compression levels for on the fly caching can put a load on the web server. #KirbyCMS #Perf #WebDev
https://github.com/getkirby/staticache/releases/tag/2.1.0 -
The Kirby Staticache plugin can now pre-compress ahead. That sounds interesting. High compression levels for on the fly caching can put a load on the web server. #KirbyCMS #Perf #WebDev
https://github.com/getkirby/staticache/releases/tag/2.1.0 -
Кэш, который нас предал: как мы ловили призраков в L3 и нашли side-effects в продакшене
Это история о том, как мы несколько недель искали странные скачки latency в продакшене и в итоге уткнулись в поведение кэша процессора. Не в аллокатор, не в GC, не в сеть. В кэш. В статье — реальные эксперименты, код, метрики, гипотезы, которые не подтвердились, и довольно неприятные выводы о том, насколько процессор может быть непредсказуемым, когда система нагружена по-взрослому.
https://habr.com/ru/articles/1003824/
#кэш_процессора #cache_miss #L3_cache #latency #perf #false_sharing #NUMA #side_effects
-
Am Perfstausee, mit Blick zum oberen (1910) und unteren Schloss Breidenstein (1714)
#Perf #Perfstausee #Stausee #Schloss #Jugendstilvilla #Breidenstein #Breidenbach #LahnDillBergland #MarburgBiedenkopf #Biedenkopf #Marburg #Mittelhessen #Hessen #HE #Germany
-
@rsc have you seen https://vitaut.net/posts/2025/smallest-dtoa/ / https://vitaut.net/posts/2025/faster-dtoa/ / https://github.com/vitaut/zmij ? It's also pretty short and fast, and I didn't see it mentioned in your references section. (I haven't read your post in detail yet, apologies if it's mentioned somewhere!)
Also, the #perf link doesn't seem to work for me.
-
@rsc have you seen https://vitaut.net/posts/2025/smallest-dtoa/ / https://vitaut.net/posts/2025/faster-dtoa/ / https://github.com/vitaut/zmij ? It's also pretty short and fast, and I didn't see it mentioned in your references section. (I haven't read your post in detail yet, apologies if it's mentioned somewhere!)
Also, the #perf link doesn't seem to work for me.
-
New Directive Clarifies Existing Use of Force Policy at U.S. Customs and Border Protection - American Immigration Council March 11, 2014
> The new directive restricts Border Patrol agents from shooting at moving vehicles merely fleeing from agents. It also states that “agents should not place themselves in the path of a moving vehicle or use their body to block a vehicle’s path.”
-
New Directive Clarifies Existing Use of Force Policy at U.S. Customs and Border Protection - American Immigration Council March 11, 2014
> The new directive restricts Border Patrol agents from shooting at moving vehicles merely fleeing from agents. It also states that “agents should not place themselves in the path of a moving vehicle or use their body to block a vehicle’s path.”
-
How to use #flamegraphs for #performance #profiling
https://runbooks.gitlab.com/tutorials/how_to_use_flamegraphs_for_perf_profiling/
Off-CPU Analysis - by Brendan Gregg:
https://www.brendangregg.com/offcpuanalysis.html- On-CPU: where threads are spending time running on-CPU
- Off-CPU: where time is spent waiting while blocked on I/O, locks, timers, paging/swapping, etc. -
How to use #flamegraphs for #performance #profiling
https://runbooks.gitlab.com/tutorials/how_to_use_flamegraphs_for_perf_profiling/
Off-CPU Analysis - by Brendan Gregg:
https://www.brendangregg.com/offcpuanalysis.html- On-CPU: where threads are spending time running on-CPU
- Off-CPU: where time is spent waiting while blocked on I/O, locks, timers, paging/swapping, etc. -
-
-
Аппаратные breakpoint’ы: для чего они нужны и как устроены в Linux
Всем привет! Наша группа занимается RISC-V Linux и загрузчиками в компании «Синтакор». Однажды перед нами возникла задача — реализовать поддержку аппаратных триггеров в ядре Linux и OpenSBI. Она стала началом исследования, в ходе которого я изучил смысл аппаратных триггеров с точки зрения отладчика, их устройство и использование для watchpoint’ов и breakpoint’ов, а также принял участие в совершенствовании поддержки аппаратных триггеров в RISC-V Linux и OpenSBI. Этими знаниями я хотел бы поделиться в статье. Покажу на примерах, как устроены breakpoint’ы и watchpoint’ы в отладчиках, сравню их программную и аппаратную реализации, покопаюсь в деталях их работы в ядре Linux. Начну с легкого способа сломать GDB, а к каким выводам он приведет, вы узнаете далее под катом. GDB хрясь!
-
Randomly thinking about the `shadowrootadoptedstylesheets` proposal today and had some thoughts about how it could support streaming use cases better.
https://github.com/MicrosoftEdge/MSEdgeExplainers/issues/1188
-
Randomly thinking about the `shadowrootadoptedstylesheets` proposal today and had some thoughts about how it could support streaming use cases better.
https://github.com/MicrosoftEdge/MSEdgeExplainers/issues/1188
-
🎙️ #92 ASICS FUJISPEED 4: La Velocità che Trasforma il Tuo Trail!
ASICS FUJISPEED 4: La Velocità che Trasforma il Tuo Trail!Amanti del trail running, preparatevi!Leggerezza incredibile, trazione da paura e un comfort che ti fa volare sui sentieri. Dimentica la fatica, concentrati sul divertimento! #Asics #Fujispeed4 #TrailRunning #ScarpeRunning #Perf...
-
Evaluating framework performance. Loren Stewart compared 10 meta-frameworks: Marko, SolidStart, SvelteKit, Qwik, and Nuxt deliver 35–39 ms FCP and 29–176 kB bundles. The takeaway is React’s architectural ceiling: even TanStack Start on React ships bundles 2× as large compared to Solid; Angular via Analog remains heavy, while Vue via Nuxt is quite competitive. MPA frameworks like Marko and HTMX ship minimal JS per page, while SPAs pay the baseline runtime cost. #js #perf
-
Одно и то же приложение 10 раз. Лорен Стюарт сравнил 10 метафреймворков: Marko, SolidStart, SvelteKit, Qwik и Nuxt дают FCP 35–39 мс и бандл от 28,8 до 176,3 КБ. Ключевой вывод — архитектурный потолок React: даже TanStack Start на React даёт вдвое больший бандл, чем тот же TanStack Start c Solid, Angular через Analog остаётся тяжёлым, а Vue через Nuxt вполне конкурентен. MPA на Marko и HTMX везут минимум JS на страницу, тогда как SPA платят базовым рантаймом. #js #perf
-
Одно и то же приложение 10 раз. Лорен Стюарт сравнил 10 метафреймворков: Marko, SolidStart, SvelteKit, Qwik и Nuxt дают FCP 35–39 мс и бандл от 28,8 до 176,3 КБ. Ключевой вывод — архитектурный потолок React: даже TanStack Start на React даёт вдвое больший бандл, чем тот же TanStack Start c Solid, Angular через Analog остаётся тяжёлым, а Vue через Nuxt вполне конкурентен. MPA на Marko и HTMX везут минимум JS на страницу, тогда как SPA платят базовым рантаймом. #js #perf
-
#TIL Random ordering in Django. order_by('?') works but can be costly on large tables. Prefer order_by(Random()) (Django’s DB function) + LIMIT.
Notes + pitfalls inside. #Django #Python #ORM #PostgreSQL #Perf
https://til.sanyamkhurana.com/#/topics/django/random-ordering-with-order-by-in-django
-
#TIL Random ordering in Django. order_by('?') works but can be costly on large tables. Prefer order_by(Random()) (Django’s DB function) + LIMIT.
Notes + pitfalls inside. #Django #Python #ORM #PostgreSQL #Perf
https://til.sanyamkhurana.com/#/topics/django/random-ordering-with-order-by-in-django
-
Tiny guide: run EXPLAIN & EXPLAIN ANALYZE from Django, read the plan, then choose fixes (index? rewrite join?). Notes + pitfalls. #Django #Postgres #Perf #TIL
https://til.sanyamkhurana.com/#/topics/django/using-explain-and-explain-analyze-for-django-querysets -
Copilot Diagnostics toolset for .NET In Visual Studio | by Harshada Hole.
#dotnet #githubcopilot #visualstudio #productivity #copilot #ai #diagnostics #perf #debugging
-
Copilot Diagnostics toolset for .NET In Visual Studio | by Harshada Hole.
#dotnet #githubcopilot #visualstudio #productivity #copilot #ai #diagnostics #perf #debugging
-
Andrea Savage Multi-Cam Comedy ‘Perf’ In Works At Fox
#News #AndreaSavage #Fox #Perfhttps://deadline.com/2025/07/andrea-savage-perf-fox-multi-camera-comedy-1236473601/
-
Andrea Savage Multi-Cam Comedy ‘Perf’ In Works At Fox
#News #AndreaSavage #Fox #Perfhttps://deadline.com/2025/07/andrea-savage-perf-fox-multi-camera-comedy-1236473601/
-
That gives us our baseline: bzip2 (in C) vs bzip2 (in Rust). But is it a fair enough comparison? I mentioned initially that I was implementing an lbzip2 "clone" (mostly a PoC for the decompression part). lbzip2 is an other program (a C binary, without a library), that can compress and decompress bzip2 files in parallel. Surely it should be slower than bzip2 since it has the parallel management overhead? 7/N
-
That gives us our baseline: bzip2 (in C) vs bzip2 (in Rust). But is it a fair enough comparison? I mentioned initially that I was implementing an lbzip2 "clone" (mostly a PoC for the decompression part). lbzip2 is an other program (a C binary, without a library), that can compress and decompress bzip2 files in parallel. Surely it should be slower than bzip2 since it has the parallel management overhead? 7/N
-
But why? Let's see what
perf stathas to say: the Rust version has less instructions, but with much less IPC (Instruction-per-clock); the Rust version also has less branches and misses in general. On the efficiency cores, we see that worse IPC and branch prediction of the Rust version give the advantage to the C version. 6/N -
But why? Let's see what
perf stathas to say: the Rust version has less instructions, but with much less IPC (Instruction-per-clock); the Rust version also has less branches and misses in general. On the efficiency cores, we see that worse IPC and branch prediction of the Rust version give the advantage to the C version. 6/N -
I don't know much about Apple/Swift but this is a great talk!
For a starter, changing from `Collection<E>` to contiguous container type`Span<E>` increase the performance 4x of a sample binary search. Just wow
https://developer.apple.com/videos/play/wwdc2025/308/ -
TIL: using hpet (instead of tsc) as #Linux clocksource apparently can cause massive slowdowns in some workloads. E.g. when reading data from a socket using select() and read(), with hpet clock the select() calls alone can reduce throughput down to 25%. #Perf profiling shows that the kernel spends lots of time in read_hpet().
Guess now I cannot postpone debugging the clock drift problems that occur with tsc clocksource on this one system :-(
-
Whose optimisation is better? #perf #dbms #postgresql https://danolivo.substack.com/p/whose-optimisation-is-better
-
Whose optimisation is better? #perf #dbms #postgresql https://danolivo.substack.com/p/whose-optimisation-is-better
-
Hotspot: Linux `perf` GUI for performance analysis
https://github.com/KDAB/hotspot
#HackerNews #Hotspot #Linux #perf #GUI #performance #analysis #KDAB #GitHub
-
Hotspot: Linux `perf` GUI for performance analysis
https://github.com/KDAB/hotspot
#HackerNews #Hotspot #Linux #perf #GUI #performance #analysis #KDAB #GitHub
-
Today's bug is a `perf` hangup bug: https://lkml.org/lkml/2025/5/5/1089
There a simple `perf record -a` / `perf report` hangs up if you happen to have a `/dev/dri/renderD128` file `mmap()`ed in any of the processes. Browsers and compositors usually do have them `mmap()`ed.
-
How we ended up rewriting NuGet Restore in .NET 9 - .NET Blog
https://devblogs.microsoft.com/dotnet/rewriting-nuget-restore-in-dotnet-9/