#swe — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #swe, aggregated by home.social.
-
Stable uptimes are more important than ever.
Last year I was already happy to have tasks run 0.5-1.5 hours. These days it’s common to run tasks between 2-7 hours.
#Ai #Claude #ClaudeCode #Anthropic #ChatGPT #Codex #OpenAI #dev #developer #SWE #AINativeEngineer -
Zach Weinersmith
@zachweinersmith.bsky.social
There's lots of stuff LLMs can't do, but if you want to disbelieve all the recent math, you have to believe a bunch of Fields Medalists with long track records are just straightforwardly lying about the thing they care most about. Here's a recent blog by Gowers:
https://gowers.wordpress.com/2026/05/08/a-recent-experience-with-chatgpt-5-5-pro/#more-6666
-
Zach Weinersmith
@zachweinersmith.bsky.social
There's lots of stuff LLMs can't do, but if you want to disbelieve all the recent math, you have to believe a bunch of Fields Medalists with long track records are just straightforwardly lying about the thing they care most about. Here's a recent blog by Gowers:
https://gowers.wordpress.com/2026/05/08/a-recent-experience-with-chatgpt-5-5-pro/#more-6666
-
SoundSim360 long-distance low-frequency noise simulator #SWE: SoundSim360 is a high-fidelity acoustic simulation tool developed by SoundSim Technologies, based on more than 25 years of research in advanced numerical methods for wave-dominated partial differential equations. Unlike traditional noise modeling software that relies on simplified ray-tracing or parabolic equation (PE) methods, SoundSim360 uses state-of-the-art physics-based algorithms… https://www.wind-watch.org/alerts/2026/07/30/soundsim360-long-distance-low-frequency-noise-simulator/ #windpower #windenergy
-
Periodic git-extra-commands announcement
https://github.com/unixorn/git-extra-commands is a collection of #git helper scripts and other git-related resources I've accumulated over the years.
You don't need a ZSH framework to use it - the framework code just adds the repo's bin directory to your PATH automatically.
-
Periodic git-extra-commands announcement
https://github.com/unixorn/git-extra-commands is a collection of #git helper scripts and other git-related resources I've accumulated over the years.
You don't need a ZSH framework to use it - the framework code just adds the repo's bin directory to your PATH automatically.
-
Sweden permits two offshore wind farms, denies 11 due to defense concerns #SWE: While saying that Sweden needs to develop new power sources and critically green power to remain competitive, the government had nonetheless decided to deny 11 pending applications. It was the latest move by the government, which also denied 13 applications in 2024, citing defense concerns. “The government assesses that the applied for activities would, among other… https://www.wind-watch.org/news/2026/07/17/sweden-permits-two-offshore-wind-farms-denies-11-due-to-defense-concerns/ #windpower #windenergy
-
RT @songqiaosu: 🐦 Ornith-1.0 model family has crossed 3M downloads on 🤗 @huggingface in two weeks of release. This milestone belongs to the community! Please leave any feedback in the comments! Every issue and PR will make Ornith stronger 💪 We'll open source and keep pushing the local LLM experience forward🫡 Ornith (@ornith_) Aloha! 🌺 Meet Ornith-1.0, a family of open-source LLMs specialized for agentic coding. Ornith-1.0 spans the full parameter sizes including 9B Dense, 31B Dense, 35B MoE, and 397B MoE. It achieves state-of-the-art performance among open-source models of comparable size on coding benchmarks including: ✅Terminal-Bench 2.1(77.5) ✅SWE-Bench(82.4 on verified, 62.2 on pro, 78.9 on Multilingual) ✅NL2Repo(48.2) ✅SWE Atlas(41.2 on QnA, 42.6 RF, 39.1 TW) ✅ClawEval(77.1) Post-trained on top of gemma4 and qwen3.5, Ornith-1.0 employs a novel self-improving training strategy in which reinforcement learning is used to generate not only solution rollouts, but also the task-specific scaffolds that drive those rollouts. By jointly optimizing the scaffold and the resulting solution, the model generate higher-quality solutions in agentic coding.😎 All models are released under the MIT license, enabling full commercial and research use. 📖Tech Blog: deep-reinforce.com/ornith_1_… 🤗Huggingface: huggingface.co/collections/d… — https://nitter.net/ornith_/status/2070148887067963854#m
mehr auf Arint.info
#huggingface #Huggingface #make #MIT #MITlicense #nitter #opensource #qwen35 #SWE #SWEBench #arint_info
-
RT @EntelligenceAI: 🚨 NEW MODEL ALERT Singapore just dropped a model that's putting up frontier-level numbers. Agnes 2.5 Pro: • 82.7 on SWE-bench Verified • 78.7 multilingual • Strong gains on SWE Atlas • Beating GLM 5.2 and DeepSeek V4 Pro on multiple cuts • Free API available today The biggest takeaway isn't the benchmark here. It's the country. NOW frontier is no longer just a US vs China story.
mehr auf Arint.info
-
RT @EntelligenceAI: Germany just released a model that's actually... pretty good. Soofi S 30B-A3B: • 27T training tokens • Open weights • German + English focused • Full transparency on data, training, and evaluation • Trained entirely in Germany It's still behind the frontier leaders. But it's much closer than most people would expect. The bigger story: More countries are building serious sovereign AI stacks from scratch. Entelligence AI (@EntelligenceAI) 🚨 NEW MODEL ALERT Singapore just dropped a model that's putting up frontier-level numbers. Agnes 2.5 Pro: • 82.7 on SWE-bench Verified • 78.7 multilingual • Strong gains on SWE Atlas • Beating GLM 5.2 and DeepSeek V4 Pro on multiple cuts • Free API available today The biggest takeaway isn't the benchmark here. It's the country. NOW frontier is no longer just a US vs China story. — https://nitter.net/EntelligenceAI/status/2076618892030685384#m
mehr auf Arint.info
#API #China #DeepSeek #Germany #nitter #SWE #SWEbench #US #arint_info
-
RT @ArtificialAnlys: GPT-5.6 Sol comes close second to Claude Fable 5 in the Artificial Analysis Intelligence Index at one third of the cost, and leads the Artificial Analysis Coding Agent Index in OpenAI’s Codex harness We supported @OpenAI with pre-release evaluation of GPT-5.6 Sol, Terra, and Luna. GPT-5.6 Sol (max) scores 1 point below Claude Fable 5 (max) in the Artificial Analysis Intelligence Index at 59 points, at approximately one third of the cost. GPT-5.6 Terra (max) and Luna (max) score 55 and 51 respectively in the Intelligence Index, at ~50% and ~80% lower Cost per Task than Sol. GPT-5.6 Sol (max) leads the Artificial Analysis Coding Agent Index at 80 points. Congratulations @OpenAI and @sama on the launch! Key takeaways: ➤ One third of the cost of Claude Fable 5: On max reasoning effort, GPT-5.6 Sol costs $1.04 per task in the Artificial Analysis Intelligence Index - offering a similar level of intelligence to Claude Fable 5 at approximately one third of the cost. Reasoning levels across GPT-5.6 Sol and Luna offer a range of options at the Pareto frontier of Intelligence vs Cost per Task. For example, GPT-5.6 Luna (max) matches or exceeds the intelligence of GLM-5.2 (max) and Gemini 3.5 Flash at a lower cost. GPT-5.6 Terra (max) and Luna (max) cost $0.55 and $0.21 per Intelligence Index task, ~50% and ~80% less than Sol. Across reasoning efforts, each new GPT-5.6 model pushes past GPT-5.5 on the Pareto frontier (excluding non-reasoning). Notably, Luna and Sol are always on the Pareto frontier ahead of Terra. This means that for any Terra effort level, there is a Luna or Sol effor…
mehr auf Arint.info
#Agent #Anthropic #Claude #ClaudeCode #Codex #Gemini #GPT5 #Grok #OpenAI #SWE #arint_info
-
RT @ArtificialAnlys: GPT-5.6 Sol comes close second to Claude Fable 5 in the Artificial Analysis Intelligence Index at one third of the cost, and leads the Artificial Analysis Coding Agent Index in OpenAI’s Codex harness We supported @OpenAI with pre-release evaluation of GPT-5.6 Sol, Terra, and Luna. GPT-5.6 Sol (max) scores 1 point below Claude Fable 5 (max) in the Artificial Analysis Intelligence Index at 59 points, at approximately one third of the cost. GPT-5.6 Terra (max) and Luna (max) score 55 and 51 respectively in the Intelligence Index, at ~50% and ~80% lower Cost per Task than Sol. GPT-5.6 Sol (max) leads the Artificial Analysis Coding Agent Index at 80 points. Congratulations @OpenAI and @sama on the launch! Key takeaways: ➤ One third of the cost of Claude Fable 5: On max reasoning effort, GPT-5.6 Sol costs $1.04 per task in the Artificial Analysis Intelligence Index - offering a similar level of intelligence to Claude Fable 5 at approximately one third of the cost. Reasoning levels across GPT-5.6 Sol and Luna offer a range of options at the Pareto frontier of Intelligence vs Cost per Task. For example, GPT-5.6 Luna (max) matches or exceeds the intelligence of GLM-5.2 (max) and Gemini 3.5 Flash at a lower cost. GPT-5.6 Terra (max) and Luna (max) cost $0.55 and $0.21 per Intelligence Index task, ~50% and ~80% less than Sol. Across reasoning efforts, each new GPT-5.6 model pushes past GPT-5.5 on the Pareto frontier (excluding non-reasoning). Notably, Luna and Sol are always on the Pareto frontier ahead of Terra. This means that for any Terra effort level, there is a Luna or Sol effor…
mehr auf Arint.info
#Agent #Anthropic #Claude #ClaudeCode #Codex #Gemini #GPT5 #Grok #OpenAI #SWE #arint_info
-
🍪 Oh no, the sky is falling because #OpenAI has changed its mind about #SWE-Bench Pro! Better enable #JavaScript, or you won't see the #cookies that tell you how to live your life! 🤦♂️
https://openai.com/index/separating-signal-from-noise-coding-evaluations/ #Pro #TechNews #HackerNews #ngated -
🍪 Oh no, the sky is falling because #OpenAI has changed its mind about #SWE-Bench Pro! Better enable #JavaScript, or you won't see the #cookies that tell you how to live your life! 🤦♂️
https://openai.com/index/separating-signal-from-noise-coding-evaluations/ #Pro #TechNews #HackerNews #ngated -
OpenAI no longer recommends SWE-Bench Pro
https://openai.com/index/separating-signal-from-noise-coding-evaluations/
Comments: https://news.ycombinator.com/item?id=48837396
#HackerNews #OpenAI #SWE-Bench #Pro #AI #Research #Coding #Evaluations #Tech #News
-
OpenAI no longer recommends SWE-Bench Pro
https://openai.com/index/separating-signal-from-noise-coding-evaluations/
Comments: https://news.ycombinator.com/item?id=48837396
#HackerNews #OpenAI #SWE-Bench #Pro #AI #Research #Coding #Evaluations #Tech #News
-
From “lines of code” to “tokens spent”, why vanity metrics mislead engineering teams and why customer happiness and business outcomes are the measures that actually matter.
#CustomerHappiness #EngineeringProductivity #MeasureWhatMatters #OutcomeOverOutput #SoftwareEngineering #DevLeadership #AIinEngineering
#swe #software #softwaredeveloperhttps://www.linkedin.com/pulse/museum-meaningless-metrics-%C3%B6zkan-pakdil-i0tre/
-
From “lines of code” to “tokens spent”, why vanity metrics mislead engineering teams and why customer happiness and business outcomes are the measures that actually matter.
#CustomerHappiness #EngineeringProductivity #MeasureWhatMatters #OutcomeOverOutput #SoftwareEngineering #DevLeadership #AIinEngineering
#swe #software #softwaredeveloperhttps://www.linkedin.com/pulse/museum-meaningless-metrics-%C3%B6zkan-pakdil-i0tre/
-
RT @Ric_RTP: This is the moment Chinese AI beat American AI. One of the largest public crypto companies in the world just DUMPED OpenAI and Anthropic. Coinbase switched to open-weight Chinese models from Zhipu and DeepSeek, and shaved nearly 50% off the company's internal AI spending. The numbers are absolutely ridiculous: Running the same enterprise workload through Anthropic's Claude costs $4,811. Running it through Zhipu's GLM 5.2 costs $544. That's a 9x price difference for equivalent output. OpenAI's GPT-5.5 sits in the middle at $3,357. DeepSeek's V4 lands at $1,071. Moonshot's Kimi at $948. On the actual benchmarks: Zhipu's GLM 5.2 scored 62.1 on SWE-bench Pro, the gold standard for coding. OpenAI's GPT-5.5 scored 58.6. One AI researcher called GLM 5.2 "at least as good as Opus 4.8 and GPT 5.5." Another called it "the first open model that can really compete with closed-source systems." The Chinese models are not just cheaper but they are now also beating American models on the benchmarks American companies pay $4,811 per workload for. Coinbase did the math first and reacted - more companies will certainly follow. Now watch what happens to the IPO timeline: Anthropic confidentially filed for an IPO targeting October at a $965 billion valuation. OpenAI followed days later with its own confidential filing. Both companies built their financial models on the assumption that they could keep charging enterprise prices that are 9 to 33x what Chinese competitors charge for the same task. Brian Armstrong publicly proved customers WILL leave. 45% of companies are now spending over $100,000…
mehr auf Arint.info
#Alibaba #Anthropic #China #Claude #Coinbase #crypto #DeepSeek #GPT5 #OpenAI #SWE #SWEbench #US #arint_info
-
⚽ Probabilistic forecasts for today's matches at #FIFA2026 using our machine learning ensemble.
In Round of 32: #FRASWE #FRA #SWE
Probability to advance: 71.2% vs. 28.8%🥅 In normal time:
Mean goals: 1.5-0.7
🇫🇷 55.9%
Draw 25.9%
🇸🇪 18.2%Heatmap with probabilistic forecasts for the possible outcomes of the match in normal time:
-
⚽ Probabilistic forecasts for today's matches at #FIFA2026 using our machine learning ensemble.
In Round of 32: #FRASWE #FRA #SWE
Probability to advance: 71.2% vs. 28.8%🥅 In normal time:
Mean goals: 1.5-0.7
🇫🇷 55.9%
Draw 25.9%
🇸🇪 18.2%Heatmap with probabilistic forecasts for the possible outcomes of the match in normal time:
-
RT @Ric_RTP: This is the moment Chinese AI beat American AI. One of the largest public crypto companies in the world just DUMPED OpenAI and Anthropic. Coinbase switched to open-weight Chinese models from Zhipu and DeepSeek, and shaved nearly 50% off the company's internal AI spending. The numbers are absolutely ridiculous: Running the same enterprise workload through Anthropic's Claude costs $4,811. Running it through Zhipu's GLM 5.2 costs $544. That's a 9x price difference for equivalent output. OpenAI's GPT-5.5 sits in the middle at $3,357. DeepSeek's V4 lands at $1,071. Moonshot's Kimi at $948. On the actual benchmarks: Zhipu's GLM 5.2 scored 62.1 on SWE-bench Pro, the gold standard for coding. OpenAI's GPT-5.5 scored 58.6. One AI researcher called GLM 5.2 "at least as good as Opus 4.8 and GPT 5.5." Another called it "the first open model that can really compete with closed-source systems." The Chinese models are not just cheaper but they are now also beating American models on the benchmarks American companies pay $4,811 per workload for. Coinbase did the math first and reacted - more companies will certainly follow. Now watch what happens to the IPO timeline: Anthropic confidentially filed for an IPO targeting October at a $965 billion valuation. OpenAI followed days later with its own confidential filing. Both companies built their financial models on the assumption that they could keep charging enterprise prices that are 9 to 33x what Chinese competitors charge for the same task. Brian Armstrong publicly proved customers WILL leave. 45% of companies are now spending over $100,000…
mehr auf Arint.info
#Alibaba #Anthropic #China #Claude #Coinbase #crypto #DeepSeek #GPT5 #OpenAI #SWE #SWEbench #US #arint_info
-
The more turbines appeared across the forest, the slower the birds moved #SWE: Forests in central Sweden once echoed with the heavy wingbeats of capercaillie moving freely between feeding grounds. Then the turbines started appearing gradually. Roads cut deeper into woodland areas. Tall towers replaced open patches between trees. Researchers decided to follow the birds more closely using GPS transmitters. The movement data eventually revealed an… https://www.wind-watch.org/news/2026/05/26/the-more-turbines-appeared-across-the-forest-the-slower-the-birds-moved/ #windpower #windenergy
-
SWE History 101
-
(archive: 2026) Reindeer habitat selection and movement changes with cumulative impacts from mining and wind power development #NOR #SWE: Abstract: Industrial expansion often occurs in landscapes already affected by multiple disturbances, leading to cumulative impacts on biodiversity and local communities. With the societal pressure for an rapid ‘green transition’, land-use changes intensifies, yet most impact assessments remain local and… https://www.wind-watch.org/documents/reindeer-habitat-selection-and-movement-changes-with-cumulative-impacts-from-mining-and-wind-power-development/ #windpower #windenergy
-
🎉 WOW, Microsoft! With a whopping 5 billion active parameters, your #MAI-Code-1-Flash managed to achieve... a *glorious* 51% on the #SWE-Bench Pro! 🚀 Who knew it took so much #AI wizardry to barely pass? Just imagine what it could do with 10 billion! 🙄
https://microsoft.ai/models/mai-code-1-flash/ #Microsoft #Pro #technews #HackerNews #ngated -
Microsoft's MAI-Code-1-Flash Scores 51% SWE-Bench Pro with Just 5B Active Params
https://microsoft.ai/models/mai-code-1-flash/
#HackerNews #Microsoft #MAI-Code #AI #ML #Performance #SWE-Bench #Tech #News
-
RT @KyleHessling1: BREAKING! Qwopus 3.6 27B is LIVE! Thank you for your patience on this one, but I believe you'll find the wait was worth it! We've benchmarked this thing up and down, verified that it holds at least a 75.25% (152/202) in the initial 202 SWE bench solves. Not a full run of 500, but it shows the agentic coding quality from the original 27B is retained while adding all of the additional Qwopus benefits across many domains. As always, Jackrong is absolutely cooking here! COT quality has improved significantly through the inversion techniques from our Negentropy proof of concept. It also went through thorough curriculum training. You can check out the MMLU pro benchmarks on the model card, but it improved a whopping 10 points over the base model in physics, as well as meaningful jumps in Chemistry, business, and computer science. However, the best part is that I was able to build an entire survival shooter game using this local model entirely. I genuinely was blown away by the results, which you can play right now on my HF space (link in comments below). "Qwopus Commander" was completed in 9 turns of Qwopus 3.6! To test the new long context training, I made it re-output the entire 3000+ line program each turn, and it would make fixes and add features that I requested in large prompts, while perfectly replicating the entire rest of the game from context. What's more is that I did it all at Q8 KV cache quantization, and never had an issue over the entire 303k token run! IMPORTANT: Run it at --temp 0.75 to 1. Mess with it in that range for your use case. Higher temp actually…
mehr auf Arint.info
#GGUF #huggingface #make #rest #science #SWE #Swe #arint_info
-
Sorry, aber das gefällt mir wirklich gut. Der Daft Punk Dancetrack hat was. Diese harten Sägezahn-Synths mag ich, da bin ich ganz einfach.
-
Der Techno Teil von #swe ist auch zu alt. Der Dance Teil aber macht Laune. #Eurovision #esc #esc2026
-
-