#glasswing — Public Fediverse posts
Live and recent posts from across the Fediverse tagged #glasswing, aggregated by home.social.
-
The Anthropic Glasswing Receipts Are Starting to Trickle In
https://www.vulncheck.com/blog/anthropic-glasswing-receipts
> Anthropic’s Project Glasswing is approaching 5 months old, and Anthropic published its Vulnerability Disclosure Ledger on May 22nd. It hadn't received an update until this past week, when it backfilled the ledger with additional findings and updates, so naturally I thought it would be worthwhile to take a look at the receipts.
-
The Anthropic Glasswing Receipts Are Starting to Trickle In
https://www.vulncheck.com/blog/anthropic-glasswing-receipts
> Anthropic’s Project Glasswing is approaching 5 months old, and Anthropic published its Vulnerability Disclosure Ledger on May 22nd. It hadn't received an update until this past week, when it backfilled the ledger with additional findings and updates, so naturally I thought it would be worthwhile to take a look at the receipts.
-
"Anthropic’s own dashboard states 421 findings patched upstream, resulting in 462 advisories (GHSA/CVEs assigned); however, the ledger itself shows only 202 fixed findings.
If we look at the GHSA ledger, CVE ledger, and main ledger (yes, there are three), the numbers also don’t add up. The CVE ledger lists 70 CVEs while the main ledger has 82 CVEs. The GHSA ledger lists 49 GHSAs, while the main ledger lists 77 GHSAs. None of which add up to the claim of 421 findings fixed upstream.
So this makes me wonder whether the ledger is AI-assisted and under-reviewed. Likely a combination of the two.
The Ledger has received only two bulk updates, so results are not published in real time. It’s worth highlighting that while the ledger has a public reveal date, that date doesn’t reflect the actual day the finding was revealed in the ledger as we saw with the recent update.
Don’t discount the reality that AI tools like Claude are incredibly valuable and useful tools for discovering vulnerabilities. AI can help accelerate the discovery of bugs and vulnerabilities in software, as the evidence I've discussed across software suppliers and the CVE program shows. This blog aims to better understand the claims that Anthropic and other frontier model providers have made about Project Glasswing, and to see if the evidence aligns with those claims. It appears there are many discrepancies in the data they've published, which is still a small fraction of the findings after five months. The receipts are starting to trickle in, they just don’t reconcile."
https://www.vulncheck.com/blog/anthropic-glasswing-receipts
#AI #Anthropic #CyberSecurity #ProjectGlasswing #Mythos #Glasswing #LLMs #GenerativeAI
-
"Anthropic’s own dashboard states 421 findings patched upstream, resulting in 462 advisories (GHSA/CVEs assigned); however, the ledger itself shows only 202 fixed findings.
If we look at the GHSA ledger, CVE ledger, and main ledger (yes, there are three), the numbers also don’t add up. The CVE ledger lists 70 CVEs while the main ledger has 82 CVEs. The GHSA ledger lists 49 GHSAs, while the main ledger lists 77 GHSAs. None of which add up to the claim of 421 findings fixed upstream.
So this makes me wonder whether the ledger is AI-assisted and under-reviewed. Likely a combination of the two.
The Ledger has received only two bulk updates, so results are not published in real time. It’s worth highlighting that while the ledger has a public reveal date, that date doesn’t reflect the actual day the finding was revealed in the ledger as we saw with the recent update.
Don’t discount the reality that AI tools like Claude are incredibly valuable and useful tools for discovering vulnerabilities. AI can help accelerate the discovery of bugs and vulnerabilities in software, as the evidence I've discussed across software suppliers and the CVE program shows. This blog aims to better understand the claims that Anthropic and other frontier model providers have made about Project Glasswing, and to see if the evidence aligns with those claims. It appears there are many discrepancies in the data they've published, which is still a small fraction of the findings after five months. The receipts are starting to trickle in, they just don’t reconcile."
https://www.vulncheck.com/blog/anthropic-glasswing-receipts
#AI #Anthropic #CyberSecurity #ProjectGlasswing #Mythos #Glasswing #LLMs #GenerativeAI
-
"Anthropic’s own dashboard states 421 findings patched upstream, resulting in 462 advisories (GHSA/CVEs assigned); however, the ledger itself shows only 202 fixed findings.
If we look at the GHSA ledger, CVE ledger, and main ledger (yes, there are three), the numbers also don’t add up. The CVE ledger lists 70 CVEs while the main ledger has 82 CVEs. The GHSA ledger lists 49 GHSAs, while the main ledger lists 77 GHSAs. None of which add up to the claim of 421 findings fixed upstream.
So this makes me wonder whether the ledger is AI-assisted and under-reviewed. Likely a combination of the two.
The Ledger has received only two bulk updates, so results are not published in real time. It’s worth highlighting that while the ledger has a public reveal date, that date doesn’t reflect the actual day the finding was revealed in the ledger as we saw with the recent update.
Don’t discount the reality that AI tools like Claude are incredibly valuable and useful tools for discovering vulnerabilities. AI can help accelerate the discovery of bugs and vulnerabilities in software, as the evidence I've discussed across software suppliers and the CVE program shows. This blog aims to better understand the claims that Anthropic and other frontier model providers have made about Project Glasswing, and to see if the evidence aligns with those claims. It appears there are many discrepancies in the data they've published, which is still a small fraction of the findings after five months. The receipts are starting to trickle in, they just don’t reconcile."
https://www.vulncheck.com/blog/anthropic-glasswing-receipts
#AI #Anthropic #CyberSecurity #ProjectGlasswing #Mythos #Glasswing #LLMs #GenerativeAI
-
"Anthropic’s own dashboard states 421 findings patched upstream, resulting in 462 advisories (GHSA/CVEs assigned); however, the ledger itself shows only 202 fixed findings.
If we look at the GHSA ledger, CVE ledger, and main ledger (yes, there are three), the numbers also don’t add up. The CVE ledger lists 70 CVEs while the main ledger has 82 CVEs. The GHSA ledger lists 49 GHSAs, while the main ledger lists 77 GHSAs. None of which add up to the claim of 421 findings fixed upstream.
So this makes me wonder whether the ledger is AI-assisted and under-reviewed. Likely a combination of the two.
The Ledger has received only two bulk updates, so results are not published in real time. It’s worth highlighting that while the ledger has a public reveal date, that date doesn’t reflect the actual day the finding was revealed in the ledger as we saw with the recent update.
Don’t discount the reality that AI tools like Claude are incredibly valuable and useful tools for discovering vulnerabilities. AI can help accelerate the discovery of bugs and vulnerabilities in software, as the evidence I've discussed across software suppliers and the CVE program shows. This blog aims to better understand the claims that Anthropic and other frontier model providers have made about Project Glasswing, and to see if the evidence aligns with those claims. It appears there are many discrepancies in the data they've published, which is still a small fraction of the findings after five months. The receipts are starting to trickle in, they just don’t reconcile."
https://www.vulncheck.com/blog/anthropic-glasswing-receipts
#AI #Anthropic #CyberSecurity #ProjectGlasswing #Mythos #Glasswing #LLMs #GenerativeAI
-
"Anthropic’s own dashboard states 421 findings patched upstream, resulting in 462 advisories (GHSA/CVEs assigned); however, the ledger itself shows only 202 fixed findings.
If we look at the GHSA ledger, CVE ledger, and main ledger (yes, there are three), the numbers also don’t add up. The CVE ledger lists 70 CVEs while the main ledger has 82 CVEs. The GHSA ledger lists 49 GHSAs, while the main ledger lists 77 GHSAs. None of which add up to the claim of 421 findings fixed upstream.
So this makes me wonder whether the ledger is AI-assisted and under-reviewed. Likely a combination of the two.
The Ledger has received only two bulk updates, so results are not published in real time. It’s worth highlighting that while the ledger has a public reveal date, that date doesn’t reflect the actual day the finding was revealed in the ledger as we saw with the recent update.
Don’t discount the reality that AI tools like Claude are incredibly valuable and useful tools for discovering vulnerabilities. AI can help accelerate the discovery of bugs and vulnerabilities in software, as the evidence I've discussed across software suppliers and the CVE program shows. This blog aims to better understand the claims that Anthropic and other frontier model providers have made about Project Glasswing, and to see if the evidence aligns with those claims. It appears there are many discrepancies in the data they've published, which is still a small fraction of the findings after five months. The receipts are starting to trickle in, they just don’t reconcile."
https://www.vulncheck.com/blog/anthropic-glasswing-receipts
#AI #Anthropic #CyberSecurity #ProjectGlasswing #Mythos #Glasswing #LLMs #GenerativeAI
-
❝ I'm assuming the other 23k findings were pure dog shit and they want you to forget that it has a 90+% false positive rate ❞
https://masto.nyc/@gbargoud/117237178321400249but… WHO DECIDED THAT?
the #MythosAI #Glasswing report isn’t just obfuscating #accountability, it is pretty much leaving us in the dark about the actual #employment of tech workers’ and the amount of work/hours evaluating the LLMs code review.
someone at #Anthropic is lying about the actual human/hours #labor involved in managing Mythos.
-
❝ I'm assuming the other 23k findings were pure dog shit and they want you to forget that it has a 90+% false positive rate ❞
https://masto.nyc/@gbargoud/117237178321400249but… WHO DECIDED THAT?
the #MythosAI #Glasswing report isn’t just obfuscating #accountability, it is pretty much leaving us in the dark about the actual #employment of tech workers’ and the amount of work/hours evaluating the LLMs code review.
someone at #Anthropic is lying about the actual human/hours #labor involved in managing Mythos.
-
❝ I'm assuming the other 23k findings were pure dog shit and they want you to forget that it has a 90+% false positive rate ❞
https://masto.nyc/@gbargoud/117237178321400249but… WHO DECIDED THAT?
the #MythosAI #Glasswing report isn’t just obfuscating #accountability, it is pretty much leaving us in the dark about the actual #employment of tech workers’ and the amount of work/hours evaluating the LLMs code review.
someone at #Anthropic is lying about the actual human/hours #labor involved in managing Mythos.
-
❝ I'm assuming the other 23k findings were pure dog shit and they want you to forget that it has a 90+% false positive rate ❞
https://masto.nyc/@gbargoud/117237178321400249but… WHO DECIDED THAT?
the #MythosAI #Glasswing report isn’t just obfuscating #accountability, it is pretty much leaving us in the dark about the actual #employment of tech workers’ and the amount of work/hours evaluating the LLMs code review.
someone at #Anthropic is lying about the actual human/hours #labor involved in managing Mythos.
-
❝ I'm assuming the other 23k findings were pure dog shit and they want you to forget that it has a 90+% false positive rate ❞
https://masto.nyc/@gbargoud/117237178321400249but… WHO DECIDED THAT?
the #MythosAI #Glasswing report isn’t just obfuscating #accountability, it is pretty much leaving us in the dark about the actual #employment of tech workers’ and the amount of work/hours evaluating the LLMs code review.
someone at #Anthropic is lying about the actual human/hours #labor involved in managing Mythos.
-
RE: https://mastodon.social/@bagder/117236076403978219
maybe am missing something, but:
so #Glasswing uses #MythosAI to review all the partners’ code bases. they get 26,153 “findings” but only 2,736 are actually reported to project maintainers.
so someone ―either at #Anthropic and/or the +200 companies in Glasswing― decided that 23,417 flags should remain “off ledger”.
WHO MADE THAT DECISION? who reviewed the code review?
is this why management types loves #AI: it completely obfuscates accountability?