conv.

All stories
SecurityQuiet 3d · day 5

Open-weight AI models trigger year-long security crisis

Cheap, abliterated LLMs like GLM 5.3-flash can find vulnerabilities faster than humans can patch them.

What to know

  • Open-weight GLM 5.3-flash, now stripped of safety guardrails by third parties, scores 84.5% on real-world vulnerability benchmarks and is runnable on a ~$9,500 Mac; security researchers say this puts dangerous AI exploit generation in reach of any attacker.
  • The industry faces a compressed deadline—estimates range from one year to already-negative (attackers are exploiting now)—to patch systemic vulnerabilities before LLM-assisted mass exploitation becomes the norm.
  • Consumer operating systems lack usable sandboxing; the conversation reveals no consensus tool exists that combines both security hardening and ease of use, leaving practitioners to choose between obscure/undocumented (Apple Sandbox), difficult to configure securely (gVisor, Kubernetes), or hard to deploy (Firecracker).
  • Long-term fix proposed: capability-based security models (already used in Android/iOS pickers) replacing identity-based process authorization, but adoption requires significant UI and cultural redesign.

The dispute Whether the threat timeline is one year (researchers' starting position) or already negative (some believe exploitation is underway but hidden from public view). · positions read across 57 posts and comments

many voices

Operating-system sandboxing is the critical missing layer; capability-based security should replace identity-based models.

  • “I would like to be able to run a process in a way that prevents it from digging around in my home directory and exfiltrating anything it finds to someone else over the internet.”

    simonw · Lobsters ↗
some voices

The one-year timeline is already too optimistic; black-hat exploitation is underway now and not visible in public reporting.

  • “We don't even have a year. I ran a security audit using GPT-6 Astra and Claude Fable 5.1 against my main open source project (Datasette) recently and then spent a full *week* fixing vulnerabilities.”

    simonw · Lobsters ↗
many voices

No existing sandboxing tool is both secure-by-default and easy to use; every option trades off security hardening against usability.

  • “My problem with sandboxes is that what I really want is an easy-to-use sandbox from a company with a dedicated security team that works on that sandbox product, and who risk millions (ideally billions) of dollars if it leaks.”

    simonw · Lobsters ↗
some voices

Urgent security response is necessary and possible; LLMs finding vulnerabilities faster than humans is actually pushing the industry toward taking security seriously.

  • “We are definitely at the time in which cheap / unaligned / locally-runnable AI models are *excellent* at finding vulnerabilities and crafting exploit scripts, and they're only going to get better. That does suck in some regards, but I…”

    apropos · Lobsters ↗

apropos (author of Datasette) Security researcher, post authorSimon Willison (simonw) Datasette maintainer, open-source security expertZ.ai Co. (formerly Zhipu AI) AI model developerjfred (commenter) Security architecture discussantDeAlignAI Model abliteration group

How it unfolded 5 developments, newest first · click a bar or a number to jump postscomments

Peak 8 pieces in one hour at Sep 19, 7 PM; 58 pieces over 5 days (1 post · 57 comments) Sep 19, 2 PM — 1 piece · 1 post — Lobsters 1Sep 19, 3 PM — 1 piece · 1 comment — Lobsters 1Sep 19, 4 PM — quietSep 19, 5 PM — 4 pieces · 4 comments — Lobsters 4Sep 19, 6 PM — 1 piece · 1 comment — Lobsters 1Sep 19, 7 PM — 8 pieces · 8 comments — Lobsters 8Sep 19, 8 PM — 1 piece · 1 comment — Lobsters 1Sep 19, 9 PM — 1 piece · 1 comment — Lobsters 1Sep 19, 10 PM — 6 pieces · 6 comments — Lobsters 6Sep 19, 11 PM — 3 pieces · 3 comments — Lobsters 3Sep 20, 12 AM — 3 pieces · 3 comments — Lobsters 3Sep 20, 1 AM — 1 piece · 1 comment — Lobsters 1Sep 20, 2 AM — quietSep 20, 3 AM — quietSep 20, 4 AM — quietSep 20, 5 AM — 2 pieces · 2 comments — Lobsters 2Sep 20, 6 AM — 2 pieces · 2 comments — Lobsters 2Sep 20, 7 AM — 1 piece · 1 comment — Lobsters 1Sep 20, 8 AM — 1 piece · 1 comment — Lobsters 1Sep 20, 9 AM — 1 piece · 1 comment — Lobsters 1Sep 20, 10 AM — quietSep 20, 11 AM — quietSep 20, 12 PM — 2 pieces · 2 comments — Lobsters 2Sep 20, 1 PM — 3 pieces · 3 comments — Lobsters 3Sep 20, 2 PM — 2 pieces · 2 comments — Lobsters 2Sep 20, 3 PM — 2 pieces · 2 comments — Lobsters 2Sep 20, 4 PM — 2 pieces · 2 comments — Lobsters 2Sep 20, 5 PM — 4 pieces · 4 comments — Lobsters 4Sep 20, 6 PM — quietSep 20, 7 PM — 1 piece · 1 comment — Lobsters 1Sep 20, 8 PM — quietSep 20, 9 PM — quietSep 20, 10 PM — quietSep 20, 11 PM — quietSep 21, 12 AM — quietSep 21, 1 AM — 1 piece · 1 comment — Lobsters 1Sep 21, 2 AM — 1 piece · 1 comment — Lobsters 1Sep 21, 3 AM — quietSep 21, 4 AM — quietSep 21, 5 AM — 2 pieces · 2 comments — Lobsters 2Sep 21, 6 AM — quietSep 21, 7 AM — quietSep 21, 8 AM — 1 piece · 1 comment — Lobsters 1Sep 21, 9 AM — quietSep 21, 10 AM — quietSep 21, 11 AM — quietSep 21, 12 PM — quietSep 21, 1 PM — quietSep 21, 2 PM — quietSep 21, 3 PM — quietSep 21, 4 PM — quietSep 21, 5 PM — quietSep 21, 6 PM — quietSep 21, 7 PM — quietSep 21, 8 PM — quietSep 21, 9 PM — quietSep 21, 10 PM — quietSep 21, 11 PM — quietSep 22, 12 AM — quietSep 22, 1 AM — quietSep 22, 2 AM — quietSep 22, 3 AM — quietSep 22, 4 AM — quietSep 22, 5 AM — quietSep 22, 6 AM — quietSep 22, 7 AM — quietSep 22, 8 AM — quietSep 22, 9 AM — quietSep 22, 10 AM — quietSep 22, 11 AM — quietSep 22, 12 PM — quietSep 22, 1 PM — quietSep 22, 2 PM — quietSep 22, 3 PM — quietSep 22, 4 PM — quietSep 22, 5 PM — quietSep 22, 6 PM — quietSep 22, 7 PM — quietSep 22, 8 PM — quietSep 22, 9 PM — quietSep 22, 10 PM — quietSep 22, 11 PM — quietYesterday, 12 AM — quietYesterday, 1 AM — quietYesterday, 2 AM — quietYesterday, 3 AM — quietYesterday, 4 AM — quietYesterday, 5 AM — quietYesterday, 6 AM — quietYesterday, 7 AM — quietYesterday, 8 AM — quietYesterday, 9 AM — quietYesterday, 10 AM — quietYesterday, 11 AM — quietYesterday, 12 PM — quietYesterday, 1 PM — quietYesterday, 2 PM — quietYesterday, 3 PM — quietYesterday, 4 PM — quietYesterday, 5 PM — quietYesterday, 6 PM — quietYesterday, 7 PM — quietYesterday, 8 PM — quietYesterday, 9 PM — quietYesterday, 10 PM — quietYesterday, 11 PM — quietToday, 12 AM — quietToday, 1 AM — quietToday, 2 AM — quiet 12–34–5
Sep 20Sep 21Sep 22yesterdaynow · 3:43 AM ET
  1. background

    Security experts debate capability-based OS architecture — Commenters propose capability-based security models (requiring processes to hold explicit authority tokens rather than inheriting user identity) as a long-term fix, though they acknowledge UI challenges. References to government compartmentalization practices (SCIF model) and emerging tools like gVisor, Firecracker, and bubblewrap suggest incremental progress.

  2. 5

    Researcher claims threat timeline may be even shorter

    A commenter argues the actual window is negative—that black-hat exploitation is already underway in the wild and simply not making headlines, suggesting the industry is not merely behind schedule but already under active attack.

    “I think it's even worse than that, it's less "we don't even have a year" and more "we have negative six months". Black hats are absolutely going ham in the wild and just not making the news.”
    — wareya
    • I think it's even worse than that, it's less "we don't even have a year" and more "we have negative six months". Black hats are absolutely going ham in the wild and just not making the news.

      wareyasecurity,vibecoding4d ago19▲view on Lobsters ↗
    2 more of the top 3 · 32 posts in this stretch
    • Attacker level of effort has been the only thing protecting anyone for the last 40 years, but if money is not an object then all other factors in "attacker level of effort" will approach *zero*. We don't need a Judgement Day scenario or genius-hacker level models for that to be disastrous.

      Qyriadsecurity,vibecoding3d ago11▲view on Lobsters ↗
    • "There is an exploitable security vulnerability in $TARGET. Please find it and give me a reproducible test" When doing so, $TARGET can be either a file name or a function, which allows you to loop / parallelize over the codebase. If you hunt for specific vulnerabilities, it's best to enable the LLM to check its work by providing a verifier. For…

      freddybsecurity,vibecoding3d ago7▲view on Lobsters ↗
    all of them →
  3. 4

    Community debate over practical sandboxing solutions emerges

    Thread identifies multiple real-world sandboxing approaches (gVisor donated to CNCF, Firecracker, Kata Containers, bubblewrap, Apple Sandbox) but reveals deep tensions: every tool trades security hardening against usability, and many are either undocumented (Apple Sandbox), hard to configure securely (gVisor, Kubernetes), or difficult to deploy (Firecracker). No consensus emerges on what practitioners should use now.

    “My problem with sandboxes is that what I really want is an easy-to-use sandbox from a company with a dedicated security team that works on that sandbox product, and who risk millions (ideally billions) of dollars if it leaks.”
    — simonw
    • I've been spending quite a lot of time fixing llm-reported vulnerabilities this year, especially in the last 3 or 4 months. Hopefully at some point, chromium will just "not have bugs", and if introduced they get found quickly before they hit stable. It's a lot of work getting there though lol.

      dmurphsecurity,vibecoding4d ago20▲view on Lobsters ↗
    2 more of the top 3 · 11 posts in this stretch
    • UI design is definitely a challenge, yeah; capabilities have been more obscure for a while and haven't had as much UI work as more common architectures. Though for what it's worth, things like the file/photo pickers in Android and iOS are capability-style powerboxes that people use every day. I think it's possible to expand on this, but it will…

      jfredsecurity,vibecoding4d ago12▲view on Lobsters ↗
    • I really love [bubblewrap](https://github.com/containers/bubblewrap) for this on Linux. It really does make it simple to run a single command with controlled permissions. You can share and unshare specific parts of the filesystem, as well as control network access with a simple wrapper script. Linux has some seriously powerful tools to do this…

      spillybonessecurity,vibecoding4d ago9▲view on Lobsters ↗
    all of them →
  4. 3

    Community identifies sandboxing as critical missing defense

    Discussion shifts to operating system sandboxing as an inadequate defense. Multiple researchers argue that consumer OSes lack usable, documented sandboxing—and that capability-based security models (used in Android and iOS pickers but rarely elsewhere) should replace the identity-based model where processes run "as you" with full authority.

    “I would like to be able to run a process in a way that prevents it from digging around in my home directory and exfiltrating anything it finds to someone else over the internet. This is way harder than it should be.”
    — simonw
    • I would like to see a whole lot more attention paid to sandboxing. It infuriates me that consumer operating systems don't ship with clearly documented, usable sandboxing features - the sandboxes they include today may as well have signs pasted on them saying "Beware of the Leopard". I would like to be able to run a process in a way that prevents…

      simonwsecurity,vibecoding4d ago32▲view on Lobsters ↗
    2 more of the top 3 · 11 posts in this stretch
    • I remain convinced that the right answer for desktop OSes in the long term (albeit not necessarily one that's easy to get to from here) will have to involve [capabilities](http://habitat-chronicles.com/2017/05/what-are-capabilities/comment-page-1/). The identity-based model of processes that run "as you" with all your authority hasn't really fit…

      jfredsecurity,vibecoding4d ago19▲view on Lobsters ↗
    • If we can figure out how to wrap that in a UI that doesn't require half a degree in cybersecurity to use safely and effectively I'm all for it!

      simonwsecurity,vibecoding4d ago9▲view on Lobsters ↗
    all of them →
  5. 2

    Researcher reports immediate vulnerability flood from LLM audits

    Datasette maintainer Simon Willison, citing a recent security audit of his project using GPT-6 Astra and Claude Fable 5.1, reports spending a full week patching extremely obscure vulnerabilities that LLMs discovered—vulnerabilities that had gone unspotted by humans for extended periods. He contradicts the one-year timeline, stating "We don't even have a year."

    “I ran a security audit using GPT-6 Astra and Claude Fable 5.1 against my main open source project (Datasette) recently and then spent a full *week* fixing vulnerabilities that they found, many of which were extremely obscure, hence why nobody had spotted them before.”
    — simonw
    • We don't even have a year. I ran a security audit using GPT-6 Astra and Claude Fable 5.1 against my main open source project (Datasette) recently and then spent a full *week* fixing vulnerabilities that they found, many of which were extremely obscure, hence why nobody had spotted them before. An obscure vulnerability is still a vulnerability…

      simonwsecurity,vibecoding4d ago30▲view on Lobsters ↗
    1 more of the top 2 · 2 posts in this stretch
    • I’m in the middle of reading a book on the stuxnet attack and one thing that really stood out to me is that at least two of the zero days had been publicly published or disclosed for a year or more.

      bradsecurity,vibecoding4d ago12▲view on Lobsters ↗
    all of them →
  6. 1

    Researcher publishes vulnerability timeline warning

    Security researcher and author of Datasette posts analysis showing that abliterated GLM 5.3-flash is affordable to run locally (30–45 tokens/second on the upcoming M5 Mac Studio for ~$9,500) and scores 84.5% on CyberGym (real-world vulnerabilities) and 54.4% on ExploitBench (exploit generation). The post warns the industry has roughly one year to fix systemic security flaws before cheap, unaligned models enable mass exploitation.

    “We have a year to fix security everywhere... We can use frontier LLMs that move faster than a human to find and fix these issues in the time we have left.”
    — apropos
    • I think this is a very interesting post and worthy of discussion. For my part, I agree with it, but only partially. We are definitely at the time in which cheap / unaligned / locally-runnable AI models are *excellent* at finding vulnerabilities and crafting exploit scripts, and they're only going to get better. That does suck in some regards, but…

      apropossecurity,vibecoding4d ago19▲view on Lobsters ↗
  7. background

    Z.ai releases GLM 5.3-flash model — Z.ai (formerly Zhipu AI) releases the General Language Model 5.3-flash as an open-weight model. When hosted by Z.ai it includes task refusals required by law, but organizations like DeAlignAI immediately release abliterated versions with safety mechanisms surgically removed.

What people are saying 11 voices from 1 site · best of 57 · verbatim

Still unanswered
  • Which sandboxing tool should practitioners use right now that is both secure and practical to deploy?
  • How much of the LLM exploit generation happening today is already being weaponized in the wild versus remaining theoretical?