Developer warns frontier AI providers are "pirates" after Navier-Stokes incident
Patrick McCanna argues that OpenAI and Anthropic's data practices make self-hosted LLMs the only privacy-safe option.
What to know
- Developer Patrick McCanna frames migrating to self-hosted LLMs as a privacy necessity, alleging OpenAI and Anthropic train on sensitive user session metadata to extract novel insights.
- McCanna accuses frontier providers of security theater and deliberately insecure practices, calling them "pirates" rather than trustworthy partners.
- Hacker News readers dispute McCanna's analysis, arguing the article conflates technical context-window limitations with privacy concerns and lacks substance on either front.
- Commenters note that 35KB prompts may represent legitimate structured data architectures, not confused design, and report real-world context usage far exceeding McCanna's claimed limits.
The dispute Whether McCanna's article addresses a real privacy risk or merely recasts a technical limitation (context window incompatibility) as a security issue. · positions read across 16 posts and comments
McCanna's privacy concerns about frontier providers are justified; self-hosted models are the responsible choice for protecting proprietary reasoning.
-
“If you want to protect your ideas, you cannot run inference on someone else's hardware.”
Patrick McCanna · patrickmccanna.net ↗
The article conflates context window technical limits with privacy claims; 35KB prompts are a design problem, not a privacy problem.
-
“if your prompt is 35kb, your prompt is confusing, unfocused, and doesn't work right on any LLM”
DiabloD3 · Hacker News ↗
Large prompts are legitimate for complex agent workloads with structured data; real-world context usage often exceeds advertised limits.
-
“It could be sets of data so the agent doesn't have to collect it every time, like program interfaces, commands, views, databases, tables, data models etc.”
gchamonlive · Hacker News ↗
The real blocker for local LLM adoption is prohibitive hardware costs, not privacy concerns or software architecture.
-
“The main gotcha for local models is insane hardware requirements. Even for $10K you get mediocre performance.”
stackedinserter · Hacker News ↗
Patrick McCanna Developer, authorOpenAI Frontier AI providerAnthropic Frontier AI provider
How it unfolded 3 developments, newest first · click a bar or a number to jump articlespostscomments
-
3
Commenters defend large prompts as structured data, not prose confusion
gchamonlive pushed back against the assumption that large prompts are verbose human text, noting that 35KB could easily be structured metadata—"program interfaces, commands, views, databases, tables, data models"—needed so agents don't have to re-fetch it. This suggests large system prompts may be a legitimate architecture pattern rather than evidence of poor design.
“You are assuming the entirety of the prompt is human prose, but it could be sets of data so the agent doesn't have to collect it every time, like program interfaces, commands, views, databases, tables, data models etc.”
— gchamonlive, Hacker News commenter · source -
> you run out of useful context that the model can accurately attend to around the 250k mark no matter how much they advertise their context size is.This was certainly true when I first tried the new models with a 1M context. After 200k things got weird pretty fast. I haven’t had that problem since Opus 4.8. I’m regularly bumping against 800k…
2 more of the top 3 · 9 posts in this stretch
-
I've been using Claude Code Extension in VSCode (no phone-home configured), backed by DwarfStar on a LAN local MBPro 128GB M5. The context bloat is horrendous, leading to 5-10 minute prefills.I've recently been exploring tools like headroom to help manage context, with some limited "success" (for some definition of success). What do others with…
-
Tbf, didn't read the article because it isn't applicable to me. I don't use system prompts or memory, I just use models stock and write the problem out.Is it really 250k? I had a long running autonomous Astra session today that got to about 600k and it finished fine with everything I asked it to do solved nicely. Opus 5 last week got to around…
-
-
2
Commenters report context performance beyond advertised limits
birdsongs reported using autonomous sessions reaching 600K+ tokens with good performance, contradicting claims that useful context maxes out around 250K. Others noted that complex tasks naturally balloon context size, suggesting real-world usage patterns differ significantly from McCanna's assumptions about prompt bloat.
-
The main gotcha for local models is insane hardware requirements.Even for $10K you get mediocre performance.
-
-
1
Hacker News commenters dispute McCanna's technical claims and scope
Readers pushed back on McCanna's article for lacking substance and technical depth. DiabloD3 argued that 35KB prompts indicate confused, unfocused system design and that practical context window limits are far lower than advertised. SyneRyder summarized the article as a straightforward context-window incompatibility (local models have smaller windows than cloud services) rather than the novel privacy insight McCanna implied. Other commenters noted missing practical solutions for the technical problems described.
“It's become evident that the frontier providers are not only untrustworthy- but actively devious. If you want to protect your ideas, you cannot run inference on someone else's hardware.”
— Patrick McCanna, Developer, author · source -
The article doesn't really describe the problem: if your prompt is 35kb, your prompt is confusing, unfocused, and doesn't work right on any LLM, and is needlessly bloating your context.At this point in time, due to how most people and companies run their inference engine, regardless of the model (yes, this includes the newest from OpenAI and…
2 more of the top 3 · 5 posts in this stretch
-
> Everyone who begins learning exploitation hits a phase of exploitability grief about 3 month into dedicated, practiced study. They hack something they didn’t think they had the skill to break into and it terrifies them. They’re smart enough to know that, relatively speaking, they are an idiot, and if an idiot can do this then nothing is safe…
-
TLDR: Local models have a smaller context window, so your 35kB prompts that worked fine against a hosted 1 Million token window, crash out when you only have a 65K (!) token window locally.I dislike being negative, but I was really hoping for more substance when reading this. It would have been an interesting topic.
-
-
background
McCanna accuses OpenAI and Anthropic of security theater — McCanna dismisses recent public statements by OpenAI and Anthropic about AI security risks as misdirection. He argues that frontier providers knowingly run unsandboxed agent fleets and have deliberately chosen not to implement serious security controls, and that their public security warnings serve a different purpose than actual protection. He brands frontier providers as "pirates" rather than legitimate partners.
-
background
McCanna characterizes frontier providers as deliberately extracting value from users — McCanna argues that the most valuable information frontier providers collect is not raw user data, but the metadata and reasoning patterns in user sessions—the "intuitions" and problem-solving approaches users apply. He alleges frontier providers intentionally train on user activity to extract novel insights, and that when questioned about retention and training practices, providers can only say "Cannot rule it out."
-
background
Developer publishes privacy concerns about frontier AI providers — Patrick McCanna published technical notes on migrating 35KB system prompts from Anthropic's Claude and OpenAI's models to self-hosted Ollama, framing the migration as motivated by privacy risks. He cites recent public drama involving mathematicians and frontier providers around potential training on user activity, and argues that frontier AI companies have become untrustworthy with user data and session metadata.
What people are saying 8 voices from 1 site · best of 16 · verbatim
- What is the actual state-of-the-art for context window performance with large prompts, and what improvements are expected in the next year?
- Can local open-weight models realistically replace cloud providers for privacy-sensitive workloads given current hardware costs?
- Do frontier providers' terms of service actually permit training on user session data, or is this speculation?
- Sep 16
-
Funny times!How many years after "public clouds" and non-local "disks" and "drives" we are ? :) And you still need to tell peoples that other have access to your private data :)Wait, no... They even have access to a thingie you just about to think about! ;) That is a superpower, no less :>And managers are firing peoples just to outsource…
- Sep 15
-
From what I'm seeing elsewhere, context size up to 128k should be possible on this hardware. It really matters for agentic workloads to push that context size headroom up. Anthropic are spoiling us with models that do 500k context and beyond.
- Sep 14
-
The "dumb zone" threshold is fuzzy but comes waay before 250k tokens. Like half that.
-
Can a 27b model even do meaningful security tasks?I thought the interesting cyber stuff is really at the edge of frontier
-
> "until the context rot and sampling problem is fixed forever"I agree, prompt adherence seems to get worse when operating on large inputs. Does anyone have some notion of the SOTA with this? Can we expect big improvements by this time next year? (hopefully in open weights)
-
> your prompt is confusing, unfocused, and doesn't work right on any LLMYou are assuming the entirety of the prompt is human prose, but it could be sets of data so the agent doesn't have to collect it every time, like program interfaces, commands, views, databases, tables, data models etc...I could see this scale to multiple kiltobytes of metadata…
-
The Claude Code system prompt was >50KB, though I think they trimmed it down heavily recently. (The newer models don't need as much hand-holding.)
-
I had hoped to get some new information out of this topic, but unfortunately found the same local "dead-ends" that I explored myself.It unfortunately feels like we will be stuck waiting for a burst bubble before local hardware can be reasonably acquired for personal LLM usage.