Commenter surprised by Pi harness efficiency matching larger models
7 Sep 17 9:53 AM · 9d ago · 2 posts · 3 comments · 3 sources · development 7 of 7
A user expresses surprise that the Pi harness achieves similar efficiency to Claude Code and Codex, and wonders whether the same holds for smaller 9-32B models that would presumably require more steering.
“The most astonishing thing to me is that Pi harness is basically as efficient as the Codex/Claude.”
Otterly99matt_d HarnessTax research author/submitterYashjain413 HN commenterSupermancho HN commenternojs HN commenterlukax HN commenter
The whole story articlespostscomments the bright band is this development · numbered dots are the others · click one to jump
What people said 11 voices · verbatim
-
I think it’s more likely that models have random behavior, and so the results in these dimensions are just random.
-
I find that in many cases, you ARE the harness. Case in point, Terence Tao utilized the simple chat interface to find the Jacobian Conjecture counterexample. This is arguable no harness at all. I also find that many people will disagree with me but most of the time, they just want a button to press that will solve the problem. If that's the work…
-
> Claude Fable 5 solves 97.8% of attempts in Claude Code, 96.7% in Codex and 96.7% in Pi, yet Claude Code costs about twice as much as Pi ($1.33 vs $0.67). If the harness tax is much lower than the subscription subsidy, it makes sense to not use custom harnesses like pi. Right? It’s a shame i would really love to use alternative, less buggy…
-
I came to conclusion that all I really care about from the harness is interoperability. I want to be able to switch providers, models at any point in any session or between sessions in a single action, and I want my transcripts to be in one single format from the beginning till the end of time.I want one place where I define project prompt, one…
-
If you use the subscription your usage is heavily subsidized over the api billing. If you use the subscription you can’t use other harnesses. So to use another harness you have to give up the subscription subsidy,
-
We really do need better benchmarks and for models too- Most people use a harness because of its subscription (most companies pay Anthropic) - All model benchmarks are biased and gamed, harness benchmarks are too few to matter - Everyone is just guessing, acting on sample sizes of 1 and trust me bro vibes
-
sub subsidy? are you talking about another adjustment that happened after their billing retraction?
-
The most astonishing thing to me is that Pi harness is basically as efficient as the Codex/Claude.I wonder if the same is true for the smaller models in the 9-32B range? I would expect that these models need more steering, but again I was not expecting this result either.
-
To clarify, this is true for Claude. A ChatGPT sub can be used in other harnesses.
-
I also test harness with "Ship Harness Bench", jcode with its browser integration gives the best results
-
The problem is that Pi, on any third party harness, cannot really compete with codex or cc due to subscriptions.
All 7 developments of HarnessTax study: how much does the harness matter for… →
Hacker NewsMastodonNewswiresLobsters