Research community discusses harness design and benchmarking trade-offs
2 Sep 18 9:06 AM · 8d ago · 1 article · 1 post · 6 comments · 2 sources · development 2 of 3
Hacker News users engage with the empirical harness study, debating the generalizability of findings across model families, the definition of terms like 'bash capable' and 'harness,' and whether existing benchmarks capture real-world performance differences.
“we definitely need more principled studies on the role of harnesses. I'd also say that there aren't too many benchmarks where the more complicated harnesses consistently outperform extremely simple agents.”
lieretHaozhe Liu Lead researcher, SoL-Pi studyRun-Ze Fan Lead researcher, harness components studyNVIDIA, NTU, MIT Collaborating institutions
The whole story articlespostscomments the bright band is this development · numbered dots are the others · click one to jump
Reported in the same hours no headline names this development itself — these 1 claim were published in its stretch
-
first by HN Best, 8d ago · also HN Frontpage, arXiv cs.AI
What people said 8 voices · verbatim
-
The conclusions:> Planning improves success at additional cost for weaker models but mainly reduces cost, with small decreases in success rate, for stronger models.> Predefined tools raise success rates for models with weak bash control, whereas bash-only yields higher success at lower cost for bash-capable models, most clearly on shell-centric…
-
Haven't gone through full PDF as its very detailed, few things have resonated with me so far.Basically if a Car A is performing better (be it speed, milage or in general sense) than Car B, then it is not necessarily because its engine. It could be because of better tires, better gearbox, lighter body, better usability of features, etc.You can…
-
Quite inline with what I had found with my Claude code sessions over the last year. I wrote about this a few months ago.https://rahulmax.com/notes/how-i-keep-the-ai-bill-down/In their case, context management pays off more the tighter your window. Their gap between managing and not managing is 35.7 points of success rate at 32k and 2.7 points at…
-
Cool study, we definitely need more principled studies on the role of harnesses. I'd also say that there aren't too many benchmarks where the more complicated harnesses consistently outperform extremely simple agents. But I'm also biased, because I wrote https://github.com/swe-agent/mini-swe-agent/ , which is probably the most minimal agent out…
-
Todo/task-tracking tools (TaskCreate/Get/Update/List, TodoWrite) are no longer available on Opus 4.8, Sonnet 5, Fable 5, Mythos 5, and newer models; set CLAUDE_CODE_ENABLE_TODO_TOOLS=1 to bring them back"Anthropic appears to agree frontier models don't need in-session planning tools.
-
As far as I can tell, the paper says "bash capable", without ever describing what that means. How would one know whether a given model is "bash capable" or not?I would have to imagine, that Luna would very much fall into the camp of "bash capable". At which point- it seems to me that adding any tools beyond just Bash requires some rigorous testing…
-
nice work and paper, seeing more harness benchmarks emerge and we definitely need more. I ran mouse on the Frontier Harness benchmark and scored the highest pass rate, however I'm not convinced the results there actually translate to meaning the "best" harness in practice.Always looking for more harness evals, although I'm going broke running them…
-
This is done on Nemotron models + mistral, so its not very relevant to the current frontier of cheap chinese models + big models from Claude/GPT. Big miss not having qwen or deepseek in this research.
All 3 developments of SoL-Pi: Researchers Cut Coding Agent Costs by a Third →
Hacker NewsNewswiresMastodon