conv.

All stories
AIQuiet 9d · day 11

Researcher trains 4B model to optimize Postgres queries 81% faster

Rohan Bansal uses reinforcement learning to teach Qwen to beat Postgres's default query plans.

What to know

  • A 4B-parameter Qwen model trained via SFT and reinforcement learning achieved 81% faster query execution than Postgres's default optimizer on tested plans.
  • Query optimization is a strong fit for LLM learning because performance is easily measurable and verifiable—faster execution is the sole optimization metric.
  • Join ordering, a core query optimization task, is NP-hard, explaining why traditional optimizers leave significant room for improvement despite decades of research.

Rohan Bansal Researcher

Researcher trains 4B model to optimize Postgres queries 81% faster
rohanbansal.com

How it unfolded 1 development · click the chart to see its coverage articlesposts

Peak 8 pieces in 3h at Sep 16, 5 PM; 25 pieces over 11 days (2 articles · 4 posts · 19 comments) Sep 16, 2 PM — 7 pieces · 2 articles · 1 post · 4 comments — Hacker News 5, Newswires 2Sep 16, 5 PM — 8 pieces · 2 posts · 6 comments — Hacker News 6, Mastodon 2Sep 16, 8 PM — 3 pieces · 3 comments — Hacker News 3Sep 16, 11 PM — 1 piece · 1 comment — Hacker News 1Sep 17, 2 AM — 2 pieces · 1 post · 1 comment — Mastodon 1, Hacker News 1Sep 17, 5 AM — 1 piece · 1 comment — Hacker News 1Sep 17, 8 AM — quietSep 17, 11 AM — 1 piece · 1 comment — Hacker News 1Sep 17, 2 PM — quietSep 17, 5 PM — quietSep 17, 8 PM — quietSep 17, 11 PM — quietSep 18, 2 AM — quietSep 18, 5 AM — 1 piece · 1 comment — Hacker News 1Sep 18, 8 AM — quietSep 18, 11 AM — 1 piece · 1 comment — Hacker News 1Sep 18, 2 PM — quietSep 18, 5 PM — quietSep 18, 8 PM — quietSep 18, 11 PM — quietSep 19, 2 AM — quietSep 19, 5 AM — quietSep 19, 8 AM — quietSep 19, 11 AM — quietSep 19, 2 PM — quietSep 19, 5 PM — quietSep 19, 8 PM — quietSep 19, 11 PM — quietSep 20, 2 AM — quietSep 20, 5 AM — quietSep 20, 8 AM — quietSep 20, 11 AM — quietSep 20, 2 PM — quietSep 20, 5 PM — quietSep 20, 8 PM — quietSep 20, 11 PM — quietSep 21, 2 AM — quietSep 21, 5 AM — quietSep 21, 8 AM — quietSep 21, 11 AM — quietSep 21, 2 PM — quietSep 21, 5 PM — quietSep 21, 8 PM — quietSep 21, 11 PM — quietSep 22, 2 AM — quietSep 22, 5 AM — quietSep 22, 8 AM — quietSep 22, 11 AM — quietSep 22, 2 PM — quietSep 22, 5 PM — quietSep 22, 8 PM — quietSep 22, 11 PM — quietSep 23, 2 AM — quietSep 23, 5 AM — quietSep 23, 8 AM — quietSep 23, 11 AM — quietSep 23, 2 PM — quietSep 23, 5 PM — quietSep 23, 8 PM — quietSep 23, 11 PM — quietSep 24, 2 AM — quietSep 24, 5 AM — quietSep 24, 8 AM — quietSep 24, 11 AM — quietSep 24, 2 PM — quietSep 24, 5 PM — quietSep 24, 8 PM — quietSep 24, 11 PM — quietSep 25, 2 AM — quietSep 25, 5 AM — quietSep 25, 8 AM — quietSep 25, 11 AM — quietSep 25, 2 PM — quietSep 25, 5 PM — quietSep 25, 8 PM — quietSep 25, 11 PM — quietYesterday, 2 AM — quietYesterday, 5 AM — quietYesterday, 8 AM — quietYesterday, 11 AM — quietYesterday, 2 PM — quietYesterday, 5 PM — quietYesterday, 8 PM — quietYesterday, 11 PM — quietToday, 2 AM — quietToday, 5 AM — quietToday, 8 AM — quietToday, 11 AM — quietToday, 2 PM — quietToday, 5 PM — quiet 1
Sep 17Sep 18Sep 19Sep 20Sep 21Sep 22Sep 23Sep 24Sep 25yesterdaynow · 8:33 PM ET
  1. 1

    Bansal publishes query optimization experiment results

    Rohan Bansal published a detailed technical article describing how he trained a 4B-parameter Qwen model to generate Postgres query plans that execute 81% faster than Postgres's default optimizer. The approach combines supervised fine-tuning with agentic reinforcement learning, where each query generates four RL rollouts, with Qwen producing candidate plans that are tested against Postgres for scalar rewards.

    “can a small, open-weights model be post-trained via supervised fine-tuning (SFT) and agentic reinforcement learning (RL) to produce Postgres query plans that beat Postgres's default plans?”
    — Rohan Bansal
    1. first by HN Best, 11d ago · also HN Frontpage

    • The end game is adaptive query plans.A big reason the initial plan isn't guaranteed to be optimal, even with all the right indexes, is that table statistics aren't perfect. For example, you might track a column's correlation (how closely the column's logical ordering matches its physical ordering in the heap), but that won't be broken down at a…

      rand_rHacker News10d agoview on Hacker News ↗
    2 more of the top 3 · 19 posts in this stretch
    • I had the similar feelings, the setup is biased for certain outcomes it feels. At times I feel like that I am in an eternal questioning mode but then again I find it to be a better choice to be critical and skeptical for technology related things.Couple things that I found interesting1. Inefficiencies/limitations of the query planner in certain…

      sandeepkdHacker News10d agoview on Hacker News ↗
    • > I paid ~$800 to rent a 2x H100 SXM node from Lambda for ~95 hours, and ~$400 in OpenAI API fees to generate the Astra trajectory demonstrations.> a tiny 4B model went from not being able to understand the harness it was wrapped in, to achieving a 1.81x geometric mean speedup and a summed latency decrease of 44.7% across a workload of join-heavy…

      SomeoneHacker News11d agoview on Hacker News ↗
    all of them →

What people are saying 16 voices from 1 site · best of 19 · verbatim