Many engineers have written a query that seemed to work, only to discover later that spend in AWS
or Snowflake unexpectedly spiked. SQL optimization looks simple until you are in the middle of it:
there are many angles to attack the problem, and the “right” answer depends on many interacting
factors.
You can throw a query at an LLM, ask it to rewrite, and ship the result, but how do you know the
optimized query actually does the same thing? This talk explores why query optimization is a
deceptively hard problem, not just computationally but mathematically. Using production examples,
we examine what “optimal” really means, what suboptimal queries cost, and why naive first
solutions do not hold up under scrutiny.
We tackle the “but couldn't we just...?” questions head-on: why you cannot reliably sample a
database and compare two queries, and why asking an LLM whether two queries are equivalent
confuses confidence with correctness. LLMs learn language patterns, not algebraic ones, and
struggle with query equivalence for similar reasons they struggle with chess-like exact reasoning.
From there, we explore more rigorous alternatives: rules-based approaches like incremental view
maintenance, the BAG algebra underneath them (more accessible than it sounds), and formal
verification methods that can actually prove equivalence. We close with a practical framework for
selecting the right technique, because sometimes a heuristic is exactly right, and sometimes only
a proof will do.