Projects

What I build

Working systems in agentic analytics and LLM evaluation — not slide-ware.

Flagship

Agentic Analytics

An LLM agent that answers questions over a real e-commerce warehouse and won't ship a number it can't ground in the data: it plans, runs read-only SQL, and passes its answer to an independent judge before replying. Backed by a 26-question eval — including traps a naive query gets wrong — that measures accuracy and groundedness separately. Wiring an agent to a database is the easy part; the work is knowing when the answer is wrong.

Read the writeup →

Tooling

More soon

Additional experiments in agentic analytics and LLM evaluation will land here.