Evaluating AI developer tools without the drama
Christina Chan |
Thank you to the organizers of LDX3 New York where I originally prestended this version of the slides!
Abstract
Have you ever inherited a technical decision that went sideways? Where teams were running their own experiments, trust was eroded, and every conversation felt like “my may vs your way”?
Last year, I took over an AI code review tool evaluation that was heading in exactly that direction. An enthusiastic rollout had backfired, developers had lost trust, and teams were championing different tools with no shared criteria for success.
Six weeks later, we had a clear winner that exceeded our targets. More importantly, we had buy-in from developers who’d been skeptical that any AI tool could work, including those whose preferred tools didn’t make the cut.
The evaluation process was featured in The Pragmatic Engineer newsletter, but this talk goes deeper into the framework that made it work. I’ll share the methodology we used to turn a contentious tool evaluation into a process everyone could trust.
You’ll leave with a reusable framework for making technical decisions where stakeholder buy-in matters as much as the metrics.
Slides

Check back later!