Switched my team from GPT-4o to Claude 3.5 Sonnet for code reviews and the bug catch rate jumped
We ran both on the same 200 pull requests from our repo last month and Sonnet flagged 47 real issues versus 29 for GPT-4o, with way fewer false alarms too. The big difference was it actually explained why something would break instead of just flagging style stuff. Anyone else seeing this or did we just get lucky with our codebase?