D
3

Fine-tuned Llama 3 vs GPT-4o for support ticket triage, the gap was huge

I ran a side-by-side test last Tuesday using 500 real tickets from our help desk, canned responses and all. The fine-tuned Llama 3 got the intent right 92% of the time, GPT-4o was at 71% and kept overcomplicating simple password resets. Has anyone else seen open-source models beat the big APIs when you actually feed them your own data?
1 comments

Log in to join the discussion

Log In
1 Comment
king.kevin
Did you run the same ticket set through both after tweaking your prompts?
5