0
Fine-tuning beats big models for my small dataset, and here's the proof
Everyone's chasing the latest 100B parameter model, but I got better results last month fine-tuning a small open-source model on my own data. I run a niche legal document classifier, only 4,200 examples total. Tried GPT-4 and Claude for two weeks, they kept missing the specific phrasing my clients use. Switched to a 7B model with LoRA on a single A100, hit 94% accuracy in 3 days of training. The big models maxed out around 88% no matter how much I prompt engineered. Anyone else found that specialized small models beat the giants for narrow use cases, or did I just get lucky with my data?
0 comments
Log in to join the discussion
Log In0 Comments
No comments yet
Be the first to share your thoughts on this discussion.