Spent 3 hours last Saturday getting an old chatbot to actually answer a simple question
I was messing around with one of the open source models on my laptop, just trying to get it to tell me what day of the week a date fell on. Sounds easy, right? Turns out the thing kept giving me the wrong day and making up fake calendars, and I burned 3 hours in my kitchen in Tucson redoing my prompt over and over before it hit me. The model was fine, my setup was the problem, I had the temperature cranked way too high so it was basically guessing instead of doing the math. Dropped it to 0.1 and it nailed it on the first try. Also learned the hard way that these models are way better at writing than at actual counting, so now I just have it hand off math to a calculator tool. Anyone else get fooled by a setting like that and waste a whole afternoon?
Temperature was probably only part of it. These models don't actually do math even at 0.1, they just predict what number usually comes next. Ask it what day a date falls on and it'll confidently say "Tuesday" because that sounds right, not because it counted. The low temp just kept it from wandering into random wrong answers, but it can still get a simple one wrong on a good day. That's why the calculator tool worked, it stops being a guessing game. Same thing happens with counting letters in a word, try asking it how many r's are in strawberry and watch it sweat.