A researcher at Stanford said AI models could hit a wall by 2027 because of one simple thing
I was reading a piece on a tech blog last Tuesday and a Stanford researcher said we might run out of good training data by 2027, which honestly weirded me out. The number that got me was this: something like 90% of the internet's useful text could already be fed to these models within a few years. So companies are now paying places like Reddit and news sites real money just to use their old posts and articles as food for AI. I guess I always figured the internet was endless, but I never thought about the good part of it running out. Does anyone know if synthetic data is actually working as a backup, or are we just headed for slower AI growth no matter what?
Did you catch that the paying part is really the whole story here? When companies start buying old forum posts and news archives, that tells you the free stuff is already picked over, and synthetic data mostly just teaches a model to copy its own habits back to itself. So my guess is we get slower growth and a lot more lawsuits over who owns the text, not a hard wall in 2027.