I’ve got a few things going through my mind at the moment in regards to the AI bubble such as the systemic financial risk of having so much debt being tied up in something that has a dubious business plan attached to it. I was watching the following video over on YouTube:
I’m going to split this blog post into two areas, the first being the financial side of it in terms of the business and the second being the technology itself. When it comes to the financial side of it I see it as questionable whether there can be a sustainable business developed around providing LLMs in the cloud given that the whole gamble is based on the same logic that SaaS and cloud infrastructure was based. The logic being, build up the infrastructure, reach economies of scale and at that point the cost of delivery plateaus then at which point you start seeing profits. That makes sense in the case of cloud computing and SaaS because there is the cost of the initial build out but once you reach the economies of scale with each additional customer the cost of servicing them is next to nothing resulting in high profit margins and low overhead costs. This is the reason why for the last almost decade and a half the SaaS industry has been pretty much a licence to print money and the wizkids over in VCs thought they could replicate the same model for the AI industry.
Let us assume that Alex Karp is correct about Sovereign AI, do hyperscalers become little more than cloud providers who simply provide the ‘bare metal’ infrastructure and then the customers load their own operating system, models, middleware etc. and pay for the capacity they need instead of paying for API tokens from Anthropic or OpenAI to access their frontier models? Then there is the question regarding running the models locally given that there has been improvements in NPUs, the mainstreaming of unified memory architecture, the models are having their memory footprint reduced without meaningfully impacting output or performance. There is is the rise of new models being developed such as Inkling from Thinking Machines and bespoke models that do one and one thing very well such as the recent announcement regarding a Cisco security model. If a significant amount of the workload can be done on device then what will happen to all that capacity that was built out only for it not be used – keeping in mind that every year that passes where the hardware isn’t be utilised is another year closer to when the hardware will need to be written off and then replaced because it’ll be out of date and less efficient.
Now there is the matter of technology itself and I’m personally sceptical of its utility outside a small number of areas where it makes sense. Is it to say that the technology is entirely worthless? no, but I do question the exaggerated claims that the ‘high priests’ of AI who have put them out there with the unquestioned relaying by the media. I also question whether it is going to be a multi trillion dollar industry given the rise in the quality of local models, the performance/price and performance of NPUs that are now becoming standard, nVidia with their RTX Spark platform. With all that being said, a lot of the hype is a solution in search of a problem – if a technology was actually useful they wouldn’t need to spend billions hiring social media influencers and celebrities in trying to convince people of the usefulness nor would there be a backlash by normies who feel as though it is something that is intrusive and not something they opted in for.
In the past the technologies that have arrived were readily apparent in what benefits it brought to the lives of ordinary people. I remember when mum bought a fax machine and rather than having to send a letter to nana she could send a fax and nana would get the document. Then email arrived and that was even better – you could attach a document or any file you want then send it through. Broadband then made it possible to have cloud based storage, work on documents collaboratively in the cloud, streaming videos on demand, stream live events such as when Apple or Microsoft had a keynote but you couldn’t physically attend because it was on the other side of the world. With each of the technologies it didn’t require billions of marketing because the benefits were obvious – it didn’t require people to convince them of a problem that didn’t exist and then sell a solution that kind of addresses the problem in an unreliable nondeterministic way.
Have there been some interesting advances? sure, LLMs have improved converting text between languages, help with code coplemention, auto correct, grammar checking along with improved reliability of voice assistants. Most recently Google through the use of AI can translate sign language into text so then a non fluent sign language person can understand it – all of that done locally. So I keep coming back to the question – is it really a trillion dollar industry once you realise that a lot of the stuff that is being marketed is actually being delivered as local models on device? If Anthropic and OpenAI got rid of their free tier and subscriptions in favour of token billing (pay as you go) could they be profitable? I guess but at that point you really have to wonder how big the market is – people are happy to overlook wasted tokens when they’re not paying for them but the moment that the incorrect response is provided thus requiring more tokens to be wasted I think you’re going to find that even businesses who were all in on ‘tokenmaxxing’ are going to question what they’re getting out of all this.
