Off to bed.

Another ‘cloud gaming’ vendor has put restrictions on the number of hours you can use their service per month and that follows nVidia doing the same thing on their service back last year in December. You’d think that these AI obsessed twits would have learned from their ‘cloud gaming’ fiasco but alas they’ve convinced themselves that something that is even more hardware and energy intensive such as running LLMs in the cloud is going to be the big winner. Microsoft is the one that has introduced restrictions (link) – the subscription service for games (aka attempting to be the Netflix of games) didn’t work for obvious reasons and the ‘cloud gaming’ is turned out to be unprofitable or not profitable enough to make it worth the amount invested into the infrastructure.

As I’ve mentioned in the past, the economics of LLMs in the cloud is based on this idea that you first build out the infrastructure then once economies of scale are reached the costs plateau, the revenue grows and voila you have a profit. You can companies taking the SaaS model and believing they can apply it to everything and anything but that simply isn’t the case – cloud gaming has shown it cannot be done unless with limits and now LLMs subscriptions have usage limits. The whole business model is a house built on sand and in the case of cloud gaming it is a solution search of a problem – the phones are already fast enough to run modern games so how about just making your games run better on the hardware instead of delivering games in the most inefficient way aka the cloud.

I think the interesting thing to see will be when the nVidia RTX Spark devices start shipping, both laptops and desktops, and whether we’ll see people move over to local models. Claud code can be use used with a local model and if you match it up with Gemma 4 model and given that the nVidia RTX Spark device can come with up to 128GB then it’ll have more than enough processing grunt and memory to run the model. Same can be said with running Gemma 4 on a Mac Studio with an M5 Ultra and 128GB – plenty of power without having to worry about on going to costs of having to pay per million tokens. Personally I think the greatest threat to LLMs in the cloud isn’t people protesting data centres but rather people realising that they can grab an open weight model, run it on a decent machine and now having to worry about on going costs other than powering the machine they’re running it on.

Leave a comment

This site uses Akismet to reduce spam. Learn how your comment data is processed.

I’m Matisyahu

An open diary of my life as I navigate the world – the terrifying lows, the dizzying highs, the creamy middles.

“When the people are being beaten with a stick, they are not much happier if it is called ‘the People’s Stick’” – Mikhail Bakunin

Let’s connect