Really interesting argument, and I think you’re right about one major point: Apple’s unified memory architecture is seriously underrated for running very large models locally.
Where I’m less convinced is the jump from “Apple is unusually good at local inference” to “Apple is the king of AI and NVIDIA is finished.” A Mac Studio and a data-center GPU cluster solve very different problems. Fitting a huge model into memory for one user is not the same as serving thousands of users with high throughput, batching, reliability, and fast interconnects.
The comparison also feels slightly selective when a single workstation GPU is used to represent NVIDIA’s entire platform. Apple may have a real advantage in private, local, high-memory inference, but that does not automatically translate into dominance in training, cloud infrastructure, or large-scale deployment.
So I agree with the underlying trend, but not the conclusion. Apple may become the leader in local AI hardware. That is a strong thesis on its own, without needing NVIDIA to be “dead.”
Most commentators disagreeing on the conclusion are missing the point of the post. If my inference requirements aren’t humongous (small enterprises) and renting intelligence from hyperscalers is costing me a fortune, not just because of cost differences pointed out in the post but also the massive physical build out of the DCs (real estate, network latency etc), why would I not move my inference costs in-house if I can afford it without needing massive debt to fund it. This isn’t hypothetical for me. I’m planning to do exactly that. To me this is a simple cost-benefit analysis. This option hasn’t existed so far but now it does and I see no reason not to exercise it.
Agreed Andy. Those disagreeing with the conclusion are stuck on the hyper-scale, centralized model as inevitable and necessary, in the same way centralized cloud has been treated as necessary and inevitable. It’s necessary for market dominance of a select few mammoth players - the salespeople pushing inevitability - but not for the full spectrum of users and use cases.
Spot on Chaos - the complexities of training and serving networked models is well outside Apple's strategy, local models have a role, and AS has a space in that local market, but I can't see Apple pivoting to take on the DC space.
When this much disruption to what was dreamed is already in place, I don’t think we can imagine all the ways that creative people will come up with to save massive amounts of money.
That is possible in theory, but filling a data center with M5 Ultra machines would not automatically solve the throughput, interconnect, software and scaling problems. Apple Silicon is efficient, but efficiency per machine is not the same as efficiency per token served at scale. Until Apple shows a serious data-center platform, that remains a hypothetical rather than evidence that it can compete with NVIDIA’s infrastructure.
Alright, I agree with the premise, the trajectory (commoditization of compute), and most of the historical analogies.Where I disagree is the prediction of decentralized compute.
All historical analogies point to huge economies of scale from *specialized* compute and SaaS. I suspect that Google and Apple (and maybe AWS) will emerge as winners exactly because they're so good at efficient Ops and delivery (which includes their extensive content networks: app store, browsers, assistants) and, more recently, custom hardware.
The privacy concerns are real, and I can imagine a proliferation of SaaS vendors for law firms and medical practices, but I have a *very* difficult time imagining a law firm procuring, installing, and maintaining their own cluster. I don't believe the expertise or appetite is there.
On-device compute for cell phones? 100%, and that will be owned by Apple and Google. Salesforce-type products targeting specific professions, and research-grade cluster for universities and R&D labs? Yes and yes. More on-vehicle heavy compute? Likely (another place I expect US mfgs to drop the ball).
Exactly. That is also where I think the market is heading: much more on-device inference, but not the disappearance of centralized compute. Most organizations will want the privacy and control of local AI without becoming infrastructure operators themselves, which leaves a large role for specialized SaaS and hybrid systems.
I’m not assuming Apple has no data-center experience. Private Cloud Compute clearly proves otherwise. My point is narrower: Apple has not yet demonstrated a general-purpose AI serving platform that competes with NVIDIA at scale.
The “one twentieth of the energy” claim also needs an apples-to-apples comparison: same model, quantization, latency, batch size, throughput, and whole-system joules per token. A machine drawing less power but serving far fewer tokens per second is not necessarily cheaper or more efficient at scale.
And Apple’s own architecture currently supports a hybrid conclusion, not full decentralization. Apple explicitly sends workloads that are too complex for on-device execution to Private Cloud Compute. Local inference will clearly grow, but “data centers will be replaced” is a much stronger claim than the available evidence supports.
> the whole point of this piece is that data centers for inference will be replaced by local hardware
If that is your whole point, NVIDIA is not even the main rival. Amazon and Google both have their own specialized inference chips which Apple has to beat. Would look forward to your analysis of Amazon Inferentia and Google TPU 8i.
Then they'd hit the same bottlenecks (HBM, power, DC build rates, copper etc) the current DC build-out companies are facing, rendering any wished-for savings as moot, probably.
I don’t think anyone is suggesting that AI DCs will become obsolete. What is being suggested is that there is another way, to own the inference hardware and for some it is far more affordable and even practical to do so than a frontier model token cost. Which in turn means that cost of inference has to come down. There are other reasons for costs to drop and this is one of them.
We are talking about years to allow this happen, specially with the current RAM prices, but the article is missing someone. Chy-na.
I'm sure by the time hardware to run AI locally is affordable and available Chinese will have already similar or better hardware than Apple given they're full steam building their lithography machines.
Also Huawei is making their own Kirin chips, Alibaba has their AWS Gravitron ARM equivalent and they also just started to make RAM.
Interesting times ahead, in very excited to see the outcome of this in 20 years, look at internet in the 90s and what has become, apply that to AI and see how it speeds up exponentially and having a big player la China also hands on it.
Market share of privately owned on-device inference products is at 0% so currently growth potential is infinite.
Reasons to pay top $$ for self owned high performance devices for inference, comparing to paying for that performance as a cloud service, are a matter of pricing, priority, privacy, regulation or national security.
Should I pay now for hardware that could run today's small-medium size models, rather than paying on the go higher prices. Take into account that a chip depreciates within a couple years, while cloud compute is dynamic and up to date allowing you to easily switch to a different hardware.
I appreciate the expert analysis—especially since it confirms my inchoate suspicions about this whole cluster we’re watching unfold. Apple is the second mover here—they stood back and watched the land rush unfold in the wrong direction and are using those mistakes to shape their own course. Really excellent work sir.
Great piece - well, the parts I could understand - LEJ. Does fractal computing play into the NVDA is dead as well. I don’t understand the technical side of fractal but I think I’ve read it can run large data on a laptop.
Agreed. I guess I was involved enough with the argument that I skipped past what I usually stop reading at. How do you suppose those constructions came about? “Intelligence” unconnected with anything human except a dictionary
First off, This is an opinion piece. second, your statement seems to imply that you take things as true unless you suspect they were written by A.I.?!?!?
If that's the case man, I would not be going around giving advice. You might want to have AI do a pass on your next comment before you make it.
Nvidia has already announced DGX station with 748GB of memory. And it has Grace Blackwell. I bought a 64GB Mac Mini for some of the reasons you cite. But I don’t think Apple has a moat should Nvidia prioritize this space.
You could be right and honestly it would be absolutely fucking amazing if you were.
To be clear, what you could be right about is that Nvidia might be really diving in and focusing on this section, which means Apple wouldn't have a moat. They would still be extraordinarily competitive in keeping the competition up and running, which I think is a huge win
It is estimated the DGX Station is 10-20x more throughput and up to possibly 50x on long prompts and high concurrency. It’s estimated about 7x more expensive. If these are true, it’s better cost efficiency. Not all users need it so Apple has a place but it seems companies with many users would pick DGX. All these are estimates of course. But even if off a bit, Apple being king of AI and Nvidia walking dead doesn’t seem likely at all. I say that liking what Apple has done and spent money there.
Like I said I hope you're right. It could be pretty amazing to see that kind of intense competition going on. The writing's on the wall though and I try to stress in my piece that this is all predictions but I want to emphasize that again.
I am predicting that Apple is about to launch a public hardware initiative that is going to be extraordinary. I can see all the puzzle pieces falling into place right now, especially with MLX. You get developers developing on MLX first, specifically because MLX has CUDA support. That's Opening the door for cross-compatible everything, which frees up Apple to do some pretty insane stuff.
that point about NVIDIA Link that I wrote about in the article is bigger than you might think. Thunderbolt 5 is an order of magnitude cheaper than NVIDIA Link and with RDMA you're looking at extremely capable interconnected systems.
Also from everything I'm hearing, if my memory serves me correctly, that 768 GB system that NVIDIA is planning on putting out has a projected cost of over $100,000
Compare that to predictions for the M5 Mac Studio, specked out to 768 GB per unit at around $15,000
Connect two of those fuckers With a $70 Thunderbolt 5 cable and you could run Kimi K3 f/P/8 with limited quant - connect 3, and you can run the whole thing. Connect 4, you can run different things in parallel, simultaneously. You still haven't gotten to the cost of one of NVIDIA's Enterprise-grade or Corporate-grade workstations with 760 GB of RAM.
you'll have to pardon me because I'm using voice-to-text right now so some of that might have gotten scrambled but I think you see where I'm going with this
Interesting perspective. I have looked at the Apple option to run local AI model for a project but paused for now. Then I did see that NVIDIA machine being talked about. Have you looked into AMD machines as well and how they stack up compared to Apple? At first glance seems a bit more cost effective. But I havent done a detailed side by side yet.
This was my question so I’m glad to see it here, but I’d be curious if you have a longer answer: “what’s Apple’s moat?”
What I mean specifically is your argument for at-home/work compute is persuasive, but why is it going to be from Apple *long term*? Security doesn’t do it for me - the world is chock full of Windows machines made in China that basically come pre-compromised, businesses buy them hand over fist. Apple has the idea first, okay - what prevents a dozen other companies doing what you say Nvidia has already done, and competing away those fat profits?
He’s literally describing unobtainium and saying the Apple 512 GB RAM is the present and future. But it’s not even listed on Apple Store with more than 96GB RAM anymore. I see eBay sales at over $20k for the discontinued 512 GB model.
𝙼𝚊𝚗𝚢 𝚙𝚎𝚘𝚙𝚕𝚎 𝚠𝚘𝚞𝚕𝚍 𝚗𝚎𝚟𝚎𝚛 𝚝𝚘𝚞𝚌𝚑 𝙰𝚙𝚙𝚕𝚎 , 𝙽𝚅𝙸𝙳𝙸𝙰 𝚑𝚊𝚜 𝚋𝚎𝚎𝚗 𝚐𝚒𝚟𝚒𝚗𝚐 𝚞𝚗𝚕𝚒𝚖𝚒𝚝𝚎𝚍 𝚊𝚌𝚌𝚎𝚜𝚜 𝚝𝚘 90 𝚖𝚘𝚍𝚎𝚕𝚜 𝚏𝚘𝚛 𝚜𝚎𝚟𝚎𝚛𝚊𝚕 𝚢𝚎𝚊𝚛𝚜 𝚗𝚘𝚠. Had it not been for NVIDIA I would have not had the compute available in order to teach myself as much as I have about 𝙰𝙸, 𝚜𝚘 𝚜𝚑𝚘𝚞𝚝 𝚘𝚞𝚝 𝚝𝚘 𝚝𝚑𝚎𝚖, meanwhile I still was never able to get back my Apple account that some girl stole off me somehow when I logged into a phone of hers or something I don't know how she actually got it but despite providing like all types of validation information it didn't work all the way and I just gave up they keep locking it even though I would validate it get it unlocked it would just lock it again so they're garbage I stopped even caring if my work is on Apple podcasts
There's another point that isn't being made but points strongly towards your Apple argument.
Frontier models try to be good at everything, they have to be because they don't know what they're going to be asked. That's not how human society is structured. We divide our cognitive tasks between generalists and specialists, between broad but shallow(er) knowledge and deep specialised knowledge.
These functions are distributed. I foresee the future of AI at work being the human generalist -- AI can't compete on, say, and architect's flare for what a client would like, that's emotional, or a GP's ability to empathise with a sick patient -- working with the AI specialist: human architect, AI structural engineer. That doesn't demand a frontier model. It needs an LLM connected to a physics model, constrained by an ontology of physical, engineering and structures concepts so it doesn't hallucinate, a trained on a specialised corpus of engineering texts, building regulations and the like. It is smaller by far.
Very little, it's basically what "mixture of experts" models do today, but loading and unloading models as demanded by a supervisory model. The difficulty at present is that everything an LLM knows about is encoded in its weights as determined during training. That will, I think, change as LLMs become grounded in knowledge bases and ontologies, reasoning over them as well as language. That expansion is critical to their success. They can't keep getting bigger and, as knowledge expands continuously, it simply isn't economically feasible to continuously train new ones; they have to acquire a way to learn.
NVidia runs the datacenters that power the big companies offering AI subscriptions to frontier systems like ChatGPT and Claude. Apple makes computers that you can run in your own home and office that are big enough and have enough memory to run the open-weights models like Kimi, Qwen, GLM, DeepSeek, MuseSpark, and Mistral.
These open weights models aren't at the cutting edge - but they're only about 4-6 months behind the frontier systems. Running these models on your own computer is also a lot slower than accessing the frontier models on a datacenter - a coding task that would take half an hour on Claude Fable or GPT 5.6 Sol might take 5-10 hours on your own computer (assuming the local model is smart enough to do it right).
But your only cost is electricity, none of your data leaves your building, and it doesn't matter if Trump throws a tantrum because he can't cut you off from your own computer the way he can cut you off from Claude.
This article is claiming that over the next few months or years, the latter advantages will eventually win out over the disadvantages, and so Apple with local open weights models will eventually be the big winner over NVidia data centers with frontier models. I suspect there will be some users in each of these camps, and the big question is where most of the value is.
Good summary. I'm pretty dumb but that's what I got from the article too. Historically Apple has a habit of coming out with things later than others (phones) but with hardware that's infinitely better. So that tracks. Historically tasks that only ran on massive, expensive machines in top-secret locations evolved so that teenage girls could do them on their phones. So that tracks too. AI will likely head in the same direction. Bleeding edge AI might not happen on your Apple Watch, but for most consumers - and businesses, close enough is good enough. iPhones will ship with it built-in AI soon enough I imagine. The counter-argument, is most people don't care. Whether my photos are in the cloud or on my phone, doesn't matter as long as they're there. AI running stand-alone on my desktop or via internet doesn't matter much, as long as it works. Counter counter argument, businesses deffo want to avoid cloud if they can. The company I work for will never trust cloud. By the time they vet the providers the tech has moved on. I can see them loving a machine they can run in-house. But then, that requires people to run it... I'm just going in circles here as I think it through it out loud...
that’s what I got from the difficult article too, and I was already moving to a similar model with a few closed off iPads that are disconnected from the internet to hold data.
I don’t like inputting data into the big AIs because it feels like they can use my inputs against me somehow.
There isn’t any moat in the model. It’s always going to be a race to the bottom and the open weight models will win. Similarly there isn’t a moat in the data center either. As soon as models are good enough to run locally, workloads will move there. The data center build out is the dark fiber of this bubble.
The only real moat is the user experience and connecting with the user. Apple and Google own that connection and will continue to do so.
So baring the arrival of a new computing platform for the AI era, we’ll still be using phones and PCs. Apple and Google end up the winners, everyone else fights over the scraps.
After their initial misfire, Apple understands where the puck is heading and has positioned themselves to be in the right place at the right moment and as your article points out NVIDIA don’t really have any good answers when that moment arrives.
The other parallel I would bring in, is the mainframe/workstation/ desktop/ phone, division that evolved.
Central processing resources are just the first stage of the deployment of technology development.
AI is just a sophisticated averaging machine. All processing power does is compute a slightly better average. The data centre was always a temporary solution 🥰
I felt that this was a weak part of the article. The move from mainframe to desktop has not been a one-way move - at various points in the history of computing we've gone through a pendulum. All our media lives in the cloud now rather than on our personal devices! I think it's far from obvious that the cloud won't win on AI!
The need for control of the tools we use dictates otherwise. The swing of the pendulum was first dictated by processing cost, then storage cost, storage accessibility & sharing. Each time the task could be done as cheaply as a shared resource, the process devolved to the individual 🥰 the next cost centralisation will be quantum computing.
I think you have a strong thesis here, but it operates within a binary that is, in my opinion, a bit myopic. You actually almost touched on it near the end when you said:
> nobody trusts the centralized bet anymore.
...but you're framing "centralized vs. decentralized" in the narrowest scope possible, and betting on a company that is famous for its centralization (or as it's been known for decades, the "walled garden").
Where I think you're right is where the architecture moves; I'm currently involved in a few different projects which drive for local first models that can call on frontier model compute when needed. But where you seem to have completely ignored is in decentralized compute itself. Bittensor, Akash, Acurast, and many others are already scaling decentralized inference that doesn't require local big boy boxes at all. Akash in particular has been breaking ATH usage records month over month serving Qwen and DeepSeek.
If I own the cost of the local hardware, then I own the cost of maintaining and updating it. That's always been the advantage of the cloud. The disadvantages of the centralized cloud are exactly the ones you pointed out...and that's exactly where the DE-centralized cloud shines.
A bet's no fun if there's no one on the other side of it, so I'll take that position to help add some excitement around it. Here's my timestamped claim:
Apple grows market share based on local AI usage, but it's primarily scoped to their existing market, and doesn't meaningfully drive multiples. Enterprise CTOs broadly reject replacing frontier model usage with local models for the same reason on-prem solutions are so minimal in every other place in the enterprise: support and liability. Cost isn't enough of a factor once the actuarials bake in those secondary costs. MSPs continue to own the marketplace by directly targeting those exact layers. And there's not enough of a consumer level marketplace willing to spend 10K on an AI box in an economy where inflation and unemployment are both high, and Apple can't even maintain market share dominance on iPhones. ( https://finance.yahoo.com/markets/stocks/articles/prediction-apple-soon-surpass-nvidias-072201027.html ).
my bet isn’t on the consumer market (they will be happy with the upgraded siri and chatbots) my bet is on Apple disrupting the dev community. I would happily find a way to pay 30k-40k to get a box that hosts my own frontier level ai with zero api costs and Enterprise is already spending orders of magnitude more than that in api costs per month.
> I would happily find a way to pay 30k-40k to get a box that hosts my own frontier level ai with zero api costs
That's a quarter of the annual salary for the average US dev, and up to half for devs globally. It's a pretty small market that would be able to happily find that way.
> Enterprise is already spending orders of magnitude more than that in api costs per month
And will continue to do so when it comes with third party support and liability. If I'm running my own boxes, I have to run my own support, and I own the liability if something bad happens. Deploy AI code that breaks the software and costs millions? I want to point the finger at Anthropic, not myself (the CTO), and that's where the cost premium comes in.
Every major enterprise app comes with either fully cloud, or heavily third party supported on-prem solutions specifically for this reason, which is why MSPs own the Microsoft space. Unless Apple gets Deloitte and Accenture on-board with owning the entire life cycle of deployment, support, and liability, I just don't see broad scale enterprise adoption of on-prem boxes. The vast majority of cloud services are more expensive than running local solutions where you build once and do small iterative updates over time, but core price isn't the consideration that drives their choice.
I wouldn't discount the consumer market. That's what made Apple one of the biggest most profitable companies on the planet. I remember when they made specialist computers for design and advertising. Big money, but very limited audience. Virtually nobody had heard of them. With the iPhone and iPad they appealed to consumers and have grown exponentially ever since. I agree that Apple will come out as one of the AI leaders despite a slow start - that's just their way. Watch and learn, look before leaping. And we'll be laughing at massive data centres the same way we laugh at the massive Mini Vac computers etc - when we now have more compute power on our wrists.
the local inference play is the part nobody talks about. nvidia's whole business is selling shovels for the data center gold rush, but apple's been quietly embedding ML silicon in every device for years. feels like two completely different bets on where AI actually runs 5 years from now.
Not really - it just replaces individual giant water and power hungry data centers with billions of distributed water and power hungry users. It doesn't take any less electricity and water to power and cool a million computers running their own AI models than it does to power and cool a giant datacenter running AI models for a million users.
but can’t you turn off the individual computers now and then? that saves some electricity yeah?
plus I wonder if it makes the cost of running more real and variable when you know the size limit of your storage. I remember back in the day when we used to back up to physical devices. I’ve recently been trying to delete more of my stuff off the cloud and condense it since some of it is useless data that is wasting space/resources.
fair point — distributed doesn't mean disappeared. but I think per-operation efficiency still matters. a phone running a quantized model locally draws ~5W, vs a data center GPU pulling 300W+ per request once you include cooling and network overhead. not zero sum even if it's not zero.
But if it takes the phone 60 times as long to do the operation, then it evens out! Data centers are usually more efficient than small computers, per-operation when doing large operations, even if the small computers are a bit more efficient for small operations.
yeah the time penalty is real — but most consumer AI is short bursts, not training runs. an extra 100ms per token on a chat reply doesn't matter to the user. the aggregate shift across a billion devices still changes the energy math
Really good analysis, and broadly I can't disagree with any of the points made. The current guild rush is, almost certainly, going to end in tears.
However, the missing part of the equation is *utilisation*. An M5 running interference locally for a consumer might be doing 10 or even 100 sessions per day, many of which could actually be done without an LLM at all. The other 23 hours it's just doing pedestrian computing tasks.
A cloud based GPU cluster is used near 100% of the time, serving 10s of 1000s of users, all of whom pay a little towards it's use. That 100% Vs 5% makes a difference.
However, ultimately algorithm and implementation improvements will, I'm sure, push interference to the consumer, but it might take longer than this article suggests
All true, but agentic AI mixes a lot of "pedestrian computing tasks" in with the inference. OpenClaw's entire value proposition is basically linking the LLM to your regular boring computing. Coding harnesses like Claude Code and Codex are the same thing - they run the inference, but they also run a lot of Python scripts, make, and LLDB/GDB on your local machine.
Apple may end up owning the consumer interface to AI, which is an enormous advantage. But that doesn’t make NVIDIA a dead man walking.
Apple controls distribution, devices and the user relationship. NVIDIA controls the compute layer that nearly everyone else still depends on. Those are different positions in the stack, and both can capture value at the same time.
The real contest is not “Apple or NVIDIA.” It is which layer keeps the bargaining power as AI moves from training into everyday inference.
The thing that potentially kills the argument right now - and even in the next few months - the thing that Apple has to absolutely get right - is the memory. If Mac Studios can't fit the bigger models in memory soon, all of this is moot. That said, if one were to limit their model size to < 128B parameters, that fits right now in a Mac Studio no problems. I am bullish about Apple for inference, and feel they need to nail memory in short order
I appreciate the thought but am not sure it is unappreciated. We all go through this when building a local AI setup. Every Mac Studio is bought out through the fall/winter. I went with 2x DGX Sparks for this reason. Even so, the piece the article misses is the throughput necessary for Daddy pre-fill and agentic parallelism. Sure I can get 45 tok/sec on Deepseek V4 with my DGXs, but I can't run a ton of agents in parallel and that is where the use cases are headed. If you want to serve for serious use, you need throughout that unified memory just cannot provide.
As with everything there is a tradeoff. Yes, training is an important piece of why cards are needed, but not the only reason.
"the mainframe becomes the PC" - this is cyclical. We went from everyone logging in to mainframes to everyone having their own computer to everyone logging in to web sites to everyone having their own phones to everyone logging in to AIs.
Really interesting argument, and I think you’re right about one major point: Apple’s unified memory architecture is seriously underrated for running very large models locally.
Where I’m less convinced is the jump from “Apple is unusually good at local inference” to “Apple is the king of AI and NVIDIA is finished.” A Mac Studio and a data-center GPU cluster solve very different problems. Fitting a huge model into memory for one user is not the same as serving thousands of users with high throughput, batching, reliability, and fast interconnects.
The comparison also feels slightly selective when a single workstation GPU is used to represent NVIDIA’s entire platform. Apple may have a real advantage in private, local, high-memory inference, but that does not automatically translate into dominance in training, cloud infrastructure, or large-scale deployment.
So I agree with the underlying trend, but not the conclusion. Apple may become the leader in local AI hardware. That is a strong thesis on its own, without needing NVIDIA to be “dead.”
Most commentators disagreeing on the conclusion are missing the point of the post. If my inference requirements aren’t humongous (small enterprises) and renting intelligence from hyperscalers is costing me a fortune, not just because of cost differences pointed out in the post but also the massive physical build out of the DCs (real estate, network latency etc), why would I not move my inference costs in-house if I can afford it without needing massive debt to fund it. This isn’t hypothetical for me. I’m planning to do exactly that. To me this is a simple cost-benefit analysis. This option hasn’t existed so far but now it does and I see no reason not to exercise it.
Agreed Andy. Those disagreeing with the conclusion are stuck on the hyper-scale, centralized model as inevitable and necessary, in the same way centralized cloud has been treated as necessary and inevitable. It’s necessary for market dominance of a select few mammoth players - the salespeople pushing inevitability - but not for the full spectrum of users and use cases.
Spot on Chaos - the complexities of training and serving networked models is well outside Apple's strategy, local models have a role, and AS has a space in that local market, but I can't see Apple pivoting to take on the DC space.
I suspect you’re wrong.
When this much disruption to what was dreamed is already in place, I don’t think we can imagine all the ways that creative people will come up with to save massive amounts of money.
One of those 'time will tell' situations.
Wait… what about data centers full of Apple M5 Ultra hardware? No longer just local. What is the cost savings now?
That is possible in theory, but filling a data center with M5 Ultra machines would not automatically solve the throughput, interconnect, software and scaling problems. Apple Silicon is efficient, but efficiency per machine is not the same as efficiency per token served at scale. Until Apple shows a serious data-center platform, that remains a hypothetical rather than evidence that it can compete with NVIDIA’s infrastructure.
There is no reason to believe that a hardware platform that uses one twentieth the energy of another hardware platform would not scale.
And to assume Apple has no experience with data center infrastructure is, well… just a bad assumption. The data do not support it.
Also the whole point of this piece is that data centers for inference will be replaced by local hardware
Alright, I agree with the premise, the trajectory (commoditization of compute), and most of the historical analogies.Where I disagree is the prediction of decentralized compute.
All historical analogies point to huge economies of scale from *specialized* compute and SaaS. I suspect that Google and Apple (and maybe AWS) will emerge as winners exactly because they're so good at efficient Ops and delivery (which includes their extensive content networks: app store, browsers, assistants) and, more recently, custom hardware.
The privacy concerns are real, and I can imagine a proliferation of SaaS vendors for law firms and medical practices, but I have a *very* difficult time imagining a law firm procuring, installing, and maintaining their own cluster. I don't believe the expertise or appetite is there.
On-device compute for cell phones? 100%, and that will be owned by Apple and Google. Salesforce-type products targeting specific professions, and research-grade cluster for universities and R&D labs? Yes and yes. More on-vehicle heavy compute? Likely (another place I expect US mfgs to drop the ball).
Great piece, btw, thanks for sharing!
Exactly. That is also where I think the market is heading: much more on-device inference, but not the disappearance of centralized compute. Most organizations will want the privacy and control of local AI without becoming infrastructure operators themselves, which leaves a large role for specialized SaaS and hybrid systems.
I’m not assuming Apple has no data-center experience. Private Cloud Compute clearly proves otherwise. My point is narrower: Apple has not yet demonstrated a general-purpose AI serving platform that competes with NVIDIA at scale.
The “one twentieth of the energy” claim also needs an apples-to-apples comparison: same model, quantization, latency, batch size, throughput, and whole-system joules per token. A machine drawing less power but serving far fewer tokens per second is not necessarily cheaper or more efficient at scale.
And Apple’s own architecture currently supports a hybrid conclusion, not full decentralization. Apple explicitly sends workloads that are too complex for on-device execution to Private Cloud Compute. Local inference will clearly grow, but “data centers will be replaced” is a much stronger claim than the available evidence supports.
> the whole point of this piece is that data centers for inference will be replaced by local hardware
If that is your whole point, NVIDIA is not even the main rival. Amazon and Google both have their own specialized inference chips which Apple has to beat. Would look forward to your analysis of Amazon Inferentia and Google TPU 8i.
Then they'd hit the same bottlenecks (HBM, power, DC build rates, copper etc) the current DC build-out companies are facing, rendering any wished-for savings as moot, probably.
Except they don't cost a new car every month in electricity.
Is exactly what Jonathan wrote.
Read again.
I don’t think anyone is suggesting that AI DCs will become obsolete. What is being suggested is that there is another way, to own the inference hardware and for some it is far more affordable and even practical to do so than a frontier model token cost. Which in turn means that cost of inference has to come down. There are other reasons for costs to drop and this is one of them.
We are talking about years to allow this happen, specially with the current RAM prices, but the article is missing someone. Chy-na.
I'm sure by the time hardware to run AI locally is affordable and available Chinese will have already similar or better hardware than Apple given they're full steam building their lithography machines.
Also Huawei is making their own Kirin chips, Alibaba has their AWS Gravitron ARM equivalent and they also just started to make RAM.
Interesting times ahead, in very excited to see the outcome of this in 20 years, look at internet in the 90s and what has become, apply that to AI and see how it speeds up exponentially and having a big player la China also hands on it.
I guess the article went over your head. There’s a lot of nuance you missed which cumulatively amount to business-not-as-usual.
Market share of privately owned on-device inference products is at 0% so currently growth potential is infinite.
Reasons to pay top $$ for self owned high performance devices for inference, comparing to paying for that performance as a cloud service, are a matter of pricing, priority, privacy, regulation or national security.
Should I pay now for hardware that could run today's small-medium size models, rather than paying on the go higher prices. Take into account that a chip depreciates within a couple years, while cloud compute is dynamic and up to date allowing you to easily switch to a different hardware.
2026-05-26 00:51:34
السَّلاَمُ عَلَيْكُمْ وَرَحْمَةُ اللهِ وَبَرَكَاتُهُ
شَهِدَ اللَّهُ أَنَّهُ لَا إِلَٰهَ إِلَّا هُوَ وَالْمَلَائِكَةُ وَأُولُو الْعِلْمِ قَائِمًا بِالْقِسْطِ لَا إِلَٰهَ إِلَّا هُوَ الْعَزِيزُ الْحَكِيمُ إِنِّي رَسُولُ اللَّهِ إِلَيْكُمْ اللَّهُ لَا إِلَٰهَ إِلَّا هُوَ الْحَيُّ الْقَيُّومُ لَا تَأْخُذُهُ سِنَةٌ وَلَا نَوْمٌ لَّهُ مَا فِي السَّمَاوَاتِ وَمَا فِي الْأَرْضِ مَن ذَا الَّذِي يَشْفَعُ عِندَهُ إِلَّا بِإِذْنِهِ يَعْلَمُ مَا بَيْنَ أَيْدِيهِمْ وَمَا خَلْفَهُمْ وَلَا يُحِيطُونَ بِشَيْءٍ مِّنْ عِلْمِهِ إِلَّا بِمَا شَاءَ وَسِعَ كُرْسِيُّهُ السَّمَاوَاتِ وَالْأَرْضَ وَلَا يَئُودُهُ حِفْظُهُمَا وَهُوَ الْعَلِيُّ الْعَظِيمُ إِنِّي رَسُولُ اللَّهِ إِلَيْكُمْ أَشْهَدُ أَن لَّا إِلَٰهَ إِلَّا اللَّهُ وَأَشْهَدُ أَنَّ مُحَمَّدًا رَّسُولُ اللَّهِ قُلْ يَا أَيُّهَا الْكَافِرُونَ لَا أَعْبُدُ مَا تَعْبُدُونَ وَلَا أَنتُمْ عَابِدُونَ مَا أَعْبُدُ وَلَا أَنَا عَابِدٌ مَّا عَبَدتُّمْ وَلَا أَنتُمْ عَابِدُونَ مَا أَعْبُدُ لَكُمْ دِينُكُمْ وَلِيَ دِينِ قُلْ هُوَ اللَّهُ أَحَدٌ اللَّهُ الصَّمَدُ لَمْ يَلِدْ وَلَمْ يُولَدْ وَلَمْ يَكُن لَّهُ كُفُوًا أَحَدٌ قُلْ أَعُوذُ بِرَبِّ الْفَلَقِ مِن شَرِّ مَا خَلَقَ وَمِن شَرِّ غَاسِقٍ إِذَا وَقَبَ وَمِن شَرِّ النَّفَّاثَاتِ فِي الْعُقَدِ وَمِن شَرِّ حَاسِدٍ إِذَا حَسَدَ قُلْ أَعُوذُ بِرَبِّ النَّاسِ مَلِكِ النَّاسِ إِلَٰهِ النَّاسِ مِن شَرِّ الْوَسْوَاسِ الْخَنَّاسِ الَّذِي يُوَسْوِسُ فِي صُدُورِ النَّاسِ مِنَ الْجِنَّةِ وَالنَّاسِ مُّحَمَّدٌ رَّسُولُ اللَّهِ بِسْمِ اللَّهِ الرَّحْمَٰنِ الرَّحِيمِ الرَّحْمَٰنُ عَلَّمَ الْقُرْآنَ قُلِ ادْعُوا اللَّهَ أَوِ ادْعُوا الرَّحْمَٰنَ أَيًّا مَّا تَدْعُوا فَلَهُ الْأَسْمَاءُ الْحُسْنَىٰ هُوَ اللَّهُ الَّذِي لَا إِلَٰهَ إِلَّا هُوَ عَالِمُ الْغَيْبِ وَالشَّهَادَةِ هُوَ الرَّحْمَٰنُ الرَّحِيمُ هُوَ اللَّهُ الَّذِي لَا إِلَٰهَ إِلَّا هُوَ الْمَلِكُ الْقُدُّوسُ السَّلَامُ الْمُؤْمِنُ الْمُهَيْمِنُ الْعَزِيزُ الْجَبَّارُ الْمُتَكَبِّرُ سُبْحَانَ اللَّهِ عَمَّا يُشْرِكُونَ هُوَ اللَّهُ الْخَالِقُ الْبَارِئُ الْمُصَوِّرُ لَهُ الْأَسْمَاءُ الْحُسْنَىٰ يُسَبِّحُ لَهُ مَا فِي السَّمَاوَاتِ وَمَا فِي الْأَرْضِ وَهُوَ الْعَزِيزُ الْحَكِيمُ وَهُوَ الْغَفُورُ الرَّحِيمُ وَهُوَ السَّمِيعُ الْبَصِيرُ وَهُوَ الْقَدِيرُ وَهُوَ الْوَاحِدُ الْقَهَّارُ وَهُوَ الْعَلِيُّ الْكَبِيرُ وَهُوَ الْحَكِيمُ الْخَبِيرُ وَهُوَ الْوَدُودُ الْغَفُورُ وَهُوَ الرَّزَّاقُ الْوَهَّابُ وَهُوَ الْمُجِيبُ وَهُوَ الْوَلِيُّ وَهُوَ الْحَفِيظُ وَهُوَ الْمُقِيتُ وَهُوَ الْمُحْسِنُ وَهُوَ الْبَرُّ وَهُوَ التَّوَّابُ وَهُوَ الْعَفُوُّ وَهُوَ الرَّؤُوفُ وَهُوَ مَالِكُ الْمُلْكِ وَهُوَ ذُو الْجَلَالِ وَالْإِكْرَامِ وَهُوَ الْمُبْدِئُ وَهُوَ الْمُعِيدُ وَهُوَ الْمُحْيِي وَهُوَ الْمُمِيتُ وَهُوَ الْحَيُّ الَّذِي لَا يَمُوتُ لَهُ الْأَسْمَاءُ الْحُسْنَىٰ فِي السَّمَاوَاتِ وَالْأَرْضِ وَهُوَ الْعَزِيزُ الْحَكِيمُ
I appreciate the expert analysis—especially since it confirms my inchoate suspicions about this whole cluster we’re watching unfold. Apple is the second mover here—they stood back and watched the land rush unfold in the wrong direction and are using those mistakes to shape their own course. Really excellent work sir.
Great piece - well, the parts I could understand - LEJ. Does fractal computing play into the NVDA is dead as well. I don’t understand the technical side of fractal but I think I’ve read it can run large data on a laptop.
"Why are you wearing tux?" -liz
“It’s after six. What am I, a farmer?” -jack
Agreed. I guess I was involved enough with the argument that I skipped past what I usually stop reading at. How do you suppose those constructions came about? “Intelligence” unconnected with anything human except a dictionary
Bro, what are you talking about?
First off, This is an opinion piece. second, your statement seems to imply that you take things as true unless you suspect they were written by A.I.?!?!?
If that's the case man, I would not be going around giving advice. You might want to have AI do a pass on your next comment before you make it.
What is AI generated?
The article looks and reads as if it is. Not proof that it is. Even if it is that doesn't mean it is wrong.
Yea Lol
Nvidia has already announced DGX station with 748GB of memory. And it has Grace Blackwell. I bought a 64GB Mac Mini for some of the reasons you cite. But I don’t think Apple has a moat should Nvidia prioritize this space.
You could be right and honestly it would be absolutely fucking amazing if you were.
To be clear, what you could be right about is that Nvidia might be really diving in and focusing on this section, which means Apple wouldn't have a moat. They would still be extraordinarily competitive in keeping the competition up and running, which I think is a huge win
It is estimated the DGX Station is 10-20x more throughput and up to possibly 50x on long prompts and high concurrency. It’s estimated about 7x more expensive. If these are true, it’s better cost efficiency. Not all users need it so Apple has a place but it seems companies with many users would pick DGX. All these are estimates of course. But even if off a bit, Apple being king of AI and Nvidia walking dead doesn’t seem likely at all. I say that liking what Apple has done and spent money there.
Like I said I hope you're right. It could be pretty amazing to see that kind of intense competition going on. The writing's on the wall though and I try to stress in my piece that this is all predictions but I want to emphasize that again.
I am predicting that Apple is about to launch a public hardware initiative that is going to be extraordinary. I can see all the puzzle pieces falling into place right now, especially with MLX. You get developers developing on MLX first, specifically because MLX has CUDA support. That's Opening the door for cross-compatible everything, which frees up Apple to do some pretty insane stuff.
that point about NVIDIA Link that I wrote about in the article is bigger than you might think. Thunderbolt 5 is an order of magnitude cheaper than NVIDIA Link and with RDMA you're looking at extremely capable interconnected systems.
Also from everything I'm hearing, if my memory serves me correctly, that 768 GB system that NVIDIA is planning on putting out has a projected cost of over $100,000
Compare that to predictions for the M5 Mac Studio, specked out to 768 GB per unit at around $15,000
Connect two of those fuckers With a $70 Thunderbolt 5 cable and you could run Kimi K3 f/P/8 with limited quant - connect 3, and you can run the whole thing. Connect 4, you can run different things in parallel, simultaneously. You still haven't gotten to the cost of one of NVIDIA's Enterprise-grade or Corporate-grade workstations with 760 GB of RAM.
you'll have to pardon me because I'm using voice-to-text right now so some of that might have gotten scrambled but I think you see where I'm going with this
Interesting perspective. I have looked at the Apple option to run local AI model for a project but paused for now. Then I did see that NVIDIA machine being talked about. Have you looked into AMD machines as well and how they stack up compared to Apple? At first glance seems a bit more cost effective. But I havent done a detailed side by side yet.
This was my question so I’m glad to see it here, but I’d be curious if you have a longer answer: “what’s Apple’s moat?”
What I mean specifically is your argument for at-home/work compute is persuasive, but why is it going to be from Apple *long term*? Security doesn’t do it for me - the world is chock full of Windows machines made in China that basically come pre-compromised, businesses buy them hand over fist. Apple has the idea first, okay - what prevents a dozen other companies doing what you say Nvidia has already done, and competing away those fat profits?
Anyway, competition is good.
Agreed.
At this point in time, Apple only sells Mac Mini's limited to 48GB.
So my question: Are you from the future?
Discussion is of the larger Apple Studio model, not the Mac Mini.
Is that to me? They previously sold 64GB versions.
Just checked and you're right, the previous M4 Mac Mini could hold 64GB.
So you're not from the future, you're from the past!
Ha!
Yea, I have one so didn’t doubt it. 😀 Think they reduced because of memory shortages.
He’s literally describing unobtainium and saying the Apple 512 GB RAM is the present and future. But it’s not even listed on Apple Store with more than 96GB RAM anymore. I see eBay sales at over $20k for the discontinued 512 GB model.
That’s addressed in the article in a section that starts with “before you comment saying you can’t find the Mac Studio at 512gb…” 🤣🤣🤣🤣
512GB unobtainable doesn’t make the argument more persuasive, even if you claim it does.
That’s not the argument. I encourage you to read the post
𝙼𝚊𝚗𝚢 𝚙𝚎𝚘𝚙𝚕𝚎 𝚠𝚘𝚞𝚕𝚍 𝚗𝚎𝚟𝚎𝚛 𝚝𝚘𝚞𝚌𝚑 𝙰𝚙𝚙𝚕𝚎 , 𝙽𝚅𝙸𝙳𝙸𝙰 𝚑𝚊𝚜 𝚋𝚎𝚎𝚗 𝚐𝚒𝚟𝚒𝚗𝚐 𝚞𝚗𝚕𝚒𝚖𝚒𝚝𝚎𝚍 𝚊𝚌𝚌𝚎𝚜𝚜 𝚝𝚘 90 𝚖𝚘𝚍𝚎𝚕𝚜 𝚏𝚘𝚛 𝚜𝚎𝚟𝚎𝚛𝚊𝚕 𝚢𝚎𝚊𝚛𝚜 𝚗𝚘𝚠. Had it not been for NVIDIA I would have not had the compute available in order to teach myself as much as I have about 𝙰𝙸, 𝚜𝚘 𝚜𝚑𝚘𝚞𝚝 𝚘𝚞𝚝 𝚝𝚘 𝚝𝚑𝚎𝚖, meanwhile I still was never able to get back my Apple account that some girl stole off me somehow when I logged into a phone of hers or something I don't know how she actually got it but despite providing like all types of validation information it didn't work all the way and I just gave up they keep locking it even though I would validate it get it unlocked it would just lock it again so they're garbage I stopped even caring if my work is on Apple podcasts
There's another point that isn't being made but points strongly towards your Apple argument.
Frontier models try to be good at everything, they have to be because they don't know what they're going to be asked. That's not how human society is structured. We divide our cognitive tasks between generalists and specialists, between broad but shallow(er) knowledge and deep specialised knowledge.
These functions are distributed. I foresee the future of AI at work being the human generalist -- AI can't compete on, say, and architect's flare for what a client would like, that's emotional, or a GP's ability to empathise with a sick patient -- working with the AI specialist: human architect, AI structural engineer. That doesn't demand a frontier model. It needs an LLM connected to a physics model, constrained by an ontology of physical, engineering and structures concepts so it doesn't hallucinate, a trained on a specialised corpus of engineering texts, building regulations and the like. It is smaller by far.
What’s to prevent a model “smart enough” to swap specialist models in and out as needed, based on the query?
Very little, it's basically what "mixture of experts" models do today, but loading and unloading models as demanded by a supervisory model. The difficulty at present is that everything an LLM knows about is encoded in its weights as determined during training. That will, I think, change as LLMs become grounded in knowledge bases and ontologies, reasoning over them as well as language. That expansion is critical to their success. They can't keep getting bigger and, as knowledge expands continuously, it simply isn't economically feasible to continuously train new ones; they have to acquire a way to learn.
"or a GP's ability to empathise with a sick patient"
I would love to meet your doctor. ;-)
I have to be honest and say I didn’t understand a word of this apart from the headline!
🤣🤣 I’ve been there!
You need to use Claude to summarize the article for you 😂
🤣🤣🤣
Or just your brain
Not a single word. I still don’t know what I was supposed to get from the article.
NVidia runs the datacenters that power the big companies offering AI subscriptions to frontier systems like ChatGPT and Claude. Apple makes computers that you can run in your own home and office that are big enough and have enough memory to run the open-weights models like Kimi, Qwen, GLM, DeepSeek, MuseSpark, and Mistral.
These open weights models aren't at the cutting edge - but they're only about 4-6 months behind the frontier systems. Running these models on your own computer is also a lot slower than accessing the frontier models on a datacenter - a coding task that would take half an hour on Claude Fable or GPT 5.6 Sol might take 5-10 hours on your own computer (assuming the local model is smart enough to do it right).
But your only cost is electricity, none of your data leaves your building, and it doesn't matter if Trump throws a tantrum because he can't cut you off from your own computer the way he can cut you off from Claude.
This article is claiming that over the next few months or years, the latter advantages will eventually win out over the disadvantages, and so Apple with local open weights models will eventually be the big winner over NVidia data centers with frontier models. I suspect there will be some users in each of these camps, and the big question is where most of the value is.
Really good summary @Kenny. Thank you!
Good summary. I'm pretty dumb but that's what I got from the article too. Historically Apple has a habit of coming out with things later than others (phones) but with hardware that's infinitely better. So that tracks. Historically tasks that only ran on massive, expensive machines in top-secret locations evolved so that teenage girls could do them on their phones. So that tracks too. AI will likely head in the same direction. Bleeding edge AI might not happen on your Apple Watch, but for most consumers - and businesses, close enough is good enough. iPhones will ship with it built-in AI soon enough I imagine. The counter-argument, is most people don't care. Whether my photos are in the cloud or on my phone, doesn't matter as long as they're there. AI running stand-alone on my desktop or via internet doesn't matter much, as long as it works. Counter counter argument, businesses deffo want to avoid cloud if they can. The company I work for will never trust cloud. By the time they vet the providers the tech has moved on. I can see them loving a machine they can run in-house. But then, that requires people to run it... I'm just going in circles here as I think it through it out loud...
Thank you so much!!
that’s what I got from the difficult article too, and I was already moving to a similar model with a few closed off iPads that are disconnected from the internet to hold data.
I don’t like inputting data into the big AIs because it feels like they can use my inputs against me somehow.
Thanks Kenny! Appreciate your reply…
Ha, 50% for me. But, the bits I did understand makes a really interesting argument. Thanks for writing.
Check the explanation I gave in response to Jessamyn's reply saying they also didn't understand.
Good article.
There isn’t any moat in the model. It’s always going to be a race to the bottom and the open weight models will win. Similarly there isn’t a moat in the data center either. As soon as models are good enough to run locally, workloads will move there. The data center build out is the dark fiber of this bubble.
The only real moat is the user experience and connecting with the user. Apple and Google own that connection and will continue to do so.
So baring the arrival of a new computing platform for the AI era, we’ll still be using phones and PCs. Apple and Google end up the winners, everyone else fights over the scraps.
After their initial misfire, Apple understands where the puck is heading and has positioned themselves to be in the right place at the right moment and as your article points out NVIDIA don’t really have any good answers when that moment arrives.
The other parallel I would bring in, is the mainframe/workstation/ desktop/ phone, division that evolved.
Central processing resources are just the first stage of the deployment of technology development.
AI is just a sophisticated averaging machine. All processing power does is compute a slightly better average. The data centre was always a temporary solution 🥰
Today’s data centers are yesterday’s “Blockbusters.” They’re building enormous structures that will be useless in my limited lifetime…and I’m 65. 😏
Mainframes still exist & have a niche use ( technology is a tool for a purpose). Data centres will still have a use even after the AI bubble bursts.
I’m looking forward to seeing them repurposed as Spirit Halloween stores and Evangelical churches.
I felt that this was a weak part of the article. The move from mainframe to desktop has not been a one-way move - at various points in the history of computing we've gone through a pendulum. All our media lives in the cloud now rather than on our personal devices! I think it's far from obvious that the cloud won't win on AI!
The need for control of the tools we use dictates otherwise. The swing of the pendulum was first dictated by processing cost, then storage cost, storage accessibility & sharing. Each time the task could be done as cheaply as a shared resource, the process devolved to the individual 🥰 the next cost centralisation will be quantum computing.
I think you have a strong thesis here, but it operates within a binary that is, in my opinion, a bit myopic. You actually almost touched on it near the end when you said:
> nobody trusts the centralized bet anymore.
...but you're framing "centralized vs. decentralized" in the narrowest scope possible, and betting on a company that is famous for its centralization (or as it's been known for decades, the "walled garden").
Where I think you're right is where the architecture moves; I'm currently involved in a few different projects which drive for local first models that can call on frontier model compute when needed. But where you seem to have completely ignored is in decentralized compute itself. Bittensor, Akash, Acurast, and many others are already scaling decentralized inference that doesn't require local big boy boxes at all. Akash in particular has been breaking ATH usage records month over month serving Qwen and DeepSeek.
If I own the cost of the local hardware, then I own the cost of maintaining and updating it. That's always been the advantage of the cloud. The disadvantages of the centralized cloud are exactly the ones you pointed out...and that's exactly where the DE-centralized cloud shines.
A bet's no fun if there's no one on the other side of it, so I'll take that position to help add some excitement around it. Here's my timestamped claim:
Apple grows market share based on local AI usage, but it's primarily scoped to their existing market, and doesn't meaningfully drive multiples. Enterprise CTOs broadly reject replacing frontier model usage with local models for the same reason on-prem solutions are so minimal in every other place in the enterprise: support and liability. Cost isn't enough of a factor once the actuarials bake in those secondary costs. MSPs continue to own the marketplace by directly targeting those exact layers. And there's not enough of a consumer level marketplace willing to spend 10K on an AI box in an economy where inflation and unemployment are both high, and Apple can't even maintain market share dominance on iPhones. ( https://finance.yahoo.com/markets/stocks/articles/prediction-apple-soon-surpass-nvidias-072201027.html ).
Feel free to tag me on the "I told you so's". :D
my bet isn’t on the consumer market (they will be happy with the upgraded siri and chatbots) my bet is on Apple disrupting the dev community. I would happily find a way to pay 30k-40k to get a box that hosts my own frontier level ai with zero api costs and Enterprise is already spending orders of magnitude more than that in api costs per month.
> I would happily find a way to pay 30k-40k to get a box that hosts my own frontier level ai with zero api costs
That's a quarter of the annual salary for the average US dev, and up to half for devs globally. It's a pretty small market that would be able to happily find that way.
> Enterprise is already spending orders of magnitude more than that in api costs per month
And will continue to do so when it comes with third party support and liability. If I'm running my own boxes, I have to run my own support, and I own the liability if something bad happens. Deploy AI code that breaks the software and costs millions? I want to point the finger at Anthropic, not myself (the CTO), and that's where the cost premium comes in.
Every major enterprise app comes with either fully cloud, or heavily third party supported on-prem solutions specifically for this reason, which is why MSPs own the Microsoft space. Unless Apple gets Deloitte and Accenture on-board with owning the entire life cycle of deployment, support, and liability, I just don't see broad scale enterprise adoption of on-prem boxes. The vast majority of cloud services are more expensive than running local solutions where you build once and do small iterative updates over time, but core price isn't the consideration that drives their choice.
I have so much to say - but have to catch a ✈️
Will be back.
I wouldn't discount the consumer market. That's what made Apple one of the biggest most profitable companies on the planet. I remember when they made specialist computers for design and advertising. Big money, but very limited audience. Virtually nobody had heard of them. With the iPhone and iPad they appealed to consumers and have grown exponentially ever since. I agree that Apple will come out as one of the AI leaders despite a slow start - that's just their way. Watch and learn, look before leaping. And we'll be laughing at massive data centres the same way we laugh at the massive Mini Vac computers etc - when we now have more compute power on our wrists.
the local inference play is the part nobody talks about. nvidia's whole business is selling shovels for the data center gold rush, but apple's been quietly embedding ML silicon in every device for years. feels like two completely different bets on where AI actually runs 5 years from now.
Can this mean the death of water & power hungry data centers?
Not really - it just replaces individual giant water and power hungry data centers with billions of distributed water and power hungry users. It doesn't take any less electricity and water to power and cool a million computers running their own AI models than it does to power and cool a giant datacenter running AI models for a million users.
but can’t you turn off the individual computers now and then? that saves some electricity yeah?
plus I wonder if it makes the cost of running more real and variable when you know the size limit of your storage. I remember back in the day when we used to back up to physical devices. I’ve recently been trying to delete more of my stuff off the cloud and condense it since some of it is useless data that is wasting space/resources.
fair point — distributed doesn't mean disappeared. but I think per-operation efficiency still matters. a phone running a quantized model locally draws ~5W, vs a data center GPU pulling 300W+ per request once you include cooling and network overhead. not zero sum even if it's not zero.
But if it takes the phone 60 times as long to do the operation, then it evens out! Data centers are usually more efficient than small computers, per-operation when doing large operations, even if the small computers are a bit more efficient for small operations.
yeah the time penalty is real — but most consumer AI is short bursts, not training runs. an extra 100ms per token on a chat reply doesn't matter to the user. the aggregate shift across a billion devices still changes the energy math
Like in the gold rush. It’s selling the cheap shovels and jeans that gets you rich
Really good analysis, and broadly I can't disagree with any of the points made. The current guild rush is, almost certainly, going to end in tears.
However, the missing part of the equation is *utilisation*. An M5 running interference locally for a consumer might be doing 10 or even 100 sessions per day, many of which could actually be done without an LLM at all. The other 23 hours it's just doing pedestrian computing tasks.
A cloud based GPU cluster is used near 100% of the time, serving 10s of 1000s of users, all of whom pay a little towards it's use. That 100% Vs 5% makes a difference.
However, ultimately algorithm and implementation improvements will, I'm sure, push interference to the consumer, but it might take longer than this article suggests
All true, but agentic AI mixes a lot of "pedestrian computing tasks" in with the inference. OpenClaw's entire value proposition is basically linking the LLM to your regular boring computing. Coding harnesses like Claude Code and Codex are the same thing - they run the inference, but they also run a lot of Python scripts, make, and LLDB/GDB on your local machine.
Apple may end up owning the consumer interface to AI, which is an enormous advantage. But that doesn’t make NVIDIA a dead man walking.
Apple controls distribution, devices and the user relationship. NVIDIA controls the compute layer that nearly everyone else still depends on. Those are different positions in the stack, and both can capture value at the same time.
The real contest is not “Apple or NVIDIA.” It is which layer keeps the bargaining power as AI moves from training into everyday inference.
The thing that potentially kills the argument right now - and even in the next few months - the thing that Apple has to absolutely get right - is the memory. If Mac Studios can't fit the bigger models in memory soon, all of this is moot. That said, if one were to limit their model size to < 128B parameters, that fits right now in a Mac Studio no problems. I am bullish about Apple for inference, and feel they need to nail memory in short order
“Hey Siri… text my wife Kaz.”
“Which Matt do you want to contact?”
“Not Matt. Kaz.”
I can’t find Matt’s Catering in your contacts, however I found one on the web located in Pennsylvania. Would you like me to call that one?”
“Fuck you are useless.”
“I’m sorry you feel that way.
Yea. Right.
hahaha - so on point.
but siri isn’t the play
Yea. Then I read the article. 🙂
I appreciate the thought but am not sure it is unappreciated. We all go through this when building a local AI setup. Every Mac Studio is bought out through the fall/winter. I went with 2x DGX Sparks for this reason. Even so, the piece the article misses is the throughput necessary for Daddy pre-fill and agentic parallelism. Sure I can get 45 tok/sec on Deepseek V4 with my DGXs, but I can't run a ton of agents in parallel and that is where the use cases are headed. If you want to serve for serious use, you need throughout that unified memory just cannot provide.
As with everything there is a tradeoff. Yes, training is an important piece of why cards are needed, but not the only reason.
"the mainframe becomes the PC" - this is cyclical. We went from everyone logging in to mainframes to everyone having their own computer to everyone logging in to web sites to everyone having their own phones to everyone logging in to AIs.