Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Isn't that the same thing you can get from ordinary 2S Epyc/Xeon servers at a similar price that have 24 memory channels (when the M3 Ultra has the equivalent of 16)?

And the reason people rarely use that for AI is that the enterprise GPUs from AMD and Nvidia are only moderately more expensive but are significantly faster because they use HBM instead of DDR5.



Yeah kind of, I think a 24 channels DDR5 works out approx 1TB/s, but the cost is astronomical, a M5 studio would probably beat that performance for around half the cost. You also get to use the GPU/NPU cores of the mac vs CPU only on the servers. M5 ultra studio with 128GB RAM could probably beat out a sever with a RTX 6000 pro at half the price.


24 channels of DDR5-6400 is 1.2TB/s, M3 ultra is 0.8TB/s. They both use DDR5-6400.

> a M5 studio would probably beat that performance for around half the cost.

A barebones 2S system with no CPUs or memory is ~$2000, a pair of 16 core CPUs another ~$1000 each, and then however much memory you want. The price seems pretty comparable. The "problem" with doing this is actually that 128GB is too little memory, because you want to populate all the channels, but even using 16GB sticks, 24x16GB is already 384GB.

> You also get to use the GPU/NPU cores of the mac vs CPU only on the servers.

You only need enough cores to make sure the bottleneck is memory bandwidth.


M3 ultra is obviously 1-2 generations behind and new the studio is expected 'any day now. Even if this was M4 Ultra it would still be ~comparable to any EPYC system in bandwidth, but get to use the GPU for compute so potentially faster than the EPYC. Total Cost of Ownership in the Epyc is going to be WAY higher because of electricity costs, the EYPC is going to be consuming probably 5X the electricity and is probably not going to sit quietly on your desk. More RAM though, but again it's more about the ratio of RAM (size) to RAM (Memory Bandwith) to Compute and you may find a model bigger than e.g 70b suddenly is bottlenecked by the CPUs or memory bandwidth and therefore the extra RAM (size) is wasted. But maybe not, different use cases will yeild different results I guess.

> A barebones 2S system with no CPUs or memory is ~$2000, a pair of 16 core CPUs another ~$1000 each, and then however much memory you want.

As you say, the thing is it's not 'however much memory you want' it's 24 sticks which at $300 a stick for 16GB is $7200, then you also need at least one NVME disk so you're looking at what $13,000?


> M3 ultra is obviously 1-2 generations behind and new the studio is expected 'any day now.

M3 Ultra uses a 1024-bit memory bus, which is a major inconvenience to Apple because they're soldering everything. In ordinary systems if you have 16 memory slots and any one of the memory chips is bad, you replace that stick. If the processor is bad, you replace the processor. If the system board is bad, you transfer the processors and memory to another one.

If any of those has a defect after you solder thousands of dollars worth of memory onto the same board as a >$1000 CPU, you're not doing well. Worse, the more memory chips you have and the more pins the CPU needs for its memory bus, the higher the chances of one of them having a defect.

In addition to that, when you get to that number of pins it starts getting harder to run the memory at the highest speeds. The M3 uses DDR5-6400 but some of the M5 line is using DDR5-9600. It may or may not be possible to do that speed when using a 1024-bit bus -- not every existing M5 even does it. If it isn't then the Ultra wouldn't be much faster than the Max since it would have to use a lower memory speed. If it is then the tolerances would have to be even tighter and increase the defect rate even more.

Which is to say, I can see why they haven't released an Ultra since the M3.

> but get to use the GPU for compute so potentially faster than the EPYC.

Only if the bottleneck is compute rather than memory bandwidth, and for LLMs it's generally memory bandwidth. And if something significant was compute bound, there are also higher core count CPUs.

> Total Cost of Ownership in the Epyc is going to be WAY higher because of electricity costs, the EYPC is going to be consuming probably 5X the electricity and is probably not going to sit quietly on your desk

The M3 Ultra has a 480W TDP. There are relevant EPYC SKUs on SP5 at 125-200W/socket.

> you may find a model bigger than e.g 70b suddenly is bottlenecked by the CPUs or memory bandwidth

The extra RAM allows you to fit the larger model in memory to begin with, without which it's pretty hopeless. The bottleneck typically is memory bandwidth for LLMs, but that's true pretty much regardless of the model size. Moreover, mixture of experts models require significantly more RAM for the same amount of compute/bandwidth.

> you also need at least one NVME disk

That's ~$100.

> As you say, the thing is it's not 'however much memory you want' it's 24 sticks which at $300 a stick for 16GB is $7200

The chips don't cost a materially different amount based on whether you solder them. Apple presumably discontinued the 256GB and 512GB versions of the M3 Ultra because they'd have had to add a similar number to the price.


> but even using 16GB sticks, 24x16GB is already 384GB.

Question: you need 16GB sticks because they're the smallest doublesided ones, which you need for maximum BW, right? Otherwise why not 8G?


If 8GB sticks of registered DDR5-6400 even exist, they're not common.


That was because memory used to be cheap. Will be curious if smaller capacities come back en vogue as a cost cutting mechanism.


Using a server system designed for having hundreds of cores and TBs of RAM for LLMs because it has a medium-high amount of memory bandwidth is a hack for enthusiasts who want to take the price/performance trade off to run big models without paying for enterprise GPUs. The market for 8GB RDIMMs would be specifically that, since the more typical server workloads that need that amount of memory bandwidth also want the larger memory sticks (e.g. database servers, systems running hundreds of VMs), or aren't bounded by memory bandwidth to begin with and then don't need to populate all the channels.

And if you wanted that market segment then what you'd really do is produce a consumer GPU with 128GB of RAM.




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: