2 used 3090 would give you 48 GB VRAM for about the same price of a used 4090, but the prompt processing, which is basically based on how many FPS in games a GPU scores, will be slower, because a 3090 has a lower gaming performance than a 4090. It depends on the AI model you want to run and your task (if e.g. big inputs and you want them to be processed faster), but usually 48 GB VRAM are (much) preferred over faster prompt processing (aka prefill), for reasons like: You are able to fit a bigger AI model or a higher quant of same model and/or bigger context.
For 2 GPUs and a consumer motherboard you may want to see if using bifurcation (PCIe 5.0 x8 (GPU 1) + PCIe 5.0 x8 (GPU 2)), instead of the usual PCIe 5.0 x16 (GPU 1) + PCIe 5.0 x4 (GPU 2), would give you a performance boost. A hassle free solution is to get a motherboard that has bifurcation built-in, instead of getting a bifurcation card, like the Asus ProArt X870E-Creator WIFI or the cheaper X870E Taichi Lite. See reddit.com/r/LocalLLaMA for more.
As for a smaller and lower power server starting point, see AMD Epyc 8004 "Siena" DDR5 6-channel, 96 PCIe Gen5 lanes, platform. Except for Qwen 27B dense, all recent AI LLM model releases have been MoE models, which fit perfectly RAM setups (not everyone can afford 4 to 8 RTX PRO 6000s), so the faster your RAM is, the faster they will run and Siena looks like the perfect balance between size, power consumption and price (on the other end would be "Venice" (Zen 6 16-channel RAM)).
For AI: 5700 bucks for only 48 GB RAM and no fast VRAM? For this kind of money you could get/build a much faster desktop/workstation/(even)server system.
An example -- Thank me later: Most/all B850 motherboards support up to 256 GB RAM and cost 180 bucks. A 64 GB DDR5 6400 MT/s CUDIMM 1.1V stick costs 1050 bucks. Add 2 sticks for dual channel and you have 128 GB RAM right there (yes, it won't be 9600 MT/s but you will be able to fit bigger AI LLM models). And to speed prompt processing and token generation/everything up, get a used 4090 for ~2500 bucks from a company so you can send it back, just in case (paying a bit less and getting it form a private individual would be too risky for me personally). Then use "nvidia-smi -q -d POWER" to see its "Min Power Limit" (the minimum power limit is usually 50% of its TBP and this is also where the power efficiency is the highest) and set the power limit: nvidia-smi -pl <YOUR_WATT_NUMBER>. You don't have to do -50% power necessarily, see e.g. "4090 almost no performance loss at 75% power limit sweet spot", there are many such tests. Search for e.g.: "4090 power scaling". Undervolting should give you better results still.
I wish non-Apple computers had that single core performance. I only have the older M4 and it is a lot faster than my Thinkpad with an Arrow Lake in compute tasks.
The MacBook Pro 16 handles the M5 Max SoC much better than the smaller 14-inch model. It is still a great multimedia laptop with a gorgeous Mini-LED screen, but the old design starts to show limitations and the 140W power adapter is insufficient.