2 used 3090 would give you 48 GB VRAM for about the same price of a used 4090, but the prompt processing, which is basically based on how many FPS in games a GPU scores, will be slower, because a 3090 has a lower gaming performance than a 4090.
It depends on the AI model you want to run and your task (if e.g. big inputs and you want them to be processed faster), but usually 48 GB VRAM are (much) preferred over faster prompt processing (aka prefill), for reasons like: You are able to fit a bigger AI model or a higher quant of same model and/or bigger context.
For 2 GPUs and a consumer motherboard you may want to see if using bifurcation (PCIe 5.0 x8 (GPU 1) + PCIe 5.0 x8 (GPU 2)), instead of the usual PCIe 5.0 x16 (GPU 1) + PCIe 5.0 x4 (GPU 2), would give you a performance boost. A hassle free solution is to get a motherboard that has bifurcation built-in, instead of getting a bifurcation card, like the Asus ProArt X870E-Creator WIFI or the cheaper X870E Taichi Lite. See reddit.com/r/LocalLLaMA for more.
As for a smaller and lower power server starting point, see AMD Epyc 8004 "Siena" DDR5 6-channel, 96 PCIe Gen5 lanes, platform. Except for Qwen 27B dense, all recent AI LLM model releases have been MoE models, which fit perfectly RAM setups (not everyone can afford 4 to 8 RTX PRO 6000s), so the faster your RAM is, the faster they will run and Siena looks like the perfect balance between size, power consumption and price (on the other end would be "Venice" (Zen 6 16-channel RAM)).