On a similar/same level to GPT-5.6 Luna is the model DeepSeek-V4-Flash-0731 and if you don't have enough VRAM + RAM, you can offloaded the rest to a SSD and for chatting the speed is still usable (depends on the speed of the SSD and how much of the model is offloaded to it).
Quote from: artificialanalysis.ai/articles/deepseek-v4-flash-0731-scores-50-on-the-artificial-analysis-intelligence-index-10-points-above-previous-deepseek-v4-flashDeepSeek V4 Flash 0731 is one Intelligence Index point behind GPT-5.6 Luna (max, 51).
After AA updated their Intelligence Index to v4.11 2 days ago, DS-V4-Flash is actually at the same 52 points now:
artificialanalysis.ai/?models=qwen3-5-122b-a10b-non-reasoning%2Cqwen3-5-122b-a10b%2Cmimo-v2-5-0424%2Cqwen3-6-27b-non-reasoning%2Cminimax-m2-7%2Cnvidia-nemotron-3-super-120b-a12b%2Cqwen3-6-35b-a3b-non-reasoning%2Cstep-3-7-flash%2Cinkling-small%2Cqwen3-6-35b-a3b%2Cqwen3-6-27b%2Cmotif-0714%2Chy3%2Cgemma-4-31b%2Cgemma-4-31b-non-reasoning%2Cmistral-medium-3-5%2Cling-3-0-flash%2Cgpt-oss-120b%2Cdeepseek-v4-flash%2Cgpt-5-6-luna&intelligence=artificial-analysis-intelligence-index
More about also SSD offloading: reddit.com/r/LocalLLaMA.