This is my upgraded rig:

  • Ryzen9 5950x with 64gb DDR4
  • Dual NVIDIA RTX A4000 (16+16GB VRAM)

to whom read my previous posts, i jumped the gun and upgraded my server, it was worthwhile and somewhat cheap given i already had the two GPUs and the DDR4 RAM.

Anyway, i am currently running Qwen3.6-35B-A3B-UD-Q5_K_XL all in VRAM with 65536 context and pretty happy with speed (80-90t/s) and overall responses (mostly chat).

I would like to experiment with something beefier, with CPU offload, that i can run with my llama.cpp. Of course t/s is not a goal here, but precision and accuracy of responses is.

I tried to find a good model with claude and gemini, but always got short. Once the model suggested fully crashed my server (guess fill up RAM and ended up in a swap loop), more then once i ended up chasing non existent models. Pretty annoying.

Considering i would only use between 32 and 48GB or system RAM, can you suggest (preferably with links to HF) some models?

I like qwen3.6, but open to anything.

    • e0qdk@reddthat.com
      link
      fedilink
      English
      arrow-up
      3
      ·
      21 days ago

      +1 for llmfan46’s heretic variants. I use one of his uncensored Qwen 3.6 35B-A3B variants as my default model.

      rpDungeon’s Luchador models (gemma derived) are also quite interesting – they tend to have better prose quality for creative writing tasks.

      I’ve been curious to try experimenting with using Rudo in particular for making more interesting NPC interactions in a text adventure for a while now, but haven’t gotten to it yet.