This is an MoE model with 1.6T-A49B

The weights were up briefly then taken down due to some issues in the repo files apparently, now they’re back up:

GGUFs are out as well:

DeepSeek published benchmarks for reference:

  • Jo Miran@lemmy.ml
    link
    fedilink
    English
    arrow-up
    2
    ·
    16 days ago

    I’m actually curious as to what’s the most I can run on an Apple M4 Max system.

        • SirDimples@programming.devOP
          link
          fedilink
          English
          arrow-up
          1
          ·
          6 days ago

          You might want to try running Ling-3.0-Flash (124B-A5.1B) it might be perfect for your M4 and should perform close to the level of DeepSeek V4 Flash preview or MiniMax M2.7

        • e0qdk@reddthat.com
          link
          fedilink
          English
          arrow-up
          2
          ·
          16 days ago

          If I understand the nature of your hardware correctly, you should be able to run the MoE models like Gemma4 26B-A4B or Qwen3.6 35B-A3B at a high quantization fairly performantly.

          You could try running some of the dense models (like today’s Qwen 3.8 27B) as well, but I expect they’ll be pretty slow (judging by my own experience with a unified RAM system that has a Strix Halo APU). Might still be useful for tasks that you can leave running on their own for a long time instead of for interactive chat style interaction though.

          You’ve got enough RAM to load larger models, but there hasn’t been much released in between the “it fits on a 24GB or 32GB GPU that a gamer might own” and the “oh god you need HOW MUCH RAM!?” scales lately…

        • lynx@sh.itjust.works
          link
          fedilink
          English
          arrow-up
          2
          ·
          14 days ago

          With Q4 everything below 150B should be fine. You can also run the -Flash variant of this model in Q1, but it is probably not usable.