

Hi, thanks for posting this guide!
I tried it on a modern laptop with 32GB of unified memory. It starts at 40 tok/s when loading the context, but the output drops to 3 tok/s.
Could I ask you which GPU you are using for these tests? What context size would a 16GB GPU give?















Same issue for me. My workaround is to remove the anonymous account and let Aurora Store fetch another one. It might take two or three attempts before you get one that works.