

1·
2 months agoWhat are you using to run the model? Llama.cpp will automatically split the model between your system ram and graphics card’s vram.
Qwen 3.6 is a mixture of experts model with only 3B parameters active at a time. Even without quantization your card could easily run that.


Here to suggest piwigo. Use it with docker, zero hassle but I think it does depend on a database, so it might be too heavy for your use case. Supports tags & galleries.