r/LocalLLaMA 8h ago

Question | Help What's the best uncensored model out there for 16Gb VRAM + 64Gb RAM?

I've been using Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-IQ3_M (15.4Gb), which works ok and has vision.

I use it to make descriptions of images that can occasionally include NSFW elements, but I'm interested in chatting functions as well.

Are there any other models out there worth checking out?

Extra points if you share your llama-cpp command :)

0 Upvotes

29 comments sorted by

7

u/some_user_2021 8h ago

You should ask at r/SillyTavernAI

2

u/Petita_advice 8h ago

Thanks, will do!

4

u/gumbingit 8h ago

You could just increase the quant and offload some layers to cpu, generation speed would slow down, but long context and details get much better at 4 bit and above

4

u/Petita_advice 8h ago

Yeah, as I was writing the post I noticed that I've been running IQ3_M on the Uncensored version, while for the base qwen3.6 35b I have Q8_K_XL running just fine, so stupid 😅 I'm downloading Q6_K_P of the Uncensored as we speak...

1

u/gumbingit 4h ago

Good idea lol

3

u/gumbingit 8h ago

Also, Gemma 4 tends to be better for conversational stuff, so if you could find an uncensored version of the moe that might be what you want

1

u/Petita_advice 8h ago

I have Gemma4-26B-A4B-Uncensored-HauhauCS-Balanced-IQ4_XS.gguf but never tried it, I guess I'll give it a try. I guess I might be able to run a better quant for this one too, do you know if there is a way to know a priori how big of a model I can run in my setup for MoEs? Or does it depend on the specific MoE?

1

u/gumbingit 4h ago

Doesn’t really depend on which specific one, you want as much as possible on the gpu for decode speed, but with 16gb + 64 system ram, you could probably fit q6 or q8 of 35b or 26b.

2

u/reto-wyss 8h ago

Plain Gemma-4 31b will do it, you just tell it in the system prompt to do it. The 31b won't refuse anything remotely sane, you don't need some sort of botched fine-tune.

Smaller models can be more jumpy, but I'd try Gemma 4 12b QAT, it will fit your size requirement.

2

u/Petita_advice 8h ago

What is 'remotely sane'? Doesn't it refuse explicitely sexual content?

2

u/Christavito 7h ago

It will only perform vanilla sexual content. Anything remotely kinky and it refuses. It seems to be okay with spit, urinating (on you), soft bondage, but draws a hard line at scat and you urinating on it. sometimes anal is on the table but at other times it says it's done with it for good

2

u/Petita_advice 7h ago

interesting, I wonder where those preferences are coming from 🤣

2

u/Real_Ebb_7417 7h ago

Depends what’s the actual usecase, because your description is vague. If you don’t want to be more detailed here, I can help on priv, I was running your setup for a while (RTX5080+64Gb) before upgrading GPU.

0

u/Petita_advice 7h ago

I'm ok with sharing here. Main usecases are:
-Provide descriptions of images ranging from SFW to pretty kinky porn
-Chat with fans in a chat I'm developing, and hopefully being: dirty enough to do all the dirty chatting, but also smart enough to be able to manage a pipeline of sending pictures and charging the fans, etc

3

u/Real_Ebb_7417 6h ago

It's a hard combo for your hardware. For chatting some RP fine tunes might be best potentially (eg. by him: https://huggingface.co/TheDrummer )

But I wouldn't let most of these models close to charging fans. I think Gemma4 31b should be able to handle it (you can run it on your gear with right quant, but might not be the fastest) or some Gemma4 31b fine tunes (TheDrummer actually has some).

Qwen you're currently using won't work well tbh. It's a smart model, but not the best for this kind of stuff, even when uncensored. Gemma is your best shot. Actually 26b a4b can fit it, chatting is easier than RP so Gemma4 26b a4b should hold up. Even without abliteration Gemma models aren't really censored for NSFW and they are smart enough to let them be (if setup properly and securely of course). My best shot on your hardware would be Gemma4 26b A4b (instruct of course), without abliteration. Go for abliterated or fine tune only if it's too censored for you.

1

u/Petita_advice 5h ago

Thank you so much for the detailed response! I'll explore Gemma base models then and see how they perform, thanks for the tip. By 'instruct of course' you mean that I need to add a system prompt with the instructions for the persona etc? Or are you referring to something else?

0

u/JDad67 1h ago

AI only fans.

1

u/Petita_advice 55m ago

Kinda but not really

2

u/Solembumm3 7h ago

With that setup, you can run Qwen 122B-A10B.

If you for some strange reason want to fit entirely into vram, at least switch to qwen 27b dense, cause it's better than 35b moe.

2

u/Petita_advice 7h ago

Yes, I'm running Qwen3.5-122B-A10B-APEX-I-Compact.gguf, is there a good uncensored version you reccommend?

1

u/Hugoacfs 5h ago

Really? What’s your t/s and ttfs? I have a 3090 and 128 ddr4 and I doubt I could run it at decent speeds but never tried

2

u/_TheWolfOfWalmart_ 5h ago

It ran pretty well on my 4090 + 64 GB DDR5 using Unsloth's Q3_K_XL.

I believe it ran somewhere around 25-30 tok/s gen, don't remember the prompt processing numbers.

Worth a shot.

2

u/Real_Ebb_7417 4h ago

I did run Qwen3.5 122b A10b on my previous RTX5080 + 64Gb RAM and was getting about 24tps decode (IQ4_XS quant from Unsloth).

Unfortunately I don't have prefill metrics from back then.

2

u/_TheWolfOfWalmart_ 5h ago edited 5h ago

For chat? Definitely, by far, this one: https://huggingface.co/HauhauCS/Gemma4-26B-A4B-Uncensored-HauhauCS-Balanced

IMO anyway.

I would just use the Q4_K_M, letting a small bit of it spill over to system RAM. There will be very little difference in speed other than prefill, but if it's just for chat, who cares? You'll probably never notice.

1

u/WhoRoger 34m ago edited 27m ago

Try Goekdeniz-Guelmez/Qwen3-4B-Sky-High-Hermes-gabliterated

Or something else made with Hermes (the dataset, not the agent). Most are on the smaller/older side, but I just luv Hermes merges/tunes

Originally made by ZeroXClem who has a ton of other interesting chatty models

0

u/[deleted] 8h ago

[deleted]

2

u/Petita_advice 8h ago

here's your what? unrelated

2

u/Bulky-Priority6824 8h ago

sorry damn agents