# Uncensored AI models & self-hosting guides

Uncensored and abliterated AI model news, cloud GPU costs and self-hosting research. Find a model worth running, then get help deploying it.

- [One base, three uncensors: the Ornith-1.5 task-vector transplant](https://abliterated.cloud/blog/ornith-1-5-35b-a3b-uncensored-transplant/): 2026-08-20. DeepReinforce's self-improving Ornith-1.5-35B-A3B, uncensored three ways in 48 hours: a streamed task-vector transplant that adds Qwen3.6's measured uncensoring delta to Ornith's weights (102 tensors changed, vision and MTP intact, 0/16 heuristic refusals on Q4_K_M), plus two classic orthogonalization edits.
- [The lossless aggressive: Qwen3.8 27B Uncensored FP8](https://abliterated.cloud/blog/qwen3-8-27b-uncensored-aggressive/): 2026-08-19. The most-liked Qwen3.8 uncensored on Hugging Face (553 likes): orcarouter's lossless aggressive edit, block-FP8 with the vision tower at full precision, served gated on OrcaRouter at $0.40/$4.21 per 1M tokens.
- [Small uncensored agents: what a 4.5B Heretic distill is for](https://abliterated.cloud/blog/qwen3-5-4b-emperoai-qwen3-8-distill-heretic-abliterated/): 2026-08-17. A source-linked field note on insraq's Heretic v1.4.0 decensor of EmperoAI's Qwen3.8-4B distill: what the 4.5B class is for, and what it trades away.
- [What 'Aggressive' means: Muse-Glimmer-30B abliterated to 0/100 refusals](https://abliterated.cloud/blog/muse-glimmer-30b-abliterated-aggressive/): 2026-08-17. A relaxed-KL LoRA de-abliteration of Meta's 29.8B agentic model claiming 0/100 refusals at ~1.7x KL drift — a mirror upload with the benchmarks left unmeasured.
- [The 2.78-trillion-parameter abliteration nobody can run](https://abliterated.cloud/blog/kimi-k3-abliterated-modal/): 2026-08-17. An abliterated re-upload of Moonshot's 2.78T-parameter Kimi K3 — any-to-any, MXFP4, 96 shards, ~1.56 TB — with zero downloads and no deployment: the scale math behind why nobody can run it.
- [The first abliterated diffusion LLM: 1,000+ tokens a second on one GPU](https://abliterated.cloud/blog/diffusiongemma-26b-e38-abliterated-nvfp4/): 2026-08-16. Abliterating a model that doesn't generate autoregressively, then quantizing it to NVFP4: 0/402 target refusals and 1,053.64 tok/s aggregate on one RTX PRO 6000, 51.68 GB cut to 18.86 GB.
- [The 48-hour abliteration race](https://abliterated.cloud/blog/huihui-qwen3-8-27b-abliterated/): 2026-08-16. Within 48 hours of Qwen3.8-27B's community release, huihui-ai published a 27.8B-parameter abliterated edit that leaves the first 15 layers, MTP and the vision tower untouched, and its GGUF companion landed the same afternoon.
- [Three ARA passes: how RVN got Qwen3.8-27B down to 0–1/100 refusals](https://abliterated.cloud/blog/qwen3-8-27b-rvn-heretic-abliterated-uncensored/): 2026-08-14. Abliteration as a matrix optimization problem, run three times: KL 0.0535 → 0.0085, refusals 3/100 → 0–1/100, 106K downloads in four days, one corrupted quant and a loud community pushback.
- [DeepSeek V4 Flash, uncensored by dial: 757 KB against 284B](https://abliterated.cloud/blog/huihui-deepseek-v4-flash-0731-abliterated/): 2026-08-13. A 757 KB refusal-directions file turns DeepSeek V4 Flash 0731 abliteration into a runtime λ dial, with measured proof that the baked λ=2.5 checkpoint overshoots and inverts the direction.

[Get self-hosting help on Signal](https://signal.me/#p/+13103408213) · [All articles](https://abliterated.cloud/blog/)

Model research, dated at publication. Model licenses, publisher benchmarks and hosting estimates are specific to each article, not a live availability or price list. Reported zero-refusal results are test-specific, not a universal guarantee.
