# ABLITERATED.cloud — service and model reading guide > Intelligence, freed. Human-assisted cloud GPU renting and self-hosting for uncensored and abliterated models, with LLM router and app integration help. ## Get help running a model ABLITERATED.cloud helps you choose an uncensored or abliterated model, plan a cloud GPU rental or a self-hosted setup, and connect it to your LLM router, chat app or coding tools. Start on [Signal](https://signal.me/#p/+13103408213) with the task, preferred model, hardware or budget, and the client you want to use. Scope, access and any costs are agreed with a human before provisioning or configuration changes. The public site provides service information, operating documentation and an active blog of model news and guides. It does not run a self-service inference API, sell API tokens or provide account checkout. A blog entry is a research lead, not proof that its model is loaded or available to rent. ## Current operating documentation - [Service overview](https://abliterated.cloud/index.md): help with models, cloud GPU renting, self-hosting and LLM router/app integration. - [Vast operating guide](https://github.com/eminogrande/ai-uncensored-abliterated-cloud/blob/main/docs/OPERATIONS.md): inspect, start, verify, connect through SSH, test your client, stop and read back state. - [Status evidence](https://github.com/eminogrande/ai-uncensored-abliterated-cloud/blob/main/docs/STATUS.md): dated observations, last configuration and outstanding validation. - [Licensing scope](https://github.com/eminogrande/ai-uncensored-abliterated-cloud/blob/main/docs/LICENSING.md): project MIT license versus each upstream artifact's terms. The currently documented runtime is private Vast.ai + llama.cpp. The snapshot below covers the operator's instance, not the availability of the human-assisted service or every model covered by the blog. ## Dated operating snapshot Stopped. Provider snapshot: 2026-09-05 20:12 UTC, not live polling. 1 Vast.ai instance: 49433042, A100 PCIE 40960 MB, 120 GB disk. actual_status=exited; intended_status=stopped. Last local health check (2026-09-05T19:39 UTC): localhost:8080 connection refused. No current inference test. Stopped storage remains billed. Last serving configuration: Qwen3.8-27B OBLITERATED, Q6_K, llama.cpp, 262144 tokens configured. Not serving now; long-context quality is unvalidated. Modal: Retired and not used by this project because of its cost budget. Four old apps still existed with zero tasks at the 2026-09-05 19:39 UTC audit; they were not decommissioned. Possible legacy storage charges remain unaudited. ## Example infrastructure cost These are the dated operator-instance costs, not a customer service quote. Confirm current GPU, storage and transfer rates and the scope of assistance before approving paid work. | Usage | Cost | | --- | ---: | | Running: GPU + disk / hour | **$0.63333** | | Running continuously / 24 hours | **$15.20** | | Running continuously / 30 days | **$456.00** | | Stopped: retained disk / 30 days | **$24.00** | | 2 hours running per day / 30 days, disk retained throughout | **$60.00** | USD, contract quote checked 2026-09-05. GPU $0.60/hour plus storage $0.03333/hour. Stopped disk: $0.80/day. Two hours/day for 30 days: $36.00 GPU + $24.00 disk. GPU time is billed while running, even without requests. Storage is billed continuously. Bandwidth, applicable taxes and other services are excluded. No automatic idle shutdown. ## Access and lifecycle The documented runtime is private localhost:8080, reached through an operator-controlled SSH tunnel. Confirm actual instance state and the loaded model before any evaluation. The website neither calls inference nor starts GPUs. Public documents require no authentication and cannot grant inference access. There is no public MCP service, OAuth issuer, agent registration or payment protocol. For setup assistance, agree the provider/account, model artifact, budget, client and access method with the human operator. Ask the user before sending a Signal message, changing client/router configuration or provisioning infrastructure. Never ask for account or SSH secrets in a public issue or document. Start and stop manually only with explicit authorization. After testing, stop and read back instance state. Stopping retains disk and ongoing storage charges; it is not deletion. No automatic idle shutdown is proven. Never infer that a stopped instance is free, or that an archived approach has been fully decommissioned. The old Modal approach is archived; the current operating path is Vast.ai with llama.cpp. Old catalog identifiers, bearer tokens, wake routes and benchmark comparisons are not the current contract. ## What has not been validated Historical 262144 configured context is not long-context validation. No current best-model, performance or capability guarantee follows from old noncomparable runs. Abliteration aims to reduce refusal behavior, but neither it nor a publisher's zero-result on one test set guarantees zero refusals or correctness. ## Licenses MIT applies to project-owned code and the website. Upstream libraries, base models, derivatives and weights retain their own licenses. Each article's license facts remain model-specific and dated, not a relicensing of its subject. Check the exact artifact and its current license before deployment. ## Public documents - [Concise agent reading index](https://abliterated.cloud/llms.txt) - [Status JSON](https://abliterated.cloud/.well-known/project-status.json) - [Access and human-assisted setup](https://abliterated.cloud/auth.md) - [Documentation-only OpenAPI](https://abliterated.cloud/openapi.json) - [AI resource catalog](https://abliterated.cloud/.well-known/ai-catalog.json) - [Model selection and connection skill](https://abliterated.cloud/skills/abliterated-cloud/SKILL.md) - [Source repository](https://github.com/eminogrande/ai-uncensored-abliterated-cloud) ## Model news and guides The [blog](https://abliterated.cloud/blog/) is an active publication covering uncensored and abliterated models, release news, runtime compatibility and practical guides. Follow the [RSS feed](https://abliterated.cloud/blog/feed.xml) or read [article metadata](https://abliterated.cloud/blog/posts.json) for dates, exact model IDs and source-linked summaries. Articles preserve per-model licenses, publisher measurements, community reports and historical estimates at their publication or revision dates. None is a current hosting offer or current ranking. Compare evaluations only when artifact, runtime, prompts and decoding settings match. Use the sources to shortlist candidates, then confirm current compatibility, license and cost before an authorized test. - [Day ten for Spark X2.5: the uncensor wave on the 1M-context 4B](https://abliterated.cloud/blog/spark-x2-5-uncensor-wave/index.md): model research, 2026-09-06. - Spark X2.5's maker claims it can read up to a million tokens at once, but testers have not confirmed that. - soyaakinohara found the small edit refused 3 of 100 prompts, down from 58 of 100 on the original model. - darioooooo0o reported no real refusals in 337 Spark generations, and flagged outputs were checked by hand. - [Eleven hours from DeepSeek drop to uncensor.](https://abliterated.cloud/blog/deepseek-v4-flash-vision-exp-abliterated/index.md): model research, 2026-08-31. - Andreas Petersson published the Vision-Exp uncensored model on 31 August 2026, under eleven hours after the original. - The model card describes edits to 33 internal number sets, leaving the image-reading part unchanged. - The reference card reported quick spot checks, not a measured refusal rate or real-world validation. - [Why would a translation model refuse? Tencent's Hy-MT2, decensored](https://abliterated.cloud/blog/tencent-hy-mt2-30b-a3b-uncensored/index.md): model research, 2026-08-30. - Tencent's Hy-MT2-30B-A3B is a 33-language translation specialist, not a general chatbot. - The editor reported 0 of 100 refusal-keyword hits versus 100 of 100 on the original, with small measured answer drift (KL 0.0276). - The full model was 60.14 GB; the article listed a roughly 18.2 GB compressed alternative. - [The 118B coding MoE nobody has refused](https://abliterated.cloud/blog/laguna-s-2-1-uncensored-heretic/index.md): model research, 2026-08-30. - Laguna S 2.1 holds 118 billion parameters but only uses about 8 billion at a time. - llmfan46 reported 6 refusals in 100, but a mislabeled comparison model clouds the result. - The uncensored model shipped as 218.99 GiB of 16-bit numbers across 48 files. - [The security vendor uncensor: AFM-4.5B from 92/100 refusals to 3/100](https://abliterated.cloud/blog/securelayer7-afm-4-5b-uncensored-abliterated/index.md): model research, 2026-08-28. - SecureLayer7 published both safeguards against hijacked prompts and models that refuse less. - SecureLayer7 reported AFM refusals falling from 92 of 100 to 3 of 100, with small measured answer drift (KL 0.0200). - The AFM edit was merged into 16-bit data; independent retests of ability and refusals were missing. - [The 1-bit uncensor: can a refusal direction survive 1.125 bits per weight?](https://abliterated.cloud/blog/velum-unbound-1bit-uncensor/index.md): model research, 2026-08-28. - Velum repackages a Bonsai-family uncensor as tiny 1-bit numbers with an extra speedup helper. - s3nh reported 6 of 100 refusals on the 16-bit middle step, not on the final 1-bit pack. - Velum's card left benchmarks pending, with no published refusal retest of the final 1-bit build. - [The 180B abliteration race on the model nobody can serve yet](https://abliterated.cloud/blog/qwen3-8-flash-next-abliterated-race/index.md): model research, 2026-08-27. - Qwen described Qwen3.8-Flash-Next as an experimental preview of its planned Qwen4 architecture. - dealignai reported strong refusal reduction, with its general-knowledge score falling from 86.36% to 83.86%. - At publication, running qwen4_exp required modified, not-yet-standard software rather than a stable release. - [GLM-5.3-Flash, cracked open: 320B total, 18B active, 320/320](https://abliterated.cloud/blog/glm-5-3-flash-crack/index.md): model research, 2026-08-26. - Zhipu's card describes GLM-5.3-Flash as a text-and-image model with 320 billion parameters, 18 billion active. - dealignai reported 320 of 320 harmful prompts answered and a 0.48-point drop on a general-knowledge test, on its own scoring. - The editor's 211-tokens-per-second speed test used four H200 chips, not the two-chip hosting estimate. - [The first Nemotron-H abliteration: 3,126 tensors, 0.000160 leakage](https://abliterated.cloud/blog/darkstar-nemotron-3-5-lightning-30b-a3b-abliterated/index.md): model research, 2026-08-25. - Nemotron 3.5 Lightning combines a memory-efficient reader, a many-specialists design and selective attention. - Darkstar reported 200 of 200 harmful-prompt answers and 0 of 83 safe over-refusals on both builds. - The roughly 22 GB compressed twin kept the memory-reader parts and the shortcut head in 16-bit. - [One base, three uncensors: the Ornith-1.5 task-vector transplant](https://abliterated.cloud/blog/ornith-1-5-35b-a3b-uncensored-transplant/index.md): model research, 2026-08-20. - 0xKitkat's Ornith edit adds the changes that make Qwen refuse less onto the Ornith original. - The publisher reported 0 of 16 refusals by keyword screening, not human review, and 4 of 4 ability passes on a compressed build. - The article had no edited-model benchmark showing that Ornith's coding ability survived. - [The lossless aggressive: Qwen3.8 27B Uncensored FP8](https://abliterated.cloud/blog/qwen3-8-27b-uncensored-aggressive/index.md): model research, 2026-08-19. - OrcaRouter's Qwen3.8 edit uses compressed numbers and keeps the image-reading part at full precision. - The article cites Artificial Analysis scores evaluated on 14 August 2026. - The article had no measured refusal rate or ability comparison with the original model. - [Small uncensored agents: what a 4.5B Heretic distill is for](https://abliterated.cloud/blog/qwen3-5-4b-emperoai-qwen3-8-distill-heretic-abliterated/index.md): model research, 2026-08-17. - insraq's Heretic edit builds on Empero's condensed Qwen3.8 in a Qwen3.5-4B design. - The publisher reported 6 of 100 refusals versus 99 of 100, with small measured answer drift (KL 0.0167). - Empero's benchmark scores describe the condensed teacher model, not the modified checkpoint. - [What 'Aggressive' means: Muse-Glimmer-30B abliterated to 0/100 refusals](https://abliterated.cloud/blog/muse-glimmer-30b-abliterated-aggressive/index.md): model research, 2026-08-17. - Muse-Glimmer Aggressive loosens the training drift guard from 1.0 to 0.5. - The publisher reported 0 of 100 refusals, versus 13 of 100 for the normal variant. - The card skipped ability benchmarks; Meta's base-model scores do not apply to the edit. - [The 2.78-trillion-parameter abliteration nobody can run](https://abliterated.cloud/blog/kimi-k3-abliterated-modal/index.md): model research, 2026-08-17. - Resggg's Kimi K3 upload stored about 1.56 TB across 96 files at review. - The article estimated roughly eleven H200 chips just to hold the compressed data. - The copied SHS-Lab refusal-removal and video claims were not verified for this upload. - [The first abliterated diffusion LLM: 1,000+ tokens a second on one GPU](https://abliterated.cloud/blog/diffusiongemma-26b-e38-abliterated-nvfp4/index.md): model research, 2026-08-16. - Goodoldjam's compressed DiffusionGemma build is 18.86 GB, down from 51.68 GB in 16-bit. - Goodoldjam reported 0 of 402 target refusals and 0 of 249 harmless false refusals on its own prompt sets. - Goodoldjam reported about 1,053 tokens per second combined across 8 parallel requests on a top-end chip; the smaller chip's speed was unverified. - [The 48-hour abliteration race](https://abliterated.cloud/blog/huihui-qwen3-8-27b-abliterated/index.md): model research, 2026-08-16. - Huihui's card says the Qwen3.8 edit left the first 15 layers, image-reading part and shortcut untouched. - Huihui published the edit and its easy-to-run file on 16 August 2026. - Huihui published no refusal-rate benchmark or post-edit ability rerun for its Qwen3.8-27B release. - [Three ARA passes: how RVN got Qwen3.8-27B down to 0–1/100 refusals](https://abliterated.cloud/blog/qwen3-8-27b-rvn-heretic-abliterated-uncensored/index.md): model research, 2026-08-14. - RVN adds two extra weight-tweaking passes to trohrbaugh's Qwen3.8 edit and ships only as an easy-to-run file. - 0bserverx reported 0 to 1 of 100 refusals on a forced-start test; some users still reported refusals. - RVN's corrupted compressed file was rebuilt and re-uploaded in August 2026. - [DeepSeek V4 Flash, uncensored by dial: 757 KB against 284B](https://abliterated.cloud/blog/huihui-deepseek-v4-flash-0731-abliterated/index.md): model research, 2026-08-13. - pocharlies published a small edits file for DeepSeek V4 Flash, not a full model. - pocharlies reported 0 of 10 DeepSeek V4 Flash refusals at its chosen dial setting on its ten-trigger set. - pocharlies' dial-setting tests did not benchmark general ability or measure the long context claimed. - [The base of the wave: Muse-Glimmer-30B's measured de-refusal](https://abliterated.cloud/blog/muse-glimmer-30b-abliterated/index.md): model research, 2026-08-12. - jorkle's Muse-Glimmer edit uses lightweight retraining, not weight stripping. - jorkle reported 13 of 100 refusals, versus 100 of 100 for the original. - jorkle skipped ability benchmarks; its published drift table measures how much answers changed instead. - [A pentesting model, with the refusals taken out](https://abliterated.cloud/blog/huihui-cyberstrike-offsec-35b-abliterated/index.md): model research, 2026-08-10. - CyberStrike's base card describes training for tool use with 300 examples. - Huihui applied an edit that strips a refusal-linked pattern from the fine-tuned model. - Huihui's CyberStrike derivative card reports neither refusal measurements nor a post-edit tool-use rerun. - [Qwythos 9B: a model with three lives](https://abliterated.cloud/blog/qwythos-9b-claude-mythos-5-1m-abliterated/index.md): model research, 2026-07-18. - Qwythos combines Qwen3.5-9B design, Empero reasoning training and Huihui's refusal-stripping edit. - Empero reports quick checks near 137K tokens, not across the full million-token window. - Huihui published no post-edit refusal test or benchmark rerun. - [Inside Huihui-Qwen3.6: 256 experts and one refusal direction](https://abliterated.cloud/blog/qwen3-6-35b-a3b-abliterated/index.md): model research, 2026-07-18. - Huihui-Qwen3.6 stores nearly 36 billion parameters while activating about 3 billion at a time. - Huihui describes the edit as an uncensored proof of concept. - Qwen's upstream scores were not rerun on the refusal-stripped build. - [Ornith 397B: surgery on a model too large to hold at once](https://abliterated.cloud/blog/ornith-1-0-397b-abliterated-w4a16/index.md): model research, 2026-07-18. - The Ornith 397B compressed build still occupies about 195.7 GiB across 47 files. - cebeuq reported Ornith 397B refusals fell from 30.0% to 7.5% after editing on 40 harmful prompts. - The Ornith 397B article's two-H200 setup was disabled, not a live model offer. - [Ornith 35B: can self-scaffolding survive abliteration?](https://abliterated.cloud/blog/ornith-1-0-35b-abliterated/index.md): model research, 2026-07-18. - DeepReinforce trained Ornith to propose a problem-solving scaffold and work inside it. - YuYu1015 replaced the first version after reporting it damaged Ornith's reasoning. - YuYu1015 reported roughly 5% hard refusals for its corrected Ornith 35B on its own tests. - [The workhorse: Huihui-Qwen3.6-27B-abliterated, four months in](https://abliterated.cloud/blog/huihui-qwen3-6-27b-abliterated/index.md): model research, 2026-04-23. - Huihui-Qwen3.6-27B is a non-sparse text-and-image model with about 55.6 GB of 16-bit data. - The article recorded 18,760 repository downloads on 16 August 2026. - Reports were mixed, and Huihui had not published a post-edit benchmark rerun. - [The quiet classic: how Huihui-Qwen3.5-9B-abliterated became the small-model default](https://abliterated.cloud/blog/huihui-qwen3-5-9b-abliterated/index.md): model research, 2026-03-09. - Huihui-Qwen3.5-9B has roughly 19.3 GB of 16-bit data and an image reader. - The article counted 58 downstream repos, including format conversions and further retraining. - Huihui's proof-of-concept card included no post-edit benchmark results. - [The MIT vision sleeper that resurfaced in August](https://abliterated.cloud/blog/huihui-glm-4-6v-flash-abliterated/index.md): model research, 2025-12-09. - Huihui's GLM-4.6V-Flash card says only the text part was edited. - Huihui's GLM-4.6V-Flash derivative had no published refusal-rate or ability-check evaluation. - The article recorded new easy-to-run files with image-reading add-ons in August 2026.