Phala: Gemma-4 26B-A4B Uncensored (Heretic)
phala/gemma-4-26b-a4b-uncensoredUncensored "Heretic" variant of google/gemma-4-26B-A4B-it created using Heretic v1.2.0 with the Arbitrary-Rank Ablation (ARA) method and row-norm preservation. Refusals drop from 100/100 to 11/100 with KL divergence 0.0499 vs the base model. The base Gemma 4 26B A4B is a Mixture-of-Experts model with 25.2B total / 3.8B active parameters (8 active / 128 total experts), 30-layer transformer with hybrid local sliding (1024) + global attention, supporting a 256K context window. Natively multimodal (text + images, variable aspect ratios). Strong on coding, reasoning, function calling, with native system prompt support across 35+ languages. Served on Phala in TDX-attested H200 enclave with end-to-end ECDSA response signing; vLLM-compatible FP8-Static quantization by cloud19 (router excluded from quantization).
input
$0.15/M
output
$0.70/M
context
262K
created
May 23, 2026
Supported API shape
input
text · image
output
text
tools
Supported
json mode
Supported
Verification
receipt
x-receipt-id
attestation
gateway report
session
attested upstream
provider
Phala
Provider
Phala
GPU TEE
input
$0.15/M
output
$0.70/M
context
262K
Performance comparison
Last 72h · UTCMore models
Other private inference routes.
Phala: Qwen3.8 27B Uncensored (Aggressive)
HauhauCS Qwen3.8-27B Uncensored Aggressive, using the pinned Q8_K_P GGUF and BF16 vision projector. Served with SGLang on a Phala TDX-attested H200, with text and image input and a 262144-token context.
context
262K
input
$0.30/M
NVIDIA: Nemotron 3.5 Lightning
NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...
context
262K
input
$0.07/M
Z.ai: GLM 5.3
GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window. Served as a text-only TEE deployment via Phala.
context
1.0M
input
$1.40/M
Qwen: Qwen3.8 27B
Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal interaction, and long-running agent tasks, with flexible thinking that can be enabled or disabled. Served on Phala in a TDX-attested enclave.
context
1M
input
$0.20/M