Models
ReadyGPU TEE

Phala: Qwen3.6 35B-A3B Uncensored (Aggressive)

Model IDphala/qwen3.6-35b-a3b-uncensored

Uncensored "Aggressive" variant of Qwen3.6-35B-A3B from Alibaba's Qwen team. The fine-tune by HauhauCS removes refusal behaviors (0/465 refusals) without modifying datasets or core capabilities. The base architecture is a 35B-parameter Mixture-of-Experts model with 256 experts routing 8 per token (~3B active params), 40 layers, and a hybrid linear+full-softmax attention mechanism (3:1 ratio). Supports a native 262K context and is natively multimodal across text, images, and video. Served on Phala in TDX-attested H200 enclave with end-to-end ECDSA response signing; FP8 quantization by lamianlbe.

input

$0.30/M

output

$1.50/M

context

131K

created

May 23, 2026

Supported API shape

input

text · image

output

text

tools

Supported

json mode

Supported

Verification

receipt

x-receipt-id

attestation

gateway report

session

attested upstream

provider

Phala

Provider

Phala

GPU TEE

input

$0.30/M

output

$1.50/M

context

131K

Performance comparison

Last 72h · UTC