ReadyGPU TEE

DeepSeek: DeepSeek V4 Pro

Model IDdeepseek/deepseek-v4-pro

DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding, and tool use.

Start building Docs

input

$1.50/M

output

$5.25/M

context

800K

created

May 26, 2026

Supported API shape

input

text

output

text

tools

Not listed

json mode

Not listed

Verification

receipt

x-receipt-id

attestation

gateway report

session

attested upstream

provider

Phala

Provider

Phala

GPU TEE

input

$1.50/M

output

$5.25/M

context

800K

Performance comparison

Last 72h · UTC

$ pip install openai

1from openai import OpenAI
2import os
3
4MODEL = "deepseek/deepseek-v4-pro"
5
6client = OpenAI(
7    base_url="https://inference.phala.com/v1",
8    api_key=os.environ["PHALA_API_KEY"],
9)
10
11response = client.chat.completions.create(
12    model=MODEL,
13    messages=[{"role": "user", "content": "Run this privately."}],
14)
15
16print(response.id)
17print(response.choices[0].message.content)

$ npm install openai

1import OpenAI from "openai"
2
3const MODEL = "deepseek/deepseek-v4-pro"
4
5const openai = new OpenAI({
6  baseURL: "https://inference.phala.com/v1",
7  apiKey: process.env.PHALA_API_KEY,
8})
9
10const response = await openai.chat.completions.create({
11  model: MODEL,
12  messages: [{ role: "user", content: "Run this privately." }],
13})
14
15console.log(response.id)
16console.log(response.choices[0].message.content)

1curl https://inference.phala.com/v1/chat/completions \
2  -H "Authorization: Bearer $PHALA_API_KEY" \
3  -H "Content-Type: application/json" \
4  -d '{
5    "model": "deepseek/deepseek-v4-pro",
6    "messages": [
7      { "role": "user", "content": "Run this privately." }
8    ]
9  }'

More models

Other private inference routes.

View catalog

encrypted

Z.ai: GLM 5.2

GLM-5.2 is Z.ai's flagship model for the era of long-horizon tasks. With a truly usable 1M-token context window, it can handle project-level engineering context and execute long-running tasks more reliably. Served as a text-only TEE deployment via Phala.

context

1.0M

input

$1.40/M

encrypted

Google: Gemma 4 31B

Gemma 4 31B Instruct is Google DeepMind's 30.7B dense model. Features a 256K token context window, configurable thinking/reasoning mode, native function calling, and strong multilingual performance. Served as a text-only TEE deployment via NEAR AI.

context

262K

input

$0.15/M

encrypted

Phala: Gemma-4 26B-A4B Uncensored (Heretic)

Uncensored "Heretic" variant of google/gemma-4-26B-A4B-it created using Heretic v1.2.0 with the Arbitrary-Rank Ablation (ARA) method and row-norm preservation. Refusals drop from 100/100 to 11/100 with KL divergence 0.0499 vs the base model. The base Gemma 4 26B A4B is a Mixture-of-Experts model with 25.2B total / 3.8B active parameters (8 active / 128 total experts), 30-layer transformer with hybrid local sliding (1024) + global attention, supporting a 256K context window. Natively multimodal (text + images, variable aspect ratios). Strong on coding, reasoning, function calling, with native system prompt support across 35+ languages. Served on Phala in TDX-attested H200 enclave with end-to-end ECDSA response signing; vLLM-compatible FP8-Static quantization by cloud19 (router excluded from quantization).

context

66K

input

$0.15/M

encrypted

Phala: Qwen3.6 35B-A3B Uncensored (Aggressive)

Uncensored "Aggressive" variant of Qwen3.6-35B-A3B from Alibaba's Qwen team. The fine-tune by HauhauCS removes refusal behaviors (0/465 refusals) without modifying datasets or core capabilities. The base architecture is a 35B-parameter Mixture-of-Experts model with 256 experts routing 8 per token (~3B active params), 40 layers, and a hybrid linear+full-softmax attention mechanism (3:1 ratio). Supports a native 262K context and is natively multimodal across text, images, and video. Served on Phala in TDX-attested H200 enclave with end-to-end ECDSA response signing; FP8 quantization by lamianlbe.

context

131K

input

$0.30/M