DeepSeek: DeepSeek V4.1 Flash
deepseek/deepseek-v4.1-flashDeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on...
input
$0.34/M
output
$1.38/M
context
1.0M
created
Sep 14, 2026
Supported API shape
input
text · image
output
text
tools
Not listed
json mode
Not listed
Verification
receipt
x-receipt-id
attestation
gateway report
session
attested upstream
provider
Phala
Provider
Phala
Intel TDX
input
$0.34/M
output
$1.38/M
context
1.0M
Performance comparison
Last 72h · UTCMore models
Other private inference routes.
NVIDIA: Nemotron 3.5 Lightning
NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...
context
262K
input
$0.08/M
Z.ai: GLM 5.3
GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window. Served as a text-only TEE deployment via Phala.
context
1.0M
input
$1.40/M
Qwen: Qwen3.8 27B
Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal interaction, and long-running agent tasks, with flexible thinking that can be enabled or disabled. Served on Phala in a TDX-attested enclave.
context
262K
input
$0.24/M
Meta: Muse Glimmer 30B
Muse Glimmer 30B is a dense, open-weight multimodal model from Meta Superintelligence Labs, distilled from Muse Spark and optimized for autonomous agents on consumer hardware. It is suited for long-horizon agentic and coding workflows, with multi-step reasoning, reliable tool use, failure recovery, image understanding, and multilingual support across more than 100 languages.
context
131K
input
$0.30/M