Qwen: Qwen3.5-122B-A10B
qwen/qwen3.5-122b-a10bQwen3.5-122B-A10B is a large Mixture-of-Experts model from Alibaba Cloud with 122B total parameters and 10B active parameters per token. Strong on reasoning, coding, and tool calling with 262K context. Served as a text-only TEE deployment via NEAR AI.
input
$0.46/M
output
$3.68/M
context
262K
created
May 26, 2026
Supported API shape
input
text · image · video
output
text
tools
Not listed
json mode
Not listed
Verification
receipt
x-receipt-id
attestation
gateway report
session
attested upstream
provider
Phala
Provider
Phala
GPU TEE
input
$0.46/M
output
$3.68/M
context
262K
Performance comparison
Last 72h · UTCMore models
Other private inference routes.
Qwen: Qwen3.8 27B
Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal interaction, and long-running agent tasks, with flexible thinking that can be enabled or disabled. Served on Phala in a TDX-attested enclave.
context
262K
input
$0.40/M
Meta: Muse Glimmer 30B
Muse Glimmer 30B is a dense, open-weight multimodal model from Meta Superintelligence Labs, distilled from Muse Spark and optimized for autonomous agents on consumer hardware. It is suited for long-horizon agentic and coding workflows, with multi-step reasoning, reliable tool use, failure recovery, image understanding, and multilingual support across more than 100 languages.
context
131K
input
$0.30/M
DeepSeek: DeepSeek V4 Flash 0731
DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows.
context
1.0M
input
$0.20/M
Z.ai: GLM 5.2
GLM-5.2 is Z.ai's flagship model for the era of long-horizon tasks. With a truly usable 1M-token context window, it can handle project-level engineering context and execute long-running tasks more reliably. Served as a text-only TEE deployment via Phala.
context
1.0M
input
$1.26/M