Qwen: Qwen3 VL 30B A3B Instruct
qwen/qwen3-vl-30b-a3b-instructQwen3-VL-30B-A3B-Instruct is a multimodal model that unifies strong text generation with visual understanding for images and videos. Its Instruct variant optimizes instruction-following for general multimodal tasks. It excels in perception of real-world/synthetic categories, 2D/3D spatial grounding, and long-form visual comprehension, achieving competitive multimodal benchmark results. For agentic use, it handles multi-image multi-turn instructions, video timeline alignments, GUI automation, and visual coding from sketches to debugged UI. Text performance matches flagship Qwen3 models, suiting document AI, OCR, UI assistance, spatial tasks, and agent research.
input
$0.20/M
output
$0.70/M
context
128K
created
Nov 28, 2025
Supported API shape
input
text · image
output
text
tools
Supported
json mode
Supported
Verification
receipt
x-receipt-id
attestation
gateway report
session
attested upstream
provider
Phala
Provider
Phala
GPU TEE
input
$0.20/M
output
$0.70/M
context
128K
Performance comparison
Last 72h · UTCMore models
Other private inference routes.
Meta: Muse Glimmer 30B
Muse Glimmer 30B is a dense, open-weight multimodal model from Meta Superintelligence Labs, distilled from Muse Spark and optimized for autonomous agents on consumer hardware. It is suited for long-horizon agentic and coding workflows, with multi-step reasoning, reliable tool use, failure recovery, image understanding, and multilingual support across more than 100 languages.
context
131K
input
$0.35/M
DeepSeek: DeepSeek V4 Flash 0731
DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows.
context
1.0M
input
$0.20/M
Z.ai: GLM 5.2
GLM-5.2 is Z.ai's flagship model for the era of long-horizon tasks. With a truly usable 1M-token context window, it can handle project-level engineering context and execute long-running tasks more reliably. Served as a text-only TEE deployment via Phala.
context
1.0M
input
$1.13/M
Google: Gemma 4 31B
Gemma 4 31B Instruct is Google DeepMind's 30.7B dense model. Features a 256K token context window, configurable thinking/reasoning mode, native function calling, and strong multilingual performance. Served as a text-only TEE deployment via NEAR AI.
context
262K
input
$0.15/M