Qwen2.5 7B Instruct
qwen/qwen-2.5-7b-instructQwen2.5 7B is the latest series of Qwen large language models. Qwen2.5 brings the following improvements upon Qwen2: - Significantly more knowledge and has greatly improved capabilities in coding and mathematics, thanks to our specialized expert models in these domains. - Significant improvements in instruction following, generating long texts (over 8K tokens), understanding structured data (e.g, tables), and generating structured outputs especially JSON. More resilient to the diversity of system prompts, enhancing role-play implementation and condition-setting for chatbots. - Long-context Support up to 128K tokens and can generate up to 8K tokens. - Multilingual support for over 29 languages, including Chinese, English, French, Spanish, Portuguese, German, Italian, Russian, Japanese, Korean, Vietnamese, Thai, Arabic, and more. Usage of this model is subject to [Tongyi Qianwen LICENSE AGREEMENT](https://huggingface.co/Qwen/Qwen1.5-110B-Chat/blob/main/LICENSE).
input
$0.10/M
output
$0.20/M
context
33K
created
Oct 3, 2025
Supported API shape
input
text
output
text
tools
Supported
json mode
Not listed
Verification
receipt
x-receipt-id
attestation
gateway report
session
attested upstream
provider
Phala
Provider
Phala
GPU TEE
input
$0.10/M
output
$0.20/M
context
33K
Performance comparison
Last 72h · UTCMore models
Other private inference routes.
Phala: Qwen3.8 27B Uncensored (Aggressive)
HauhauCS Qwen3.8-27B Uncensored Aggressive, using the pinned Q8_K_P GGUF and BF16 vision projector. Served with SGLang on a Phala TDX-attested H200, with text and image input and a 262144-token context.
context
262K
input
$0.30/M
NVIDIA: Nemotron 3.5 Lightning
NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that...
context
262K
input
$0.07/M
Z.ai: GLM 5.3
GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window. Served as a text-only TEE deployment via Phala.
context
1.0M
input
$1.40/M
Qwen: Qwen3.8 27B
Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal interaction, and long-running agent tasks, with flexible thinking that can be enabled or disabled. Served on Phala in a TDX-attested enclave.
context
262K
input
$0.24/M