ReadyGPU TEE

Qwen: Qwen3 Coder 480B A35B

Model IDqwen/qwen3-coder-480b-a35b-instruct

Qwen3-Coder-480B-A35B-Instruct is a Mixture-of-Experts (MoE) code generation model developed by the Qwen team. It is optimized for agentic coding tasks such as function calling, tool use, and long-context reasoning over repositories. The model features 480 billion total parameters, with 35 billion active per forward pass (8 out of 160 experts).

Start building Docs

input

$2.00/M

output

$2.00/M

context

262K

created

Nov 28, 2025

Supported API shape

input

text

output

text

tools

Not listed

json mode

Not listed

Verification

signature

response ID

attestation

GPU TEE

provider

1 routes

Providers

tinfoil

GPU TEE

input

$2.00/M

output

$2.00/M

context

262K

$ pip install openai

1from openai import OpenAI
2import os
3
4MODEL = "qwen/qwen3-coder-480b-a35b-instruct"
5
6client = OpenAI(
7    base_url="https://api.redpill.ai/v1",
8    api_key=os.environ["REDPILL_API_KEY"],
9)
10
11response = client.chat.completions.create(
12    model=MODEL,
13    messages=[{"role": "user", "content": "Run this privately."}],
14)
15
16print(response.id)
17print(response.choices[0].message.content)

$ npm install openai

1import OpenAI from "openai"
2
3const MODEL = "qwen/qwen3-coder-480b-a35b-instruct"
4
5const openai = new OpenAI({
6  baseURL: "https://api.redpill.ai/v1",
7  apiKey: process.env.REDPILL_API_KEY,
8})
9
10const response = await openai.chat.completions.create({
11  model: MODEL,
12  messages: [{ role: "user", content: "Run this privately." }],
13})
14
15console.log(response.id)
16console.log(response.choices[0].message.content)

1curl https://api.redpill.ai/v1/chat/completions \
2  -H "Authorization: Bearer $REDPILL_API_KEY" \
3  -H "Content-Type: application/json" \
4  -d '{
5    "model": "qwen/qwen3-coder-480b-a35b-instruct",
6    "messages": [
7      { "role": "user", "content": "Run this privately." }
8    ]
9  }'

More models

Other private inference routes.

View catalog

encrypted

Qwen: Qwen3.5-27B

The Qwen3.5 27B native vision-language Dense model incorporates a linear attention mechanism, delivering fast response times while balancing inference speed and performance. Its overall capabilities are comparable to those of the Qwen3.5-122B-A10B.

context

262K

input

$0.30/M

encrypted

Qwen: Qwen3 Embedding 8B

The Qwen3 Embedding model series is the latest proprietary model of the Qwen family, specifically designed for text embedding and ranking tasks. This series inherits the exceptional multilingual capabilities, long-text understanding, and reasoning skills of its foundational model. The Qwen3 Embedding series represents significant advancements in multiple text embedding and ranking tasks, including text retrieval, code retrieval, text classification, text clustering, and bitext mining.

context

33K

input

$0.01/M

encrypted

Qwen: Qwen3 VL 30B A3B Instruct

Qwen3-VL-30B-A3B-Instruct is a multimodal model that unifies strong text generation with visual understanding for images and videos. Its Instruct variant optimizes instruction-following for general multimodal tasks. It excels in perception of real-world/synthetic categories, 2D/3D spatial grounding, and long-form visual comprehension, achieving competitive multimodal benchmark results. For agentic use, it handles multi-image multi-turn instructions, video timeline alignments, GUI automation, and visual coding from sketches to debugged UI. Text performance matches flagship Qwen3 models, suiting document AI, OCR, UI assistance, spatial tasks, and agent research.

context

128K

input

$0.20/M

encrypted

Qwen2.5 7B Instruct

Qwen2.5 7B is the latest series of Qwen large language models. Qwen2.5 brings the following improvements upon Qwen2: - Significantly more knowledge and has greatly improved capabilities in coding and mathematics, thanks to our specialized expert models in these domains. - Significant improvements in instruction following, generating long texts (over 8K tokens), understanding structured data (e.g, tables), and generating structured outputs especially JSON. More resilient to the diversity of system prompts, enhancing role-play implementation and condition-setting for chatbots. - Long-context Support up to 128K tokens and can generate up to 8K tokens. - Multilingual support for over 29 languages, including Chinese, English, French, Spanish, Portuguese, German, Italian, Russian, Japanese, Korean, Vietnamese, Thai, Arabic, and more. Usage of this model is subject to [Tongyi Qianwen LICENSE AGREEMENT](https://huggingface.co/Qwen/Qwen1.5-110B-Chat/blob/main/LICENSE).

context

33K

input

$0.04/M