Nemotron 3 Ultra with Jev
Nemotron 3 Ultra is a model from NVIDIA, released 2026-06-04. It costs $0.600 per million input tokens and $2.40 per million output tokens, reads up to 262K tokens of context, and writes up to 182,520 tokens in one reply. Through llm11, set model to nvidia/nemotron-3-ultra-550b-a55b.
- Input price
- $0.600 / 1M
- Output price
- $2.40 / 1M
- Context window
- 262K tokens
- Max output
- 183K tokens
- Cached input
- $0.120 / 1M
- Released
- 2026-06-04
- Accepts
- text
- Intelligence index
- 22.9
Read from the provider's prompt cache
Artificial Analysis, via the catalogue
Read from the live catalogue and refreshed daily. At blended list price, 39% of the 332 priced models we can call cost more.
Using Nemotron 3 Ultra with Jev
Jev is TypeSafe AI's System One model. It reads each request first and decides which model in your pool answers, with a calibrated confidence on that call. Nemotron 3 Ultra does the answering when Jev sends it work, or when you name it yourself. At list price Nemotron 3 Ultra falls in the llm11-balanced band, but no pack routes to it because NVIDIA is not one of the labs the packs draw from. You can still pin it: pass nvidia/nemotron-3-ultra-550b-a55b as model and routing steps aside while verification still runs.
Jev is new, so we have not published results for Nemotron 3 Ultra with and without it, and this page does not invent any. Every response comes with a receipt that names the model that answered and what it cost against your most expensive model, so you can measure the pairing on your own traffic.
What a request costs
List price arithmetic, the same sum a receipt does. llm11 adds nothing per request; the only fee is on buying credits.
| Request | Tokens in / out | One | A thousand |
|---|---|---|---|
| Short chat turn | 500 / 200 | $0.00078 | $0.780 |
| Question over retrieved documents | 4,000 / 500 | $0.00360 | $3.60 |
| Long document summary | 30,000 / 1,000 | $0.020 | $20.40 |
Cheaper models to try
Nearest in price and below Nemotron 3 Ultra, each within a few points of it on the intelligence index.
| Model | Input / 1M | Output / 1M | Context | Index | Released |
|---|---|---|---|---|---|
| Qwen Qwen3.6 27B | $0.320 | $3.20 | 262K | 21.4 | |
| Google Gemini 3.5 Flash Lite | $0.300 | $2.50 | 1.05M | 22.2 | |
| Qwen Qwen3.7 Plus | $0.320 | $1.28 | 1M | 25.2 | |
| DeepSeek DeepSeek V4.1 Flash | $0.300 | $1.20 | 1.05M | 39.5 |
Call it
Same OpenAI request shape. Naming the model pins it, so nothing is routed, and the answer is still checked.
python
from openai import OpenAI
client = OpenAI(base_url="https://www.llm11.com/v1", api_key="llm11_live_...")
res = client.chat.completions.create(
model="nvidia/nemotron-3-ultra-550b-a55b",
messages=[{"role": "user", "content": "Hello"}],
)More from NVIDIA
| Model | Input / 1M | Output / 1M | Context | Released |
|---|---|---|---|---|
| Nemotron 3.5 Lightning | $0.060 | $0.160 | 262K | |
| Nemotron 3.5 Content Safety | $0.200 | $0.200 | 131K | |
| Nemotron 3 Super | $0.080 | $0.450 | 262K | |
| Nemotron 3 Nano 30B A3B | $0.050 | $0.200 | 262K |
Questions
- What is Nemotron 3 Ultra with Jev?
- Jev is TypeSafe AI's System One model. It reads each request first and decides which model in your pool answers, with a calibrated confidence on that call. Nemotron 3 Ultra does the answering when Jev sends it work, or when you name it yourself. At list price Nemotron 3 Ultra falls in the llm11-balanced band, but no pack routes to it because NVIDIA is not one of the labs the packs draw from. You can still pin it: pass nvidia/nemotron-3-ultra-550b-a55b as model and routing steps aside while verification still runs. Routing decisions are made by Jev, TypeSafe AI's System One model.
- Is Nemotron 3 Ultra a System One model?
- No. Nemotron 3 Ultra answers requests. The System One model is Jev, which decides which model answers each one, so the two do different jobs and llm11 uses both.
- How much does Nemotron 3 Ultra cost?
- $0.600 per million input tokens and $2.40 per million output tokens. A 4,000 token prompt with a 500 token answer costs about $0.00360, so a thousand of them is about $3.60. llm11 passes provider list price through and charges 5% when you buy credits.
- What is the Nemotron 3 Ultra context window?
- 262,144 tokens, with up to 182,520 tokens in a single reply.
- Does Nemotron 3 Ultra support tool calling and structured output?
- The catalogue lists tool calling, structured output and reasoning controls.
- Can I use Nemotron 3 Ultra through llm11?
- Yes. Set model to nvidia/nemotron-3-ultra-550b-a55b on the OpenAI-compatible endpoint and the request goes to it directly, with verification still running.
- Does Jev route to Nemotron 3 Ultra?
- Routing decisions are made by Jev, TypeSafe AI's System One model. At list price Nemotron 3 Ultra falls in the llm11-balanced band, but no pack routes to it because NVIDIA is not one of the labs the packs draw from. You can still pin it: pass nvidia/nemotron-3-ultra-550b-a55b as model and routing steps aside while verification still runs.
- What is a cheaper alternative to Nemotron 3 Ultra?
- Nearest in price and below it: Qwen Qwen3.6 27B at $0.320 in and $3.20 out, Google Gemini 3.5 Flash Lite at $0.300 in and $2.50 out and Qwen Qwen3.7 Plus at $0.320 in and $1.28 out. Nearest in price and below Nemotron 3 Ultra, each within a few points of it on the intelligence index.