Llama Guard 4 12B with Jev
Llama Guard 4 12B is a model from Meta, released 2025-04-30. It costs $0.180 per million input tokens and $0.180 per million output tokens, reads up to 164K tokens of context, and writes up to 16,384 tokens in one reply. Through llm11, set model to meta-llama/llama-guard-4-12b.
- Input price
- $0.180 / 1M
- Output price
- $0.180 / 1M
- Context window
- 164K tokens
- Max output
- 16K tokens
- Released
- 2025-04-30
- Accepts
- image, text
Read from the live catalogue and refreshed daily. At blended list price, 79% of the 332 priced models we can call cost more.
Using Llama Guard 4 12B with Jev
Jev is TypeSafe AI's System One model. It reads each request first and decides which model in your pool answers, with a calibrated confidence on that call. Llama Guard 4 12B does the answering when Jev sends it work, or when you name it yourself. At list price Llama Guard 4 12B falls in the llm11-fast band, but no pack routes to it because it does not advertise structured output, which the verification checks rely on. You can still pin it: pass meta-llama/llama-guard-4-12b as model and routing steps aside while verification still runs.
Jev is new, so we have not published results for Llama Guard 4 12B with and without it, and this page does not invent any. Every response comes with a receipt that names the model that answered and what it cost against your most expensive model, so you can measure the pairing on your own traffic.
What a request costs
List price arithmetic, the same sum a receipt does. llm11 adds nothing per request; the only fee is on buying credits.
| Request | Tokens in / out | One | A thousand |
|---|---|---|---|
| Short chat turn | 500 / 200 | $0.00013 | $0.126 |
| Question over retrieved documents | 4,000 / 500 | $0.00081 | $0.810 |
| Long document summary | 30,000 / 1,000 | $0.00558 | $5.58 |
Cheaper models to try
The catalogue has no benchmark score for Llama Guard 4 12B, so this is a price ordering. It does not claim these answer as well.
| Model | Input / 1M | Output / 1M | Context | Released |
|---|---|---|---|---|
| DeepSeek DeepSeek V4 Flash 0423 | $0.140 | $0.280 | 1.05M | |
| Google Gemini 2.5 Flash LiteRetires 2026-10-20 | $0.100 | $0.400 | 1.05M | |
| OpenAI GPT-4.1 Nano | $0.100 | $0.400 | 1.05M | |
| Google Gemma 3 27B | $0.080 | $0.450 | 131K |
Call it
Same OpenAI request shape. Naming the model pins it, so nothing is routed, and the answer is still checked.
python
from openai import OpenAI
client = OpenAI(base_url="https://www.llm11.com/v1", api_key="llm11_live_...")
res = client.chat.completions.create(
model="meta-llama/llama-guard-4-12b",
messages=[{"role": "user", "content": "Hello"}],
)More from Meta
| Model | Input / 1M | Output / 1M | Context | Released |
|---|---|---|---|---|
| Llama 4 Maverick | $0.188 | $0.652 | 1.05M | |
| Llama 4 Scout | $0.100 | $0.300 | 1.31M | |
| Llama 3.3 70B Instruct | $0.100 | $0.320 | 131K | |
| Llama 3.2 1B Instruct | $0.027 | $0.201 | 60K | |
| Llama 3.2 3B Instruct | $0.050 | $0.330 | 131K | |
| Llama 3.1 70B Instruct | $0.400 | $0.400 | 131K |
Questions
- What is Llama Guard 4 12B with Jev?
- Jev is TypeSafe AI's System One model. It reads each request first and decides which model in your pool answers, with a calibrated confidence on that call. Llama Guard 4 12B does the answering when Jev sends it work, or when you name it yourself. At list price Llama Guard 4 12B falls in the llm11-fast band, but no pack routes to it because it does not advertise structured output, which the verification checks rely on. You can still pin it: pass meta-llama/llama-guard-4-12b as model and routing steps aside while verification still runs. Routing decisions are made by Jev, TypeSafe AI's System One model.
- Is Llama Guard 4 12B a System One model?
- No. Llama Guard 4 12B answers requests. The System One model is Jev, which decides which model answers each one, so the two do different jobs and llm11 uses both.
- How much does Llama Guard 4 12B cost?
- $0.180 per million input tokens and $0.180 per million output tokens. A 4,000 token prompt with a 500 token answer costs about $0.00081, so a thousand of them is about $0.810. llm11 passes provider list price through and charges 5% when you buy credits.
- What is the Llama Guard 4 12B context window?
- 163,840 tokens, with up to 16,384 tokens in a single reply.
- Does Llama Guard 4 12B support tool calling and structured output?
- The catalogue does not list tool calling, structured output and reasoning controls.
- Can I use Llama Guard 4 12B through llm11?
- Yes. Set model to meta-llama/llama-guard-4-12b on the OpenAI-compatible endpoint and the request goes to it directly, with verification still running.
- Does Jev route to Llama Guard 4 12B?
- Routing decisions are made by Jev, TypeSafe AI's System One model. At list price Llama Guard 4 12B falls in the llm11-fast band, but no pack routes to it because it does not advertise structured output, which the verification checks rely on. You can still pin it: pass meta-llama/llama-guard-4-12b as model and routing steps aside while verification still runs.
- What is a cheaper alternative to Llama Guard 4 12B?
- Nearest in price and below it: DeepSeek DeepSeek V4 Flash 0423 at $0.140 in and $0.280 out, Google Gemini 2.5 Flash Lite at $0.100 in and $0.400 out and OpenAI GPT-4.1 Nano at $0.100 in and $0.400 out. The catalogue has no benchmark score for Llama Guard 4 12B, so this is a price ordering. It does not claim these answer as well.