DeepSeek-V4.1-Flash (TEE)
Open weights, run inside a TEE. Inference happens sealed, so the prompt stays inside the enclave.
- Tool calling: yes
- Image input: yes
Released
Sep 2026Modalities
Text · Tool calling · Image inputBest reported price on SayGM
-10.0%In
Cached
Out
per Mtok*
Call DeepSeek-V4.1-Flash (TEE) from the API
Create an API key and set it as SAYGM_API_KEY in your terminal.
curl "https://api.saygm.com/v1/chat/completions" \
-H "Authorization: Bearer $SAYGM_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek-v4.1-flash-tee",
"messages": [
{
"role": "user",
"content": "Hello!"
}
]
}'SayGM vs. list price
SayGM's best reported rate for deepseek-v4.1-flash-tee next to the published list price — 10.0% off.
Compare every DeepSeek model across providers in our DeepSeek pricing breakdown.
| Dimension | SayGM price | List price |
|---|---|---|
| Input | $0.27per Mtok | $0.30 |
| Output | $1.08per Mtok | $1.20 |
| Cache read | $0.0054per Mtok | $0.006 |
Recent effective price
What buyers have actually paid on this model recently, across every provider and cache hits.
Input
$0.27
$0.30Cached input
$0.005
$0.006Output
$1.08
$1.20How the effective price is calculated
Requests are load balanced across every provider serving this model, each offering its own rate, so what you pay is a blend rather than a single number. These are the blended rates actually paid — priced at what a provider offered, capped at list.
Measured over 6 windows: 395 requests and 6,298,855 input-side tokens. Most recent window closed 2026-10-10 03:15 UTC.
Who's serving this model
Every provider currently serving this model, ranked by discount.
| Provider | Discount | In | Out |
|---|---|---|---|
| Provider 1 | 10.0% off | $0.27 | $1.08 |
Performance
Measured on the requests SayGM routed to this model.
- Throughput
- 51.1 tok/sover 672 requests
- Time to first token
- 2.3 sover 664 requests
- Request success rate
- 100.0% over 32 requests
Benchmarks
Fixed-seed suites SayGM runs against this model.
More from DeepSeek
Explore the DeepSeek lab and compare related models.
Frequently asked questions
How much does deepseek-v4.1-flash-tee cost on SayGM?
The best reported rate for deepseek-v4.1-flash-tee on SayGM is $0.27 per million input tokens and $1.08 per million output tokens.
How much cheaper is SayGM than list price?
SayGM's best reported rate for deepseek-v4.1-flash-tee is 10.0% off list price.
Is deepseek-v4.1-flash-tee OpenAI-compatible on SayGM?
Yes. deepseek-v4.1-flash-tee is served on SayGM's OpenAI-compatible chat/completions surface, so an unmodified OpenAI SDK pointed at SayGM's base URL works with it.
Related reading
What is private AI inference? →
- Top 5 OpenRouter Alternatives (With Real Pricing)OpenRouter earned its spot as the default multi-model gateway, but its 5.5% platform fee and policy-based privacy model have left room for real competition. This breakdown compares five alternatives, including SayGm's TEE-verified routing, with real, current pricing for each.
- DeepSeek API Pricing 2026: Peak, Off-Peak and ProvidersDeepSeek is the only major lab that charges by time of day, and most traffic is still running on a version that's been superseded. Here's what peak and off-peak actually cost, which model you should be on, and the cheapest provider for each.
- NEAR AI's Confidential Models, Now on SayGMSayGM now routes requests to NEAR AI Cloud, bringing DeepSeek, GLM, Kimi, and Qwen into the same gateway that already reaches Claude, GPT, and Gemini. Here's what NEAR AI's confidential compute actually guarantees, and why it's a notable addition to SayGM's confidential tier.