I like that v4 flash is so fast! I run it on both FireWorks.ai in the US and bought some tokens directly from DeepSeek as an experiment. I only work on Open Source projects, so I don’t have to worry about my work being used to train models - I welcome AI’s being trained on my open content books and code (but not my conventionally published books: I am a party to the copyright suit against Anthropic).
Yea, Flash is quite fast, though looking at model data on Open Router some of the other models are quite fast (Muse Spark, Grok, etc). I’m sure all these models have been trained on my conventionally published books as well, but I don’t care.
The model is fantastic. And costs almost nothing. The only problem I see is that they will train on your data.
There are zero-data-retention providers of DeepSeek models, of which I have used openrouter (with zdr guardrails), and fireworks. But these are 3x to 5x more expensive than directly using DeepSeek, possibly due to poor caching. Thats the price to pay for zdr.
I use these guys. https://crof.ai/tos. Their prices match deepseek, pretty dang fast, and support ZDR. Maybe they serve quantized models but I havent seen a drop in quality from my evals using them vs Deepseek or even for other models.
These models were trained on massive scale copyright infringement, I really don't think they're drawing the line at training on the requests you send them
I use it from pi.dev as well through the OpenCode Go $10 subscription ($5 first month).
Used more than 20M tokens at a cost of ~$20 (up to $60 is included in the $5 plan)
Out of which deepseek pro had ~200 messages which is around 1.5M tokens (10+M cached)
BTW the quotas for Go have very recently changed, now only $15 for some models instead of $60. Which is not actually a difference for DS4 Pro, because they lowered the token pricing 4x at the same time (to match the change in official pricing from DeepSeek months ago)
It doesn't matter if it's cheaper, specially if it consumes more resources to do the same task as the competition
Besides, in a few days, they'll change their pricing, doubling it during their peak hours, so, realistically:
- It will be 2x more expensive if you live in their time zone
- It will be 1.5x more expensive if you live in a time zone that is adjacent to theirs
- It will be the same price IF you use it while they sleep (during offpeak hours)
It's still cheap, but the price/performance ratio is not that good
DeepSeek V4 didn't produce the same impact as V3, and Huawei dropping the ball is making it worse
They had promised massive price cuts for July, so now (Huawei chips), but they had to rush the cuts because lack of momumtum (they advertised them as promotion), and are now backtracking by introducing this peak hours pricing
Trump decided to help them a little by allowing them to buy more NVIDIA chips, so what exactly is China's role in all of this?
We are supposed to blindly pat them in the back while praising them, all while handing them over our data? I thought they were dangerous competition threatening our model of society
I was not referring to the input/output price, but the cost of doing a specific tasks, in practice it is ~10x cheaper than GLM-5.2 for example, to accomplish the same task (for the tasks it can do).
I have been happily using DeepSeek V4 Flash for the last couple of months now. I tried GLM-5.2 for a while, but it was too slow and verbose compare to DeepSeek V4 Flash. If I have a basic skill I need to execute, DeepSeek V4 flash is still the best model for it.
Over the past few weeks while using pro from them directly I have had an increasing number of responses that are obviously from a much, much better model. It is so good that the closed model dog and pony show is already spinning fud about "dark routing" and "stolen directly from fable"
Even at their new pricing it is a genuinely ridiculous amount of value. If you are the type of person who, very reasonably, does not have time to be trying out every model, and just want to use what seems to be the best currently... don't try it. You will be sick to your stomach with buyers remorse as you start to internalize just how much more you could have accomplished had you spent the first six months of the year giving them $1200 instead of OpenAI.
Of course, but why be logical and think about the situation critically when you're pushing very hard for regulatory capture against competitors that give their weights away and provide services that are more reliable, offer a better value, and, this is the worst part, they're from CHINA.
Some of the accusations were going so far as to imply that they were outright routing your requests to anthropic and logging its/your response. It's kind of pathetic