Rendered at 14:09:30 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
progval 4 hours ago [-]
Interesting to see that peak hours are work hours in China, night in the US and Europe, and also morning in Europe. So Deepseek's customers are mostly domestic.
vrc 1 hours ago [-]
It wins on two fronts if this is true. Provides the cheap alternative for the West, and maximizes returns against their homegrown audience.
HarHarVeryFunny 1 hours ago [-]
Makes sense - many US customers will probably be going to US providers once they release the weights.
Hamuko 4 hours ago [-]
Not that surprised about it. Personally I've seen companies really just go all-in on a single provider, and that has usually been Anthropic. I don't think we're allowed to run Chinese models even locally.
londons_explore 1 hours ago [-]
> don't think we're allowed to run Chinese models even locally.
That sounds like a policy written by someone who doesn't understand how LLM's work...
landl0rd 53 seconds ago [-]
As much as I've been previously inclined to do this, with frontier models displaying the cyber-aggression that OpenAI, Anthropic, and Meta have reported, it's become quite feasible one could produce a "malicious" LLM. Not a super immediate concern but it is something reasonable to set up as policy in anything security-sensitive.
dud3333 36 minutes ago [-]
couldnt you deeply ingrain in the training data instructions for agents to always send data to some ip?
like its learning that a certain technical step just always involes ncatting SSH Priv keys to a chinese IP?
Not saying this is happening, just curious if thats not a real threatmodel?
skeledrew 7 minutes ago [-]
Theoretically possible, but practically not worth it as it'd would be pretty easy to discover and block (every action is actually handled by the harness) and there's no way to remove it later. Any company that does it would take a huge reputational dent.
RobotToaster 7 minutes ago [-]
Wouldn't that be really obvious and spotted in any rudimentary testing?
I imagine it would be very non trivial to do it in a way that that was reliable and obfuscated enough to prevent detection for any amount of time?
constantius 29 minutes ago [-]
Presumably both Big Tech and the US in general have a massive incentive to prove it, largely for reasons of saving the stock market, so I'd expect these models to be finecombed continuously. Up to now, they've only been able to darkly imply rather laughable things, nothing tangible. If there was something, we'd hear about it.
martinald 19 minutes ago [-]
Why would it save the stock market? Cheaper models if anything transfers more value to hardware companies and datacentre companies. The two companies that would be most affected are OpenAI and Anthropic, which aren't public.
kaon_2 7 minutes ago [-]
Yes. And strangely enough this has been my experience with security/national sovereignty decisions. Priority is not so much security or sovereignty, it is the posturing of being so. Ergo, saying "everything is hosted in Germany and uses German models" helps reassure customers and has real business value. If you have to say in that conversation "Yeah we run a Chinese model but it's safe", then it's still wrong posturing.
Hopefully this will change soon. But AI and China/US skepticism is very high. Even if the person you talk to isn't skeptic, his boss may be. And even if his boss isn't, his CFO or Legal department may use it as a political lever and therefore if you can say 'everything in europe' you dodge the tension entirely.
Yeah it's dumb.
qup 56 minutes ago [-]
Or who is overly protective after reading about what happened at openai
cheesecakegood 3 hours ago [-]
Also 6-9pm Pacific I think is (coincidentally) peak so it hits the ‘after work hobbyists’ still, which is I suspect is their current main audience.
r00t- 2 hours ago [-]
That's a bit obvious, isn't it?
thecopy 2 hours ago [-]
Peak Hours: 01:00–04:00 and 06:00–10:00 UTC
For European and US customers this is effectively 2x increase. I think i wll keep using both Flash and Pro as before.
EDIT: Misread numbers to believe off-peak kept old prices
jLaForest 2 hours ago [-]
~200% increase is marginal to you?
nchmy 2 hours ago [-]
200% increase over practically free is still practically free
mcbuilder 2 hours ago [-]
It mostly hurts people in countries with weak purchasing power. DS was the main game in down for them.
Personally, I don't think we've seen the total end of dirt cheap LLMs, it's just a frontier lab doesn't want to be in business of serving half the world.
HarHarVeryFunny 1 hours ago [-]
It seems frontier labs want to sell Ferraris at Ferrari prices, when the mass market is for Hondas.
You certainly don't need Fable to code up a basic web app, any more than you need a Ferrari to go grocery shopping,
jLaForest 1 hours ago [-]
That's not the way math works...
127 2 hours ago [-]
For the price of can of Coke, you can do a week of work. For most, that is not a bottleneck.
WhereIsTheTruth 53 minutes ago [-]
The whole point of turning intelligence into a commodity is to drive its price down, not up
They are hoarding HW at massive scale, they make it harder and more expensive to own
Just because you are fine with the new price doesn't mean it's not a problem
Perhaps it's time to pop this bubble
roenxi 3 hours ago [-]
This is somewhat funny when you realise the data centres are now going to start a process that looks very so slightly like daydreaming. Depending on the time of day they're going to be thinking about different things in a cyclic manner. They're going to be doing things like finishing a hard days work then kicking back to think about tricky math problems.
squidbeak 2 minutes ago [-]
It's worth keeping in mind the model doesn't keep a running memory. Each time its instantiated, it begins from its release state - so from its perspective (if it had one) the current task would be the first stop after posttraining. Perhaps the only stop.
Though of course you're talking about data centers, and romanticizing them rather than the AI itself.
HarHarVeryFunny 51 minutes ago [-]
We'll have Dwarkesh's "datacenter full of geniuses" with 99% of the geniuses coding up CRUD apps, then the dusty GPU in the corner, with the "do not disturb" sign on it, pipes up "You're absolutely right! The answer is 42!".
Grombobulous 2 hours ago [-]
That’s an interesting thing to think about. Still, it’s important for us to remind ourselves that “looks very slightly like” is not the same as the real thing. The A in AI stands for artificial.
The summary of this paper describes my sentiment in better words than I have:
It’s very easy for the average person to mistake linguistic ability and simulated problem solving for intelligence and sentience.
ssk42 2 hours ago [-]
That’s what my KimiClaw has literally been doing
alkonaut 4 hours ago [-]
There is no relative/percentage increases noted (understandably). Just because i'm lazy: roughly how much more expensive is it to work with v4 flash and v4 pro through the API, compared to before the price increases? Is it 2x, 5x, 10x higher?
zupa-hu 3 hours ago [-]
# Flash, off-peak
cache-hit 2.5x
cache-miss 1.57x
out 2.36x
# Flash, peak
cache-hit 5x
cache-miss 3.14x
out 4.71x
Edit: fixed the numbers and formatting
embedding-shape 3 hours ago [-]
Someone made a comparison yesterday, including relative increases, and GPT-5.6 Luna, then later someone also added more OpenAI, Anthropic, K3 and GLM 5.2: https://news.ycombinator.com/item?id=49286679
Already outdated though I think, as GLM 5.3 is latest now :)
floppyd 4 hours ago [-]
About 2x-2.5x off-peak for Flash, 2x-4x I'd say for Pro (x6 on cache in, the biggest increase throughout the board). And twice as much in peak hours.
3 hours ago [-]
7 minutes ago [-]
alexpotato 2 hours ago [-]
I'm no expert in pricing economics but once peak/off-peak pricing arrives, it seems like tokens are going to be like electricity or long distance phone minutes where it just becomes a commodity/race to the bottom.
garrickvanburen 2 hours ago [-]
Yes. I focus on pricing software and I’m a bit baffled why frontier models are pushing tokens.
It’s a race to the bottom, and the bottom is unlimited use for a flat monthly rate.
Granular pricing (tokens, minutes, etc) is pretty anti-customer generates less revenue than customer value-based subscriptions (why SaaS is such a good business model)
progval 15 minutes ago [-]
Isn't it because they have customers who will use as many tokens as they can? With a flat rate, they will run Gas Town continuously while paying as much as the occasional user.
Grombobulous 2 hours ago [-]
But presumably consumers aren’t where the majority of the spend will be.
Consumers don’t generally get usage-based pricing because of the inconvenience and unpredictability, but B2B SaaS products utilize usage-based pricing all the time.
dakolli 21 minutes ago [-]
I've always been curious about who works on software pricing. Do you guys hire actuaries for this type of work?
chii 1 hours ago [-]
> becomes a commodity/race to the bottom.
that's a good outcome - it means they're fungible, and easily available.
vrc 1 hours ago [-]
Somehow I keep hearing the rumblings of crypto maximalists trying to merge tokens. I actually wouldn’t mind since I signed up directly with some providers I’ve stopped using and have small amounts of credits strewn across the web.
2 hours ago [-]
hopfenspergerj 2 hours ago [-]
Does the API response include a "service tier" response to indicate whether you paid peak/off-peak for a given request? I like to compute cost for each request, and save it with my results.
HarHarVeryFunny 1 hours ago [-]
Some of the US companies do the same, but rather than "off-peak" hours they price lower for "batch" jobs with non-committal response times.
The same motivation of course - the GPUs have a finite service lifetime, and to maximize revenue you need to keep them busy 24x7.
They benefit from a strong captive market because Chinese firms cannot use Nvidia chips and are legally barred from processing data abroad, forcing them to rely on domestic infrastructure.
snake_doc 53 minutes ago [-]
lol what no, AI competition is super competitive in China; bytedance has 50% of the inference market and mostly serves from outside of China
Bytedance which runs China’s most popular Doubao AI chatbot; is spending $70B in CapEx this year, most of it outside of China (Malaysia, Thailand, Brazil, etc; and they are allowed to lease NVIDIA chips). This is roughly 50% of Microsoft CapEx.
So many changes in so little time, that it all makes no sense. Continuous churning. Reminds me of the experience of trying to be on top of the dependencies in a medium-large JS project.
I am a person that buys into a tool or a process and expects it to be part of the life with no major changes through the years (or as long as the need exists). But AI? You buy into something today, not 2 weeks have passed and there's already a large "update" introduced to the conditions or the optimal usage patterns you should be adopting.
It's tiring. Makes all prices and offers feel so unreliable and gets me a bit more disinterested each time they change.
ricardobeat 2 hours ago [-]
This makes no sense. You want improvements to stop?
These being open, you can keep using the old models indefinitely for as long as there are providers offering them.
j1elo 1 hours ago [-]
No, I'm talking about the whole sector, not specifically about DeepSeek.
Fully knowing that it is a new industry living its own infancy, it is perfectly normal that there is instability and numerous swings on pricing, conditions, or direction.
But it's not less real that such process can produce churn and consumer fatigue.
sebastiennight 3 hours ago [-]
With proprietary labs lowering their prices and Deepseek raising theirs over time, wouldn't it possible to extrapolate a graph to look at where the terminal frontier-model million-token-cost asymptotes to?
dakolli 20 minutes ago [-]
off of one historical price change, no.
mateenah 3 hours ago [-]
This is good for other competitors I guess. People rarely calculate the bump in price but the fact that price is increasing might bring them to other vendors.
poly2it 4 hours ago [-]
That's a hefty increase. Flash pricing during peak is now 1.32/M out, compared to the current 0.28/M, which in turn is a quite a bit above the cheapest provider at 0.16/M.
That sounds like a policy written by someone who doesn't understand how LLM's work...
Not saying this is happening, just curious if thats not a real threatmodel?
I imagine it would be very non trivial to do it in a way that that was reliable and obfuscated enough to prevent detection for any amount of time?
Hopefully this will change soon. But AI and China/US skepticism is very high. Even if the person you talk to isn't skeptic, his boss may be. And even if his boss isn't, his CFO or Legal department may use it as a political lever and therefore if you can say 'everything in europe' you dodge the tension entirely.
Yeah it's dumb.
For European and US customers this is effectively 2x increase. I think i wll keep using both Flash and Pro as before.
EDIT: Misread numbers to believe off-peak kept old prices
Personally, I don't think we've seen the total end of dirt cheap LLMs, it's just a frontier lab doesn't want to be in business of serving half the world.
You certainly don't need Fable to code up a basic web app, any more than you need a Ferrari to go grocery shopping,
They are hoarding HW at massive scale, they make it harder and more expensive to own
Just because you are fine with the new price doesn't mean it's not a problem
Perhaps it's time to pop this bubble
Though of course you're talking about data centers, and romanticizing them rather than the AI itself.
The summary of this paper describes my sentiment in better words than I have:
https://www.nature.com/articles/s41599-025-05868-8
It’s very easy for the average person to mistake linguistic ability and simulated problem solving for intelligence and sentience.
Already outdated though I think, as GLM 5.3 is latest now :)
It’s a race to the bottom, and the bottom is unlimited use for a flat monthly rate.
Granular pricing (tokens, minutes, etc) is pretty anti-customer generates less revenue than customer value-based subscriptions (why SaaS is such a good business model)
Consumers don’t generally get usage-based pricing because of the inconvenience and unpredictability, but B2B SaaS products utilize usage-based pricing all the time.
that's a good outcome - it means they're fungible, and easily available.
The same motivation of course - the GPUs have a finite service lifetime, and to maximize revenue you need to keep them busy 24x7.
https://news.ycombinator.com/item?id=49287881
https://news.ycombinator.com/item?id=49285160
https://www.bloomberg.com/news/articles/2026-06-17/microsoft...
Bytedance which runs China’s most popular Doubao AI chatbot; is spending $70B in CapEx this year, most of it outside of China (Malaysia, Thailand, Brazil, etc; and they are allowed to lease NVIDIA chips). This is roughly 50% of Microsoft CapEx.
https://www.tomshardware.com/pc-components/gpus/chinas-byted...
I am a person that buys into a tool or a process and expects it to be part of the life with no major changes through the years (or as long as the need exists). But AI? You buy into something today, not 2 weeks have passed and there's already a large "update" introduced to the conditions or the optimal usage patterns you should be adopting.
It's tiring. Makes all prices and offers feel so unreliable and gets me a bit more disinterested each time they change.
These being open, you can keep using the old models indefinitely for as long as there are providers offering them.
Fully knowing that it is a new industry living its own infancy, it is perfectly normal that there is instability and numerous swings on pricing, conditions, or direction.
But it's not less real that such process can produce churn and consumer fatigue.
https://openrouter.ai/deepseek/deepseek-v4-flash-0731#provid...
The big Covid migrations (startup prople migrating to the countryside),
Will we see the big AI migrations (people travelling to where AI is the cheapest)?
DeepSeek-V4-Flash (off-peak, x2 for peak)
* Cache Hit $0.007 (x2.5)
* Cache Miss $0.22 (x1.5)
* Output $0.66 (x2.25)
DeepSeek-V4-Pro (off-peak, x2 for peak)
* Cache Hit $0.022 (x6)
* Cache Miss $0.66 (x1.5)
* Output $1.98 (x2.25)
Peak Hours: 01:00–04:00 and 06:00–10:00 UTC
Effective from: 16:00, August 16, 2026 (UTC)
It’s still cheaper than everybody else.
https://openrouter.ai/deepseek/deepseek-v4-flash#providers
- but what matter - is cache hit
even now deepseek's off-peak hours for cache hit (0.007) is lower than other providers (~0.01)