Back to News
Advertisement
Advertisement

⚡ Community Insights

Discussion Sentiment

44% Positive

Analyzed from 836 words in the discussion.

Trending Topics

#cpu#limits#pods#using#requests#scheduler#pod#don#more#core

Discussion (12 Comments)Read Original on HackerNews

conradludgate•about 2 hours ago
While I agree that CPU limits tend to make your performance worse, I don't think the delivery of the post is all too convincing (and is pretty heavy on the LLM-isms that it's putting me off from reading).

It mentions that a cpu request is a guarantee, but how is that enforced? If I have 32 pods running on a 32 core machine, each with 1cpu requested, what stops one of those pods using an unfair share? I assume we just rely on the Linux scheduler. If I have 16 pods with 1cpu and 1 pod with 16cpu, does the Linux scheduler make sure to give the 16cpu pod more time? Or are we back to using cgroups.

croemer•about 2 hours ago
Requests are guaranteed no matter what.

> If I have 16 pods with 1cpu and 1 pod with 16cpu, does the Linux scheduler make sure to give the 16cpu pod more time?

Yes if there's CPU pressure the 16cpu will get 16x more than 1cpu.

Sayrus•about 2 hours ago
Each pod has weight. These are written to your cgroup (cpu.weight in CGroups V2). CFS is scheduling based on these weights.

There's a part about it: https://github.com/inevolin/k8s-cpu-limits-analyzed#3-withou...

toast0•about 2 hours ago
> If I have 32 pods running on a 32 core machine, each with 1cpu requested, what stops one of those pods using an unfair share?

I don't use all this fancy stuff, but wouldn't you set that up as each of the 32 pods is limited to running on a specific core? No sense letting them each run on all cores, because all of the X per core will get too big.

tbrownaw•about 2 hours ago
Limits are what give consistency when your pod gets scheduled on nodes with different amounts of load.
acuteaura•about 2 hours ago
Some workloads will also consume all the resources you hand them without being latency sensitive at all.

I've handled outages of CPU time available getting suddenly compressed (we were running pod priorities with staging/prod on one cluster and up to 70% spots in 2019) and then learning that some very important applications outgrew their original requests, gone unnoticed because limits were removed a year or so prior. You can fix this with monitoring/right-sizing tools, but that requires your org to not be dysfunctional, and my style of platform engineering usually has to account for the org being very dysfunctional.

aairey•about 2 hours ago
How is this “news”?

It was already the case in 2018.

Also, no mention of the scheduler overhead. And the maintenance overhead is the worst.

websap•about 2 hours ago
Scheduler overhead for CPU limits? Are you talking about the Linux Scheduler or the k8s scheduler?
acuteaura•about 2 hours ago
It's not as much an "overhead" as it will mess up your latency. Limits don't stop you from using available resources until you hit the relative allowance in a CFS window, so a 1 CPU limit on a 32 CPU machine at worst gives you 32 cores for 3.3ms every 100ms.
denysvitali•about 1 hour ago
Looks like a bad copy of: https://home.robusta.dev/blog/stop-using-cpu-limits

"For the Love of God, Stop Using CPU Limits on Kubernetes" literally same title

ofjcihen•about 2 hours ago
Many orgs are just now discovering K8s at scale. Seriously.

I’m not sure why that is but a large number of the F100s I contract with are suddenly deploying 4 times the number of containers they had before.

stanac•about 2 hours ago
Could it be that they're using coding agents to develop applications they previously wouldn't spent time on developing? Like internal tools, or experimental builds.
websap•about 2 hours ago
For the love of god - care about other pods on the node, especially in a multi-tenant setup.

Sorry for the cheeky response.

CPU Limits have a place, you don't want a bad change for 1 deployment object affect all neighbors by taking all the CPU. You need to be able to constrain the blast radius. This doc gives me strong AI vibes. Setting CPU limits isn't free. You still need to care about how the programming language that you use discovers those limits, and correctly handles them. For e.g. if you spin up a 100 Java threads, but only have 1 cpu as the limit, that's bad design.

mystifyingpoi•about 1 hour ago
Exactly on point. Shit happens, performance bugs appear, someone messes up Kafka config and it starts consuming from the beginning of the world, etc. Limiting CPU is a must. I could see maybe if someone has a super good monitoring + oncall response team, then letting things go loose for a bit is a lesser evil than working out limits, but still.
inigyou•about 2 hours ago
That's what CPU requests are for.
websap•about 2 hours ago
CPU requests are cgroup weights.
ksbd-pls-finish•about 1 hour ago
But they also affect scheduling, right? If you set them too high, you will waste resources.
crymer11•about 2 hours ago
I don’t think you understand how CPU limits and the Linux CPU scheduler work. CPU limits don’t protect you from something taking all the CPU; that’s what CPU requests do. Limits throttle your pods even if the CPU is idle/free to do work.
ozyschmozy•about 2 hours ago
From the docs another comment linked:

> The CPU request typically defines a weighting. If several different containers (cgroups) want to run on a contended system, workloads with larger CPU requests are allocated more CPU time than workloads with small requests.

So I guess limits _would_ protect other pods to a degree. Though I agree that it doesn't seem worth the tradeoff of your pod getting constantly interrupted while the rest of the box is sitting idle

da_chicken•about 2 hours ago
That's literally not what the documentation says.

https://kubernetes.io/docs/concepts/configuration/manage-res...

acuteaura•about 2 hours ago
Where do you think the contradiction is? A CPU limit of 1 does not prevent an application from using 8 cores worth of compute at once, it just limits the compute to 1 core per 100ms slice on average, and usually that means getting unscheduled for 7/8 of the slice, if the app is using all available resources and nothing else contests them (weighted by requests).
websap•about 2 hours ago
Hahaha! Thanks for the laugh.
rimworld•about 2 hours ago
tbh I've never known search a 'feature' as limits and requests coupled with health probes to cause more problems in production than anything else.....
callamdelaney•about 2 hours ago
Hilarious, another kubernetes footgun - the gift that keeps on giving.
johanj•about 2 hours ago
One of my former colleagues wrote this on how Uber approached the same CPU-quota throttling problem, but with dedicated CPUs as the solution: https://www.uber.com/dk/en/blog/avoiding-cpu-throttling-in-a...

Essentially, they avoided CFS quota throttling by assigning exclusive CPU cores via cpusets. That sacrifices some burstability and packing efficiency in exchange for stronger, more predictable CPU isolation.