Back to News
Advertisement
Advertisement

⚡ Community Insights

Discussion Sentiment

45% Positive

Analyzed from 2897 words in the discussion.

Trending Topics

#loop#function#infinite#compiler#code#std#why#call#undefined#should

Discussion (84 Comments)Read Original on HackerNews

JoshTriplettabout 1 hour ago
> When both conditions are met, the loop body is replaced with a call to std::this_thread::yield().

Insert screaming here.

An infinite loop, with no library calls whatsoever, gets a system call inserted. That's a horrible surprise waiting to happen.

The entire concept of the "forward progress guarantee" is broken. An infinite loop should compile to an infinite loop. Nothing more, nothing less.

ozgrakkurt41 minutes ago
But the compilers have to optimize the crap code in big tech codebases by 0.5%, it saves a lot of money.

Also performance doesn't matter that much and developer time is more important btw, keep using react.

muvlon6 minutes ago
It's not even about optimizing some big tech codebase by 0.5%. The progress guarantees in particular are in place s.t. Nvidia can choose a certain implementation strategy in Cuda C++ that has "surprising" consequences for users (one thread getting stuck in an infinite loop that never yields can livelock its entire warp) but still get to claim "full C++ standards compliance".
saghm12 minutes ago
I guess given that it was UB before, the compiler was already allowed to put a system call here if it wanted for some reason
ibobev23 minutes ago
> An infinite loop should compile to an infinite loop.

I think that a compiler option should control this. It can be a nice optimization, but the programmer should be able to opt out.

ameliaquiningabout 1 hour ago
I'm curious, what exactly do you imagine going wrong here?
rfgplk13 minutes ago
The language is already littered with these "the compiler shall insert" and then a reference to the STANDARD LIBRARY FEATURE N.X. Which means if you're compiling in a freestanding environment half the time you'll get linker errors such as "couldn't find symbol whatever". And what's worse the compiler inserts a call to a function that is LITERALLY STD NAMESPACED. Meaning you have to provide that signature yourself. See how std vector is hardcoded into compare/meta and I can't remember what else.

This then forces developers to create undefined behaviour because according to the standard you can't namespace std your own functions even though it's required to get it to work.

rcxdudeabout 1 hour ago
The biggest headache will probably be it getting emitted in inappropriate contexts: where there is no actual means to sched_yield for whatever reason (bare metal, kernel, whatever). The second is just that the behaviour of the infinite loop changes: suddenly you're getting a bunch of extra system calls from your spinning thread instead of just a high CPU usage, which could disguise the issue or perhaps cause problems for other parts of the system. I don't see a good reason for the transformation: pretty much any time you are writing a bare infinite loop like this you don't want anything else to happen (it's also silly that it only happens with a particular spelling of an infinite loop, keeping the others still undefined).
JoshTriplettabout 1 hour ago
"Emitted in inappropriate contexts" is very much one of the shapes I would expect unpleasant surprises to take, yeah. If you're writing code in C, you often need a lot of control over exactly what's happening. You might, for instance, be writing a .so for use with LD_PRELOAD, where it's important that you know everything being called so you can't accidentally recurse. You might be writing code for a sandbox, where you have an allowlist of permitted syscalls.
rcxdudeabout 1 hour ago
Yeah, this is almost the worst way they could choose to 'fix' the problem.
wahernabout 1 hour ago
> When both conditions are met, the loop body is replaced with a call to std::this_thread::yield(). This gives execution of the loop the forward-progress semantics it previously lacked.

That's the epitome of the hidden code downside that Linus and many others dislike about C++. For constructors and destructors it's somewhat unavoidable and not so random, though Rust does better at limiting the blast radius of non-local code, at least in the drop case.

If they didn't want to adopt the C11 rule, the C++ committee should've explored a rule that required the compiler to emit a diagnostic or error for trivial loops (whether as defined by C11 or otherwise), requiring the programmer to explicitly insert ::yield or similar. No hidden code, and less opportunity for the compiler to do surprising things.

The C committee has been rigorously enumerating UB cases in the standard and addressing each case in turn, often by requiring a diagnostic, error, or by turning it into implemention defined behavior. But inserting code like that would be unthinkable.

rfgplk7 minutes ago
It's catastrophic actually. Like disastrously catastrophic. It started with C++20 mostly, and has only kept getting worse from then. See zero initializing variables by default (WHY?) compare/meta including half the STL and HARDCODING those symbols, std::initializer_list being in the std namespace (if you don't include <initializer_list> you literally can't use it, and there is no such thing as a __initializer_list or some internal symbol), the entire coroutine library where you MUST provide coroutine_handle, noop_coroutine, suspends et al (coroutines aren't that bad because they're not necessarily spaghetti).

<meta> is the single WORST OFFENDER, where they hardcode std::vector (literally std::vector in the std namespace) std::ranges std::allocator.

IsTomabout 1 hour ago
> should've explored a rule that required the compiler to emit a diagnostic or error for trivial loops (whether as defined by C11 or otherwise), requiring the programmer to explicitly insert ::yield or similar

It wouldn't work when this kind of loop is generated by macros/templates in some unreachable case left after const folding.

rcxdudeabout 1 hour ago
If it's truly unreachable then it's not likely to be a problem. If it is reachable and it's emerging from some macros and templates then I would be more inclined want a warning for it.
IsTom40 minutes ago
Yeah, but then you need compiler to somehow know if it's truly unreachable to know when to emit the warning and when to not do that.
Aurornisabout 2 hours ago
I never would have guessed that the unreachable() function would get executed in that example. Probably not something you’d encounter in practice, though I have seen some weird things happen with layers of #ifdef
saghm11 minutes ago
That's kind of what you get with UB; the compiler doesn't need to do what you expect.
omoikaneabout 1 hour ago
> The loop must be a trivially empty iteration statement -- meaning its body is literally empty

This seems to say that the loop body can not be "continue". Indeed, I just tried -std=c++26 with ";" and got an infinite loop as promised, but "continue" restores the undefined behavior:

- "while(true);" -> https://godbolt.org/z/T65o51crx

- "while(true) continue;" -> https://godbolt.org/z/Pj9raEcnP

This is unfortunate since I know of one style guide that prefers "continue" over single semicolons. I guess all those code will be doing "while(true) {}" from now on.

https://google.github.io/styleguide/cppguide.html#Formatting...

ameliaquiningabout 1 hour ago
The article, most unfortunately, doesn't explain why anyone would want infinite loops to be UB in the first place. I found this explanation: https://www.open-std.org/jtc1/sc22/wg14/www/docs/n1528.htm
omoikaneabout 1 hour ago
The article mentions it's a halt-on-error pattern:

https://www.sandordargo.com/blog/2026/09/16/cpp26-trivial-in...

Edit: sorry, missed the UB bit.

Lvl999Noobabout 1 hour ago
That says why they don't want it to be UB. The question, I believe, was why they want statically-known-infinite non-trivial loops to continue being UB.
rcxdudeabout 1 hour ago
That's more why you would want them to be defined in the first place.
peterusabout 1 hour ago
There are valid use cases for the infinite while(1) loop in microcontroller programming (contrary to popular belief it seems). Autogenerated HAL code for the stm32 uses it for error handlers, and they support C++ so I am surprised this was UB.

I only use it for error handling and of course it is a bad idea to use this to wait/stall in power sensitive applications, in that case use wake from interrupt.

As an aside, I like to include a software breakpoint in my error handlers. It makes debugging easier without wasting a hardware breakpoint (which are physically limited by the microcontroller):

  __BKPT();
  while (1)
    ;
rfgplk4 minutes ago
UB according to the standard committee is "we didn't think of it". It's not literal UB it's well known what it compiles down to, every time. (.loop: jmp .loop)
MiroslavPokornyabout 15 hours ago
Breadcrumbs for "blog", "year", "month" etc are broken and give 404s :(

One can browse other blog entries so it really doesnt matter too much.

aabolfazlabout 1 hour ago
Sometimes while (true){} doesn't mean anything clever. It just means the system is broken stay here.
adzmabout 2 hours ago
i have never before thought that a function could 'fall through' to another function. why does this behavior even exist?
kzrdudeabout 2 hours ago
Well you leave the C++ realm (execution model), as you should with UB and it depends on implementation. The implementation of the compiler was such that the two functions are placed after each other in the machine code; and if the first function doesn't return, then you continue executing into the code for the next function.
Someoneabout 1 hour ago
But the compiler assumes the function will make forward progress. If the function does that, it will return, so why doesn’t the compiler emit a function epilogue?
muvlon3 minutes ago
The compiler can assume that the function will return, but it can also statically deduce that the function cannot return. That's a contradiction, so the compiler deduces that the function is simply UB when called, i.e. no need to emit an epilogue. It's the logical principle of explosion in compiler format, basically.
raverbashingabout 1 hour ago
This makes no sense to me

If I think about asm:

function1:

    (do stuff)

    jp function1

    ret

function2:

    (other stuff)

    ret

main:

    call function1

    call function2

the 2nd call might happen internally due to branch prediction but in practice it shouldn't and the processor fixes this

Oh yeah and TFA also goes with:

> The funny bit is that C got this right.(...) but C included one more rule: loops whose controlling expression is a constant expression may not be assumed to terminate.

Well, duh! A broken clock is right twice a day it seems

rcxdudeabout 1 hour ago
With UB the compiler has no particular requirement to emit the 'ret'. (or, in the example, anything at all for the function)
apple1417about 1 hour ago
The assembly gives a bit of a hint as to what's happening.

    main:
    
    unreachable():
            push    rbx
            ...
Due to the undefined behavior, it decides calling main must be impossible, so the easiest thing to do is just give up, don't bother defining the rest of it. You can also do the same with std::unreachable(). But the label for the function still sticks around for some reason, so when you jump to it, it falls through. Which leads to the really stupid fact that reordering the functions changes the behavior.

I assume there are good reasons they can't just completely delete the label. Maybe it would screw linking, or with cases where you deliberately have multiple labels for the same function. And if the effect is only visible due to undefined behavior, it's not technically wrong. But I have always thought this is such a stupid case, surely it can't be that complex to add a trap instruction, even in an optimized build you shouldn't really care if it slows down a function that's "never called".

rcxdudeabout 1 hour ago
I suspect it's more a chain of: emitting the ret is unnecessary because the infinite loop will never return -> emitting the infinite loop is unnecessary because there's no side effects within it and it's undefined behaviour -> emitting any setup for the function is necessary because it's doing nothing else (all probably decisions from different stages of the compiler).
rcxdudeabout 1 hour ago
The CPU doesn't really see functions, it just sees instructions. Functions are a convention on top of the machine code. What happens in this case is the compiler emits essentially a malformed function: it ends without performing a return, so execution just continues into the next function in memory. You can get the same behaviour by missing a 'return' statement from a function that needs one (though in that case I've also seen kind of the opposite: the function returns into the function two slots up in the stack, essentially returning from the function that called it! Undefined behaviour can utterly destroy normal control flow).

Probably the process was one optimization pass saw that the function will never return due to an infinite loop, and removed the function return from the IR of the function, then a later pass saw that the infinite loop was a no-op and undefined so removed that as well, leaving a function that basically did nothing, not even return.

echoangle32 minutes ago
> The CPU doesn't really see functions, it just sees instructions. Functions are a convention on top of the machine code.

Not really true, most instructions set have instructions specifically to implement functions as found in normal programming languages. x86 has CALL and RET for example.

https://en.wikipedia.org/wiki/X86_calling_conventions

Of course the compiler can stil optimize by inlining etc., but functions still mostly exist at the assembly level.

rcxdude21 minutes ago
they have instructions for implementing them, but the important point here is that functions are still only defined by instructions that are executing between a call and ret instruction (or their equivalent more spelled-out equivalent operations), and not only can these not match up with what the compiler considers a function (for useful reasons like tail-calls as well as not-useful reasons like compiler bugs and UB), it might not be statically obvious exactly what instructions these are. So the CPU in practice has only a rough guess of where the function boundaries are (it might use these guesses for things like branch prediction, but they don't define the visible execution of the code beyond the nuts and bolts of what those instructions actually do).
echoangleabout 1 hour ago
I'm also confused that an uncalled function is even compiled and linked, wouldn't it make sense to remove it entirely if the compiler can detect that it's never called?
rcxdudeabout 1 hour ago
If it's declared as static, maybe (well, usually, in my experience. You'll also usually get an unused warning). Otherwise the compiler can't assume some other compilation unit won't want it. Linkers can perform a garbage collection pass but they don't often do it by default and they often need finer grained information from the compiler (see the gcc arguments --ffunction-sections and -Wl,--gc-sections)
lou1306about 1 hour ago
I can understand adding the 'unreachable' function to the object file, I can even understand plugging it into the final executable, what I (and most other people) object to is making it the de-facto entry point.

This is literally the opposite behaviour compared to what is written in the source code, even when you "assume the infinite loop terminates".

account42about 2 hours ago
Unfortunate. There isn't ever a good reason to have an infinite loop so concerned compilers could have just diagnosed this as a warning.
echoangleabout 1 hour ago
The article mentions a use case for that:

> What I found is that this is common in embedded and kernel code as a halt-on-error pattern. When a fatal error occurs and there’s no operating system to exit to, you simply stop:

pdonisabout 1 hour ago
If this is a genuine use case, I wonder why the language can't just introduce a built-in function for it. For example, std::get_stuck_here(). Then the compiler would know not to optimize this away. The implementation under the hood could still be an infinite loop, but the compiler would not have to guess why it's there.
account42about 1 hour ago
Low level code can and should use assembly to get the precise effect they desire in these cases.
echoangleabout 1 hour ago
Why not just allow infinite loops instead of having me write assembly for it though?
mdspanabout 1 hour ago
That would be pretty cumbersome though. If you're targeting N different architectures, you would have to write N different assembly blocks.
rcxdudeabout 1 hour ago
I shouldn't need to drop to assembly to get an infinite loop that works!
sumtechguyabout 1 hour ago
> There isn't ever a good reason to have an infinite loop

That seems to be a very broad statement. For example in a system where interrupts mostly control things this sort of 'do not close the program' could be useful.

A guy I worked with had one I never would think of because I do not work in that field.

But yeah a warning would probably be useful.

weinzierlabout 1 hour ago
For Rust the infinite loop is important enough to have its own keyword.
mdspanabout 1 hour ago
Compilers can still diagnose something as a warning even if it's not UB.
not_the_fdaabout 1 hour ago
Interrupt driven super loops are very common on bare metal systems.
shmerlabout 1 hour ago
Why does the loop mean halt in that embedded case example?
rcxdudeabout 1 hour ago
It just spins the CPU in the loop, stopping execution from progressing. Technically, whether this fully halts the system depends on what else is going on: you might need to fully disable interrupts before entering the loop to get a full halt. OTOH you can design your system so that everything happens in interrupts (with modern interrupt controllers the common wisdom of doing as little as possible in interrupts no longer applies and it can be a good way to get a predictable and low-latency system) and so you finish your setup code with an infinite loop to stop the CPU running off the end of your function when it's not executing one of the interrupts.

In a lot of cases, you might insert some 'wait-for-interrupt' type instruction in the loop that halts the CPU more 'cleanly' (and in a lower power mode), and usually this will appear as a side-effect and keep the behaviour defined. But this is not always desirable or possible.

Advertisement
oleganzaabout 2 hours ago
Why is null-terminated C string considered a "billion dollar mistake", but UB isn't?
Guvanteabout 2 hours ago
Null terminated strings were an intentional compromise, known to be inferior for execution but superior for memory

Null being an "allowed" value for pointers is the mistake e.g. what became nullptr. "Allowed" because garbage values are garbage.

account42about 1 hour ago
Think of the alternative where we'd be dealing with endless issues because someone though 255 or 2^16-1 characters ought to be enough for everyone.
DonaldPShimodaabout 2 hours ago
The "billion-dollar mistake" was about implicitly nullable values, i.e., allowing a variable with type `T` to also be set to `null`, not null-terminated strings.

Anyway, one argument is that UB is fundamentally useful in languages that are insufficiently type-safe, like C and C++. The "holes" in the specification allow for regions where the compiler can optimize the code in ways you may not expect.

As we have developed more advanced type systems, the utility of undefined behavior has lessened considerably.

returningfory2about 1 hour ago
Agreed that this is why a lot of people support the current UB situation, but the history of UB makes this feel wrong:

> As far as I can tell, C89 did not use performance as a justification for any of its undefined behaviors. They were non-portabilities, like signed overflow and null pointer dereferences, or they were outright bugs, like use-after-free. But now experts like Chris Lattner and Hans Boehm point to optimization potential, not portability, as justification for undefined behaviors. I conclude that the rationales really have shifted from the mid-1980s to today: an idea that meant to capture non-portability has been preserved for performance, trumping concerns like correctness and debuggability.

https://research.swtch.com/ub

ameliaquiningabout 1 hour ago
I recommend this explanation of why UB is good and necessary (but C and C++ are doing it wrong, defining some things as UB that really shouldn't be): https://www.ralfj.de/blog/2021/11/18/ub-good-idea.html
nicoburnsabout 1 hour ago
Probably because null-terminated strings are completely avoidable, whereas some amount of UB is all but required for performance (albeit C and C++ have far too much).
adzmabout 2 hours ago
as an aside, i've always preferred the zoidberg for (;;) to while(true)
glouwbugabout 1 hour ago
I think you're thinking of (;,,;)
BobbyTables2about 2 hours ago
TLDR: For almost 1/6 of a century, the C++ standards broke the simplest infinite loop and only just recently fixed it.

Idiots!

Don’t they really that people write real programs to solve real problems? This isn’t a theoretical academic exercise!

Sharlinabout 2 hours ago
The argument is that an infinite loop without side effects isn't a real program. It's not useful for anything except wasting cycles.
nh2about 1 hour ago
Of course the infinite loop should run as expected.

It breaks the most fundamental debugging expectations (such as "delete code until problem disappears") if the fundamental, minimal building blocks of a language, when on their own, do random rubbish.

To understand a program that does something, better first understand a program that does nothing.

As a fan of sensible analogies:

You put a salad bowl with vinegar into the fridge and notice that when you do that, the fridge stinks afterwards. You try again without the vinegar, then without the salad. In C++ world, upon receiving the empty bowl, the fridge detonates ("it is not useful"), blowing up your house. That is not OK.

Jaxanabout 1 hour ago
But if you program a for loop computing the sum from 1 to n, this also gets replaced by a constant (unless you build in debug mode). Why would an empty loop be different?
kibwenabout 1 hour ago
And unfortunately that argument would be incorrect, because not only is there a realistic chance of hitting this on embedded systems, the fact that LLVM baked this into its low-level semantics resulted in miscompilations in Rust for a time, where `loop {}` is a valid way to implement a diverging function: https://github.com/rust-lang/rust/issues/28728
lou1306about 1 hour ago
Yeah the argument here is clear, also rather silly. Either you must accept that your language allows for completely useless computation, or, if the compiler is so good at detecting "unreal programs" it should also refuse to compile them.
bryanlarsenabout 2 hours ago
They also realized that people choose compilers based on performance benchmarks, and that insane optimizations let them win.
vlovich123about 2 hours ago
Until Rust proved actually you can get really good or better performance if the language itself is better. I really don’t know how C++ digs itself out of the UB hole it has dug.
bryanlarsenabout 2 hours ago
Probably by working together with Rust. Eliminating undefined behavior from unsafe Rust is a big deal for the Rust community at the moment. And given that most unsafe rust code exists to call into C or C++, concepts like pointer provenance need to be extended. And proper pointer provenance guarantees can both decrease UB and increase optimization potential.

IIUC, my understanding is shallow.

bluGillabout 2 hours ago
You are an idiot if you write an infinite loop. An infinite loop is a waste of CPU cycles and energy when run.

If it wasn't so hard to detect (the trivial cases are easy, but it gets hard quickly) I'd say the program should fail to compile.

echoangleabout 1 hour ago
And how would you generate assembly to keep a microcontroller idle then?
bluGillabout 1 hour ago
You call the CPU halt instruction.