Back to News
Advertisement
Advertisement

⚡ Community Insights

Discussion Sentiment

100% Positive

Analyzed from 329 words in the discussion.

Trending Topics

#html#regex#parse#finite#length#stackoverflow#parsing#documents#answer#why

Discussion (18 Comments)Read Original on HackerNews

danlittabout 2 hours ago
You can't parse [X]HTML with regex. Because HTML can't be parsed by regex. Regex is not a tool that can be used to correctly parse HTML.
layer8about 1 hour ago
Some commenters are missing that this is a reference to https://stackoverflow.com/a/1732454.
chrisandchris20 minutes ago
> Locked. There are disputes about this answer’s content being resolved at this time [sic]

And also

> Nov, 2020

And then, StackOverflow asks itself why it looses users.

magicalist10 minutes ago
It's a 17 year old joke, maybe one of the most famous on stackoverflow for the aging engineers out there. No one wants new "funny" edits on it.
layer818 minutes ago
The answer is actually from 2009. As long as HTML/XML doesn’t suddenly become a regular language, I think it’s pretty timeless.
pwdisswordfishq42 minutes ago
Good ol' appeal to emotion devoid of technical argument.
tyhoabout 1 hour ago
You can absolutely parse HTML with regex, so long as the document is finite in length. Every finite language is regular, hence can be parsed with regex's.
magicalist12 minutes ago
Zalgo and situational subset parsing aside:

> You can absolutely parse HTML with regex, so long as the document is finite in length

This isn't sufficient, unless I'm misinterpreting what you're saying. It's not enough to have documents of finite length (all documents are finite in length), you need documents with a max length, so you have a finite number of possible documents to parse.

gpvosabout 1 hour ago
Do a web search for the parent of your comment, read the Stackoverflow answer. It's a classic. Learn about Zalgo and Tony the pony, he comes.
rokkamokkaabout 2 hours ago
You can parse a subset of it though, like if you're in control of the html yourself and avoid certain structures
zarzavatabout 2 hours ago
Parsing HTML with a regex is never a good option, but it's sometimes the only option.
pkalabout 2 hours ago
In the example from the article it certainly is an option. In Python you could either use a "soup" library or you could play around with a tool like https://www.w3.org/Tools/HTML-XML-utils/man1/hxpipe.html.

The more fundamental question for me is why the author didn't decide to make make code blocks non-breaking by default, or just add the class annotations when he writes the HTML?

prmoustache43 minutes ago
I do not count substitution as "parsing"
timedudeabout 1 hour ago
What do you mean you can't. I do it all the time
throw123456789141 minutes ago
You do some parts all the time.
matheusmoreiraabout 2 hours ago
Tony the pony, he comes.
pwdisswordfishq44 minutes ago
Why even bother with the CSS class? Just apply text-wrap: nowrap; to all <code> elements and dispense with the fragile regex parsing.
pimlottc28 minutes ago
Unfortunately this makes it much harder to read on narrow-width screens (e.g. mobile), where the use must scroll horizontally