Re^4: Wierd behaviour with HTML::Entities::decode

in reply to Re^3: Wierd behaviour with HTML::Entities::decode_entities()
in thread Wierd behaviour with HTML::Entities::decode_entities()

This is very particular to the ordering I mentioned earlier (first decimal, then hexadecimal, then named entities are expanded)

No, it isn't

Different nesting order:

>perl -MHTML::Entities -le"print decode_entities '&#x26;amp;quot;';
&amp;quot;

>perl -MHTML::Entities -le"print decode_entities '&amp;#x26;quot;';
&#x26;quot;
[download]

Different sibling order:

[ Can't find a valid example ]

Or I still don't understand. Please given an example where ordering matters.

Comment on Re^4: Wierd behaviour with HTML::Entities::decode_entities() Download Code

Replies are listed 'Best First'.
Re^5: Wierd behaviour with HTML::Entities::decode_entities() by JadeNB (Chaplain) on Dec 14, 2009 at 17:02 UTC
The source for `decode_entities_old`, which I thought was just a pure-Perl version of `decode_entities`, looks roughly like this (I've stripped out some context dependence): `sub decode_entities_old { my $array = [ @_ ]; my $c; for (@$array) { s/(&\#(\d+);?)/$2 < 256 ? chr($2) : $1/eg; s/(&\#[xX]([0-9a-fA-F]+);?)/$c = hex($2); $c < 256 ? chr($c) : $1/ +eg; s/(&(\w+);?)/$entity2char{$2} \|\| $1/eg; } return @$array; }` [download] With this code, and the under-populated hash `my %entity2char = ( amp => '&', quot => '"' )`, we have `say decode_entities_old '&amp;quot;'; => " say decode_entities_old '&#x26;quot;'; => &quot;` [download] We would get (essentially) the opposite behaviour if we switched the second and third substitutions; that's what I mean by “This is very particular to the ordering”. Of course, your example shows that `decode_entities` (which I guess is implemented in XS—I couldn't find it in the source) doesn't exhibit this buggy behaviour, so I guess that there's a reason that the sub I quoted has `_old` postpended. :-) UPDATE: I just noticed that I'd mangled my intended bug-inducer, `&quot;`, above. I wonder if `decode_entities` (not `decode_entities_old`) handles it correctly?	[reply] [d/l] [select]
Re^6: Wierd behaviour with HTML::Entities::decode_entities() by ikegami (Patriarch) on Dec 14, 2009 at 18:05 UTC
If HTML::Entities were to decode `&amp;quot;` to `"`, it would be buggy. I did understand you correctly. Does what I said make more sense now? I wonder if decode_entities (not decode_entities_old) handles it correctly? I don't know if that's legal in SGML/HTML — unescaped ampersand — but yes. `$ perl -MHTML::Entities -le'print decode_entities "&quot;"' "` [download]	[reply] [d/l] [select]
Re^7: Wierd behaviour with HTML::Entities::decode_entities() by Baz (Friar) on Dec 14, 2009 at 19:05 UTC
You can test with this - `<!DOCTYPE html PUBLIC "-//W3C//DTD XHTML 1.0 Strict//EN" "http://www.w +3.org/TR/xhtml1/DTD/xhtml1-strict.dtd"> <html xmlns="http://www.w3.org/1999/xhtml" xml:lang="sv" lang="sv"> <head> <meta http-equiv="Content-Type" content="text/html;charset=iso-885 +9-1"/> </head> <body> Mic i vÂr replokal &quot;The Dungeon&quot; </body> </html>` [download] Now it very clear what the problem is...	[reply] [d/l]
Re^8: Wierd behaviour with HTML::Entities::decode_entities() by ikegami (Patriarch) on Dec 14, 2009 at 19:53 UTC

In Section Seekers of Perl Wisdom