Overly idiomatic but this was for fun, not production :P–
use XML::LibXML;
my $doc = XML::LibXML->load_html( location => "example.html",
{ recover => 1 } );
my @ids2text = map { [ $_->value, $_->getOwnerElement->textContent ] }
$doc->findnodes('//@id');
$_->[1] =~ s/\W+//g for @ids2text;
print join ", ", map sprintf("%s=%s", @$_), @ids2text;
While this happens to be XHTML
Sidenote on that. I am sure you know the sample is not XHTML but I thought I'd call it out for the sake of readers.
Update: I missed the "transitional" part of the XHTML declaration. It is indeed, shockingly, valid transitional XHTML. Goes to show how on point haukex is on this matter.
Update 2: updated node title per LanX. Pulled strict/warnings to shorten post. Plus link to module: XML::LibXML
-
Are you posting in the right place? Check out Where do I post X? to know for sure.
-
Posts may use any of the Perl Monks Approved HTML tags. Currently these include the following:
<code> <a> <b> <big>
<blockquote> <br /> <dd>
<dl> <dt> <em> <font>
<h1> <h2> <h3> <h4>
<h5> <h6> <hr /> <i>
<li> <nbsp> <ol> <p>
<small> <strike> <strong>
<sub> <sup> <table>
<td> <th> <tr> <tt>
<u> <ul>
-
Snippets of code should be wrapped in
<code> tags not
<pre> tags. In fact, <pre>
tags should generally be avoided. If they must
be used, extreme care should be
taken to ensure that their contents do not
have long lines (<70 chars), in order to prevent
horizontal scrolling (and possible janitor
intervention).
-
Want more info? How to link
or How to display code and escape characters
are good places to start.
|