Beefy Boxes and Bandwidth Generously Provided by pair Networks
Perl-Sensitive Sunglasses

comment on

( #3333=superdoc: print w/replies, xml ) Need Help??
I like the code from roboticus. A slight reformulation is below.

You can capture all of the names from the "match global" in one expression.

For most of my "web scraping" code, very fast performance is not that important. Neither is being super general purpose. If I can write a short regex in 5 minutes that gets me what I want, then I go with it and if the web page changes in a year, then I write another 5 minute regex.

What makes sense in your application has to do with what you are "scraping", how often the page format changes, what the impact of that will be (maybe boss calling you at midnight - or just some thing that you have to get "done this week"). Mileage varies.

There are some very fine HTML parsing modules and they can be used to make much more general solutions. However, I often write one regex to get to a hunk of html that has what I want and then write another regex like below to extract what I want from that hunk of stuff. Write as much code as you need, but don't write more than you have to. And no matter what you do, this HTML stuff is a very "fragile" interface - meaning that your code will break at the whim of the web developer.

#!/usr/bin/perl -w use strict; $/=undef; my $data =<DATA>; $data =~ tr/\n/ /; #turn \n's into spaces my (@names) = $data =~ m/<a[^>]*>(.*?)<\/a>/g; foreach (@names) { print "$_\n"; } =prints Jon.Martinez Mary Jones Rob Oticus Joe Blow =cut __DATA__ <a href="foo">Jon.Martinez</a><li>gabba, gabba, hey!</li><a href=bar>Mary Jones</a><p>Gazebo!</p><a href="baz">Rob Oticus</a><a>Joe Blow</a>

In reply to Re: A regex question by Marshall
in thread A regex question by emelianenko

Use:  <p> text here (a paragraph) </p>
and:  <code> code here </code>
to format your post; it's "PerlMonks-approved HTML":

  • Are you posting in the right place? Check out Where do I post X? to know for sure.
  • Posts may use any of the Perl Monks Approved HTML tags. Currently these include the following:
    <code> <a> <b> <big> <blockquote> <br /> <dd> <dl> <dt> <em> <font> <h1> <h2> <h3> <h4> <h5> <h6> <hr /> <i> <li> <nbsp> <ol> <p> <small> <strike> <strong> <sub> <sup> <table> <td> <th> <tr> <tt> <u> <ul>
  • Snippets of code should be wrapped in <code> tags not <pre> tags. In fact, <pre> tags should generally be avoided. If they must be used, extreme care should be taken to ensure that their contents do not have long lines (<70 chars), in order to prevent horizontal scrolling (and possible janitor intervention).
  • Want more info? How to link or How to display code and escape characters are good places to start.
Log In?

What's my password?
Create A New User
Domain Nodelet?
and the web crawler heard nothing...

How do I use this? | Other CB clients
Other Users?
Others pondering the Monastery: (5)
As of 2023-02-02 14:19 GMT
Find Nodes?
    Voting Booth?
    I prefer not to run the latest version of Perl because:

    Results (19 votes). Check out past polls.