Scraping Javascript page using perl

Bpl has asked for the wisdom of the Perl Monks concerning the following question:

Replies are listed 'Best First'.
Re: Scraping Javascript page using perl by Corion (Patriarch) on Nov 21, 2020 at 14:55 UTC
WWW::Mechanize::PhantomJS is somewhat undermaintained, but as you don't tell us how it fails, it's hard to give you concrete advice. Maybe WWW::Mechanize::Chrome works better.	[reply]
Re^2: Scraping Javascript page using perl by Bpl (Scribe) on Nov 21, 2020 at 15:00 UTC
Hi Corion, the problem with WWW::Mechanize::PhantomJS is that its 'get' function return only the classical HTML code (the page source), not the loaded page with JS. Regards, Edoardo Mantovani, 2020	[reply]
Re: Scraping Javascript page using perl by bliako (Monsignor) on Nov 21, 2020 at 17:23 UTC
Quick fix: Open Firefox's Developer Tools, go to the Network tab and observe all the transactions which happen during loading. One of them is requesting the HTML for the actual contents of that page, the 1892 theses (alas only mathematics is immortal), something like this: `http://operedigitali.lincei.it/rendicontiFMN/rol/visart.php?lang=it&type=mat&serie=5&anno=1892&volume=1` The longer way which is the "proper" way IMO is to do what Corion suggested and use WWW::Mechanize::Chrome (I am not acquainted with WWW::Mechanize::PhantomJS). This instructs google chrome browser to get the web-page and then asks it to provide Perl with the DOM. Then it's straight forward with XPath selectors to reach and suck out the desired div's contents. bw, bliako	[reply] [d/l]
Re^2: Scraping Javascript page using perl by Bpl (Scribe) on Nov 21, 2020 at 19:12 UTC
Hi Bliako, nice to see you again, Many thanks for the help! The project is growing, you'll have more info next days. Still thanks for everything!	[reply]
Re^3: Scraping Javascript page using perl by bliako (Monsignor) on Nov 21, 2020 at 19:23 UTC
Glad you never give up! All the best	[reply]
Re: Scraping Javascript page using perl by markong (Pilgrim) on Nov 21, 2020 at 16:50 UTC
Hi Edoardo, a quick reply with some pointers in directions worth exploring. First things first: if I look at the source of that page, I don't see any JS. I see pretty basic HTML. What I see is also a bunch of frames, and you should look at the source of the pages loaded in those frames if you already get what you're looking for (sorry have no time to do this for you, also you don't specify what you're after). If you're also interested in easily executing/parsing some JS (ECMA Script v3) code, I would give something like JE a try! Why? Because I see in the source of the page this: `<meta name="GENERATOR" content="Microsoft FrontPage 3.0">` [download] and I bet the JS wouldn't be an Angular module, if any! Saluti	[reply] [d/l]
Re^2: Scraping Javascript page using perl by Bpl (Scribe) on Nov 21, 2020 at 17:06 UTC
Hi, I am currently fixing this using https://metacpan.org/pod/WWW::Mechanize::Frames The idea is to download every pdf file from the site Grazie per L'interessamento. Edoardo	[reply]


go ahead... be a heretic
	PerlMonks